News >
Recent advances in deep learning video coding within international standardization
Introduction
After the latest video coding standard VVC was finalized in 2020, the international Joint Video Experts Team (JVET) began exploring neural network-based video coding (NNVC) in preparation for the next generation of video coding standards. Competition in NNVC research has been intense, with participants including Tencent, ByteDance, Qualcomm, Ericsson, Nokia, and OPPO.
Tencent Media Lab has made significant contributions to NNVC, including standardizing the technology adoption process, developing reference software, and building several fundamental tools. It also holds key roles including NNVC subgroup chair, reference software maintenance chair, and algorithm description editor chair. At the April 2023 meeting, Tencent Media Lab outperformed competing proposals in the highly competitive in-loop filtering category, achieving the best coding performance. Its single-model filter design was adopted for the unified solution. Tencent is also currently the only organization with major technologies adopted in different modules, namely in-loop filtering and super-resolution. This article briefly introduces Tencent Media Lab's main contributions at recent JVET meetings.
Background
Video is a major carrier of information on today's internet and accounts for a large share of internet data. The vast volume of video presents significant storage and transmission challenges, making video compression essential. Video coding aims to reduce data size or real-time transmission bitrate while maintaining comparable subjective quality. In-loop filtering reduces distortion in the current reconstructed image and supplies higher-quality references for subsequent frames, improving coding efficiency. Super-resolution can also improve efficiency: a lower-resolution version of the source image is encoded to save bits, then restored to high resolution using super-resolution technology.
At the April 2023 JVET meeting, the working group decided to combine existing technologies into a unified, high-complexity neural network-based in-loop filter to advance NNVC. Tencent, ByteDance, Qualcomm, Ericsson, OPPO, and others participated in the discussions and competing proposals. During several days of deliberation, Tencent Media Lab secured the adoption of multiple in-loop filtering technologies into the standards reference software and proposed a unified filter solution that provided an effective reference for the meeting's decisions. Tencent Media Lab was also the only organization with major technologies adopted across both in-loop filtering and super-resolution modules.
Neural network-based in-loop filtering
At the April 2023 JVET meeting, Tencent Media Lab submitted proposal JVET-AD0166, introducing a neural network-based in-loop filter combining transformers and convolutional neural networks (CNNs). As shown in Figure 1, its filter achieved the best average coding performance among the competing designs, with a bitrate reduction of 15.5%.
This section describes the technologies adopted into the standards reference software in terms of network architecture, training, and the inference interface.
1. Network architecture
The neural network-based in-loop filter in JVET-AD0166 uses a single model to process both the luma (Y) and chroma (UV) components of YUV video. As shown in Figure 2, the network receives the luma and chroma components of the reconstructed image (rec_yuv) and the predicted image [1] (pred_yuv), and outputs filtered luma and chroma components together.
Tencent Media Lab also submitted JVET-AD0379 at the April 2023 JVET meeting, proposing a unified filter design ahead of ByteDance. The architecture shown in Figure 3 provided a reliable reference for the JVET working group. Offering a better balance between coding performance and model storage, Tencent's single-model design gained support from hardware experts over the two-model approach supported by ByteDance, Qualcomm, and Ericsson, and was ultimately adopted in the unified filter.
2. Network training and inference interface
To improve consistency between network training and testing, Tencent Media Lab proposed an iterative training method [2]that improves the generalization of neural network filters during testing. The lab's TVD dataset [3] was also incorporated into the standards reference software.
The lab's strategy combining adaptive and fixed weighting was adopted to blend conventional filter outputs with neural network filter outputs, improving subjective quality.
Neural network-based adaptive super-resolution
At the January 2023 JVET meeting, Tencent Media Lab submitted proposal JVET-AC0196, and its neural network-based adaptive super-resolution tool was adopted. The algorithm adaptively chooses the coding resolution, either original or quarter resolution, at the group-of-pictures (GOP) level. When images are coded at quarter resolution, the proposed neural network super-resolution filter upsamples them. Compared with the international standards organization's latest codec software, the tool achieves an average coding efficiency improvement of 5.34% on mainstream 4K sequences.
Summary
During exploration and research for the next generation of video coding standards, multiple Tencent Media Lab technologies, including neural network-based in-loop filters and super-resolution tools, have been adopted into the standards software. Tencent Media Lab is also the only organization with major technologies adopted across different coding modules. The lab will continue its long-term investment in frontier deep learning-based coding research, including practical applications, while maintaining its leadership in this field within international standards organizations.
References
[1] H. Zhu, X. Xu, S. Liu. "Deep learning in-loop filtering tool with QP and residual distribution information," Doc. AVS-M5654, August 2020.
[2] L. Wang, X. Xu, S. Liu, "Optimize neural network based in-loop filters through iterative training,"Picture Coding Symposium, 2022.
[3] X. Xu, S. Liu and Z. Li, "A Video Dataset for Learning-based Visual Data Compression and Analysis,"IEEE International Conference on Visual Communications and Image Processing,2021.