A system-level processing chip, an image processing method and an electronic device

CN122554637APending Publication Date: 2026-08-11SMARTER SILICON (SHANGHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-27
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,现有基于SOC平台的显示画质增强方案,普遍采用全帧统一强制增强的处理模式,无且全程依靠CPU、GPU等通用处理器运行高强度画质增强算法并实时计算画质增强目标参数

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554637A_ABST
    Figure CN122554637A_ABST
Patent Text Reader

Abstract

This application discloses a system-level processing chip, an image processing method, and an electronic device. The system-level processing chip includes: a first processing unit, configured to: receive a first video frame layer generated by processing decoded video frames; perform scene switching detection of video content based on the first video frame layer to obtain a detection result; extract image features from the first video frame layer in response to the detection result indicating a scene switching of video content; and determine target parameters based on the detection result and the image features; and a second processing unit, configured to perform image quality enhancement processing on the first video frame layer according to the target parameters, and output the image quality enhanced first video frame layer to a display component.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a system-level processing chip, an image processing method, and an electronic device. Background Technology

[0002] Currently, display quality enhancement functions are generally integrated into system-on-a-chip (SoC). However, existing display quality enhancement solutions based on SoC platforms generally adopt a uniform, forced enhancement processing mode across the entire frame, relying entirely on general-purpose processors such as CPUs and GPUs to run high-intensity quality enhancement algorithms and calculate the target parameters for quality enhancement in real time. This approach suffers from high power consumption and significant processing latency. Summary of the Invention

[0003] The technical solution provided in this application is as follows:

[0004] The first aspect of this application provides a system-level processing chip, comprising:

[0005] The first processing unit is used for:

[0006] Receive the first video frame layer generated after processing the decoded video frames;

[0007] Based on the first video frame layer, video scene content switching detection is performed to obtain the detection results;

[0008] In response to the detection result indicating a scene change in the video content, image features are extracted from the first video frame layer;

[0009] Based on the detection results, the target parameters are determined according to the image features;

[0010] The second processing unit is used to perform image quality enhancement processing on the first video frame layer according to the target parameters, and output the image quality enhanced first video frame layer to the display component.

[0011] In one possible implementation, the first processing unit performs video scene switching detection based on the first video frame layer to obtain detection results, including:

[0012] Determine the first frame global feature distribution data of the first video frame layer; the first frame global feature distribution data characterizes the global attributes of the first video frame layer;

[0013] Obtain the global feature distribution data of the second screen of the previous video frame of the first video frame layer and the global feature distribution data of the third screen of the starting frame after the last scene switch;

[0014] Determine the first similarity between the global feature distribution data of the first screen and the global feature data of the second screen;

[0015] Determine the second similarity between the global feature distribution data of the first screen and the global feature distribution data of the third screen;

[0016] In response to both the first similarity and the second similarity satisfying the set similarity threshold, it is determined that a scene change has occurred in the video content, and a detection result representing the scene change is obtained.

[0017] In one possible implementation, the first processing unit performs video scene switching detection based on the first video frame layer to obtain detection results, including:

[0018] The first processing unit obtains a second video frame layer based on the first video frame layer; the resolution of the second video frame layer is smaller than the resolution of the first video frame layer.

[0019] The first processing unit performs scene switching detection of video content based on the second video frame layer and obtains the detection result;

[0020] The first processing unit extracts image features from the first video frame layer, including:

[0021] The first processing unit extracts image features from the second video frame layer.

[0022] In one possible implementation, the system-level processing chip further includes:

[0023] The third processing unit is used for:

[0024] Obtain the service scenario information corresponding to the decoded video frame; the service scenario information includes image size specifications.

[0025] Determine whether the image size specification meets the preset resolution threshold;

[0026] In response to the image size specification meeting the preset resolution threshold, super-resolution processing is performed on the decoded video frame to obtain a super-resolution processed video frame.

[0027] The first processing unit receives a first video frame layer generated after processing the decoded video frames, including:

[0028] The first processing unit receives a first video frame layer generated after processing the super-resolution video frames.

[0029] In one possible implementation, the business scenario information further includes: application type;

[0030] The third processing unit performs super-resolution processing on the decoded video frames, including:

[0031] The third processing unit performs super-resolution processing on the decoded video frames based on a target model adapted to the application type.

[0032] In one possible implementation, the second processing unit includes a plurality of target units; the target units are used to perform image quality enhancement processing;

[0033] Extracting image features from the first video frame layer includes:

[0034] The first video frame is divided into multiple image blocks according to the image size that matches the operation array of the target unit;

[0035] Image features are extracted from each of the image blocks.

[0036] In one possible implementation, the second processing unit performs image quality enhancement processing on the first video frame layer according to the target parameters, including:

[0037] Based on the multiple target units, each image block is subjected to image quality enhancement processing in parallel according to the target parameters corresponding to each image block.

[0038] In another aspect, this application provides an image processing method, comprising:

[0039] The first video frame layer is generated by processing the decoded video frames received by the first processing unit.

[0040] Based on the first processing unit's first video frame layer, video scene content switching detection is performed to obtain the detection result;

[0041] Based on the first processing unit's response to the detection result indicating a scene change in the video frame content, image features are extracted from the first video frame layer;

[0042] Based on the detection results obtained by the first processing unit, target parameters are determined according to the image features.

[0043] The second processing unit performs image quality enhancement processing on the first video frame layer according to the target parameters, and outputs the image quality enhanced first video frame layer to the display component.

[0044] In one possible implementation, the step of the first processing unit performing video scene switching detection based on the first video frame layer to obtain the detection result includes:

[0045] Determine the first frame global feature distribution data of the first video frame layer; the first frame global feature distribution data characterizes the global attributes of the first video frame layer;

[0046] Obtain the global feature distribution data of the second screen of the previous video frame of the first video frame layer and the global feature distribution data of the third screen of the starting frame after the last scene switch;

[0047] Determine the first similarity between the global feature distribution data of the first screen and the global feature data of the second screen;

[0048] Determine the second similarity between the global feature distribution data of the first screen and the global feature distribution data of the third screen;

[0049] In response to both the first similarity and the second similarity satisfying the set similarity threshold, it is determined that a scene change has occurred in the video content, and a detection result representing the scene change is obtained.

[0050] In a third aspect of this application, an electronic device is provided, comprising:

[0051] The system-level processing chip includes a first processing unit and a second processing unit;

[0052] The first processing unit is configured to:

[0053] Receive the first video frame layer generated after processing the decoded video frames;

[0054] Based on the first video frame layer, video scene content switching detection is performed to obtain the detection results;

[0055] In response to the detection result indicating a scene change in the video content, image features are extracted from the first video frame layer;

[0056] Based on the detection results, the target parameters are determined according to the image features;

[0057] The second processing unit is used to perform image quality enhancement processing on the first video frame layer according to the target parameters, and output the image quality enhanced first video frame layer to the display component;

[0058] The display component is connected to the second processing unit and is used to display the first video frame layer after the image quality enhancement processing. Attached Figure Description

[0059] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0060] Figure 1 This is a schematic diagram of a processing data flow of a system-level processing chip provided in Embodiment 1 of this application;

[0061] Figure 2 This is a schematic diagram of the data flow for a conventional image enhancement processing scheme.

[0062] Figure 3 This is a schematic diagram of the structure of a first processing unit provided in Embodiment 1 of this application;

[0063] Figure 4 This is a schematic diagram of the processing procedure of a VDSP and a CPU provided in Embodiment 1 of this application;

[0064] Figure 5 A schematic diagram illustrating the scene switching detection and image feature extraction process of a first processing unit provided in this application;

[0065] Figure 6 This is a schematic diagram of the processing data flow of a system-level processing chip provided in Embodiment 4 of this application;

[0066] Figure 7 A schematic diagram of the processing flow of a third processing unit provided in this application;

[0067] Figure 8 This is a flowchart illustrating an image processing method provided in Embodiment 8 of this application. Detailed Implementation

[0068] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0069] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0070] The terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0071] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0072] Reference Figure 1 This is a schematic diagram of the processing data flow of a system-level processing chip provided in Embodiment 1 of this application, as shown below. Figure 1 As shown, the system-level processing chip may include, but is not limited to, a video decoding unit 100, a first processing unit 200, and a second processing unit 300.

[0073] In this embodiment, electronic devices (such as mobile phones, tablets, laptops, smart TVs, etc.) can receive externally input video content to meet the user's need to watch video content.

[0074] To adapt to real-world application scenarios with limited device storage capacity and fluctuating network bandwidth, externally received video content needs to be compressed and encoded (e.g., using mainstream encoding formats such as H.264 and H.265). Compression can significantly reduce the size of video files, making them easier to store locally and improving network transmission efficiency, thus avoiding storage pressure and transmission stuttering caused by excessively large video files.

[0075] However, in order to achieve efficient compression during the compression encoding process, it is inevitable that video image details will be lost, colors will be distorted, contrast will be insufficient, and noise will be superimposed. If such compressed videos are displayed directly, it will seriously affect the user's viewing experience. Therefore, image quality enhancement can be performed through a system-on-a-chip (SOC). The first step in image quality enhancement is to decode the compressed video and convert it into decoded video frames that can be processed later.

[0076] Based on this, the video decoding unit (VPU decoder) 100 can receive compressed video streams (including locally stored video files, online video streams transmitted over the network, etc.) from external sources through the dedicated interface built into the SOC. Then, it starts the dedicated decoding circuit and sequentially completes a series of hardware-level parallel decoding operations such as stream parsing, entropy decoding, inverse quantization, motion compensation, and pixel reconstruction, converting the compressed video stream (which cannot be directly displayed and has image quality defects) into decoded video frames.

[0077] After obtaining the decoded video frames, current conventional image enhancement processing solutions adopt a front-end centralized processing architecture. The overall image quality optimization process is completed before the display software pipeline framework is processed, without relying on the display link's native hardware to adapt to the needs of fine-tuning in the later stages.

[0078] For example, such as Figure 2 As shown, current conventional image enhancement processing solutions, after video frame decoding, directly schedule general-purpose computing processors such as CPUs (Central Processing Units) and GPUs (Graphics Processing Units) within the SOC as the core image enhancement carriers. The entire process relies on the general computing power resources of the general-purpose processors to run various complex and computationally intensive global image enhancement algorithms. For each decoded original video frame, a full-chain image quality optimization operation is continuously performed, including full-dimensional image feature statistical analysis, real-time calculation of image enhancement target parameters, global color adjustment, detail sharpening, noise reduction and repair. After all high-intensity image enhancement operations are completed, the enhanced video frame is then sent to the display software pipeline framework for subsequent formatting, color calibration and layer encapsulation preprocessing. The subsequent display software pipeline framework will directly send the processed standardized video frame layers to the DPU (Display Processing Unit) inside the SOC. The DPU will uniformly complete the basic processing tasks such as display timing adaptation, layer compositing, and display driver adaptation, and finally output the processed video image to the display component of the terminal device for display.

[0079] This processing mode, which relies on the CPU / GPU general-purpose processor for forced full-frame enhancement, does not perform differentiated on-demand scheduling based on changes in video scene. It maintains a state of high-intensity full-frame computation at all times. The general-purpose processor runs under high load for a long time and data interaction scheduling is not only the core reason for the high overall power consumption and large cumulative processing latency of this type of pre-enhancement processing solution, but also fundamentally differs from the processing flow and hardware calling method of this application, which relies on the display software pipeline to complete layer preprocessing and then the second processing unit 300 performs post-enhancement targeted enhancement in accordance with target parameters.

[0080] In this embodiment, the decoded video frames can be processed by a display software pipeline framework (e.g., Android frameworks pipeline, DRM / KMS (DirectRendering Manager / kernel mode settings) under Linux system) to generate the first video frame layer. Specific processing steps may include, but are not limited to:

[0081] Step S11: Format the decoded video frames.

[0082] The VPU decoder typically outputs raw YUV frames, but different display components and subsequent processing modules have different requirements for pixel format. The Android frameworks pipeline converts the raw YUV frames into a unified pixel format, such as RGBA or YUV_420_SP. This unified format conversion ensures smooth flow of video frame data between different components and modules, avoiding processing errors or performance degradation caused by format incompatibility.

[0083] Step S12: Perform color calibration on the decoded video frames after formatting.

[0084] Display components have different color characteristics. In order to present accurate and consistent color effects on different devices, it is necessary to convert the color standard according to the characteristics of the display components. Common color standards include BT.601, BT.709, and BT.2020, which are suitable for display scenarios with different resolutions and color gamuts.

[0085] The display software pipeline framework can perform color space mapping on video frames based on the color standard of the target display component, converting the original color information to the target color space. For example, when the video originally uses the BT.601 color standard encoding, but the display component supports the BT.709 color standard, corresponding color space conversion is required to ensure accurate color reproduction.

[0086] Step S13: Encapsulate the layer attributes of the decoded video frames after color calibration to obtain the first video frame layer.

[0087] During video display, in addition to the content of the video frames themselves, layer composition information also needs to be considered, such as z-order (layer order), alpha transparency, and rotation / scaling matrices. z-order determines the front-to-back order of different layers during display, ensuring that video layers are correctly displayed above or below other interface elements; alpha transparency controls the transparency of layers, achieving blending effects between layers; and rotation / scaling matrices can perform geometric transformations such as rotation and scaling on video frames to meet different display requirements.

[0088] The display software pipeline framework can encapsulate these layer attributes into the decoded video frame to obtain the first video frame layer.

[0089] In this embodiment, the first processing unit 200 can be placed after the display software pipeline framework, and the first processing unit 200 receives the first video frame layer generated after processing the decoded video frames.

[0090] Since the display software pipeline framework has completed preprocessing operations such as format standardization, color space mapping, and layer attribute encapsulation, the generated first video frame layer has a unified format, standard color, and complete attributes. The first processing unit 200 does not need to perform redundant operations such as format conversion and color calibration. After receiving the data, it can directly enter the core processing stage, which simplifies hardware design, reduces design costs, and significantly shortens processing latency.

[0091] Furthermore, the first processing unit 200 can be used to perform video scene content switching detection based on the first video frame layer to obtain detection results.

[0092] In practical video processing scenarios, image enhancement processing to improve picture quality is crucial. Since video images contain many elements that affect visual perception, achieving high-quality image enhancement often requires complex and powerful algorithms to deeply mine and optimize various details in the image.

[0093] High-resolution videos contain richer and more diverse details, which places higher demands on image enhancement algorithms. Consequently, these algorithms require enormous computing resources. If a uniform and high-intensity image enhancement algorithm is applied to every frame during video processing, the system-level processing chip will face immense computational pressure. This leads to a series of problems: significantly increased processing latency causes stuttering and choppy playback, severely impacting the user's viewing experience; device power consumption also increases dramatically, accelerating battery drain, shortening battery life, and potentially causing performance degradation, system instability, or even hardware damage due to overheating.

[0094] However, in reality, videos exhibit continuity and stability during playback, with infrequent changes in the scene for most of the time. For example, in a video of an indoor dialogue scene, the background environment and lighting conditions remain largely unchanged throughout the conversation, resulting in relatively stable image content. In such cases, continuously processing each frame with a high-intensity algorithm would involve much redundant computation, leading to a waste of computing resources.

[0095] In this embodiment, the first processing unit 200 is positioned after the display software pipeline framework. This design makes the latency constraints more stringent. Because the display software pipeline framework has already completed a series of preprocessing operations, subsequent processing needs to be completed in a shorter time to meet real-time requirements. Scene switching detection is a crucial means of addressing this stringent latency constraint, effectively reducing processing latency.

[0096] From a power consumption perspective, if every frame requires enhancement processing, the system-level processing chip will always be operating under high load, continuously consuming a large amount of power to maintain the computation of complex algorithms, resulting in persistently high power consumption. For example, when playing a one-hour high-resolution video, if each frame undergoes high-intensity image enhancement processing, the chip needs to continuously perform a large number of data calculations and storage accesses, all of which consume power, leading to a significant increase in the overall power consumption of the device.

[0097] The first processing unit 200 can accurately identify scene changes in the video frame through scene switching detection. When no scene change is detected, it means that the current frame and the previous frame have high similarity in image features. At this time, a relatively simple processing method can be used, or the image enhancement parameters determined in the previous scene can be directly used. This method avoids repeating complex and high-intensity calculations, greatly reducing unnecessary computation. Due to the reduced computation, the load on the system-level processing chip is also reduced, and the power consumption is correspondingly reduced. For example, during the aforementioned 1-hour video playback, the scene is stable for most of the time, and the first processing unit 200 only needs to perform simple processing or use the parameters. The chip does not need to perform a large number of complex calculations, and the power consumption is significantly reduced.

[0098] Once a scene change is detected, such as switching from an indoor scene to an outdoor scene, significant changes may occur in the color distribution, texture features, and lighting conditions of the image. At this time, the first processing unit 200 promptly re-determines suitable image enhancement parameters for the new scene and applies corresponding algorithms to process the video frames in the new scene accordingly. Although some complex calculations are required when switching to a new scene, this accounts for a relatively small percentage of the overall video playback. This ensures good image enhancement in the new scene while ensuring reasonable allocation of computing resources when the scene is stable, thus reducing the overall power consumption of the device. The power consumption mode of the first processing unit 200, triggered by the scene, greatly improves the energy efficiency of the device and extends its battery life.

[0099] Furthermore, the first processing unit 200 can be used to extract image features from the first video frame layer in response to a scene change in the video frame content represented by the detection result.

[0100] In real-world video playback scenarios, the detection result of scene transitions in video content is crucial information, reflecting whether the video has transitioned from the current scene to a new one. When the detection result obtained by the first processing unit 200 clearly indicates that a scene transition has occurred in the video content, it means that the video has entered a completely new scene context, at which point it is necessary to re-extract image features for image quality enhancement.

[0101] The image features extracted by the first processing unit 200 can be based on the first video frame layer output by the display software pipeline framework. Since the first video frame layer has completed format standardization and color space mapping, and has a unified format, standard color and complete attributes, the extracted image features can accurately reflect the real image quality characteristics of the new scene, avoiding the feature extraction distortion problem caused by non-standardized data.

[0102] The extracted image features may include, but are not limited to, at least one of the following: color distribution features, texture features, and lighting condition features.

[0103] Color distribution characteristics can intuitively reflect the color composition of an image and are a core basis for judging scene differences and optimizing color distortion. Taking a landscape video frame as an example, the sky may be predominantly blue, with the blue area occupying a large proportion. At the same time, the saturation and brightness of blue will change with different weather conditions (high saturation and brightness on sunny days, and low saturation and brightness on cloudy days). The ground may contain different elements such as green vegetation and brown soil. The color distribution characteristics of each element are different. Green vegetation usually has high saturation, while brown soil tends to be a warm color with low saturation. These differences in color distribution characteristics can accurately distinguish different scenes.

[0104] Texture features reflect the texture details of different objects in an image, including texture roughness, detail density, directionality, and distribution patterns. Different objects have significantly different surface textures. By extracting texture features, the composition of objects in the image can be clearly distinguished, while also compensating for detail loss caused by compression encoding. For example, in a video frame containing buildings and streets, the walls of buildings may have regular brick textures with uniform directionality and periodicity; the street surface may have a rough asphalt texture with a chaotic texture and high detail density; and objects such as clothing and wood will also exhibit their own unique texture features. By capturing these texture details, the first processing unit 200 can accurately grasp the detailed characteristics of the image, providing a basis for subsequent detail compensation and sharpness optimization.

[0105] Lighting conditions reflect the intensity, direction, color, and distribution of light and shadow in an image. Differences in lighting conditions directly affect the brightness, contrast, and color reproduction of an image, serving as a key basis for optimizing uneven brightness and color deviations. For example, in indoor scenes, there is usually a superposition of natural light from windows and indoor lighting. Natural light is often cool-toned and bright, while indoor lighting is often warm-toned and relatively soft. The superposition of these two will create different shadows and highlights on the surface of objects. In outdoor scenes, lighting conditions vary with time. At midday, direct sunlight results in high light intensity, short shadows on object surfaces, and sharp edges. In the evening, oblique sunlight reduces light intensity and warms the tone, resulting in longer shadows on object surfaces with blurred edges.

[0106] Furthermore, the first processing unit 200 can be used to determine target parameters based on the image features using the detection results.

[0107] In this embodiment, the first processing unit 200 can determine the target parameters based on the premise that the scene has been switched, combined with the extracted image features such as color distribution, texture, and lighting conditions, through a preset adaptive algorithm. The adaptive algorithm can automatically adjust the calculation weight according to the differences in image features of different scenes, ensuring that the target parameters are highly matched with the image quality characteristics and defects of the new scene, and avoiding the problem of poor enhancement effect caused by general parameters.

[0108] In this embodiment, saturation optimization parameters (i.e., an implementation of a target parameter) can be generated according to the color distribution characteristics to improve the color distortion and dullness caused by compression encoding. For example, in the case of insufficient saturation of the blue sky in a landscape frame, the parameter will specifically increase the saturation of the blue channel.

[0109] For texture features, corresponding detail compensation parameters (i.e., an implementation of a target parameter) can be generated to compensate for the loss of texture details caused by compression. For example, for the case of blurred texture of building bricks, the parameters will strengthen the texture edges and improve the clarity of details.

[0110] Based on the characteristics of lighting conditions, corresponding brightness adjustment parameters and contrast calibration parameters (i.e., an implementation method of target parameters) can be generated to optimize uneven lighting and brightness deviation. For example, for insufficient brightness caused by soft indoor lighting, the parameters will appropriately increase the local brightness. For contrast imbalance caused by strong outdoor midday light, the parameters will calibrate the brightness difference between bright and dark areas.

[0111] In addition, target parameters may include noise suppression parameters, which are used to suppress noise superposition caused by compression encoding and further improve image purity.

[0112] When performing image quality enhancement and configuring various target parameters, it is essential to fully consider the differentiated image quality requirements of different video scenarios, and to set the parameters flexibly and precisely based on this. These target parameters can be presented in various forms:

[0113] On one hand, it can be set as an image quality adjustment gain coefficient and a threshold parameter. The image quality adjustment gain coefficient can directly amplify or reduce a certain characteristic of the image quality. For example, adjusting the contrast gain coefficient of the image can change the degree of difference between bright and dark areas in the image. The threshold parameter is used to define the triggering conditions for specific image quality processing operations. For example, setting a threshold for noise removal will activate the corresponding noise reduction algorithm only when the noise level in the image exceeds the threshold.

[0114] On the other hand, various parametric curves can also be used. Brightness transformation curves can finely adjust the brightness distribution of an image; by changing the position of different points on the curve, local brightness can be increased or decreased, making the image's tonal gradations richer and more natural. Saturation transformation curves focus on adjusting the image's color saturation, making the colors in the picture more vibrant or softer to meet the color performance requirements of different scenes. Furthermore, brightness and saturation transformation curves can be integrated to form a custom HDR transformation curve. This comprehensive curve can simultaneously optimize both the brightness and color of an image, achieving more complex and precise image quality adjustments.

[0115] The specific rules for the values ​​of the above-mentioned image adjustment parameters, the curve fitting methods, and the algorithm configuration logic can all be implemented with reference to existing conventional image processing technologies in this field. This application does not impose any specific limitations on these aspects.

[0116] Once the target parameters are determined, the first processing unit 200 can quickly transmit them to the second processing unit 300 through the dedicated data transmission link built into the SOC. The transmission process does not require intermediate conversion, ensuring the efficiency and accuracy of parameter transmission and providing timely support for the real-time image quality enhancement processing of the second processing unit 300, while avoiding the impact of parameter transmission delay on the overall processing efficiency.

[0117] The second processing unit 300 is used to perform image quality enhancement processing on the first video frame layer according to the target parameters, and output the image quality enhanced first video frame layer to the display component.

[0118] The second processing unit 300 employs a parallel receiving mechanism, simultaneously receiving the first video frame layer output by the display software pipeline framework and the target parameters transmitted by the first processing unit 200. This parallel receiving design eliminates the need to wait for the first processing unit 200 to fully complete the target parameter calculation; the second processing unit 300 can prepare for processing immediately after receiving the first video frame layer, effectively shortening data latency and further optimizing overall processing efficiency. This addresses the stringent latency constraints imposed by placing the first processing unit 200 after the display software pipeline framework. Furthermore, since the first video frame layer has already undergone standardization, the second processing unit 300 does not need to perform redundant operations such as format conversion and color calibration, and can directly perform enhancement processing based on this layer, further saving processing time.

[0119] The second processing unit 300 can use the saturation optimization parameters determined by the first processing unit 200 based on the color distribution characteristics to make targeted adjustments to the color saturation of the first video frame layer. For example, it can increase the saturation of the blue sky and green vegetation in the landscape frame, correct the warm tone deviation of the brown land, improve the color distortion and dullness caused by compression encoding, and make the picture colors more vivid and realistic, in line with the color standard of the display component.

[0120] The second processing unit 300 can use the detail compensation parameters determined by the first processing unit 200 based on the texture features to enhance the texture details of various objects in the image, such as sharpening the brick texture of buildings and clarifying the asphalt texture of streets, to make up for the loss of details caused by compression encoding, and to make the image details richer and clearer.

[0121] The second processing unit 300 can use the brightness adjustment parameters and contrast calibration parameters determined by the first processing unit 200 based on the characteristics of the lighting conditions to optimize the brightness and contrast of the image locally or globally according to the lighting characteristics of the scene. For example, it can improve the overall brightness of the indoor scene and calibrate the contrast of the outdoor midday scene to solve the problems of uneven lighting and brightness deviation, making the image brightness more uniform and the sense of layering stronger.

[0122] The second processing unit 300 can also suppress noise superposition caused by compression encoding in a targeted manner according to the noise suppression parameters, reduce problems such as image blurring and noise, and improve image purity.

[0123] After the second processing unit 300 completes the image enhancement processing, it can directly output the processed first video frame layer to the display component of the electronic device (such as a mobile phone screen, tablet computer screen, laptop screen, smart TV screen, etc.), and the display component will display the image enhancement processed first video frame layer.

[0124] In this embodiment, the second processing unit 300 does not require a completely new, dedicated hardware unit design. It can be implemented using a dedicated display processing hardware unit (e.g., a DPU) integrated within the SOC chip at the front end of the display link and upstream of the display components. In this embodiment, the dedicated display processing hardware unit (e.g., the DPU) has comprehensive basic display processing functions, mainly including basic image processing and display driving related tasks such as video layer compositing, image geometric rotation and scaling processing, basic color calibration and adaptation, display format compatibility adaptation, and precise control of display timing. It is a core standard hardware component that ensures the stable and normal output of various video images and system interfaces to the terminal display components.

[0125] Dedicated display processing hardware units (e.g., DPUs) possess dedicated hardware-level parallel computing capabilities, and their inherent computing resources are fully capable of meeting the image enhancement processing requirements of this application. The key improvement of this application does not lie in modifying or upgrading the hardware architecture, computing circuits, and basic processing logic of the second processing unit 300 (e.g., the DPU), but rather in integrating the image enhancement processing function into the second processing unit 300 without altering its basic display functions. The second processing unit 300 then uniformly undertakes both basic display processing and image enhancement processing tasks. This design, which reuses existing hardware, can optimize power consumption and latency, and its specific advantages include:

[0126] First, power consumption is significantly reduced, approaching the system power consumption level without display quality enhancement. Since the second processing unit 300 is a hardware unit in the SOC that needs to run continuously, it is already in a stable operating state in order to complete the basic display driving tasks. The embodiment only adds the quality enhancement processing task on its existing operation, without the need to start other processors such as CPU and GPU to perform the quality enhancement operation. This avoids the redundant power consumption caused by starting and running additional processors, and is basically close to the system power consumption level without any quality enhancement processing.

[0127] Secondly, the image enhancement processing is directly completed by the second processing unit 300 hardware, without adding any additional processing latency. The second processing unit 300 itself has a hardware architecture specifically designed for display-related processing, integrating a large number of dedicated image processing units. While completing its own basic display processing tasks, it can perform image enhancement operations in parallel, without the need for complex software scheduling processes or data transfer and switching between different processors. The entire image enhancement process is synchronized with the basic display processing, without adding any additional latency to the overall video processing.

[0128] In this application, the type of the second processing unit 300 is not limited, and may include, but is not limited to, DPU (display processing unit), ISP (image signal processor), etc.

[0129] In this embodiment, as Figure 3 As shown, the first processing unit 200 can be composed of a VDSP (Video Digital Signal Processor) and a CPU (Central Processing Unit). Based on the differences in hardware characteristics between the VDSP and the CPU, it is specifically adapted to the processing requirements of each task to realize the extraction of image features and the driving of target parameters.

[0130] The following steps can all be performed by VDSP: receiving the first video frame layer generated after processing by the display software pipeline framework, detecting video scene switching based on the first video frame layer and obtaining the detection result, and extracting image features from the first video frame layer in response to the scene switching represented by the detection result.

[0131] The task of determining the target parameters by combining the detection results with the extracted image features can be performed by the CPU.

[0132] On the one hand, VDSP, as a dedicated video digital signal processor, has a hardware architecture tailored for video image processing scenarios. It has extremely strong parallel computing capabilities and dedicated computing units, requiring no additional software scheduling overhead. It excels at handling high-frequency, high-intensity, and repetitive real-time image processing tasks such as scene switching detection and image feature extraction.

[0133] Compared to CPUs, VDSPs offer a significant speed advantage in image data processing. CPUs, as general-purpose processors, must handle various logical operations and task scheduling within the system. When processing specialized tasks like image feature extraction and scene detection, they require software-level scheduling, resulting in scheduling latency and low computational efficiency. In contrast, VDSPs have built-in dedicated image processing modules and data caching units, allowing direct access to standardized first video frame layers. Without complex scheduling, they can rapidly complete layer reception, feature extraction, and scene detection at hardware-level parallel processing speeds, significantly reducing the time required for such real-time processing tasks. This is a crucial supplementary measure to address the more stringent latency constraints imposed by placing the first processing unit 200 behind the display software pipeline framework. For example, when extracting image features, VDSPs can simultaneously perform parallel operations on three types of features: color distribution, texture, and lighting conditions. Compared to the serial processing of CPUs, this can improve processing efficiency several times over, effectively reducing single-frame processing latency.

[0134] On the other hand, as a general-purpose processor, the CPU excels at handling highly flexible tasks that require dynamic adjustment, such as logical judgment, algorithm adaptation, and parameter calculation. The task of determining target parameters based on detection results and image features precisely requires dynamically adjusting computational logic based on scene differences and adapting to image quality defects in different scenes, making it suitable for CPU execution. If this task were executed by the VDSP, it would require additional complex algorithm programming adaptation for the VDSP, increasing hardware design complexity and consuming VDSP computing resources, thus affecting the efficiency of its core real-time image processing tasks. Assigning this task to the CPU, however, fully leverages the CPU's logical operation advantages, quickly completing adaptive algorithm calculations, target parameter calibration, and output, while avoiding the consumption of VDSP resources. This ensures that the VDSP can focus on core real-time tasks such as layer reception, scene detection, and feature extraction, achieving parallel collaboration and leveraging each other's strengths.

[0135] In summary, the first processing unit 200 adopts a collaborative design between the VDSP and the CPU, which avoids the inefficiency and latency of the CPU handling dedicated image processing tasks, and also avoids the waste of resources in the VDSP handling logic operation tasks. This further reduces the overall processing latency of the first processing unit 200. Together with the preprocessing optimization of the display software pipeline framework and the reduction of redundant operations in scene switching detection, it forms a triple synergy, further reducing power consumption and latency, while improving the processing stability and reliability of the first processing unit 200.

[0136] The processing procedures of VDSP and CPU can be found in [reference needed]. Figure 4 ,like Figure 4As shown, the VDSP can receive the first video frame layer generated after processing by the display software pipeline framework, perform video scene switching detection based on the first video frame layer and obtain the detection result, and extract image features from the first video frame layer in response to the scene switching represented by the detection result.

[0137] In some implementations, the specific process by which the VDSP extracts image features from the first video frame layer may include, but is not limited to, the following steps: First, the VDSP converts the first video frame layer from Image(T) format to HSV format and denotes the converted image as Image_hsv; then, the VDSP performs feature extraction on Image_hsv, calculating its S-channel (saturation channel) and V-channel (luminance channel) histograms respectively. The S-channel histogram is denoted as Hist_s(T) (i.e., color saturation histogram), and the V-channel histogram is denoted as Hist_v(T) (i.e., grayscale histogram). The two histograms together constitute the core image features for subsequent CPU calculation of target parameters.

[0138] After feature extraction is complete, the VDSP sends the extracted image features (Hist_s(T) and Hist_v(T)) along with the scene transition detection results to the CPU. The scene transition detection results are represented by a flag, bSceneChangeFlag, specifically defined as follows: when bSceneChangeFlag=0, it indicates that no scene transition has occurred in the current video frame; when bSceneChangeFlag=1, it indicates that a scene transition has occurred in the current video frame.

[0139] After receiving the image features and detection results sent by the VDSP, the core task of the CPU is to calculate various target parameters adapted to different scenarios based on the above information (e.g., a custom HDR transformation curve integrating brightness transformation curve and saturation transformation curve), and complete the configuration of DPU-related registers. The specific processing logic is as follows:

[0140] The CPU first checks the value of the bSceneChangeFlag flag. If bSceneChangeFlag = 0, indicating that no scene change has occurred, the CPU does not perform any processing operations and exits the current processing flow directly. If bSceneChangeFlag = 1, indicating that a scene change has occurred, the CPU performs the following two operations:

[0141] Firstly, for the V-channel histogram Hist_v(T) extracted by VDSP, the CPU processes it using an adaptive local region stretching algorithm to obtain the custom brightness transformation curve TransCurse_v required by the DPU, and writes the transformation curve into the corresponding register of the DPU to provide parameter support for the subsequent brightness optimization processing of the DPU.

[0142] Secondly, for the S-channel histogram Hist_s(T) extracted by VDSP, the CPU first calculates the average saturation value of the histogram, and then compares the average value with a preset threshold. If the average saturation value is lower than the preset threshold, each saturation value in the image is weighted separately (if block processing is used, the weighting coefficients of each block can be the same or different according to actual needs). The weighted processing is used to obtain the saturation transformation curve TransCurse_s required by DPU, and the transformation curve is written into the corresponding register of DPU to provide parameter support for the subsequent saturation optimization processing of DPU.

[0143] After receiving the custom HDR transformation curves (TransCurse_v and TransCurse_s) configured by the CPU, the DPU performs image enhancement processing on the first video frame layer based on the custom HDR transformation curves, completes its own DPUpipeline (display processing pipeline) process, and finally outputs the display image after image enhancement processing, which is then transmitted to the display component of the electronic device for display.

[0144] In this embodiment, the first processing unit 200 is placed after the display software pipeline framework. With the help of the preprocessing operations already completed by the framework, such as format standardization, color space mapping, and layer attribute encapsulation, the first processing unit 200 does not need to perform redundant operations such as format conversion and color calibration. It can directly enter the core processing links such as scene switching detection and image feature extraction, which greatly shortens the processing latency and simplifies hardware design and reduces design costs.

[0145] During the core processing, the first processing unit 200 uses the original parameters and reduces redundant calculations when the scene is not switched. When the scene is switched, it accurately extracts image features and determines the target parameters for adaptation. This ensures the targeted nature of the image enhancement effect and avoids the problems of increased power consumption and latency caused by high-intensity processing of the whole frame, thus realizing the reasonable allocation of computing resources.

[0146] Furthermore, the dedicated hardware unit already existing in the front end of the display link in the SOC can be reused as the second processing unit 300 without adding any hardware costs. The image enhancement function can be integrated without changing its basic display function. The basic display and image enhancement processing can be completed in parallel based on its hardware architecture. There is no need to start the CPU, GPU and other processors, which significantly reduces the system power consumption to a level close to that of no display image enhancement. Moreover, the image enhancement processing and the basic display processing are carried out synchronously without adding any additional processing latency.

[0147] Furthermore, the first processing unit 200 adopts a collaborative design between VDSP and CPU. Relying on the hardware advantages of VDSP dedicated video image processing, it can quickly complete high-frequency and high-intensity real-time tasks such as receiving the first video frame layer, scene switching detection, and image feature extraction. This avoids the inefficiency and latency of CPU processing such dedicated tasks. At the same time, the CPU undertakes highly flexible tasks that require dynamic adjustment, such as target parameter calculation, avoiding the occupation of VDSP computing resources. This further reduces the overall processing latency of the first processing unit 200 and improves processing stability and reliability.

[0148] This application is not limited to the combination of DSP and CPU processors. In practical applications, one or more other processor combinations with corresponding computing capabilities can be used as alternatives, as long as they can fully carry out all core processing tasks such as scene detection, feature extraction, and parameter calculation.

[0149] It should be noted that by placing the first processing unit 200 after the display software pipeline framework, the first video frame layer it receives has already undergone unified preprocessing. Whether it originates from hardware decoding (triggered by the video decoding unit 100 VPU decoder), software decoding (completed by the CPU through software algorithms), or non-video decoding scenarios such as camera capture, game rendering, and system UI display, it can be normally recognized and processed by the first processing unit 200. This means the image quality enhancement function no longer relies on the trigger signal of the video decoding unit 100, breaking through the limitation of image quality enhancement being a dedicated function customized only for video playback, and upgrading to a universal image processing function applicable to all display content. This universal design can significantly expand the application whitelist. Whether it's a video app, gallery, camera, game, or system interface, as long as its display content can be processed by the display software pipeline framework to generate the first video frame layer and ultimately be composited for display, it can enjoy the corresponding image quality enhancement effect through scene detection and parameter calculation by the first processing unit 200 and enhancement processing by the second processing unit 300, greatly improving the application adaptability and practicality of the solution.

[0150] In another embodiment of this application, a system-level processing chip is provided in Embodiment 2 of this application. This embodiment is mainly an implementation method for the first processing unit 200 to perform video scene switching detection based on the first video frame layer and obtain the detection result. Specifically, it may include, but is not limited to:

[0151] The first processing unit 200 determines the first frame global feature distribution data of the first video frame layer; the first frame global feature distribution data characterizes the global attributes of the first video frame layer.

[0152] In this embodiment, the first processing unit 200 can statistically analyze the frequency of different colors appearing in the first video frame layer to obtain a color histogram, and use the color histogram as global feature distribution data of the first frame.

[0153] For example, when the first video frame layer is a landscape image, the color histogram can clearly show the distribution ratio of the main colors such as blue (sky), green (vegetation), and brown (soil) in the image. By determining the color histogram, the first processing unit 200 can obtain the global color information of the first video frame layer, providing basic data for subsequent scene transition detection.

[0154] Additionally, the first processing unit 200 acquires the second-view global feature distribution data of the previous video frame of the first video frame layer and the third-view global feature distribution data of the starting frame after the last scene switch.

[0155] The preceding video frame and the first video frame layer are temporally adjacent, and their content typically exhibits high similarity. In this embodiment, by comparing the global feature distribution data of the first video frame layer and the preceding video frame, a preliminary determination can be made as to whether significant changes have occurred in the image.

[0156] The starting frame after the last scene transition represents the beginning of a scene, which may differ from the scene to which the current frame belongs. By comparing the global feature distribution data of the first video frame layer and the starting frame, it can be further determined whether the current frame has entered a new scene.

[0157] In addition, the first processing unit 200 determines a first similarity between the first screen global feature distribution data and the second screen global feature distribution data.

[0158] The first similarity score reflects the degree of similarity between the current frame and the previous frame in terms of global color features. If the first similarity score is high, it means that the content of the current frame and the previous frame has changed little and may still belong to the same scene; if the first similarity score is low, it means that the content of the current frame and the previous frame has changed significantly and a scene change may have occurred.

[0159] In this embodiment, the first processing unit 200 can determine the first cosine distance or the first Euclidean distance between the first screen global feature distribution data and the second screen global feature distribution data as the first similarity.

[0160] The first cosine distance can range from [0,1]. The closer the first cosine distance is to 0, the higher the similarity between the two sets of feature data and the smaller the difference in the image content. The closer the first cosine distance is to 1, the lower the similarity between the two sets of feature data and the greater the difference in the image content.

[0161] The first Euclidean distance can be in the range of [0, +∞). The smaller the Euclidean distance value, the higher the similarity between the two sets of feature data and the smaller the difference in the image content; the larger the first Euclidean distance value, the lower the similarity between the two sets of feature data and the greater the difference in the image content.

[0162] In addition, the first processing unit 200 determines a second similarity between the global feature distribution data of the first screen and the global feature distribution data of the third screen.

[0163] In this embodiment, the first processing unit 200 can also determine the second cosine distance or the second Euclidean distance between the global feature distribution data of the first screen and the global feature distribution data of the third screen as the second similarity.

[0164] The second similarity score reflects the degree of similarity between the current frame and the starting frame after the last scene change in terms of global color features. If the second similarity score is high, it means that the content of the current frame and the starting frame are quite similar and may still belong to the same scene; if the second similarity score is low, it means that the content of the current frame and the starting frame are quite different and may have entered a new scene.

[0165] Furthermore, in response to the first similarity and the second similarity both satisfying the set similarity threshold, the first processing unit 200 determines that a scene change has occurred in the video content and obtains a detection result representing the scene change.

[0166] Corresponding to the implementation method that uses cosine distance as similarity, the similarity threshold can be set as the cosine distance threshold. When both the first cosine distance and the second cosine distance are greater than the cosine distance threshold, it means that the short-term difference between the current frame and the immediately preceding frame, and the long-term difference with the starting frame of the previous scene, have reached the scene switching judgment criteria, that is, the current scene has transitioned from the original scene to a completely new scene.

[0167] Corresponding to the implementation method that uses Euclidean distance as the similarity metric, the similarity threshold can be set as the Euclidean distance threshold. When both the first Euclidean distance and the second Euclidean distance are greater than the set Euclidean distance threshold, it is determined that the video content has changed scene.

[0168] In this embodiment, by introducing a dual feature comparison between the previous frame and the starting frame of the previous scene, the first similarity and the second similarity are calculated respectively, taking into account both the judgment of short-term fluctuations in the image and long-term scene differences. Combined with the quantization method of cosine distance or Euclidean distance, the accuracy of scene switching detection is greatly improved, effectively avoiding misjudgment caused by small fluctuations in a single frame and missed judgment caused by the omission of cross-scene differences, thus ensuring the reliability of scene switching detection results.

[0169] Furthermore, using color histograms as global feature distribution data for the screen, this feature can accurately reflect the global color attributes of video frames, and has low computational complexity. It is compatible with the hardware computing advantages of VDSP, and can quickly complete feature extraction and similarity calculation without occupying a large amount of computing resources. This effectively shortens the processing latency of scene switching detection and adapts to the stringent latency constraints of the first processing unit 200 being set after the display software pipeline framework.

[0170] In another embodiment of this application, a system-level processing chip is provided in Embodiment 3 of this application. This embodiment is mainly an implementation method for the first processing unit 200 to perform video scene switching detection based on the first video frame layer and obtain the detection result. Specifically, it may include, but is not limited to:

[0171] The first processing unit 200 obtains a second video frame layer based on the first video frame layer; the resolution of the second video frame layer is less than the resolution of the first video frame layer.

[0172] In this embodiment, the first processing unit 200 can perform a downsampling operation on the first video frame layer to obtain a second video frame layer with a resolution lower than that of the first video frame layer.

[0173] Downsampling can be achieved in various ways, such as average pooling and max pooling. Taking average pooling as an example, assuming the first video frame layer is divided into multiple non-overlapping small regions (pooling windows), for each small region, the average value of all pixel values ​​within that region is calculated, and this average value is used as the pixel value at the corresponding position in the downsampled image. In this way, the image resolution is reduced while retaining the main feature information of the image.

[0174] The first processing unit 200 performs video scene content switching detection based on the second video frame layer and obtains the detection result.

[0175] In this embodiment, the same operation process can be performed on the second video frame layer, referring to the steps, logic and judgment method for scene switching detection of the first video frame layer described in detail above, in order to determine whether the video content corresponding to the second video frame layer has undergone scene switching. The specific process will not be repeated here.

[0176] In this embodiment, the first processing unit 200 extracts image features from the first video frame layer, which may specifically include, but is not limited to:

[0177] The first processing unit 200 extracts image features from the second video frame layer.

[0178] In this embodiment, the first processing unit 200 may use the same method as extracting image features from the first video frame layer to obtain the image features corresponding to the second video frame layer.

[0179] The process by which the first processing unit 200 performs video scene content switching detection and image feature extraction based on the second video frame image can be found in [reference]. Figure 5 ,like Figure 5 As shown, firstly, the VDSP within the first processing unit 200 can perform downsampling processing on the first video frame layer output by the display software pipeline framework, reducing the resolution of the high-resolution layer to below 256P, to obtain a smaller target layer with a lower resolution, namely the second video frame layer, which is uniformly denoted as Image(T).

[0180] After downsampling is completed, VDSP still performs scene change detection based on the small-sized second video frame layer Image(T). After the detection is completed, it will determine whether a scene change has occurred and execute the corresponding branch operation:

[0181] If the detection result determines that the scene has not changed, VDSP immediately terminates all subsequent image feature extraction processes and simultaneously sets the scene switching flag bSceneChangeFlag to 0. It does not need to transmit any image feature data to the CPU and directly completes this round of processing.

[0182] If the detection result indicates a scene change, VDSP continues to perform image feature extraction: first, the small-sized layer Image(T) is converted from its original format to HSV format to obtain the converted target image Image_hsv; then, the S-channel (saturation channel) histogram and the V-channel (brightness channel) histogram are extracted from Image_hsv respectively, and the two sets of histograms are denoted as Hist_s (i.e., color saturation histogram) and Hist_v (i.e., grayscale histogram).

[0183] Finally, the extracted Hist_s and Hist_v image features are output and transmitted to the CPU for subsequent target parameter calculation.

[0184] In this embodiment, due to the reduced resolution of the second video frame layer, its data volume is greatly reduced, which significantly reduces the computing resources required for subsequent operations such as scene switching detection and image feature extraction, effectively shortening the processing time and meeting the stringent latency requirements of the first processing unit 200 being located after the display software pipeline framework.

[0185] Furthermore, lower resolution images have less impact on the representation of some global features, such as color distribution and overall texture, and can still provide sufficient information for scene transition detection. This can effectively avoid feature loss or misjudgment caused by reduced resolution, ensuring the accuracy and reliability of scene transition detection results.

[0186] In another embodiment of this application, reference is made to Figure 6 This is a schematic diagram of the processing data flow of a system-level processing chip provided in Embodiment 4 of this application, as shown below. Figure 6 As shown, the system-level processing chip may also include: a third processing unit 400.

[0187] The third processing unit 400 can be used to obtain the service scenario information corresponding to the decoded video frame; the service scenario information includes image size specifications.

[0188] In practical applications, business scenario information can be obtained through multiple channels, not limited to a single source. The appropriate acquisition method can be flexibly selected based on the specific business scenario. For example, in a video playback scenario, image size specifications can be directly extracted from the video file's metadata. The video file's metadata usually pre-records core parameters such as video resolution and aspect ratio, which can be quickly obtained without additional computation. Alternatively, it can be extracted from the decoding parameters output by the video decoding unit 100. After completing the decoding operation of the compressed video stream, the video decoding unit 100 synchronously outputs the decoded video frames (original pixel data) and corresponding decoding parameters. Image size specifications, as one of the core contents of the decoding parameters, can be directly extracted from this data.

[0189] In this embodiment, it is not necessary to extract the image size specifications from the decoded video frames because the decoded video frames are only a set of raw pixels. If the image size specifications are to be extracted from them, additional operations such as pixel parsing and row and column counting are required, which will increase the computational overhead of the third processing unit 400 and introduce additional processing latency, which is inconsistent with the design goals of low power consumption and low latency.

[0190] Whether the image is obtained from the video file metadata or extracted from the decoding parameters of the video decoding unit 100, the third processing unit 400 does not need to perform additional calculations. This allows for the rapid acquisition of the required image size specifications while avoiding additional power consumption and latency losses, thus ensuring processing efficiency.

[0191] Additionally, the third processing unit 400 can be used to determine whether the image size specification meets a preset resolution threshold.

[0192] The preset resolution threshold can be pre-set according to the characteristics of the display component and the image quality requirements. Its setting logic is adapted to the native resolution of the display component. For example, if the target display component is 1080P resolution, the preset resolution threshold can be set to 1080P; if the display component is 4K resolution, the preset resolution threshold can be set to 4K. The specific threshold can be flexibly adjusted according to the actual application scenario.

[0193] In this embodiment, the image size (original resolution) of the decoded video frame is compared with a preset resolution threshold. If the original resolution is lower than the preset resolution threshold, it means that the video frame is a low-resolution frame with defects such as blurry image and insufficient detail, and needs to be over-resolution processed. If the original resolution is not lower than the preset resolution threshold, it means that the resolution of the video frame meets the display requirements and does not need to be over-resolution processed. It can directly enter the subsequent display software pipeline framework for processing.

[0194] Furthermore, the third processing unit 400 can be used to perform super-resolution processing on the decoded video frame in response to the image size specification meeting the preset resolution threshold, so as to obtain a super-resolution processed video frame.

[0195] In this embodiment, if the third processing unit 400 determines that the image size of the decoded video frame is lower than a preset resolution threshold, this indicates that the video frame is a low-resolution frame.

[0196] Low-resolution frames often suffer from blurry images and loss of detail when displayed, severely impacting the user's viewing experience. To address this, the third processing unit 400 can perform super-resolution processing on the video frame. Specifically, super-resolution algorithms (such as SRCNN, ESPCN, EDSR, etc.) can be used to interpolate and reconstruct the low-resolution frame, increasing the number of pixels and optimizing the image structure. This results in a clearer presentation of the video content and reduced blurring and jagged edges.

[0197] After super-resolution processing, the original low-resolution video frames have a significantly improved image quality, which can better meet the requirements of display components and provide users with a better visual experience.

[0198] In this embodiment, the first processing unit 200 receives a first video frame layer generated after processing the decoded video frames, which may specifically include:

[0199] The first processing unit 200 receives a first video frame layer generated after processing the super-resolution video frames.

[0200] In this embodiment, the process of the first processing unit 200 receiving the first video frame layer will be adjusted accordingly with the addition of the third processing unit 400: if the third processing unit 400 performs super-resolution processing on the decoded video frame, the display software pipeline framework will perform the formatting processing, color calibration processing, and layer attribute encapsulation processing described above on the super-resolution processed video frame to generate the first video frame layer, and then transmit it to the first processing unit 200.

[0201] If the decoded video frame does not meet the preset resolution threshold (no over-resolution required), the display software pipeline framework directly performs the above preprocessing on the original decoded video frame to generate the first video frame layer and transmit it to the first processing unit 200.

[0202] In this embodiment, the third processing unit 400 presets a reasonable resolution threshold based on the characteristics of the display component and the actual image quality requirements. It then accurately judges and filters out low-resolution video frames and uses super-resolution algorithms such as SRCNN and ESPCN that are adapted to the hardware to process them. This effectively improves the resolution and image clarity of the low-resolution frames, compensates for the image quality defects caused by compression encoding, and enables the video frames to better adapt to the native resolution of the target display component, thus solving the problems of blurry and insufficient detail in low-resolution frame display.

[0203] Subsequently, the display software pipeline framework performs preprocessing operations such as formatting, color calibration, and layer attribute encapsulation on the super-resolution video frames to generate a standardized first video frame layer, which is then transmitted to the first processing unit 200. Upon receiving this first video frame layer, the first processing unit 200 performs scene detection, image feature extraction, and parameter calculation to output image quality enhancement target parameters adapted to the current scene. Upon receiving these target parameters, the second processing unit 300, leveraging its own hardware parallel computing advantages, performs targeted image quality enhancement processing on the video frames, further optimizing image color and detail. Throughout this process, the super-resolution optimization of the third processing unit 400, the preprocessing of the display software pipeline, the parameter calculation of the first processing unit 200, and the image quality enhancement of the second processing unit 300 work together efficiently. This not only strictly controls overall computational overhead and processing latency, reducing system power consumption, but also effectively improves video display quality, providing users with a clearer and smoother visual experience.

[0204] In another embodiment of this application, a system-level processing chip is provided in Embodiment 5 of this application. This embodiment is mainly an implementation method for super-resolution processing of the decoded video frames by the third processing unit 400. In this embodiment, the service scenario information may also include: application type.

[0205] Application type can reflect the differences in application requirements for video quality and real-time processing.

[0206] Application types can be categorized based on actual application scenarios, including but not limited to the following: video playback applications (such as online video apps and local video players), camera applications (such as system cameras and third-party photo / video apps), game applications (such as mobile games and console emulators), and gallery applications (such as local gallery and image browsing tools).

[0207] Different types of applications have different super-resolution requirements. For example, video playback applications prioritize a balance between image clarity and playback smoothness, needing to consider both super-resolution effects and processing latency; camera applications (especially real-time preview and recording scenarios) prioritize real-time processing, needing to minimize processing latency while ensuring basic super-resolution effects; game applications prioritize image detail and dynamic smoothness, needing to adapt to the rapid processing of dynamic video frames while improving image texture details; gallery applications prioritize the ultimate improvement of static image quality, are less sensitive to processing latency, and can use high-precision super-resolution models.

[0208] In this embodiment, the method for obtaining the application type is consistent with the logic for obtaining the image size specifications, without requiring additional computational overhead for the third processing unit 400. Specifically, the third processing unit 400 can obtain the currently running application type identifier from the application management module of the electronic device through the system interface built into the SOC. The application management module of the electronic device records the category information of currently active applications in real time, and the third processing unit 400 can quickly read the identifier through a dedicated data link without performing complex recognition calculations.

[0209] In addition, it can also be extracted from the additional information of the video frames output by the application. Some applications will carry their own type identifier when outputting video frames. The third processing unit 400 can directly parse the additional information to obtain the application type, ensuring that the acquisition process is efficient and has low latency.

[0210] The third processing unit 400 performs super-resolution processing on the decoded video frames, which may include, but is not limited to:

[0211] The third processing unit 400 performs super-resolution processing on the decoded video frames based on a target model adapted to the application type.

[0212] In this embodiment, the third processing unit 400 can pre-store a variety of commonly used super-resolution models. Each super-resolution model is optimized for the needs of a specific application type. This method does not rely on the network, has a fast calling speed, and no network latency. It is suitable for scenarios with unstable networks, no network coverage, or extremely high real-time requirements (such as offline video playback and real-time camera preview).

[0213] Of course, in scenarios where network conditions permit (such as when the device is connected to a stable network and the real-time requirements can be appropriately relaxed), the third processing unit 400 can also call the super-resolution model from the cloud server through the network interface built into the SOC. The cloud can store richer and more cutting-edge super-resolution models (including customized models optimized for specific application types), and the cloud models can be updated and iterated in real time. Without the need for hardware upgrades or local model updates to the third processing unit 400, it can adapt to new application types or optimize super-resolution effects, further improving the scalability and adaptability of the solution.

[0214] Super-resolution models adapted to different types of applications can include, but are not limited to:

[0215] For video playback applications, a lightweight super-resolution model (such as the optimized ESPCN model) can be adapted. This model simplifies the calculation process and reduces the amount of computation while ensuring the basic super-resolution effect (improving resolution and restoring details). It balances image quality and real-time performance, avoids playback stuttering caused by super-resolution processing, and is suitable for scenarios requiring continuous video playback.

[0216] For camera applications (real-time preview / recording scenarios), it can be adapted to high-speed super-resolution models (such as the lightweight SRCNN model). This model focuses on fast computation, which greatly reduces the latency of single-frame super-resolution processing, ensuring smooth and lag-free camera preview, while making up for the image quality defects of low-resolution preview frames and improving the preview and recording quality.

[0217] For gaming applications, dynamic super-resolution models (such as the EDSR lightweight variant model) can be adapted. This model optimizes motion blur suppression logic for the characteristics of dynamic game visuals, improving screen resolution while reducing jagged edges and blurring in dynamic frames, balancing dynamic smoothness and detail, and meeting the high frame rate requirements of gaming scenarios.

[0218] For image gallery applications, high-precision super-resolution models (such as the full version of the EDSR model) can be adapted. This model has higher computational complexity, but better super-resolution effect. It can accurately restore the texture details of static images and optimize color performance. Since image gallery scenes do not require real-time processing and have low latency requirements, the advantages of high-precision models can be fully utilized to achieve the ultimate improvement in image quality of static images.

[0219] In this embodiment, the third processing unit 400 may include, but is not limited to, an NPU (Neural Processing Unit) or a DSP (Digital Signal Processor). Both can adapt to the computational requirements of super-resolution models and are suitable for different hardware design scenarios. An NPU is adept at parallel computation of neural network models and is suitable for high-precision, multi-model super-resolution processing scenarios; while a DSP has efficient digital signal processing capabilities and can be adapted to lightweight super-resolution models through hardware-level optimization, balancing computational efficiency and power consumption control.

[0220] In this embodiment, combined with Figure 7 The processing flow of the third processing unit 400 is described below. Figure 7 As shown, after obtaining the business scenario information (including image size specifications and application type), the third processing unit 400 first performs a resolution judgment step, that is, it judges whether the image size specifications of the decoded video frame are greater than 1080P (i.e., an implementation of a preset resolution threshold).

[0221] If the judgment result is yes, it means that the resolution of the video frame has met the display requirements and there is no need for super-resolution processing. The third processing unit 400 directly outputs the original decoded video frame and transmits it to the subsequent display software pipeline framework to enter the subsequent preprocessing process.

[0222] If the judgment result is negative, it means that the video frame is a low-resolution frame and needs to be super-resolution processed. At this time, the third processing unit 400 selects the appropriate target model from multiple super-resolution models according to the obtained application type, uses the target model to perform super-resolution processing on the decoded video frame, obtains the super-resolution processed video frame, and finally outputs the super-resolution processed video frame to the subsequent display software pipeline framework.

[0223] In this embodiment, super-resolution processing of decoded video frames is performed by combining application type. Different types of applications are adapted to specific optimized super-resolution models, taking into account the differentiated requirements of various applications for image quality and real-time performance. At the same time, the application type is efficiently obtained through the SOC interface or video frame additional information, reducing computational overhead. It supports both local pre-stored models and dynamic cloud calls, which not only ensures processing efficiency in offline or high real-time scenarios, but also expands adaptability and update capabilities with the help of cloud models. Finally, through resolution judgment and model matching process, accurate super-resolution of low-resolution frames and direct output of high-resolution frames are achieved, effectively improving the flexibility, efficiency and image quality performance of super-resolution processing.

[0224] In another embodiment of this application, a system-level processing chip is provided in Embodiment 6 of this application. This embodiment is mainly an implementation of the first processing unit 200 extracting image features from the first video frame layer. In this embodiment, the second processing unit 300 may include multiple target units; the target units are used to perform image quality enhancement processing.

[0225] Each target unit has an independent computing array, which adopts a fixed-size parallel computing architecture and can efficiently process image data of a specific size.

[0226] The target unit may specifically include: the ACAD module (Adaptive Color and Detail Enhancement Module), which is the core computing unit in the second processing unit 300 responsible for performing specific image quality enhancement operations.

[0227] Extracting image features from the first video frame layer may specifically include:

[0228] The first video frame is divided into multiple image blocks according to the image size that matches the operation array of the target unit;

[0229] Image features are extracted from each of the image blocks.

[0230] The fixed size of the computing array determines that it performs best for image data of a specific size. During design, the hardware structure and computing logic of the computing array are optimized for image blocks of specific sizes, such as 8×8, 16×16, and 128×128 pixel blocks. The layout of its internal computing units, data transmission paths, and the adaptability of its computing logic are all customized around this fixed size. This enables precise capture and efficient processing of each pixel information within an image block of that size, maximizing the preservation of local image details and restoring image quality characteristics.

[0231] Specifically, for image blocks of a preset fixed size, the computing array can achieve pixel-level parallel computing without the need for additional splitting, splicing, or size conversion of image data. It can directly adapt to the feature distribution of the image block and accurately respond to local image quality enhancement needs. For example, in operations such as texture detail compensation, color saturation adjustment, and local brightness optimization, it can achieve more delicate and accurate enhancement processing based on the features of the image block of this size, avoiding enhancement deviation, loss of detail, or image quality distortion caused by size mismatch.

[0232] Conversely, if the size of the input image patch does not match the fixed size of the computing array, even if it is adjusted to a suitable size through scaling, cropping, or other means, it will destroy the original local feature correlation of the image, causing the computing array to fail to fully utilize its hardware optimization advantages, thereby affecting the accuracy of image enhancement and failing to achieve the best enhancement effect.

[0233] In this embodiment, the first processing unit 200 divides the image blocks according to the fixed size of the computing array. This is to ensure that each image block can be adapted to the hardware optimization design of the computing array, so that the computing array can give full play to its processing advantages for specific sizes. This provides a better processing foundation for the target unit (ACAD module) of the second processing unit 300, ensuring that each image block can obtain accurate and efficient image quality enhancement, and ultimately improving the image quality performance of the entire video frame layer.

[0234] For each image block obtained after segmentation, image features can be extracted separately using the same method as in the embodiments described above. Since the foregoing embodiments have already described this in detail, it will not be repeated here.

[0235] The first processing unit 200 can determine the target parameters corresponding to each image block based on the image features of each image block. In this embodiment, the target parameters corresponding to each image block can be determined in a manner consistent with the embodiments described above. Since the foregoing embodiments have already provided a detailed description of this, it will not be repeated here.

[0236] After receiving the target parameters corresponding to each image block transmitted by the first processing unit 200, the second processing unit 300 can allocate the target parameters corresponding to each image block to the corresponding target unit (one target unit processes one image block). After receiving the corresponding image block and target parameters, each target unit does not need to perform redundant operations such as size scaling and data conversion, but directly performs targeted enhancement based on the target parameters corresponding to the image block.

[0237] In this embodiment, there are no restrictions on the processing methods between the target units. For example, the target units can be processed serially, that is, multiple target units process their respective image blocks in a preset order. After the previous target unit completes the enhancement processing of its own image block, the next target unit starts the processing flow. Even with serial processing, each target unit still focuses only on its corresponding image block, relying on the hardware optimization advantages of its own computing array and its exclusive target parameters to perform targeted enhancement processing. The enhancement accuracy of a single image block will not be affected by the order of processing.

[0238] After all target units have completed the enhancement processing of their respective image blocks, the second processing unit 300 will stitch together all the enhanced image blocks to restore the complete enhanced first video frame layer.

[0239] In this embodiment, because the computational array of the target unit of the second processing unit 300 is optimized for image blocks of a specific size, the first video frame is divided into blocks to extract features. After the blocks are divided, the first processing unit 200 can determine the target parameters, enabling the computational array to accurately and effectively process each image block according to the target parameters, thus ensuring the image quality enhancement effect. Finally, the enhanced image blocks are stitched together to restore the complete frame, ensuring the integrity of the output video frame and achieving a better overall image quality enhancement effect.

[0240] In another embodiment of this application, a system-level processing chip is provided in Embodiment 7 of this application. This embodiment is mainly an implementation method for the second processing unit 300 to perform image quality enhancement processing on the first video frame layer according to the target parameters. Specifically, it may include, but is not limited to:

[0241] The second processing unit 300 performs image quality enhancement processing on each of the multiple target units in parallel according to the target parameters corresponding to each of the image blocks.

[0242] The multiple target units within the second processing unit 300 are each independently responsible for the image quality enhancement operation of one image block. They do not interfere with each other or wait for each other. Based on the hardware advantages of each target unit, they complete the targeted enhancement of multiple image blocks in parallel.

[0243] After all target units have completed the enhancement processing of their respective image blocks, the second processing unit 300 stitches together all the enhanced image blocks to restore the complete enhanced first video frame layer.

[0244] During the stitching process, the edge transitions of adjacent image blocks can be optimized to avoid stitching marks, color banding, and other issues, ensuring that the stitched video frame layers are coherent and have consistent image quality. Finally, the complete layer is output to the display components of electronic devices to complete the entire image enhancement process.

[0245] In this embodiment, under parallel processing, each target unit focuses only on enhancing a single image patch without having to consider other image patches. This allows for a full focus on the image quality requirements of that image patch, achieving precise enhancement based on the hardware advantages of the computing array. This avoids local enhancement deviations that may occur with global uniform enhancement or sequential processing, further improving the enhancement effect of each image patch.

[0246] The image processing method provided in this application will be described below. The image processing method described below can be referred to in correspondence with the system-level processing chip described above.

[0247] Reference Figure 8 This is a flowchart illustrating an image processing method provided in Embodiment 8 of this application, as shown below. Figure 8As shown, the method may include, but is not limited to, the following steps:

[0248] Step S101: The first processing unit receives the first video frame layer generated after processing the decoded video frames.

[0249] Step S102: Based on the first video frame layer, the first processing unit performs video scene switching detection to obtain the detection result.

[0250] Step S103: Based on the first processing unit's response to the detection result indicating a scene change in the video frame content, extract image features from the first video frame layer.

[0251] Step S104: Based on the detection results obtained by the first processing unit, determine the target parameters according to the image features.

[0252] Step S105: Based on the target parameters, the second processing unit performs image quality enhancement processing on the first video frame layer and outputs the image quality enhanced first video frame layer to the display component.

[0253] The detailed execution process of steps S101 to S105 can be referred to the relevant functional descriptions of the corresponding processing units in the system-level processing chip above. The specific implementation details of each step are consistent with those described above, and will not be repeated here.

[0254] As another optional embodiment of this application, a text recognition method provided in Embodiment 9 of this application, this embodiment is mainly an implementation of the above-mentioned step S102, which may include but is not limited to:

[0255] Step S1021: Determine the first frame global feature distribution data of the first video frame layer; the first frame global feature distribution data characterizes the global attributes of the first video frame layer;

[0256] Step S1022: Obtain the global feature distribution data of the second screen of the previous video frame of the first video frame layer and the global feature distribution data of the third screen of the starting frame after the last scene switch;

[0257] Step S1023: Determine the first similarity between the global feature distribution data of the first screen and the global feature distribution data of the second screen;

[0258] Step S1024: Determine the second similarity between the global feature distribution data of the first screen and the global feature distribution data of the third screen;

[0259] Step S1025: In response to the fact that both the first similarity and the second similarity satisfy the set similarity threshold, it is determined that the video screen content has changed scene, and a detection result representing the scene change is obtained.

[0260] The detailed execution process of steps S1021 to S1025 can be referred to the relevant description in Embodiment 2 above. The specific implementation details of each step are consistent with those described in Embodiment 2, and will not be repeated here.

[0261] As another optional embodiment of this application, a text recognition method provided in Embodiment 10 of this application, this embodiment is mainly an implementation of the above-mentioned step S102, which may include but is not limited to:

[0262] Step S1026: The first processing unit obtains a second video frame layer based on the first video frame layer; the resolution of the second video frame layer is less than the resolution of the first video frame layer.

[0263] Step S1027: The first processing unit performs video scene switching detection based on the second video frame layer to obtain the detection result.

[0264] The detailed process of steps S1026-S1027 can be referred to the relevant description in the aforementioned embodiment 3. The specific implementation details of each step are consistent with those described in embodiment 3, and will not be repeated here.

[0265] In this embodiment, the first processing unit extracts image features from the first video frame layer, which may specifically include:

[0266] Step S1031: The first processing unit extracts image features from the second video frame layer.

[0267] The detailed process of step S1031 can be referred to the relevant description in the aforementioned embodiment 3. The specific implementation details of each step are consistent with those described in embodiment 3, and will not be repeated here.

[0268] As another optional embodiment of this application, a text recognition method is provided in Embodiment 11 of this application. This method is based on a system-level processing chip, which includes: a video decoding unit, a first processing unit, a second processing unit, and a third processing unit. The method may include, but is not limited to:

[0269] Step S201: Obtain the service scenario information corresponding to the decoded video frame based on the third processing unit; the service scenario information includes image size specifications.

[0270] Step S202: Based on the third processing unit, determine whether the image size specification meets the preset resolution threshold.

[0271] Step S203: Based on the fact that the image size specification meets the preset resolution threshold, the third processing unit performs super-resolution processing on the decoded video frame to obtain the super-resolution processed video frame.

[0272] Step S204: The first processing unit receives the first video frame layer generated after processing the super-resolution video frame.

[0273] Step S205: Based on the first video frame layer, the first processing unit performs video scene switching detection to obtain the detection result.

[0274] Step S206: Based on the first processing unit's response to the detection result indicating a scene change in the video frame content, extract image features from the first video frame layer.

[0275] Step S207: Based on the detection results obtained by the first processing unit, determine the target parameters according to the image features.

[0276] Step S208: Based on the target parameters, the second processing unit performs image quality enhancement processing on the first video frame layer and outputs the image quality enhanced first video frame layer to the display component.

[0277] The detailed process of steps S201-S208 can be referred to the relevant description in the aforementioned embodiment 4. The specific implementation details of each step are consistent with those described in embodiment 4, and will not be repeated here.

[0278] As another optional embodiment of this application, a text recognition method provided in embodiment 12 of this application, in this embodiment, the business scenario information may further include the application type. This embodiment is mainly an implementation of the above-mentioned super-resolution processing of the decoded video frames, specifically including:

[0279] Step S2031: Perform super-resolution processing on the decoded video frames based on the target model adapted to the application type.

[0280] The detailed process of step S2031 can be referred to the relevant description in the aforementioned embodiment 5. The specific implementation details of each step are consistent with those described in embodiment 5, and will not be repeated here.

[0281] As another optional embodiment of this application, a text recognition method provided in embodiment 13 of this application, in this embodiment, the second processing unit includes a plurality of target units; the target units are used for image quality enhancement processing. This embodiment is mainly an implementation of the above-mentioned extraction of image features from the first video frame layer, and may specifically include:

[0282] Step S1032: Divide the first video frame into multiple image blocks according to the image size that matches the computational array of the target unit.

[0283] Step S1033: Extract image features from each of the image blocks.

[0284] The detailed process of steps S1032-S1033 can be referred to the relevant description in the aforementioned embodiment 6. The specific implementation details of each step are consistent with those described in embodiment 6, and will not be repeated here.

[0285] As another optional embodiment of this application, a text recognition method provided in embodiment 14 of this application is mainly an implementation of the above-mentioned image quality enhancement processing of the first video frame layer according to the target parameters, and may specifically include:

[0286] Step S1051: Based on the multiple target units, each image block is subjected to image quality enhancement processing in parallel according to the target parameters corresponding to each image block.

[0287] The detailed process of step S1051 can be referred to the relevant description in the aforementioned embodiment 7. The specific implementation details of each step are consistent with those described in embodiment 7, and will not be repeated here.

[0288] In another embodiment of this application, an electronic device is provided, which may specifically include, but is not limited to:

[0289] The system-level processing chip includes a first processing unit and a second processing unit.

[0290] The first processing unit is configured to:

[0291] Receive the first video frame layer generated after processing the decoded video frames;

[0292] Based on the first video frame layer, video scene content switching detection is performed to obtain the detection results;

[0293] In response to the detection result indicating a scene change in the video content, image features are extracted from the first video frame layer;

[0294] Based on the detection results, target parameters are determined according to the image features.

[0295] The second processing unit is used to perform image quality enhancement processing on the first video frame layer according to the target parameters, and output the image quality enhanced first video frame layer to the display component.

[0296] The display component is connected to the second processing unit and is used to display the first video frame layer after the image quality enhancement processing.

[0297] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0298] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0299] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0300] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A system-level processing chip, comprising: The first processing unit is used for: Receive the first video frame layer generated after processing the decoded video frames; Based on the first video frame layer, video scene content switching detection is performed to obtain the detection results; In response to the detection result indicating a scene change in the video content, image features are extracted from the first video frame layer; Based on the detection results, the target parameters are determined according to the image features; The second processing unit is used to perform image quality enhancement processing on the first video frame layer according to the target parameters, and output the image quality enhanced first video frame layer to the display component.

2. The system-level processing chip according to claim 1, wherein the first processing unit performs video scene content switching detection based on the first video frame layer to obtain a detection result, including: Determine the global feature distribution data of the first frame of the first video frame layer; The global feature distribution data of the first frame represents the global attributes of the first video frame layer; Obtain the global feature distribution data of the second screen of the previous video frame of the first video frame layer and the global feature distribution data of the third screen of the starting frame after the last scene switch; Determine the first similarity between the global feature distribution data of the first screen and the global feature data of the second screen; Determine the second similarity between the global feature distribution data of the first screen and the global feature distribution data of the third screen; In response to both the first similarity and the second similarity satisfying the set similarity threshold, it is determined that a scene change has occurred in the video content, and a detection result representing the scene change is obtained.

3. The system-level processing chip according to claim 1, wherein the first processing unit performs video scene content switching detection based on the first video frame layer to obtain a detection result, including: The first processing unit obtains a second video frame layer based on the first video frame layer; The resolution of the second video frame layer is smaller than the resolution of the first video frame layer; The first processing unit performs scene switching detection of video content based on the second video frame layer and obtains the detection result; The first processing unit extracts image features from the first video frame layer, including: The first processing unit extracts image features from the second video frame layer.

4. The system-level processing chip according to claim 1, further comprising: The third processing unit is used for: Obtain the service scenario information corresponding to the decoded video frame; The business scenario information includes image size specifications; Determine whether the image size specification meets the preset resolution threshold; In response to the image size specification meeting the preset resolution threshold, super-resolution processing is performed on the decoded video frame to obtain a super-resolution processed video frame. The first processing unit receives a first video frame layer generated after processing the decoded video frames, including: The first processing unit receives a first video frame layer generated after processing the super-resolution video frames.

5. The system level processing chip of claim 4, the traffic scenario information further comprising: Application type; The third processing unit performs super-resolution processing on the decoded video frames, including: The third processing unit performs super-resolution processing on the decoded video frames based on a target model adapted to the application type.

6. The system level processing chip of claim 1, the second processing unit comprising a plurality of target units; The target unit is used for image quality enhancement processing; Extracting image features from the first video frame layer includes: The first video frame is divided into multiple image blocks according to the image size that matches the operation array of the target unit; Image features are extracted from each of the image blocks.

7. The system-level processing chip according to claim 6, wherein the second processing unit performs image quality enhancement processing on the first video frame layer according to the target parameters, comprising: Based on the multiple target units, each image block is subjected to image quality enhancement processing in parallel according to the target parameters corresponding to each image block.

8. An image processing method, comprising: The first video frame layer is generated by processing the decoded video frames received by the first processing unit. Based on the first processing unit's first video frame layer, video scene content switching detection is performed to obtain the detection result; Based on the first processing unit's response to the detection result indicating a scene change in the video frame content, image features are extracted from the first video frame layer; Based on the detection results obtained by the first processing unit, target parameters are determined according to the image features. The second processing unit performs image quality enhancement processing on the first video frame layer according to the target parameters, and outputs the image quality enhanced first video frame layer to the display component.

9. The image processing method according to claim 8, wherein the step of the first processing unit performing video scene content switching detection based on the first video frame layer to obtain a detection result includes: Determine the global feature distribution data of the first frame of the first video frame layer; The global feature distribution data of the first frame represents the global attributes of the first video frame layer; Obtain the global feature distribution data of the second screen of the previous video frame of the first video frame layer and the global feature distribution data of the third screen of the starting frame after the last scene switch; Determine the first similarity between the global feature distribution data of the first screen and the global feature data of the second screen; Determine the second similarity between the global feature distribution data of the first screen and the global feature distribution data of the third screen; In response to both the first similarity and the second similarity satisfying the set similarity threshold, it is determined that a scene change has occurred in the video content, and a detection result representing the scene change is obtained.

10. An electronic device, comprising: The system-level processing chip includes a first processing unit and a second processing unit; The first processing unit is configured to: Receive the first video frame layer generated after processing the decoded video frames; Based on the first video frame layer, video scene content switching detection is performed to obtain the detection results; In response to the detection result indicating a scene change in the video content, image features are extracted from the first video frame layer; Based on the detection results, the target parameters are determined according to the image features; The second processing unit is used to perform image quality enhancement processing on the first video frame layer according to the target parameters, and output the image quality enhanced first video frame layer to the display component; The display component is connected to the second processing unit and is used to display the first video frame layer after the image quality enhancement processing.