Method and system for video display optimization on mobile phone screen based on vision processor
By analyzing video content with a visual processor, segmenting frame regions, and optimizing display strategies by combining environmental and user data, the problem of insufficient accuracy in visual quality enhancement in existing technologies is solved, thereby improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN HONGSHENGTONG PHOTOELECTRIC CO LTD
- Filing Date
- 2025-09-30
- Publication Date
- 2026-04-10
AI Technical Summary
Existing methods for optimizing mobile phone screen video display cannot be personalized based on the characteristics of video content, user environment, and focus, resulting in insufficient accuracy in enhancing visual quality and failing to provide a high-quality user experience.
By using a visual processor to analyze video content, a multidimensional dataset is built, video frame regions are segmented, environmental sensors and gaze sensors are activated, attention markers and environmental datasets are built, and video enhancement strategies are corrected through feedback stimulation to optimize display effects.
It improves the accuracy of visual quality enhancement, enhances the user viewing experience, and enables personalized display optimization based on video content and user needs.
Smart Images

Figure CN121170680B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of display optimization, in particular to a mobile phone screen video display optimization method and system based on a vision processor. BACKGROUND
[0002] With the rapid growth of mobile video consumption, the video display quality of mobile phone screens has become a key factor affecting user experience. At present, for mobile phone screen video display optimization, methods based on general video encoding algorithms and fixed display parameter adjustment are mainly used. Through unified encoding standards, videos are processed, and simple adjustments are made to the display effect according to preset fixed parameters, such as uniform adjustment of brightness, contrast, etc. However, the existing method does not fully consider the characteristic differences of video content, and cannot perform targeted optimization for videos of different types of content, resulting in problems such as loss of picture quality and stuttering when processing complex videos. At the same time, fixed display parameter adjustment cannot be dynamically adjusted in real time according to the environment in which the user is located and the user's focus, so the display effect cannot be well adapted to different use scenarios and user needs, and personalized high-quality display experience cannot be provided.
[0003] At the present stage, the mobile phone screen video display optimization in the related art has the technical problem of insufficient precision of visual quality enhancement. SUMMARY
[0004] The present application provides a mobile phone screen video display optimization method and system based on a vision processor, which analyzes video content using a vision processor, evaluates frame complexity and establishes a multidimensional dataset, thereby segmenting video frame regions, and performs enhanced processing and analysis on the region segmentation results and the multidimensional dataset to establish a first video enhancement strategy. An environment sensor and a gaze sensor are activated, a focus identifier and an environment dataset are established, the first video enhancement strategy is simulated with the focus identifier and the environment dataset, a feedback incentive is established, the feedback incentive is used to correct the strategy, a second video enhancement strategy is established and used for optimization, and other technical means. The technical problem of insufficient precision of visual quality enhancement in the existing mobile phone screen video display optimization is solved, and the technical effects of improving the precision of visual quality enhancement and improving user viewing experience are achieved.
[0005] The application provides a visual processor-based mobile phone screen video display optimization method, which comprises the following steps: performing real-time analysis on video content by using a visual processor, identifying dynamic elements, static backgrounds and text information, and performing complexity evaluation on each frame of the video content, and establishing a multidimensional data set according to the complexity evaluation result and the identification result; performing video frame region segmentation by using the identification result in the multidimensional data set, and establishing a region segmentation result; performing enhancement processing analysis on the region segmentation result and the multidimensional data set, and establishing a first video enhancement strategy; activating an environment sensor and a gaze sensor, and establishing a focus identifier and an environment data set; performing enhancement simulation on the first video enhancement strategy, the focus identifier and the environment data set, and establishing a feedback incentive; performing strategy correction on the first video enhancement strategy by using the feedback incentive, establishing a second video enhancement strategy, and performing display optimization management according to the second video enhancement strategy.
[0006] In a possible implementation, the region segmentation result and the multidimensional data set are subjected to enhancement processing analysis to establish a first video enhancement strategy, and the following processing is performed: after the region segmentation result is identified according to the identification result in the multidimensional data set, multi-scale convolutional neural network multi-level feature extraction under the identification constraint is performed on the region segmentation result to establish a feature extraction result; the power consumption mode of the mobile phone is read, and a power consumption constraint is established according to the power consumption mode and the complexity evaluation result; the power consumption constraint is taken as a whole power consumption target, feature enhancement optimization under the identification constraint is performed according to the feature extraction result, and the first video enhancement strategy is established according to the feature enhancement optimization result.
[0007] In a possible implementation, the power consumption constraint is taken as a whole power consumption target, and feature enhancement optimization under the identification constraint is performed according to the feature extraction result, and the following processing is performed: a smoothing feedback layer is established, the smoothing feedback layer is an adaptive expansion frame interval with the current video frame as a center point; frame display transition smoothing analysis in the feature enhancement optimization process is performed by using the smoothing feedback layer to establish an abnormal distribution; a transition abnormal feedback is configured by using the abnormal distribution, and guiding compensation management of the feature enhancement optimization is performed according to the transition abnormal feedback.
[0008] In a possible implementation, guiding compensation management of the feature enhancement optimization is performed according to the transition abnormal feedback, and the following processing is performed: a balance verification window is set according to the transition abnormal feedback; after iteration of the guiding compensation management is performed, global balance effect verification of the feature enhancement optimization is performed by using the balance verification window; if the global balance effect fails to meet a balance threshold within a predetermined iteration period, a power consumption breaking instruction is generated; after the whole power consumption target is adjusted according to the power consumption breaking instruction, the feature enhancement optimization management is continued.
[0009] In a possible implementation, the first video enhancement strategy and the attention identifier and the environment data set are subjected to enhancement simulation, feedback incentives are established, and the following processing is performed: attention stability analysis is performed on the attention identifier, and an attention stability coefficient is established; the position importance of the video frame is reconstructed using the attention identifier and the attention stability coefficient; attention mechanism analysis is performed using the position importance and the content of the video frame, and an attention importance is established; the first video enhancement strategy is simulated based on the attention importance and the environment data set, and feedback incentives are established.
[0010] In a possible implementation, the second video enhancement strategy is used for display optimization management, and the following processing is performed: real power consumption mapped by the second video enhancement strategy is obtained; mode matching analysis is performed according to the real power consumption and the power consumption mode, and a matching abnormality early warning is established; the matching abnormality early warning is subjected to power consumption early warning display prompting, and prompt feedback is received; and the second video enhancement strategy is adjusted according to the prompt feedback.
[0011] In a possible implementation, the region segmentation result and the multi-dimensional data set are subjected to enhancement processing analysis, and the first video enhancement strategy is established, and the following processing is further performed: video scene classification is performed on the video content, and a scene classification identifier is established; a scene enhancement template is configured according to the scene classification identifier, and enhancement optimization is performed based on the scene enhancement template as a basic strategy, and the first video enhancement strategy is established.
[0012] In a possible implementation, the second video enhancement strategy is used for display optimization management, and the following processing is performed: a user evaluation data set mapped by the second video enhancement strategy is established; the user evaluation data set, the second video enhancement strategy, and an enhancement scene are stored in association as a feedback optimization data set, and the enhancement scene includes a video scene, an environment scene, and an attention scene; an enhancement preference of a user is configured according to the feedback optimization data set, and subsequent video enhancement management is performed according to the enhancement preference.
[0013] In a possible implementation, the first video enhancement strategy and the attention identifier and the environment data set are subjected to enhancement simulation, feedback incentives are established, and the following processing is further performed: a first attention subject, a second attention subject, and other subjects are determined according to the attention identifier; enhancement simulation under environment data set adaptation is performed using the first attention subject, the second attention subject, and the other subjects, to establish feedback incentives.
[0014] The application also provides a mobile phone screen video display optimization system based on a vision processor, comprising: a multidimensional data set establishment module, configured to perform real-time analysis on video content by using the vision processor, identify dynamic elements, static backgrounds and text information, and perform complexity evaluation on each frame of the video content, and establish a multidimensional data set according to the complexity evaluation result and the identification result; a video frame region segmentation module, configured to perform video frame region segmentation by using the identification result in the multidimensional data set, and establish a region segmentation result; an enhancement processing analysis module, configured to perform enhancement processing analysis on the region segmentation result and the multidimensional data set, and establish a first video enhancement strategy; a sensor activator, configured to activate an environmental sensor and a gaze sensor, and establish a focus identifier and an environmental data set; an enhancement simulation module, configured to perform enhancement simulation on the first video enhancement strategy, the focus identifier and the environmental data set, and establish a feedback incentive; and a display optimization management module, configured to perform strategy correction of the first video enhancement strategy by using the feedback incentive, establish a second video enhancement strategy, and perform display optimization management according to the second video enhancement strategy.
[0015] The mobile phone screen video display optimization method and system based on a vision processor provided in the application first perform real-time analysis on video content by using the vision processor, identify dynamic elements, static backgrounds and text information, perform complexity evaluation on each frame of the video content, establish a multidimensional data set according to the complexity evaluation result and the identification result, then perform video frame region segmentation by using the identification result in the multidimensional data set, establish a region segmentation result, perform enhancement processing analysis on the region segmentation result and the multidimensional data set, establish a first video enhancement strategy, activate an environmental sensor and a gaze sensor, establish a focus identifier and an environmental data set, then perform enhancement simulation on the first video enhancement strategy, the focus identifier and the environmental data set, establish a feedback incentive, finally perform strategy correction of the first video enhancement strategy by using the feedback incentive, establish a second video enhancement strategy, and perform display optimization management according to the second video enhancement strategy. The technical effect of improving the accuracy of visual quality enhancement and improving the user viewing experience is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings of the embodiments of the application will be briefly introduced below. In the application, a flowchart is used to illustrate the operations performed by the system according to the embodiments of the application. It should be understood that the foregoing or the following operations are not necessarily performed in sequence. On the contrary, various steps can be processed in reverse order or simultaneously according to needs. Meanwhile, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0017] Figure 1A flowchart of a method for optimizing video display on a mobile phone screen based on a visual processor is provided in the embodiments of the present application.
[0018] Figure 2 A structural diagram of a system for optimizing video display on a mobile phone screen based on a visual processor is provided in the embodiments of the present application.
[0019] Legend: multidimensional data set establishment module 10, video frame region segmentation module 20, enhancement processing analysis module 30, sensor activation module 40, enhancement simulation module 50, and display optimization management module 60. DETAILED DESCRIPTION
[0020] To further illustrate the technical means and effects adopted by the present application to achieve the predetermined object, the specific embodiments, structures, features and effects according to the present application are described in detail below in combination with the drawings and preferred embodiments.
[0021] The embodiments of the present application provide a method for optimizing video display on a mobile phone screen based on a visual processor, as shown in Figure 1 The method comprises the following steps:
[0022] In step S100, a visual processor is used to perform real-time analysis on video content, identify dynamic elements, static backgrounds and text information, and perform complexity evaluation on each frame of the video content, and a multidimensional data set is established based on the complexity evaluation results and the identification results.
[0023] Specifically, the visual processor is a hardware module dedicated to image and video processing, such as an ISP, NPU or DSP in a mobile phone SoC. The visual processor performs pixel-level semantic segmentation and motion estimation on the input video frame by running a pre-trained deep learning model. For example, a target detection and segmentation model such as YOLOv5 or Mask R-CNN is used to identify moving objects, static background regions and text regions in the video. Among them, dynamic elements refer to regions that change in position, shape or brightness in the video frame, such as people and vehicles. Static backgrounds refer to background regions that do not change much in the video, such as the sky and buildings. The text region includes subtitles, interface text, etc. The complexity of each frame is calculated based on image entropy, motion vector amplitude, edge density, etc. to quantify the difficulty of frame processing, and finally the identification results and complexity values are integrated into structured data to form a multidimensional data set, which is stored as a JSON or time series database record.
[0024] A specific implementation method of video frame complexity evaluation is to use a hybrid calculation model based on image entropy and motion vector. This method converts the video frame from the RGB color space to the YUV color space and extracts the luminance component. The image entropy of the luminance component is calculated by counting the probability of each luminance value appearing in the entire frame, and then substituting these probability values into the information entropy calculation formula, i.e. summing the negative of each probability value multiplied by the logarithm of the probability value to the base of two. The image entropy measures the complexity of the texture within the frame. At the same time, the optical flow method is used to calculate the motion vector field between the current frame and the previous frame, and the average value and standard deviation of the amplitude are counted to quantify the intensity and consistency of the inter-frame motion. Finally, the overall complexity of the frame is obtained by a weighted formula, which linearly combines the image entropy value with the average value and standard deviation of the motion vector according to pre-set weight coefficients. The weight coefficients are empirical values obtained by training a large number of video samples. This comprehensive index can accurately reflect both the static information content and the dynamic change intensity of the frame.
[0025] Step S200, using the recognition results in the multi-dimensional data set to perform video frame region segmentation, and establishing a region segmentation result.
[0026] Specifically, region segmentation refers to dividing an image into multiple regions with different semantic or functional attributes. Based on the recognition results of S100, the watershed algorithm or U-Net segmentation network is used to divide the video frame into multiple semantic regions. For example, dynamic regions are marked as high refresh areas, static backgrounds are marked as low refresh areas, and text regions are marked as high sharpening areas. The region segmentation result is saved in the form of a mask layer, aligned with the original frame coordinates, and each pixel value represents a region category.
[0027] Step S300, performing enhancement processing analysis on the region segmentation result and the multi-dimensional data set to establish a first video enhancement strategy.
[0028] Specifically, combined with the region attributes and complexity data, a differentiated enhancement strategy is formulated, and image processing algorithms are selected and configured according to the region characteristics. For example, super-resolution reconstruction is applied to text regions, motion compensation and frame interpolation are implemented for dynamic regions, and noise reduction and color enhancement are used for static backgrounds. The first video enhancement strategy is an initial set of video processing parameters, stored in the form of a parameter configuration file, including the sharpening intensity, brightness gain, frame rate control, etc. of each region.
[0029] In a possible implementation, the region segmentation result and the multi-dimensional data set are subjected to enhancement processing analysis to establish a first video enhancement strategy. Step S300 further includes step S310 of identifying the region segmentation result according to the recognition result in the multi-dimensional data set, performing multi-scale convolutional neural network multi-level feature extraction on the region segmentation result under the constraint of the identification, and establishing a feature extraction result. Specifically, the multi-scale convolutional neural network refers to a deep learning model that can capture multi-level features through convolution kernels of different sizes and can capture both local details and global structures of an image. The multi-level feature extraction refers to extracting features including edge, texture, object, and the like from the output of different levels of CNN. A multi-scale CNN such as ResNet or EfficientNet is used to extract shallow edge features and deep semantic features for each identified region. The input is frame data with region labels, and the feature maps are extracted at different levels by the multi-scale convolution kernel, and the feature vectors are stored in association with the region labels. For example, stroke detail features are extracted for text regions, and motion trajectory features are extracted for dynamic regions.
[0030] Step S320, read the power consumption mode of the mobile phone, and establish a power consumption constraint according to the power consumption mode and the complexity evaluation result. Specifically, the current power consumption mode is obtained through a system API. The power consumption mode is an energy consumption management strategy set by the mobile phone system, which affects the upper limit of hardware performance, such as power saving mode and performance mode. Combined with frame complexity data, the maximum allowed power threshold is calculated. For example, in the power saving mode, the GPU frequency and NPU load are limited, and frames with high complexity are allocated with a lower power budget.
[0031] In the process of calculating the maximum allowed power threshold, the video frame complexity score in step S100 is obtained, and the score is normalized to the range of zero to one. Then, a corresponding reference power value is selected according to the current power consumption mode, wherein the power saving mode corresponds to a lower reference value, the balanced mode takes an intermediate value, and the performance mode adopts a higher reference value. A mapping relationship between complexity and power adjustment coefficient is established, and a piecewise linear function is used for processing. A lower adjustment coefficient is allocated in a low complexity interval, and the adjustment coefficient is improved in a high complexity interval, but an upper limit is set to prevent excessive consumption. The final maximum allowed power threshold is obtained by multiplying the reference power value and the adjustment coefficient. At the same time, combined with the real-time state parameters of the device, such as the remaining battery capacity and temperature readings, the influence factors of these parameters are adjusted. For example, when the battery capacity is lower than 20%, an additional decay coefficient of 0.8 is applied to ensure that the power consumption is further limited in the low power condition. The threshold value calculated is transmitted to the feature enhancement optimization process as a hard constraint to ensure that all video processing operations are performed within the set power consumption range.
[0032] Step S330, according to the feature extraction result, perform feature enhancement optimization under the constraint of identification, and establish a first video enhancement strategy according to the feature enhancement optimization result, wherein the feature enhancement optimization refers to a process of searching for optimal image processing parameters under a limitation condition, and the constraint of identification refers to a differentiated limitation of enhancement parameters based on a region segmentation result. Reinforcement learning or genetic algorithm is used to search for an optimal combination of enhancement parameters under the constraint of power consumption. For example, a NSGA-II algorithm is used to solve a Pareto frontier with a peak signal-to-noise ratio (PSNR) and power consumption as double target optimization, and an optimal enhancement parameter meeting the constraint of power consumption is selected.
[0033] In a possible implementation, according to the feature extraction result, perform feature enhancement optimization under the constraint of identification, and step S330 further includes step S331 of establishing a smooth feedback layer, wherein the smooth feedback layer is an adaptive expansion frame interval with a current video frame as a center point. Specifically, the smooth feedback layer is a multi-frame buffer for analyzing inter-frame visual coherence. The adaptive expansion frame interval refers to a frame sequence with a dynamically adjusted buffer size according to a motion amplitude, and an interval length is determined according to inter-frame similarity, and the interval is reduced to reduce calculation amount when the similarity is high. The smooth feedback layer is composed of the current frame and N frames before and after the current frame, for example, N=5. A Farneback optical flow algorithm is used to analyze motion vectors of continuous frames, and a motion discontinuous region is marked as a potential flicker region.
[0034] Step S332, use the smooth feedback layer to perform inter-frame display transition smoothing analysis in the feature enhancement optimization process, and establish an abnormal distribution. Specifically, a difference between frames in the interval is calculated, a difference matrix of luminance and chrominance after enhancement of adjacent frames is calculated, and an abnormal jump pixel is detected through a threshold value. For example, a pixel with a luminance change of more than 50 nit / frame is marked as an abnormal pixel, and a binary abnormal distribution map is formed.
[0035] Step S333, use the abnormal distribution to configure a transition abnormal feedback, and perform guiding compensation management of the feature enhancement optimization according to the transition abnormal feedback. Specifically, the transition abnormal feedback is an enhancement parameter correction signal based on inter-frame discontinuity detection. The guiding compensation management is a mechanism of dynamically adjusting an optimization direction according to the feedback. The weight of the enhancement parameter is adjusted according to the abnormal distribution. For example, the sharpening intensity is reduced in an abnormal region, motion blur compensation is increased, and the parameter is iteratively optimized through a feedback loop until the proportion of abnormal pixels is lower than a threshold value.
[0036] In a possible implementation, the step S333 further includes a step S3331 of setting a balance verification window according to the transition abnormal feedback. Specifically, the balance verification window is a time segment for evaluating stability of the enhancement effect. The balance threshold is an upper limit of fluctuation of the power consumption and the quality. The balance verification window is a fixed-length video segment, such as 1 second, and the fluctuation of the power consumption and the visual quality standard deviation of the frames after enhancement are counted in the window. For example, the power consumption difference of each frame in the window cannot exceed 10%, and the PSNR fluctuation cannot exceed 2 dB.
[0037] The step S3332 includes a step of verifying the global balance effect of the feature enhancement optimization by using the balance verification window after each iteration of the guided compensation management. Specifically, after each parameter adjustment, all frames in the verification window are re-rendered, and the variances of the power consumption and the quality indicators are calculated. If the variances exceed the threshold, the iteration optimization is continued.
[0038] The step S3333 includes a step of generating a power consumption breaking instruction if the global balance effect fails to meet the balance threshold within a predetermined iteration period. Specifically, when the number of iterations exceeds the predetermined iteration period, such as 10 times, and the balance condition is still not met, the instruction is generated to allow temporarily increasing the upper limit of the power consumption, such as increasing by 5%, to stabilize the visual quality. The power consumption breaking instruction is a system instruction for temporarily relaxing the power consumption limit.
[0039] The step S3334 includes a step of continuing the feature enhancement optimization management according to the power consumption breaking instruction. Specifically, the optimization target function is adjusted according to the power consumption breaking instruction, and the optimization algorithm is re-run. For example, the power consumption constraint value is increased from the original value to 105%, and the NSGA-II optimization is continued until the balance condition is met.
[0040] In a possible implementation, the region segmentation result and the multi-dimensional data set are subjected to enhancement processing and analysis to establish a first video enhancement strategy. The step S300 further includes a step S340 of classifying video scenes of the video content and establishing a scene classification identifier. Specifically, a pre-trained scene classification model, such as MobileNetV3, is used to classify the video frames and identify categories such as night scene, motion, portrait, and text. For example, the scene probability distribution of each frame is output, and the maximum probability category is taken as the scene identifier.
[0041] Step S350, according to the scene classification, configure the scene enhancement template, and perform enhancement optimization based on the scene enhancement template to establish the first video enhancement strategy. Specifically, the scene enhancement template is a set of image processing parameters preset for a specific scene. The basic strategy is a general enhancement scheme without individual adjustment. For each type of scene, an enhancement template is predefined, such as a night scene template that prioritizes noise reduction and brightness enhancement, and a motion template that prioritizes frame interpolation and dynamic contrast. The template is used as the initial parameter, and then fine-tuned based on the feature extraction result to form the final strategy.
[0042] Step S400, activate the environment sensor and the gaze sensor, and establish the attention identifier and the environment data set.
[0043] Specifically, the environment sensor is a hardware module that collects environmental light, distance, and other data. The gaze sensor is a software module that detects the user's gaze point and the attention area. The light sensor and the distance sensor are called to obtain the environmental light intensity and the viewing distance; the user's gaze point is detected through the front camera and the eye tracker. For example, the MediaPipe Iris model is used to estimate the gaze point, and the sensor data is combined to generate a JSON format environment and attention data set.
[0044] Step S500, perform enhancement simulation on the first video enhancement strategy and the attention identifier and the environment data set to establish a feedback incentive.
[0045] Specifically, enhancement simulation refers to the pre-rehearsal of enhancement effects before output. The first video enhancement strategy is applied in the virtual rendering pipeline, and the visual experience is simulated in combination with the gaze point data. For example, GAN is used to generate simulated enhancement effects, and a scoring model is used to output a score as a feedback incentive based on the gaze dwell time and environmental light adaptability. The feedback incentive is a scoring signal that quantifies the pros and cons of the simulation effect. The specific implementation of the scoring model can be based on a pre-trained multi-task deep learning network that takes the video frames after simulation enhancement, the user's gaze point heat map, and the environmental light data as joint input. The first output branch performs visual quality evaluation, calculates the multi-scale structural similarity index between the enhanced frames and the ideal reference frames in the gaze area; the second output branch performs comfort prediction, analyzes the adaptability between environmental light intensity and screen brightness, and if the environmental light is very strong and the screen brightness is insufficient, or the environmental light is very dark and the screen brightness is too high, a negative score will be generated; the third output branch performs attention effectiveness analysis, models the gaze dwell time sequence through a recurrent neural network model, and judges whether the gaze behavior is effective interest attention or ineffective wandering. Finally, the scoring model weights and fuses the outputs of the three branches through a fully connected layer to generate a comprehensive feedback incentive score to quantify the comprehensive performance of the first video enhancement strategy in a specific environment and user attention.
[0046] In a possible implementation, the first video enhancement strategy is simulated based on the attention mark and the environment data set to establish the feedback incentive, and step S500 further includes step S510 of performing attention stability analysis on the attention mark to establish an attention stability coefficient. Specifically, the attention stability coefficient is an index quantifying the stability degree of the user's gaze. The dwell time and the jitter amplitude of the gaze point in a specific area are counted. For example, the standard deviation of the gaze point coordinates in the past 1 second is calculated, and if the standard deviation is less than 5 pixels, the stability coefficient is 1, otherwise the stability coefficient is attenuated in proportion.
[0047] Step S520, reconstructing the position importance of the video frame by using the attention mark and the attention stability coefficient. Specifically, the position importance is a weight map indicating the visual importance of each area in the frame. A Gaussian weight map is generated with the gaze point as the center, and the weight value is positively correlated with the stability coefficient. For example, when the stability coefficient is high, the standard deviation of the Gaussian kernel is reduced, and the focus area is more concentrated.
[0048] Step S530, performing attention mechanism analysis by using the position importance and the content of the video frame to establish the attention importance. Specifically, the position importance is fused with the semantic importance, such as the face and text detection scores. For example, a weighted sum is used to obtain a final attention importance map, which is used to guide the allocation of enhancement resources.
[0049] Step S540, performing adaptive simulation of the first video enhancement strategy based on the attention importance and the environment data set to establish the feedback incentive. Specifically, the enhancement parameter priority is adjusted according to the attention importance, such as preferentially improving the brightness of the high-attention area in a strong light environment. The simulation effect score is calculated by a scoring model as the feedback incentive.
[0050] In a possible implementation, the first video enhancement strategy is simulated based on the attention mark and the environment data set to establish the feedback incentive, and step S500 further includes step S550 of determining a first attention subject, a second attention subject and other subjects according to the attention mark. Specifically, the first attention subject is an object with the longest gaze time of the user, such as a human face. The second attention subject is a secondary gaze object, such as a gesture. The main gaze object, the secondary gaze object and other objects are identified through target detection and gaze trajectory analysis. For example, the subject priority is determined based on the YOLOv5 detection result and the gaze point overlap rate.
[0051] Step S560, performing enhancement simulation under the adaptation of the environment data set by using the first attention subject, the second attention subject and the other subjects to establish the feedback incentive. Specifically, different enhancement is implemented for different subjects, such as super-resolution for the first subject, color enhancement for the second subject, and original processing for the other subjects. The global color tone is adjusted in combination with the environment light data, and finally the feedback score is output by a scoring model.
[0052] Step S600, performing policy modification of the first video enhancement strategy using the feedback incentive, establishing a second video enhancement strategy, and performing display optimization management according to the second video enhancement strategy.
[0053] Specifically, the policy modification refers to a process of adjusting the enhancement parameters according to the feedback. The second video enhancement strategy is the final enhancement scheme after personalized optimization. Based on the feedback incentive, the enhancement parameters are modified using gradient descent or Bayesian optimization. For example, if the feedback shows that the dark details are lost, the gamma value of the shadow area is increased. The modified parameter set is delivered to the rendering pipeline as the second video enhancement strategy.
[0054] In a possible implementation, the display optimization management is performed according to the second video enhancement strategy, and step S600 further includes step S610 of acquiring real power consumption mapped by the second video enhancement strategy. Specifically, the real power consumption refers to the energy consumption value of the hardware actually executing the enhancement strategy. The GPU and NPU power consumption is monitored in real time by a hardware counter, and the energy consumption of a single frame processing is accumulated. For example, the power consumption data is collected using the Android BatteryManager and the Perfetto tool chain.
[0055] Step S620, mode matching analysis is performed according to the real power consumption and the power consumption mode, and a matching abnormality early warning is established. Specifically, it is verified whether the actual power consumption meets the mode requirement, and the measured power consumption is compared with the mode limit. If the power consumption exceeds the limit in the power saving mode, an early warning is triggered. The early warning information includes the exceeding amplitude and the duration.
[0056] Step S630, the matching abnormality early warning is displayed and prompted for power consumption early warning, and the prompt feedback is received. Specifically, the user is visually reported of the power consumption abnormality, a power consumption warning icon is displayed in the corner of the screen, and a “reduce image quality” or “continue” option is provided. After the user selects by touch, the system records the feedback.
[0057] Step S640, the second video enhancement strategy is adjusted according to the prompt feedback. Specifically, the enhancement parameters are modified according to the user feedback. If the user selects “reduce image quality”, the resolution and frame rate parameters are proportionally adjusted downward, and the optimization strategy is regenerated. Otherwise, the original strategy is maintained.
[0058] In a possible implementation, the display optimization management is performed according to the second video enhancement strategy, and step S600 further includes step S650 of establishing a user evaluation data set mapped with the second video enhancement strategy. Specifically, the 1-5 star evaluation of the user on the image quality and smoothness is collected through a built-in scoring interface, and is stored as time series data.
[0059] Step S660, the user evaluation data set, the second video enhancement strategy and the enhancement scene are associated and stored as a feedback optimization data set, and the enhancement scene includes a video scene, an environment scene and a focus scene. Specifically, the triplets including the video enhancement strategy parameters, the user score and the scene label are associated and stored using the SQLite database to form a historical optimization record for iterative optimization of the model. Among them, the enhancement scene is a combination description of the video, the environment and the focus context.
[0060] Step S670, configuring the enhancement preference of the user according to the feedback optimization data set, and managing the subsequent video enhancement according to the enhancement preference. Specifically, the user preference is learned from the historical data by clustering algorithm such as K-Means analysis of the user score mode, that is, the user preferred image processing style features, such as a user preferring high saturation and low sharpening. The preference configuration is automatically loaded to initialize the enhancement parameters when the new video is played.
[0061] The embodiment of the application utilizes the visual processor to analyze the video content, evaluates the frame complexity and establishes a multi-dimensional data set, according to which the video frame area is segmented, the region segmentation result and the multi-dimensional data set are enhanced and processed, the first video enhancement strategy is established, the environment sensor and the gaze sensor are activated, the focus identifier and the environment data set are established, the first video enhancement strategy is simulated with the focus identifier and the environment data set, the feedback incentive is established, the strategy is corrected by using the feedback incentive, the second video enhancement strategy is established and the display is optimized accordingly, which solves the technical problem of insufficient precision of visual quality enhancement in the existing mobile phone screen video display optimization, and achieves the technical effect of improving the precision of visual quality enhancement and improving the user viewing experience.
[0062] In the foregoing, the mobile phone screen video display optimization method based on the visual processor according to the embodiment of the application is described in detail. Next, the mobile phone screen video display optimization system based on the visual processor according to the embodiment of the application will be described with reference to the drawings. Figure 1 The mobile phone screen video display optimization system based on the visual processor according to the embodiment of the application is used to solve the technical problem of insufficient precision of visual quality enhancement in the existing mobile phone screen video display optimization, and achieves the technical effect of improving the precision of visual quality enhancement and improving the user viewing experience. The mobile phone screen video display optimization system based on the visual processor includes a multi-dimensional data set establishment module 10, a video frame area segmentation module 20, an enhancement processing analysis module 30, a sensor activation module 40, an enhancement simulation module 50 and a display optimization management module 60. Figure 2
[0063] The mobile phone screen video display optimization system based on the visual processor according to the embodiment of the application is used to solve the technical problem of insufficient precision of visual quality enhancement in the existing mobile phone screen video display optimization, and achieves the technical effect of improving the precision of visual quality enhancement and improving the user viewing experience. The mobile phone screen video display optimization system based on the visual processor includes a multi-dimensional data set establishment module 10, a video frame area segmentation module 20, an enhancement processing analysis module 30, a sensor activation module 40, an enhancement simulation module 50 and a display optimization management module 60.
[0064] The multi-dimensional data set establishing module 10 is configured to perform real-time analysis on video content by using a visual processor, identify dynamic elements, static backgrounds and text information, and perform complexity evaluation on each frame of the video content, and establish a multi-dimensional data set according to the complexity evaluation result and the identification result; the video frame region segmentation module 20 is configured to perform video frame region segmentation by using the identification result in the multi-dimensional data set, and establish a region segmentation result; the enhancement processing analysis module 30 is configured to perform enhancement processing analysis on the region segmentation result and the multi-dimensional data set, and establish a first video enhancement strategy; the sensor activator module 40 is configured to activate an environmental sensor and a gaze sensor, and establish a focus identifier and an environmental data set; the enhancement simulation module 50 is configured to perform enhancement simulation on the first video enhancement strategy, the focus identifier and the environmental data set, and establish a feedback incentive; and the display optimization management module 60 is configured to perform strategy correction on the first video enhancement strategy by using the feedback incentive, establish a second video enhancement strategy, and perform display optimization management according to the second video enhancement strategy.
[0065] The enhancement processing analysis module 30 is configured to perform enhancement processing analysis on the region segmentation result and the multi-dimensional data set, and establish a first video enhancement strategy, and can further include: a feature extraction unit configured to identify the region segmentation result according to the identification result in the multi-dimensional data set, perform multi-scale convolutional neural network multi-level feature extraction on the region segmentation result under the identification constraint, and establish a feature extraction result; a power consumption constraint establishing unit configured to read a power consumption mode of the mobile phone, establish a power consumption constraint according to the power consumption mode and the complexity evaluation result; and a feature enhancement optimization unit configured to take the power consumption constraint as a whole power consumption target, perform feature enhancement optimization according to the feature extraction result under the identification constraint, and establish the first video enhancement strategy according to the feature enhancement optimization result.
[0066] The feature enhancement optimization unit is configured to take the power consumption constraint as a whole power consumption target, perform feature enhancement optimization according to the feature extraction result under the identification constraint, and can further include: a smooth feedback layer establishing subunit configured to establish a smooth feedback layer, the smooth feedback layer being an adaptive expansion frame interval with a current video frame as a center point; an abnormal distribution establishing subunit configured to perform inter-frame display transition smoothing analysis in the feature enhancement optimization process by using the smooth feedback layer, and establish an abnormal distribution; and a guide compensation management subunit configured to configure a transition abnormal feedback by using the abnormal distribution, and perform guide compensation management of the feature enhancement optimization according to the transition abnormal feedback.
[0067] The guiding compensation management of the feature enhancement optimization according to the transition abnormal feedback can further include: a balance verification window setting component configured to set a balance verification window according to the transition abnormal feedback; a global balance effect verification component configured to perform global balance effect verification of the feature enhancement optimization using the balance verification window after an iteration of the guiding compensation management; a power consumption breaking instruction generating component configured to generate a power consumption breaking instruction if the global balance effect fails to meet a balance threshold within a predetermined iteration period; and a power consumption target adjusting component configured to continue the feature enhancement optimization management after adjusting the overall power consumption target according to the power consumption breaking instruction.
[0068] The detailed description of the specific configuration of the enhancement simulation module 50 is explained as follows: as described above, the first video enhancement strategy and the attention identifier and the environment data set are simulated for enhancement, and feedback incentives are established. The enhancement simulation module 50 can further include: an attention stability analysis unit configured to perform attention stability analysis on the attention identifier to establish an attention stability coefficient; a position importance reconstruction unit configured to reconstruct the position importance of the video frame using the attention identifier and the attention stability coefficient; an attention mechanism analysis unit configured to perform attention mechanism analysis using the position importance and the content of the video frame to establish an attention importance; and an adaptive simulation unit configured to perform adaptive simulation of the first video enhancement strategy based on the attention importance and the environment data set to establish feedback incentives.
[0069] The detailed description of the specific configuration of the display optimization management module 60 is explained as follows: as described above, the display optimization management is performed according to the second video enhancement strategy. The display optimization management module 60 can further include: a real power consumption obtaining unit configured to obtain the real power consumption mapped by the second video enhancement strategy; a mode matching analysis unit configured to perform mode matching analysis according to the real power consumption and the power consumption mode to establish a matching abnormality early warning; a power consumption early warning display prompting unit configured to perform power consumption early warning display prompting on the matching abnormality early warning and receive a prompt feedback; and a second video enhancement strategy adjusting unit configured to adjust the second video enhancement strategy according to the prompt feedback.
[0070] The detailed description of the specific configuration of the display optimization management module 60 is explained as follows: as described above, the display optimization management is performed according to the second video enhancement strategy. The display optimization management module 60 can further include: a real power consumption obtaining unit configured to obtain the real power consumption mapped by the second video enhancement strategy; a mode matching analysis unit configured to perform mode matching analysis according to the real power consumption and the power consumption mode to establish a matching abnormality early warning; a power consumption early warning display prompting unit configured to perform power consumption early warning display prompting on the matching abnormality early warning and receive a prompt feedback; and a second video enhancement strategy adjusting unit configured to adjust the second video enhancement strategy according to the prompt feedback.
[0071] The display optimization management module 60 can further include a user evaluation dataset establishing unit configured to establish a user evaluation dataset mapped with the second video enhancement strategy; and a feedback optimization dataset storage unit configured to store the user evaluation dataset, the second video enhancement strategy and an enhancement scene, including a video scene, an environment scene and a focus scene, as a feedback optimization dataset.
[0072] The first video enhancement strategy and the focus identifier and the environment dataset are simulated for enhancement, and feedback incentives are established. The enhancement simulation module 50 can further include a focus subject determining unit configured to determine a first focus subject, a second focus subject and other subjects according to the focus identifier; and an enhancement simulation unit configured to simulate enhancement under environment dataset adaptation using the first focus subject, the second focus subject and the other subjects to establish feedback incentives.
[0073] The mobile phone screen video display optimization system based on a visual processor provided by the embodiments of the present application can execute the mobile phone screen video display optimization method based on a visual processor provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0074] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or server, and the various units and modules are only divided according to the functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for easy mutual differentiation, and do not limit the protection scope of the present application.
[0075] The above is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with the preferred embodiment, however, it is not intended to limit the present application, and any person skilled in the art can make some changes or modifications to the above disclosed technical content to make equivalent embodiments with equivalent changes, as long as they do not deviate from the technical solution of the present application. Any modification, equivalent change and modification of the above embodiments according to the technical essence of the present application are still within the scope of the technical solution of the present application.
Claims
1. A method for optimizing video display on a mobile phone screen based on a vision processor, characterized in that, The method comprises: real-time video content analysis is performed by using a visual processor to identify dynamic elements, static backgrounds and text information, and to perform complexity evaluation of each frame of the video content, and a multidimensional dataset is established according to the complexity evaluation results and the identification results; video frame region segmentation is performed by using the identification results in the multidimensional dataset to establish a region segmentation result; the region segmentation result and the multidimensional dataset are subjected to enhancement processing analysis to establish a first video enhancement strategy; environmental sensors and gaze sensors are activated to establish a focus identifier and an environmental dataset; the first video enhancement strategy and the focus identifier and the environmental dataset are subjected to enhancement simulation to establish a feedback incentive; the feedback incentive is used to perform strategy correction of the first video enhancement strategy to establish a second video enhancement strategy, and display optimization management is performed according to the second video enhancement strategy; the region segmentation result and the multidimensional dataset are subjected to enhancement processing analysis to establish a first video enhancement strategy, comprising: after the region segmentation result is identified according to the identification results in the multidimensional dataset, multi-scale convolutional neural network multi-level feature extraction under identification constraint is performed on the region segmentation result to establish a feature extraction result; a power consumption mode of a mobile phone is read, and a power consumption constraint is established according to the power consumption mode and the complexity evaluation results; the power consumption constraint is taken as a whole power consumption target, and feature enhancement optimization under identification constraint is performed according to the feature extraction result, and a first video enhancement strategy is established according to the feature enhancement optimization result; the power consumption constraint is taken as a whole power consumption target, and feature enhancement optimization under identification constraint is performed according to the feature extraction result, comprising: a smoothing feedback layer is established, and the smoothing feedback layer is an adaptive expansion frame interval with a current video frame as a center point; frame display transition smoothing analysis in the feature enhancement optimization process is performed by using the smoothing feedback layer to establish an abnormal distribution; transition abnormal feedback is configured by using the abnormal distribution, and guiding compensation management of the feature enhancement optimization is performed according to the transition abnormal feedback; the guiding compensation management of the feature enhancement optimization according to the transition abnormal feedback comprises: a balance verification window is set according to the transition abnormal feedback; after iteration of the guiding compensation management is performed, global balance effect verification of the feature enhancement optimization is performed by using the balance verification window; if the global balance effect fails to meet a balance threshold within a predetermined iteration period, a power consumption breaking instruction is generated; after whole power consumption target adjustment is performed according to the power consumption breaking instruction, feature enhancement optimization management is continued.
2. The visual processor based mobile phone screen video display optimization method of claim 1, wherein, the first video enhancement strategy and the focus identifier and the environmental dataset are subjected to enhancement simulation to establish a feedback incentive, comprising: focus stability analysis is performed on the focus identifier to establish a focus stability coefficient; position importance of a video frame is reconstructed by using the focus identifier and the focus stability coefficient; attention mechanism analysis is performed by using the position importance and the content of the video frame to establish focus importance; based on the focus importance and the environmental dataset, adaptive simulation of the first video enhancement strategy is performed to establish the feedback incentive.
3. The visual processor based mobile phone screen video display optimization method of claim 1, wherein, display optimization management is performed according to the second video enhancement strategy, comprising: real power consumption mapped by the second video enhancement strategy is obtained; According to the real power consumption and power consumption mode, a mode matching analysis is performed to establish a matching abnormality early warning; The matching abnormality early warning is displayed for power consumption early warning and prompt, and prompt feedback is received; The second video enhancement strategy is adjusted according to the prompt feedback.
4. The visual processor based mobile phone screen video display optimization method of claim 1, wherein, The region segmentation result and the multi-dimensional data set are subjected to enhancement processing analysis to establish a first video enhancement strategy, and the method further comprises: The video content is subjected to video scene classification to establish a scene classification identifier; According to the scene classification identifier, a scene enhancement template is configured, and the scene enhancement template is used as a basis for performing enhancement optimization to establish the first video enhancement strategy.
5. The visual processor based mobile phone screen video display optimization method of claim 1, wherein, According to the second video enhancement strategy, display optimization management is performed, including: A user evaluation data set is established in mapping with the second video enhancement strategy; The user evaluation data set, the second video enhancement strategy and an enhancement scene are associated and stored as a feedback optimization data set, and the enhancement scene includes a video scene, an environmental scene and a focus scene; According to the feedback optimization data set, an enhancement preference of the user is configured, and subsequent video enhancement management is performed according to the enhancement preference.
6. The visual processor based mobile phone screen video display optimization method of claim 1, wherein, The first video enhancement strategy and the focus identifier and the environmental data set are subjected to enhancement simulation to establish a feedback incentive, and the method further comprises: According to the focus identifier, a first focus subject, a second focus subject and other subjects are determined; The first focus subject, the second focus subject and the other subjects are used for enhancement simulation under environmental data set adaptation to establish a feedback incentive.
7. A visual processor based mobile phone screen video display optimization system characterized by, The system is used to implement the visual processor-based mobile phone screen video display optimization method according to any one of claims 1-6, and the system comprises: A multi-dimensional data set establishment module is configured to use a visual processor to perform real-time analysis on video content, identify dynamic elements, static backgrounds and text information, perform complexity evaluation on each frame of the video content, and establish a multi-dimensional data set according to the complexity evaluation result and the identification result; A video frame region segmentation module is configured to use the identification result in the multi-dimensional data set to perform video frame region segmentation and establish a region segmentation result; An enhancement processing analysis module is configured to perform enhancement processing analysis on the region segmentation result and the multi-dimensional data set to establish a first video enhancement strategy; A sensor activator is configured to activate an environmental sensor and a gaze sensor to establish a focus identifier and an environmental data set; An enhancement simulation module is configured to perform enhancement simulation on the first video enhancement strategy and the focus identifier and the environmental data set to establish a feedback incentive; A display optimization management module is configured to use the feedback incentive to perform strategy correction of the first video enhancement strategy, establish a second video enhancement strategy, and perform display optimization management according to the second video enhancement strategy.
Citation Information
Patent Citations
Video processing method, device, equipment and system
CN115941979A
High-definition camera video processing method and system
CN117689577A