Video conference image enhancement processing method and system based on environment perception
By generating environmental and enhancement strategy portraits and performing intelligent image enhancement processing, combined with skin color correction through deep learning, the problem of optimizing facial areas in complex lighting environments during video conferencing is solved, achieving high-quality video conferencing image enhancement.
Patent Information
- Application Number
- CN202510942655.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video conferencing image enhancement methods are difficult to adapt to complex and changing lighting environments and cannot specifically optimize the facial area, resulting in poor image quality, especially distortion of facial skin color, which affects the visual experience.
By generating a portrait of the environment and enhancement strategy, intelligently selecting the enhancement module and performing parameterized configuration, the original image is enhanced, and a refined skin color verification and correction process is introduced, using a deep learning fine-grained color contrast analysis network for skin color correction.
Provides a stable, high-quality video conferencing visual experience in various environments, ensuring natural and realistic facial skin tones, and improving the overall image appearance and clarity of the facial area.
Smart Images

Figure CN120765485A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image enhancement, and more specifically, to a method and system for video conferencing image enhancement processing based on environment perception. Background Art
[0002] With the prevalence of remote work and online collaboration, video conferencing has become an indispensable means of communication in modern work and life. However, the image quality of video conferencing is often significantly affected by the lighting conditions of the user's environment, such as dark or bright lighting, backlighting, or uneven lighting. These complex environmental factors can lead to poor quality image frames captured by the camera, blurring of facial areas, loss of details, and distorted skin tones, seriously affecting the visual experience and communication efficiency of participants. Traditional image enhancement methods typically adopt global adjustment strategies or optimize only for a single image issue. These methods are difficult to adapt to the complex and changing lighting environments of video conferencing, and cannot specifically optimize the face, a critical visual area. They may even cause unnatural deviations in facial skin tone while enhancing the overall image.
[0003] Therefore, an optimized video conferencing image enhancement processing solution is desired, which can intelligently perceive environmental changes and perform targeted and refined enhancement processing on video conferencing images, especially face areas, which is of great significance for improving the overall video conferencing experience. Summary of the Invention
[0004] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a video conferencing image enhancement processing method and system based on environmental perception, which first generates a dynamic "environment and enhancement strategy portrait" by jointly evaluating the overall illumination of the original video frame and the local illumination conditions of the face area, and uses this as the basis for subsequent intelligent decision-making. Subsequently, based on this portrait, the system intelligently selects appropriate modules from the enhancement module library and performs parameterized configuration to perform preliminary, environmentally adaptive enhancement processing on the original image, aiming to improve the overall look and feel of the image and highlight the clarity of the face area. In order to ensure the naturalness and authenticity of the enhancement effect, especially the accuracy of the facial skin color, the concept further introduces a refined skin color verification and correction link. By using a fine-grained color contrast analysis network based on deep learning, the enhanced facial skin color is accurately compared with the standard skin color to accurately determine whether there is any abnormality in the skin color. If a skin color abnormality is detected, color correction is automatically performed to ensure that the final output facial skin color is natural and realistic. This closed-loop processing mechanism of "perception-decision-execution-feedback optimization" can effectively overcome the shortcomings of existing technologies in adaptability to complex lighting, targeted optimization of facial areas, and skin color fidelity, thereby providing a stable and high-quality video conferencing visual experience in various environments.
[0005] According to one aspect of the present application, a method for video conferencing image enhancement processing based on environment perception is provided, which includes:
[0006] Obtaining original video conference image frames captured by a video conference camera;
[0007] Detecting a face region from the original video conference image frame to obtain a face ROI image;
[0008] Generate an environment and enhancement strategy portrait based on the original video conference image frame and the face ROI image;
[0009] Based on the environment and enhancement strategy portrait, performing image enhancement processing on the original video conference image frame to obtain a preliminarily enhanced video conference image frame;
[0010] Detecting a face region from the preliminarily enhanced video conference image frame to obtain an enhanced video conference face ROI image;
[0011] The enhanced video conference face ROI image is compared with a preset standard facial skin color image to determine whether there is any abnormality in the skin color, and the optimized video conference image frame is output to the video conference interface for display.
[0012] According to another aspect of the present application, a video conferencing image enhancement processing system based on environment perception is provided, comprising:
[0013] The original video frame acquisition module permanently obtains the original video conference image frames captured by the video conference camera;
[0014] A face region detection module, configured to detect a face region from the original video conference image frame to obtain a face ROI image;
[0015] An environment and enhancement strategy portrait generation module, configured to generate an environment and enhancement strategy portrait based on the original video conference image frame and the face ROI image;
[0016] A video conference image frame enhancement module is used to perform image enhancement processing on the original video conference image frame based on the environment and enhancement strategy portrait to obtain a preliminarily enhanced video conference image frame;
[0017] An enhanced image frame face detection module is configured to permanently detect a face region from the initially enhanced video conference image frame to obtain an enhanced video conference face ROI image;
[0018] The image display module is used to compare the enhanced video conference face ROI image with the preset standard facial skin color image to determine whether there is any abnormality in the skin color, and output the optimized video conference image frame to the video conference interface for display.
[0019] Compared with the existing technology, the present application provides a video conferencing image enhancement processing method and system based on environmental perception. It first generates a dynamic "environment and enhancement strategy portrait" by jointly evaluating the overall lighting conditions of the original video frame and the local lighting conditions of the face area, which serves as the basis for subsequent intelligent decision-making. Subsequently, based on this portrait, the system intelligently selects appropriate modules from the enhancement module library and performs parameterized configuration to perform preliminary, environmentally adaptive enhancement processing on the original image, aiming to improve the overall look and feel of the image and highlight the clarity of the face area. In order to ensure the naturalness and authenticity of the enhancement effect, especially the accuracy of the facial skin color, the concept further introduces a refined skin color verification and correction link. By using a fine-grained color contrast analysis network based on deep learning, the enhanced facial skin color is accurately compared with the standard skin color to accurately determine whether there is any abnormality in the skin color. If a skin color abnormality is detected, color correction is automatically performed to ensure that the final output facial skin color is natural and realistic. This closed-loop processing mechanism of "perception-decision-execution-feedback optimization" can effectively overcome the shortcomings of existing technologies in adaptability to complex lighting, targeted optimization of facial areas, and skin color fidelity, thereby providing a stable and high-quality video conferencing visual experience in various environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0021] Figure 1 Flowchart of a method for video conferencing image enhancement based on environment perception according to an embodiment of the present application;
[0022] Figure 2 Schematic diagram of data flow of a method for enhancing video conferencing images based on environment perception according to an embodiment of the present application;
[0023] Figure 3 A flowchart of a method for enhancing a video conference image based on environment perception according to an embodiment of the present application, wherein the enhanced video frame face ROI image is compared with a preset standard face skin color image to determine whether the skin color is abnormal;
[0024] Figure 4A flowchart of a method for video conferencing image enhancement based on environment perception according to an embodiment of the present application for processing a face ROI image color feature vector of the enhanced video frame and a face standard skin color feature vector through a fine-grained color contrast analysis network to obtain a face image color feature contrast response encoding vector;
[0025] Figure 5 This is a block diagram of a video conferencing image enhancement processing system based on environment perception according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0027] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0028] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.
[0029] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0030] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0031] Currently, some video conferencing software or image processing applications have integrated enhancement features such as automatic brightness and contrast adjustment, which can be considered preliminary attempts at environmental awareness. However, these existing solutions are often relatively crude. For example, they may only adjust based on the overall brightness of the image, lacking specific attention and optimization for the facial region. This leads to poor enhancement results under complex lighting conditions (such as strong background light or uneven facial lighting), and may even degrade facial details or distort skin tone. Furthermore, these methods typically use fixed enhancement models or limited parameter adjustments, making it difficult to flexibly select and configure the optimal enhancement strategy based on diverse environmental characteristics and specific lighting conditions in the facial region. They also lack a detailed description and utilization of the deep connection between the environment and the enhancement strategy, thus limiting the adaptability and ultimate effectiveness of image enhancement.
[0032] In order to effectively address the negative impact of complex and changing lighting environments on image quality during video conferencing and to specifically improve the visual presentation of facial areas, the technical concept of this application is to first accurately understand the specific environmental characteristics of the current video conferencing scene, and to form a dynamic "environment and enhancement strategy portrait" by jointly evaluating the overall lighting of the original image frame and the local lighting of the facial area. Based on this portrait, the system can intelligently make decisions and call the corresponding image enhancement module for parameterized configuration, applying preliminary, environmentally adaptive enhancement processing to the original image to improve the overall appearance of the image and the clarity of the facial area. Furthermore, considering the potential impact that image enhancement operations may have on the naturalness of facial skin color, special emphasis is placed on the refined verification and correction of the enhanced facial skin color. By introducing a fine-grained color contrast analysis mechanism based on deep learning, the enhanced facial skin color is accurately compared with the standard skin color to accurately determine whether there are any abnormalities, and targeted color correction is performed when abnormalities are found. This closed-loop processing method of "perception-decision-execution-feedback optimization" can effectively solve problems such as the poor adaptability of traditional methods to complex lighting, insufficient attention to facial areas, and skin color distortion after enhancement. It can thus output clear, natural, high-quality video conferencing images in various environments.
[0033] In the technical solution of the present application, a video conferencing image enhancement processing method based on environment perception is proposed. Figure 1 The flowchart of the video conferencing image enhancement processing method based on environment perception according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the video conferencing image enhancement processing method based on environment perception according to an embodiment of the present application. Figure 1 and Figure 2As shown, the video conferencing image enhancement processing method based on environment perception according to the embodiment of the present application includes the steps of: S100, obtaining the original video conferencing image frame captured by the video conferencing camera; S200, detecting the face area from the original video conferencing image frame to obtain the face ROI image; S300, generating the environment and enhancement strategy portrait based on the original video conferencing image frame and the face ROI image; S400, based on the environment and enhancement strategy portrait, performing image enhancement processing on the original video conferencing image frame to obtain a preliminary enhanced video conferencing image frame; S500, detecting the face area from the preliminary enhanced video conferencing image frame to obtain an enhanced video conferencing face ROI image; S600, comparing the enhanced video conferencing face ROI image with the preset standard face skin color image to determine whether there is any abnormality in the skin color, and outputting the optimized video conferencing image frame to the video conferencing interface for display.
[0034] Specifically, in step S100, raw video conferencing image frames captured by the video conferencing camera are obtained. It is worth noting that these raw video conferencing image frames are the direct target of all subsequent analysis, processing, and optimization, carrying the most direct, unmodified visual information of the current conference scene. Without this input, operations such as face detection, environmental assessment, image enhancement, and skin tone correction cannot be performed.
[0035] Specifically, in one specific example of this application, the first step is device initialization and selection. The system enumerates currently available video input devices (e.g., built-in cameras, USB external cameras) and determines the target capture device based on pre-set or user selection. It then establishes a communication connection with the selected camera through the operating system's driver and application programming interface (API, such as DirectShow and Media Foundation on Windows platforms, V4L2 on Linux platforms, or cross-platform libraries such as OpenCV). Capture parameters such as resolution, frame rate, and color space can be configured. The second step is data stream capture. Once the connection is established and configured, the system issues a command to the camera to begin transmitting the video data stream. The camera continuously converts the captured optical signal into digital image data, packages it into continuous image frames at the set frame rate, and transmits it to the computer system via a data interface (e.g., USB). The third step is frame data reception and decoding. The application receives raw data frame by frame from the video data stream through the aforementioned API. This data may be in an uncompressed format (e.g., YUV, RGB) or a lightly hardware-compressed format. Upon receipt, if the data is in a compressed format, it must be decoded and restored to a pixel matrix that can be directly processed. The fourth step is data temporary storage and transmission. Each frame of successfully received and decoded image data, namely the "original video conferencing image frame", will be temporarily stored in the buffer of the system memory and passed to the subsequent image processing module, such as the face detection module, according to the requirements of the processing flow, thereby starting the entire video conferencing image enhancement processing flow based on environmental perception.
[0036] Specifically, in step S200, the face area is detected from the original video conference image frame to obtain a face ROI image. It should be understood that the core of video conferencing is communication between people. The face is the main carrier for conveying expressions, emotions and identity information, and its clarity and naturalness directly affect the communication effect. If the face area is not specifically detected and optimized, the global image enhancement may ignore the particularity of the face, resulting in improper lighting in the face area, blurred details or distorted skin color, which in turn reduces the visual experience. Therefore, by accurately identifying the face area, subsequent analysis and enhancement resources can be more effectively focused on this key area to achieve refined processing, such as independently evaluating the lighting of the face area, ensuring natural skin color, enhancing facial details, etc., thereby significantly improving the effectiveness of video conferencing and user satisfaction.
[0037] Specifically, in one example of this application, the first stage is image preprocessing. To improve detection accuracy and robustness, the raw video conferencing image frames may be preprocessed, such as converting them to grayscale to reduce computational effort, scaling them to fit the input size of the detection model, or using methods like histogram equalization to initially improve image contrast for easier detection. The second stage is feature extraction and candidate region generation. In this stage, the algorithm scans the image, searching for regions that match facial features. Common methods include knowledge-based approaches (such as detecting the relative positions of components like eyes, nose, and mouth), feature-based approaches (such as Haar-like features, LBP (Local Binary Patterns) features, and HOG (Histogram of Oriented Gradients) features, which effectively describe facial texture and edge information), and the recently widely used deep learning-based methods (such as using convolutional neural networks (CNNs) to automatically learn and extract deep discriminative features of faces). These methods generate a large number of candidate regions within the image that potentially contain faces. The third stage is classification and localization: The candidate regions generated in the previous stage are classified to determine whether they are truly human faces. For example, when using Haar-like features, an AdaBoost cascade classifier is often used to quickly eliminate non-face regions through a combination of weak classifiers. Deep learning-based methods use a classification layer (such as Softmax) at the end of the network to output a confidence score for each candidate region being a face. Simultaneously, a regression layer accurately predicts the coordinates of the face's bounding box (such as the x, y coordinates of the upper left corner, as well as the width and height). The fourth stage is ROI image extraction: Once a face region is successfully detected and localized by a bounding box, the system uses these coordinates to crop the corresponding face region from the original video conference image frame, generating a face ROI image. This ROI image is then used for subsequent processing steps such as face region illumination assessment, skin color analysis, and possible local enhancement. Sometimes, to include the full facial information or to allow for more flexibility in subsequent processing, the extracted ROI region may be slightly larger than the precisely detected bounding box.
[0038] Specifically, in step S300, an environment and enhancement strategy portrait is generated based on the original video conference image frame and the face ROI image. It should be understood that due to the complexity and variability of the lighting in the video conference environment, and the core position of the face as the visual focus, traditional image enhancement methods often lack a detailed perception of the current specific environment, resulting in poor enhancement effects or unnatural side effects. Therefore, in the technical solution of the present application, an environment and enhancement strategy portrait is further generated based on the original video conference image frame and the face ROI image. If the environment is not accurately assessed first, and special attention is paid to the lighting conditions in the face area, the subsequent enhancement processing is like a blind man touching an elephant. It is difficult to be targeted and cannot effectively solve problems such as too dark light, too bright light, backlight or skin color distortion.
[0039] More specifically, in an embodiment of the present application, based on the original video conference image frame and the face ROI image, an environment and enhancement strategy portrait is generated, including: performing an overall light level assessment on the original video conference image frame to obtain an overall light level assessment result; performing a face region light level assessment on the face ROI image to obtain a face light level assessment result; performing a relative light comparison based on the original video conference image frame and the face ROI image to determine a relative light assessment result; and generating the environment and enhancement strategy portrait based on the overall light level assessment result, the face light level assessment result, and the relative light assessment result. Specifically, by separately assessing the overall light level of the original image frame and the local light level of the face ROI image, the overall brightness and darkness of the scene and the direct lighting conditions of the face region can be understood; and by comparing the relative light levels between the two, complex situations such as backlighting (overall light, face dark), facial overexposure (overall dark, face bright), or light balance can be further revealed. Combining the results of these three assessments forms a structured environmental portrait that can be used for subsequent decision-making. It not only describes the current environmental state, but also implies the direction of enhancement strategies that should be adopted in response to this state, thus providing key basis for the "perception-decision-making" link proposed in the technical concept.
[0040] Specifically, in step S400, based on the environment and enhancement strategy portrait, the original video conference image frame is enhanced to obtain a preliminarily enhanced video conference image frame. It should be understood that the previous step has generated a precise assessment of the lighting characteristics and potential image issues of the current video conference scene—the "environment and enhancement strategy portrait." Without this portrait as a guide, directly performing image enhancement processing may repeat the mistakes of traditional methods, failing to specifically address issues caused by complex lighting environments (such as excessive darkness, excessive brightness, and backlighting), making it difficult to prioritize the visual quality of the facial area, and even potentially introducing new distortion due to improper enhancement. Therefore, based on the environment and enhancement strategy portrait, the original video conference image frame is further enhanced to obtain a preliminarily enhanced video conference image frame. This allows for preliminary, targeted enhancement of the original video conference image frame to improve its overall visual quality, with particular attention to the rendering of the facial area.
[0041] More specifically, in an embodiment of the present application, based on the environment and enhancement strategy profile, image enhancement processing is performed on the original video conference image frame to obtain a preliminarily enhanced video conference image frame. The process includes: selecting an application module from an enhancement module library based on the environment and enhancement strategy profile, configuring parameters of the application module to obtain a configured application module; and performing image enhancement processing on the original video conference image frame using the configured application module to obtain the preliminarily enhanced video conference image frame. Specifically, the process utilizes the overall lighting, facial area lighting, and relative lighting assessment results contained in the environment and enhancement strategy profile to intelligently select one or more application modules that best suit the current scenario from a pre-set enhancement module library (e.g., a collection of modules including brightness adjustment, contrast enhancement, denoising, sharpening, HDR processing, and other functions). Key parameters of these modules are then refined based on cues provided by the image. For example, if the image indicates overall underexposure and underexposure in the facial area, a brightness boost module may be selected with a higher gain setting. If backlighting is indicated (bright background, dark face), a dynamic range compression or local illumination compensation module may be selected. The ultimate goal is to generate a frame of "preliminary enhanced video conferencing image frame", which has significant improvements in clarity, contrast, brightness, etc. compared to the original frame, laying a good foundation for possible subsequent skin color correction.
[0042] Specifically, in step S500, the facial region is detected from the preliminarily enhanced video conference image frame to obtain an enhanced video conference face ROI image. It should be understood that although the original image frame has undergone a preliminarily enhanced process based on the environment and enhancement strategy, improving the overall visual effect and visibility of the facial region, the next critical step is the refined verification and correction of facial skin color. To accurately assess and adjust skin color, the facial region must first be relocated within the preliminarily enhanced image. The preliminarily enhanced process may slightly alter image features and, in extreme cases, may even cause a slight shift in facial position or a change in detection difficulty. Therefore, the facial region is further detected from the preliminarily enhanced video conference image frame to obtain an enhanced video conference face ROI image. Re-performing face detection on the enhanced image ensures the accuracy of subsequent skin color analysis, directly applying it to facial pixels in the current visual state. It is worth noting that the enhanced video conference face ROI image will serve as direct input for subsequent skin color anomaly detection and color correction. Compared to the face ROI in the original frame, this enhanced ROI typically has better brightness and contrast, closer to the final state presented to the user. Therefore, skin color analysis and adjustment based on this ROI will better reflect and influence the final visual perception. Ensuring that the face is repositioned on the enhanced image can avoid bias in the analysis area caused by any geometric or appearance changes that may be introduced during the image processing pipeline.
[0043] Specifically, in step S600, the enhanced video conference face ROI image is compared with the preset standard facial skin color image to determine whether there is any abnormality in the skin color, and the optimized video conference image frame is output to the video conference interface for display. It should be understood that even after the preliminary image enhancement based on environmental perception, the naturalness and accuracy of facial skin color are still the core evaluation criteria for video conferencing image quality, and are easily affected by complex lighting and side effects of enhancement algorithms, thereby deviating from the normal range and causing distortion such as reddish, greenish, too white or too dark, which seriously affects the user experience. Traditional methods are often difficult to ensure the authenticity and naturalness of skin color. Therefore, after the preliminary enhancement of the image, it is necessary to further detect and ensure the quality of facial skin color, and to perform refined "feedback optimization" on the enhancement effect.
[0044] Figure 3 The flowchart of the method for enhancing the video conference image based on environment perception according to the embodiment of the present application is to compare the face ROI image of the enhanced video frame with the preset standard skin color image of the face to determine whether there is any abnormality in the skin color. Figure 3As shown, according to the environment perception-based video conferencing image enhancement processing method of an embodiment of the present application, step S600 includes: S610, extracting color features from the enhanced video frame face ROI image and the face standard skin color image respectively to obtain the enhanced video frame face ROI image color histogram and the face standard skin color image color histogram; S620, performing dual-branch color implicit feature mining on the enhanced video frame face ROI image color histogram and the face standard skin color image color histogram to obtain the enhanced video frame face ROI image color feature vector and the face standard skin color image color feature vector; S630, passing the enhanced video frame face ROI image color feature vector and the face standard skin color image color feature vector through a fine-grained color contrast analysis network to obtain a face image color feature contrast response coding vector; S640, determining whether there is any skin color abnormality based on the face image color feature contrast response coding vector.
[0045] Specifically, in step S610, color features are extracted from the enhanced video frame face ROI image and the standard skin color image to obtain a color histogram of the enhanced video frame face ROI image and a color histogram of the standard skin color image. It should be understood that directly comparing the raw pixel values between the enhanced video frame face ROI image and the standard skin color image is too detailed for skin color analysis and is easily affected by noise, making it difficult to capture macroscopic color trends and anomalies. Therefore, in the technical solution of the present application, color features are further extracted from the enhanced video frame face ROI image and the standard skin color image to obtain a color histogram of the enhanced video frame face ROI image and a color histogram of the standard skin color image. By extracting the color histograms of these two images, the color information of the enhanced face ROI and standard skin color images can be converted from a complex pixel matrix into a concise statistical representation, namely, the distribution of the number of pixels in each color interval. This representation method not only summarizes the overall hue and color composition of the image but also lays the foundation for subsequent feature learning and comparative analysis using complex neural network models. The ultimate goal is to obtain a feature representation that can effectively reflect the essential characteristics of skin color and facilitate difference measurement.
[0046] Specifically, in step S620, dual-branch color implicit feature mining is performed on the enhanced video frame face ROI image color histogram and the standard facial skin color image color histogram to obtain the enhanced video frame face ROI image color feature vector and the standard facial skin color image color feature vector. It should be understood that although the enhanced video frame face ROI image color histogram and the standard facial skin color image color histogram provide a quantitative representation of color distribution, they are essentially relatively rudimentary and shallow features and may not be able to capture subtle, nonlinear changes in skin color and the complex dependencies between different color channels. These subtleties are crucial for accurately determining whether skin color is abnormal. In addition, directly comparing color histograms may be overly sensitive or not robust to factors such as lighting changes and individual skin color differences. That is to say, due to the complexity of skin color distortion, in order to perform more refined processing and comparison to determine whether the skin color is normal or not, in the technical solution of the present application, the color histogram of the enhanced video frame face ROI image and the color histogram of the standard face skin color image are further subjected to dual-branch color implicit feature mining to obtain the color feature vector of the enhanced video frame face ROI image and the color feature vector of the standard face skin color image.
[0047] More specifically, in an embodiment of the present application, dual-branch color implicit feature mining is performed on the color histogram of the enhanced video frame face ROI image and the color histogram of the standard skin color image to obtain a color feature vector of the enhanced video frame face ROI image and a color feature vector of the standard skin color image, including: passing the color histogram of the enhanced video frame face ROI image and the color histogram of the standard skin color image through a dual-branch feature extractor based on a dilated convolutional neural network model to obtain the color feature vector of the enhanced video frame face ROI image and the color feature vector of the standard skin color image. In particular, by processing the dual-branch feature extractor based on the dilated convolutional neural network model, more advanced and robust color feature vectors can be extracted independently but in parallel from the two input color histograms (one from the enhanced face ROI image and the standard skin color image, respectively). The application of dilated convolutions expands the receptive field, enabling the network to capture a wider range of contextual information and correlations between different color intervals in the color histogram without increasing the number of parameters or computational effort. This allows the network to learn deeper implicit patterns that better reflect the essential characteristics of skin color. The dual-branch structure ensures that the same feature extraction logic is used for both enhanced and standard facial skin tones, ensuring fairness and consistency in subsequent comparisons. The ultimate goal is to convert the original, potentially high-dimensional and information-redundant color histogram into a feature vector with more compact dimensions, more refined information, and better suited for subsequent fine-grained comparative analysis.
[0048] Specifically, in step S630, the enhanced video frame face ROI image color feature vector and the standard facial skin color feature vector are passed through a fine-grained color contrast analysis network to obtain a facial image color feature contrast response encoding vector. It should be understood that while the enhanced video frame face ROI image color feature vector and the standard facial skin color feature vector extracted in the previous step are more advanced than color histograms, a more in-depth and detailed comparison mechanism is required to accurately determine whether there is "abnormality" or "distortion" between the two skin color features. Simple vector distance or similarity calculations may not capture the complex, nonlinear differences between skin colors, especially those subtle deviations that have a significant impact on human visual perception but may not be numerically significant. Therefore, in the technical solution of the present application, the enhanced video frame face ROI image color feature vector and the standard facial skin color feature vector are further passed through a fine-grained color contrast analysis network to obtain a facial image color feature contrast response encoding vector. Through processing by a fine-grained color contrast analysis network, a deep, structured comparative analysis is performed between the enhanced facial ROI image color feature vector and the standard skin color feature vector, ultimately generating a single, highly condensed "face image color feature contrast response encoding vector." This response encoding vector does not simply indicate whether the two are identical; rather, it encapsulates complex interactive response information between them, such as their local similarities, differences, alignment relationships, and conditional dependencies across different feature intensity ranges. The network focuses on the intrinsic strength of features through ordered arrangement, performs local comparison through equal-granularity segmentation, models deep interactions between local features through transfer response inference units, and finally integrates global information through sequence transfer and aggregation. The ultimate goal is to obtain an encoding that comprehensively and compactly represents the overall, deep-level interactive response characteristics between the two original color feature vectors. This encoding will be directly used to subsequently determine whether skin color anomalies exist and guide possible color correction.
[0049] Figure 4 This is a flowchart of the method for video conferencing image enhancement based on environment perception according to an embodiment of the present application, which processes the enhanced video frame face ROI image color feature vector and the face standard skin color image color feature vector through a fine-grained color contrast analysis network to obtain a face image color feature contrast response encoding vector. Figure 4As shown, according to the environment perception-based video conferencing image enhancement processing method of an embodiment of the present application, step S630 includes: S631, performing ordered arrangement of the enhanced video frame face ROI image color feature vector and the face standard skin color image color feature vector based on the eigenvalue size to obtain the enhanced video frame face ROI image color feature ordered arrangement coding vector and the face standard skin color image color feature ordered arrangement coding vector; S632, performing equal-granularity feature segmentation on the enhanced video frame face ROI image color feature ordered arrangement coding vector and the face standard skin color image color feature ordered arrangement coding vector to obtain a sequence of enhanced video frame face ROI image color local feature ordered coding vectors and a sequence of face standard skin color image color local feature ordered coding vectors; S633, performing feature response inference information transfer on the sequence of enhanced video frame face ROI image color local feature ordered coding vectors and the sequence of face standard skin color image color local feature ordered coding vectors to obtain the face image color feature contrast response coding vector.
[0050] More specifically, in step S631, the enhanced video frame face ROI image color feature vector and the face standard skin color feature vector are arranged in an ordered manner based on the eigenvalue size to obtain an enhanced video frame face ROI image color feature ordered arrangement coding vector and a face standard skin color feature ordered arrangement coding vector, which are expressed as follows:
[0051]
[0052]
[0053] in, and Respectively represent the enhanced video frame face ROI image color feature vector and the face standard skin color image color feature vector, Indicates sorting of vector elements. and They respectively represent the ordered arrangement coding vector of the color features of the face ROI image of the enhanced video frame and the ordered arrangement coding vector of the color features of the face standard skin color image.
[0054] It should be understood that the original order of the elements in the enhanced video frame face ROI image color feature vector and the standard facial skin color image color feature vector may not contain semantic information, or their order may be arbitrary with respect to the characteristics of skin color itself. If this fixed-order vector is directly input into the subsequent deep comparison network, the network may learn patterns related to the original order that are not essential skin color characteristics, thus affecting the accuracy and robustness of the comparison. Therefore, it is necessary to perform an ordered permutation of the enhanced video frame face ROI image color feature vector and the standard facial skin color image color feature vector based on the magnitude of their eigenvalues. Specifically, the ordered permutation reorders the elements within the vectors according to the magnitude of each eigenvalue, generating two new ordered permutation encoding vectors. The core idea is that this ordering ensures that the feature components with larger values (i.e., higher intensities) in the enhanced video frame face ROI image color feature vector and the standard facial skin color image color feature vector are arranged together, while the smaller components are arranged at the other end. This approach aims to eliminate the potential arbitrariness introduced by the original feature order, endowing the network with a degree of invariance or equivariance to input permutations. The ultimate goal is to enable subsequent network analysis to more consistently focus on and compare characteristics across different feature value ranges, making it easier to identify which color features exhibit significant differences or similarities between two skin color samples, thus laying the foundation for subsequent equal-granularity segmentation and local interaction analysis.
[0055] More specifically, step S632 performs equal-granularity feature segmentation on the ordered arrangement coding vector of the color features of the enhanced video frame face ROI image and the ordered arrangement coding vector of the color features of the standard facial skin color image to obtain a sequence of ordered coding vectors of the local color features of the enhanced video frame face ROI image and a sequence of ordered coding vectors of the local color features of the standard facial skin color image, which can be expressed as follows:
[0056]
[0057] ;
[0058] in, represents the feature segmentation function, They represent the first, second, and third ordered encoding vectors of the color local features of the face ROI image of the enhanced video frame. and Enhanced video frame face ROI image color local feature ordered coding vector, They represent the first, second, and third in the sequence of ordered coding vectors of the local color features of the standard skin color image of the face. and Ordered encoding vector of local color features of a standard skin color image of a human face.
[0059] It should be understood that while a global comparison between the ordered permutation encoding vectors of the color features of the enhanced video frame face ROI image and the ordered permutation encoding vectors of the color features of the standard skin color image can provide a sense of overall similarity, it often fails to capture subtle differences and complex interaction patterns within local regions. These local details are crucial for accurately determining whether skin color is natural and whether specific types of distortion (for example, deviations within a specific color intensity range) exist. In other words, skin color distortion is often not globally consistent and may only exhibit abnormalities in certain color dimensions or intensity ranges. Therefore, a method is needed to decompose the overall features into local segments for refined analysis to capture subtle differences in skin color. Based on this, the technical solution of the present application further performs equal-granularity feature segmentation on the ordered permutation encoding vectors of the color features of the enhanced video frame face ROI image and the ordered permutation encoding vectors of the color features of the standard skin color image. "Equal granularity" means that each segmented local feature segment has the same length (dimension), which ensures scale consistency when subsequently comparing and modeling the local color features of each set of corresponding enhanced video frame face ROI image and standard skin color image. This decomposition allows subsequent analysis to focus on comparing the degree of match and interaction between the two skin tones within different local regions, such as "high-intensity feature regions," "medium-intensity feature regions," and "low-intensity feature regions." By deeply analyzing these local feature fragments one by one, the network can capture subtle differences that are difficult to detect in a global comparison. For example, the overall skin tone may appear acceptable, but there may be unnatural saturation or color cast in a specific color intensity region. This provides more targeted input for the subsequent transfer response inference unit and is a key step in achieving fine-grained skin tone comparison and accurate anomaly detection.
[0060] Accordingly, according to an embodiment of the present application, feature response inference information transfer is performed on the sequence of the enhanced video frame face ROI image color local feature ordered coding vectors and the sequence of the face standard skin color image color local feature ordered coding vectors to obtain the face image color feature contrast response coding vector, including: inputting each group of corresponding enhanced video frame face ROI image color local feature ordered coding vectors and the sequence of the face standard skin color image color local feature ordered coding vectors into a transfer response inference unit to obtain a sequence of face image color local transfer response coding matrices; and performing transfer response inference sequence transfer on the sequence of the face image color local transfer response coding matrices to obtain the face image color feature contrast response coding vector.
[0061] More specifically, each corresponding set of the sequence of the enhanced video frame face ROI image color local feature ordered coding vectors and the sequence of the face standard skin color image color local feature ordered coding vectors is input into the transfer response inference unit to obtain a sequence of face image color local transfer response coding matrices, which can be expressed as follows:
[0062]
[0063] in, for and The local color transfer response encoding matrix of the face image between is matrix multiplication, for activation function, is the trainable modulation weight matrix.
[0064] It's understandable that while previously segmenting the global ordered color feature vector into sequences of local features laid the foundation for refined comparison, these local segments alone are insufficient to reveal the complex and potentially nonlinear interactions between the color of the enhanced video frame's facial ROI image and the color of the standard facial skin color image. Subtle anomalies in skin color often manifest within specific color intensity ranges, and a powerful modeling tool is needed to capture how these two skin color features influence, respond to, or deviate from each other. Through the processing of the transfer response inference unit, it is possible to quantify and encode multiple possible interaction patterns between each corresponding set of ordered color feature encoding vectors for the enhanced video frame's facial ROI image and the standard facial skin color image, such as their local similarity, difference, alignment, conditional dependency, and even co-activation or inhibition within specific feature value ranges. The output "sequence of local color transfer response encoding matrices for facial images" is designed to encapsulate rich and detailed interaction information within the corresponding local region in a structured manner. For example, it captures element-by-element correlation strength rather than a simple scalar score. This preserves more detailed color interactions between the enhanced video frame's face ROI image and the standard skin color image, providing high-quality local insights for subsequent global information integration. In other words, the representation of the local color transfer response encoding matrix for facial images retains richer and more structured information than directly comparing local vectors or outputting a single similarity score. This means the system now possesses a "deep portrait" of how two skin color features interact at different levels of granularity and intensity. The detailed encoding of these local interactions enables subsequent sequence transfer and aggregation modules to more comprehensively understand and determine whether there are any anomalies in the overall skin color based on these highly informative intermediate results. They are particularly adept at detecting subtle skin color distortions that are confined to specific local regions, thereby improving the accuracy and robustness of skin color anomaly detection.
[0065] More specifically, the sequence of the facial image color local transfer response coding matrix is transferred by transfer response inference sequence to obtain the facial image color feature contrast response coding vector, which is expressed as follows:
[0066]
[0067]
[0068] ;
[0069] in, Indicates flattening the matrix into vector processing, is the face image color local transfer response encoding vector, for of norm, is the value of the natural exponential function with the natural constant e as the base, To transfer the response inference weighted operation, for encoder, The color feature contrast response encoding vector of the face image.
[0070] It should be understood that although the previous steps have generated detailed local color transfer response encoding matrices for enhanced and standard facial skin tones across various local feature intensity ranges, these insights remain fragmented and local. Accurately determining whether overall skin tone is abnormal requires integrating this local information and understanding how it evolves and influences each other as feature intensity changes. In other words, the naturalness of skin tone is a holistic perception; the match or mismatch of a single local region cannot fully determine the final conclusion. The dependencies between local features and overall trends are crucial. Based on this, the sequence of local color transfer response encoding matrices for facial images is further processed using transfer response inference sequence transfer. By applying sequence modeling techniques (such as recurrent neural networks (RNNs), long short-term memory (LSTMs), or transformers), the generated sequence of local color transfer response encoding matrices for facial images is processed to integrate local interaction information distributed across different eigenvalue ranges, critically considering the sequential relationships between them (derived from the ordered arrangement of eigenvalues). This process aims to learn how these local color transfer response encoding matrices for facial images evolve as feature intensity changes, capturing dependencies across feature intensity ranges and global contextual information. The ultimate goal is to form a single, concise and information-rich "face image color feature contrast response encoding vector". This vector can comprehensively and compactly represent the overall and deep interactive response characteristics between the original two input color feature vectors, and can more comprehensively and meticulously reflect the true degree and nature of the difference between the enhanced facial skin color and the standard skin color, providing highly optimized input for downstream skin color abnormality judgment, which enables subsequent skin color abnormality discrimination models to make more accurate and robust judgments based on this vector.
[0071] Preferably, here, the segmentation granularity of the feature segmentation process not only adjusts the local area division range of the local color transfer response coding matrix of the facial image, but also directly affects the presentation of the linkage relationship between the corresponding enhanced video frame facial ROI image color local feature ordered coding vector and the facial standard skin color image color local feature ordered coding vector.
[0072] Since the local transfer response encoding matrix of the face image color is the transfer response between two local areas, it actually describes the spatial measurement of the interactive response domain based on its row vector, and the row vector dimension is also the representation of the local area size mentioned above, if the local area size, that is, the row vector length As a local size strength constraint condition, the low-dimensional structure metric of the local transfer response encoding matrix of the face image color is presented, namely the Frobenius norm Should follow a pattern similar to a Poisson distribution:
[0073] ;
[0074] That is, the row vector length The low-dimensional space metric count imposed on the local region interaction response domain as a local size strength constraint is Second-rate, express The factorial of . From this, we can solve the parameter .
[0075] Thus, the transfer-response interaction is introduced by and the mean expectation is In the case of a Poisson-like process, and The edge connection representation describing the spatial measure further determines the possibility of interactive response linkage between the two local regions as follows:
[0076]
[0077] in, represents the two-norm of the vector, is an absolute value.
[0078] Then, based on the interactive response linkage possibility To adjust the size of the local area Iteratively adjust to obtain :
[0079] ;
[0080] That is, under the strict guarantee that the expected degree is In this case, the normalization processing of the overall spatial metric is determined by the normalization limit formed by the fluctuation of the average expected connection possibility of each row, aiming to enable the structured linkage information within the local spatial metric to avoid overfitting of the local area, thereby improving the overall representation efficiency of the sequence of the local transfer response encoding matrix of the facial image color.
[0081] Specifically, in step S640, based on the facial image color feature contrast response encoding vector, a determination is made as to whether the skin color is abnormal. It should be understood that the previous, complex, fine-grained color contrast analysis has successfully encoded the deep interactions and subtle differences between the enhanced facial skin color and the standard skin color into a highly condensed "facial image color feature contrast response encoding vector." However, this vector itself remains a high-dimensional, complex feature representation; it describes "what the difference looks like" but does not directly determine "whether this difference constitutes a visual abnormality." Therefore, in the technical solution of the present application, further determination as to whether the skin color is abnormal is required. More specifically, in an embodiment of the present application, determining whether the skin color is abnormal based on the facial image color feature contrast response encoding vector includes: passing the facial image color feature contrast response encoding vector through a classifier-based skin color abnormality discriminator to obtain a discrimination result, which indicates whether the skin color is abnormal. In other words, the facial image color feature contrast response encoding vector is converted into a clear, binary "discrimination result" by a specially trained "skin color abnormality discriminator" (typically a classifier such as a support vector machine, logistic regression, or a small neural network). This judgment directly indicates whether the enhanced facial skin color, when compared to the standard skin color, is deemed visually unacceptable by the system. The ultimate goal is to provide a clear basis for subsequent image processing decisions: if the skin color is abnormal, the color correction module needs to be activated; if the skin color is normal, the initially enhanced image can be directly output or subjected to other non-color-related processing.
[0082] In particular, in an embodiment of the present application, the enhanced video conference face ROI image is compared with a preset standard facial skin color image to determine whether there is any abnormality in the skin color, and the optimized video conference image frame is output to the video conference interface for display, including: when there is an abnormality in the skin color, the color correction is performed on the preliminary enhanced video conference image frame to obtain the optimized video conference image frame, and the optimized video conference image frame is output to the video conference interface for display; when the skin color is normal, the preliminary enhanced video frame is output to the video conference interface for display. In other words, if there is an abnormality in the skin color, the preliminary enhanced video conference image frame needs to be subjected to targeted color correction to generate and display the "optimized video conference image frame"; if the skin color is normal, it indicates that the preliminary enhancement effect has met the standards in terms of skin color, and the "preliminary enhanced video frame" can be directly output. Its fundamental goal is to ensure that no matter what the initial environment is or how the preliminary enhancement is performed, the facial skin color ultimately presented to the participants is natural, healthy, and comfortable.
[0083] In summary, according to the embodiment of the present application, the method for enhancing the video conference image based on environmental perception is explained. It first generates a dynamic "environment and enhancement strategy portrait" by jointly evaluating the overall illumination of the original video frame and the local illumination conditions of the face area, which serves as the basis for subsequent intelligent decision-making. Subsequently, based on this portrait, the system intelligently selects the appropriate module from the enhancement module library and performs parameterized configuration to perform preliminary, environmentally adaptive enhancement processing on the original image, aiming to improve the overall look and feel of the image and highlight the clarity of the face area. In order to ensure the naturalness and authenticity of the enhancement effect, especially the accuracy of the facial skin color, the concept further introduces a refined skin color verification and correction link. By using a fine-grained color contrast analysis network based on deep learning, the enhanced facial skin color is accurately compared with the standard skin color to accurately determine whether there is any abnormality in the skin color. If a skin color abnormality is detected, color correction is automatically performed to ensure that the final output facial skin color is natural and realistic. This closed-loop processing mechanism of "perception-decision-execution-feedback optimization" can effectively overcome the shortcomings of existing technologies in adaptability to complex lighting, targeted optimization of facial areas, and skin color fidelity, thereby providing a stable and high-quality video conferencing visual experience in various environments.
[0084] Furthermore, a video conferencing image enhancement processing system based on environment perception is also provided.
[0085] Figure 5 FIG is a block diagram of a video conferencing image enhancement processing system based on environment perception according to an embodiment of the present application. Figure 5 As shown, the video conferencing image enhancement processing system 500 based on environment perception according to the embodiment of the present application includes: an original video frame acquisition module 510, which is used to obtain the original video conferencing image frame captured by the video conferencing camera; a face area detection module 520, which is used to detect the face area from the original video conferencing image frame to obtain a face ROI image; an environment and enhancement strategy portrait generation module 530, which is used to generate an environment and enhancement strategy portrait based on the original video conferencing image frame and the face ROI image; a video conferencing image frame enhancement module 540, which is used to perform image enhancement processing on the original video conferencing image frame based on the environment and enhancement strategy portrait to obtain a preliminary enhanced video conferencing image frame; an enhanced image frame face detection module 550, which is used to detect the face area from the preliminary enhanced video conferencing image frame to obtain an enhanced video conferencing face ROI image; an image display module 560, which is used to compare the enhanced video conferencing face ROI image with a preset face standard skin color image to determine whether there is any abnormality in the skin color, and output the optimized video conferencing image frame to the video conferencing interface for display.
[0086] As mentioned above, the environment-aware video conference image enhancement processing system 500 according to embodiments of the present application can be implemented in various wireless terminals, such as a server with environment-aware video conference image enhancement processing algorithms, etc. In one possible implementation, the environment-aware video conference image enhancement processing system 500 according to embodiments of the present application can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the environment-aware video conference image enhancement processing system 500 can be a software module in the operating system of the wireless terminal, or can be an application developed for the wireless terminal; of course, the environment-aware video conference image enhancement processing system 500 can also be one of the many hardware modules of the wireless terminal.
[0087] Alternatively, in another example, the environment-aware video conference image enhancement processing system 500 can also be a separate device from the wireless terminal, and the environment-aware video conference image enhancement processing system 500 can be connected to the wireless terminal through a wired and / or wireless network, and transmit interactive information in an agreed data format.
[0088] Embodiments of the present disclosure have been described above, the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles, practical applications, or improvements to the technology in the market of the embodiments, or to enable other ordinary skilled in the art to understand the embodiments disclosed herein.
Claims
1. A video conferencing image enhancement processing method based on environment perception, characterized in that: include: Obtaining original video conference image frames captured by a video conference camera; Detecting a face region from the original video conference image frame to obtain a face ROI image; Generate an environment and enhancement strategy portrait based on the original video conference image frame and the face ROI image; Based on the environment and enhancement strategy portrait, performing image enhancement processing on the original video conference image frame to obtain a preliminarily enhanced video conference image frame; Detecting a face region from the preliminarily enhanced video conference image frame to obtain an enhanced video conference face ROI image; The enhanced video conference face ROI image is compared with a preset standard facial skin color image to determine whether there is any abnormality in the skin color, and the optimized video conference image frame is output to the video conference interface for display.
2. The method for video conferencing image enhancement processing based on environment perception according to claim 1, characterized in that: Generate an environment and enhancement strategy portrait based on the original video conference image frame and the face ROI image, including: Performing an overall light level assessment on the original video conference image frame to obtain an overall light level assessment result; Performing a facial region illumination level assessment on the facial ROI image to obtain a facial illumination level assessment result; Performing relative illumination comparison based on the original video conference image frame and the face ROI image to determine a relative illumination evaluation result; The environment and enhancement strategy portrait is generated based on the overall light level assessment result, the face light level assessment result and the relative light assessment result.
3. The method for video conferencing image enhancement processing based on environment perception according to claim 2, characterized in that: Based on the environment and enhancement strategy portrait, performing image enhancement processing on the original video conference image frame to obtain a preliminarily enhanced video conference image frame, including: Based on the environment and enhancement policy profile, select an application module from an enhancement module library, and configure parameters of the application module to obtain a configured application module; The original video conference image frame is subjected to image enhancement processing by the configured application module to obtain the preliminarily enhanced video conference image frame.
4. The method for video conferencing image enhancement processing based on environment perception according to claim 3 is characterized in that: Comparing the enhanced video frame face ROI image with a preset face standard skin color image to determine whether the skin color is abnormal, including: Extracting color features from the enhanced video frame face ROI image and the face standard skin color image respectively to obtain an enhanced video frame face ROI image color histogram and a face standard skin color image color histogram; Performing dual-branch color implicit feature mining on the enhanced video frame face ROI image color histogram and the face standard skin color image color histogram to obtain an enhanced video frame face ROI image color feature vector and a face standard skin color image color feature vector; The enhanced video frame face ROI image color feature vector and the face standard skin color feature vector are passed through a fine-grained color contrast analysis network to obtain a face image color feature contrast response encoding vector; Based on the color feature comparison of the facial image and the response coding vector, it is determined whether there is an abnormality in the skin color.
5. The method for video conferencing image enhancement processing based on environment perception according to claim 4 is characterized in that: The method comprises the steps of: performing dual-branch color implicit feature mining on the color histogram of the enhanced video frame face ROI image and the color histogram of the standard face skin color image to obtain a color feature vector of the enhanced video frame face ROI image and a color feature vector of the standard face skin color image, comprising: passing the color histogram of the enhanced video frame face ROI image and the color histogram of the standard face skin color image through a dual-branch feature extractor based on a void convolutional neural network model to obtain the color feature vector of the enhanced video frame face ROI image and the color feature vector of the standard face skin color image.
6. The method for video conferencing image enhancement processing based on environment perception according to claim 5, characterized in that: The enhanced video frame face ROI image color feature vector and the face standard skin color feature vector are subjected to a fine-grained color contrast analysis network to obtain a face image color feature contrast response encoding vector, including: The enhanced video frame face ROI image color feature vector and the face standard skin color image color feature vector are arranged in an ordered manner based on the eigenvalue size to obtain an enhanced video frame face ROI image color feature ordered arrangement coding vector and a face standard skin color feature ordered arrangement coding vector; Performing equal-granularity feature segmentation on the ordered arrangement coding vector of color features of the enhanced video frame face ROI image and the ordered arrangement coding vector of color features of the standard face skin color image to obtain a sequence of ordered coding vectors of local color features of the enhanced video frame face ROI image and a sequence of ordered coding vectors of local color features of the standard face skin color image; Feature response inference information transfer is performed on the sequence of ordered color feature coding vectors of the enhanced video frame face ROI image and the sequence of ordered color feature coding vectors of the standard skin color image to obtain the face image color feature contrast response coding vector.
7. The method for video conferencing image enhancement processing based on environment perception according to claim 6, characterized in that: Performing feature response inference information transfer on the sequence of ordered color feature coding vectors of the enhanced video frame human face ROI image and the sequence of ordered color feature coding vectors of the human face standard skin color image to obtain the human face image color feature contrast response coding vector, including: Inputting each corresponding set of the sequence of the enhanced video frame face ROI image color local feature ordered coding vectors and the sequence of the face standard skin color image color local feature ordered coding vectors into a transfer response inference unit to obtain a sequence of face image color local transfer response coding matrices; The sequence of the facial image color local transfer response coding matrices is transferred using a transfer response inference sequence to obtain the facial image color feature contrast response coding vector.
8. The method for video conferencing image enhancement processing based on environment perception according to claim 1, characterized in that: Based on the facial image color feature contrast response coding vector, determining whether the skin color is abnormal includes: passing the facial image color feature contrast response coding vector through a classifier-based skin color abnormality discriminator to obtain a discrimination result, and the discrimination result is used to indicate whether the skin color is abnormal.
9. The method for video conferencing image enhancement processing based on environment perception according to claim 8, characterized in that: Comparing the enhanced video conference face ROI image with a preset standard facial skin color image to determine whether the skin color is abnormal, and outputting the optimized video conference image frame to the video conference interface for display, including: When the skin color is abnormal, color correction is performed on the preliminarily enhanced video conference image frame to obtain the optimized video conference image frame, and the optimized video conference image frame is output to the video conference interface for display; When the skin color is normal, the preliminarily enhanced video frame is output to the video conferencing interface for display.
10. A video conferencing image enhancement processing system based on environment perception, characterized in that: include: The original video frame acquisition module permanently obtains the original video conference image frames captured by the video conference camera; A face region detection module, configured to detect a face region from the original video conference image frame to obtain a face ROI image; An environment and enhancement strategy portrait generation module, configured to generate an environment and enhancement strategy portrait based on the original video conference image frame and the face ROI image; A video conference image frame enhancement module, configured to perform image enhancement processing on the original video conference image frame based on the environment and enhancement strategy portrait to obtain a preliminarily enhanced video conference image frame; An enhanced image frame face detection module is configured to permanently detect a face region from the initially enhanced video conference image frame to obtain an enhanced video conference face ROI image; The image display module is used to compare the enhanced video conference face ROI image with the preset standard facial skin color image to determine whether there is any abnormality in the skin color, and output the optimized video conference image frame to the video conference interface for display.