Image quality evaluation and optimization method and device, electronic equipment and storage medium

CN121482478BActive Publication Date: 2026-09-29ZKTECO CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511681162.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-09-29
Estimated Expiration
2045-11-17

AI Technical Summary

Technical Problem

[0005]本发明提供了一种图像质量评估及优化方法、装置、电子设备及存储介质,用于解决或部分解决当前图像质量评估方法评估全面性不足、场景适配性差以及存在被动判断局限的技术问题

Benefits of technology

[0021]提供了一种图像质量评估及优化方法。首先获取评估任务类型及待评估的多模态图像,以用于后续图像质量评估的基础数据。对多模态图像进行基于模态协同的多维特征提取,获得多维度特征,从而通过多模态协同,特征提取稳定性较单模态评估大幅提升,可适配低光、复杂背景、多干扰等恶劣场景,从而解决单模态评估在复杂环境下的鲁棒性不足问题,解决了当前评估全面性不足的问题。基于多维度特征及评估任务类型构建动态关联矩阵,并对动态关联矩阵进行评估权重的实时动态优化,获得图像评估结果,从而在前述特征提取流程基础上,提出一种任务驱动的动态智能评估架构,通过构建任务→特征→权重的动态联动评估架构,实现评估标准与业务需求的精准匹配。其中,通过构建动态关联矩阵,可以实现评估标准与任务需求的精准匹配,解决了传统一刀切的评估局限性,任务适配性显著提升。根据图像评估结果判断多模态图像是否合格,若否,则对多模态图像进行优化;若是,则输出并保存图像评估结果。从而通过设置图像优化机制,以将被动质量判断升级为主动质量优化,打破当前图像质量评估的被动判断局限。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482478B_ABST
    Figure CN121482478B_ABST
Patent Text Reader

Abstract

The application discloses an image quality evaluation and optimization method and device, electronic equipment and a storage medium, and is used for solving the technical problems that current image quality evaluation methods are insufficient in comprehensive evaluation, poor in scene adaptability and limited in passive judgment. The method comprises the following steps: acquiring an evaluation task type and a multi-modal image to be evaluated; performing multi-dimensional feature extraction based on modal cooperation on the multi-modal image to obtain multi-dimensional features; constructing a dynamic correlation matrix based on the multi-dimensional features and the evaluation task type, and performing real-time dynamic optimization of the evaluation weight of the dynamic correlation matrix to obtain an image evaluation result; judging whether the multi-modal image is qualified according to the image evaluation result, if not, optimizing the multi-modal image; and if yes, outputting and saving the image evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image quality assessment and optimization method, apparatus, electronic device, and storage medium. Background Technology

[0002] Image quality assessment (IQA) is an important research area in computer vision and image processing, aiming to objectively or subjectively evaluate the quality of images. With the rapid development of digital image processing technology, IQA plays a crucial role in medical image analysis, remote sensing image processing, image compression and transmission, and image enhancement. Its core objective is to quantify image quality, thereby providing a basis for optimizing and improving image processing algorithms.

[0003] Traditional image quality assessment methods are primarily based on the statistical or structural features of images, such as mean squared error, peak signal-to-noise ratio, and structural similarity index. These methods evaluate quality by calculating the differences between a reference image and the image to be evaluated, or by extracting features such as edges and textures. However, these technical solutions generally suffer from three core technical bottlenecks, making it difficult to meet the high-precision application requirements in complex scenarios.

[0004] On the one hand, traditional methods mostly rely on the basic visual features of single-modal images for evaluation, resulting in limited information dimensions and insufficient comprehensiveness. On the other hand, traditional schemes use fixed thresholds and static weights for quality judgment, leading to a significant drop in accuracy when the same evaluation standard is applied to multiple scenarios, rigid evaluation logic, and poor scenario adaptability. Furthermore, current technologies only output a binary quality level of "pass" or "fail," failing to guide users or systems to make targeted adjustments. This results in high resampling rates for low-quality images, low processing efficiency, and limitations in passive judgment, offering poor guidance for optimization. Summary of the Invention

[0005] This invention provides an image quality assessment and optimization method, apparatus, electronic device, and storage medium to solve or partially solve the technical problems of insufficient comprehensiveness, poor scene adaptability, and passive judgment limitations in current image quality assessment methods.

[0006] This invention provides an image quality assessment and optimization method, the method comprising:

[0007] Acquire the evaluation task type and the multimodal images to be evaluated;

[0008] Multi-dimensional feature extraction based on modal collaboration is performed on the multimodal image to obtain multi-dimensional features;

[0009] A dynamic correlation matrix is ​​constructed based on the multi-dimensional features and the evaluation task type, and the evaluation weights of the dynamic correlation matrix are dynamically optimized in real time to obtain the image evaluation result.

[0010] Based on the image evaluation results, determine whether the multimodal image is qualified. If not, optimize the multimodal image; if yes, output and save the image evaluation results.

[0011] The present invention also provides an image quality assessment and optimization apparatus, comprising:

[0012] The data acquisition unit is used to acquire the evaluation task type and the multimodal images to be evaluated;

[0013] The feature extraction unit is used to perform multi-dimensional feature extraction based on modal collaboration on the multimodal image to obtain multi-dimensional features;

[0014] The real-time dynamic optimization unit is used to construct a dynamic correlation matrix based on the multi-dimensional features and the evaluation task type, and to perform real-time dynamic optimization of the evaluation weights of the dynamic correlation matrix to obtain the image evaluation result.

[0015] The image quality judgment unit is used to determine whether the multimodal image is qualified based on the image evaluation result. If not, the multimodal image is optimized; if so, the image evaluation result is output and saved.

[0016] The present invention also provides an electronic device, the device comprising a processor and a memory:

[0017] The memory is used to store program code and transmit the program code to the processor;

[0018] The processor is used to execute the image quality assessment and optimization method as described above, according to the instructions in the program code.

[0019] The present invention also provides a computer-readable storage medium for storing program code for performing the image quality assessment and optimization method as described in any of the preceding claims.

[0020] As can be seen from the above technical solutions, the present invention has the following advantages:

[0021] This paper presents an image quality assessment and optimization method. First, it acquires the assessment task type and the multimodal images to be assessed, serving as the foundational data for subsequent image quality assessment. Multimodal image multidimensional feature extraction based on modal collaboration is performed to obtain multidimensional features. Through multimodal collaboration, the stability of feature extraction is significantly improved compared to single-modal assessment, making it adaptable to harsh scenarios such as low light, complex backgrounds, and multiple interferences. This addresses the insufficient robustness of single-modal assessment in complex environments and solves the problem of insufficient comprehensiveness in current assessments. A dynamic correlation matrix is ​​constructed based on the multidimensional features and assessment task type, and the assessment weights are dynamically optimized in real time to obtain the image assessment results. Based on the aforementioned feature extraction process, a task-driven dynamic intelligent assessment architecture is proposed. By constructing a dynamic linkage assessment architecture of task → feature → weight, precise matching between assessment standards and business requirements is achieved. Specifically, by constructing the dynamic correlation matrix, precise matching between assessment standards and task requirements is achieved, overcoming the limitations of traditional one-size-fits-all assessments and significantly improving task adaptability. The image evaluation results determine whether the multimodal image is qualified. If not, the multimodal image is optimized; if so, the image evaluation results are output and saved. This image optimization mechanism upgrades passive quality judgment to active quality optimization, breaking the limitations of the current passive image quality assessment. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 A flowchart illustrating the steps of an image quality assessment and optimization method;

[0024] Figure 2 This is a schematic diagram of the overall process of an image quality assessment and optimization method.

[0025] Figure 3 This is a structural block diagram of an image quality assessment and optimization device. Detailed Implementation

[0026] This invention provides an image quality assessment and optimization method, apparatus, electronic device, and storage medium to solve or partially solve the technical problems of insufficient comprehensiveness, poor scene adaptability, and passive judgment limitations in current image quality assessment methods.

[0027] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0028] As an example, traditional image quality assessment methods are mainly based on the statistical or structural features of images, such as mean squared error, peak signal-to-noise ratio, and structural similarity index. These methods evaluate quality by calculating the differences between a reference image and the image to be evaluated, or by extracting features such as edges and textures. However, these technical solutions generally suffer from three core technical bottlenecks, making it difficult to meet the high-precision application requirements in complex scenarios.

[0029] On the one hand, traditional methods mostly rely on the basic visual features of single-modal images for evaluation, resulting in limited information dimensions and insufficient comprehensiveness. On the other hand, traditional schemes use fixed thresholds and static weights for quality judgment, leading to a significant drop in accuracy when the same evaluation standard is applied to multiple scenarios, rigid evaluation logic, and poor scenario adaptability. Furthermore, current technologies only output a binary quality level of "pass" or "fail," failing to guide users or systems to make targeted adjustments. This results in high resampling rates for low-quality images, low processing efficiency, and limitations in passive judgment, offering poor guidance for optimization.

[0030] Therefore, one of the core inventive points of this invention is to propose an image quality assessment and optimization method. First, through a three-stage progressive process of pre-assessment, dynamic adjustment, and fine assessment, the invalid acquisition rate can be significantly reduced, making it particularly suitable for real-time acquisition scenarios on mobile devices. Second, through multimodal collaboration, the stability of feature extraction is significantly improved compared to single-modal assessment, adapting to harsh scenarios such as low light, complex backgrounds, and multiple interferences, thus solving the problem of insufficient robustness of single-modal assessment in complex environments. Based on the aforementioned feature extraction process, a task-driven dynamic intelligent assessment architecture is proposed. That is, a dynamic linkage assessment architecture of task → feature → weight is constructed to achieve precise matching between assessment standards and business requirements. Specifically, by constructing a dynamic correlation matrix, precise matching between assessment standards and task requirements can be achieved, overcoming the limitations of traditional one-size-fits-all assessments and significantly improving task adaptability. Furthermore, an AR (Augmented Reality) guided defect localization → operation guidance → acquisition optimization integrated AR feedback mechanism is proposed to upgrade passive quality judgment to proactive quality optimization. Meanwhile, based on actual task evaluation needs, a flexible and selectable threshold back-reasoning mechanism is proposed. For scenarios with high task evaluation requirements, the threshold back-reasoning mechanism can achieve a deep binding between quality evaluation and business objectives, avoiding the problems of over-evaluation or under-evaluation.

[0031] Reference Figure 1 The diagram illustrates a flowchart of an image quality assessment and optimization method provided by an embodiment of the present invention, which may specifically include the following steps:

[0032] Step 101: Obtain the evaluation task type and the multimodal image to be evaluated;

[0033] When image quality assessment is required, it is necessary to first obtain the assessment task type and the multimodal images to be assessed. The assessment task type can be any task that requires image quality assessment, such as face recognition, pedestrian detection, vehicle detection, or face anti-spoofing.

[0034] To reduce the invalid acquisition rate and thus decrease unnecessary image quality assessment calculations, this invention proposes a progressive acquisition optimization process. Its core lies in maximizing acquisition efficiency while ensuring quality through a three-stage progressive acquisition and evaluation process consisting of pre-assessment, dynamic adjustment, and fine-tuning.

[0035] Specifically, the first step is low-resolution pre-assessment, i.e., rapid screening. When the image acquisition device is triggered for the first time, it first acquires a low-resolution image (such as 640×480 pixels) and only preliminarily assesses core features (such as target integrity and global illumination entropy). This process takes less than 300ms.

[0036] If the pre-assessment score is less than a certain preset assessment threshold, such as a pre-assessment score < 0.5, it is judged as seriously unqualified. At this time, it will immediately provide feedback that the current scene does not meet the acquisition conditions and suggest basic adjustments (such as placing the target in a well-lit area) to avoid invalid high-definition acquisition.

[0037] If the pre-evaluation score is greater than or equal to the preset evaluation threshold, such as a pre-evaluation score ≥ 0.5, it is determined to have the potential for data collection, and then proceeds to the next step of dynamic parameter adjustment.

[0038] The second step is dynamic adaptation of the acquired parameters. Specifically, based on the pre-evaluation results, the acquired parameters of the equipment are automatically adjusted to reduce quality defects in the subsequent fine-tuning evaluation.

[0039] Regarding lighting, if the pre-assessment indicates that the overall lighting is too dark, the camera's ISO will be automatically increased to 800 (maximum 1600, to avoid a surge in noise) and the exposure time will be extended to a safe threshold (e.g., ≤1 / 30s for handheld shooting).

[0040] Regarding resolution, it dynamically adjusts based on the size of the target. If the target occupies less than 30% of the area, it automatically switches to the highest resolution (such as 4K); if the target occupies more than 70% of the area, it maintains 1080P resolution to balance quality and storage.

[0041] In terms of modal coordination, if the infrared modal pre-evaluation is better in low-light environments, the gain of the infrared sensor will be automatically increased and the weight of the RGB sensor will be reduced.

[0042] The third step is high-quality precision evaluation, which is the final verification process in subsequent steps 102 to 104.

[0043] After acquiring high-quality images as multimodal images to be evaluated, a full-dimensional evaluation (covering sharpness, illumination entropy, noise, modal consistency, etc.) is performed, generating a detailed quality report. If the fine evaluation score is greater than or equal to the target threshold, a qualified result is output and the image is saved. If the score does not meet the target but is close to the threshold (e.g., target 0.8, actual 0.78), a lightweight optimization algorithm (e.g., local sharpening, illumination compensation) is initiated for secondary processing. If the score meets the target after processing, it is deemed qualified. If the score difference is large (e.g., fine evaluation score < 0.7), AR-guided re-acquisition is triggered, and a more accurate defect analysis is performed based on the fine evaluation results, improving adjustment efficiency by 30% compared to the first time.

[0044] Based on the foregoing description, in some embodiments, when the image acquisition device held by the user triggers the first acquisition action, a low-resolution image can be acquired. Then, an image pre-evaluation is performed based on the low-resolution image to obtain a pre-evaluation score. In one case, when the pre-evaluation score is less than a preset evaluation threshold, an acquisition failure message and basic adjustment suggestions are sent to the acquisition interface of the image acquisition device. In another case, when the pre-evaluation score is greater than or equal to the preset evaluation threshold, the acquisition parameters of the image acquisition device are dynamically adjusted based on the pre-evaluation result, and after the parameter adjustment, the image acquisition process is started to obtain the multimodal image to be evaluated.

[0045] This gradual process can reduce the invalid data collection rate by more than 50% and the average data collection time by 40%. It is particularly suitable for real-time data collection scenarios on mobile devices (such as mobile ID card photography and remote identity verification).

[0046] Step 102: Perform multi-dimensional feature extraction based on modal collaboration on the multimodal image to obtain multi-dimensional features;

[0047] To address the issue of insufficient comprehensiveness in traditional technology assessments, this invention proposes a feature extraction architecture that combines modal collaboration with regional differentiation. Through multi-dimensional and refined feature quantification, it achieves a comprehensive assessment of image quality.

[0048] First, there is a multi-scale, region-adaptive sharpness assessment for the sharpness dimension. Addressing the technical challenge of varying importance across different image regions, this invention constructs a region-differentiated sharpness assessment model based on image semantic segmentation results. Specifically, it first provides a region division and assessment strategy, and then presents the calculation process for the comprehensive sharpness index.

[0049] Specifically, a lightweight semantic segmentation network (such as MobileNet+U-Net) is first used to divide the multimodal image to be evaluated into high-information regions (such as core regions of target objects like faces and hands) and background regions (such as non-critical regions like the sky and walls). Differentiated evaluation strategies are then adopted for the two types of regions.

[0050] For high-information areas, multi-scale Laplacian operators (1×3, 3×1, 3×3) are used to calculate gradient variance. Among them, the 1×3 and 3×1 directional operators are used to capture subtle texture blurs in the horizontal and vertical directions, while the 3×3 standard operator is used to evaluate overall sharpness. Multi-scale fusion ensures that no key details are blurred.

[0051] For the background area, a single-scale (5×5) Gaussian gradient operator is used to calculate the edge intensity, allowing for moderate blurring and avoiding non-critical areas from affecting the overall quality judgment.

[0052] For the calculation process of the comprehensive sharpness index, a dynamic weighting of regional importance coefficients is introduced, adjusting the coefficient values ​​according to the task scenario corresponding to the evaluation task type. For example, in scenarios centered on target recognition (such as face recognition and license plate recognition), the coefficient for high-information areas is set to 0.8-0.9, and the coefficient for background areas is set to 0.1-0.2. Similarly, in scenarios centered on scene reconstruction (such as surveillance scenarios), the coefficient for high-information areas is set to 0.7-0.8, and the coefficient for background areas is set to 0.2-0.3. Comprehensive Sharpness Features The calculation formula is as follows:

[0053]

[0054] in, It is the regional importance coefficient; It is the mean of the multi-scale gradient variance in the high-information region; It is the mean edge intensity of the Gaussian gradient in the background region.

[0055] The above process can effectively identify low-quality images with blurred high-information areas but clear backgrounds, solving the problem that traditional global sharpness assessment cannot distinguish the importance of regions. The accuracy of the overall sharpness assessment is improved by more than 35% compared with the traditional global assessment.

[0056] Secondly, there is the assessment of the dynamic distribution entropy of illumination in the illumination dimension, which includes two main processes: quantization and fusion.

[0057] To address the shortcomings of traditional brightness assessment methods that fail to consider both global and local conditions, this invention designs a dynamic light distribution entropy index that integrates global distribution uniformity with the effectiveness of key local areas.

[0058] The entropy quantization process for dynamic illumination distribution mainly implements the following three-layer illumination feature extraction:

[0059] Global Illumination Entropy The information entropy of the image's grayscale histogram (levels 0-255) is calculated as follows:

[0060]

[0061] in, It is the first The percentage of pixels at one gray level, with an original value range of 0-8. When all pixels are concentrated at one gray level... When evenly distributed across 256 gray levels, .

[0062] Regional Illumination Deviation Divide the image into a 16×16 uniform grid and calculate the mean gray value of each grid. Compared with the global mean deviation rate With the first Taking one grid as an example, the calculation process is as follows:

[0063]

[0064] , The larger the number, the more likely it is to be the first. The more significant the difference between the lighting of individual grids and the global lighting (e.g., local overexposure / underexposure); The deviation rate is for all 256 grids. The statistical analysis yielded the percentage of grid cells with abnormal lighting. Among them, The calculation process is as follows:

[0065]

[0066] The deviation rate is defined here. A mesh with more than 20% of the mesh is considered an anomalous mesh. (The mesh was determined to be an anomalous lighting mesh). This represents the number of abnormal grid cells. The original value range is 0-1 (when there are no abnormal grids). When all meshes are abnormal ).

[0067] High information area illumination contrast entropy Calculate the grayscale contrast (absolute value of the grayscale difference between adjacent pixels) distribution entropy of pixels within the high-information area. The calculation process is as follows:

[0068]

[0069] in, It is the first The pixel percentage of contrast level, with an original value range of 0-8, when the contrast is fully concentrated. When uniformly distributed, .

[0070] For the dynamic distribution entropy fusion process of illumination, since the original value ranges of the three illumination features are different ( and It ranges from 0 to 8. The values ​​(0-1) need to be normalized to the 0-1 range to ensure fairness in weighted averages. The specific normalization process follows the logic tailored to the specific scenario:

[0071] Global illumination entropy normalization : The original maximum value is 8. According to the maximum value normalization method, then... After normalization The closer the value is to 1, the more uniform the global illumination distribution.

[0072] Regional illumination deviation normalization : The original range was 0-1, but the deviation rate A higher value indicates worse lighting quality, requiring a reverse mapping (converting the degree of anomalousness to a quality level). After normalization The closer the value is to 1, the less local lighting abnormality there is.

[0073] Entropy normalization of illumination contrast in high-information areas :and The normalization logic is consistent; normalization is performed based on the maximum value. Therefore... After normalization The closer the value is to 1, the more balanced the illumination contrast in the high-information area and the stronger the recognition of key features.

[0074] Based on the rule that the weights are adjusted according to the scene's lighting sensitivity, the normalized three-layer lighting features are weighted and summed to obtain the final dynamic lighting distribution entropy feature. The calculation process is as follows:

[0075]

[0076] in, , , It changes dynamically with changes in scene lighting. For example, by default, in a normal lighting scene (50 lux < light intensity < 5000 lux), Set to 0.35, Set to 0.35, Set to 0.3. Adjust the weights incrementally based on the scene (step size ±0.05), starting with low-light scenes (light intensity <50 lux). →Normal lighting scene →High-light scene The global illumination weight gradually decreases. And in low-light scenes... →Normal lighting scene →High-light scene The weight of local lighting gradually increases, forming a pattern where the stronger the scene lighting, the higher the attention given to local anomalies.

[0077] Entropy index of dynamic distribution of illumination A value closer to 1 indicates better light quality. Definition For excellent lighting, For the lighting to be adequate, Due to insufficient light, This indicates a severe lighting anomaly.

[0078] The above-described dynamic light distribution entropy extraction process solves the problem of misjudging brightness in traditional brightness assessments where the overall brightness is normal but local areas are overexposed or underexposed. The light anomaly recognition rate is improved by 40%, and it can effectively adapt to complex lighting environments.

[0079] Modal consistency assessment features and noise intensity assessment features can be extracted through a multimodal feature complementarity enhancement process. Specifically, for the assessment needs of multimodal images (such as RGB+infrared, RGB+depth, infrared+depth, RGB+infrared+depth, etc.), this invention constructs an integrated feature extraction model of "modal collaboration + complementarity + denoising," overcoming the robustness limitations of single-modal assessment. Modal consistency assessment features can be obtained by implementing intermodal consistency verification and dynamic complementary weight allocation. Noise intensity assessment features can be obtained by implementing a cross-modal noise suppression process.

[0080] The principle of the intermodal consistency verification process is to calculate the feature mapping error of the same high-information region under different modes in order to verify the reliability of the modal data.

[0081] Specifically, in RGB+infrared images, the correlation between temperature distribution and grayscale distribution in high-information areas can be calculated, or an attention weight can be assigned to each pixel or region based on the color, texture, and other information of the RGB image and the temperature information of the infrared image. If the correlation is less than 0.3 (e.g., the infrared image shows a high-temperature area but there is no corresponding target in the visible light), and / or the difference between the RGB and infrared attention weights of a certain region exceeds 30%, it is marked as a modal conflict.

[0082] In the event of a modal conflict, the more reliable modal feature is prioritized. For example, in scenarios with extreme lighting (strong backlight / ultra-low light) or those relying on temperature differences for identification, infrared features are preferred. However, for scenarios with high latency requirements, such as normal lighting, tasks requiring fine texture recognition, real-time inspection, and real-time tracking in video surveillance, RGB features are preferred to balance efficiency and accuracy. Specific examples are as follows:

[0083] Example 1 illustrates a low-light scene conflict: Nighttime rural road security monitoring (illuminance <20 lux). Infrared imagery detects a 37°C thermal response area on the roadside (identified as a pedestrian), but the RGB imagery, due to insufficient light, only shows a blurry dark patch without a clear human outline. The correlation between temperature distribution and grayscale distribution in this area is calculated to be 0.23, while the RGB attention weight is 0.35 and the infrared attention weight is 0.72, a difference of 37%, triggering a modal conflict. Because this is an ultra-low-light scene, infrared features are prioritized for pedestrian tracking, while the RGB weight for this area is reduced by 40%.

[0084] Example 2 illustrates a texture recognition conflict under normal lighting conditions: During daytime shopping mall product traceability (500 lux illumination), the RGB image clearly displays the QR code texture on the product packaging (sharp edges, complete details), with an attention weight of 0.81. The corresponding area in the infrared image exhibits a flat thermal response (temperature difference from the packaging material <1℃), with an attention weight of 0.46, representing a 35% difference. The correlation between the two is calculated to be 0.28, which is marked as a modal conflict. Since texture recognition of the QR code is required, RGB features are prioritized to ensure traceability accuracy.

[0085] In an RGB+depth image, the matching degree between the RGB texture edges and depth contours in high-information areas is compared. If the matching error exceeds 15% (e.g., the RGB displays the object's edge but the depth has no corresponding contour), the modal data in that region is determined to be abnormal, and the evaluation weight of the abnormal modality is reduced. A specific example is as follows:

[0086] Example 1 illustrates depth anomalies caused by glass reflection: In an indoor smart office scenario, the RGB image clearly displays the rectangular edges of desktop documents (continuous contours extracted using the Canny operator), but the depth image shows chaotic and fragmented depth values ​​in the corresponding areas due to reflections from the document's surface coating. Calculations show a matching accuracy of 79% and an error of 21%, indicating a depth modality anomaly. In this case, the weight of the depth region is reduced by 35%, prioritizing the use of RGB texture edges for document shape recognition.

[0087] Example 2 illustrates depth blurring due to excessive distance: In a park security system for long-distance personnel detection (10 meters away), the RGB image, when magnified, can identify the edges of the human shoulder and head. However, due to limitations in ranging accuracy, the depth image lacks a clear three-dimensional outline in the corresponding area, only presenting a diffuse depth distribution. Calculations yielded a matching degree of 73% with an error of 27%, indicating a depth modality anomaly. In this case, the depth weight was reduced by 40%, and RGB edge features were used to assist in personnel counting.

[0088] In infrared + depth images, the overlap between the infrared thermal gradient contour and the depth geometric contour of high-information areas (such as faces, object cores) is calculated as (intersection area / union area of ​​the two modal contours × 100%). If the overlap is less than 65% (e.g., infrared detects a human thermal contour but depth shows no 3D structure, or depth shows an object contour but infrared shows no thermal response), it is marked as a modal conflict. In this case, the modality that matches the scene characteristics is prioritized (infrared is prioritized in the absence of light sources at night, and depth is prioritized in complex occlusion). The angle between the infrared thermal gradient direction of high-information areas (e.g., the temperature decrease direction from the human torso to the limbs) and the depth normal direction (e.g., the 3D orientation of an object surface) is calculated. If the average angle exceeds 30° (e.g., infrared shows the thermal gradient of an arm is downward but the depth normal direction is upward), it is marked as a modal data anomaly, and the weight of the modality with the largest deviation between the thermal gradient and the normal is reduced. A specific example is as follows:

[0089] Example 1 illustrates a conflict scenario in the absence of ambient light at night: During a nighttime search and rescue operation in the wilderness (without ambient light), infrared images capture the thermal gradient contours of a human body (torso 36.5℃, limbs 33℃, complete contours). However, depth images show broken three-dimensional structures in the corresponding areas due to vegetation obstruction. The overlap is calculated to be 61%. Given this is a nighttime scenario with no ambient light, infrared features are prioritized to pinpoint the location of the trapped person, and the weight of the depth image in this area is reduced by 30%.

[0090] Example 2 illustrates a complex occlusion depth-first scenario: For inventory counting on warehouse shelves (goods partially obscured by dust covers), the depth image detects the three-dimensional structure of the cuboid goods on the shelves through spatial contour detection (continuous depth gradient). The infrared image shows no significant thermal response in the corresponding area due to the heat insulation of the dust cover. Calculations show an overlap of only 59%. In this complex occlusion scenario, depth features are prioritized for calculating the goods volume, reducing the infrared weight by 50%.

[0091] In an RGB+infrared+depth trimodal image, the trimodal features in high-information areas are converted into binary hash values ​​using a feature hashing algorithm. Then, the hash similarities between RGB-infrared, RGB-depth, and infrared-depth are calculated. If any two sets of similarities are less than 0.5 (e.g., the trimodal features have no intersection), a global modality conflict is identified, and a backup single-modal evaluation is initiated (e.g., prioritizing the preservation of spatial contour features of the depth modality). A specific example is as follows:

[0092] Example 1 illustrates a global modal conflict during heavy rain: For vehicle detection on urban roads during heavy rain, RGB images exhibit blurry and trailing effects due to rain and fog scattering, infrared images suffer from fragmented thermal contours due to raindrop thermal interference, and depth images display significant speckle noise due to water vapor reflection. Calculating the hash similarity of the high-information region (vehicle core area) yields RGB-IR 0.38, RGB-Depth 0.41, and Infrared-Depth 0.35, all below 0.5. This indicates a global modal conflict, prompting the activation of a backup scheme. Prioritizing the preservation of spatial contour features from the depth modality for vehicle localization, the weights of the other two modalities are reduced by 60%.

[0093] Example 2 illustrates a global conflict in a dense smoke environment: In fire scene rescue, RGB images are completely obscured by dense smoke, lacking effective features; infrared images exhibit chaotic thermal outlines (no clear target shape) due to high-temperature radiation from the fire; and depth images show drastic changes in depth values ​​due to smoke scattering. The similarity scores of the three modal hashes are 0.29, 0.34, and 0.27, respectively, triggering a global conflict. Prioritizing the preservation of relatively concentrated thermal response areas in the infrared modality (excluding isolated high-temperature points), this helps to roughly locate the area of ​​trapped personnel, while reducing the weights of depth and RGB by 70%.

[0094] By verifying intermodal consistency, the evaluation results for the entire region are prevented from being affected by a single modal failure. Retaining the weights of normal modes maximizes the use of effective data and ensures the robustness of the evaluation.

[0095] The core of the dynamic complementary weight allocation process lies in adjusting the evaluation weights of each modality in real time based on modal characteristics and scenario requirements.

[0096] For low-light environments (light intensity <50 lux), the infrared modal weight is increased from 0.3 to 0.6, and the RGB modal weight is decreased from 0.7 to 0.4, taking advantage of infrared night vision to compensate for the loss of RGB image quality.

[0097] For complex background scenarios (such as densely populated areas), the depth modality weight is increased from 0.2 to 0.4, filtering background interference through depth information and focusing on the evaluation of the main target.

[0098] For normal lighting scenarios, the RGB modal weight is kept at 0.6-0.7, and the weights of other modalities are 0.3-0.4, in order to balance evaluation accuracy and computational efficiency.

[0099] For strong backlighting scenes (light intensity > 10000 lux), the RGB mode weight is reduced from 0.5 to 0.2, the infrared mode weight is 0.45, and the depth mode weight is 0.35. Infrared mode is used to avoid overexposure, and depth mode is used to supplement contour details.

[0100] For fine-grained detection scenarios (such as face anti-spoofing and object material recognition), the RGB modality weight is 0.4, the infrared modality weight is 0.3, and the depth modality weight is 0.3. RGB is responsible for texture details, infrared verifies liveness detection and thermal distribution, and depth provides 3D morphology verification.

[0101] For severe weather scenarios (rain, fog, dust), the infrared + depth fusion weight is 0.7 (infrared 0.4, depth 0.3), and the RGB modality weight is 0.3. Infrared penetration through occlusion and depth filtering of interference improve the robustness of the evaluation.

[0102] The principle behind cross-modal noise suppression is to use the correlation between modes to distinguish between real features and noise.

[0103] In the RGB+depth image, the edges of high-information areas in the RGB image are extracted using the Canny operator, while median filtering (3×3 window) is applied to the depth image to remove speckle interference. The variance of depth values ​​in the same region is calculated. If the variance is <0.05 (no significant change in depth value) and the RGB grayscale value drops sharply by ≥40% compared to adjacent non-shaded areas, or if the brightness V <0.2 and saturation S >0.6 after conversion to HSV space (meeting the color space characteristics of shadows), it is identified as shadow noise. For subsequent optimization of shadow noise, morphological closing operations (kernel size 5×5) can be used to optimize the contour of the shadow region, reducing the brightness evaluation weight of this region by 50%. Simultaneously, referring to the spatial structure information of the depth image, local brightness compensation is performed on the RGB shadow region (compensation coefficient = average brightness of adjacent non-shaded areas × 0.3) to avoid misjudging shadows as being too dark.

[0104] In RGB+infrared images, an adaptive thresholding algorithm (threshold = local area grayscale mean + 2 standard deviations) is used to extract highlight regions from the RGB image, while simultaneously calculating the thermal response value (unit: ℃) of the corresponding region in the infrared image. If the difference between the infrared thermal response value of the highlight region and the thermal response of the surrounding environment is < 2℃ (no corresponding thermal feature), and the area of ​​the highlight region is < 50 pixels (excluding real bright targets), it is identified as highlight noise such as light reflection or metallic reflection. For subsequent optimization of highlight noise, guided filtering (radius = 6, variance = 0.01) can be used to locally suppress the RGB highlight region, preserving target edge details while reducing brightness. The brightness evaluation weight of this region is reduced by 40%, and the local contrast of the RGB image is corrected based on the thermal distribution of the infrared image to avoid highlight noise interfering with the overall brightness evaluation.

[0105] In infrared + depth images, after smoothing the infrared image using Gaussian filtering (σ=1.5), isolated high-temperature pixels (pixels whose grayscale value is ≥30% higher than the average of the surrounding 8 pixels) are detected using the eight-neighbor method. Simultaneously, the spatial contour of the corresponding location in the depth image is extracted (depth gradient calculated using the Sobel operator). If the depth gradient value is <10 (no corresponding spatial contour) and the high-temperature pixel has no continuous thermal distribution area (connected region area <30 pixels), it is identified as environmental thermal noise (such as distant heat source radiation or sensor thermal noise). For subsequent optimization of environmental thermal noise, bilateral filtering (spatial standard deviation = 5, grayscale standard deviation = 20) can be used to smooth the infrared noise region while preserving its edges. The infrared evaluation weight of this region is reduced by 60%, and isolated hot pixels with no spatial correspondence in the infrared image are removed based on the spatial contour of the depth image, avoiding interference from heat sources in target detection.

[0106] In trimodal (RGB + infrared + depth) images, trimodal cross-verification rules are established. For RGB reflective regions, the grayscale value is ≥240 (8-bit image) and saturation <0.1, while the corresponding infrared region has a thermal response difference <1.5℃ and a depth gradient <8 (no contour support). For depth speckle noise, the depth value jump is ≥50% (compared to adjacent regions), and the corresponding RGB region has no edge features and no infrared thermal response. For infrared false thermal targets, the thermal response value is ≥5℃ higher than the environmental mean, but there are no corresponding visual features in RGB and no spatial occupation in depth. For RGB reflective noise, the color evaluation weight of this region can be reduced by 50%, and the RGB texture features can be corrected by referring to the infrared thermal distribution and depth contour. For depth speckle noise, Gaussian-Laplace filtering can be used for smoothing, and abnormal depth values ​​can be replaced with the neighborhood mean. For infrared false thermal targets, the contribution of thermal features in this region can be directly blocked, and the joint verification results of RGB and depth can be given priority to achieve accurate filtering of multi-source noise.

[0107] Based on the preceding content, the implementation process for extracting multidimensional features from multimodal images using modal collaboration to obtain multidimensional features can include:

[0108] To extract sharpness assessment features, the multimodal image can first be segmented to obtain high-information regions and background regions. Then, the gradient variance of the high-information regions and the edge intensity of the background regions are calculated separately. A region importance coefficient is introduced for dynamic weighting, and the coefficient values ​​are adjusted in combination with the assessment task type. Based on the gradient variance and edge intensity, adaptive sharpness assessment of multi-scale regions is performed to obtain comprehensive sharpness features.

[0109] To extract the dynamic distribution entropy evaluation features of illumination, the dynamic distribution entropy of illumination is first quantized in the multimodal image to obtain the global illumination entropy, regional illumination deviation, and illumination contrast entropy of the high-information area. Then, the dynamic distribution entropy of illumination is fused based on the global illumination entropy, regional illumination deviation, and illumination contrast entropy of the high-information area to obtain the dynamic distribution entropy features of illumination.

[0110] For the extraction of modal consistency evaluation features, intermodal consistency verification can be performed on high-information regions of multimodal images. After verification, dynamic complementary weight allocation between modalities can be performed based on modal characteristics to obtain modal consistency features.

[0111] To extract noise intensity assessment features, the correlation between modes can be considered, and cross-modal noise suppression can be performed on multimodal images to obtain noise intensity features.

[0112] Thus, through multimodal collaboration, the stability of feature extraction is improved by 50% compared to single-modal evaluation, and it can adapt to harsh scenarios such as low light, complex backgrounds, and multiple interferences. This solves the problem of insufficient robustness of single-modal evaluation in complex environments.

[0113] Step 103: Construct a dynamic correlation matrix based on the multi-dimensional features and the evaluation task type, and perform real-time dynamic optimization of the evaluation weights on the dynamic correlation matrix to obtain the image evaluation result;

[0114] To overcome the limitations of static and rigid evaluation methods in traditional image quality assessment and address the poor adaptability of traditional technologies to various scenarios, this invention proposes a task-driven dynamic intelligent evaluation architecture based on the aforementioned feature extraction process. Specifically, it constructs a dynamic, interconnected evaluation architecture involving tasks, features, and weights to achieve precise matching between evaluation standards and business requirements.

[0115] In the preliminary work of image quality assessment, a large dataset can be collected, and based on large-scale labeled data and machine learning models, the influence weight of different assessment features on various tasks can be quantified. A dynamically adjustable dynamic correlation matrix that associates task type with assessment features can be constructed. Based on this dynamic correlation matrix, a set of influence weights can be obtained as the initial weights for constructing the dynamic correlation matrix in the subsequent actual image quality assessment.

[0116] The first step is to construct the basic dataset and train the model.

[0117] For the construction of the basic dataset, a labeled dataset covering 10+ task types (face recognition, pedestrian detection, vehicle detection, face anti-spoofing, etc.) and 500,000+ images can be built. Each image is labeled with its quality level, contribution of key features, and task performance results (such as recognition accuracy and detection recall).

[0118] For training the weight quantization model, a gradient boosting tree model can be used. The task performance results are used as labels, and evaluation features such as sharpness, illumination, noise intensity, and modal consistency are used as inputs to train a feature importance model to output the influence weights (i.e., influence coefficients, with values ​​ranging from 0 to 1) of each feature on different tasks.

[0119] Secondly, the construction of a dynamic association matrix structure and an online update mechanism is required.

[0120] For constructing the dynamic association matrix, the number of rows corresponds to the evaluation features (such as sharpness, dynamic illumination entropy, noise intensity, and modal consistency), the number of columns corresponds to the evaluation task type (such as recognition, detection, etc.), and the matrix elements are the weights of the evaluation features for that task type (which can also be understood as importance coefficients). For example, sharpness has a coefficient of 0.8 for text recognition (text recognition is sensitive to texture) and a coefficient of 0.5 for vehicle detection (vehicle detection is sensitive to contours). Dynamic illumination entropy has a coefficient of 0.7 for face recognition (faces have high requirements for illumination uniformity) and a coefficient of 0.3 for infrared target detection (infrared is not sensitive to illumination).

[0121] For the online update mechanism, it can be set to automatically trigger model retraining and update the weight coefficients of each element in the dynamic correlation matrix every time a preset number of newly labeled images are accumulated (e.g., 10,000 new labeled images). It is understood that the online update mechanism can be used continuously after the image quality assessment process begins. That is, the online update mechanism can be triggered not only during the model training phase, but also during subsequent actual image quality assessments, after each preset number of multimodal images that have been evaluated, to ensure that the matrix remains synchronized with the latest business data and scenario requirements, avoiding evaluation bias caused by model aging.

[0122] By constructing a dynamic correlation matrix, we can achieve a precise match between evaluation criteria and task requirements, overcoming the limitations of traditional one-size-fits-all evaluation and significantly improving task adaptability.

[0123] Building upon the constructed dynamic association matrix, this invention further optimizes the influence weights of matrix elements using a deep reinforcement learning framework, resulting in more accurate influence weights. Specifically, based on the initial weights of the dynamic association matrix (obtained through training), a deep reinforcement learning model is introduced to achieve real-time dynamic optimization of the evaluation weights, thereby further improving evaluation accuracy.

[0124] The design of deep reinforcement learning models mainly includes state space, action space, and reward function.

[0125] The state space mainly contains three parts of information: multi-dimensional feature values ​​of the current image (such as sharpness, illumination entropy, and noise intensity), the current evaluation task type (such as face recognition), and historical evaluation accuracy (such as the evaluation accuracy of the last 100 images), to ensure that the model can perceive the current evaluation scenario and historical performance.

[0126] The action space defines the adjustment step size for the weights of each evaluation feature. Its range is -0.1 to +0.1 (step size 0.01). For example, the sharpness weight can be adjusted from 0.8 to 0.85, and the illumination weight from 0.7 to 0.68.

[0127] The reward function uses the current task performance metric as the reward signal, aiming to balance task performance improvement with evaluation efficiency. The calculation process is as follows:

[0128]

[0129] in, This indicates the performance improvement corresponding to the current evaluation result (e.g., an increase in recognition accuracy from 90% to 93%). (3%) This indicates the performance weight, set to 0.8; This indicates the increase in evaluation time (e.g., from 5ms to 7ms, then T is 2ms). This represents the efficiency weight, set to 0.2.

[0130] The model receives a positive reward when the adjusted evaluation accuracy improves and the increase in processing time remains manageable. Conversely, the model receives a negative reward when accuracy decreases or processing time surges. This guides the weights towards optimization for higher accuracy and efficiency.

[0131] Through dynamic optimization using deep reinforcement learning, the evaluation accuracy in different task scenarios is improved by 25% to 40% compared to static weights. In face recognition tasks, the accuracy is increased from 85% to 97%, while the evaluation time is controlled within 10ms, meeting the requirements of real-time applications.

[0132] Based on the foregoing, the multi-dimensional features proposed in this embodiment of the invention can include multiple evaluation features of different dimensions. These multiple evaluation features of different dimensions can further include comprehensive sharpness features, dynamic illumination distribution entropy features, modal consistency features, and noise intensity features. In a specific implementation, the process of constructing a dynamic correlation matrix based on the multi-dimensional features and the evaluation task type, and dynamically optimizing the evaluation weights of the dynamic correlation matrix in real time to obtain the image evaluation result, can include:

[0133] First, a dynamic correlation matrix is ​​constructed based on the assessment task type and each assessment feature. In the dynamic correlation matrix, each matrix row corresponds to a different assessment feature, the matrix column corresponds to the assessment task type, the matrix element represents the influence weight of the assessment feature at the element's position on the assessment task type, and each matrix element corresponds to an initial weight.

[0134] Next, based on the evaluation task type and various evaluation features, and combined with historical evaluation accuracy, a pre-trained deep reinforcement learning model is used to dynamically optimize each initial weight in real time, thereby obtaining the optimized influence weight of each evaluation feature on the evaluation task type.

[0135] Then, based on the optimized influence weights, the quality assessment prediction score is calculated. The quality assessment prediction score and the optimized influence weights are integrated to obtain the image assessment result for the multimodal image. The quality assessment prediction score can be derived by multiplying the optimized influence weights of each assessment feature by the normalized index value of that feature. For example, if the optimized influence weights of the sharpness feature, illumination dynamic distribution entropy feature, modal consistency feature, and noise intensity feature are A, B, C, and D respectively, and the normalized values ​​of the assessment features are a, b, c, and d respectively, then the final quality assessment prediction score is: A*a + B*b + C*c + D*d.

[0136] In some embodiments, the initial weights are obtained by training a weight quantization model. The weight quantization model is used to quantify the influence weights of different evaluation features on various tasks. After completing the image quality evaluation, the multimodal images and image evaluation results can be integrated into training images and stored in the model training set. When the model training set successfully accumulates a preset number of new training images (e.g., every 10,000 images), an online update mechanism is automatically triggered to retrain the weight quantization model based on the accumulated new training images, or all training images, and update the initial weights according to the retraining results.

[0137] It's understandable that online update mechanisms and real-time dynamic weight optimization are completely different processes. Online update mechanisms refer to periodically updating the initial weight baseline of the dynamic correlation matrix in batches. Real-time dynamic weight optimization, on the other hand, involves fine-tuning the weights evaluated in each frame in real-time. The latter is a secondary optimization based on the former, forming a two-layer optimization logic of batch calibration and real-time adaptation.

[0138] Specifically, the online update mechanism first establishes a reliable initial weight baseline. For example, after retraining with 10,000 newly accumulated data images, it is found that text recognition has a higher requirement for clarity. At this point, the initial weight of "clarity-text recognition" in the dynamic correlation matrix can be updated from 0.8 to 0.82. This new baseline value will serve as the starting point for evaluating relevant features in all subsequent frames.

[0139] Dynamic weight optimization involves making personalized fine-tuning adjustments based on a baseline. For example, if a text image in a low-light environment is evaluated with an initial sharpness weight of 0.82, and the accuracy is found to be low, the reinforcement learning model will fine-tune the sharpness weight for that frame in real time, such as adjusting it to 0.87, while simultaneously reducing the illumination weight to ensure accurate evaluation of a single frame. However, this fine-tuning is only effective for the current frame. If the next frame image returns to normal, the initial weight will return to near the baseline of 0.82.

[0140] Step 104: Determine whether the multimodal image is qualified based on the image evaluation result. If not, optimize the multimodal image; if yes, output and save the image evaluation result.

[0141] Finally, the image evaluation results can be used to determine whether the multimodal image is qualified. If not, the multimodal image can be further optimized; if so, the image evaluation results can be directly output and saved to complete the image quality evaluation process.

[0142] To overcome the limitations of passive judgment in traditional technologies and to address the poor guidance for current technology optimization, this invention proposes an AR-guided feedback optimization mechanism for images with excessively low quality assessment scores. Specifically, this invention designs an integrated AR feedback mechanism encompassing defect localization, operation guidance, and data acquisition optimization, thereby upgrading passive quality judgment to proactive quality optimization.

[0143] The first step is defect visualization and annotation. Computer vision algorithms are used to locate, classify, and quantify the severity of defects in low-quality images, providing a precise basis for subsequent optimization.

[0144] Defect localization and classification in images mainly include the following scenarios:

[0145] For blur defects, a heat map of blurry areas is generated based on the multi-scale sharpness assessment results (red indicates severe blur and yellow indicates slight blur), and the blur type is marked (such as blurry texture in high-information areas, excessive background blur, etc.).

[0146] For lighting defects, based on the grid deviation rate of the dynamic distribution entropy of lighting, mark the areas of abnormal lighting (such as overexposure in the upper left corner and underexposure in the lower right corner), and analyze the causes of the abnormalities (such as uneven ambient light, lens backlight, etc.).

[0147] For noise defects, the noise type (such as reflection noise in RGB images and speckle noise in depth images) is identified by the results of multimodal noise suppression, and the noise concentration area is marked.

[0148] For modal defects, based on modal consistency verification, regions of modal conflict (such as RGB display targets but no depth contours) are marked, indicating abnormal modal data.

[0149] Based on defect localization, the severity of each defect area needs to be quantified.

[0150] Specifically, this invention introduces a defect severity index (0-100). The defect severity index is calculated by comprehensively considering the size of the defect area and its impact on the task (impact weight). For example, a severely blurred area of ​​20% of the high-information area corresponds to an index of 80 (severe defect), while a slightly blurred area of ​​10% of the background area corresponds to an index of 20 (minor defect). Priorities are assigned based on the index (0-30: low priority, no immediate adjustment required; 31-70: medium priority, optimization recommended; 71-100: high priority, adjustment required) to guide users to prioritize the handling of critical defects.

[0151] Through visual annotation, users can intuitively understand the location, type, and severity of low-quality image problems, solving the problem of only knowing the result but not the cause in traditional assessments.

[0152] Based on the defect type and scenario characteristics, it can accurately generate actionable optimization suggestions and present them through AR visualization to guide users to optimize the image acquisition process based on the optimization suggestions.

[0153] The following are some operational suggestions and examples for different defect categories.

[0154] Lighting defect guidance is mainly achieved by calculating the position of the fill light and adjusting its intensity. For example, assuming the upper left corner is overexposed, the fill light adjustment parameters are calculated (such as increasing the fill light on the right by 30% and decreasing the fill light on the left by 15%), and the AR interface displays the fill light direction arrow and intensity value. If the overall light is too dark, it can prompt "Increase the ambient light intensity to 200 lux" or "Adjust the camera exposure time from 1 / 100s to 1 / 50s".

[0155] The blur defect guidance primarily indicates the focus area and the target for sharpness adjustment. For example, assuming underfocus causes overall blur, it suggests adjusting the camera focal length to 2.5m or enabling autofocus. If a high-information area is locally blurry, the target focus area is marked, such as "Please focus the lens on the center person's face; sharpness needs to be improved by 20%," with a focus box and progress bar overlaid on the AR interface.

[0156] Posture / position defect guidance is achieved by generating rotation angle and translation distance instructions. For example, assuming the target is tilted, causing feature distortion (such as an angle deviation in an ID card photo), the rotation adjustment amount is calculated (such as a 3° clockwise rotation), and the AR displays a rotation arrow and angle value. If the target is too far away, resulting in insufficient resolution, a prompt is given to "Move closer to the shooting device by 50cm to ensure the target occupies ≥ 60% of the image area," and the AR displays distance scale lines and the target area bounding box.

[0157] For modal coordination defect guidance, if there is a conflict between infrared and RGB modes, the system will prompt "adjust the infrared light source power to 5W" or "switch to multimodal fusion mode" to ensure the consistency of dual-modal features.

[0158] AR visualization can be achieved using a dual guidance approach combining natural language and graphical annotations. Details are as follows:

[0159] A semi-transparent AR layer is overlaid on the mobile device screen in real time, with colored arrows indicating the direction for adjustment. For example, a red arrow points to the fill light position, and a blue arrow indicates the rotation direction.

[0160] Key parameters (such as distance, angle, and brightness values) are displayed as floating numeric labels in their respective areas and are updated in real time. For example, the distance value changes dynamically as the user adjusts the settings.

[0161] Complex operation steps (such as rotating first and then focusing) are demonstrated with step-by-step animations to reduce the user's understanding cost.

[0162] By setting up an AR-guided feedback optimization mechanism, user adjustment efficiency has been improved by 60%, and the number of times low-quality images are re-sampling has been reduced from an average of 3 times to 1.2 times, significantly optimizing the user experience.

[0163] Based on the preceding discussion, the image evaluation results include the quality assessment prediction score of the multimodal image and the optimized influence weights of each evaluation feature on the evaluation task type. In some embodiments, the process of determining whether the multimodal image is qualified based on the image evaluation results, and optimizing the multimodal image if not, may specifically include:

[0164] Determine if the quality assessment prediction score is greater than or equal to a preset target threshold; if not, the multimodal image is deemed unqualified. In one scenario where the multimodal image is deemed unqualified, if the quality assessment prediction score is greater than or equal to a preset mild threshold but less than the preset target threshold, a lightweight optimization algorithm is initiated to perform secondary optimization on the multimodal image.

[0165] In another scenario, when the quality assessment prediction score is less than a preset severity threshold, the multimodal image is first localized based on multiple assessment features of different dimensions to obtain at least one defect region; wherein the preset severity threshold is less than a preset mild threshold; for each defect region, a defect severity index is calculated based on the region size and the optimized impact weight corresponding to the defect region, and the severity priority of the defect region is determined according to the defect severity index; then, the defect regions with severity priorities in the highest priority range are identified as target defect regions, and operation optimization suggestions for the target defect regions are generated; then, the operation optimization suggestions are fed back to the image acquisition device held by the user in an AR visualization manner to guide the user to optimize the image acquisition process based on the operation optimization suggestions.

[0166] To achieve quality standards based on business objectives, this embodiment of the invention further constructs a regression model between quality score and task performance, namely a quality-performance mapping prediction model, which is used to support the inference of the required quality threshold based on the target performance.

[0167] The model architecture can employ a lightweight neural network (such as a 2-layer fully connected network + ReLU activation function), with the number of parameters controlled to within 1 million, ensuring fast inference (time <10ms). The model input is a multi-dimensional quality feature vector (such as sharpness, illumination entropy, noise intensity, and normalized values ​​of modal consistency), and the output is a predicted value for task performance (such as face recognition accuracy and pedestrian detection recall). A mean squared error (MSE) loss function is used, combined with an adaptive learning rate optimizer (such as AdamW), and training and optimization are performed on over 100,000 labeled images, achieving a goodness-of-fit R² ≥ 0.92 to ensure prediction accuracy.

[0168] In practical applications, the user inputs the target task performance value (e.g., face recognition accuracy ≥ 99%), and the model calculates the minimum required quality assessment score threshold through inverse mapping. For example, in a face recognition task, if the model predicts that a quality assessment score of 0.9 corresponds to a recognition accuracy of 99%, then the quality assessment score threshold is automatically set to 0.9, allowing only images with a score ≥ 0.9 to enter the subsequent recognition process.

[0169] By using a threshold-based reverse calculation mechanism, a deep alignment between quality assessment and business objectives can be achieved. This approach automatically derives the required quality threshold based on target performance, enabling a technical solution that sets quality standards according to performance targets. This avoids the problems of over-assessment (pursuing excessively high quality leading to decreased efficiency) or under-assessment (failing to meet quality standards leading to business failure).

[0170] Based on the preceding discussion, after outputting and saving the image evaluation results, it is possible to further determine whether the threshold back-calculation mechanism has been triggered.

[0171] In one scenario, the threshold back-inference mechanism is not enabled, and subsequent image recognition processes can be directly performed based on the multimodal image.

[0172] In another scenario, a threshold-based reverse estimation mechanism is enabled. This involves: first, obtaining the target task performance value input by the user; then, performing an inverse mapping calculation on the target task performance value using a pre-built quality-performance mapping prediction model to obtain the minimum quality assessment score required for the target task performance value; next, extracting the quality assessment prediction score of the multimodal image from the image assessment results and comparing it with the minimum quality assessment score; if the quality assessment prediction score is greater than or equal to the minimum quality assessment score, then executing the subsequent image recognition process based on the multimodal image; if the quality assessment prediction score is less than the minimum quality assessment score, then not executing the subsequent image recognition process for the multimodal image.

[0173] Understandably, threshold back-calculation is an optional, rather than mandatory, pre-processing step in image recognition for precise screening. Even without threshold back-calculation, the previously provided image quality assessment workflow still constitutes a complete solution. The difference between the two solutions lies in their applicable scenarios and optimization priorities.

[0174] The main purpose of threshold back-engineering is to align with business performance goals, rather than simply performing pre-processing filtering. For example, if a user requires a face recognition accuracy rate of ≥99%, threshold back-engineering calculates a quality assessment score threshold (e.g., 0.9) that just meets that accuracy requirement, allowing only images with a quality score ≥0.9 to proceed to subsequent recognition. Essentially, it uses business results to deduce quality standards, avoiding issues such as acceptable quality but substandard business performance or excessively pursuing high quality and wasting computing power.

[0175] Assuming no threshold-based back-factor step is set, the initial screening can be completed by the preceding evaluation process. That is, the quality evaluation score optimized by dynamic weights (e.g., setting ≥0.8 as acceptable) directly determines whether an image enters the subsequent recognition process. In this case, the acceptance standard is a general quality standard, rather than a precise standard tied to business objectives. For scenarios with low evaluation requirements (such as ordinary surveillance image acquisition), unclear business objectives (no fixed accuracy requirements), or those pursuing the highest pass rate (such as non-critical information detection), a threshold-based back-factor step can be omitted.

[0176] This invention proposes an image quality assessment and optimization method. First, a three-stage progressive process—pre-assessment, dynamic adjustment, and fine-tuning—significantly reduces the invalid acquisition rate, making it particularly suitable for real-time mobile acquisition scenarios. Second, through multimodal collaboration, feature extraction stability is greatly improved compared to single-modal assessment, adapting to harsh scenarios such as low light, complex backgrounds, and multiple interferences, thus solving the problem of insufficient robustness of single-modal assessment in complex environments. Based on the aforementioned feature extraction process, a task-driven dynamic intelligent assessment architecture is proposed. This involves constructing a dynamic linkage assessment architecture of task → feature → weight, achieving precise matching between assessment standards and business requirements. Specifically, by constructing a dynamic correlation matrix, precise matching between assessment standards and task requirements can be achieved, overcoming the limitations of traditional one-size-fits-all assessments and significantly improving task adaptability. Furthermore, an AR-guided integrated AR feedback mechanism—defect localization → operation guidance → acquisition optimization—is proposed to upgrade passive quality judgment to proactive quality optimization. Meanwhile, based on actual task evaluation needs, a flexible and selectable threshold back-reasoning mechanism is proposed. For scenarios with high task evaluation requirements, the threshold back-reasoning mechanism can achieve a deep binding between quality evaluation and business objectives, avoiding the problems of over-evaluation or under-evaluation.

[0177] For better explanation, refer to Figure 2 This diagram illustrates the overall flow of an image quality assessment and optimization method provided by an embodiment of the present invention. It should be noted that this embodiment only provides a brief description of the general flow of image quality assessment and optimization. The specific implementation process of each step can be understood by referring to the relevant content in the foregoing embodiments, and will not be elaborated upon here. It is understood that the present invention does not impose any limitations on this.

[0178] Step 201: The image acquisition device held by the user is detected to have triggered the first acquisition action. A low-resolution image is acquired, and a pre-evaluation of the image is performed based on the low-resolution image to obtain a pre-evaluation score. It is then determined whether the pre-evaluation score is less than a preset evaluation threshold. If yes, proceed to step 202; otherwise, proceed to step 203.

[0179] Step 202: Send an acquisition failure message and basic adjustment suggestions to the acquisition interface of the image acquisition device;

[0180] Step 203: Based on the pre-evaluation results, dynamically adjust the acquisition parameters of the image acquisition device, and after the parameter adjustment, start the image acquisition process to obtain the multimodal image to be evaluated;

[0181] Step 204: Obtain the evaluation task type, perform multi-dimensional feature extraction based on modal collaboration on the multimodal image, and obtain multi-dimensional features;

[0182] Step 205: Construct a dynamic correlation matrix based on multi-dimensional features and evaluation task type, perform real-time dynamic optimization of evaluation weights on the dynamic correlation matrix, obtain image evaluation results, and determine whether the multimodal image is qualified based on the image evaluation results; if not, proceed to step 206; if yes, proceed to step 207 after outputting and saving the image evaluation results.

[0183] Step 206: Optimize the multimodal image based on the AR-guided feedback optimization mechanism, generate operation optimization suggestions, and feed them back to the image acquisition device held by the user;

[0184] Step 207: Determine whether the threshold back-pushing mechanism has been triggered; if not, proceed to step 208; if yes, proceed to step 209.

[0185] Step 208: Directly execute subsequent image recognition processes based on the multimodal images;

[0186] Step 209: Determine the minimum quality assessment score based on the threshold back-inference mechanism, and determine whether to execute the subsequent image recognition process for multimodal images based on the comparison between the minimum quality assessment score and the image assessment results.

[0187] Reference Figure 3 The diagram illustrates a structural block diagram of an image quality assessment and optimization device provided in an embodiment of the present invention, which may specifically include:

[0188] The data acquisition unit 301 is used to acquire the evaluation task type and the multimodal image to be evaluated;

[0189] Feature extraction unit 302 is used to perform multi-dimensional feature extraction based on modal collaboration on the multimodal image to obtain multi-dimensional features;

[0190] The real-time dynamic optimization unit 303 is used to construct a dynamic correlation matrix based on the multi-dimensional features and the evaluation task type, and to perform real-time dynamic optimization of the evaluation weights of the dynamic correlation matrix to obtain the image evaluation result.

[0191] The image quality judgment unit 304 is used to judge whether the multimodal image is qualified based on the image evaluation result. If not, the multimodal image is optimized; if yes, the image evaluation result is output and saved.

[0192] In one optional embodiment, the feature extraction unit 302 includes:

[0193] The image segmentation unit is used to segment the multimodal image to obtain high-information regions and background regions;

[0194] A region information calculation unit is used to calculate the gradient variance of the high-information region and the edge intensity of the background region, respectively.

[0195] An adaptive sharpness evaluation unit is used to introduce a dynamic weighted regional importance coefficient, and adjust the coefficient value in combination with the evaluation task type. Based on the gradient variance and the edge intensity, an adaptive sharpness evaluation of multi-scale regions is performed to obtain comprehensive sharpness features.

[0196] The illumination dynamic distribution entropy quantization processing unit is used to perform illumination dynamic distribution entropy quantization processing on the multimodal image to obtain global illumination entropy, regional illumination deviation, and illumination contrast entropy of high information area.

[0197] The dynamic distribution entropy fusion unit is used to fuse the dynamic distribution entropy of illumination based on the global illumination entropy, the regional illumination deviation and the illumination contrast entropy of the high information area to obtain the dynamic distribution entropy features of illumination.

[0198] The intermodal consistency verification unit is used to perform intermodal consistency verification on the high-information regions of the multimodal image, and after the verification is completed, to perform dynamic complementary weight allocation between modalities based on modal characteristics to obtain modal consistency features.

[0199] A cross-modal noise suppression unit is used to consider the correlation between modes and perform cross-modal noise suppression on the multimodal image to obtain noise intensity features.

[0200] In one optional embodiment, the multi-dimensional features include multiple evaluation features of different dimensions; the real-time dynamic optimization unit 303 includes:

[0201] A dynamic correlation matrix construction unit is used to construct a dynamic correlation matrix based on the evaluation task type and each of the evaluation features. Each row of the dynamic correlation matrix corresponds to a different evaluation feature, the column corresponds to the evaluation task type, and the matrix element represents the influence weight of the evaluation feature at the element's position on the evaluation task type. Each matrix element corresponds to an initial weight. The multiple evaluation features of different dimensions include comprehensive sharpness features, dynamic light distribution entropy features, modal consistency features, and noise intensity features.

[0202] The initial weight optimization unit is used to dynamically optimize each of the initial weights in real time based on the evaluation task type and each of the evaluation features, combined with historical evaluation accuracy, using a pre-trained deep reinforcement learning model, to obtain the optimized influence weight of each of the evaluation features on the evaluation task type.

[0203] The quality assessment prediction score calculation unit is used to calculate the quality assessment prediction score based on each of the optimized influence weights.

[0204] The image evaluation result integration unit is used to integrate the quality evaluation prediction score and each of the optimized influence weights to obtain the image evaluation result of the multimodal image.

[0205] In one optional embodiment, the initial weights are obtained by training a weight quantization model; the weight quantization model is used to quantify the influence weights of different evaluation features on various tasks; the apparatus further includes:

[0206] The training image storage unit is used to integrate the multimodal images and the image evaluation results into training images after completing the image quality assessment, and store them in the model training set.

[0207] The model retraining unit is used to automatically trigger an online update mechanism when the model training set successfully accumulates a preset number of new training images, to retrain the weight quantization model based on the accumulated new training images, or all training images, and to update the initial weights according to the retraining results.

[0208] In one optional embodiment, the multi-dimensional features include multiple evaluation features of different dimensions; the image evaluation result includes the quality evaluation prediction score of the multimodal image, and the optimized influence weight of each evaluation feature on the evaluation task type; the image quality judgment unit 304 includes:

[0209] A quality assessment score determination unit is used to determine whether the predicted quality assessment score is greater than or equal to a preset target threshold.

[0210] An image defect determination unit is used to determine that the multimodal image is defective;

[0211] A lightweight optimization unit is used to activate a lightweight optimization algorithm to perform secondary optimization on the multimodal image when the quality assessment prediction score is greater than or equal to a preset lightweight threshold and less than a preset target threshold.

[0212] A defect localization unit is used to locate defects in the multimodal image based on the multiple different dimensions of evaluation features when the quality assessment prediction score is less than a preset severe threshold, thereby obtaining at least one defect region; the preset severe threshold is less than the preset mild threshold.

[0213] The severity priority determination unit is used to calculate the severity index for each defect region based on the region size of the defect region and the optimized influence weight corresponding to the defect region, and to determine the severity priority of the defect region according to the severity index.

[0214] An operation optimization suggestion generation unit is used to identify defect areas with severity priorities in the highest priority range as target defect areas and generate operation optimization suggestions for the target defect areas.

[0215] The AR feedback guidance unit is used to provide the operation optimization suggestions to the user's image acquisition device in an AR visualization manner, so as to guide the user to optimize the image acquisition process based on the operation optimization suggestions.

[0216] In one alternative embodiment, the device further includes:

[0217] The low-resolution image acquisition unit is used to acquire a low-resolution image when the user-held image acquisition device triggers the first acquisition action.

[0218] An image pre-evaluation unit is used to perform image pre-evaluation based on the low-resolution image and obtain a pre-evaluation score;

[0219] The basic adjustment suggestion sending unit is used to send a collection failure prompt message and basic adjustment suggestions to the collection interface of the image acquisition device when the pre-evaluation score is less than the preset evaluation threshold.

[0220] The acquisition parameter dynamic adjustment unit is used to dynamically adjust the acquisition parameters of the image acquisition device based on the pre-evaluation result when the pre-evaluation score is greater than or equal to the preset evaluation threshold, and to start the image acquisition process after the parameter adjustment to obtain the multimodal image to be evaluated.

[0221] In one alternative embodiment, the device further includes:

[0222] The threshold back-calculation mechanism trigger judgment unit is used to determine whether the threshold back-calculation mechanism is triggered;

[0223] The threshold back-reasoning mechanism does not execute the unit; it is used to directly execute the subsequent image recognition process based on the multimodal image.

[0224] The inverse mapping calculation unit is used to obtain the target task performance value input by the user, and perform inverse mapping calculation on the target task performance value through a pre-built quality-performance mapping prediction model to obtain the minimum quality assessment score required for the target task performance value.

[0225] The quality assessment score comparison unit is used to extract the quality assessment prediction score of the multimodal image from the image assessment result, and compare the quality assessment prediction score with the minimum quality assessment score;

[0226] An image recognition process execution initiation unit is used to execute subsequent image recognition processes based on the multimodal image when the quality assessment prediction score is greater than or equal to the minimum quality assessment score.

[0227] The image recognition process skip unit is used to prevent the subsequent image recognition process of the multimodal image from being executed when the quality assessment prediction score is less than the minimum quality assessment score.

[0228] As the device embodiment is basically similar to the method embodiment, it is described in a relatively simple way. For relevant details, please refer to the description of the method embodiment above.

[0229] This invention also provides an electronic device, which includes a processor and a memory:

[0230] The memory is used to store program code and transfer the program code to the processor;

[0231] The processor is used to execute the image quality assessment and optimization method of any embodiment of the present invention according to the instructions in the program code.

[0232] This invention also provides a computer-readable storage medium for storing program code for executing the image quality assessment and optimization method of any embodiment of this invention.

[0233] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0234] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0235] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units through some interfaces, and may be electrical, mechanical, or other forms.

[0236] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0237] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0238] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0239] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image quality assessment and optimization method, characterized in that, include: Acquire the evaluation task type and the multimodal images to be evaluated; Multi-dimensional feature extraction based on modal collaboration is performed on the multimodal image to obtain multi-dimensional features; A dynamic correlation matrix is ​​constructed based on the multi-dimensional features and the evaluation task type, and the evaluation weights of the dynamic correlation matrix are dynamically optimized in real time to obtain the image evaluation result. Based on the image evaluation results, determine whether the multimodal image is qualified. If not, optimize the multimodal image; if yes, output and save the image evaluation results. The step of performing multi-dimensional feature extraction based on modal collaboration on the multimodal image to obtain multi-dimensional features includes: The multimodal image is segmented to obtain high-information regions and background regions; Calculate the gradient variance of the high-information region and the edge intensity of the background region, respectively. A dynamic weighted regional importance coefficient is introduced, and the coefficient value is adjusted in conjunction with the evaluation task type. Based on the gradient variance and the edge strength, an adaptive sharpness evaluation of multi-scale regions is performed to obtain comprehensive sharpness features. The multimodal image is subjected to dynamic distribution entropy quantization processing to obtain global illumination entropy, regional illumination deviation, and illumination contrast entropy of high information area; The dynamic distribution entropy of illumination is fused based on the global illumination entropy, the regional illumination deviation, and the illumination contrast entropy of the high-information area to obtain the dynamic distribution entropy feature of illumination. Modal consistency verification is performed on the high-information regions of the multimodal image, and after the verification is completed, dynamic complementary weight allocation between modalities is performed based on modal characteristics to obtain modal consistency features; Considering the correlation between modes, cross-modal noise suppression is performed on the multimodal image to obtain noise intensity features.

2. The image quality assessment and optimization method according to claim 1, characterized in that, The multi-dimensional features include multiple evaluation features of different dimensions; the process of constructing a dynamic correlation matrix based on the multi-dimensional features and the evaluation task type, and dynamically optimizing the evaluation weights of the dynamic correlation matrix in real time to obtain the image evaluation result includes: A dynamic correlation matrix is ​​constructed based on the assessment task type and each assessment feature. Each row of the dynamic correlation matrix corresponds to a different assessment feature, each column corresponds to an assessment task type, and each matrix element represents the influence weight of the assessment feature at the element's position on the assessment task type. Each matrix element corresponds to an initial weight. The multiple assessment features of different dimensions include comprehensive sharpness features, dynamic light distribution entropy features, modal consistency features, and noise intensity features. Based on the evaluation task type and each of the evaluation features, and combined with historical evaluation accuracy, a pre-trained deep reinforcement learning model is used to dynamically optimize each of the initial weights in real time, so as to obtain the optimized influence weight of each of the evaluation features on the evaluation task type. Calculate the quality assessment prediction score based on each of the optimized influence weights. By integrating the quality assessment prediction scores and the optimized influence weights, the image evaluation results of the multimodal image are obtained.

3. The image quality assessment and optimization method according to claim 2, characterized in that, The initial weights are obtained by training a weight quantization model; the weight quantization model is used to quantify the influence weights of different evaluation features on various tasks; the method further includes: After completing the image quality assessment, the multimodal images and the image assessment results are integrated into training images and stored in the model training set; When the model training set successfully accumulates a preset number of new training images, an online update mechanism is automatically triggered to retrain the weight quantization model based on the accumulated new training images, or all training images, and to update the initial weights according to the retraining results.

4. The image quality assessment and optimization method according to claim 1, characterized in that, The multi-dimensional features include multiple evaluation features of different dimensions; the image evaluation result includes the quality evaluation prediction score of the multimodal image, and the optimized influence weight of each evaluation feature on the evaluation task type. The step of determining whether the multimodal image is qualified based on the image evaluation result, and optimizing the multimodal image if not, includes: Determine whether the predicted score of the quality assessment is greater than or equal to a preset target threshold; If not, the multimodal image is deemed unqualified; When the quality assessment prediction score is greater than or equal to a preset mild threshold and less than a preset target threshold, a lightweight optimization algorithm is activated to perform secondary optimization on the multimodal image. When the quality assessment prediction score is less than a preset severe threshold, the multimodal image is localized based on the assessment features of multiple different dimensions to obtain at least one defect region; the preset severe threshold is less than the preset mild threshold. For each defect region, a defect severity index is calculated based on the region size and the optimized impact weight corresponding to the defect region, and the severity priority of the defect region is determined according to the defect severity index. Defect regions with severity priorities in the highest priority range are identified as target defect regions, and operational optimization suggestions are generated for the target defect regions. The operation optimization suggestions are fed back to the user's image acquisition device in an AR visualization manner to guide the user to optimize the image acquisition process based on the operation optimization suggestions.

5. The image quality assessment and optimization method according to claim 1, characterized in that, Also includes: When the image acquisition device held by the user is detected to have triggered the first acquisition action, a low-resolution image is acquired. Based on the low-resolution image, perform image pre-evaluation to obtain a pre-evaluation score; If the pre-evaluation score is less than the preset evaluation threshold, a collection failure prompt message and basic adjustment suggestions are sent to the collection interface of the image acquisition device. If the pre-evaluation score is greater than or equal to the preset evaluation threshold, the acquisition parameters of the image acquisition device are dynamically adjusted based on the pre-evaluation result, and the image acquisition process is started after the parameter adjustment to obtain the multimodal image to be evaluated.

6. The image quality assessment and optimization method according to any one of claims 1 to 5, characterized in that, After outputting and saving the image evaluation results, the method further includes: Determine whether the threshold back-calculation mechanism has been triggered; If not, then the subsequent image recognition process is directly performed based on the multimodal image; If so, the target task performance value input by the user is obtained, and the target task performance value is inversely mapped and calculated using a pre-built quality-performance mapping prediction model to obtain the minimum quality assessment score required for the target task performance value. Extract the quality assessment prediction score of the multimodal image from the image assessment results, and compare the quality assessment prediction score with the minimum quality assessment score; If the predicted quality assessment score is greater than or equal to the minimum quality assessment score, then the subsequent image recognition process is executed based on the multimodal image; If the predicted quality assessment score is less than the minimum quality assessment score, the subsequent image recognition process for the multimodal image will not be executed.

7. An image quality assessment and optimization device, characterized in that, include: The data acquisition unit is used to acquire the evaluation task type and the multimodal images to be evaluated; The feature extraction unit is used to perform multi-dimensional feature extraction based on modal collaboration on the multimodal image to obtain multi-dimensional features; The real-time dynamic optimization unit is used to construct a dynamic correlation matrix based on the multi-dimensional features and the evaluation task type, and to perform real-time dynamic optimization of the evaluation weights of the dynamic correlation matrix to obtain the image evaluation result. An image quality judgment unit is used to determine whether the multimodal image is qualified based on the image evaluation result; if not, the multimodal image is optimized. If so, output and save the image evaluation result; The feature extraction unit includes: The image segmentation unit is used to segment the multimodal image to obtain high-information regions and background regions; A region information calculation unit is used to calculate the gradient variance of the high-information region and the edge intensity of the background region, respectively. An adaptive sharpness evaluation unit is used to introduce a dynamic weighted regional importance coefficient, and adjust the coefficient value in combination with the evaluation task type. Based on the gradient variance and the edge intensity, an adaptive sharpness evaluation of multi-scale regions is performed to obtain comprehensive sharpness features. The illumination dynamic distribution entropy quantization processing unit is used to perform illumination dynamic distribution entropy quantization processing on the multimodal image to obtain global illumination entropy, regional illumination deviation, and illumination contrast entropy of high information area. The dynamic distribution entropy fusion unit is used to fuse the dynamic distribution entropy of illumination based on the global illumination entropy, the regional illumination deviation and the illumination contrast entropy of the high information area to obtain the dynamic distribution entropy features of illumination. The intermodal consistency verification unit is used to perform intermodal consistency verification on the high-information regions of the multimodal image, and after the verification is completed, to perform dynamic complementary weight allocation between modalities based on modal characteristics to obtain modal consistency features. A cross-modal noise suppression unit is used to consider the correlation between modes and perform cross-modal noise suppression on the multimodal image to obtain noise intensity features.

8. An electronic device, characterized in that, The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the image quality assessment and optimization method according to any one of claims 1-6 according to the instructions in the program code.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code for executing the image quality assessment and optimization method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Image quality evaluation method, device and system

    CN113344843A

  • Non-contact multi-modal fusion biological recognition system, method and device

    CN119007252A