Quality marking method and system for machine vision system, medium and equipment
By determining the reference/distorted image pairs in the machine vision system, using machine subjects to perform a variety of downstream tasks, divide and integrate scores, the problem of lack of image quality evaluation in the machine vision system in the prior art is solved, a machine preference database is established, data consistency and reliability are ensured, and the development of image quality evaluation technology is promoted.
Patent Information
- Application Number
- CN202411980771.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-09-02
AI Technical Summary
It is difficult for the prior art to establish a database that can correspond images to scores one by one, representing the overall preference of machines for images, resulting in the inability to obtain objective and comprehensive end-to-end evaluation of image processing algorithms for machine vision.
By determining the reference/distorted image pair, the machine subjects were used to perform regression, classification, visual question and answer, labeling, segmentation, detection and retrieval tasks under preset temperature parameters, to determine the machine perceived quality scores corresponding to each task, and to perform dimension divisions, and finally determine the average opinion score as the quality label.
A quality labeling method for machine vision systems has been established to ensure data consistency and reliability, drive the future development of image quality evaluation technology, and build a large-scale machine preference database.
Smart Images

Figure CN120583218A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of machine vision evaluation, and in particular to a quality labeling method, system, medium, and device for a machine vision system. Background Art
[0002] Over the past decade, the rapid rise of smart cities and the booming Internet of Things (IoT) have fundamentally transformed the architecture of applications. According to research, the number of machine-to-machine (M2M) connections will surpass machine-to-human (M2H) connections for the first time in 2023, reaching 14.7 billion. Machines have gradually replaced humans as the primary consumer of image and video data. On the application side, in the field of image processing, the primary goal is to improve the quality of processed images, and this quality should be consistent with perception. In the past, the performance of a compression, super-resolution, reconstruction, or generation algorithm was often determined by human perception. This technology, known as Image Quality Assessment (IQA), aims to comprehensively model the human visual system (HVS) and thus predict human subjective preferences for images. Therefore, with the evolution of applications, image quality assessment (IQA) needs to move beyond the human visual system (HVS) and further serve the machine visual system (MVS).
[0003] Unfortunately, there are significant differences in the perception mechanisms of human and machine vision systems. The human eye focuses on similarity between reference and distorted images, such as texture, structure, and color; whereas machines focus on consistency in results across downstream tasks. Distortion caused by factors like encoding and quantization can cause noticeable degradation in quality for the human eye, while machine vision tasks like segmentation and detection remain largely unaffected. Conversely, subtle perturbations imperceptible to the human eye can significantly disrupt machine output. This makes it difficult to universally apply image quality assessment algorithms for both human and machine vision systems.
[0004] What kind of images do machines prefer? This remains an open question. Image quality assessment for human visual systems has made significant progress in the past 20 years. Hundreds of fine-grained databases are capable of driving image quality assessment algorithms that are highly consistent with human perception. However, the development of machine vision systems has been much slower. Not to mention end-to-end machine vision system algorithms, there is currently no database that can map images to scores and characterize the machine's overall image preference. This means that image processing algorithms for machine vision can only be verified through the effectiveness of specific tasks and two or three models, making objective and comprehensive end-to-end evaluation impossible. Summary of the Invention
[0005] In view of the defects in the prior art, the purpose of the present disclosure is to provide a quality labeling method, system, medium and equipment for machine vision systems.
[0006] To achieve the above objectives, according to one aspect of the present disclosure, a quality labeling method for a machine vision system is provided, comprising:
[0007] determining a reference / distorted image pair;
[0008] Under preset temperature parameters, using a machine subject to perform a preset downstream task on the reference / distorted image pair to determine a machine perception quality score corresponding to each of the preset downstream tasks, the preset downstream tasks including a regression task, a classification task, a visual question answering task, a labeling task, a segmentation task, a detection task, and a retrieval task;
[0009] Dividing the machine perception quality score corresponding to each of the preset downstream tasks into dimensions to determine a score corresponding to each dimension;
[0010] According to the scores corresponding to each dimension, an average opinion score is determined, and a quality label is determined. The average opinion score represents a comprehensive representation of the machine subject's preference for the same distorted image.
[0011] Optionally, the step of using a machine subject to perform a preset downstream task on the reference / distorted image pair under preset temperature parameters and determining a machine perception quality score corresponding to each preset downstream task includes:
[0012] Using the machine subject to perform the regression task on the reference / distorted image pair to determine a machine perception quality score corresponding to the regression task, wherein the regression task is in the form of a true / false question;
[0013] Using the machine subject to perform the classification task on the reference / distorted image pair to determine a machine perception quality score corresponding to the classification task, wherein the classification task is in the form of a multiple-choice question;
[0014] performing the visual question answering task on the reference / distorted image pair using the machine subject to determine a machine perception quality score corresponding to the visual question answering task;
[0015] performing the labeling task on the reference / distorted image pair using the machine subject to determine a machine perception quality score corresponding to the labeling task;
[0016] performing the segmentation task on the reference / distorted image pair using the machine subject, and determining a machine perception quality score corresponding to the segmentation task;
[0017] performing the detection task on the reference / distorted image pair using the machine subject to determine a machine perception quality score corresponding to the detection task;
[0018] The machine subject is employed to perform the retrieval task on the reference / distorted image pair, and a machine perception quality score corresponding to the retrieval task is determined.
[0019] Optionally, dividing the machine perception quality score corresponding to each of the preset downstream tasks into dimensions and determining the score corresponding to each dimension includes:
[0020] Dividing the regression task into a first dimension, using a general multimodal large model as the machine subject, and determining a score corresponding to the first dimension based on an average evaluation result of the general multimodal large model performing the regression task;
[0021] Dividing the classification task into a second dimension, using the universal multimodal large model as the machine subject, and determining a score corresponding to the second dimension based on an average evaluation result of the universal multimodal large model performing the classification task;
[0022] Dividing the visual question answering task into a third dimension, using the universal multimodal large model as the machine subject, and determining a score corresponding to the third dimension based on an average evaluation result of the universal multimodal large model performing the visual question answering task;
[0023] Dividing the labeling task into a fourth dimension, using the universal multimodal large model as the machine subject, and determining a score corresponding to the fourth dimension based on an average evaluation result of the universal multimodal large model performing the labeling task;
[0024] The segmentation task, the detection task, and the retrieval task are uniformly divided into a fifth dimension, and a dedicated computer vision model is used as the machine subject. The average evaluation results of the segmentation task, the detection task, and the retrieval task are respectively performed according to the dedicated computer vision model to determine the score corresponding to the fifth dimension.
[0025] Optionally, the step of uniformly dividing the segmentation task, the detection task, and the retrieval task into a fifth dimension, using a dedicated computer vision model as the machine subject, and performing the segmentation task, the detection task, and the retrieval task respectively based on the average evaluation results of the dedicated computer vision model to determine the score corresponding to the fifth dimension includes:
[0026] Using the dedicated computer vision model to perform the segmentation task, and determining an average evaluation result of the dedicated computer vision model performing the segmentation task;
[0027] Using the dedicated computer vision model to perform the detection task, and determining an average evaluation result of the dedicated computer vision model performing the detection task;
[0028] Using the dedicated computer vision model to perform the retrieval task, and determining an average evaluation result of the dedicated computer vision model performing the retrieval task;
[0029] The average evaluation result of the segmentation task, the average evaluation result of the detection task, and the average evaluation result of the retrieval task are weighted averaged according to preset weights to determine the score corresponding to the fifth dimension.
[0030] Optionally, the method further includes:
[0031] The scores corresponding to each dimension are regularized to determine the scores corresponding to each dimension after regularization.
[0032] Optionally, determining a mean opinion score and a quality label based on the scores corresponding to each dimension includes:
[0033] summing the regularized score corresponding to the first dimension, the regularized score corresponding to the second dimension, the regularized score corresponding to the third dimension, the regularized score corresponding to the fourth dimension, and the regularized score corresponding to the fifth dimension to determine the mean opinion score;
[0034] The mean opinion score is determined as the quality label.
[0035] According to a second aspect of the present disclosure, a quality labeling system for a machine vision system is provided, comprising:
[0036] an image determination module for determining a reference / distorted image pair;
[0037] a task execution module, configured to perform preset downstream tasks on the reference / distorted image pair using a machine subject under preset temperature parameters, and determine a machine perception quality score corresponding to each of the preset downstream tasks, wherein the preset downstream tasks include regression tasks, classification tasks, visual question answering tasks, labeling tasks, segmentation tasks, detection tasks, and retrieval tasks;
[0038] A dimension division module, configured to divide the machine perception quality score corresponding to each of the preset downstream tasks into dimensions and determine a score corresponding to each dimension;
[0039] The mean opinion score determination module is used to determine the mean opinion score and the quality label according to the score corresponding to each dimension, wherein the mean opinion score represents a comprehensive representation of the preference of the machine subject for the same distorted image.
[0040] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the method provided in the first aspect of the present disclosure are implemented.
[0041] According to a fourth aspect of the present disclosure, there is provided an electronic device, including:
[0042] a memory having a computer program stored thereon;
[0043] A processor is used to execute the computer program in the memory to implement the steps of the method provided in the first aspect of the present disclosure.
[0044] Compared with the prior art, the embodiments of the present disclosure have at least one of the following beneficial effects:
[0045] Through the above technical solution, under preset temperature parameters, machine subjects are used to perform preset downstream tasks, including regression tasks, classification tasks, visual question-answering tasks, labeling tasks, segmentation tasks, detection tasks and retrieval tasks, and the machine perception quality scores corresponding to each preset downstream task are divided into dimensions. The scores of each dimension are then combined to determine the average opinion score, which represents a comprehensive representation of the machine subjects' preferences for the same distorted image. Based on the average opinion score, as a quality label for machine vision systems, it is conducive to establishing future machine-oriented image quality evaluation datasets, ensuring data consistency and reliability, and establishing a large-scale machine preference database, which effectively drives the development of future machine-oriented image quality evaluation technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Other features, objects and advantages of the present disclosure will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0047] Figure 1 The present invention is a flowchart of a quality labeling method for a machine vision system according to an exemplary embodiment.
[0048] Figure 2 It is a schematic diagram showing a machine subject performing a preset downstream task according to an exemplary embodiment.
[0049] Figure 3 The figure shows a relationship between a mean opinion score and a machine perception quality score corresponding to each preset downstream task, and a score distribution diagram of each preset downstream task, according to an exemplary embodiment.
[0050] Figure 4 The present invention is a block diagram of a quality labeling system for a machine vision system according to an exemplary embodiment. DETAILED DESCRIPTION
[0051] The present disclosure is described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art further understand the present disclosure, but are not intended to limit the present disclosure in any way. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the scope of the present disclosure. These modifications and improvements are all within the scope of protection of the present disclosure.
[0052] The present disclosure provides a quality labeling method for machine vision systems, which is used to extend the concept of image quality assessment (IQA) from the human visual system (HVS) to the machine visual system (MVS) and to standardize and mathematically define machine preferences.
[0053] Figure 1 The present invention is a flowchart of a quality labeling method for a machine vision system according to an exemplary embodiment.
[0054] like Figure 1 As shown, the present disclosure provides a quality labeling method for a machine vision system, including S11 to S14.
[0055] S11, determining a reference / distorted image pair.
[0056] S12, using a machine subject to perform a preset downstream task on a reference / distorted image pair under preset temperature parameters, and determining a machine perception quality score corresponding to each preset downstream task.
[0057] In the present disclosure, the preset temperature parameter is 0, and the preset temperature parameter can be adaptively adjusted according to specific needs.
[0058] The pre-set downstream tasks include regression, classification, visual question answering, labeling, segmentation, detection, and retrieval. Regression, classification, visual question answering, and labeling are multimodal tasks associated with large multimodal models (LMMs), while segmentation, detection, and retrieval are classic computer vision (CV) tasks.
[0059] When the machine subject performs the above seven downstream tasks, it uses different rules to calculate the similarity of the execution results of the reference / distorted images, and fuses them into a machine preference score, which is an indicator to measure the machine preference. The machine preference score is used as the machine perception quality score for each preset downstream task.
[0060] S13, dividing the machine perception quality score corresponding to each preset downstream task into dimensions, and determining the score corresponding to each dimension.
[0061] S14: Determine the average opinion score based on the score corresponding to each dimension.
[0062] Among them, the Mean Opinion Score (MOS) represents the comprehensive representation of the machine subjects' preferences for the same distorted image.
[0063] Through the above technical solution, under preset temperature parameters, machine subjects are used to perform preset downstream tasks, including regression tasks, classification tasks, visual question-answering tasks, labeling tasks, segmentation tasks, detection tasks and retrieval tasks, and the machine perception quality scores corresponding to each preset downstream task are divided into dimensions. The scores of each dimension are then combined to determine the average opinion score, which represents a comprehensive representation of the machine subjects' preferences for the same distorted image. Based on the average opinion score, as a quality label for machine vision systems, it is conducive to establishing future machine-oriented image quality evaluation datasets, ensuring data consistency and reliability, and establishing a large-scale machine preference database, which effectively drives the development of future machine-oriented image quality evaluation technologies.
[0064] In a possible embodiment, S11, determining a reference / distorted image pair, may include:
[0065] A preset number of distortion mechanisms and a preset type of distortion strength are defined, and interference is added to a reference image to obtain a distorted image, thereby determining a reference / distorted image pair.
[0066] As an example, the present disclosure defines 30 distortion mechanisms and 5 distortion intensities to add interference to a high-quality reference image to obtain a reference / distorted image pair.
[0067] Image processing algorithms in machine vision primarily focus on seven tasks: regression, classification, visual question answering (VQA), captioning (CAP), segmentation (SEG), detection (DET), and retrieval (RET). The first four tasks—regression, classification, VQA, and captioning—can be accomplished using a general-purpose large multimodal model (LMM). Segmentation, detection, and retrieval require specialized models for their respective fields, specifically computer vision models.
[0068] Moreover, from the perspective of the large model LMM, the above-mentioned regression task corresponds to a judgment question (Yes-or-No, YoN), and the classification task corresponds to a multiple-choice question (MCQ) from the perspective of the large model LMM. Experts can be hired to design challenging questions and answers for the reference image, such as the judgment question YoN, the visual question answering task VQA, and the annotation task CAP.
[0069] This disclosure defines the "machine preference score" as the fidelity of the downstream task. Machine subjects are used to perform the above seven downstream tasks on the reference image and the distorted image respectively, and the consistency of the two task execution results is measured.
[0070] Figure 2 It is a schematic diagram showing a machine subject performing a preset downstream task according to an exemplary embodiment.
[0071] As shown in Figure 2, in a possible embodiment, S12, under preset temperature parameters, a machine subject is used to perform a preset downstream task on a reference / distorted image pair to determine a machine perception quality score corresponding to each preset downstream task, which may include S21 to S27.
[0072] S21, use machine subjects to perform a regression task on the reference / distorted image pair to determine the machine perception quality score corresponding to the regression task.
[0073] The regression task is in the form of a judgment question. In the judgment question, a confusion option is introduced to accurately measure the degree of image quality degradation perceived by the machine subject, that is, the score S YoN .
[0074] The machine perception quality score corresponding to the regression task is expressed as:
[0075] S YoN =|σ(P dis (Yes,No))-σ(P ref (Yes,No))|
[0076] Among them, S YoN represents the machine perception quality score corresponding to the regression task, dis represents the distorted image, ref represents the reference image, P dis (Yes, No) represents the probability that the machine subject outputs “YES” or “No” when viewing the distorted image, P ref (Yes, No) represents the probability that the machine subject outputs “YES” or “No” when viewing the reference image pair, and σ(·) represents the softmax function.
[0077] It is normalized by the softmax function σ(·).
[0078] S22, using machine subjects to perform a classification task on the reference / distorted image pairs, and determining a machine perception quality score corresponding to the classification task.
[0079] The classification task is a multiple-choice question. It consists of one correct answer and three confusing answers (usually no more than four). The probability of each answer is output as a four-element vector. The distance between the two four-element vectors derived from the reference / distorted image is used as the machine perception quality score for the classification task. The distance between the two four-element vectors in space can be measured using cosine similarity.
[0080] The machine perception quality score corresponding to the classification task is:
[0081] S MCQ =cos(P dis (A,B,C,D)-P ref (A,B,C,D))
[0082] Among them, S MCQ Represents the machine perception quality score corresponding to the classification task, P dis(A, B, C, D) represents the probability that the machine subject outputs the four options "A, B, C, D" when viewing the distorted image, P ref (A, B, C, D) represents the probability that the machine subject outputs the four options “A, B, C, D” when viewing the reference image, and cos(·) represents the cosine similarity.
[0083] S23, using machine subjects to perform a visual question answering task on the reference / distorted image pairs to determine a machine perception quality score corresponding to the visual question answering task.
[0084] Among them, for visual question answering tasks, we can go beyond the character level and perform sentence level evaluation. The large model LMM gives answers within five English words for the preset questions. The semantic difference between the two answers to the reference / distorted image pair is used as the machine perception quality score S corresponding to the visual question answering task. VQA .
[0085] The machine perception quality score corresponding to the visual question answering task is:
[0086] S VQA = CLIP(T dis ,T ref )
[0087] Among them, S VQA represents the machine perception quality score corresponding to the visual question answering task, T dis represents the specific text content output by the machine subject when viewing the distorted image, T ref represents the specific text content output by the machine subject when viewing the reference image, and CLIP(·) represents the text encoder.
[0088] S24, using a machine subject to perform a labeling task on the reference / distorted image pair to determine a machine perception quality score corresponding to the labeling task.
[0089] Among them, for the labeling task, we can go beyond the character level and perform sentence level evaluation. The large model LMM outputs an answer of about 40 words to the preset question, describing the main theme, action, scene and any significant objects or features of the image, and uses the index eval∈(BLEU, CIDEr, SPICE) specifically for image description to analyze the difference between the two long sentences of the two answers of the reference / distorted image pair as the machine perception quality score S corresponding to the labeling task. CAP .
[0090] The machine perception quality score corresponding to the labeling task is expressed as:
[0091]
[0092] Among them, SCAP represents the machine perception quality score corresponding to the labeling task, T dis represents the specific text content output by the machine subject when viewing the distorted image, T ref represents the specific text content output by the machine subject when viewing the reference image, and eval(·) represents the indicator used for image description.
[0093] S25, using a machine subject to perform a segmentation task on the reference / distorted image pair, and determining a machine perception quality score corresponding to the segmentation task.
[0094] Among them, for the segmentation task, the most common evaluation criterion is adopted, namely Intersection-over-Union (IoU) as the machine perception quality score S corresponding to the segmentation task SEG .
[0095] The machine perception quality score corresponding to the segmentation task is:
[0096] S SEG =IoU(M dis ,M ref )
[0097] Among them, S SEG represents the machine perception quality score corresponding to the segmentation task, M dis represents the mask region output by the machine subject when viewing the distorted image, M ref represents the mask region output by the machine subject when viewing the reference image, and IoU(·) represents the intersection over union.
[0098] S26, using a machine subject to perform a detection task on the reference / distorted image pair to determine a machine perception quality score corresponding to the detection task.
[0099] Among them, for the detection task, the detection result includes spatial information and category labels. When the categories are different, the detection is considered to have failed. When the categories are consistent, the intersection-over-union ratio is further calculated as the machine perception quality score S corresponding to the detection task. DET .
[0100] When the detected categories are consistent, the machine perception quality score corresponding to the detection task is:
[0101] S DET =Acc1(C dis ,C ref )·IoU(B dis ,B ref )
[0102] Among them, S RET represents the machine perception quality score corresponding to the retrieval task, Cdis represents the category label output by the machine subject when viewing the distorted image, C ref represents the category label output by the machine subject when viewing the reference image, and Acc1(·) represents the Top-1 accuracy in the classification task.
[0103] S27, using machine subjects to perform a retrieval task on the reference / distorted image pairs, and determining a machine perception quality score corresponding to the retrieval task.
[0104] Among them, for the retrieval task, the three common image-to-text indicators, namely Top-1 / 5 / 10 accuracy, are combined to evaluate the retrieval sequence as the machine perception quality score S corresponding to the retrieval task RET .
[0105] The machine-perceived quality score corresponding to the retrieval task is expressed as:
[0106]
[0107] Among them, S DET represents the machine perception quality score corresponding to the detection task, C dis represents the category label output by the machine subject when viewing the distorted image, C ref represents the category label output by the machine subject when viewing the reference image, B dis represents the bounding box area output by the machine subject when viewing the distorted image pair, represents the bounding box area output by the machine subject when viewing the reference image, Acc i (·) represents the Top-i accuracy in the classification task, and IoU(·) represents the intersection-over-union ratio.
[0108] Through steps S21 to S27, the preferences of the large model LMM in the four dimensions of regression task, classification task, visual question answering task and labeling task are obtained, as well as the preferences of each computer vision model in segmentation task, detection task and retrieval task.
[0109] The same settings were used to obtain the mean opinion score (MOS score) of the machine vision system according to the International Telecommunication Union (ITU) standard for human subjective annotation.
[0110] Setting the temperature parameter to 0 prevents unstable output and ensures that all perceived quality degradation comes from distortion.
[0111] When the machine vision system performs the regression task, classification task, visual question answering task, and labeling task in steps S21 to S24 above, 15 common large models LMMs are used for inference, namely DeepseekVL, InstructBLIP, InternVL2, InternLM-XComposer2, LLaVA1.5, LLaVANext, LLaVA-OneVision, Mantis, Mini-InternVL, MPlugOwl3, Ovis, Phi3.5, Panega, Qwen2-VL, and Yi1.5.
[0112] When the machine vision system performs the segmentation, detection, and retrieval tasks in steps S25 to S27 above, five dedicated models are used in each computer vision task, taking into account popular trends and performance. The segmentation task uses: BoxInst, ConvNext, Mask-RCNN, SCNet, and RTMDet (segmentation mode); the detection task uses Dino, Mask-RCNN, RTMDet (detection mode), ViTDet, and Yolo-X; and the retrieval task uses DIME, InternVL (retrieval-based fine-tuning), LAVIS (retrieval-based fine-tuning), NAAF, and VSE++.
[0113] In a possible embodiment, S13, dividing the machine perception quality score corresponding to each preset downstream task into dimensions and determining the score corresponding to each dimension, includes S31 to S35.
[0114] S31, divide the regression task into a first dimension, take the general multimodal large model as the machine subject, and determine the score corresponding to the first dimension based on the average evaluation result of the general multimodal large model performing the regression task.
[0115] Among them, the average evaluation result of the general multimodal large model performing the regression task is the average score of the 15 general large model LMMs performing the regression task in step S21. The 15 general large model LMMs are DeepseekVL, InstructBLIP, InternVL2, InternLM-XComposer2, LLaVA1.5, LLaVANext, LLaVA-OneVision, Mantis, Mini-InternVL, MPlugOwl3, Ovis, Phi3.5, Panega, Qwen2-VL and Yi1.5.
[0116] S32, dividing the classification task into a second dimension, taking the general multimodal large model as the machine subject, and determining the score corresponding to the second dimension based on the average evaluation result of the general multimodal large model performing the classification task.
[0117] The average evaluation result of the general multimodal large model performing the classification task is the average score of the 15 general large models LMMs performing the classification task in step S22.
[0118] S33, divide the visual question answering task into the third dimension, take the general multimodal large model as the machine subject, and determine the score corresponding to the third dimension based on the average evaluation results of the general multimodal large model performing the visual question answering task.
[0119] The average evaluation result of the general multimodal large model performing the visual question answering task is the average score of the 15 general large models LMMs performing the visual question answering task in step S23.
[0120] S34, divide the labeling task into a fourth dimension, use the general multimodal large model as the machine subject, and determine the score corresponding to the fourth dimension based on the average evaluation result of the general multimodal large model performing the labeling task.
[0121] The average evaluation result of the general multimodal large model performing the labeling task is the average score of the labeling task performed by 15 general large models LMMs in step S24.
[0122] S35, the segmentation task, detection task, and retrieval task are uniformly divided into the fifth dimension, and a dedicated computer vision model is used as a machine subject. The average evaluation results of the segmentation task, detection task, and retrieval task performed by the dedicated computer vision model are used to determine the score corresponding to the fifth dimension.
[0123] Among them, the average evaluation results of the dedicated computer vision models performing the segmentation task, detection task, and retrieval task respectively are expressed as the average scores of using 5 dedicated models to perform the segmentation task, detection task, and retrieval task in steps S25 to S27, respectively. The 5 dedicated models are BoxInst, ConvNext, Mask-RCNN, SCNet, and RTMDet (segmentation mode).
[0124] In the above steps S31 to S35 , 15 machine subjects are ensured in each dimension.
[0125] In a possible embodiment, S35, the segmentation task, detection task, and retrieval task are uniformly divided into a fifth dimension, a dedicated computer vision model is used as a machine subject, and the average evaluation results of the segmentation task, detection task, and retrieval task are respectively performed according to the dedicated computer vision model to determine the score corresponding to the fifth dimension, which may include S351 to S354.
[0126] S351: Use a dedicated computer vision model to perform the segmentation task, and determine an average evaluation result of the dedicated computer vision model performing the segmentation task.
[0127] S352: Use a dedicated computer vision model to perform the detection task, and determine an average evaluation result of the dedicated computer vision model performing the detection task.
[0128] S353: Use a dedicated computer vision model to perform the retrieval task, and determine an average evaluation result of the dedicated computer vision model performing the retrieval task.
[0129] S354: Perform weighted averaging on the average evaluation result of the segmentation task, the average evaluation result of the detection task, and the average evaluation result of the retrieval task according to preset weights to determine the score corresponding to the fifth dimension.
[0130] Among them, in the present disclosure, the preset weights use the same weight, that is, the average evaluation result of the segmentation task, the average evaluation result of the detection task, and the average evaluation result of the retrieval task use the same weight.
[0131] In a possible embodiment, a quality marking system for a machine vision system may further include S16.
[0132] S16, regularizing the score corresponding to each dimension to determine the score corresponding to each dimension after regularization.
[0133] Step S16 of the present disclosure is performed between step S13 and step S14 .
[0134] By regularizing the scores corresponding to each dimension, we can prevent the situation where the score of one dimension is too high and the score of another dimension is too low. For example, the visual question answering task is relatively simple and its score is generally high, while the labeling task score is low. Regularization is used to ensure that the overall weight of the mean opinion score (MOS) is fair and reliable.
[0135] In a possible embodiment, S14, determining a mean opinion score and a quality label according to the score corresponding to each dimension, may include S41 to S42.
[0136] S41, summing up the scores corresponding to the regularized first dimension, the scores corresponding to the regularized second dimension, the scores corresponding to the regularized third dimension, the scores corresponding to the regularized fourth dimension, and the scores corresponding to the regularized fifth dimension to determine the average opinion score.
[0137] Among them, the score corresponding to the regularized first dimension, the score corresponding to the regularized second dimension, the score corresponding to the regularized third dimension, the score corresponding to the regularized fourth dimension, and the score corresponding to the regularized fifth dimension are all in the range of (0, 1), and the range of the mean opinion score MOS is determined to be (0, 5).
[0138] S42, determining the mean opinion score as a quality label.
[0139] By providing the mean opinion score (MOS), the machine preference score, i.e., the standardized data collection paradigm of machine MOS, it facilitates the construction of image quality assessment (IQA) data for machine vision in the future.
[0140] Figure 3 The figure shows a relationship between a mean opinion score and a machine perception quality score corresponding to each preset downstream task, and a score distribution diagram of each preset downstream task, according to an exemplary embodiment.
[0141] in Figure 3 (a) is the relationship between the mean opinion score and the machine perception quality score corresponding to each preset downstream task, Figure 3 (b) The score distribution of each pre-defined downstream task.
[0142] like Figure 3 (a) and Figure 3 As shown in (b), the present disclosure adopts a quality labeling method for machine vision systems to obtain an average opinion score that has a strong correlation with the machine perception quality score corresponding to each task, so the average opinion score can be used as the same representation of perceptual quality.
[0143] To mimic the multi-dimensional scoring mechanism of humans, comprehensively consider the performance of machines on downstream tasks to obtain quality labels. A large-scale machine preference database (MPD) was established. Its preference scores are highly versatile, and each score can comprehensively represent preferences across different machines and tasks. Using 15 general-purpose large multi-modal models (LMMs) and 15 specialized computer vision (CV) models, 30,000 reference / distorted image pairs and 2.25 million mean opinion scores (meaning preference data) were collected, respectively. This established the first Image Quality Assessment (IQA) dataset for machine vision.
[0144] Figure 4 The present invention is a block diagram of a quality labeling system for a machine vision system according to an exemplary embodiment.
[0145] Based on the same concept, Figure 4 As shown, the present disclosure further provides a quality labeling system 100 for a machine vision system, comprising: an image determination module 110 , a task execution module 120 , a dimension division module 130 , and a mean opinion score determination module 140 .
[0146] An image determination module 110 for determining a reference / distorted image pair;
[0147] a task execution module 120 for performing, on the reference / distorted image pair, a preset downstream task using a machine subject under preset temperature parameters, and determining a machine perception quality score corresponding to each of the preset downstream tasks, wherein the preset downstream tasks include a regression task, a classification task, a visual question answering task, a labeling task, a segmentation task, a detection task, and a retrieval task;
[0148] A dimension division module 130 is configured to divide the machine perception quality score corresponding to each of the preset downstream tasks into dimensions and determine a score corresponding to each dimension;
[0149] The mean opinion score determination module 140 is configured to determine a mean opinion score based on the scores corresponding to each dimension, where the mean opinion score represents a comprehensive representation of the machine subject's preference for the same distorted image.
[0150] Through the above technical solution, under preset temperature parameters, machine subjects are used to perform preset downstream tasks, including regression tasks, classification tasks, visual question-answering tasks, labeling tasks, segmentation tasks, detection tasks and retrieval tasks, and the machine perception quality scores corresponding to each preset downstream task are divided into dimensions. The scores of each dimension are then combined to determine the average opinion score, which represents a comprehensive representation of the machine subjects' preferences for the same distorted image. Based on the average opinion score, as a quality label for machine vision systems, it is conducive to establishing future machine-oriented image quality evaluation datasets, ensuring data consistency and reliability, and establishing a large-scale machine preference database, which effectively drives the development of future machine-oriented image quality evaluation technologies.
[0151] Regarding the embodiment of the above system, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0152] Based on the same concept as above, in another embodiment of the present disclosure, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor is used to execute a quality labeling method for a machine vision system when executing the program.
[0153] Optionally, the memory is used to store programs; the memory may include volatile memory (English: volatile memory), such as random-access memory (English: random-access memory, abbreviated: RAM), such as static random-access memory (English: static random-access memory, abbreviated: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviated: DDR SDRAM), etc.; the memory may also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory is used to store computer programs (such as applications, functional modules, etc. that implement the above-mentioned methods), computer instructions, etc., and the above-mentioned computer programs, computer instructions, etc. can be partitioned and stored in one or more memories. In addition, the above-mentioned computer programs, computer instructions, data, etc. can be called by the processor.
[0154] The aforementioned computer programs, computer instructions, etc. may be partitioned and stored in one or more memories, and the aforementioned computer programs, computer instructions, data, etc. may be called by a processor.
[0155] The processor is configured to execute the computer program stored in the memory to implement the various steps of the method involved in the above embodiment. For details, please refer to the relevant description in the above method embodiment.
[0156] The processor and memory can be independent structures or integrated structures. When the processor and memory are independent structures, the memory and processor can be coupled via a bus.
[0157] In an embodiment of the present disclosure, a non-temporary computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of a quality labeling method for a machine vision system in any of the above embodiments are implemented.
[0158] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0159] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0160] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0162] Although the preferred embodiments of the present disclosure have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present disclosure.
[0163] Obviously, those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.
Claims
1. A quality labeling method for a machine vision system, characterized in that: include: determining a reference / distorted image pair; Under preset temperature parameters, using a machine subject to perform a preset downstream task on the reference / distorted image pair to determine a machine perception quality score corresponding to each of the preset downstream tasks, the preset downstream tasks including a regression task, a classification task, a visual question answering task, a labeling task, a segmentation task, a detection task, and a retrieval task; Dividing the machine perception quality score corresponding to each of the preset downstream tasks into dimensions to determine a score corresponding to each dimension; According to the scores corresponding to each dimension, an average opinion score is determined, and a quality label is determined. The average opinion score represents a comprehensive representation of the machine subject's preference for the same distorted image.
2. The method according to claim 1, characterized in that The method further comprises: performing a preset downstream task on the reference / distorted image pair using a machine subject under a preset temperature parameter, and determining a machine perception quality score corresponding to each preset downstream task. Using the machine subject to perform the regression task on the reference / distorted image pair to determine a machine perception quality score corresponding to the regression task, wherein the regression task is in the form of a true / false question; Using the machine subject to perform the classification task on the reference / distorted image pair to determine a machine perception quality score corresponding to the classification task, wherein the classification task is in the form of a multiple-choice question; performing the visual question answering task on the reference / distorted image pair using the machine subject to determine a machine perception quality score corresponding to the visual question answering task; performing the labeling task on the reference / distorted image pair using the machine subject to determine a machine perception quality score corresponding to the labeling task; performing the segmentation task on the reference / distorted image pair using the machine subject, and determining a machine perception quality score corresponding to the segmentation task; performing the detection task on the reference / distorted image pair using the machine subject to determine a machine perception quality score corresponding to the detection task; The machine subject is employed to perform the retrieval task on the reference / distorted image pair, and a machine perception quality score corresponding to the retrieval task is determined.
3. The method according to claim 2, characterized in that The machine perception quality score corresponding to the regression task is expressed as: S YoN =|σ(P dis (Yes,No))-σ(P ref (Yes,No))| Among them, S YoN represents the machine perception quality score corresponding to the regression task, dis represents the distorted image, ref represents the reference image, P dis (Yes, No) represents the probability that the machine subject outputs "YES" or "No" when viewing the distorted image, P ref (Yes, No) represents the probability that the machine subject outputs "YES" or "No" when viewing the reference image pair, and σ(·) represents the softmax function; The machine perception quality score corresponding to the classification task is expressed as: S MCQ =cos(P dis (A,B,C,D)-P ref (A,B,C,D)) Among them, S MCQ represents the machine perception quality score corresponding to the classification task, P dis (A, B, C, D) represents the probability that the machine subject outputs the four options "A, B, C, D" when viewing the distorted image, P ref (A, B, C, D) represents the probability that the machine subject outputs the four options "A, B, C, D" when viewing the reference image, and cos(·) represents the cosine similarity; The machine perception quality score corresponding to the visual question answering task is expressed as: S VQA =CLIP(T dis ,T ref ) Among them, S VQA represents the machine perception quality score corresponding to the visual question answering task, T dis represents the specific text content output by the machine subject when viewing the distorted image, T ref represents the specific text content output by the machine subject when viewing the reference image, and CLIP(·) represents a text encoder; The machine perception quality score corresponding to the labeling task is expressed as: Among them, S CAP represents the machine perception quality score corresponding to the labeling task, T dis represents the specific text content output by the machine subject when viewing the distorted image, T ref represents the specific text content output by the machine subject when viewing the reference image; eval(·) represents an indicator for image description; The machine perception quality score corresponding to the segmentation task is expressed as: S SEG =IoU(M dis ,M ref ) Among them, S SEG represents the machine perception quality score corresponding to the segmentation task, M dis represents the mask area output by the machine subject when viewing the distorted image, M ref represents the mask area output by the machine subject when viewing the reference image, and IoU(·) represents the intersection over union ratio; The machine perception quality score corresponding to the retrieval task is expressed as: Among them, S RET represents the machine perception quality score corresponding to the retrieval task, C dis represents the category label output by the machine subject when viewing the distorted image, C ref represents the category label output by the machine subject when viewing the reference image, Acc i (·) represents the Top-i accuracy in the classification task; The machine perception quality score corresponding to the detection task is expressed as: S DET =Acc1(C dis ,C ref )·IoU(B dis ,B ref ) Among them, S DET represents the machine perception quality score corresponding to the detection task, C dis represents the category label output by the machine subject when viewing the distorted image, C ref represents the category label output by the machine subject when viewing the reference image, B dis represents the bounding box area output by the machine subject when viewing the distorted image pair, B ref represents the bounding box area output by the machine subject when viewing the reference image, Acc1(·) represents the Top-1 accuracy in the classification task, and IoU(·) represents the intersection-over-union ratio.
4. The method according to claim 1, wherein Dividing the machine perception quality score corresponding to each of the preset downstream tasks into dimensions and determining the score corresponding to each dimension includes: Dividing the regression task into a first dimension, using a general multimodal large model as the machine subject, and determining a score corresponding to the first dimension based on an average evaluation result of the general multimodal large model performing the regression task; Dividing the classification task into a second dimension, using the universal multimodal large model as the machine subject, and determining a score corresponding to the second dimension based on an average evaluation result of the universal multimodal large model performing the classification task; Dividing the visual question answering task into a third dimension, using the universal multimodal large model as the machine subject, and determining a score corresponding to the third dimension based on an average evaluation result of the universal multimodal large model performing the visual question answering task; Dividing the labeling task into a fourth dimension, using the universal multimodal large model as the machine subject, and determining a score corresponding to the fourth dimension based on an average evaluation result of the universal multimodal large model performing the labeling task; The segmentation task, the detection task, and the retrieval task are uniformly divided into a fifth dimension, and a dedicated computer vision model is used as the machine subject. The average evaluation results of the segmentation task, the detection task, and the retrieval task are respectively performed according to the dedicated computer vision model to determine the score corresponding to the fifth dimension.
5. The method according to claim 4, characterized in that The step of uniformly dividing the segmentation task, the detection task, and the retrieval task into a fifth dimension, using a dedicated computer vision model as the machine subject, and performing the segmentation task, the detection task, and the retrieval task respectively based on the average evaluation results of the dedicated computer vision model to determine a score corresponding to the fifth dimension includes: Using the dedicated computer vision model to perform the segmentation task, and determining an average evaluation result of the dedicated computer vision model performing the segmentation task; Using the dedicated computer vision model to perform the detection task, and determining an average evaluation result of the dedicated computer vision model performing the detection task; Using the dedicated computer vision model to perform the retrieval task, and determining an average evaluation result of the dedicated computer vision model performing the retrieval task; The average evaluation result of the segmentation task, the average evaluation result of the detection task, and the average evaluation result of the retrieval task are weighted averaged according to preset weights to determine the score corresponding to the fifth dimension.
6. The method according to claim 4, characterized in that The method further comprises: The scores corresponding to each dimension are regularized to determine the scores corresponding to each dimension after regularization.
7. The method according to claim 6, characterized in that Determining the mean opinion score and the quality label based on the scores corresponding to each dimension includes: summing the regularized score corresponding to the first dimension, the regularized score corresponding to the second dimension, the regularized score corresponding to the third dimension, the regularized score corresponding to the fourth dimension, and the regularized score corresponding to the fifth dimension to determine the mean opinion score; The mean opinion score is determined as the quality label.
8. A quality labeling system for machine vision systems, characterized in that: include: an image determination module for determining a reference / distorted image pair; a task execution module, configured to perform preset downstream tasks on the reference / distorted image pair using a machine subject under preset temperature parameters, and determine a machine perception quality score corresponding to each of the preset downstream tasks, wherein the preset downstream tasks include regression tasks, classification tasks, visual question answering tasks, labeling tasks, segmentation tasks, detection tasks, and retrieval tasks; A dimension division module, configured to divide the machine perception quality score corresponding to each of the preset downstream tasks into dimensions and determine a score corresponding to each dimension; The mean opinion score determination module is used to determine the mean opinion score and the quality label according to the score corresponding to each dimension, wherein the mean opinion score represents a comprehensive representation of the preference of the machine subject for the same distorted image.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 7.