An image quality classification and processing method and system based on a multimodal large model
By identifying the target attributes of the inspection image and filtering the associated modal data, and combining information with the identifiable and unrecognizable ranges for pairing judgment, the problem of difficult to guarantee the matching degree, synchronization and consistency of modal data in multimodal large models is solved, and the quality of model input data and the accuracy of analysis results are improved.
Patent Information
- Application Number
- CN202510422151.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-07
AI Technical Summary
When the prior art uses a multimodal large model to fuse data from different modes for output, it cannot guarantee the matching, synchronization and consistency of the modal data, resulting in a decrease in the accuracy of model training and analysis results.
By identifying the target attributes of the inspection image, filtering the associated modal data sets, and type division. Then, based on the identifiable range and unrecognizable range of the inspection image, effective information and fuzzy information are extracted, and paired judgments are made to ensure the matching and consistency of the modal data.
The quality of the input image data of multimodal large model is improved, ensuring high consistency and accuracy of the data, thereby improving the analysis and judgment results of the model.
Smart Images

Figure CN119919746B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power inspection image processing. Specifically, it relates to an image quality classification processing method and system based on a multimodal large model. Background Art
[0002] A multimodal large model refers to a machine learning model that can process and understand multiple types of data or different forms of the same type of data. By using a multimodal large model, power inspection can achieve a high degree of automation and intelligence. Moreover, the combination of multimodal data helps to more accurately locate and classify problems, and can provide more accurate diagnostic results than a single modality, which is of great significance in the application of power inspection. In the process of using a multimodal large model to fuse different modalities of data for output, it depends on the correlation matching and synchronous consistency of these data. If the matching and synchronization cannot be guaranteed, it is possible that the multimodal large model deviates in the training and learning method, or the accuracy of the judgment and analysis results is greatly reduced. Therefore, how to ensure that various modality data have complementary correlation characteristics, time synchronization characteristics, space consistency characteristics, etc. when applying the multimodal large model has become the key point.
[0003] In the prior art, most multimodal large models are applied to image data enhancement. It uses image data combined with input data from different shooting sources or angles to improve and enrich image processing tasks, that is, uses information from other modalities to provide a deeper understanding and support, and fuses and analyzes the information from other modalities after standardizing it (normalizing, resizing, encoding conversion). In the whole process, for different modality information, only simple matching and alignment operations are performed relying on information such as timestamps and spatial coordinates, without further considering the matching degree between different modality information, how to more reliably classify and process the information with defects, and how to effectively annotate different modality information, resulting in low quality of the fused image synthesis data, with large noise, distortion or other forms of interference, affecting the accuracy of the final analysis output results.
[0004] In view of this, the present application is specifically proposed. Summary of the Invention
[0005] The purpose of the present invention is to provide an image quality classification processing method and system based on a multimodal large model. After extracting other modality data sets associated with the inspection images, the processing method and system perform matching judgment according to each associated modality data, and can perform pairing separately according to the recognizable range of defective inspection images, and perform associated annotation processing according to the final pairing results to ensure the quality of the synthesized data after fusing different modality information.
[0006] The embodiments of the present invention are implemented as follows:
[0007] First aspect, an image quality classification processing method based on a multimodal large model, comprising the following steps: identifying the target attributes of the inspection image, performing a preliminary analysis on the inspection image based on the target attributes, and determining the modal data matching strategy for the inspection image; screening the associated modal data set from the inspection database based on the modal data matching strategy, classifying the associated modal data set by type to obtain multiple types of associated modal data; performing a matching judgment on each type of the associated modal data and the inspection image to obtain the matching degree of this type of associated modal data, and taking the associated modal data with a matching degree higher than the preset matching value as the preliminarily screened modal data, wherein the matching judgment includes inspection time judgment, inspection object judgment, and inspection location judgment; delimiting the recognizable range and unrecognizable range of the inspection image, extracting effective information according to the recognizable range, extracting fuzzy information according to the unrecognizable range, pairing each type of the preliminarily screened modal data with the effective information and the fuzzy information respectively to obtain a first pairing result and a second pairing result; judging whether the preliminarily screened modal data is used as the fused modal data based on the first pairing result and the second pairing result, and performing an associated annotation process on the fused modal data and the inspection image, wherein the associated annotation process includes annotating the associated area, annotating the associated object, and annotating the associated form.
[0008] In some optional implementation manners, the delimiting the recognizable range and unrecognizable range of the inspection image, extracting effective information according to the recognizable range, and extracting fuzzy information according to the unrecognizable range includes the following steps: using a training model to analyze and judge the pixel quality of the inspection image to identify the normal pixel point coordinates and abnormal pixel point coordinates; using all the normal pixel point coordinates to construct the recognizable range, and using all the abnormal pixel point coordinates to construct the unrecognizable range; performing a preliminary recognition on the pixel points within the recognizable range to obtain the covered target content, and obtaining the effective information based on the covered target content; performing a preliminary recognition on the pixel points within the unrecognizable range to obtain the abnormal type content, and obtaining the fuzzy information based on the abnormal type content.
[0009] In some optional embodiments, the steps of constructing the recognizable range by using all the normal pixel point coordinates and constructing the unrecognizable range by using all the abnormal pixel point coordinates include the following: Assigning a morphological mask value to each normal pixel point coordinate and each abnormal pixel point coordinate of the inspection image; respectively constructing a first positive mask value distribution map and a first negative mask value distribution map according to the morphological mask values of all the normal pixel point coordinates, and respectively constructing a second positive mask value distribution map and a second negative mask value distribution map according to the morphological mask values of all the abnormal pixel point coordinates; determining the recognizable range by using the first positive mask value distribution map and the second negative mask value distribution map, and determining the unrecognizable range by using the first negative mask value distribution map and the second positive mask value distribution map.
[0010] In some optional embodiments, the steps of preliminarily identifying the pixel points within the recognizable range and obtaining the covered target content include the following: determining a first coverage area of the recognizable range, extracting features from the mask value distribution map of the first coverage area, detecting and identifying the extracted features to obtain the covered target content; the steps of preliminarily identifying the pixel points within the unrecognizable range to obtain the abnormal type content include the following: determining a second coverage area of the unrecognizable range and the intersection area where the unrecognizable range overlaps with the recognizable range; quantifying the blur degree of the mask value distribution within the second coverage area, performing clustering analysis on all the quantified blur degree information to generate a blur degree visualization distribution map; performing a preliminary screening of the abnormal type on the blur degree visualization distribution map to obtain a preliminary screening result of the abnormality; determining the first coverage area associated with the intersection area, and re-screening the preliminary screening result of the abnormality according to the associated first coverage area to obtain the abnormal type content.
[0011] In some optional embodiments, the steps of pairing each of the preliminary screening modal data with the valid information and the fuzzy information to obtain a first pairing result and a second pairing result include the following: obtaining the format conversion information of the preliminary screening modal data; performing a correlation analysis between the format conversion information and the mask value features within the first coverage area to obtain a first feature matching correlation degree, and calculating the first pairing result according to the first feature matching correlation degree; performing a correlation analysis between the format conversion information and the mask value features within the second coverage area to obtain a second feature matching correlation degree, and calculating the second pairing result according to the second feature matching correlation degree; wherein, after obtaining the second feature matching correlation degree, it further includes: performing a correlation analysis between the format conversion information and the mask value features within the intersection area to obtain a third feature matching correlation degree, and calculating the second pairing result according to the second feature matching correlation degree and the third feature matching correlation degree.
[0012] In some optional embodiments, calculating the second pairing result according to the second feature matching correlation degree and the third feature matching correlation degree includes the following steps: determining a first weight and a second weight; comparing the difference between the second feature matching correlation degree and the third feature matching correlation degree, and adjusting the first weight and / or the second weight according to the difference; assigning the adjusted first weight to the second feature matching correlation degree and / or the adjusted second weight to the third feature matching correlation degree.
[0013] In some optional embodiments, associating and annotating the fused modal data with the inspection image includes the following steps: determining the association area position information of the inspection image and the fused modal data; locking the association objects of the inspection image and the fused modal data respectively based on the association area position information; judging the type of the association form based on the association objects of the inspection image and the fused modal data; wherein, the type of the association form includes one of full association, strong association and weak association.
[0014] In some optional embodiments, judging the type of the association form based on the association objects of the inspection image and the fused modal data includes the following steps: calculating the similarity of any association object of the inspection image and the fused modal data, merging the similarities of all association objects of the inspection image and the fused modal data to obtain a total similarity, and determining the type of the association form based on the interval comparison of the total similarity.
[0015] In some optional embodiments, it further includes the step of adjusting the type of the association form: tracing the second pairing result of the fused modal data, calculating the association adjustment parameter determined by the second pairing result, assigning the association adjustment parameter to the total similarity to obtain an adjusted similarity, and re-determining the type of the association form based on the adjusted similarity.
[0016] Second aspect, an image quality classification processing system based on a multi-modal large model, includes:
[0017] A first recognition unit, which is used to recognize the target attributes of the inspection image, perform a preliminary analysis on the inspection image based on the target attributes, and determine the modal data matching strategy of the inspection image;
[0018] A second recognition unit, which is used to screen an associated modal data set from the inspection database based on the modal data matching strategy, classify the associated modal data set to obtain multiple associated modal data;
[0019] A first processing unit, which is used to perform a matching judgment on each type of the associated modal data and the inspection image to obtain the matching degree of this type of associated modal data, and use the associated modal data with a matching degree higher than a preset matching value as the preliminarily screened modal data, wherein the matching judgment includes inspection time judgment, inspection object judgment, and inspection location judgment;
[0020] A second processing unit, which is used to delimit the recognizable range and unrecognizable range of the inspection image, extract effective information according to the recognizable range, extract fuzzy information according to the unrecognizable range, and pair each type of the preliminarily screened modal data with the effective information and the fuzzy information respectively to obtain a first pairing result and a second pairing result;
[0021] A third processing unit, which is used to judge whether the preliminarily screened modal data is used as the fused modal data based on the first pairing result and the second pairing result, and perform an associated annotation process on the fused modal data and the inspection image, wherein the associated annotation process includes annotating the associated area, annotating the associated object, and annotating the associated form.
[0022] The beneficial effects of the embodiments of the present invention are:
[0023] The image quality classification processing method and system based on a multi-modal large model provided by the embodiments of the present invention can, by initially identifying the target attributes of the inspection image, determine the type of modal data that can be matched based on the target attributes, and directly extract different categories of associated modal data from the inspection database according to the corresponding matching strategy of the type of modal data that can be matched. This associated modal data has the basis for data fusion analysis with the inspection image; then, by judging the matching degree between each type of associated modal data and the inspection image, the matching degree judgment operation between different modal information is carried out from multiple aspects such as whether the inspection time is consistent, whether the inspection object is consistent, and whether the inspection location is consistent. On this basis, it can further analyze whether there are defects in the unrecognizable range of the inspection image, that is, by aiming at the pairing degree between the associated modal data and the recognizable range and the unrecognizable range respectively to deal with problems such as occlusion, overexposure, and blurred shooting that may exist in the inspection image, so as to more accurately pair the inspection image in this defective situation with the preliminarily screened modal data, and use the preliminarily screened modal data with a higher final pairing degree as the fused modal data, making the fused modal data and the inspection image more conducive to subsequent training or analysis and judgment by the multi-modal large model, thereby ensuring the quality of the input image data of the multi-modal large model and obtaining a more accurate output result.
[0024] Generally speaking, the image quality classification processing method and system based on the multi-modal large model provided by the embodiments of the present invention can analyze and judge from the situations of whether the associated modal data is initially matched and whether it can be highly matched with the defect inspection image, so as to perform associated annotation processing on the finally selected fused modal data and the inspection image, so as to achieve different levels of classification judgment processing from the initial associated modal data to the final fused modal data, so as to ensure the data quality of the input image of the final multi-modal large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a flowchart of the main steps of the processing method provided by the embodiments of the present invention;
[0027] Figure 2 is Figure 1 The first part of the flowchart of one of the main steps S400 shown;
[0028] Figure 3 is Figure 2 The sub-step flowchart of S420 in the first part of the steps shown;
[0029] Figure 4 is Figure 2 The sub-step flowchart of S430 and S440 in the first part of the steps shown;
[0030] Figure 5 is Figure 1 The second part of the flowchart of one of the main steps S400 shown;
[0031] Figure 6 is Figure 1 The flowchart of one of the main steps S500 shown;
[0032] Figure 7 is Figure 6 The sub-step flowchart of S530 in the step S500 shown;
[0033] Figure 8 It is a modular schematic diagram of the processing system provided by the embodiments of the present invention.
[0034] Icons: 600 - Processing system; 610 - First recognition unit; 620 - Second recognition unit; 630 - First processing unit; 640 - Second processing unit; 650 - Third processing unit. Detailed implementation manners
[0035] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Components of the embodiments of the present invention described and illustrated in the drawings here usually can be arranged and designed in various different configurations.
[0036] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0037] It should be understood that the "system", "device" and / or "module" used in the present invention is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the said words can be replaced by other expressions.
[0038] As shown in the present invention and the claims, unless the context clearly indicates an exception, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0039] Flowcharts are used in the present invention to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations before or after do not necessarily need to be executed precisely in sequence. On the contrary, they can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several steps can be removed from these processes.
[0040] Example: When preprocessing the input data of a multi-modal large model, different-modal data is classified and screened, and then data fusion is performed as the model input. In this process, how to classify and screen different-modal data has become a key issue. Among them, different-modal data can be different forms of the same type of data (such as image data with different shooting angles and shooting conditions), or different types of data (such as image data, voice data, and sensing data, etc.). Currently, the most widely used is image data. Multi-modal learning or analysis and judgment are carried out based on the image data transmitted back by the inspection robot (ground robot or aerial drone), that is, it means using different-modal image data for classification and fusion as the training data set or analysis and judgment data set of the multi-modal large model.
[0041] In the foregoing operation mode, we perform classification processing according to the shooting source of the captured image data and the device identification number of the shooting target, and then form fusion data from the same type of data as the model input training set and test set, which can ensure a certain judgment accuracy, but there will still be some misjudgment situations. After in-depth analysis, it is found that the quality of the input fusion data is the guarantee basis for the accuracy of the final analysis result, that is, it is necessary to consider how to ensure the quality of the fusion data with higher matching. Therefore, it is necessary to further optimize how to classify and screen multi-modal data in the early stage and fusion annotation in the later stage, that is, consider aspects such as the matching degree, consistency, and annotation relevance between different-modal data. On this basis, we also found that once there are defects in the captured image data (such as target occlusion, unstable lighting, blurred shooting, etc.), the formed fusion data will also greatly reduce the accuracy of the subsequent analysis results. In response to this situation, it is also one of the key issues to consider how to ensure that the key defective image data can also be classified and screened with higher matching degree, consistency, and annotation relevance. In response to the above-mentioned related problems, this embodiment provides an image quality classification processing method based on a multi-modal large model. This processing method can perform a strong correlation match between the mainly used image data and other different-modal data, and also consider the correlation match of defective image data during this matching process, so that the selected other-modal data can be effectively annotated with this image data, so as to ensure the quality of the fused synthetic data.
[0042] Specifically, please refer to Figure 1 , an image quality classification processing method based on a multi-modal large model provided by this embodiment includes the following steps:
[0043] S100: Identify the target attributes of the inspection image, perform a preliminary analysis on the inspection image based on the target attributes, and determine the modal data matching strategy for the inspection image. This step means identifying the target attributes of the inspection image that is to be used as one of the bases for the training set or test set. The target attribute identification mainly focuses on the attributes (type, address, power-on status, etc.) of the target device to be monitored in the inspection image for preliminary identification. The identified target attribute data can be uploaded to the background server by the inspection device together with the image data when uploading, so as to facilitate the background server to extract strategies or plans based on the target attributes, and thus determine the modal data matching strategy for the inspection image. The modal data matching strategy mainly refers to the strategic plan formed by the type of modal data that can be matched, the data source address, and the target, etc. This strategic plan can be pre-configured or formed by combining pre-configuration and later dynamic adjustment, which is not limited here.
[0044] S200: Filter the associated modal data set from the inspection database based on the modal data matching strategy, and classify the associated modal data set by type to obtain multiple associated modal data. This step means filtering out all the data that can be associated with the target attributes of the inspection image from the inspection database received by the background server through the modal data matching strategy obtained in the above step as the associated modal data set, and then classifying the associated modal data set by type. Specifically, taking this associated modal data set as an image data set (the same below) as an example, due to the differences in the shooting time, shooting location, shooting angle, shooting target coverage, and shooting method of different inspection devices, there will be different modal image data categories for the same power equipment. It is necessary to classify these different modal image data categories, for example, classify them according to the identity of the shooting time, shooting location, shooting angle, shooting target coverage, and shooting method, so as to obtain multiple associated modal data. There may be differences in at least one of the shooting time, shooting location, shooting angle, shooting target coverage, and shooting method between different associated modal data.
[0045] S300: Judge the matching degree between each of the associated modality data and the inspection image to obtain the matching degree of this type of associated modality data. Take the associated modality data with a matching degree higher than the preset matching value as the preliminarily screened modality data. Among them, the matching judgment includes inspection time judgment, inspection object judgment, and inspection location judgment; this step means respectively judging the matching degree between all different types of initially extracted associated modality data and the inspection image, mainly judging whether the inspection time is consistent, whether the inspection object is consistent, and whether the inspection location is consistent (in different implementation manners, it can also be further subdivided into judgments such as whether the inspection angle is consistent and whether the inspection method is consistent). After judging the matching of the inspection time, inspection object, and inspection location, obtain the matching degree of each type of associated modality data, compare this matching degree with the preset matching value, and take the associated modality data with a matching degree higher than the preset matching value as the preliminarily screened modality data, that is, as the object of preliminary matching, it can be subjected to subsequent classification processing steps.
[0046] S400: Define the recognizable range and unrecognizable range of the inspection image, extract effective information according to the recognizable range, extract fuzzy information according to the unrecognizable range, and pair each of the preliminarily screened modality data with the effective information and the fuzzy information respectively to obtain a first pairing result and a second pairing result; this step means pre - considering by judging whether the inspection image has defects and whether it will have a greater impact when making a matching judgment with the associated modality data to cope with the problem of relatively reduced quality of the fused data in the later stage. That is, it means respectively defining the recognizable range and unrecognizable range on the inspection image. If there is an unrecognizable range, it is considered that the inspection image has defects (factors such as occlusion, over - exposure, and blurring). Then, when pairing with the corresponding preliminarily screened modality data, it is necessary to consider whether the recognizable range is highly paired and whether the unrecognizable range is highly paired, so as to ensure that the two can be used as the basis for fused modality data with high matching and consistency.
[0047] Specifically, it is necessary to extract effective information from the recognizable range of the inspection image, extract fuzzy information from the unrecognizable range, and then pair each of the preliminarily screened modality data with the effective information and the fuzzy information respectively to judge whether the shooting content consistency is relatively high, so as to obtain a first pairing result corresponding to the effective information and a second pairing result corresponding to the fuzzy information respectively.
[0048] S500: Determine whether the preliminarily screened modal data is used as fused modal data based on the first pairing result and the second pairing result, and perform associated annotation processing on the fused modal data and the inspection image, where the associated annotation processing includes annotating the associated region, the associated object, and the associated form; this step means comprehensively considering the obtained first pairing result and the second pairing result, and then judging whether the preliminarily screened modal data is used as fused modal data based on the combined result of the two. The fused modal data refers to the basis (with high consistency) that can be classified into the same category as the inspection image and can be used for fused data synthesis, so that the obtained fused modal data can ensure that the synthesized image data has high quality and can be used as the input data for the training or analysis and judgment of the multi-modal large model. It should be noted that before being used as input data, the fused modal data and the inspection image need to be subjected to associated annotation processing, mainly including the annotation of information such as the associated region (coordinate region in the image), the associated object (target object in the image), and the associated form (degree of association). Combining the image data and accurate annotation information can improve the accuracy of the output result of the multi-modal large model.
[0049] Through the above technical solutions, in order to ensure the quality and relevance of the image data in the multi-modal large model and thus improve the accuracy of the training, analysis, and judgment of the multi-modal large model. First, identify the target attributes of the inspection image, and based on this, determine the modal data matching strategy to ensure the pertinence of the preliminary screening. Then, screen out the modal data set related to the target attributes from the database and perform type division to ensure the effective separation of different modal data. Through precise matching judgment (including consistency checks of time, object, and location), screen out the modal data with high matching degree as the preliminary screening result. On this basis, for possible image defects, delimit the recognizable range and unrecognizable range of the inspection image, and pair the preliminarily screened modal data with it to ensure a high matching degree and consistency even in the case of image quality problems. Finally, comprehensively consider the pairing results to determine which preliminarily screened modal data can be used as fused modal data and perform associated annotation processing on it to ensure the quality and usability of the final synthesized data. This not only optimizes the classification and screening of multi-modal data but also effectively solves the problem of data quality decline caused by image defects, provides high-quality training and test data for the multi-modal large model, and thus improves the reliability of the model output.
[0050] On the basis of the above technical solutions, the process of reflecting whether there are defects in the inspection image by delimiting the recognizable range and unrecognizable range of the inspection image is mainly achieved through the processing of pixel points. For details, please refer to Figure 2, delimiting the recognizable range and unrecognizable range of the inspection image, extracting effective information according to the recognizable range, and extracting fuzzy information according to the unrecognizable range include the following steps:
[0051] S410: Use a trained model to analyze and judge the pixel quality of the inspection image, and identify the coordinates of normal pixel points and abnormal pixel points; this step means using a preliminary judgment model to preliminarily analyze the quality of the pixel points of the inspection image. For example, for each inspection image of a convolutional neural network model, analyze each pixel or small block of pixels one by one through a sliding window method to obtain the feature vector at each position. On the basis of feature extraction, add a classification layer to distinguish normal pixels and abnormal pixels, and determine which pixel points are marked as abnormal and which are marked as normal through a set threshold, so as to obtain the coordinates of normal pixel points and abnormal pixel points.
[0052] S420: Use all the coordinates of the normal pixel points to construct the recognizable range, and use all the coordinates of the abnormal pixel points to construct the unrecognizable range; this step means concatenating all the obtained coordinates of the normal pixel points to construct the pixel area of the recognizable range, and similarly using all the coordinates of the abnormal pixel points concatenated to construct the pixel area of the unrecognizable range. It should be noted that the recognizable range and the unrecognizable range can be a whole connected area or multiple segmented areas. Subsequently, each area of the recognizable range or unrecognizable range will be taken as an example for analysis and explanation (the same applies to the remaining areas).
[0053] S430: Conduct a preliminary identification of the pixel points within the recognizable range and obtain the content covering the target, and obtain the effective information based on the content covering the target; this step means obtaining the content covering the target within the recognizable range based on object detection (using the model to output the position information and category label of each detected target) and semantic segmentation technology (obtaining the contour and category information of each detected target), and the target type, contour and position of this target content are used as the effective information.
[0054] S440: Conduct a preliminary identification of the pixel points within the unrecognizable range to obtain the content of the abnormal type, and obtain the fuzzy information based on the content of the abnormal type; this step means using object detection technology to extract the feature contour and position of the unrecognizable range, and then the possible target types (fuzzy body, occluder, overexposed body, etc.) can be output through domain judgment (context judgment) as the content of the abnormal type.
[0055] Through the above technical solution, the training model is used to carefully analyze the pixel quality of the inspection image, distinguish the coordinates of normal pixel points and abnormal pixel points, separate the clear part and the defective part in the image, and provide a basis for targeted analysis. Then, for the pixel points within the recognizable range, object detection and semantic segmentation technologies are used to accurately identify and obtain the information covering the target content, ensuring a detailed understanding of the key elements in the image; while for the pixel points outside the recognizable range, the abnormal type content is determined through preliminary recognition combined with context judgment, and the information in the blurred or occluded area is parsed as much as possible, so as to provide a more detailed pairing comparison basis for the subsequent preliminary screening of modal data.
[0056] Considering that there may be three judgment situations when using the training model to analyze and judge the pixel quality of the inspection image, that is, the coordinate of the pixel point is a normal pixel point coordinate, an abnormal pixel point coordinate, or a pixel point coordinate that cannot be judged. At this time, there will be a division blind area (neither belonging to the recognizable range nor belonging to the unrecognizable range) in the process of delimiting the recognizable range and the unrecognizable range. To address this situation and ensure that the information (texture features) contained in each pixel point can be utilized, especially in the subsequent pairing process of the preliminary screening modal data, in this embodiment, the recognizable range and the unrecognizable range are delimited by the method of positive and negative union. For details, please refer to Figure 3 using all the normal pixel point coordinates to construct the recognizable range and using all the abnormal pixel point coordinates to construct the unrecognizable range includes the following steps:
[0057] S421: Assign a morphological mask value to each normal pixel point coordinate and each abnormal pixel point coordinate of the inspection image; this step means assigning recognizable numerical values to each normal pixel point coordinate and each abnormal pixel point coordinate as the morphological mask value. This morphological mask value is generally of two positive and negative types, which can be 0 or 1, or 1 and -1, or other numerical combinations, etc. The purpose is to distinguish the types of normal pixel point coordinates and abnormal pixel point coordinates.
[0058] S422: Construct a first positive mask value distribution map and a first negative mask value distribution map respectively according to the morphological mask values of all normal pixel point coordinates, and construct a second positive mask value distribution map and a second negative mask value distribution map respectively according to the morphological mask values of all abnormal pixel point coordinates; this step means constructing distribution maps of the morphological mask values of all pixel point coordinates on the inspection image. If it is a normal pixel point coordinate, it is assigned a first positive mask value (1), and the rest of the pixel point coordinates are assigned a first negative mask value (-1), so as to construct a first positive mask value distribution map and a first negative mask value distribution map, and the two are opposite complementary morphological mask value distribution maps.
[0059] Similarly, if it is determined to be the abnormal pixel point coordinates, a second positive mask value (-1) is assigned, and for the remaining pixel point coordinates, a second negative mask value (1) is assigned, thereby constructing a second positive mask value distribution map and a second negative mask value distribution map, and both are also opposite complementary morphological mask value distribution maps.
[0060] S423: Determine the recognizable range by using the first positive mask value distribution map and the second negative mask value distribution map, and determine the unrecognizable range by using the first negative mask value distribution map and the second positive mask value distribution map; this step means taking the union of the first positive mask value distribution map (the distribution map of value 1) and the second negative mask value distribution map (the distribution map of value 1) to determine the recognizable range. Similarly, take the union of the first negative mask value distribution map (the distribution map of value -1) and the second positive mask value distribution map (the distribution map of value -1) to determine the unrecognizable range. By this positive and negative union method, the situation of the mask values of the pixel point coordinates that cannot be judged can be eliminated, thus facilitating subsequent pixel feature matching.
[0061] Through the above technical solution, the combined area of the recognizable range and the unrecognizable range obtained can completely cover the entire inspection image, and for the junction of the recognizable range and the unrecognizable range, it is the coordinate position where the above-mentioned pixel point types that cannot be judged are located. These pixel points alone cannot accurately judge their texture features, but by combining the surrounding neighboring pixel points, their texture features can be further speculated, so as to more reliably judge the abnormal types implied by the blurred texture features of the pixel points in the unrecognizable range. Therefore, when extracting fuzzy information, the available information included at the junction of the aforementioned recognizable range and unrecognizable range needs to be considered. For details, please refer to Figure 4 , the preliminary recognition of the pixel points within the recognizable range and obtaining the covered target content includes the following steps:
[0062] S431: Determine the first coverage area of the recognizable range, extract features from the mask value distribution map of this first coverage area, and detect and recognize the extracted features to obtain the covered target content; this step means determining the mask values of all pixel points within the pixel point coverage area of the recognizable range (bounding box determination), and then extracting, detecting and recognizing features according to the distribution of mask values (such as object detection models and semantic segmentation models), and taking the extracted feature information and class labels as the covered target content.
[0063] Please refer to again Figure 4 , the preliminary recognition of the pixel points within the unrecognizable range to obtain the abnormal type content includes the following steps:
[0064] S441: Determine the second coverage area of the unrecognizable range and the intersection area where the unrecognizable range overlaps with the recognizable range; this step means separately determining the pixel point coverage area of the unrecognizable range for demarcation (as the second coverage area) and the pixel point coverage area where the recognizable range overlaps with the recognizable range for demarcation (as the intersection area). Since the above-mentioned positive and negative union method is used to construct the recognizable range and the unrecognizable range, there is an intersection range between the two, and these intersection ranges are composed of the pixel point coordinates that cannot accurately judge the type.
[0065] S442: Perform fuzziness quantification on the mask value distribution in the second coverage area, and perform cluster analysis on all the quantified fuzziness information to generate a fuzziness visualization distribution map; this step means performing numerical quantification of fuzziness on the mask value distribution map in the second coverage area. The fuzzy value can be any value between 0 and 1, or data in other ranges. The purpose is to convert and represent it according to the confidence value in the process of judging the pixel point type, obtain the specific confidence value of each pixel point whose type cannot be accurately judged, and convert and represent the fuzziness value of this pixel point. Then perform cluster analysis on the fuzziness information of all pixel point coordinates to generate a fuzziness visualization distribution map, so that the fuzzy distribution of all pixel points in the second coverage area can be intuitively obtained.
[0066] S443: Perform preliminary screening of abnormal types on the fuzziness visualization distribution map to obtain a preliminary screening result of abnormalities; this step means performing a preliminary screening operation on possible abnormal types according to the obtained visualization distribution map. For example, judge whether there is occlusion (the overall fuzzy value is concentrated and there is occlusion texture), exposure (the overall value is concentrated and there is no obvious texture), or fuzziness (the fuzzy value is scattered and there are multiple textures) according to the concentration of pixel point fuzziness and the range of fuzzy values, etc., so as to obtain the possible situation of the preliminary screening result of abnormalities.
[0067] S444: Determine the first coverage area associated with the intersection area, and perform re-screening on the preliminary screening result of abnormalities according to the associated first coverage area to obtain the content of the abnormal type. This step means assisting the re-screening judgment of the preliminary screening result of abnormalities by analyzing the covering target content of the first coverage area associated (directly adjacent) to the intersection area, further excluding the possibility of occlusion, exposure, or fuzziness, and finally obtaining the relatively certain content of the abnormal type.
[0068] Through the above technical solutions, combined with the analysis of the recognizable range and the unrecognizable range, for the target area that can be clearly recognized (the first coverage area), means such as feature extraction, detection, and recognition are used to obtain the target content; while for the fuzzy area that cannot be directly recognized (the second coverage area) and the junction between the two, means such as fuzziness quantification, clustering analysis, and visualization distribution maps are introduced to evaluate the abnormal types of pixel points in these areas. Among them, by associating the information of the first coverage area in the intersection area, the judgment accuracy of the abnormal types in the unrecognizable range is further optimized, which not only ensures the effective analysis of the entire inspection image, but also specifically for the pixel points that are difficult to judge, uses their neighborhood information for auxiliary decision-making to improve the identification ability for situations such as occlusion, poor exposure, or real fuzziness.
[0069] By obtaining the recognizable range and its effective information and the unrecognizable range and its fuzzy information of the inspection image, it can be further paired separately with the preliminary screening modal data, so as to avoid reducing the quality of the subsequent fusion data due to the fact that the preliminary screening modal data cannot have a high matching degree with the defective area because there is a defective area in the inspection image. Therefore, it is necessary to consider pairing the preliminary screening modal data with the effective information and the fuzzy information respectively, and then further consider whether to use it as the supporting fusion data of the inspection image after both pairing results meet the requirements. For details, please refer to Figure 5 The step of pairing each of the preliminary screening modal data with the effective information and the fuzzy information respectively to obtain the first pairing result and the second pairing result includes the following steps:
[0070] S450: Obtain the format conversion information of the preliminary screening modal data; this step means first converting the format of the preliminary screening modal data. Taking this preliminary screening modal data as image data as an example, format conversion includes steps such as image format standard conversion, size adjustment (scaling, cropping), spatial conversion processing (color, pixel value standardization), and metadata processing, so as to obtain the format conversion information through the above format conversion processing. If this preliminary screening modal data is voice data, format conversion includes steps such as file format standardization, encoding quality adjustment, feature generation and its normalization, and time-space synchronization processing of voice files (with image data) to obtain the format conversion information. This embodiment mainly takes the preliminary screening modal data as image data as an example to illustrate by classifying and analyzing other modal image data with different shooting perspectives, shooting methods, and shooting times with the inspection image.
[0071] S460: Perform a correlation analysis on the format conversion information and the mask value features within the first coverage area to obtain a first feature matching correlation degree, and calculate the first pairing result based on this first feature matching correlation degree; this step means performing feature extraction (including metadata recording) on the obtained format conversion information and performing a correlation analysis with the mask value features within the first coverage area (local feature extraction based on bounding boxes), using feature alignment methods (ensuring that the image features after format conversion and the original mask value features are compared in the same coordinate system) and similarity measurement methods to calculate the similarity of all aligned features, so as to obtain the first feature matching correlation degree (such as a weighted scoring calculation method), and finally calculate the first pairing result based on the first feature matching correlation degree.
[0072] S470: Perform a correlation analysis on the format conversion information and the mask value features within the second coverage area to obtain a second feature matching correlation degree, and calculate the second pairing result based on this second feature matching correlation degree; this step means performing feature extraction (including metadata recording) on the obtained format conversion information and performing a correlation analysis with the mask value features within the second coverage area (local feature extraction based on bounding boxes), and performing feature alignment and similarity measurement on the mask value features that can be extracted within the second coverage area and the extraction features of the format conversion information, so as to obtain the second feature matching correlation degree, and then calculate the second pairing result through the second feature matching correlation degree. Since the mask value features that can be extracted within the second coverage area may be limited, there may be a situation where the judgment accuracy is insufficient when calculating the second pairing result only based on the mask value features within the second coverage area. Therefore, on this basis, it is also necessary to further consider performing a correlation analysis with the mask value features within the cross-region, that is, performing the following steps.
[0073] S480: After obtaining the second feature matching correlation degree, it further includes: performing a correlation analysis on the format conversion information and the mask value features within the cross-region to obtain a third feature matching correlation degree, and calculating the second pairing result based on the second feature matching correlation degree and the third feature matching correlation degree; this step means performing feature extraction (including metadata recording) on the obtained format conversion information and performing a correlation analysis with the mask value features within the cross-region, so as to obtain the third feature matching correlation degree. It should be noted that this third feature matching correlation degree covers clearer texture features and can assist in judging whether the mask value features within the second coverage area are highly correlated and paired. Therefore, calculating the second pairing result based on the second feature matching correlation degree and the third feature matching correlation degree makes the second pairing result more reasonable in judgment and can avoid the problem that the pairing result cannot be accurately calculated due to unclear mask value features within the second coverage area.
[0074] Through the above technical solution, format conversion processing is performed on the preliminarily screened modal data to obtain its format conversion information, and then these format conversion information are respectively subjected to correlation analysis with the mask value features in the first coverage area (valid information) and the second coverage area (blurred information) defined in the inspection image, so as to calculate the first and second feature matching correlation degrees, and then obtain the corresponding pairing results respectively. Considering that there may be problems with unclear features in the second coverage area, the mask value features of the cross area are further introduced for auxiliary analysis to improve the accuracy of the second pairing result, which not only enhances the matching accuracy between different modal data, but also ensures that reliable data fusion quality can be obtained even in the case of defects.
[0075] On this basis, if the texture of the mask value features in the cross area is clearer, it can better assist in the high-precision calculation of the second pairing result. Therefore, when calculating the second pairing result according to the second feature matching correlation degree and the third feature matching correlation degree, it is necessary to consider the numerical data represented by the second feature matching correlation degree and the third feature matching correlation degree respectively, and the one with a greater similarity can be given a greater weight in the calculation. Specifically, the calculating the second pairing result according to the second feature matching correlation degree and the third feature matching correlation degree includes the following steps:
[0076] Determine the first weight and the second weight. These first weight and second weight can be initial assigned values and can be obtained according to experience; then compare the differences between the second feature matching correlation degree and the third feature matching correlation degree to clarify the magnitude of the differences between the two, and then adjust the first weight and / or the second weight according to this difference. For example, the greater the value of one of them, the greater the upward adjustment amplitude. This purpose can be achieved by adjusting one of them, or by adjusting both of them at the same time; finally, assign the adjusted first weight to the second feature matching correlation degree and / or assign the second weight to the third feature matching correlation degree, that is, it means that if one of them is adjusted, the adjusted first weight is assigned to the adjusted one, and thus the final second pairing result is calculated.
[0077] Through the above series of operation steps, the first pairing result and the second pairing result of the preliminary screening modal data with the recognizable range and the unrecognizable range of the inspection image can be obtained. Then, according to the respective pairing results of the first pairing result and the second pairing result, it is comprehensively judged whether the preliminary screening modal data is used as the fusion modal data. If both the first pairing result and the second pairing result meet the requirements, it can be determined whether the preliminary screening modal data is used as the fusion modal data. It should be noted that due to defects (such as occlusion, overexposure, blur, etc.) in the unrecognizable range of the inspection image, if the preliminary screening modal data (such as image data) also covers the same area, but the features of the preliminary screening modal data in this area are clearer and can also be highly paired with the features that may originally exist in the unrecognizable range (using the features in the above cross-region for auxiliary judgment), it means that the corresponding second pairing result is more likely to meet the requirements, and the preliminary screening modal data can be used as the fusion modal data.
[0078] After obtaining the fusion modal data, it indicates that the original associated modal data and the inspection image are highly matched in terms of matching and consistency. When synthesizing the two data, in order to facilitate the multi-modal large model to also recognize the semantic label of the feature during feature detection, so as to obtain a more accurate final output analysis result, it is necessary to perform associated annotation processing on the two before data synthesis. For details, please refer to Figure 6 , the associated annotation processing of the fusion modal data and the inspection image includes the following steps:
[0079] S510: Determine the position information of the associated area between the inspection image and the fusion modal data; this step means first locating the specific position of the associated area of the inspection image in the fusion modal data, then locating the area position information covering the same detection target, and then further performing target associated annotation, that is, performing step S520: Lock the associated objects of the inspection image and the fusion modal data respectively based on the position information of the associated area; this step means locking the object targets associated with each other in the associated area position, which can be one or more, and performing associated annotation on the associated objects respectively. Finally, perform step S530: Judge the type of the associated form based on the associated objects of the inspection image and the fusion modal data; this step means further annotating the associated form on each group of associated objects to indicate the degree of their association. The type of this associated form includes one of full association, strong association, and weak association. Different types of associated forms indicate the strength of the association degree between the corresponding associated objects, so as to facilitate differentiation during subsequent multi-modal large model analysis.
[0080] Through the above technical solution, after determining that the preliminary screening modal data can be used as the fusion modal data, further perform correlation annotation processing to ensure that the synthesized data has clear semantic tags. By locating the position information of the correlation region, locking the correlation object, and judging and annotating the type of the correlation form (such as full correlation, strong correlation or weak correlation), structured data that is easy to parse and process can be provided for the multi-modal large model. On the basis of the above technical solution, if there are multiple groups of correlation objects in the same correlation region, the correlation degree of each group needs to be considered separately, so as to finally judge the type of the correlation form between the fusion modal data and the inspection image with respect to the correlation region. For details, please refer to Figure 7 , the method for judging the type of the correlation form based on the correlation object between the inspection image and the fusion modal data includes the following steps:
[0081] S531: Calculate the similarity between any correlation object in the inspection image and the fusion modal data, and merge the similarities of all correlation objects in the inspection image and the fusion modal data to obtain the total similarity; this step represents judging the similarity of each group of correlation objects between the inspection image and the fusion modal data (such as using the feature similarity measurement method), and then merging the similarities of all groups of correlation objects (such as taking the mean value or taking the weighted mean value), and finally obtaining the total similarity between the inspection image and the fusion modal data with respect to the correlation objects in the corresponding correlation region. Then perform step S532: Based on this total similarity, perform interval comparison (configure an interval comparison table), and use the total similarity to determine the type of the correlation form of the corresponding correlation region, whether it is full correlation, strong correlation or weak correlation.
[0082] Considering that when determining the correlation region in the inspection image, the correlation region may exist in the unrecognizable range of the inspection image. At this time, there is a certain obstacle in judging the similarity of the correlation objects in this region, and it is impossible to accurately judge the corresponding type of the correlation form. At this time, the second pairing result can be used to assist in judging the degree of similarity, that is, after the step of judging the type of the correlation form, there is also a step of adjusting the type of the correlation form:
[0083] S533: Trace the second pairing result of the fused modal data, calculate the association adjustment parameter determined by this second pairing result, assign the association adjustment parameter to the total similarity to obtain an adjusted similarity, and re-determine the type of association form based on this adjusted similarity; this step means that by tracing the second pairing result calculated in the previous judgment step, a corresponding association adjustment parameter is determined according to the value of the second pairing result (such as determined by experience), then the association adjustment parameter is assigned to the total similarity (such as in a summation manner), and finally the adjusted similarity is obtained. This adjusted similarity can assist in the judgment result of the similarity of associated objects within the unrecognizable range, making it closer to the true analysis result, avoiding the problem of inaccurate calculation, and making the re-determination of the type of association form based on the finally obtained adjusted similarity more reasonable and reliable.
[0084] Through the above technical solution, for the situation where there are multiple groups of associated objects in the same association area, the similarity of each group of associated objects can be calculated and merged to obtain the total similarity, and then the type of association form (such as full association, strong association, or weak association) can be determined based on the total similarity, so as to provide structured and easily analyzable data for the multi-modal large model. Secondly, for the associated objects within the unrecognizable range in the inspection image, the second pairing result is introduced as an auxiliary judgment means. By calculating the association adjustment parameter and applying it to the total similarity, the adjusted similarity is obtained, and the type of association form is re-determined based on this. This not only overcomes the problem of similarity judgment caused by image defects, but also improves the accuracy and reliability of the association form judgment, ensuring the authenticity and reasonableness of the final analysis result.
[0085] In this embodiment, an image quality classification processing system based on a multi-modal large model (hereinafter referred to as processing system 600) is also provided. Please refer to Figure 8 the modular schematic diagram of this processing system 600, which is mainly used to divide the functional modules of the processing system 600 according to the embodiments of the above method. For example, each functional module can be divided, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software function modules. It should be noted that the division of modules in the present invention is schematic, only a logical function division, and there can be other division methods in actual implementation. For example, in the case of dividing each functional module corresponding to each function, Figure 8 only a system / device schematic diagram is shown. Among them, the processing system 600 can include a first recognition unit 610, a second recognition unit 620, a first processing unit 630, a second processing unit 640, and a third processing unit 650. The functions of each unit module will be described below.
[0086] The first recognition unit 610 is configured to recognize the target attributes of the inspection image, perform a preliminary analysis on the inspection image based on the target attributes, and determine the modal data matching strategy for the inspection image;
[0087] The second recognition unit 620 is configured to screen the associated modal data set from the inspection database based on the modal data matching strategy, classify the types of the associated modal data set, and obtain multiple associated modal data;
[0088] The first processing unit 630 is configured to perform a matching judgment between each of the associated modal data and the inspection image, obtain the matching degree of the associated modal data, and use the associated modal data with a matching degree higher than a preset matching value as the preliminarily screened modal data, wherein the matching judgment includes inspection time judgment, inspection object judgment, and inspection location judgment;
[0089] A second processing unit 640, which is used to delimit the recognizable range and unrecognizable range of the inspection image, extract valid information according to the recognizable range, extract fuzzy information according to the unrecognizable range, and pair each of the preliminary screening modality data with the valid information and the fuzzy information respectively to obtain a first pairing result and a second pairing result; In some embodiments, the second processing unit 640 is further used to analyze and judge the pixel quality of the inspection image by using a training model, and identify the coordinates of normal pixel points and abnormal pixel points; construct the recognizable range by using all the coordinates of the normal pixel points, and construct the unrecognizable range by using all the coordinates of the abnormal pixel points; perform preliminary identification on the pixel points within the recognizable range and obtain the covered target content, and obtain the valid information based on the covered target content; perform preliminary identification on the pixel points within the unrecognizable range to obtain the abnormal type content, and obtain the fuzzy information based on the abnormal type content; and is used to assign a morphological mask value to each normal pixel point coordinate and each abnormal pixel point coordinate of the inspection image; respectively construct a first positive mask value distribution map and a first negative mask value distribution map according to the morphological mask values of all the normal pixel point coordinates, and respectively construct a second positive mask value distribution map and a second negative mask value distribution map according to the morphological mask values of all the abnormal pixel point coordinates; determine the recognizable range by using the first positive mask value distribution map and the second negative mask value distribution map, and determine the unrecognizable range by using the first negative mask value distribution map and the second positive mask value distribution map; and is used to determine the first coverage area of the recognizable range, extract features from the mask value distribution map of this first coverage area, detect and identify the extracted features to obtain the covered target content; determine the second coverage area of the unrecognizable range and the intersection area where the unrecognizable range overlaps with the recognizable range; perform fuzziness quantification on the mask value distribution within the second coverage area, perform cluster analysis on all the quantified fuzziness information to generate a fuzziness visualization distribution map; perform preliminary screening of abnormal types on this fuzziness visualization distribution map to obtain a preliminary abnormal screening result; determine the first coverage area associated with the intersection area, and rescreen the preliminary abnormal screening result according to the associated first coverage area to obtain the abnormal type content; is also used to obtain the format conversion information of the preliminary screening modality data; perform correlation analysis between the format conversion information and the mask value features within the first coverage area to obtain a first feature matching correlation degree, and calculate the first pairing result according to this first feature matching correlation degree; perform correlation analysis between the format conversion information and the mask value features within the second coverage area to obtain a second feature matching correlation degree, and calculate the second pairing result according to this second feature matching correlation degree;After obtaining the second feature matching correlation degree, it further includes: performing correlation analysis on the format conversion information and the mask value features in the intersection area to obtain a third feature matching correlation degree, calculating the second pairing result according to the second feature matching correlation degree and the third feature matching correlation degree; and is also used to determine a first weight and a second weight; comparing the difference between the second feature matching correlation degree and the third feature matching correlation degree, and adjusting the first weight and / or the second weight according to this difference; assigning the adjusted first weight to the second feature matching correlation degree and / or the second weight to the third feature matching correlation degree.;
[0090] A third processing unit 650, which is used to judge whether the preliminary screening modal data is used as the fusion modal data based on the first pairing result and the second pairing result, and perform associated annotation processing on the fusion modal data and the inspection image, where the associated annotation processing includes annotating the associated area, annotating the associated object, and annotating the associated form; in some embodiments, the third processing unit 650 is further used to determine the position information of the associated area between the inspection image and the fusion modal data; locking the associated objects of the inspection image and the fusion modal data respectively based on the position information of the associated area; judging the type of the associated form based on the associated objects of the inspection image and the fusion modal data; and is used to calculate the similarity of any associated object between the inspection image and the fusion modal data, merge the similarities of all the associated objects of the inspection image and the fusion modal data to obtain a total similarity, determine the type of the associated form based on interval comparison of this total similarity; and is used to trace the second pairing result of the fusion modal data, calculate the associated adjustment parameter determined by this second pairing result, assign the associated adjustment parameter to the total similarity to obtain an adjusted similarity, and re-determine the type of the associated form based on this adjusted similarity.
[0091] In the above embodiments, for the more specific working processes of each functional unit, reference may be made to the corresponding content disclosed in the foregoing method embodiments. In addition, each functional unit may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0092] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0093] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 steps for the functions specified in one block or multiple blocks.
[0095] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A method for image quality classification based on a multimodal large model, characterized in that: The steps include: Identify the target attributes of the inspection image, perform preliminary analysis on the inspection image based on the target attributes, and determine a modal data matching strategy for the inspection image; Based on the modal data matching strategy, a related modal data set is selected from the inspection database, and the related modal data set is classified into types to obtain a plurality of related modal data; Performing a matching judgment between each of the associated modal data and the inspection image to obtain a matching degree of the associated modal data, and using the associated modal data with a matching degree higher than a preset matching value as the primary screening modal data, wherein the matching judgment includes inspection time judgment, inspection object judgment and inspection location judgment; Delineating a recognizable range and an unrecognizable range of the inspection image, extracting valid information according to the recognizable range, extracting fuzzy information according to the unrecognizable range, and pairing each of the primary screening modal data with the valid information and the fuzzy information, respectively, to obtain a first pairing result and a second pairing result; Based on the first pairing result and the second pairing result, it is determined whether the initial screening modal data is used as fused modal data, and the fused modal data is associated with the inspection image and annotated, wherein the associated annotation processing includes annotating associated areas, annotating associated objects, and annotating associated forms.
2. The image quality classification processing method based on multimodal large model according to claim 1 is characterized in that: Defining the recognizable range and the unrecognizable range of the inspection image, extracting valid information according to the recognizable range, and extracting fuzzy information according to the unrecognizable range includes the following steps: The training model is used to analyze and judge the pixel quality of the inspection image, and the normal pixel coordinates and the abnormal pixel coordinates are identified; the recognizable range is constructed using all the normal pixel coordinates, and the unrecognizable range is constructed using all the abnormal pixel coordinates; the pixels within the recognizable range are preliminarily identified and the covered target content is obtained, and the valid information is obtained based on the covered target content; the pixels within the unrecognizable range are preliminarily identified to obtain the abnormal type content, and the fuzzy information is obtained based on the abnormal type content.
3. The image quality classification processing method based on multimodal large model according to claim 2 is characterized in that: The steps of constructing the recognizable range by using the coordinates of all the normal pixel points and constructing the unrecognizable range by using the coordinates of all the abnormal pixel points include: Assigning morphological mask values to each normal pixel coordinate and each abnormal pixel coordinate of the inspection image; constructing a first positive mask value distribution map and a first reverse mask value distribution map according to the morphological mask values of all normal pixel coordinates, and constructing a second positive mask value distribution map and a second reverse mask value distribution map according to the morphological mask values of all abnormal pixel coordinates; The recognizable range is determined using the first positive mask value distribution map and the second inverse mask value distribution map, and the unrecognizable range is determined using the first inverse mask value distribution map and the second positive mask value distribution map.
4. The image quality classification processing method based on multimodal large model according to claim 3 is characterized in that: The preliminary identification of the pixels within the identifiable range and obtaining the covered target content comprises the following steps: determining a first coverage area of the identifiable range, performing feature extraction on a mask value distribution map of the first coverage area, detecting and identifying the extracted features, and obtaining the covered target content; The step of performing preliminary identification on the pixels within the unrecognizable range to obtain abnormal type content comprises the following steps: determining a second coverage area of the unrecognizable range and an intersection area where the unrecognizable range overlaps with the recognizable range; Fuzzy quantification is performed on the mask value distribution within the second coverage area, and cluster analysis is performed on all quantified fuzziness information to generate a fuzziness visualization distribution map; the fuzziness visualization distribution map is preliminarily screened for abnormal types to obtain a preliminarily screened abnormality result; A first coverage area associated with the intersection area is determined, and the abnormal initial screening result is rescreened according to the associated first coverage area to obtain the abnormal type content.
5. The image quality classification processing method based on multimodal large model according to claim 4 is characterized in that: The step of pairing each of the primary screening modal data with the valid information and the fuzzy information to obtain a first pairing result and a second pairing result comprises the following steps: Acquire format conversion information of the primary screening modal data; perform correlation analysis on the format conversion information and the mask value feature in the first coverage area to obtain a first feature matching correlation degree, and calculate the first pairing result according to the first feature matching correlation degree; perform correlation analysis on the format conversion information and the mask value feature in the second coverage area to obtain a second feature matching correlation degree, and calculate the second pairing result according to the second feature matching correlation degree; Among them, after obtaining the second feature matching correlation degree, it also includes: performing correlation analysis on the format conversion information and the mask value feature in the intersection area to obtain a third feature matching correlation degree, and calculating the second pairing result according to the second feature matching correlation degree and the third feature matching correlation degree.
6. The image quality classification processing method based on multimodal large model according to claim 5 is characterized in that: The step of calculating the second pairing result based on the second feature matching correlation degree and the third feature matching correlation degree includes the following steps: determining a first weight and a second weight; comparing the difference between the second feature matching correlation degree and the third feature matching correlation degree, and adjusting the first weight and / or the second weight based on the difference; assigning the adjusted first weight to the second feature matching correlation degree and / or assigning the second weight to the third feature matching correlation degree.
7. The image quality classification processing method based on multimodal large model according to claim 1 is characterized in that: The process of associating and annotating the fused modality data with the inspection image comprises the following steps: Determine the associated area position information of the inspection image and the fused modal data; based on the associated area position information, respectively lock the associated objects of the inspection image and the fused modal data; based on the associated objects of the inspection image and the fused modal data, determine the type of the associated form; wherein the type of the associated form includes one of full association, strong association and weak association.
8. The image quality classification processing method based on multimodal large model according to claim 7 is characterized in that: The method of determining the type of association form based on the associated objects of the inspection image and the fused modal data includes the following steps: calculating the similarity between the inspection image and any associated object of the fused modal data, merging the similarities between the inspection image and all associated objects of the fused modal data to obtain a total similarity, performing interval comparison based on the total similarity, and determining the type of association form.
9. The image quality classification processing method based on multimodal large model according to claim 8 is characterized in that: The method also includes the step of adjusting the type of the association form: tracing back the second pairing result of the fused modal data, calculating the second pairing result to determine an association adjustment parameter, assigning the association adjustment parameter to the total similarity to obtain an adjusted similarity, and redetermining the type of the association form based on the adjusted similarity.
10. An image quality classification processing system based on a multimodal large model, characterized in that: include: A first recognition unit, which is used to recognize the target attribute of the inspection image, perform a preliminary analysis on the inspection image based on the target attribute, and determine a modal data matching strategy for the inspection image; A second identification unit, which is used to screen a related modal data set from the inspection database based on the modal data matching strategy, classify the related modal data set into types, and obtain a plurality of related modal data; A first processing unit, which is used to perform matching judgment between each type of the associated modal data and the inspection image, obtain the matching degree of the associated modal data, and use the associated modal data with a matching degree higher than a preset matching value as the primary screening modal data, wherein the matching judgment includes inspection time judgment, inspection object judgment and inspection location judgment; a second processing unit, which is used to define a recognizable range and an unrecognizable range of the inspection image, extract valid information according to the recognizable range, extract fuzzy information according to the unrecognizable range, and pair each of the primary screening modal data with the valid information and the fuzzy information, respectively, to obtain a first pairing result and a second pairing result; A third processing unit is used to determine whether the initial screening modal data is used as fused modal data based on the first pairing result and the second pairing result, and to associate and annotate the fused modal data with the inspection image, wherein the association annotation processing includes annotating associated areas, annotating associated objects, and annotating associated forms.
Citation Information
Patent Citations
Cross-modality image-label relevance learning method facing social image
CN104899253A
Data processing method, computer equipment and readable storage medium
CN113642536A