Multi-model collaborative image detection method, device and equipment and storage medium

By using a first network model with a smaller structure and scale in multi-model collaborative image detection for preliminary detection, and using a larger model for verification when the confidence is low, combined with the same feature space processing, the problems of high resource consumption and low accuracy are solved, and efficient and accurate image detection is achieved.

CN120388165APending Publication Date: 2025-07-29TRANSWARP TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510475383.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

There are problems of high resource consumption and low accuracy of collaborative processing in existing multi-model collaborative image detection, especially in complex environments such as industrial quality inspection and medical diagnosis.

Method used

The first network model with a smaller structure and a smaller scale is used for preliminary detection, and the result is directly determined when the confidence is higher than the threshold. When the confidence is lower than the threshold, the second network model with a larger structure and a larger scale is used for verification processing, and the coordination of multiple models is ensured through the setting of the same feature space.

Benefits of technology

It effectively reduces the resource consumption of image detection, improves the accuracy and reliability of detection results, and avoids the difficulty of collaborative processing caused by inconsistent data format and update frequency among multiple models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388165A_ABST
    Figure CN120388165A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-model collaborative image detection method and device, equipment and a storage medium. The method comprises the steps of obtaining a to-be-detected image, inputting the to-be-detected image as input data into a first network model, and obtaining a first detection result output by the first network model; if the confidence coefficient of the first detection result is greater than or equal to a confidence coefficient threshold value, determining the first detection result as a target detection result of the to-be-detected image; if the first detection result is smaller than a confidence coefficient threshold value, checking the first detection result by adopting a second network model, and determining a target detection result of the to-be-detected image according to an output checking result; the first network model and the second network model have different network structures and network parameters, and the first network model and the second network model adopt the same feature space. By using the method, dynamic routing collaboration of multiple models is realized, resource consumption of image detection is reduced, and inference energy consumption in multi-model processing is also reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer application technologies, and particularly to an image detection method, device, equipment and storage medium for multi-model collaboration. Background Art

[0002] In complex and variable actual application environments, such as industrial quality inspection and medical diagnosis scenarios, there are requirements for anomaly detection of the involved images, so as to effectively promote the relevant operations in various application fields through anomaly detection.

[0003] One of the existing anomaly detection implementation methods is to complete image detection through the collaboration of multiple network models. In the existing implementation of multi-model collaborative image detection, the main problems are as follows: 1) Multiple models are often used to detect images in parallel, that is, multiple models run in parallel, thus occupying more video memory space and greatly increasing the resource consumption of image detection. 2) There are differences in the input format, data scale, update frequency, etc. of the data used by each model, thus affecting the accuracy of collaborative processing between multiple models. Summary of the Invention

[0004] The present invention provides an image detection method, device, equipment and storage medium for multi-model collaboration, which effectively reduces the resource consumption of image detection.

[0005] According to a first aspect of the present invention, there is provided an image detection method for multi-model collaboration, including:

[0006] Obtain an image to be detected, and input the image to be detected as input data into a first network model, and obtain a first detection result output by the first network model;

[0007] If the confidence level of the first detection result is greater than or equal to the confidence level threshold, determine the first detection result as the target detection result of the image to be detected;

[0008] If the first detection result is less than the confidence level threshold, use a second network model to perform verification processing on the first detection result, and determine the output verification processing result as the target detection result of the image to be detected;

[0009] Wherein, the network structure scale of the first network model is smaller than that of the second network model and the network parameters are different, and at the same time, the first network model and the second network model use the same feature space for image detection.

[0010] According to a second aspect of the present invention, there is provided an image detection device for multi-model collaboration, including:

[0011] An acquisition module, configured to acquire an image to be detected, input the image to be detected as input data into a first network model, and obtain a first detection result output by the first network model;

[0012] A first execution module, configured to, when the confidence of the first detection result is greater than or equal to the confidence threshold, determine the first detection result as the target detection result of the image to be detected;

[0013] A second execution module, configured to, when the first detection result is less than the confidence threshold, perform a verification process on the first detection result by using a second network model, and determine the output verification result as the target detection result of the image to be detected;

[0014] Wherein, the network structure scale of the first network model is smaller than that of the second network model and the network parameters are different, and at the same time, the first network model and the second network model perform image detection by using the same feature space.

[0015] According to a third aspect of the present invention, there is provided an electronic device, where the electronic device includes:

[0016] At least one processor; and a memory communicatively connected to the at least one processor;

[0017] Wherein, the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor can execute the multi-model collaborative image detection method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the multi-model collaborative image detection method according to any embodiment of the present invention is implemented.

[0019] In the technical solution of the embodiment of the present invention, first, an image to be detected is obtained, and the image to be detected is input as input data into a first network model, and a first detection result output by the first network model is obtained; then, when the confidence of the first detection result is greater than or equal to the confidence threshold, the first detection result is determined as the target detection result of the image to be detected; when the first detection result is less than the confidence threshold, the second network model is used to perform a verification process on the first detection result, and the output verification process result is determined as the target detection result of the image to be detected; wherein, the network structure scale of the first network model is smaller than that of the second network model and the network parameters are different, and at the same time, the first network model and the second network model use the same feature space for image detection. In the above technical solution of this embodiment, for the image to be detected, first consider using the first network model with relatively small structure and scale to process, and when the confidence of the detection result is higher than the set threshold, directly use the detection result as the final detection result. Then, for the confidence lower than the set threshold, consider using the second network model with relatively large structure and scale to detect the image to be detected again, and determine the detection result as the final detection result. This image detection method better realizes the dynamic collaborative detection of multiple models, that is, it ensures the accuracy of the image detection result, effectively reduces the resource consumption of image detection implementation and reduces the inference energy consumption in network model processing. At the same time, in the above technical solution of this embodiment, by setting the same feature space for multiple models, the unity of data features involved in multiple model processing is ensured, thereby effectively avoiding the difficulties brought by inconsistencies in data format, data scale, update frequency, etc. between multiple models to collaborative processing. The technical solution of this embodiment better ensures the adaptability of multiple model collaboration.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 is a flowchart of an image detection method with multi-model collaboration provided according to an embodiment of the present invention;

[0023] Figure 2It is a schematic structural diagram of an image detection device with multi - model collaboration according to an embodiment of the present invention;

[0024] Figure 3 It is a schematic structural diagram of an electronic device for implementing an image detection method with multi - model collaboration according to an embodiment of the present invention. Detailed implementation manners

[0025] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above - mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] Figure 1 It is a flowchart of an image detection method with multi - model collaboration according to an embodiment of the present invention. This embodiment is applicable to the situation of image detection. This method can be executed by an image detection device with multi - model collaboration. The image detection device with multi - model collaboration can be implemented in the form of hardware and / or software, and the image detection device with multi - model collaboration can be configured in a computer device. As Figure 1 shown, the method includes:

[0028] S101. Obtain an image to be detected, and input the image to be detected as input data into a first network model, and obtain a first detection result output by the first network model.

[0029] In this embodiment, one of the applicable image detection scenarios can be considered as the detection of abnormal objects in an image. For example, in the industrial field, images containing corresponding industrial devices can be used for quality inspection of industrial devices, or in the medical field, images containing lesions can be used for detecting lesion areas, etc. In this embodiment, the image that needs to be detected is recorded as the image to be detected.

[0030] In this embodiment, the specific implementation of image detection can be carried out through deep integration with technologies such as visual perception and knowledge graph for intelligent detection. Technologies such as visual perception and knowledge graph can be specifically reflected through the application of network models. To better ensure the image detection results, this embodiment considers using the method of multi-model collaboration to participate in the detection of image anomalies.

[0031] In this embodiment, a first network model is first introduced to process the image to be detected. The first network model can be a single forward propagation type network model with a relatively simple network structure and a relatively small scale (such as the YOLOv10 model). The first network model can achieve rapid detection and screening of the image to be detected, and thus quickly obtain the detection result of the image. The detection result obtained in this embodiment is recorded as the first detection result, and this first detection result can be used as the basis data for subsequent determination.

[0032] In an implementation of image detection, this step can preprocess the image to be detected, such as adjusting the image size and matching the image resolution, adjusting the image brightness, contrast, and saturation, etc., and can also perform random perturbation on the image to be detected through the HSV color space. Then, the preprocessed image to be detected is input into the first network model. The first network model detects the abnormal area in the image through a single forward propagation processing method, and finally can output the determined abnormal area in the form of a detection box, class label, and corresponding confidence level, forming a detection table composed of a detection box, class label, and corresponding confidence level. The formed detection table can be used as the first detection result.

[0033] S102. If the confidence level of the first detection result is greater than or equal to the confidence level threshold, then determine the first detection result as the target detection result of the image to be detected.

[0034] In this embodiment, a confidence level threshold for evaluating the reliability of the first detection result can also be obtained. At the same time, one way to determine the confidence level of the first detection result can be to sum and average the confidence levels corresponding to each detection box that makes up the first detection result, and the average value can be used as the confidence level of the first detection result for comparison with the confidence level threshold; another way can be to regard the confidence levels corresponding to each detection box in the first detection result as the confidence level of the first detection result, and the confidence levels corresponding to each detection box can be compared with the confidence level threshold respectively.

[0035] In this embodiment, when the comparison result obtained by the above confidence level comparison satisfies that the confidence level of the first detection result is greater than or equal to the confidence level threshold, the relevant logic of this step can be executed. Thus, the first detection result determined in the above step can be directly determined as the final target detection result of the image to be detected. Among them, assuming that the confidence level of the first detection result is the average value of the confidence levels of all detection frames, the execution condition of this step is satisfied when the average value is greater than or equal to the confidence level threshold; assuming that the confidence level of the first detection result is the confidence levels of all detection frames, the execution condition of this step is satisfied when all confidence levels are greater than or equal to the confidence level threshold.

[0036] It can be known that the target detection result determined in this step can be presented to relevant personnel in a clearer and more structured manner. For example, a dictionary of detection results can be constructed, which contains three key-value pairs: detection frame coordinates, class labels, and confidence levels, and the results can be output in a structured format, such as JSON format. In addition, the determined target detection result can also be directly presented on the image to be detected. For example, the detection frame is drawn on the image to be detected according to the detection frame coordinates, and the class label and confidence level are marked.

[0037] S103. If the first detection result is less than the confidence level threshold, the second network model is used to perform verification processing on the first detection result, and the output verification processing result is determined as the target detection result of the image to be detected.

[0038] In this embodiment, when the comparison result obtained by the above confidence level comparison satisfies that the confidence level of the first detection result is less than the confidence level threshold, the relevant logic of this step can be specified. At this time, it can be considered that the accuracy of the result of image detection by the above first network model is limited, and the first detection result needs to be further verified. This embodiment introduces a network structure with multi-modal input and a second network model with a relatively complex scale. It can be known that the network structure scale of the first network model in this embodiment is smaller than that of the second network model and the network parameters are different. At the same time, the first network model and the second network model use the same feature space for image detection.

[0039] In this embodiment, as a network model adopted, the second network model may include a segmentation network structure for image grid segmentation, a multi-layer perception network structure, and a language network structure model, so as to perform a more in-depth analysis of the first detection result of the image to be detected through the second network model with a more refined network structure.

[0040] Exemplarily, an image can be segmented into multiple grid regions of different sizes through image grid segmentation, and local and global features of the image can be extracted for each grid region. Then, the extracted features can be non-linearly transformed through a multi-layer perceptron network layer to enhance the expression ability of the extracted features. Then, a language network model can be used to combine the pre-given description text and the previously determined feature information to determine detection targets (such as abnormal regions or abnormal components) from the image to be detected, and provide detailed descriptions and processing suggestions for the detection targets, etc.

[0041] In this embodiment, it can be considered that the second network model finally outputs the detailed descriptions and related suggestion information of each detection target determined from the image to be detected. The output information can show the verification processing result output by the second network model, and this verification processing result can be used as the target detection result of the image to be detected.

[0042] In this embodiment, alignment processing of the same feature space is performed on the first network model and the second network model in advance, so as to ensure that the output of the first network model can be received and processed by the second network model without obstacles. This step can use the first detection result as the input of the second network model. Since the first network model and the second network model have the same feature control, the first detection result can be adapted to the input of the second network model. Based on the first detection result, the second network model combines multi-modal information to perform more comprehensive and detailed semantic understanding and analysis of the image, and excavates potential abnormal information hidden in the image, thereby improving the accuracy and reliability of the detection result.

[0043] In the above technical solution of this embodiment, for the image to be detected, first consider using the first network model with relatively small structure and scale to process. When the confidence of the detection result is higher than the set threshold, directly use the detection result as the final detection result. Then, for the confidence lower than the set threshold, consider using the second network model with relatively large structure and scale to detect the image to be detected again, and determine the detection result as the final detection result. This image detection method better realizes the dynamic collaborative detection of multiple models, that is, it ensures the accuracy of the image detection result, and also effectively reduces the resource consumption of image detection implementation and reduces the inference energy consumption in network model processing. At the same time, in the above technical solution of this embodiment, by setting the same feature space for multiple models, the unity of data features involved in multi-model processing is ensured, thereby effectively avoiding the difficulties brought by inconsistencies in data format, data scale, and update frequency, etc. among multiple models to collaborative processing. The technical solution of this embodiment better ensures the adaptability of multi-model collaboration.

[0044] As a first alternative embodiment of this embodiment, on the basis of the above embodiment, the verification process of the first detection result using the second network model can be further specified as the following steps:

[0045] a1) Input the first detection result and the image to be detected into the second network model.

[0046] In this embodiment, the data information included in the first detection result can be directly used as the input data of the image to be detected. Since the second network model and the first network model have been pre-processed for feature space alignment, the first detection result can be directly input into the second network model without secondary conversion.

[0047] b1) Through the second network model, determine the detection regions corresponding to the detection data in the first detection result in the image to be detected, and extract the image features of each detection region.

[0048] In this embodiment, taking the second network model including network modules for image segmentation, multi-layer perception, and semantic understanding as an example, this step can first segment the image containing the first detection result according to the detection boxes in the first detection result, so as to determine each segmented detection box region as a detection region. Then, the feature extraction of each detection region can also be realized through the network module of this image segmentation, thereby obtaining corresponding image features for each detection region.

[0049] c1) Generate detection description data for each detection region according to the image features and the given text description information, and use each detection description data as the verification result.

[0050] In this embodiment, the image features obtained in the above steps can be processed again through the multi-layer perception network module in the second network model, so as to obtain image features with enhanced features after non-linear transformation of the image features corresponding to each detection region. The enhanced image features are then passed through the language network module. By also inputting the pre-written text description information into the language network module, the language network model combines the pre-constructed knowledge graph, etc. for more accurate semantic analysis, so as to determine the detection description data for each detection region.

[0051] In this embodiment, the detection description data can be specifically used to describe in detail the content included in the detection region. Through the detection description data, it can be further determined whether the object to be detected is included in the detection region. Therefore, the detection description data corresponding to the detection region can be directly determined as the verification result.

[0052] The technical solution described above in this embodiment provides a specific implementation for using a second network model to conduct in-depth verification analysis on low-confidence first detection results. This technical solution is equivalent to combining multimodal information to conduct a more comprehensive and detailed semantic understanding and analysis of the image, thereby unearthing potential abnormal information hidden in the image and improving the accuracy and reliability of the detection results.

[0053] As a second optional embodiment of this embodiment, based on the above embodiment, the method further includes: performing feature comparative learning on the first network model and the second network model according to the sample image set, and controlling the first network model and the second network model to have the same feature space.

[0054] In this embodiment, to better support the collaborative processing of the first network model and the second network model, it is necessary to ensure in advance that the first network model and the second network model have the same feature space. In this embodiment, feature comparison learning training can be performed on the first network model and the second network model through this step. Through training, the network parameters and feature representation of each network model are adjusted to control the first network model and the second network model to achieve feature space alignment and have the same feature space.

[0055] Exemplarily, this embodiment can implement comparative learning of the first network model and the second network model through image detection processing of sample images in the sample image set. Feature information of the sample images in the sample image set processed by the first network model can be obtained, and feature information of the sample images in the sample image set processed by the second network model can also be obtained. Subsequently, comparative learning and alignment of the two obtained feature information can be performed to achieve the same feature space for the two network models.

[0056] Based on this second optional embodiment, as an implementation method, feature comparison learning can be performed on the first network model and the second network model based on the sample image set, and controlling the first network model and the second network model to have the same feature space can be specifically implemented as follows:

[0057] a2) Obtaining sample images from the sample image set, and inputting the sample images into the first network model and the second network model respectively.

[0058] In this embodiment, each sample image in the sample image set can be used as input data and input into the first network model and the second network model, and each sample image can be distinguished by an image identifier.

[0059] b2) forming a detection feature based on a detection result output by the first network model relative to the sample image.

[0060] In this embodiment, it can be considered that the first network model includes a feature extraction module. This step can effectively collect and save the important feature information generated by the feature extraction module. The obtained important feature information can be regarded as the expression of the detection result output by the first network model relative to the sample image at the feature layer, specifically including the position of the detection box, the category feature, and the feature expression of the corresponding confidence level. This step can present the above feature expression as a feature map, and can convert the feature map into a one-dimensional vector through global average pooling operation. The extracted one-dimensional vector can be used as the detection feature of the corresponding sample image in this step.

[0061] To facilitate better subsequent feature comparison and learning, this embodiment can also store the detection features corresponding to each sample image. The storage method can be the associated storage of the image identifier of the sample image and the determined detection feature in the database.

[0062] c2) According to the processing of the sample image by the second network model, extract the image semantic feature and the text semantic feature, and form a semantic feature according to the image semantic feature and the text semantic feature.

[0063] In this embodiment, it can also be considered that the second network model includes relevant network layers for feature extraction. Through the relevant network layers for feature extraction, valuable semantic information in the sample image can be extracted. Specifically, the extracted semantic information can be the image semantic feature and the text semantic feature extracted after the second network model processes the sample image. This step can merge the image semantic feature and the text semantic feature to form a feature vector representing the semantic feature. Different dimensions in the feature vector are used to represent the image semantic feature and the text semantic feature respectively.

[0064] Similarly, to facilitate better subsequent feature comparison and learning, this embodiment can also store the semantic feature of the sample image. The storage method can be the associated storage of the image representation of the sample image and the determined semantic feature in the database.

[0065] d2) According to the constructed contrastive learning framework, combine the detection feature and the semantic feature to adjust the model parameters of the first network model and the second network model.

[0066] In this embodiment, a contrastive learning framework can be pre-constructed. The contrastive learning framework may include a query encoder and a key encoder. The two encoders can be considered to have the same network architecture, and the main difference lies in the set network parameters. Under the contrastive learning framework, the detection features and semantic features of the sample image can be used as input data, and the two encoders in the contrastive learning framework are used to achieve the set learning goal. The learning goal can be to maximize the similarity between the detection features and semantic features of the same sample image, and minimize the similarity between the detection features of different sample images and the similarity between the semantic features.

[0067] In this embodiment, the constructed contrastive learning framework is equivalent to using the contrastive learning algorithm to align the features provided by the first network model and the second network model effectively. During the process of achieving the learning goal, the contrastive learning framework can adjust the model parameters involved in the encoders in the contrastive learning framework through the set loss function, optimizer, etc. The adjusted contrastive learning framework is equivalent to realizing the mapping of the features of the first network model and the second network model to the same feature space.

[0068] e2) Return and re-execute the operation of obtaining and inputting the sample image until the first network model and the second network model are in the same feature space, and obtain the same knowledge sharing library formed by the first network model and the second network model in the same feature space.

[0069] It should be noted that the above steps in this embodiment are a process of iterative loop learning. Each iteration can ensure that the feature representations of the first network model and the second network model are closer to being in the same feature space. In the next iteration implementation, the operation of obtaining the sample image and inputting it to the first network model and the second network model can be re-executed, thereby triggering the adjustment of the model parameters in the contrastive learning framework. Thus, the loop iterates until finally the features of the first network model and the second network model are mapped to the same feature space.

[0070] At the same time, it can be known that for the first network model and the second network model in the same feature space, the detection features and semantic features obtained by processing the sample image can be used as the knowledge reserve of the network model, constituting a unified knowledge sharing library for practical applications in image detection.

[0071] The above technical solution of this embodiment realizes the feature space object of the first network model and the second network model, and provides basic support for the cross-model barrier-free processing of the features of the first network model and the second network model in image detection processing.

[0072] As a third alternative embodiment of this embodiment, based on the above embodiment, the method may further include: generating a model optimization label according to the verification processing result output by the second network model, and optimizing the first network model according to the first detection result and the model optimization label.

[0073] In this embodiment, the first network model can be considered as a model that can be optimized in real time. Considering that the output result of the second network model is better than the output result of the first network model, therefore, this embodiment proposes a technical implementation for optimizing the first network model based on the output result of the second network model. And the optimization of the first network model can be specifically achieved through the execution logic provided in this embodiment.

[0074] Specifically, through the newly added steps of this embodiment, the verification processing result after the second network model verifies the first detection result can be obtained, and a model optimization label for optimizing the first network model can be generated based on the verification processing result. After that, the network parameters of the first network model can be adjusted again by combining the model optimization label with a set loss function, and finally, the first network model obtained after the iteration ends is used as the optimized network model.

[0075] It should be noted that this embodiment can store the network models before and after optimizing the first network model in actual application to form a backup model set of the first network model.

[0076] Based on this third alternative embodiment, one implementation manner of model optimization can specifically optimize generating a model optimization label according to the verification processing result output by the second network model, and optimizing the first network model according to the first detection result and the model optimization label as follows:

[0077] a3) Obtain the second detection result included in the verification processing result, and obtain the pre-set mapping relationship between models.

[0078] In this embodiment, through the description of the processing process of verifying the first detection result by the above second network model, it can be known that the verification processing result output by the second network model includes a depth analysis detection of the detection target in the image to be detected. Thus, it is equivalent to obtaining the second detection result of the second network model for processing the image to be detected, and the second detection result may include the detection description information of the detection target in the image to be detected.

[0079] It can be known that the representation methods of the first detection result involved in the first network model and the second detection result involved in the second network model are different. However, considering that both are descriptions of the detection target in the image to be detected, the difference in the model data results between the models can be obtained, and the mapping relationship of the output results between the models can be established according to the difference, so as to form the mapping relationship between the models. The mapping relationship between the models can be determined in advance, and this step can directly obtain the preset mapping relationship between the models.

[0080] b3) According to the mapping relationship between the models, convert the second detection result into a model optimization label corresponding to the first network model.

[0081] In this embodiment, the second detection result can be converted into a second detection result represented in the output result style of the first network model through the mapping relationship between the models. The converted second detection result can be considered to be also represented in the form of a detection box, a class label, and a confidence level. In this embodiment, the converted second detection result can be determined as the model optimization label corresponding to the first network model.

[0082] c3) According to the model optimization label and the first detection result, combined with a preset loss function, determine the loss function value.

[0083] In this embodiment, the model optimization label can be considered as a detection result with relatively high accuracy. The first detection result can be considered as the detection result obtained before the model optimization of the first network model, and its accuracy is generally lower than that of the model optimization label, which is equivalent to a difference between the first detection result and the model optimization label.

[0084] In this embodiment, the difference between the first detection result output by the first network model and the model optimization label can be measured by the set loss function, and the difference can be reflected by the loss function value.

[0085] d3) According to the loss function value, use the set gradient descent method to adjust the network parameters of the first network model, and determine the adjusted first network model as the optimized first network model.

[0086] In this embodiment, after determining the loss function value through the above steps, in order to further optimize the first network model, the set gradient descent method can be used to adjust the network parameters of the first network model, and the adjusted first network model can be recorded as the optimized first network model.

[0087] It can be understood that the above steps of this embodiment can be performed after both the first network model and the second network model have participated in the detection and processing of the image to be detected, and each time the output result of the second network model can be used to optimize the first network model to achieve real-time optimization of the first network model in practical applications, thereby ensuring the output accuracy of the first network model, and gradually increasing the proportion of the confidence level of the first detection result output by the first network model that meets the confidence level threshold, thereby effectively reducing the resource consumption and inference energy consumption brought by the participation of the second network model in the calculation.

[0088] As a fourth alternative embodiment of this embodiment, on the basis of the above embodiment, the method may further include:

[0089] a4) When it is detected that the current model evaluation condition is met, the performance evaluation information of the first network model is determined according to the set model evaluation script in combination with the verification image set.

[0090] In this embodiment, during the actual application of the first network model, the performance of the first network model is also periodically evaluated to monitor the running state of the first network model to ensure the processing performance of the first network model. This step can detect the interval duration from the end of the previous model evaluation to the current moment, and when it is monitored that the duration from the end time of the previous model evaluation to the current moment meets the set periodic duration, it is considered that the model evaluation condition is met.

[0091] In this embodiment, the model evaluation script can be regarded as a script executed for evaluating and calculating the first network model. Through this model evaluation script, the performance indicators of the first network model can be determined during the process of the first network model performing image detection on images. The performance indicators can include recall rate and average precision rate, etc. The error rate of the first network model in image detection can be characterized by the index values of the performance indicators, and the false detection and missed detection situations of the first network model can also be characterized.

[0092] It can be known that in this embodiment, the performance of the first network model is evaluated during the process of the first network model performing image detection on the verification images in the verification image set. Among them, the characteristics of the verification images in the verification image set can be that they contain the image to be detected and the correct detection information of the detection objects existing in the image to be detected, such as the detection frames representing the detection objects and the class labels, etc. The correct detection information it has can be used as the evaluation basis for the performance indicators in the performance evaluation.

[0093] In this embodiment, through this step, the performance evaluation information possessed by the first network model after a performance evaluation can be obtained, and this performance evaluation information can be saved in the form of a log.

[0094] b4) When it is determined according to the performance evaluation information that the first network model meets the switching condition, determine a target network model from the model set to replace the first network model.

[0095] In this embodiment, the average precision rate and recall rate of the first network model can be obtained from the performance evaluation information. One switching determination method can be to compare the result of the current performance evaluation with the result of the previous performance evaluation. If the current performance evaluation result is lower than the previous one, it can be considered that the performance of the first network model has declined and the model switching condition is met.

[0096] In this embodiment, after the first network model meets the switching condition, this step can switch the first network model. Specifically, a target network model can be selected from the model set to switch the first network model. Among them, the models in the model set can all be considered as models with the same network structure as the first network model, and the models included can be considered as models with adjusted network parameters obtained by optimizing the first network model in actual applications. The models in the model set can be recorded by different version numbers.

[0097] In this embodiment, according to the above description, the performance evaluation information recorded for each first network model in the model set can be obtained, and the target network model with the highest performance index value can be selected as the first network model.

[0098] The above technical solution of this embodiment can switch the first network model when it is monitored that the performance of the first network model declines or a failure occurs, so as to switch the first network model to the standby model with the best performance. This technical solution better realizes the real-time monitoring of the health status of the network model and ensures the effectiveness of image detection.

[0099] As the fifth optional embodiment of this embodiment, on the basis of the above embodiment, the method may further include: determining a secondary verification result corresponding to the image to be detected according to the first detection result and the verification processing result, and determining the secondary verification result as the target detection result of the image to be detected.

[0100] In this embodiment, after the verification processing result is determined by the second network model, the verification processing result can also be verified twice to ensure the reliability and accuracy of the final output result. For the specific implementation of the secondary verification, it can compare the detection objects included in the first detection result and the verification processing result for consistency. If the determined detection objects are the same, it can be considered that the verification processing result and the first detection result have a higher reliability.

[0101] Continuing with the above description, if it is determined that the detection objects in the consistency comparison are inconsistent, this step can be used to further analyze the reasons for the differences. Specifically, a lightweight network model (such as MobileNet) can be used to quickly verify the conflicting results in the first detection result and the verification processing result. In this embodiment, the reasons for the differences between the first detection result and the verification processing result can be compared through a lightweight network model. It can quickly verify the conflicting results through the form of large model fine-tuning, and use the output result of the verification as the secondary verification result.

[0102] In this embodiment, the secondary verification result includes the detection information that can optimize the detection result of the image to be detected and is verified correctly. Therefore, the secondary verification result can be determined as the final target detection result of the image to be detected.

[0103] The above technical solution of this embodiment better ensures the reliability and accuracy of the detection results output by the image detection method provided in this embodiment.

[0104] To facilitate a better understanding of the multi-model collaborative image detection method provided in this embodiment, this embodiment describes the implementation of multi-model collaborative image detection through an example.

[0105] Exemplarily, the method provided in this embodiment can be divided into three stages.

[0106] Among them, the first stage of the method can be regarded as the basic detection stage of the image. The implementation steps of this stage can be described as:

[0107] S1. Input the image into a small-scale network model.

[0108] S2. The small-scale network model performs anomaly detection on the image to obtain the first detection result.

[0109] S3. Whether the confidence level of the first detection result is less than the confidence threshold. If yes, execute S4; if not, execute S5.

[0110] S4. Output and display the first detection result as the final detection result.

[0111] S5. Use the large-scale network model to verify the first detection result, obtain the verification detection result, and output and display the verification detection result as the final detection result.

[0112] At the same time, the second stage of the method can be regarded as the stage of unifying the feature spaces of the small-scale network model and the large-scale network model to form a unified knowledge sharing library, and the stage of spatial alignment and model optimization for optimizing the small-scale network model based on the detection results of the large-scale network model. The implementation steps of this stage can be described as:

[0113] S6. Obtain the detection features stored in the relatively small-scale network model and the semantic features stored in the relatively large-scale network model.

[0114] S7. Map the features of the small-scale network model and the large-scale network model to the same feature space by combining the detection features, semantic features, and a contrastive learning framework.

[0115] S8. Convert the detection features and semantic features into the same feature space to form knowledge sharing information, and use the knowledge sharing information to optimize the small-scale network model.

[0116] S9. Optimize the small-scale network model according to the verification processing result of the large-scale network model, the first detection result of the small-scale network model, the knowledge sharing information, and the loss function to form an updated small-scale network model.

[0117] Among them, the verification processing result is used as the label content in the optimization of the small-scale network model.

[0118] In addition, the third stage of this method can be regarded as the monitoring switching of the network model and the fault tolerance processing stage of the model. The implementation steps of this stage can be described as:

[0119] S10. When it is detected that the current moment meets the model evaluation conditions, determine the performance evaluation information of the currently used small-scale network model according to the set model evaluation script and the verification image set.

[0120] S11. When it is determined that the performance of the currently used small-scale network model has decreased according to the performance evaluation information, determine that the model switching conditions are met, and select a new small-scale network model from the set of standby models.

[0121] Among them, the set of standby models includes multiple historical versions of the small-scale network model.

[0122] S12. Determine the secondary verification result corresponding to the input image according to the first detection result and the verification processing result, and determine the secondary verification result as the final detection result of the input image.

[0123] The basic solution implemented by the method provided in this embodiment has better improvements in terms of computing resource consumption, joint training and model optimization among multiple models, and model monitoring and fault tolerance methods compared with the prior art.

[0124] Embodiment 3

[0125] Figure 2 It is a schematic structural diagram of an image detection device for multi-model collaboration provided according to Embodiment 3 of the present invention. As Figure 2As shown in the figure, the device includes: an acquisition module 21, a first execution module 22, and a second execution module 23.

[0126] The acquisition module 21 is configured to acquire an image to be detected, input the image to be detected as input data into a first network model, and obtain a first detection result output by the first network model.

[0127] The first execution module 22 is configured to, when the confidence of the first detection result is greater than or equal to the confidence threshold, determine the first detection result as the target detection result of the image to be detected.

[0128] The second execution module 23 is configured to, when the first detection result is less than the confidence threshold, perform a verification process on the first detection result by using a second network model, and determine the output verification process result as the target detection result of the image to be detected.

[0129] Wherein, the network structure scale of the first network model is smaller than that of the second network model and the network parameters are different. At the same time, the first network model and the second network model perform image detection using the same feature space.

[0130] An image detection device with multi-model collaboration provided in this embodiment first considers using a first network model with relatively small structure and scale to process the image to be detected, and when the confidence of the detection result is higher than the set threshold, directly uses the detection result as the final detection result. Then, for the case where the confidence is lower than the set threshold, it further considers using a second network model with relatively large structure and scale to detect the image to be detected again, and determines the detection result as the final detection result. This image detection method better realizes the dynamic collaborative detection of multiple models, which not only ensures the accuracy of the image detection result, but also effectively reduces the resource consumption of image detection implementation and reduces the inference energy consumption in network model processing. At the same time, in the above technical solution of this embodiment, by setting the same feature space for multiple models, the unity of data features involved in multi-model processing is ensured, thereby effectively avoiding the difficulties brought by inconsistencies in data format, data scale, update frequency, etc. between multiple models to collaborative processing. The technical solution of this embodiment better ensures the adaptability of multi-model collaboration.

[0131] Further, the second execution module may specifically be configured to: input the first detection result and the image to be detected into the second network model.

[0132] Through the second network model, determine the detection regions corresponding to each detection data in the first detection result in the image to be detected, and extract the image features of each detection region.

[0133] Generate detection description data for each of the detection regions according to the respective image features and the given text description information, and use the detection description data as the verification processing result.

[0134] Further, the device further includes: a feature space learning module, configured to perform feature contrast learning on the first network model and the second network model according to a sample image set, and control the first network model and the second network model to have the same feature space.

[0135] Further, the feature space learning module may specifically be configured to:

[0136] Obtain sample images in the sample image set, and input the sample images into the first network model and the second network model respectively;

[0137] Form detection features according to the detection results output by the first network model relative to the sample images;

[0138] According to the processing of the sample images by the second network model, extract image semantic features and text semantic features, and form semantic features according to the image semantic features and the text semantic features;

[0139] According to the constructed contrast learning framework, combine the detection features and the semantic features to adjust the model parameters of the first network model and the second network model;

[0140] Return to re-execute the operation of obtaining and inputting the sample images until the first network model and the second network model are in the same feature space, and obtain the same knowledge sharing library formed by the first network model and the second network model in the same feature space.

[0141] Further, the device may further include: a model optimization module, configured to generate a model optimization label according to the verification processing result output by the second network model, and optimize the first network model according to the first detection result and the model optimization label.

[0142] Further, the model optimization module may specifically be configured to:

[0143] Obtain the second detection result included in the verification processing result, and obtain a preset mapping relationship between models;

[0144] According to the mapping relationship between models, convert the second detection result into a model optimization label corresponding to the first network model;

[0145] According to the model optimization label and the first detection result, combine a preset loss function to determine the loss function value;

[0146] According to the loss function value, the network parameters of the first network model are adjusted by using a set gradient descent method, and the adjusted first network model is determined as the optimized first network model.

[0147] Further, the device may further include:

[0148] A performance evaluation module, configured to determine the performance evaluation information of the first network model according to a set model evaluation script in combination with a verification image set when it is detected that the current model evaluation condition is satisfied;

[0149] A model switching module, configured to determine a target network model from a model set to replace the first network model when it is determined according to the performance evaluation information that the first network model meets the switching condition.

[0150] Further, the device may further include:

[0151] According to the first detection result and the verification processing result, determine the secondary verification result corresponding to the image to be detected, and determine the secondary verification result as the target detection result of the image to be detected.

[0152] The image detection device with multi-model collaboration provided by the embodiments of the present invention can execute the image detection method with multi-model collaboration provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0153] Figure 3 FIG. shows a schematic structural diagram of an electronic device 30 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processing, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0154] As Figure 3As shown, the electronic device 30 includes at least one processor 31 and a memory communicatively connected to the at least one processor 31, such as a read-only memory (ROM) 32, a random access memory (RAM) 33, etc. The memory stores a computer program executable by the at least one processor. The processor 31 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 32 or the computer program loaded from the storage unit 38 into the random access memory (RAM) 33. In the RAM 33, various programs and data required for the operation of the electronic device 30 can also be stored. The processor 31, the ROM 32, and the RAM 33 are connected to each other through a bus 31. The input / output (I / O) interface 35 is also connected to the bus 31.

[0155] Multiple components in the electronic device 30 are connected to the I / O interface 35, including: an input unit 36, such as a keyboard, a mouse, etc.; an output unit 37, such as various types of displays, speakers, etc.; a storage unit 38, such as a disk, an optical disc, etc.; and a communication unit 39, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 39 allows the electronic device 30 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0156] The processor 31 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 31 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 31 executes the various methods and processes described above, such as the multi-model collaborative image detection method.

[0157] In some embodiments, the multi-model collaborative image detection method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 38. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 30 via the ROM 32 and / or the communication unit 39. When the computer program is loaded into the RAM 33 and executed by the processor 31, one or more steps of the multi-model collaborative image detection method described above can be executed. Alternatively, in other embodiments, the processor 31 can be configured as the multi-model collaborative image detection method by any other appropriate means (e.g., by means of firmware).

[0158] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0159] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0160] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0161] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0162] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0163] A computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0164] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0165] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image detection method with multi-model collaboration, characterized in that, include: Acquire an image to be detected, input the image to be detected as input data into a first network model, and obtain a first detection result output by the first network model; If the confidence level of the first detection result is greater than or equal to the confidence level threshold, determining the first detection result as the target detection result of the image to be detected; If the first detection result is less than the confidence threshold, verifying the first detection result using a second network model, and determining the output verification result as the target detection result of the image to be detected; Among them, the network structure scale of the first network model is smaller than that of the second network model and the network parameters are different. At the same time, the first network model and the second network model use the same feature space for image detection.

2. The method according to claim 1, wherein The verifying the first detection result using the second network model includes: Inputting the first detection result and the image to be detected into the second network model; Determine, by means of the second network model, a detection area corresponding to each detection data in the first detection result in the image to be detected, and extract image features of each detection area; According to each of the image features and the given text description information, detection description data of each of the detection areas is generated, and each of the detection description data is used as a verification processing result.

3. The method according to claim 1, characterized in that Also includes: Perform feature comparison learning on the first network model and the second network model according to the sample image set, and control the first network model and the second network model to have the same feature space.

4. The method according to claim 3, characterized in that: The performing feature comparison learning on the first network model and the second network model according to the sample image set, and controlling the first network model and the second network model to have the same feature space, includes: Obtaining sample images from a sample image set, and inputting the sample images into the first network model and the second network model respectively; forming a detection feature according to a detection result output by the first network model relative to the sample image; Extracting image semantic features and text semantic features based on the processing of the sample image by the second network model, and forming semantic features based on the image semantic features and text semantic features; According to the constructed contrastive learning framework, in combination with the detection features and the semantic features, adjusting model parameters of the first network model and the second network model; The acquisition and input operations of the sample image are returned and re-executed until the first network model and the second network model are in the same feature space, and the same knowledge sharing library formed by the first network model and the second network model in the same feature space is obtained.

5. The method according to claim 1, wherein Also includes: A model optimization label is generated according to the verification processing result output by the second network model, and the first network model is optimized according to the first detection result and the model optimization label.

6. The method according to claim 5, characterized in that, Generating a model optimization label according to the verification processing result output by the second network model, and optimizing the first network model according to the first detection result and the model optimization label, includes: Obtaining the second detection result included in the verification processing result, and obtaining a pre-set mapping relationship between models; Converting the second detection result into a model optimization label corresponding to the first network model according to the mapping relationship between the models; Determining a loss function value based on the model optimization label and the first detection result in combination with a pre-set loss function; According to the loss function value, a set gradient descent method is used to adjust the network parameters of the first network model, and the adjusted first network model is determined as the optimized first network model.

7. The method according to claim 1, characterized in that, Also includes: When it is detected that the model evaluation condition is currently met, performance evaluation information of the first network model is determined according to the set model evaluation script combined with the verification image set; When it is determined according to the performance evaluation information that the first network model meets the switching condition, a target network model is determined from a model set to replace the first network model.

8. The method according to claim 1, wherein Also includes: A secondary verification result corresponding to the image to be detected is determined according to the first detection result and the verification processing result, and the secondary verification result is determined as the target detection result of the image to be detected.

9. An image detection device with multi-model collaboration, characterized in that include: an acquisition module, configured to acquire an image to be detected, input the image to be detected as input data into a first network model, and obtain a first detection result output by the first network model; a first execution module, configured to determine the first detection result as the target detection result of the image to be detected when the confidence level of the first detection result is greater than or equal to the confidence level threshold; a second execution module, configured to, when the first detection result is less than the confidence threshold, use a second network model to perform verification processing on the first detection result, and determine the output verification processing result as the target detection result of the image to be detected; Among them, the network structure scale of the first network model is smaller than that of the second network model and the network parameters are different. At the same time, the first network model and the second network model use the same feature space for image detection.

10. An electronic device, characterized in that, The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the multi-model collaborative image detection method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the multi-model collaborative image detection method according to any one of claims 1 to 8 when executed.

12. A computer program product, characterized in that, The computer program product includes a computer program, which, when executed by a processor, implements the multi-model collaborative image detection method according to any one of 1-8.

Citation Information

Cited By

  • Analysis method, device and equipment based on model collaboration

    CN121686007A