A training method, recognition method and system for a bi-optical image model
By training the dual-light image model, pseudo-labels are generated using the labeled visible light training set and the unlabeled infrared light training set to perform label fusion, which solves the problem of high data annotation and registration costs in dual-light image recognition, and improves the generalization ability and accuracy of recognition.
Patent Information
- Application Number
- CN202410889371.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-04
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-07-04
AI Technical Summary
The prior art requires a large amount of double-cursive annotation data in dual-ray image recognition, and the data annotation registration cost is high. The deep learning network model has poor generalization ability and accuracy on unseen data, and is greatly affected by the quality of dual-ray data and label quality.
By obtaining the labeled first visible light training set, the unlabeled second visible light training set and the infrared light training set, the visible and infrared light network framework is trained, pseudo-labels are generated, and the tag fusion is performed. The dual-light image model is trained based on the fusion tag set, reducing the workload of data annotation and registration, and improving the generalization ability and accuracy of recognition.
It effectively reduces the cost of dual-optical data annotation and registration, improves the generalization ability and accuracy of dual-optical image recognition, and reduces the dependence on dual-optical data quality and label quality.
Smart Images

Figure CN118898758B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a training method, a recognition method and a system for a bi-optical image model. Background Art
[0002] In recent years, the method of using visible light images and infrared light images for target recognition has attracted widespread attention because it can achieve good accuracy.
[0003] At present, the traditional method of target recognition based on visible light images and infrared light images is usually implemented using supervised learning. During the training stage, the deep learning network model requires a large amount of dual-light annotated data, and the dual-light annotated data needs to be aligned in advance, which has a high cost. In addition, the traditional method usually performs well on training data, but performs poorly on unseen data, and is affected by the data quality and label quality of the dual-light annotated data. The generalization ability and accuracy of the deep learning network model for identifying dual-light images (i.e., visible light images and infrared images) are not satisfactory.
[0004] Therefore, the problems existing in the prior art still need to be solved and optimized. Summary of the invention
[0005] The purpose of the present invention is to solve one of the technical problems existing in the related art to at least a certain extent.
[0006] To this end, an object of an embodiment of the present invention is to provide a training method for a bi-optical image model, which can effectively reduce the cost of data annotation and registration and improve the generalization ability and accuracy of bi-optical image recognition.
[0007] Another object of an embodiment of the present application is to provide a training system for a bi-optical image model.
[0008] In order to achieve the above technical objectives, the technical solutions adopted by the embodiments of the present application include:
[0009] In a first aspect, an embodiment of the present application provides a training method for a dual-light image model, wherein the dual-light image model includes a visible light network framework and an infrared light network framework, and the training method includes:
[0010] Obtaining a labeled first visible light training set, an unlabeled second visible light training set, and an unlabeled infrared light training set;
[0011] Inputting the first visible light training set into the initialized visible light network framework for training to obtain a trained visible light network framework;
[0012] Inputting the second visible light training set into the trained visible light network framework to obtain a visible light detection result, and inputting the infrared light training set into the initialized infrared light network framework to obtain an infrared light detection result, wherein the visible light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled visible light feature image, the infrared light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled infrared light feature image, and the weight parameters of the initialized infrared light network framework are the same as the weight parameters of the trained visible light network framework;
[0013] Performing label fusion on the visible light detection result and the infrared light detection result to obtain a fused label set;
[0014] Based on the fused label set, the first visible light training set, the corresponding second visible light training set and the corresponding infrared light training set are input into the initialized bi-optical image model for training to obtain a trained bi-optical image model.
[0015] In addition, the training method according to the above embodiment of the present application may also have the following additional technical features:
[0016] Further, in one embodiment of the present application, the step of inputting the first visible light training set into the initialized visible light network framework for training to obtain a trained visible light network framework includes:
[0017] Preprocessing the first visible light training set to obtain a preprocessed first visible light training set;
[0018] Calculating the loss weight of the preprocessed first visible light training set input into the initialized visible light network framework;
[0019] According to the loss weight, the parameters of the initialized visible light network framework are updated to obtain the trained visible light network framework.
[0020] Further, in one embodiment of the present application, the initialized infrared optical network framework is used to perform the following steps:
[0021] Performing infrared feature extraction on the infrared light training set to obtain an infrared feature training set;
[0022] Residual feature extraction is performed on the infrared feature training set to obtain the infrared light detection result.
[0023] Furthermore, in one embodiment of the present application, the label fusion of the visible light detection result and the infrared light detection result to obtain a fused label set includes:
[0024] Acquire a first feature candidate frame in the visible light detection result and a second feature candidate frame in the infrared light detection result;
[0025] According to the first feature candidate frame, performing candidate frame fusion on the second feature candidate frame to obtain a fused candidate frame;
[0026] Extract pseudo-label information from the fused candidate frame to generate the fused label set.
[0027] Further, in one embodiment of the present application, the step of performing candidate frame fusion on the second feature candidate frame according to the first feature candidate frame to obtain a fused candidate frame includes:
[0028] Get the preset confidence threshold;
[0029] According to the confidence threshold, performing a first confidence screening on the first feature candidate box to obtain a screened first feature candidate box;
[0030] According to the confidence threshold, performing a second confidence screening on the second feature candidate box to obtain a screened second feature candidate box;
[0031] The filtered first feature candidate frame and the filtered second feature candidate frame are weightedly fused to obtain the fused candidate frame.
[0032] Furthermore, in the embodiment of the present application, the training method further includes:
[0033] According to the first feature candidate box, performing a candidate box offset evaluation on the corresponding second feature candidate box to obtain candidate box offset data;
[0034] Based on the candidate box offset data, the first feature candidate box and the second feature candidate box are offset aligned to obtain an aligned data set, where the aligned data set includes the aligned first feature candidate box and the aligned second feature candidate box.
[0035] In a second aspect, the present application provides a method for recognizing a bi-optical image model, comprising:
[0036] Acquire an infrared image and a visible light image to be identified;
[0037] The infrared image and the visible light image are input into the trained dual-light image model to obtain a recognition result.
[0038] In a third aspect, an embodiment of the present application provides a training system for a dual-light image model, wherein the dual-light image model includes a visible light network framework and an infrared light network framework, and the training system includes:
[0039] An acquisition module, used to acquire a labeled first visible light training set, an unlabeled second visible light training set, and an unlabeled infrared light training set;
[0040] A first training module, used for inputting the first visible light training set into an initialized visible light network framework for training to obtain a trained visible light network framework;
[0041] A detection module, used to input the second visible light training set into the trained visible light network framework to obtain a visible light detection result, and to input the infrared light training set into the initialized infrared light network framework to obtain an infrared light detection result, wherein the visible light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled visible light feature image, the infrared light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled infrared light feature image, and the weight parameters of the initialized infrared light network framework are the same as the weight parameters of the trained visible light network framework;
[0042] A fusion module, used to perform label fusion on the visible light detection result and the infrared light detection result to obtain a fused label set;
[0043] The second training module is used to input the first visible light training set, the corresponding second visible light training set and the corresponding infrared light training set into the initialized bi-optical image model for training based on the fused label set to obtain a trained bi-optical image model.
[0044] In a fourth aspect, an embodiment of the present application further provides an electronic device, including:
[0045] at least one processor;
[0046] at least one memory for storing at least one program;
[0047] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0048] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores a program executable by a processor, and the program executable by the processor is used to implement the above method when executed by the processor.
[0049] The advantages and benefits of the present application will be partially given in the following description, and partially become apparent from the following description, or be understood through the practice of the present application:
[0050] The present application discloses a training method, a recognition method and a system for a dual-light image model, wherein the training method obtains a labeled first visible light training set, an unlabeled second visible light training set and an unlabeled infrared light training set; inputs the first visible light training set into an initialized visible light network framework for training to obtain a trained visible light network framework; inputs the second visible light training set into the trained visible light network framework to obtain a visible light detection result, and inputs the infrared light training set into the initialized infrared light network framework to obtain an infrared light detection result, wherein the visible light detection result is A set of feature candidate frames corresponding to the pseudo-labels of an unlabeled visible light feature image, the infrared light detection result is a set of feature candidate frames corresponding to the pseudo-labels of an unlabeled infrared light feature image, and the weight parameters of the initialized infrared light network framework are the same as the weight parameters of the trained visible light network framework; performing label fusion on the visible light detection result and the infrared light detection result to obtain a fused label set; based on the fused label set, inputting the first visible light training set, the corresponding second visible light training set and the corresponding infrared light training set into an initialized dual-light image model for training to obtain a trained dual-light image model. The training method obtains a trained visible light network framework based on a small amount of first visible light training set. By loading the weight parameters of the trained visible light network framework and inferring the second visible light training set and the infrared light training set, the corresponding pseudo labels can be obtained, which effectively reduces the workload of labeling and registering dual-light data and reduces the cost of data labeling and registration. In addition, the training method is based on the initialized infrared light network framework, the trained infrared light network framework and label fusion, which can better discover the intrinsic structure and association of feature image data, effectively reduce the impact of poor data quality and label quality of dual-light annotation data, and effectively improve the generalization ability and accuracy of dual-light image recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the embodiments of the present application or the drawings of the related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly expressing some embodiments of the technical solutions of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0052] Figure 1 A schematic diagram of a flow chart of a method for training a bi-optical image model provided in an embodiment of the present application;
[0053] Figure 2 A schematic diagram of the structure of a training system for a bi-optical image model provided in an embodiment of the present application;
[0054] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limitations on the present application. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0057] At present, the traditional method of target recognition based on visible light images and infrared light images is usually implemented using supervised learning. During the training phase, the deep learning network model requires a large amount of dual-light annotation data. The annotation process of dual-light data consumes a lot of time and is costly. In addition, the deep learning network model usually needs to align the dual-light data in advance to facilitate subsequent feature fusion or input layer fusion. The alignment process is relatively cumbersome, inconvenient to implement, and costly.
[0058] In addition, traditional methods usually perform well on training data but poorly on unseen data. Specifically, they may overfit specific features in the bi-photonic annotated data and their generalization capabilities are limited. In practical applications, the quality and distribution of bi-photonic annotated data often deviate from the true distribution, and the accuracy of bi-photonic image recognition is often unsatisfactory.
[0059] In view of this, an embodiment of the present invention provides a training method for a dual-light image model. The training method obtains a trained visible light network framework based on a small amount of first visible light training set. By loading the weight parameters of the trained visible light network framework, the second visible light training set and the infrared light training set are inferred to obtain corresponding pseudo labels, which effectively reduces the workload of labeling and aligning dual-light data and reduces the cost of data labeling and registration. In addition, the training method is based on an initialized infrared light network framework, a trained infrared light network framework and label fusion, which can better discover the intrinsic structure and association of feature image data, effectively reduce the impact of poor data quality and label quality of dual-light annotation data, and effectively improve the generalization ability and accuracy of dual-light image recognition.
[0060] Reference Figure 1 In an embodiment of the present application, a training method for a dual-light image model is provided, wherein the dual-light image model includes a visible light network framework and an infrared light network framework, and the training method includes:
[0061] Step 110, obtaining a labeled first visible light training set, an unlabeled second visible light training set, and an unlabeled infrared light training set;
[0062] In an embodiment of the present application, the first visible light training set may be a publicly available data set of visible light images, and the first visible light training set may be labeled; the second visible light training set and the infrared light training set may be obtained based on real-time images captured by visual sensors (including infrared light sensors and visible light sensors) from multiple perspectives within the working area.
[0063] Step 120: input the first visible light training set into the initialized visible light network framework for training to obtain a trained visible light network framework;
[0064] In some embodiments, the step 120 of inputting the first visible light training set into the initialized visible light network framework for training to obtain a trained visible light network framework includes:
[0065] A1. preprocessing the first visible light training set to obtain a preprocessed first visible light training set;
[0066] A2. Calculate the loss weight of the preprocessed first visible light training set input into the initialized visible light network framework;
[0067] A3. According to the loss weight, the parameters of the initialized visible light network framework are updated to obtain the trained visible light network framework.
[0068] In an embodiment of the present application, preprocessing the first visible light training set may include performing noise filtering, binarization processing, image enhancement processing, noise reduction processing, etc. on each visible light image data in the first visible light training set, wherein the image enhancement processing includes random cropping, rotation, translation, etc.
[0069] It is understandable that the visible light network framework can be a simple convolutional neural network (CNN). For the visible light network framework, the accuracy of the prediction results of the visible light network framework can be measured by the loss function. The loss function is defined on a single training data and is used to measure the prediction error of a training data. Specifically, the loss value of the training data is determined by the label of a single training data and the prediction result of the visible light network framework on the training data. In actual training, a training data set has a lot of training data, so the cost function is generally used to measure the overall error of the training data set. The cost function is defined on the entire training data set and is used to calculate the average value of the prediction error of all training data, which can better measure the prediction effect of the visible light network framework. For a general visible light network framework, based on the aforementioned cost function, plus a regularization term that measures the complexity of the model, it can be used as the objective function of the training, and the loss value of the entire training data set can be calculated based on the objective function. There are many types of commonly used loss functions, such as 0-1 loss function, square loss function, absolute loss function, logarithmic loss function, cross entropy loss function, etc., which can all be used as loss functions of the visible light network framework, which will not be elaborated one by one here. In the embodiment of the present application, any loss function can be selected to determine the loss value of training, such as the cross entropy loss function. In addition, the first visible light training set is used as the training data of the visible light network framework; since the loss function includes the loss weight, based on the loss value of the training, the back propagation algorithm can be used to update the parameters of the model, and the trained visible light network framework model can be obtained after several rounds of iteration. The specific number of iterations can be preset, or the training is considered to be completed when the test set meets the accuracy requirements.
[0070] Step 130: input the second visible light training set into the trained visible light network framework to obtain a visible light detection result, and input the infrared light training set into the initialized infrared light network framework to obtain an infrared light detection result, wherein the visible light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled visible light feature image, the infrared light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled infrared light feature image, and the weight parameters of the initialized infrared light network framework are the same as the weight parameters of the trained visible light network framework;
[0071] In some embodiments, the initialized infrared optical network framework in step 130 is used to perform the following steps:
[0072] B1. Extracting infrared features from the infrared light training set to obtain an infrared feature training set;
[0073] B2. Perform residual feature extraction on the infrared feature training set to obtain the infrared light detection result.
[0074] In an embodiment of the present application, after obtaining a trained visible light network framework, the trained visible light network framework can be used as a first branch backbone feature extraction network to perform target detection on a second visible light training set without labels, thereby obtaining target detection results of the second visible light training set on the visible light RGB channel (i.e., visible light detection results). For a certain visible light detection result, the visible light detection result is used to characterize a feature candidate frame of the corresponding visible light feature image, and the feature candidate frame records a pseudo-label corresponding to the target, as well as a confidence level corresponding to the pseudo-label, etc.
[0075] The infrared light network framework in the embodiment of the present application serves as a second branch backbone feature extraction network, which can be constructed based on the Laplacian pyramid (LP) and a deep residual network. The deep residual network can be any one of Resnet18, Resnet50, etc. Specifically, before the infrared light training set is input into the infrared light network framework for training, the network parameters of the visible light network framework can be reused and loaded into the initial infrared light network framework to obtain an initialized infrared light network framework. Based on the hierarchical approximation of the Laplacian pyramid and the difference of the detail layer, the image features can be represented at multiple different resolution levels, thereby capturing detail features of different scales; based on the residual feature extraction in step B2 performed by the deep residual network, the infrared light detection result of the infrared light training set on the infrared TIR channel is obtained. The infrared light detection result is similar to the aforementioned visible light detection result and can be simply deduced by analogy.
[0076] Step 140: fusing labels of the visible light detection result and the infrared light detection result to obtain a fused label set;
[0077] In some embodiments, the step 140 of fusing labels of the visible light detection result and the infrared light detection result to obtain a fused label set includes:
[0078] C1. Obtaining a first feature candidate frame in the visible light detection result and a second feature candidate frame in the infrared light detection result;
[0079] C2. Based on the first feature candidate box, fuse the second feature candidate box to obtain a fused candidate box;
[0080] Furthermore, the step C2, performing candidate frame fusion on the second feature candidate frame according to the first feature candidate frame to obtain a fused candidate frame, includes:
[0081] C21, obtain the preset confidence threshold;
[0082] C22. Performing a first confidence screening on the first feature candidate box according to the confidence threshold to obtain a screened first feature candidate box;
[0083] C23. Perform a second confidence screening on the second feature candidate box according to the confidence threshold to obtain a screened second feature candidate box;
[0084] C24. Perform weighted frame fusion on the filtered first feature candidate frame and the filtered second feature candidate frame to obtain the fused candidate frame.
[0085] In an embodiment of the present application, for a certain image feature, there is a first feature candidate box in the visible light detection result and a second feature candidate box in the infrared light detection result. Each first feature candidate box corresponds to a second feature candidate box. The label fusion in step 140 is used to filter out a unique label for the image feature.
[0086] It is understandable that step C21 may firstly obtain a preset confidence threshold, which may be any one of 0.8, 0.85, 0.9, 0.95, etc. The embodiment of the present application takes the confidence threshold of 0.9 as an example. Step C22 may be to compare the confidence of each first feature candidate frame record with the preset confidence threshold, thereby selecting a number of first feature candidate frames with a confidence greater than or equal to 0.9 from all first feature candidate frames as the selected first feature candidate frames.
[0087] It should be noted that the second confidence screening in step C23 is similar to the first confidence screening in step C22 and can be simply deduced. For step C24, all the filtered first feature candidate boxes and the filtered feature candidate boxes can be fused based on the weighted bounding box fusion (Weighted Boxes Fusion: combining boxes for object detection models, WBF) algorithm to obtain a fused candidate box.
[0088] C3. Extract pseudo-label information in the fused candidate frame to generate the fused label set.
[0089] In the embodiment of the present application, the fused candidate frame obtained by step C2 can effectively integrate the target detection results of the infrared image and the target detection results of the visible light image, and effectively improve the quality of the pseudo-label. In addition, for a certain fused candidate frame, step C3 can be directly extracted based on the pseudo-label recorded in the fused candidate frame, and the rest of the fused candidate frames are similar, and the fused label set is obtained by integrating the pseudo-label information of each fused candidate frame.
[0090] Step 150: Based on the fused label set, the first visible light training set, the corresponding second visible light training set and the corresponding infrared light training set are input into the initialized bi-optical image model for training to obtain a trained bi-optical image model.
[0091] In an embodiment of the present application, the corresponding second visible light training set can be based on the fused label set, and the original second visible light training set is screened to screen the visible light image data corresponding to each pseudo-label information, and the screened visible light image data is used as the corresponding second visible light training set. Each visible light image data in the corresponding second visible light training set has a pseudo-label information with high confidence; the corresponding infrared light training set can be simply deduced in the same way.
[0092] It is understandable that the training of the dual-light image model in step 150 is similar to the training content of the visible light network framework in the aforementioned step 120, and can be simply deduced by analogy, so this application will not go into details here.
[0093] In some embodiments, the training method further comprises:
[0094] Step 160: Based on the first feature candidate box, perform candidate box offset evaluation on the corresponding second feature candidate box to obtain candidate box offset data;
[0095] Step 170: Based on the candidate box offset data, offset-align the first feature candidate box and the second feature candidate box to obtain an aligned data set, wherein the aligned data set includes the aligned first feature candidate box and the aligned second feature candidate box.
[0096] In an embodiment of the present application, after the visible light network framework outputs the visible light detection result, and the infrared light network framework outputs the infrared light detection result, the corresponding first feature candidate frame and the corresponding second feature candidate frame can be aligned. The specific alignment can be based on the Manhattan distance matching algorithm, according to the center position of the first feature candidate frame and the center position of the second feature candidate frame, determine the mean of the offset of the first feature candidate frame and the corresponding second feature candidate frame in the x-direction and y-direction (i.e., candidate frame offset data); then according to the obtained candidate frame offset data, move the first feature candidate frame, or the corresponding second feature candidate frame, so as to obtain the first feature candidate frame and the second feature candidate frame that are aligned with each other (i.e., alignment data set). It should be noted that in the process of offset alignment, the blank spaces that appear can be filled with zero grayscale values, which will not be repeated in this application.
[0097] In an embodiment of the present application, a method for recognizing a bi-optical image model includes:
[0098] Acquire an infrared image and a visible light image to be identified;
[0099] The infrared image and the visible light image are input into the bi-optical image model trained as described above to obtain a recognition result.
[0100] In an embodiment of the present application, the infrared image and visible light image to be identified can be collected by a visual sensor at the same time, and the obtained infrared image and visible light image are input into a trained dual-light image model to obtain the recognition result predicted by the dual-light image model.
[0101] A training system for a bi-optical image model proposed according to an embodiment of the present application is described in detail below with reference to the accompanying drawings.
[0102] Reference Figure 2 , a training system for a dual-light image model proposed in an embodiment of the present application, the dual-light image model includes a visible light network framework and an infrared light network framework, and the training system includes:
[0103] An acquisition module 101 is used to acquire a labeled first visible light training set, an unlabeled second visible light training set, and an unlabeled infrared light training set;
[0104] A first training module 102 is used to input the first visible light training set into an initialized visible light network framework for training to obtain a trained visible light network framework;
[0105] A detection module 103 is used to input the second visible light training set into the trained visible light network framework to obtain a visible light detection result, and to input the infrared light training set into the initialized infrared light network framework to obtain an infrared light detection result, wherein the visible light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled visible light feature image, the infrared light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled infrared light feature image, and the weight parameters of the initialized infrared light network framework are the same as the weight parameters of the trained visible light network framework;
[0106] A fusion module 104 is used to perform label fusion on the visible light detection result and the infrared light detection result to obtain a fused label set;
[0107] The second training module 105 is used to input the first visible light training set, the corresponding second visible light training set and the corresponding infrared light training set into the initialized bi-optical image model for training based on the fused label set to obtain a trained bi-optical image model.
[0108] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0109] Reference Figure 3 , the embodiment of the present application further provides an electronic device, including:
[0110] at least one processor 201;
[0111] At least one memory 202, used to store at least one program;
[0112] When the at least one program is executed by the at least one processor 201, the at least one processor 201 implements the above method embodiment.
[0113] Similarly, it can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0114] The embodiment of the present application further provides a computer-readable storage medium, in which a program executable by the processor 201 is stored. The program executable by the processor 201 is used to implement the above-mentioned method embodiment when executed by the processor 201.
[0115] Similarly, the contents of the above method embodiments are all applicable to the computer-readable storage medium embodiments. The functions specifically implemented by the computer-readable storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0116] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the application is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part described as a larger operation is performed independently.
[0117] In addition, although the present application is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present application. More specifically, in view of the properties, functions, and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional techniques of the engineer. Therefore, those skilled in the art can implement the present application set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the attached claims and their equivalents.
[0118] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method of the embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0119] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0120] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0121] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0122] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0123] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.
[0124] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the embodiments. Technical personnel familiar with the field can make various equivalent modifications or substitutions without violating the spirit of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. A method for training a bi-optical image model, characterized in that: The dual-light image model includes a visible light network framework and an infrared light network framework, and the training method includes: Obtaining a labeled first visible light training set, an unlabeled second visible light training set, and an unlabeled infrared light training set; Inputting the first visible light training set into the initialized visible light network framework for training to obtain a trained visible light network framework; Inputting the second visible light training set into the trained visible light network framework to obtain a visible light detection result, and inputting the infrared light training set into the initialized infrared light network framework to obtain an infrared light detection result, wherein the visible light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled visible light feature image, the infrared light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled infrared light feature image, and the weight parameters of the initialized infrared light network framework are the same as the weight parameters of the trained visible light network framework; Performing label fusion on the visible light detection result and the infrared light detection result to obtain a fused label set; Based on the fused label set, inputting the first visible light training set, the corresponding second visible light training set, and the corresponding infrared light training set into an initialized bi-optical image model for training to obtain a trained bi-optical image model; The step of fusing labels of the visible light detection result and the infrared light detection result to obtain a fused label set includes: Acquire a first feature candidate frame in the visible light detection result and a second feature candidate frame in the infrared light detection result; According to the first feature candidate frame, performing candidate frame fusion on the second feature candidate frame to obtain a fused candidate frame; Extract pseudo-label information from the fused candidate frame to generate the fused label set.
2. The training method according to claim 1, characterized in that: The step of inputting the first visible light training set into the initialized visible light network framework for training to obtain a trained visible light network framework includes: Preprocessing the first visible light training set to obtain a preprocessed first visible light training set; Calculating the loss weight of the preprocessed first visible light training set input into the initialized visible light network framework; According to the loss weight, the parameters of the initialized visible light network framework are updated to obtain the trained visible light network framework.
3. The training method according to claim 1, characterized in that: The initialized infrared optical network framework is used to perform the following steps: Performing infrared feature extraction on the infrared light training set to obtain an infrared feature training set; Residual feature extraction is performed on the infrared feature training set to obtain the infrared light detection result.
4. The training method according to claim 1, characterized in that: The step of performing candidate frame fusion on the second feature candidate frame according to the first feature candidate frame to obtain a fused candidate frame includes: Get the preset confidence threshold; According to the confidence threshold, performing a first confidence screening on the first feature candidate box to obtain a screened first feature candidate box; According to the confidence threshold, performing a second confidence screening on the second feature candidate box to obtain a screened second feature candidate box; The filtered first feature candidate frame and the filtered second feature candidate frame are weightedly fused to obtain the fused candidate frame.
5. The training method according to claim 1, characterized in that: The training method further comprises: According to the first feature candidate box, performing a candidate box offset evaluation on the corresponding second feature candidate box to obtain candidate box offset data; Based on the candidate box offset data, the first feature candidate box and the second feature candidate box are offset aligned to obtain an aligned data set, where the aligned data set includes the aligned first feature candidate box and the aligned second feature candidate box.
6. A method for recognizing a bi-optical image model, characterized in that: include: Acquire an infrared image and a visible light image to be identified; The infrared image and the visible light image are input into the trained bi-optical image model as described in any one of claims 1 to 5 to obtain a recognition result.
7. A training system for a bi-optical image model, characterized in that: The dual-light image model includes a visible light network framework and an infrared light network framework, and the training system includes: An acquisition module, used to acquire a labeled first visible light training set, an unlabeled second visible light training set, and an unlabeled infrared light training set; A first training module, used for inputting the first visible light training set into an initialized visible light network framework for training to obtain a trained visible light network framework; A detection module, used to input the second visible light training set into the trained visible light network framework to obtain a visible light detection result, and to input the infrared light training set into the initialized infrared light network framework to obtain an infrared light detection result, wherein the visible light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled visible light feature image, the infrared light detection result is a feature candidate frame set corresponding to a pseudo-label of an unlabeled infrared light feature image, and the weight parameters of the initialized infrared light network framework are the same as the weight parameters of the trained visible light network framework; A fusion module, used to perform label fusion on the visible light detection result and the infrared light detection result to obtain a fused label set; a second training module, configured to input the first visible light training set, the corresponding second visible light training set, and the corresponding infrared light training set into an initialized bi-optical image model for training based on the fused label set, so as to obtain a trained bi-optical image model; The step of fusing labels of the visible light detection result and the infrared light detection result to obtain a fused label set includes: Acquire a first feature candidate frame in the visible light detection result and a second feature candidate frame in the infrared light detection result; According to the first feature candidate frame, performing candidate frame fusion on the second feature candidate frame to obtain a fused candidate frame; Extract pseudo-label information from the fused candidate frame to generate the fused label set.
8. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to implement the method according to any one of claims 1 to 6 when executed by the processor.
Citation Information
Patent Citations
Target detection method and device, electronic equipment and storage medium
CN116189015A
Method for training underwater recognition model, underwater archaeological method and related equipment
CN117953331A