Method for realizing target detection by balancing field diversity and invariance
By collecting and preprocessing data in object detection, extracting key features, and using diversity enhancement and alignment loss adjustment models to balance field diversity and invariance, the problems of insufficient robustness among fields and limited model performance in the prior art are solved, and higher adaptability and generalization performance are achieved.
Patent Information
- Application Number
- CN202510073285.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-03
AI Technical Summary
The existing single-domain object detection method is insufficient in robustness among different fields, and it is impossible to effectively balance the learning of domain-specific features and cross-domain invariant features, resulting in limited model performance and insufficient domain bias and generalization performance.
By collecting and preprocessing model data, extracting key features, and adjusting the model through diversity enhancement model and alignment loss, optimizing the overall goal, balancing the diversity and invariance of the field, thereby improving the model's adaptability and generalization performance.
It significantly enhances the adaptability to unknown target domains, eliminates semantic differences between different domains, improves the model's performance on target tasks, and overcomes the lack of learning ability of domain-specific features.
Smart Images

Figure CN120088450A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer target detection. More specifically, the present invention relates to a method for achieving target detection by balancing domain diversity and invariance. Background Art
[0002] With the wide application and deep integration of artificial intelligence technology in the field of computer vision, target detection, as one of the key technologies in computer vision, is particularly important for realizing the automatic recognition and positioning of specific targets in images and videos.
[0003] Existing single-domain target detection methods include the single-domain generalization (S-DGOD) model, which mainly focuses on learning feature invariance to improve the robustness of the model across different domains; another single-domain target detection method is the IBN-Net model, which is a convolutional neural network model that combines instance normalization and batch normalization. IBN-Net combines IN and BN to improve the generalization ability of the model.
[0004] However, in actual use, there are still some disadvantages, such as insufficient robustness across different domains, that is, insufficient domain generalization performance, resulting in performance degradation; inability to effectively balance the learning of domain-specific features and cross-domain invariant features, resulting in limited model performance; over-reliance on specific source domain data, leading to domain bias and affecting the detection performance of the target domain; and failure to fully utilize the diversity of domain-specific information, affecting the efficiency and accuracy of feature extraction. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, the present invention provides a method for achieving target detection by balancing domain diversity and invariance, through the following solutions, to solve the problems raised in the above background art.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for achieving target detection by balancing domain diversity and invariance, characterized by comprising:
[0008] S1: Collect model data: In response to the corresponding learning model in the balanced domain detection system, obtain the first learning objective of the target learning model;
[0009] S2: Preprocess the model data: Perform a preprocessing operation on the first learning objective, and the preprocessing operation is used to obtain the first key objective group corresponding to the first learning objective;
[0010] S3: Extract key features: Perform a feature extraction operation on the first key objective group, and the feature extraction operation is used to obtain the first target feature corresponding to the first learning objective;
[0011] S4: Enhance Diversity: Obtain a diversity enhancement model, and based on the diversity enhancement model and through the first target feature, obtain a second learning objective;
[0012] S5: Dynamic Weight Alignment: Obtain an alignment loss adjustment model, and based on the alignment loss adjustment model and through the first target feature, obtain a third learning objective;
[0013] S6: Optimize the Overall Objective: Perform an overall optimization operation based on the second learning objective and the third learning objective, where the overall optimization operation is used to obtain a second target feature;
[0014] S7: Evaluate and Detect the Target Learning Model: Based on the corresponding learning model in the balanced domain detection system and through the second target feature, perform a target detection operation on the corresponding output result of the learning model, where the target detection operation is used to obtain a detection report of the learning model;
[0015] Preferably, for S2, obtaining the first key target group corresponding to the first learning objective specifically includes:
[0016] Perform a data augmentation operation on the single-domain data corresponding to the first learning objective. The data augmentation operation is a target transformation operation on the single-domain data corresponding to the first learning objective, and the target transformation operation includes a random cropping operation, a color transformation operation, and a geometric transformation operation;
[0017] A401: Random Cropping Operation: Randomly crop the single-domain image corresponding to the first learning objective according to a preset ratio, and the cropped image still contains the target key feature part;
[0018] A402: Color Transformation Operation: Adjust the color channels of the single-domain image corresponding to the first learning objective by a preset ratio to simulate the color performance of the target object under different illuminations and different environmental hues;
[0019] A403: Geometric Transformation Operation: Using the image center as the rotation point, randomly rotate the single-domain image corresponding to the first learning objective within a preset angle range, and synchronously transform and adjust the annotation of the target feature; at the same time, perform horizontal flipping, vertical flipping, and horizontal-vertical combined flipping operations on the single-domain image corresponding to the first learning objective.
[0020] Preferably, for S3, the feature extraction operation is to input the first key target group into a feature extraction network to generate a feature map corresponding to the first learning objective and a feature map corresponding to the first key target group;
[0021] The feature extraction network consists of a preset basic detector and a learning network.
[0022] Preferably, for S3, obtaining the first target feature corresponding to the first learning objective specifically includes:
[0023] Based on the preset mapping and sampling rules, convert the key feature regions corresponding to multiple anchor point mechanisms into the feature PI representation of the fixed anchor point mechanism to obtain the target feature z corresponding to the first learning objective s and the target feature z corresponding to the first key objective group a ;
[0024] Obtain the specific representation of the target feature corresponding to the first learning objective based on the feature map Fs corresponding to the first learning objective and the feature map Fa corresponding to the first key objective group, which is specifically expressed as:
[0025] z s = RP(F s , PI),
[0026] where RP represents the extraction function existing in the RoI-Pooling layer, and PI represents the feature of the fixed anchor point mechanism corresponding to the key feature region;
[0027] The specific representation of the target feature corresponding to the first key objective group is:
[0028] z a = RP(F a , PI),
[0029] where RP represents the extraction function existing in the RoI-Pooling layer, and PI represents the feature of the fixed anchor point mechanism corresponding to the key feature region.
[0030] Preferably, in step S4, obtaining the second learning objective specifically includes:
[0031] Separate the unique feature z from the first target feature, and add the target classification loss L d , the maximum entropy loss L c , and the feature diversity loss L H , and calculate the diversity loss balance value L FD of the unique feature and the target feature, which is specifically expressed as: DLM
[0032] L DLM = L C + L H + β 1 × L FD ,
[0033] where β 1 represents the hyperparameter for balancing the loss.
[0034] Preferably, in step S5, obtaining the third learning objective specifically includes:
[0035] Introduce the cosine similarity losses L fa and L fs, calculate the weighted loss balance value L of the target feature corresponding to the first learning objective and the target feature corresponding to the first key objective group WAM , specifically expressed as:
[0036] L WAM =β 2 ×L fa +L fs ,
[0037] Among them, β 2 represents the weighted parameter, and its value is the feature distribution similarity between the target feature corresponding to the first learning objective and the target feature corresponding to the first key objective group.
[0038] Preferably, for step S5 of obtaining the third learning objective, it specifically includes:
[0039] The dynamic adjustment of the weighted parameter takes the class prediction alignment loss between the target feature corresponding to the first learning objective and the target feature corresponding to the first key objective group, and the bounding box prediction alignment loss between the target feature corresponding to the first learning objective and the target feature corresponding to the first key objective group as the adjustment index of the weighted parameter;
[0040] Based on the influence degree of feature differences on the detection target, combined with the adjustment index of the weighted parameter, adaptively adjust the size, and the value ranges are all between [0, 1]. Select the value close to 1 as the value of the weighted parameter.
[0041] Preferably, for step S6 of obtaining the second target feature, it specifically includes:
[0042] Based on the detection loss L det of the convolutional neural network, the diversity loss balance value L DLM , and the weighted loss balance value L WAM , calculate the overall optimization loss value L t , specifically expressed as:
[0043] L t =L det +α×ln(L DLM )+β×(L WAM / e),
[0044] Among them, α represents the diversity loss balance weight, β represents the weighted loss balance weight, and e represents a constant with a value of 2.718.
[0045] To achieve the above object, the present invention provides the following technical solution: A system for realizing object detection by balancing domain diversity and invariance, including a system operation database, a system central processing module, and a user information terminal. Implementing the above method for realizing object detection by balancing domain diversity and invariance includes:
[0046] Image data acquisition module: used to obtain the first learning objective of the target learning model in response to the corresponding learning model in the balance domain detection system;
[0047] Image data preprocessing module: used to perform preprocessing operations on the first learning objective transmitted by the image data acquisition module, and the preprocessing operations are used to obtain the first key objective group corresponding to the first learning objective;
[0048] Feature extraction module: used to perform feature extraction operations on the first key objective group transmitted by the image data preprocessing module, and the feature extraction operations are used to obtain the first objective feature corresponding to the first learning objective;
[0049] Diversity learning module: used to obtain a diversity enhancement model, and obtain the second learning objective according to the first objective feature transmitted by the feature extraction module through the diversity enhancement model;
[0050] Weighted alignment module: used to obtain an alignment loss adjustment model, and obtain the third learning objective according to the first objective feature transmitted by the feature extraction module through the alignment loss adjustment model;
[0051] Target overall optimization module: perform overall optimization operations based on the second learning objective transmitted by the diversity learning module and the third learning objective transmitted by the weighted alignment module, and the overall optimization operations are used to obtain the second objective feature;
[0052] Model evaluation and detection module: used to perform target detection operations on the corresponding output results of the learning model according to the second objective feature transmitted by the target overall optimization module through the balance domain detection system, and the target detection operations are used to obtain the detection report of the learning model;
[0053] The system operation database includes all data texts of the balance domain detection system, and real-time collects the information texts output by each module. The system central processing module is used to control the information text instructions output by each module in the control method, and the user information terminal is an information output device for receiving the balance domain detection system.
[0054] Technical effects and advantages of the present invention:
[0055] 1. The present invention constructs a feature extraction network through S3 to balance the diversity and invariance within the balance domain, significantly enhancing the adaptability to unknown target domains;
[0056] 2. The present invention uses the diversity enhancement model of S4 to eliminate semantic differences between different domains, which helps to learn and utilize domain-specific features and improve performance on target tasks;
[0057] 3. The present invention uses the alignment loss adjustment model of S5 to learn domain-invariant features, thereby overcoming the situation of insufficient domain-specific feature learning ability. Brief Description of the Drawings
[0058] Figure 1 It is a flowchart of the method steps of the present invention.
[0059] Figure 2 It is a flowchart of the method of the present invention.
[0060] Figure 3 It is a schematic structural diagram of the method of the present invention. Detailed Embodiments
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0062] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification of the present application, the singular forms "a", "an", "the", "above-mentioned", "the foregoing", "this" are also intended to include the plural forms, unless clearly indicated to the contrary in the context. It should also be understood that the term " / and / " used in the present application refers to and includes any or all possible combinations of one or more of the listed items.
[0063] Hereinafter, the terms "first", "second", and "third" are only used for descriptive purposes, and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0064] As shown in the attached Figure 1 A method for achieving object detection by balancing domain diversity and invariance, including S1: collecting model data, S2: preprocessing model data, S3: extracting key features, S4: enhancing diversity, S5: dynamic weight alignment, S6: optimizing the overall object, and S7: evaluating and detecting the object learning model.
[0065] S1: Collecting model data: In response to the corresponding learning model in the balanced domain detection system, obtain the first learning objective of the target learning model.
[0066] Specifically, the balanced domain detection system determines the demand objectives and target research fields through the target learning model. The demand objectives include, but are not limited to, improving model performance, improving target detection accuracy, increasing the diversity of learning models, etc. The balanced domain detection system collects test tasks through the target learning model. The test task is the test data input into the target learning model. The balanced domain detection system determines the data type and data source of the test task to obtain the first learning objective.
[0067] It should be noted that in response to the data type of the test task being mainly image data, including images containing target requirements under different scenarios and different environmental conditions. The data sources include, but are not limited to, publicly available large-scale datasets, datasets rendered through specific methods, datasets generated through simulation, etc.
[0068] In this embodiment, the first learning objective includes five different weather conditions, which are clear day, clear night, rainy dusk, rainy night, and foggy day. The data is from the BDD-100K dataset, including 27,708 images under clear day conditions, 26,158 images under clear night conditions, 3,775 images under foggy conditions, 3,501 images on rainy dusk days, and 2,494 images on rainy nights. The demand objective is to improve the target detection accuracy in the image data of the model under different weather conditions, especially to improve the detection accuracy for bicycles, cars, and pedestrians.
[0069] S2: Preprocess model data: Perform preprocessing operations on the first learning objective. The preprocessing operations are used to obtain the first key objective group corresponding to the first learning objective.
[0070] Specifically, after the balanced domain detection system obtains the first learning objective, through preprocessing operations, it can obtain multiple key information corresponding to the first learning objective, that is, the first key objective group.
[0071] In a possible implementation manner, the steps of preprocessing are as follows:
[0072] A1: Remove noise: Use an image filtering algorithm to remove the noise existing in the first learning objective. The image filtering algorithm includes, but is not limited to, Gaussian filtering, median filtering, etc. In this embodiment, the filter kernel is selected according to the noise level corresponding to the first learning objective. When the noise is slight, a 3×3 filter kernel is selected, and when the noise is more serious, a 5×5 filter kernel is used.
[0073] A2: Correct mislabeling: By combining manual sampling inspection with automated verification, confirm and correct abnormal situations in the first learning objective where the labeled category is incorrect or the labeled position has a large deviation. In this embodiment, the abnormal situation where the labeled category is incorrect in the first learning objective is mislabeling a bicycle as a motorcycle, and the abnormal situation where the labeled position has a large deviation is that there is a blank area between the actual boundary and the labeled boundary of the target object.
[0074] A3: Data normalization: Use an image scaling algorithm to uniformly adjust the corresponding image data in the first learning objective to a fixed size. The image scaling algorithm includes, but is not limited to, bilinear interpolation, nearest neighbor interpolation, etc. At the same time, linearly map the pixel value range of each pixel point in the corresponding image data of the first learning objective from [0, 255] to the interval [0, 1].
[0075] A4: Data augmentation: Perform data augmentation operations on the single-domain data corresponding to the first learning objective. The data augmentation operation is a target transformation operation on the single-domain data corresponding to the first learning objective. The target transformation operation includes random cropping operations, color transformation operations, and geometric transformation operations.
[0076] A401: Random cropping operation: Randomly crop the single-domain image corresponding to the first learning objective according to a preset ratio, and the cropped image still contains the key feature part of the target. In this embodiment, the preset cropping ratio range is (10%, 30%).
[0077] A402: Color transformation operation: Adjust the color channels of the single-domain image corresponding to the first learning objective by a preset ratio to simulate the color performance of the target object under different illuminations and different environmental hues. In this embodiment, multiply the pixel points of the original image by a random coefficient in the red channel value, and the range of the random coefficient is set to (0.8, 1.2).
[0078] A403: Geometric transformation operation: Use the center of the image as the rotation point, randomly rotate the single-domain image corresponding to the first learning objective within a preset angle range, and synchronously transform and adjust the annotation of the target feature. At the same time, perform horizontal flipping, vertical flipping, and horizontal-vertical combined flipping operations on the single-domain image corresponding to the first learning objective to obtain various situations corresponding to the single-domain image of the first learning objective.
[0079] S3: Extract key features: Perform a feature extraction operation on the first key target group. The feature extraction operation is used to obtain the first target feature corresponding to the first learning objective.
[0080] Specifically, the feature extraction operation is to input the first key target group into the feature extraction network and generate the feature map F corresponding to the first learning objective through the convolution and pooling operations of the feature extraction network. s and the feature map F corresponding to the first key target group a, the feature extraction network consists of a preset basic detector and a learning network; the feature map F corresponding to the first learning objective is s sent into the RPN, that is, the Region Proposal Network, to obtain multiple key features P; based on the feature map F corresponding to the first learning objective s , the feature map F corresponding to the first key objective group a , and multiple key features P to obtain multiple target features through the RoI-Pooling layer, that is, the first target features.
[0081] In a possible implementation manner, obtaining the first target features corresponding to the first learning objective includes: selecting a feature extraction network architecture; in this embodiment, the feature extraction network architecture uses Faster R-CNN as the basic detector and ResNet101 as the learning network; inputting the first key objective group into the feature extraction network to generate a feature map and multiple key features.
[0082] It should be noted that the learning network extracts features through the alternating operations of multiple convolutional layers and pooling layers. In the convolutional layer, a sliding convolutional operation is performed on the first key objective group through multiple convolutional kernels of different sizes to gradually extract local features and continuously integrate them as the network depth increases; in the pooling layer, downsampling is performed on the feature map output by the convolutional layer to reduce the feature map size and strengthen the key feature information; obtaining the feature map F s corresponding to the first learning objective and the feature map F a corresponding to the first key objective group; inputting the feature map F s into the Region Proposal Network, and the Region Proposal Network scans and predicts multiple key features P containing target requirements through preset convolutional operations and the anchor mechanism. The anchor mechanism is multiple rectangular boxes with preset sizes and aspect ratios in object detection;
[0083] In a possible implementation manner, obtaining the first target features corresponding to the first learning objective includes: performing a pooling operation on the feature map F s corresponding to the first learning objective, the feature map F a corresponding to the first key objective group, and multiple key features P through the RoI-Pooling layer to obtain multiple target features of the first key objective group.
[0084] It should be noted that based on the preset mapping and sampling rules, the key feature regions corresponding to multiple anchor mechanisms are transformed into the feature PI representation of the fixed anchor mechanism to obtain the target feature z s corresponding to the first learning objective and the target feature z a corresponding to the first key objective group; among them, the target feature corresponding to the first learning objective is specifically represented as:
[0085] z s = RP(Fs , PI),
[0086] where RP represents the extraction function existing in the RoI-Pooling layer, F s represents the feature map corresponding to the first learning objective, and PI represents the feature of the fixed anchor mechanism corresponding to the key feature region;
[0087] The target feature corresponding to the first key target group is specifically represented as:
[0088] z a = RP(F a , PI),
[0089] where RP represents the extraction function existing in the RoI-Pooling layer, F a represents the feature map corresponding to the first key target group, and PI represents the feature of the fixed anchor mechanism corresponding to the key feature region.
[0090] S4: Enhance diversity: Obtain a diversity enhancement model, and obtain a second learning objective according to the diversity enhancement model through the first target feature.
[0091] Specifically, the diversity enhancement model is a pre-constructed learning model. By inputting the first target feature into the diversity enhancement model, the diversity enhancement model obtains the second learning objective according to the first target feature.
[0092] In a possible implementation manner, obtaining the second learning objective includes: separating the unique feature z d from the first target feature, where the unique feature is the key information specific to the target domain, and adding the target classification loss L c , the maximum entropy loss L H , and the feature diversity loss L FD , calculating the diversity loss balance value L DLM of the unique feature and the target feature, which is specifically represented as:
[0093] L DLM = L C + L H + β 1 × L FD ,
[0094] where β 1 represents the hyperparameter for balancing the loss;
[0095] It should be noted that the hyperparameter β 1 of the balanced loss determines the weight of the feature diversity loss in the entire total loss, and thus affects the degree of attention of the model to diversity enhancement; the maximum entropy loss L HAim to reduce the class-level semantic information; Feature diversity loss L FD Used to measure the unique feature z from the numerical distribution and structural features of the features d Degree of diversity.
[0096] Specifically, after minimizing the diversity loss balance value between the unique feature and the target feature, input the data corresponding to the unique feature into the diversity enhancement model, and update the parameters based on the loss through backpropagation. After convergence is achieved through multiple rounds of simulation, integrate and adjust the features and perform standardization processing to obtain the second learning objective.
[0097] S5: Dynamic weight alignment: Obtain an alignment loss adjustment model, and obtain the third learning objective according to the alignment loss adjustment model through the first target feature.
[0098] Specifically, the alignment loss adjustment model is a pre-constructed learning model. By inputting the first target feature into the alignment loss adjustment model, the alignment loss adjustment model obtains the third learning objective according to the first target feature.
[0099] In a possible implementation manner, obtaining the third learning objective includes: aligning the target feature corresponding to the first learning objective and the target feature corresponding to the first key target group through linear transformation;
[0100] It should be noted that the alignment process is to transform the target feature corresponding to the first key target group through a linear transformation layer composed of a weight matrix and a bias vector to ensure that the feature distribution similarity between the target feature corresponding to the first learning objective and the target feature corresponding to the first key target group is maximized;
[0101] Define a loss function for the difference between the target feature corresponding to the first learning objective and the target feature corresponding to the first key target group, and introduce the cosine similarity loss L fa and L fs , calculate the weighted loss balance value L WAM between the target feature corresponding to the first learning objective and the target feature corresponding to the first key target group, specifically expressed as:
[0102] L WAM =β 2 ×L fa +L fs ,
[0103] where, β 2 represents the weighted parameter, and its value is the feature distribution similarity between the target feature corresponding to the first learning objective and the target feature corresponding to the first key target group;
[0104] It should be noted that the weighted parameter β 2The dynamic adjustment rule includes using the class prediction alignment loss, i.e., KL divergence, between the target features corresponding to the first learning objective and the target features corresponding to the first key target group, and the box prediction alignment loss, i.e., L1 regularization, between the target features corresponding to the first learning objective and the target features corresponding to the first key target group, as the adjustment index for the weighting parameter; adaptively adjusting the size based on the influence degree of the feature difference on the detection target, combined with the adjustment index of the weighting parameter, and the value range is between [0, 1], and select the value close to 1 as the weighting parameter β 2 value.
[0105] Specifically, minimize the weighted loss balance value between the target features corresponding to the first learning objective and the target features corresponding to the first key target group, and input the data corresponding to the target features corresponding to the first learning objective into the alignment loss adjustment model to obtain the third learning objective.
[0106] S6: Optimize the overall objective: perform an overall optimization operation based on the second learning objective and the third learning objective, and the overall optimization operation is used to obtain the second target feature.
[0107] Specifically, perform splicing and integration along the feature dimension direction based on the second learning objective and the third learning objective; expand the feature extraction network architecture, add multiple convolutional layers and fully connected layers, and perform feature extraction and fusion on the spliced and integrated data, and convert the features into fixed-length feature vectors through the global average pooling layer to obtain the second target feature.
[0108] It should be noted that the second learning objective is to introduce the maximum entropy loss and the feature diversity loss and use a specific loss function to guide the fusion of relevant feature information, and the third learning objective is to dynamically adjust the alignment loss weight based on the feature similarity according to the weighting parameter, and use the alignment loss adjustment model to process and fuse the original features and the enhanced features.
[0109] In a possible implementation manner, obtaining the second target feature includes: based on the detection loss L of the convolutional neural network det of the convolutional neural network, the diversity loss balance value L DLM and the weighted loss balance value L WAM , calculate the overall optimization loss value L t , which is specifically expressed as:
[0110] L t = L det + α × ln(L DLM ) + β × (L WAM / e),
[0111] where α represents the diversity loss balance weight, β represents the weighted loss balance weight, and e represents a constant with a value of 2.718.
[0112] It should be noted that the feature data obtained by splicing and integrating the second learning objective and the third learning objective along the feature dimension direction is batch-operated for overall optimization according to a preset batch size, and the initial second target feature is obtained through forward propagation; based on the overall optimization loss value L t Update the corresponding learning model in the balanced domain detection system through the backpropagation algorithm, and through repeated training for multiple epochs until the overall optimization loss value L t is minimized, and the obtained feature is the second target feature.
[0113] S7: Evaluate and detect the target learning model: According to the second target feature through the corresponding learning model in the balanced domain detection system, perform target detection operations on the corresponding output results of the learning model, and the target detection operations are used to obtain the detection report of the learning model.
[0114] Specifically, input the second target feature into the corresponding learning model in the balanced domain detection system. The corresponding learning model in the balanced domain detection system performs target detection operations on the second target feature. The target detection operations are to obtain the data category label, feature information, and confidence score corresponding to the output result of the learning model; and calculate the evaluation index value of the corresponding output result of the learning model. The evaluation index value includes the average precision of each data category and the corresponding quantities of true positives, false positives, and false negatives to evaluate the generalization ability of the corresponding learning model in the balanced domain detection system in different domains; integrate the detection results of the learning model data and the calculated evaluation index values to obtain the detection report of the learning model.
[0115] It should be noted that the detection report of the learning model includes the class prediction, confidence, location coordinate information detected by the learning model, and the calculated evaluation index values.
[0116] As shown in the appendix Figure 2 A system for achieving target detection by balancing domain diversity and invariance, including a system operation database, a system central processing module, and a user information terminal, further including: an image data acquisition module, an image data preprocessing module, a feature extraction module, a diversity learning module, a weighted alignment module, a target overall optimization module, and a model evaluation and detection module.
[0117] Image data acquisition module: Used to respond to the corresponding learning model in the balanced domain detection system to obtain the first learning objective of the target learning model;
[0118] Image data preprocessing module: Used to perform preprocessing operations on the first learning objective transmitted by the image data acquisition module, and the preprocessing operations are used to obtain the first key target group corresponding to the first learning objective;
[0119] Feature extraction module: used to perform feature extraction operations on the first key target group transmitted by the image data preprocessing module, and the feature extraction operations are used to obtain the first target feature corresponding to the first learning target;
[0120] Diversity learning module: used to obtain a diversity enhancement model, and based on the diversity enhancement model and the first target feature transmitted by the feature extraction module, obtain the second learning target;
[0121] Weighted alignment module: used to obtain an alignment loss adjustment model, and based on the alignment loss adjustment model and the first target feature transmitted by the feature extraction module, obtain the third learning target;
[0122] Overall target optimization module: perform overall optimization operations based on the second learning target transmitted by the diversity learning module and the third learning target transmitted by the weighted alignment module, and the overall optimization operations are used to obtain the second target feature;
[0123] Model evaluation and detection module: used to perform target detection operations on the output results corresponding to the learning model according to the second target feature transmitted by the overall target optimization module in the balanced domain detection system, and the target detection operations are used to obtain the detection report of the learning model;
[0124] The system operation database includes all data texts of the balanced domain detection system, and real-time collects the information texts output by each module. The system central processing module is used to control the information text instructions output by each module in the control method, and the user information terminal is an information output device for receiving the balanced domain detection system.
[0125] Secondly: In the attached drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments of the present disclosure are involved. Other structures can refer to the general design. Without conflict, the same embodiment and different embodiments of the present invention can be combined with each other;
[0126] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for object detection by balancing domain diversity and invariance, characterized in that: include: S1: Collecting model data: In response to the corresponding learning model in the balance field detection system, obtaining a first learning target of the target learning model; S2: preprocessing model data: performing a preprocessing operation on the first learning objective, where the preprocessing operation is used to obtain a first key objective group corresponding to the first learning objective; S3: Extract key features: perform a feature extraction operation on the first key target group, where the feature extraction operation is used to obtain a first target feature corresponding to the first learning target; S4: Enhance diversity: Obtain a diversity enhancement model, and obtain a second learning target through the first target feature according to the diversity enhancement model; S5: Dynamic weight alignment: obtain an alignment loss adjustment model, and obtain a third learning target through the first target feature according to the alignment loss adjustment model; S6: Optimize the whole target: perform a whole optimization operation based on the second learning target and the third learning target, and the whole optimization operation is used to obtain the second target feature; S7: Evaluate and detect the target learning model: perform a target detection operation on the corresponding output result of the learning model through the second target feature according to the corresponding learning model in the balance field detection system, and the target detection operation is used to obtain a detection report of the learning model; 2. The method for realizing target detection by balancing domain diversity and invariance according to claim 1, characterized in that: The step S2, obtaining a first key target group corresponding to the first learning target, specifically includes: Performing a data enhancement operation on the single-domain data corresponding to the first learning target, where the data enhancement operation is a target conversion operation of the single-domain data corresponding to the first learning target, and the target conversion operation includes a random cropping operation, a color conversion operation, and a geometric transformation operation; A401: Random cropping operation: Randomly crop the single domain image corresponding to the first learning target according to a preset ratio, and the cropped image still contains the key feature part of the target; A402: Color transformation operation: the color channel of the single-domain image corresponding to the first learning target is increased or decreased according to a preset ratio to simulate the color performance of the target object under different lighting and different environmental tones; A403: Geometric transformation operation: With the image center as the rotation point, randomly rotate the single-domain image corresponding to the first learning target within a preset angle range, and synchronously transform and adjust the annotations of the target features; at the same time, perform horizontal flipping, vertical flipping, and horizontal and vertical combined flipping operations on the single-domain image corresponding to the first learning target.
3. The method for realizing target detection by balancing domain diversity and invariance according to claim 1, characterized in that: The feature extraction operation in S3 is to input the first key target group into the feature extraction network to generate a feature graph corresponding to the first learning target and a feature graph corresponding to the first key target group; The feature extraction network consists of a preset base detector and a learning network.
4. The method for realizing target detection by balancing domain diversity and invariance according to claim 1, characterized in that: The step S3, obtaining a first target feature corresponding to the first learning target, specifically includes: Based on the preset mapping and sampling rules, the key feature regions corresponding to the multiple anchor point mechanisms are converted into the feature PI representation of the fixed anchor point mechanism to obtain the target feature z corresponding to the first learning goal s The target feature z corresponding to the first key target group a ; The target feature corresponding to the first learning target is obtained based on the feature graph Fs corresponding to the first learning target and the feature graph Fa corresponding to the first key target group. Specifically, it is expressed as: With s =RP(F s ,PI), Among them, RP represents the extraction function existing in the RoI-Pooling layer, and PI represents the characteristics of the fixed anchor mechanism corresponding to the key feature area; The target characteristics corresponding to the first key target group are specifically expressed as: With a =RP(F a ,PI), Among them, RP represents the extraction function existing in the RoI-Pooling layer, and PI represents the characteristics of the fixed anchor mechanism corresponding to the key feature area.
5. The method for realizing target detection by balancing domain diversity and invariance according to claim 1, characterized in that: The step S4, obtaining the second learning objective, specifically includes: Separate the unique feature z from the first target feature d , and increase the target classification loss L c , maximum entropy loss L H , and feature diversity loss L FD , calculate the diversity loss balance value L of the unique feature and the target feature DLM , specifically expressed as: L DLM =L C +L H +β1×L FD , Among them, β1 represents the hyperparameter of the balance loss.
6. The method for realizing target detection by balancing domain diversity and invariance according to claim 1, characterized in that: The step S5, obtaining the third learning objective, specifically includes: Introducing cosine similarity loss L fa and L fs , calculate the weighted loss balance value L of the target feature corresponding to the first learning objective and the target feature corresponding to the first key target group WAM , specifically expressed as: L WAM =β2×L fa +L fs , Among them, β2 is represented as a weighting parameter, and its value is the feature distribution similarity between the target feature corresponding to the first learning target and the target feature corresponding to the first key target group.
7. The method for realizing target detection by balancing domain diversity and invariance according to claim 1, characterized in that: The step S5, obtaining the third learning objective, specifically includes: Dynamic adjustment of weighted parameters: taking the class prediction alignment loss of the target feature corresponding to the first learning objective and the target feature corresponding to the first key target group, and the box prediction alignment loss of the target feature corresponding to the first learning objective and the target feature corresponding to the first key target group as adjustment indicators of the weighted parameters; Based on the influence of feature differences on the detected target, combined with the adjustment index of the weighted parameter, the size is adaptively adjusted, the value range is between [0, 1], and the value close to 1 is selected as the value of the weighted parameter.
8. The method for realizing target detection by balancing domain diversity and invariance according to claim 1, characterized in that: The step S6, obtaining the second target feature, specifically includes: Detection loss L based on convolutional neural network det , Diversity loss balance value L DLM , and the weighted loss balance value L WAM , calculate the overall optimization loss value L t , specifically expressed as: L t =L det +α×ln(L DLM )+β×(L WAM / e), Among them, α represents the diversity loss balance weight, β represents the weighted loss balance weight, and e represents a constant with a value of 2.718.