Precise image recognition system based on deep learning driving

Through generative adversarial networks and cross-domain transfer learning technology, the problem of the deep learning image recognition model performing poorly in the target domain is solved, and the efficient adaptability and robustness of the model under different environments and conditions is achieved.

CN120070994APending Publication Date: 2025-05-30TIANHE COLLEGE GUANGDONG POLYTECHNIC NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510153809.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The deep learning image recognition model performs poorly in the target domain, mainly due to the difference in feature distribution between the source domain and the target domain, which affects the generalization ability and accuracy of the model.

Method used

A cross-domain adaptive feature mapping module based on a generative adversarial network is adopted to generate a shared feature space between the source domain and the target domain through domain discrimination and adversarial training mechanisms. Then, the multi-level domain discriminant network and hybrid domain feature adjustment network are used in the cross-domain transfer learning module to optimize the feature representation in the target domain.

Benefits of technology

Effectively eliminate the distribution differences between the source domain and the target domain, improve the adaptability and robustness of the deep learning model in the target domain, and significantly improve the performance of the target recognition system under different environments and conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070994A_ABST
    Figure CN120070994A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to an accurate image recognition system based on deep learning driving, which comprises a cross-domain adaptive feature mapping module, a cross-domain transfer learning module and a target classification and detection module, mapping the source domain feature space to a target domain feature space through a feature mapping mode, and generating a shared feature space between the source domain and the target domain through an adversarial training mode based on a generative adversarial network; the cross-domain transfer learning module applies the obtained feature mapping in the shared feature space to identification of target domain data; and the target classification and detection module performs target classification and detection based on the feature representation optimized by the cross-domain transfer learning module, and performs target classification and positioning in a target domain image. According to the invention, the cross-domain mapping mechanism significantly improves the robustness of the target recognition system in different environments and different conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image precise recognition system driven by deep learning. Background Art

[0002] With the rapid development of deep learning technology, image recognition technology has been widely applied in multiple fields. In these applications, as one of the core tasks, image precise recognition faces challenges such as extracting meaningful information from images, classifying and locating target objects, etc. Traditional image recognition methods usually rely on a large amount of labeled data for training. However, in practical applications, due to changes in factors such as scenes, environments, lighting, and perspectives, the difference in feature distributions between the source domain and the target domain will seriously affect the generalization ability and accuracy of the model.

[0003] In image recognition tasks, there are often differences in feature distributions between the source domain (i.e., the training set) and the target domain (i.e., the test set or the data in practical applications), resulting in poor performance of deep learning models in the target domain. For example, the training data may come from a specific lighting condition, angle, or background, while the test data comes from different environments or conditions, which makes the performance of the source domain model unable to be directly transferred to the target domain. Therefore, cross-domain adaptation has become an important research direction in deep learning image recognition. Especially in the case of lacking labeled data in the target domain, how to improve the recognition accuracy in the target domain has become a key technical problem.

[0004] Although the current mainstream deep learning methods can be effectively trained in the source domain through a large amount of labeled data, how to transfer the source domain model to the target domain, especially when the data distribution in the target domain varies greatly, is still an urgent problem to be solved. Common feature extraction methods (such as convolutional neural networks) can extract effective features from source domain data when dealing with cross-domain problems. However, these features are often not directly applicable to the target domain, resulting in a decline in the transfer performance of the model. Summary of the Invention

[0005] The present invention provides an image precise recognition system driven by deep learning.

[0006] An image precise recognition system driven by deep learning includes a cross-domain adaptive feature mapping module, a cross-domain transfer learning module, and a target classification and detection module, wherein:

[0007] The cross-domain adaptive feature mapping module extracts image features from different source domain image data and maps them to the target domain feature space through a feature mapping method. Based on a generative adversarial network, a shared feature space between the source domain and the target domain is generated through an adversarial training method, specifically including:

[0008] Domain discrimination, which is used to determine whether the input image features come from the source domain or the target domain;

[0009] An adversarial training mechanism that gradually makes the feature spaces of the source domain and the target domain closer by optimizing the game process between the discriminative network and the generative network;

[0010] The cross - domain transfer learning module applies the feature mapping in the obtained shared feature space to the recognition of target - domain data, and optimizes the feature representation in the target domain through a multi - level domain discrimination network and a hybrid domain feature adjustment network;

[0011] The target classification and detection module performs target classification and detection based on the feature representation optimized by the cross - domain transfer learning module, classifies and locates the target in the target - domain image, and outputs the final classification label and the target detection box.

[0012] Optionally, the cross - domain adaptive feature mapping module extracts high - dimensional image features from the source - domain image data through a convolutional neural network. The image features include edges, textures, shapes, and color distributions. Taking the source - domain image I S as the input, after passing through the convolutional layer, the feature F S is obtained: F S = CNN(I S ; θ C ), where I S ∈R H×W×C represents the input source - domain image, H is the image height, W is the image width, C is the number of image channels, θ C represents the parameters (weights) of the convolutional neural network, and F S ∈R H′×W′×C′ is the feature output by the convolutional layer, where H′, W′ are the spatial dimensions of the feature, and C′ is the number of channels of the feature.

[0013] Optionally, the goal of the domain discrimination is to output a discrimination result according to the input feature F S and the feature F T of the target - domain image I T , which is expressed as:

[0014]

[0015]

[0016] where D(X; θ D ) represents the domain discrimination network, θ D is the parameter of the domain discrimination network, X includes F S , F T , and It is the discriminant result of the output, indicating that the input image features are respectively from the source domain y = 0 or the target domain y = 1;

[0017] The discriminant network classifies the input image features F S or F T to output a binary label and indicating that they belong to the source domain or the target domain respectively.

[0018] Optionally, the objective of the adversarial training mechanism is to optimize the game process between the generator network and the discriminant network through the generative adversarial network, making the feature representations of the source domain and the target domain gradually approach. The generator network G is used to generate feature mappings, and the discriminant network D judges its source, and domain adaptation is achieved by minimizing the adversarial loss function;

[0019] The loss function of the generative adversarial network is defined as: the loss of the discriminant network the loss of the generator network

[0020] Optionally, through the adversarial training mechanism, the generator network G generates a shared feature space between the source domain and the target domain. The generator network maps the source domain feature F S to the target domain feature F T ′, and the following relationship is expected:

[0021] F T ′ = G(F S ; θ G ), where F T ′ is the generated target domain feature, θ G is the parameter of the generator network. After adversarial training, the generated feature F T ′ will have a similar distribution to the target domain feature F T so that the feature spaces of the source domain and the target domain can be shared.

[0022] Optionally, the multi-level domain discriminant network, through the multi-level learning strategy, identifies the differences between the source domain and the target domain at multiple feature levels and performs feature alignment. The multi-level learning strategy is expressed as:

[0023] where and respectively represent the feature representations of the source domain and the target domain at the l-th layer, D l (·) represents the domain discriminant network at the l-th layer, and outputs the probability that the input feature belongs to the source domain or the target domain. is the multi-level domain discriminant loss, which measures the difference between the source domain and the target domain at each layer, and L is the number of layers of the feature extraction network.

[0024] Optionally, the hybrid domain feature adjustment network adjusts features at different levels according to the domain differences at each level to ensure that the feature representation can adapt to the task requirements of the target domain.

[0025] Optionally, the hybrid domain feature adjustment network is expressed as:

[0026] where represents the source domain features at the l-th layer and the target domain features through the feature adjustment network for the result of fusion and adjustment, the adjustment includes weighted fusion, alignment or balance of the scale and distribution of features, and α l is the weight of each layer of feature adjustment, indicating the importance of each layer in the adjustment process, is the target domain feature representation after adjustment and optimization, and is finally used for task learning in the target domain, including target classification or detection.

[0027] Optionally, the target classification and detection module specifically includes:

[0028] Feature input: The target classification and detection module receives the optimized feature representation from the cross-domain transfer learning module as input;

[0029] Target classification and localization: In the target classification and detection module, the input feature passes through a fully connected layer or a convolutional layer, and target classification and target localization are respectively performed through a classification head (for classification tasks) and a regression head (for localization tasks), expressed as:

[0030]

[0031] where y class is the target classification result obtained through the Softmax function, representing the probability distribution of the target class, and y bbox is the regression output, representing the predicted coordinates of the target box, using four values: the upper left corner coordinates (x min , y min ) and the lower right corner coordinates (x max , y max ), is the target domain feature representation output from the cross-domain transfer learning module, and W c and b c are the weights and biases of the classification head, and W b and b b are the weights and biases of the regression head;

[0032] Output: Based on the optimized feature representation of the target domain image, output the class label of each detected target and the corresponding detection box.

[0033] Optionally, the output is expressed as: where N is the number of targets, class i is the classification label of the i-th target, (x min,i , y min,i , x max,i , y max,i ) are the coordinates of the detection box of the i-th target.

[0034] Advantages of the present invention:

[0035] In the present invention, based on the generative adversarial network framework, a shared feature space between the source domain and the target domain is automatically generated through an adversarial training mechanism, which can effectively eliminate the distribution difference between the source domain and the target domain. Through the game process of using the domain discriminant network and the generative network, the feature spaces of the source domain and the target domain gradually approach, thereby improving the adaptability of the deep learning model in the target domain. The cross-domain mapping mechanism significantly enhances the robustness of the target recognition system under different environments and conditions, and avoids the performance degradation caused by the data distribution difference in the traditional method.

[0036] In the present invention, through the multi-level domain discriminant network and the hybrid domain feature adjustment network in the cross-domain transfer learning module, the system can accurately adjust and align features at different levels. The multi-level domain discriminant network adjusts the features in a fine-grained manner by identifying the differences between the source domain and the target domain at multiple feature levels, thereby enhancing the expression ability of the target domain features. At the same time, the hybrid domain feature adjustment network adjusts the learning strategy of the feature layer according to the differences between different levels of domains to ensure that the feature representation can better adapt to the task requirements of the target domain. The network design significantly enhances the learning effect of the target domain task, especially in the target detection and classification tasks, improving the precision and accuracy of the model.

[0037] In the present invention, by using the optimized feature representation of the target domain for target classification and detection, the target recognition system can accurately classify and locate targets in the target domain image. The target domain features after cross-domain transfer learning and feature adjustment have higher expression ability, and the classification head and regression head greatly improve the effect of classification and positioning on this basis. Compared with the traditional method, the present invention can reduce the feature difference between the source domain and the target domain, reduce the negative transfer phenomenon in transfer learning, and thus improve the precision and recall rate of target detection. Description of the Drawings

[0038] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0039] Figure 1 Schematic diagram of the system function modules of the embodiment of the present invention;

[0040] Figure 2 Schematic diagram of the generative adversarial network of the embodiment of the present invention. Specific implementation manners

[0041] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the accompanying drawings are only for more specifically describing the embodiments and are not intended to specifically limit the present invention.

[0042] It should be pointed out that in the specification, it is mentioned that "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc. indicate that the described embodiment may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. Additionally, when combining an embodiment to describe a specific feature, structure or characteristic, implementing such feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge scope of those skilled in the relevant art.

[0043] Generally, terms can be understood at least in part from their use in the context. For example, at least in part depending on the context, the term "one or more" used herein can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily aiming to convey a set of exclusive factors, but rather, at least in part depending on the context, allowing for the existence of other factors that may not be explicitly described.

[0044] As Figure 1 - Figure 2 shown, an image precise recognition system driven by deep learning includes a cross-domain adaptive feature mapping module, a cross-domain transfer learning module, and a target classification and detection module, wherein:

[0045] The cross - domain adaptive feature mapping module extracts image features from image data of different source domains (sources), and maps them to the target domain feature space through feature mapping to address the distribution difference problem between the source domain and the target domain. Based on the generative adversarial network, a shared feature space between the source domain and the target domain is generated through adversarial training, specifically including:

[0046] Domain discrimination, which is used to determine whether the input image features come from the source domain or the target domain;

[0047] The adversarial training mechanism makes the feature spaces of the source domain and the target domain gradually approach by optimizing the game process between the discriminative network and the generative network, improving the adaptability of the model in the target domain;

[0048] The cross - domain transfer learning module applies the feature mapping in the obtained shared feature space to the recognition of target domain data, optimizes the feature representation in the target domain through a multi - level domain discriminative network and a hybrid domain feature adjustment network, and enhances the learning effect of the target domain task;

[0049] The target classification and detection module performs target classification and detection based on the feature representation optimized by the cross - domain transfer learning module, conducts target classification and localization in the target domain image, and outputs the final classification label and target detection box.

[0050] The cross - domain adaptive feature mapping module extracts high - dimensional image features from source domain image data through a convolutional neural network. The image features include edges, textures, shapes, and color distributions. Taking the source domain image I S as the input, after passing through the convolutional layer, the feature map F S is obtained: F S = CNN(I S ; θ C ), where I S ∈R H×W×C represents the input source domain image, H is the image height, W is the image width, C is the number of image channels, θ C represents the parameters (weights) of the convolutional neural network, and F S ∈R H′×W′×C′ is the feature map output by the convolutional layer, where H′, W′ are the spatial dimensions of the feature map, and C′ is the number of channels of the feature map.

[0051] The goal of domain discrimination is to output a discrimination result based on the input feature F S and the feature F T of the target domain image I T , expressed as:

[0052]

[0053]

[0054] Among them, D(X; θ D ) represents the discriminative network of the domain, and θ D is the parameter of the domain discriminative network. X includes F S and F T . And are the output discriminative results, indicating that the input image features are respectively from the source domain y = 0 or the target domain y = 1;

[0055] The discriminative network classifies the input image features F S or F T to output a binary label and indicating that they respectively belong to the source domain or the target domain.

[0056] The goal of the adversarial training mechanism is to optimize the game process between the generator network and the discriminator network through the generative adversarial network, making the feature representations of the source domain and the target domain gradually approach. The generator network G is used to generate feature mappings, and the discriminator network D judges its source, and realizes domain adaptation by minimizing the adversarial loss function;

[0057] The loss function of the generative adversarial network is defined as: the loss of the discriminator network the loss of the generator network

[0058] The loss function of the discriminator network is expressed as:

[0059] Among them, p S and p T respectively represent the feature distributions of the source domain and the target domain. logD(F S ) is the predicted probability of the discriminator network for the source domain features, indicating the probability that it predicts as the source domain, and log(1 - D(F T )) is the predicted probability of the discriminator network for the target domain features, indicating the probability that it predicts as the target domain;

[0060] The loss function of the generator network is expressed as: Among them, G(F S ) represents the source domain feature mapping generated by the generator network, which will be close to the features of the target domain after optimization, represents the expectation operation.

[0061] The goal of the generator network is to make the discriminator network unable to distinguish between source domain features and target domain features, and finally make the feature distributions of the source domain and the target domain similar.

[0062] Through the adversarial training mechanism, the generator network G generates a shared feature space between the source domain and the target domain. The generator network maps the source domain features FS Mapped to the target domain feature F T ′, the following relationship is expected:

[0063] F T ′ = G(F S ; θ G ), where F T ′ is the generated target domain feature, θ G is the parameter of the generation network. After adversarial training, the generated feature F T ′ will have a similar distribution to the target domain feature F T , enabling the sharing of the feature spaces of the source domain and the target domain.

[0064] The multi-level domain discriminant network, through a multi-level learning strategy, identifies the differences between the source domain and the target domain at multiple feature levels and performs feature alignment. The goal of the multi-level domain discriminant network is to further align the feature spaces of the source domain and the target domain by discriminating features at different levels, reducing the distribution difference between the source domain and the target domain. The multi-level learning strategy is expressed as:

[0065] where and respectively represent the feature representations of the source domain and the target domain at the l-th layer. D l (·) represents the domain discriminant network at the l-th layer, outputting the probability that the input feature belongs to the source domain or the target domain. is the multi-level domain discriminant loss, measuring the difference between the source domain and the target domain at each layer. L is the number of layers of the feature extraction network.

[0066] The hybrid domain feature adjustment network adjusts the features at different levels according to the domain differences at each level, ensuring that the feature representation can adapt to the task requirements of the target domain. The hybrid domain feature adjustment network further adjusts the features of the source domain and the target domain by fusing information at different levels, avoiding the performance degradation of the model in the target domain caused by excessive domain differences.

[0067] The hybrid domain feature adjustment network is expressed as:

[0068] where represents the result of fusing and adjusting the source domain feature and the target domain feature through the feature adjustment network at the l-th layer. This adjustment includes weighted fusion, aligning or balancing the scales and distributions of the features. α l is the weight of the feature adjustment for each layer, indicating the importance of each layer in the adjustment process. It is the target domain feature representation after adjustment and optimization, and is finally used for task learning in the target domain, including object classification or detection.

[0069] The object classification and detection module specifically includes:

[0070] Feature input: The object classification and detection module receives the optimized feature representation from the cross-domain transfer learning module as input. This feature representation has been optimized by the multi-level domain discriminative network and the hybrid domain feature adjustment network, and has adapted to the task requirements of the target domain;

[0071] Object classification and localization: In the object classification and detection module, the input feature passes through a fully connected layer or a convolutional layer, and object classification and object localization are respectively performed through a classification head (for classification tasks) and a regression head (for localization tasks), expressed as:

[0072]

[0073] where y class is the object classification result obtained through the Softmax function, representing the probability distribution of the object category, and y bbox is the regression output, representing the predicted coordinates of the object bounding box, using four values: the upper left corner coordinates (x min , y min ) and the lower right corner coordinates (x max , y max ), is the target domain feature representation output from the cross-domain transfer learning module, W c and b c are the weights and biases of the classification head, and W b and b b are the weights and biases of the regression head;

[0074] Output: According to the optimized feature representation of the target domain image, output the class label of each detected object and the corresponding detection bounding box.

[0075] The output is expressed as: where N is the number of objects, class i is the classification label of the i-th object, and (x min,i , y min,i , x max,i , y max,i ) are the coordinates of the detection bounding box of the i-th object.

[0076] The present invention encompasses any alternatives, modifications, equivalent methods, and solutions that are made to the essence and scope of the present invention. For the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention even without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0077] The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A deep learning-driven image recognition system, characterized in that: It includes a cross-domain adaptive feature mapping module, a cross-domain transfer learning module, and a target classification and detection module, among which: The cross-domain adaptive feature mapping module extracts image features from different source domain image data, and maps them to the target domain feature space through feature mapping, and generates a shared feature space between the source domain and the target domain through adversarial training based on a generative adversarial network, specifically including: Domain discrimination, used to determine whether the input image features come from the source domain or the target domain; The adversarial training mechanism optimizes the game process between the discriminant network and the generative network, so that the feature space of the source domain and the target domain gradually approach each other; The cross-domain transfer learning module applies the obtained feature mapping in the shared feature space to the recognition of the target domain data, and optimizes the feature representation in the target domain through a multi-level domain discrimination network and a mixed domain feature adjustment network; The target classification and detection module performs target classification and detection based on the feature representation optimized by the cross-domain transfer learning module, performs target classification and positioning in the target domain image, and outputs the final classification label and target detection frame.

2. According to claim 1, the image accurate recognition system based on deep learning is characterized in that: The cross-domain adaptive feature mapping module extracts high-dimensional image features from the source domain image data through a convolutional neural network. The image features include edge, texture, shape and color distribution. S As input, after the convolution layer, the feature F is obtained S : F S =CNN(I S θ C ), where I S ∈R H×W×C represents the input source domain image, H is the image height, W is the image width, C is the number of image channels, θ C represents the parameters of the convolutional neural network, F S ∈R H′×W′×C′ is the feature output of the convolutional layer, where H′, W′ are the spatial dimensions of the feature, and C′ is the number of channels of the feature.

3. The image accurate recognition system based on deep learning drive according to claim 2 is characterized in that: The goal of the domain discrimination is to identify the input features F S and the target domain image I T Features F T The output discrimination result is expressed as: Where D(X;θ D ) represents the discriminant network of the domain, θ D are the parameters of the domain discrimination network, X includes F S 、F T , and is the output discrimination result, indicating that the input image features come from the source domain y=0 or the target domain y=1; The discriminant network uses the input image feature F S or F T Classify and output a binary label and Indicates that it belongs to the source domain or the target domain respectively.

4. The image precision recognition system based on deep learning drive according to claim 3 is characterized in that: The adversarial training mechanism aims to optimize the game process between the generative network and the discriminative network through the generative adversarial network, so that the feature representations of the source domain and the target domain gradually approach each other, use the generative network G to generate feature maps, and the discriminative network D to determine its source, and achieve domain adaptation by minimizing the adversarial loss function; The loss function of the generative adversarial network is defined as: the loss of the discriminant network The loss of the generator network 5. The image precision recognition system based on deep learning drive according to claim 4 is characterized in that: Through the adversarial training mechanism, the generative network G generates a shared feature space between the source domain and the target domain. The generative network converts the source domain feature F S Mapped to the target domain feature F T ′, the following relationship is expected: F T ′=G(F S θ G ), where F T ′ is the generated target domain feature, θ G is the parameter of the generated network. After adversarial training, the generated feature F T ′ will be combined with the target domain feature F T Having similar distribution enables the feature space of the source domain and the target domain to be shared.

6. The image precision recognition system based on deep learning drive according to claim 1, characterized in that: The multi-level domain discrimination network identifies the differences between the source domain and the target domain at multiple feature levels and performs feature alignment through a multi-level learning strategy. The multi-level learning strategy is expressed as: in, and Denote the feature representation of the source domain and the target domain at the lth layer, respectively, l (·) represents the domain discrimination network at layer l, which outputs the probability that the input feature belongs to the source domain or the target domain. It is a multi-level domain discrimination loss that measures the difference between the source domain and the target domain at each layer, and L is the number of layers of the feature extraction network.

7. The image accurate recognition system based on deep learning drive according to claim 6, characterized in that: The hybrid domain feature adjustment network adjusts features at different levels according to the differences in domains at each level to ensure that the feature representation can adapt to the task requirements of the target domain.

8. The image precision recognition system based on deep learning drive according to claim 7, characterized in that: The mixed domain feature conditioning network is expressed as: in, Indicates that at the lth layer, the source domain features and target domain features The result of fusion and adjustment through the feature adjustment network, which includes weighted fusion, alignment or balancing the scale and distribution of features, α l is the weight of feature adjustment of each layer, indicating the importance of each layer in the adjustment process. It is the feature representation of the target domain after adjustment and optimization, and is ultimately used for task learning in the target domain, including target classification or detection.

9. The image precision recognition system based on deep learning drive according to claim 8, characterized in that: The target classification and detection module specifically includes: Feature input: The target classification and detection module receives the optimized feature representation from the cross-domain transfer learning module. As input; Target classification and localization: In the target classification and detection module, input features After a fully connected layer or convolutional layer, the classification head and regression head are used to perform target classification and target positioning respectively, which can be expressed as: Among them, y class is the target classification result obtained by the Softmax function, which represents the probability distribution of the target category, y bbox It is the regression output, which represents the predicted coordinates of the target box, using four values: the upper left corner coordinate (x min ,y min ) and the lower right corner coordinate (x max ,y max ), is the target domain feature representation output from the cross-domain transfer learning module, W c and b c are the weights and biases of the classification head, W b and b b are the weights and biases of the regression head; Output: Based on the optimized feature representation of the target domain image, the category label of each detected target and the corresponding detection box are output.

10. The image precision recognition system based on deep learning drive according to claim 8, characterized in that: The output is represented by: Where N is the number of targets, class i is the classification label of the i-th target, (x min,i ,y min,i ,x max,i ,y max,i ) is the detection box coordinate of the i-th target.

Citation Information

Cited By

  • Character recognition and matching method and system for CAD (Computer Aided Design) terminal diagram and bushing image

    CN121074940A

  • A cad terminal map and sleeve map image character recognition and matching method and system

    CN121074940B