Image recognition method and device, model training method and device, equipment and medium
By dividing the image into multiple blocks and separating the causal and non-causal features, the problem of the reduction in accuracy of deep learning models in new scenarios is solved, and higher recognition stability and accuracy are achieved.
Patent Information
- Application Number
- CN202510841028.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Deep learning models are susceptible to context bias, resulting in reduced image recognition accuracy in new scenarios or new task data, making it difficult to identify causal and irrelevant areas in the image.
The image to be identified is divided into multiple image blocks, the image block features are extracted separately, and the causal and non-causal features are separated through the feature decoupling network, the causal and non-causal features are identified using the causal and effect features, and the non-causal and effect feature interference is eliminated, and the image recognition model is trained using a comprehensive loss function.
It improves the stability and accuracy of image recognition in new scenarios or complex environments, and enhances the ability to capture fine-grained information in local areas.
Smart Images

Figure CN120375098A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and more particularly, to an image recognition method, a model training method, a device, a device, and a medium. Background Art
[0002] With the rapid development of artificial intelligence, the image classification method based on deep learning can extract the visual features of images and achieve efficient classification, and can be widely applied to scenarios such as medical diagnosis, industrial inspection, security monitoring, and autonomous driving.
[0003] In the related art, since the deep learning model is easily affected by context bias and learns false associations, and the model trained through the labels of the entire image is difficult to identify which regions in the image are causally related or irrelevant. In practical applications, when the model faces new scenarios or new task data, the accuracy of image recognition will decrease. Summary of the Invention
[0004] In view of this, this application provides an image recognition method, a model training method, a device, a device, and a medium.
[0005] One aspect of this application provides an image recognition method, including: respectively performing feature extraction processing on multiple image patches of the image to be recognized to obtain multiple image patch features of the image to be recognized; respectively performing feature decoupling on the multiple image patch features to obtain the causal features and non-causal features of each of the multiple image patch features, where the causal features represent the features of the object to be recognized included in the image patch features, and the non-causal features represent the features irrelevant to the object to be recognized in the image patch features; based on the multiple causal features, performing recognition on the image to be recognized to obtain a recognition result.
[0006] Another aspect of this application provides a method for training an image recognition model, including: obtaining a sample training set, where the sample training set includes multiple sample images; respectively performing feature extraction processing on multiple sample image patches of the sample images to obtain multiple sample image patch features of the sample images; inputting the multiple sample image patch features into the feature decoupling network of the initial image recognition model to obtain the corresponding sample causal features and sample non-causal features of each of the multiple sample image patch features, where the sample causal features represent the features of the sample object included in the sample image patch features, and the non-causal features represent the features irrelevant to the sample object in the sample image patch features; inputting the multiple sample causal features into the aggregation classification network of the initial image recognition model to obtain a sample recognition result; training the initial image recognition model based on a comprehensive loss to obtain a trained image recognition model, where the comprehensive loss includes a first loss and a second loss, the first loss represents the deviation between the sample causal features and the sample non-causal features, and the second loss represents the deviation between the sample recognition result and the sample image label.
[0007] Another aspect of the present application provides an image recognition device, including: a feature extraction module, configured to perform feature extraction processing on multiple image patches of the image to be recognized respectively, to obtain multiple image patch features of the image to be recognized, where the multiple image patches represent different regions of the image to be recognized; a feature decoupling network, configured to perform feature decoupling on the multiple image patch features respectively, to obtain the causal feature and non-causal feature of each of the multiple image patch features, where the causal feature represents the feature of the object to be recognized included in the image patch feature, and the non-causal feature represents the feature irrelevant to the object to be recognized in the image patch feature; and a recognition module, configured to recognize the image to be recognized based on the multiple causal features, to obtain a recognition result.
[0008] Another aspect of the present application provides a training device for an image recognition model, including: an acquisition module, configured to acquire a sample training set, where the sample training set includes multiple sample images; a sample feature extraction module, configured to perform feature extraction processing on multiple sample image patches of the sample images respectively, to obtain multiple sample image patch features of the sample images; a sample feature decoupling network, configured to input the multiple sample image patch features into the feature decoupling network of the initial image recognition model, to obtain the corresponding sample causal feature and sample non-causal feature of each of the multiple sample image patch features, where the sample causal feature represents the feature of the sample object included in the sample image patch feature, and the non-causal feature represents the feature irrelevant to the sample object in the sample image patch feature; a sample recognition module, configured to input the multiple sample causal features into the aggregation classification network of the initial image recognition model, to obtain a sample recognition result; and a training module, configured to train the initial image recognition model based on a comprehensive loss, to obtain a trained image recognition model, where the comprehensive loss includes a first loss and a second loss, the first loss represents the deviation between the sample causal feature and the sample non-causal feature, and the second loss represents the deviation between the sample recognition result and the sample image label.
[0009] Another aspect of the present application provides an electronic device, including: one or more processors; and a memory, configured to store one or more computer programs, where the above one or more processors execute the above one or more computer programs to implement the steps of the above method.
[0010] Another aspect of the present application further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0011] Another aspect of the present application further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0012] According to the technical solution of the present application, by dividing the image to be recognized into multiple image blocks and extracting the features of the image blocks respectively, it is convenient to capture the fine-grained information of the local area of the image to be recognized. By decoupling the features of each image block and separating multiple causal features and non-causal features, it is convenient to identify the areas directly related to causality in the image to be recognized. Based on the causal features of multiple image blocks for recognition, the interference of non-causal features is excluded, and the stability and accuracy of image recognition in the face of new scenarios or complex environments are improved. Description of the Drawings
[0013] Through the following description of the embodiments of the present application with reference to the accompanying drawings, the above content and other objects, features and advantages of the present application will become clearer.
[0014] Figure 1 An exemplary system architecture to which the image recognition method and device according to the embodiments of the present application can be applied is shown.
[0015] Figure 2 A flowchart of the image recognition method according to the embodiments of the present application is shown.
[0016] Figure 3 A flowchart of the image recognition method according to another embodiment of the present application is shown.
[0017] Figure 4 A network architecture diagram of the image recognition model according to the embodiments of the present application is shown.
[0018] Figure 5 A flowchart of the training method of the image recognition model according to the embodiments of the present application is shown.
[0019] Figure 6 A flowchart of the training of the image recognition model according to another embodiment of the present application is shown.
[0020] Figure 7 A module diagram of the image recognition device according to the embodiments of the present application is shown.
[0021] Figure 8 A module diagram of the training device of the image recognition model according to the embodiments of the present application is shown.
[0022] Figure 9 A block diagram of an electronic device suitable for implementing the method described above according to the embodiments of the present application is shown. Detailed Embodiments
[0023] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0024] The terms used herein are merely for describing specific embodiments and are not intended to limit the present application. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0026] In cases where expressions such as "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0027] Figure 1 An exemplary system architecture to which an image recognition method and apparatus according to an embodiment of the present application can be applied is shown.
[0028] It should be noted that Figure 1 What is shown is only an example of a system architecture to which the embodiments of the present application can be applied, to help those skilled in the art understand the technical content of the present application, but does not mean that the embodiments of the present application cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which an image recognition method and apparatus can be applied may include a terminal device, but the terminal device can implement the image recognition method and apparatus provided by the embodiments of the present application without interacting with a server.
[0029] As Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0030] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0031] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0032] The server 105 may be a server providing various services, such as a background management server (for example only) that supports the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data, etc. obtained or generated according to user requests) to the terminal device.
[0033] It should be noted that the image recognition method provided by the embodiments of the present application can generally be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103. Correspondingly, the image recognition device provided by the embodiments of the present application can also be set in the first terminal device 101, the second terminal device 102, and the third terminal device 103.
[0034] Alternatively, the image recognition method provided by the embodiments of the present application can generally also be executed by the server 105. Correspondingly, the image recognition device provided by the embodiments of the present application can generally be disposed in the server 105. The image recognition method provided by the embodiments of the present application can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the image recognition device provided by the embodiments of the present application can also be disposed in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0035] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0036] In the technical solution of the present application, the collection, storage, use, processing, transmission, provision, disclosure, and application, etc., of the user's personal information all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.
[0037] In the technical solution of the present application, before obtaining or collecting the user's personal information, the authorization or consent of the user has been obtained.
[0038] It should be noted that the sequence numbers of the various operations in the following methods are only used as representations of the operations for description, and should not be regarded as indicating the execution order of the various operations. Unless explicitly stated, the method does not need to be executed exactly in the order shown.
[0039] Figure 2 shows a flowchart of the image recognition method according to an embodiment of the present application.
[0040] As Figure 2 shown, the method includes: operations S210 to S230.
[0041] In operation S210, feature extraction processing is respectively performed on multiple image blocks of the image to be recognized, and multiple image block features of the image to be recognized are obtained.
[0042] In operation S220, feature decoupling is respectively performed on the multiple image block features to obtain the causal features and non-causal features of each of the multiple image block features. The causal features represent the features of the object to be recognized included in the image block features, and the non-causal features represent the features irrelevant to the object to be recognized in the image block features.
[0043] In operation S230, based on the multiple causal features, the image to be recognized is recognized to obtain a recognition result.
[0044] Exemplarily, the image to be recognized can be divided into grids of a fixed size based on a preset rule to obtain a plurality of image patches. This application is not limited thereto, and multiple image patches can also be extracted from the image to be recognized by sliding a window with a fixed step size and size; or the image to be recognized can be dynamically divided into multiple image patches based on the image content. By dividing the image to be recognized into multiple image patches and extracting the features of the image patches respectively, it is convenient to capture the fine-grained information of the local area of the image to be recognized.
[0045] Feature decoupling refers to decomposing the image patch feature into multiple independent or low-coupled sub-features; such as causal features and non-causal features, to separate the core information related to the object to be recognized and the irrelevant information in the image patch feature.
[0046] Causal features refer to the part of the image patch feature that is directly related to the object to be recognized, which can reflect the essential attributes of the object (such as shape, structure), and remain stable under different environments or interferences (such as light changes, background changes). Non-causal features refer to the part of the image patch feature that is irrelevant to the object to be recognized, which can be introduced by environmental factors, noise, or biases in the data acquisition process.
[0047] Exemplarily, complementary masks can be generated based on a spatial attention mechanism to extract causal and non-causal features respectively. This application is not limited thereto, and a dual-channel adversarial network can also be constructed to separate causal features and non-causal features using adversarial training.
[0048] Exemplarily, multiple causal features can be fused and then input into a pre-trained classifier to obtain the recognition result of the image.
[0049] According to the embodiments of the present application, by dividing the image to be recognized into multiple image patches and extracting the features of the image patches respectively, it is convenient to capture the fine-grained information of the local area of the image to be recognized. By decoupling each image patch feature and separating multiple causal features and non-causal features, it is convenient to identify the area directly related to the cause in the image to be recognized. Based on the causal features of multiple image patches for recognition, excluding the interference of non-causal features, the stability and accuracy of image recognition when facing new scenarios or complex environments are improved.
[0050] According to the embodiments of the present application, performing feature extraction processing on multiple image patches of the image to be recognized to obtain multiple image features of the image to be recognized may include: respectively performing encoding processing on multiple image patches of the image to be recognized to obtain encoding vectors of each of the multiple image patches; mapping the encoding vectors of each of the multiple image patches into vectors of a target dimension to obtain image patch features.
[0051] Exemplarily, a Vision Transformer (ViT) can be utilized to encode the image blocks to be recognized, obtaining multiple image block features. For example, assume the image is , whose height is , width is , and number of channels is . Using ViT, the image is segmented into non-overlapping image blocks. The size of each image block can be, for example, , being the size of the image block. Each image block is flattened into a vector with a length of , and then through a linear projection layer, the vector is mapped to the feature space of the target dimension, obtaining multiple image block features . This application is not limited thereto.
[0052] Alternatively, a Convolutional Neural Network (CNN) can also be used to process the image to be recognized into multiple image blocks. CNN can be used to extract the features of the image to be recognized, obtaining a feature map . Then the feature map is divided into image blocks. Each image block undergoes a max pooling operation to obtain a vector of the target dimension, obtaining multiple image block features .
[0053] According to the embodiments of the present application, by independently encoding each image block and then mapping the encoded vector to the target dimension, such as a higher-dimensional feature space, it is possible to provide a richer feature representation ability. In the high-dimensional feature space, the discrimination between features may be greater, and it can better represent the fine-grained features of the image block.
[0054] According to the embodiments of the present application, based on the respective causal features of multiple image block features, the image to be recognized is recognized, and the obtained recognition result may include: obtaining multiple target causal features from the respective causal features corresponding to multiple image block features based on the first weight information. The first weight information characterizes the proportion degree of the causal feature in the corresponding image block feature. Based on the multiple target causal features, the image to be recognized is recognized to obtain the recognition result.
[0055] Exemplarily, an attention mechanism can be used to automatically assign a first weight information to each causal feature. The greater the weight of the first weight information, the greater the contribution of the causal feature to the image block feature.
[0056] According to the first weight information, the causal features corresponding to each image block feature are screened to obtain multiple target causal features.
[0057] Optionally, a weight threshold can be set. If the first weight information is greater than the weight threshold, the causal feature is regarded as the target causal feature. If the first weight information is less than the weight threshold, the causal feature is not regarded as the target causal feature and is not used for subsequent image recognition.
[0058] According to the embodiments of the present application, since the first weight information can represent the proportion of the causal feature in the corresponding image patch feature, the image patches closely related to the causal relationship in the image can be screened according to the first weight information, and the important causal contribution regions can be located. The regions that do not contain the object to be recognized are filtered, improving the subsequent processing efficiency and reducing interference.
[0059] According to the embodiments of the present application, performing feature decoupling on multiple image patch features respectively to obtain the causal features and non-causal features of each of the multiple image patch features may include: processing the multiple image patch features by using the feature decoupling network of the image recognition model to obtain the causal features and non-causal features of each of the multiple image patch features. The feature decoupling network includes a first attention network and a second attention network. The first attention network is used to extract the causal features of each of the multiple image patches and the first weight information corresponding to each causal feature. The second attention network is used to extract the non-causal features of each of the multiple image patches.
[0060] The first attention network is used to identify the feature dimensions in each image patch that have a causal association with the object to be recognized, while filtering out the interference features related to the environment. The second attention network focuses on separating the surface features that have no causal association with the object to be recognized, such as texture details, light reflections, background interference, etc.
[0061] The first attention network and the second attention network may be two networks with similar structures but different parameters.
[0062] Exemplarily, two lightweight attention networks can be used to extract the causal features and non-causal features , respectively, to achieve feature decoupling. See formulas (1) and (2):
[0063] (1)
[0064] (2)
[0065] Wherein, represents the causal feature of the i-th image patch, represents the non-causal feature of the i-th image patch, represents the sigmoid function, represents the first attention network, represents the second attention network, represents the input image patch feature, Represents the initial features obtained by the representative image block through the feature extractor, and ⊙ represents element-wise multiplication.
[0066] Through causal feature decoupling, the causal features can focus on the important regions stably perceived by the first attention network, while the non-causal features capture deviation information such as the background.
[0067] According to the embodiments of the present application, by processing multiple image block features in parallel through the first attention network and the second attention network, accurate separation of the causal features and non-causal features of the image block features is achieved, interference of irrelevant features to the recognition process is avoided, and the recognition accuracy of the image is improved.
[0068] Figure 3 Shows a flowchart of an image recognition method according to another embodiment of the present application.
[0069] As Figure 3 shown, multiple image blocks of the image to be recognized 301 are respectively subjected to feature extraction to obtain multiple image block features 302. The multiple image block features 302 are respectively input into the first attention network 303 and the second attention network 304. The non-causal features 307 are output through the second attention network 304. The causal features 305 and the first weight information 306 are respectively output through the first attention network 303. Based on the first weight information 306, the target causal features 308 are obtained from the multiple causal features, and based on the target causal features 308, the recognition result 309 is obtained. Thus, accurate separation of the causal features and non-causal features of the image block features is achieved, interference of irrelevant features to the recognition process is avoided, and the recognition accuracy of the image is improved.
[0070] According to the embodiments of the present application, based on multiple target causal features, recognizing the image to be recognized to obtain a recognition result may include: performing weighted fusion on the multiple target causal features based on the respective second weight information of the multiple target causal features to obtain a causal fusion feature; the second weight information represents the importance degree of the target causal feature to the object to be recognized in the image; based on the causal fusion feature, recognizing the image to be recognized to obtain a recognition result.
[0071] The roles of different target causal features in recognizing the object to be recognized in the image are different. For example, when recognizing an animal image, the head feature and the leg feature of the animal may both be causal features, but the importance of the head feature for accurately recognizing the animal species may be higher, so the second weight information corresponding to the head feature will be larger than the second weight information corresponding to the leg feature.
[0072] Exemplarily, the second weight information can be determined by collecting a large amount of labeled image data and using machine learning algorithms (such as decision trees, neural networks, etc.) to learn the degree of association between each target causal feature and the object to be recognized. This application is not limited to this, and the corresponding second weight information can also be assigned to each target causal feature through the understanding and experience of domain experts.
[0073] After determining the second weight information of each of the multiple target causal features, a causal fusion feature can be obtained by using a linear weighted fusion method. In addition, a non-linear function can be used to process the target causal features and weights, and then fusion can be performed. The non-linear function can introduce more complex feature interaction relationships and improve the expression ability of the fusion feature.
[0074] According to the embodiments of the present application, by performing weighted fusion according to the second weight information to obtain a causal fusion feature, key information closely related to the object to be recognized can be captured more accurately, and then a more accurate and reliable recognition result can be obtained.
[0075] In addition, by screening image patches through the first weight information to determine the key area, and then identifying the importance degree of each target causal feature through the second weight information, the combination of the two forms a collaborative recognition from spatial rough selection to feature fine screening, which can effectively improve the recognition efficiency and accuracy.
[0076] According to the embodiments of the present application, based on multiple causal features, the image to be recognized is recognized, and the obtained recognition result may include: processing the multiple causal features by using the aggregation classification network of the image recognition model to obtain the recognition result. The aggregation classification network includes an aggregation sub-network and a classification sub-network. The aggregation sub-network is used to determine the second weight information of the multiple target causal features and perform weighted fusion on the multiple target causal features, and the classification sub-network is used to process the causal fusion feature to obtain the recognition result.
[0077] Exemplarily, the aggregation sub-network can receive multiple target causal features as inputs, and through the internal attention mechanism, determine the association strength between each target causal feature and the object to be recognized, and assign the second weight information to each target causal feature. The normalization assignment of the weights can also be realized through a fully connected layer combined with the Softmax (soft maximum) function to ensure that the sum of the weights is 1, so as to reflect the importance ratio of different features.
[0078] Using the attention mechanism to perform weighted aggregation on the causal features of all instances, the weight of the target causal feature Referring to formula (3) is:
[0079] (3)
[0080] Wherein, Parameters for learning Represents the causal feature of the i-th image patch Represents the causal feature of the j-th image patch
[0081] The causal fusion feature obtained by weighted fusion of multiple target causal features See formula (4):
[0082] (4)
[0083] The classification sub-network can be a Softmax classifier or other suitable classifier, which maps the causal fusion feature to the class probability space and outputs the recognition result. See formula (5):
[0084] (5)
[0085] Wherein Can represent a probability distribution vector, where each probability corresponds to the probability value of a class, and the class with the highest probability can be selected as the recognition result of the image to be recognized
[0086] Figure 4 Shows the network architecture diagram of the image recognition model according to the embodiment of the present application
[0087] As Figure 4 Shown, the image recognition model includes ViT410, a feature decoupling network 420, and an aggregation classification network 430. The feature decoupling network 420 includes a first attention network 303 and a second attention network 304. The aggregation classification network 430 includes an aggregation sub-network 431 and a classification sub-network 432
[0088] Input the image 301 to be recognized into ViT410, and multiple image patch features can be obtained. Input the multiple image patch features into the feature decoupling network 420, and the first attention network 303 can be used to extract multiple causal features and corresponding first weight information, and the second attention network 304 can be used to extract non-causal features. Based on the first weight information, multiple target causal features can be determined from the multiple causal features. Input the multiple target causal features into the aggregation classification network 430, use the aggregation sub-network 431 to determine the second weight information of the multiple target causal features, and perform weighted aggregation on the multiple target causal features to obtain a causal fusion feature. Input the causal fusion feature into the classification sub-network 432 to obtain the recognition result 309
[0089] According to an embodiment of the present application, by the collaborative work of the aggregation sub-network and the classification sub-network, the importance of causal features is first evaluated and fused, and then accurate classification is performed based on the fused features, thereby effectively improving the accuracy and reliability of image recognition.
[0090] Figure 5 FIG. 4 shows a flowchart of a method for training an image recognition model according to an embodiment of the present application.
[0091] This method includes operations S510 to S550.
[0092] In operation S510, a sample training set is obtained, and the sample training set includes a plurality of sample images.
[0093] In operation S520, feature extraction processing is respectively performed on a plurality of sample image patches of the sample images to obtain a plurality of sample image patch features of the sample images.
[0094] In operation S530, the plurality of sample image patch features are input into the feature decoupling network of the initial image recognition model to obtain respective sample causal features and sample non-causal features corresponding to the plurality of sample image patch features. The sample causal features represent the features of the sample objects included in the sample image patch features, and the non-causal features represent the features irrelevant to the sample objects in the sample image patch features;
[0095] In operation S540, the plurality of sample causal features are input into the aggregation classification network of the initial image recognition model to obtain a sample recognition result;
[0096] In operation S550, the initial image recognition model is trained based on the comprehensive loss to obtain a trained image recognition model; the comprehensive loss includes a first loss and a second loss. The first loss represents the deviation between the sample causal features and the sample non-causal features, and the second loss represents the deviation between the sample recognition result and the sample image label.
[0097] Exemplarily, the first loss can be determined according to the cosine similarity between the sample causal features and the sample non-causal features. By training, the difference between the sample causal features and the sample non-causal features is maximized, thereby realizing the decoupling of the causal features and the non-causal features.
[0098] The formula of the first loss function is shown in formula (6):
[0099] (6)
[0100] Wherein, represents the first loss, n represents the number of image patches, represents the causal feature of the i-th image patch, represents the non-causal feature of the i-th image patch.
[0101] Exemplarily, the second loss can be determined according to the cross-entropy loss between the sample recognition result and the sample image label. The loss function of the second loss can be seen in formula (7):
[0102]
[0103] represents the second loss; C is the total number of categories, that is, the number of categories that the model needs to predict, and c represents the c-th category. represents the sample image label. If the c-th category is the true category, then = 1; otherwise, = 0, represents the prediction probability of the c-th category.
[0104] According to the embodiments of the present application, by comprehensively training the initial image recognition model with the first loss and the second loss, the model can be optimized from two key aspects: feature extraction decoupling and classification recognition, so that the trained image recognition model can more effectively extract causal features and make accurate classifications.
[0105] According to the embodiments of the present application, after obtaining the sample causal features and sample non-causal features corresponding to each of the multiple sample image patch features, it may further include: updating the non-causal features of at least one sample image patch to obtain updated non-causal features. Based on the multiple sample causal features and the updated non-causal features, the sample image is recognized to obtain an updated sample recognition result. Based on the updated sample recognition result and the sample recognition result, a third loss is determined. Based on the first loss, the second loss, and the third loss, a comprehensive loss is obtained. Based on the comprehensive loss, the initial image recognition model is trained to obtain a trained image recognition model.
[0106] Exemplarily, a causal intervention network can be introduced into the initial recognition model. After parsing the image patch features to obtain causal features and non-causal features, the non-causal features are input into the causal intervention network, and the non-causal features are updated to obtain updated non-causal features. The causal features and the updated non-causal features are input into the aggregation classification network together to output an updated sample recognition result, and based on the difference of the image recognition model before and after updating the non-causal features, it is verified whether the output result of the aggregation classification network only depends on the causal features.
[0107] The non-causal features can be updated in various ways. For example, replacement operations, masking operations, and adding random noise operations, etc. can be performed on the non-causal features of at least one sample image patch.
[0108] Assume that the current feature of a sample image patch , the feature representation after feature decoupling See Equation (8):
[0109] (8)
[0110] where represents the causal feature of the i-th image patch, represents the non-causal feature of the i-th image patch.
[0111] For perform the update operation on the non-causal feature, and use to represent the updated image patch feature. Among them, the replacement operation is to replace with another randomly selected non-causal feature See Equation (9)
[0112] (9)
[0113] The masking operation is to replace with a zero vector, see Equation (10);
[0114] (10)
[0115] The perturbation operation is to add a random noise vector to , see Equation (11):
[0116] (11)
[0117] According to the embodiments of the present application, since the causal intervention in this example is directly performed on the non-causal features obtained after decoupling, compared with the method of implementing causal intervention by modifying the image, there is no need to re-input the modified image patch into the feature extraction module, but directly operate in the feature space, which can effectively optimize the calculation efficiency and ensure that no additional variables are introduced.
[0118] This module uses the intervention loss to measure the degree of change in the model prediction after the intervention operation.
[0119] (12)
[0120] represents the third loss; represents the expectation, and the calculation is the average loss under all possible and combinations; represents the update operation performed on the non-causal feature; represents that given the input feature the predicted probability distribution of the model for the output Y; represents the predicted probability distribution of the model for output Y in the case of updated features ; KL represents the Kullback-Leibler (relative entropy) divergence, which is used to measure the divergence from the distribution .
[0121] Among them, the calculation formula of the KL divergence is shown in formula (13):
[0122]
[0123] where is the original instance set, represents the KL divergence, C represents the total number of categories, c represents the c-th category, p is the prediction result of the original image instance, and q is the prediction result of the intervened image instance
[0124] Exemplarily, in the overall training process, the model can be jointly optimized by combining the first loss, the second loss, and the third loss: see formula (14)
[0125] (14)
[0126] where represents the comprehensive loss, is the first loss (i.e., the decoupling loss), and its function is to constrain the causal features and the confounding features from interfering with each other As the second loss (i.e., the classification loss), its function is to supervise the classification result, so that the initial image recognition model learns to predict the correct category from the causal region of the entire image; is the third loss (i.e., the intervention loss), and its function is to measure the impact of the instance after intervention on the prediction, so as to guide the model to identify which regions have a true causal relationship with the label and reduce the dependence on the background or false relevant features; among them , is used to control the weights of different losses
[0127] According to the embodiments of the present application, by updating the non-causal features, a causal intervention mechanism can be introduced, which can effectively block the context dependence or false correlation path formed in the training stage, enable the image recognition model to focus on the image regions that truly have causal contributions, and thus improve the performance of the image recognition model in maintaining stability when facing out-of-distribution test data
[0128] According to an embodiment of the present application, the sample training set includes a first sample training set and a second sample training set, and the categories of the sample objects in the first sample training set and the second sample training set are the same; the method further includes: using the first sample training set to train an initial image recognition model, and obtaining an intermediate image recognition model when the comprehensive loss meets a first preset condition; using the second sample training set to train the intermediate image recognition model, and obtaining an image recognition model when the comprehensive loss meets a second preset condition.
[0129] Exemplarily, the first sample training set is obtained from multiple sample object data, and the second sample training set is obtained by combining multiple sample objects and multiple sample environment objects. The categories of the sample objects in the first sample training set and the second sample training set are the same, but the environments of the sample objects are inconsistent. The sample object environment includes the shooting angle, lighting conditions, background environment, etc. of the image. The second sample training set can simulate out-of-distribution data to improve the fast adaptation ability of the image recognition model when facing new environments and new data.
[0130] The first preset condition means that the comprehensive loss meets the convergence condition during the training with the first sample training set, and the second preset condition means that the comprehensive loss meets the convergence condition during the training with the second sample training set.
[0131] First, use the first sample training set to train the initial image recognition model. Input the images of the first training set into the initial image recognition model to obtain the comprehensive loss, and update the parameters of the initial image recognition model by the gradient descent method. See formula (15):
[0132] (15)
[0133] Where represents the first sample training set, represents the initial parameters of the model, is the step size, represents taking the gradient of the initial parameters with respect to, represents the loss obtained after the current first sample training set passes through the initial image recognition model. is the updated parameter vector.
[0134] Update using one or more gradient descents on the first sample training set calculated with the current loss to obtain an intermediate image recognition model. Next, use the second sample training set to train the intermediate image recognition model, obtain the comprehensive loss, and use the comprehensive loss to optimize the intermediate image recognition model. See formula (16):
[0135]
[0136] Where is the step size, represents the second sample training set, represents the loss obtained after the current second sample training set passes through the intermediate image recognition model. By iterating and continuously minimizing the comprehensive loss, an image recognition model is obtained. t refers to the number of tasks, and i represents the i-th task
[0137] Figure 6 Shows a flowchart of training an image recognition model according to another embodiment of the present application.
[0138] As Figure 6 shown, the sample training set 610 is divided into a first sample training set 620 and a second sample training set 630 with the same sample object categories but different sample object environments. The initial image recognition model 640 is trained using the first sample training set. Based on the comprehensive loss, the parameters of the initial image recognition model 640 are updated using the gradient descent method until the comprehensive loss meets the first preset condition, obtaining an intermediate image recognition model 650. The intermediate image recognition model 650 is trained using the second sample training set 630. Based on the comprehensive loss, the parameters of the intermediate image recognition model are updated using the gradient descent method until the comprehensive loss meets the second preset condition, obtaining an image recognition model 660.
[0139] According to the embodiments of the present application, since the image recognition model is trained using the first sample training set and the second sample training set in different environments, it performs more stably, can quickly adapt to the data characteristics in the new environment, and reduces the influence of environmental characteristics on image classification. It can accurately extract causal features and improve the generalization ability of the model when facing data with new distributions. Furthermore, the technical effect of improving the accuracy of image recognition is achieved.
[0140] Figure 7 Shows a module diagram of an image recognition device according to an embodiment of the present application.
[0141] The image recognition device 700 includes an extraction module 710, a decoupling module 720, and a recognition module 730.
[0142] The extraction module 710 is configured to perform feature extraction processing on multiple image blocks of the image to be recognized respectively, obtaining multiple image block features of the image to be recognized. The multiple image blocks represent different regions of the image to be recognized;
[0143] The decoupling module 720 is configured to perform feature decoupling on the multiple image block features respectively, obtaining the causal features and non-causal features of each of the multiple image block features. The causal features represent the features of the object to be recognized included in the image block features, and the non-causal features represent the features unrelated to the object to be recognized in the image block features;
[0144] The recognition module 730 is configured to recognize the image to be recognized based on multiple causal features and obtain a recognition result.
[0145] According to an embodiment of the present application, the recognition module 730 includes a screening sub-module and a recognition sub-module.
[0146] The screening sub-module is configured to obtain multiple target causal features from the causal features corresponding to each of the multiple image patch features based on first weight information; the first weight information represents the proportion degree of the causal feature in the corresponding image patch feature.
[0147] The first recognition sub-module is configured to recognize the image to be recognized based on the multiple target causal features and obtain a recognition result.
[0148] According to an embodiment of the present application, the recognition sub-module includes a fusion unit and a recognition unit.
[0149] The fusion unit is configured to perform weighted fusion on the multiple target causal features based on second weight information of each of the multiple target causal features to obtain a causal fusion feature; the second weight information represents the importance degree of the target causal feature to the object to be recognized.
[0150] The recognition unit is configured to recognize the image to be recognized based on the causal fusion feature and obtain a recognition result.
[0151] The encoding sub-module is configured to perform encoding processing on multiple image patches of the image to be recognized respectively according to an embodiment of the present application to obtain encoding vectors of each of the multiple image patches.
[0152] The mapping sub-module is configured to map the encoding vectors of each of the multiple image patches into vectors of a target dimension to obtain image patch features.
[0153] According to an embodiment of the present application, the decoupling module 720 includes a feature decoupling sub-module.
[0154] The feature decoupling sub-module is configured to process multiple image patch features by using a feature decoupling network of an image recognition model to obtain causal features and non-causal features of each of the multiple image patch features. Among them, the feature decoupling network includes a first attention network and a second attention network. The first attention network is used to extract causal features of each of the multiple image patches and first weight information corresponding to each causal feature; the second attention network is used to extract non-causal features of each of the multiple image patches.
[0155] According to an embodiment of the present application, the recognition module 730 includes a second recognition sub-module.
[0156] A second recognition sub-module, configured to process multiple causal features by using an aggregation classification network of an image recognition model to obtain a recognition result; wherein, the aggregation classification network includes an aggregation sub-network and a classification sub-network, the aggregation sub-network is configured to determine second weight information of multiple target causal features and perform weighted fusion on the target causal features, and the classification sub-network is configured to process the causal fusion features to obtain a recognition result.
[0157] Figure 8 FIG. shows a module diagram of a training device for an image recognition model according to an embodiment of the present application.
[0158] The training device 800 for the image recognition model includes an acquisition module 810, a sample extraction module 820, a sample decoupling module 830, a sample recognition module 840, and a training module 850.
[0159] The acquisition module 810 is configured to acquire a sample training set, and the sample training set includes multiple sample images.
[0160] The sample extraction module 820 is configured to perform feature extraction processing on multiple sample image patches of the sample image respectively to obtain multiple sample image patch features of the sample image, and the multiple sample image patches represent different regions of the sample image;
[0161] The sample decoupling module 830 is configured to input the multiple sample image patch features into a feature decoupling network of the initial image recognition model to obtain respective sample causal features and sample non-causal features corresponding to the multiple sample image patch features, the sample causal features represent features of a sample object included in the sample image patch features, and the non-causal features represent features irrelevant to the sample object in the sample image patch features;
[0162] The sample recognition module 840 is configured to input the multiple sample causal features into an aggregation classification network of the initial image recognition model to obtain a sample recognition result;
[0163] The training module 850 trains the initial image recognition model based on a comprehensive loss to obtain a trained image recognition model; the comprehensive loss includes a first loss and a second loss, the first loss represents the deviation between the sample causal features and the sample non-causal features, and the second loss represents the deviation between the sample recognition result and the sample image label.
[0164] According to an embodiment of the present application, the training module 850 includes an update sub-module, a third recognition sub-module, a loss determination sub-module, and a training sub-module.
[0165] The update sub-module is configured to update the non-causal features of at least one sample image patch to obtain updated non-causal features.
[0166] A third recognition sub-module, configured to recognize a sample image based on multiple sample causal features and updated non-causal features, and obtain an updated sample recognition result;
[0167] A loss determination sub-module, configured to determine a third loss based on the updated sample recognition result and the sample recognition result;
[0168] A training sub-module, configured to train an initial image recognition model based on the first loss, the second loss, and the third loss, and obtain a trained image recognition model.
[0169] According to an embodiment of the present application, an update sub-module is configured to perform at least one of a replacement operation, a masking operation, and an adding random noise operation on non-causal features of at least one sample image block.
[0170] According to an implementation of the present application, the sample training set includes a first sample training set and a second sample training set, and the categories of the sample objects in the first sample training set and the second sample training set are the same.
[0171] The training apparatus 800 of the image recognition model further includes a first training module and a second training module.
[0172] The first training module uses the first sample training set to train the initial image recognition model, and obtains an intermediate image recognition model when the comprehensive loss meets a first preset condition.
[0173] The second training module uses the second sample training set to train the intermediate image recognition model, and obtains an image recognition model when the comprehensive loss meets a second preset condition.
[0174] According to embodiments of the present application, any plurality of modules among the extraction module 710, the decoupling module 720, the recognition module 730, the acquisition module 810, the sample extraction module 820, the sample decoupling module 830, the sample recognition module 840, and the training module 850 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to embodiments of the present application, at least one of the extraction module 710, the decoupling module 720, the recognition module 730, the acquisition module 810, the sample extraction module 820, the sample decoupling module 830, the sample recognition module 840, and the training module 850 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the extraction module 710, the decoupling module 720, the recognition module 730, the acquisition module 810, the sample extraction module 820, the sample decoupling module 830, the sample recognition module 840, and the training module 850 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.
[0175] Figure 9 FIG. shows a block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present application. Figure 9 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0176] The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described herein and / or claimed.
[0177] As Figure 9As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 902 or computer programs loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0178] A plurality of components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0179] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as an image recognition method, a model training method. For example, in some embodiments, the image recognition method, the model training method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the image recognition method, the model training method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the image recognition method, the model training method by any other appropriate means (e.g., by means of firmware).
[0180] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0181] This application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps of the above method.
[0182] The program code for implementing the method of this application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable test device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program codes can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0183] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0184] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, speech input, or tactile input).
[0185] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0186] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.
[0187] Those skilled in the art will understand that the features recited in the various embodiments of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present application. In particular, without departing from the spirit and teachings of the present application, the features recited in the various embodiments of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application.
[0188] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although the embodiments have been described separately above, this does not mean that the measures in the respective embodiments cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present application.
Claims
1. An image recognition method, characterized in that, Including: Performing feature extraction processing on multiple image patches of the image to be recognized respectively to obtain multiple image patch features of the image to be recognized; Performing feature decoupling on the multiple image patch features respectively to obtain the causal features and non-causal features of each of the multiple image patch features, where the causal features represent the features of the object to be recognized included in the image patch features, and the non-causal features represent the features unrelated to the object to be recognized in the image patch features; Based on the multiple causal features, recognizing the image to be recognized to obtain a recognition result.
2. The method according to claim 1, characterized in that, Based on the multiple causal features, recognizing the image to be recognized to obtain a recognition result, including: Based on first weight information, obtaining multiple target causal features from the causal features corresponding to each of the multiple image patch features; the first weight information represents the proportion degree of the causal features in the corresponding image patch features; Based on the multiple target causal features, recognizing the image to be recognized to obtain a recognition result.
3. The method according to claim 2, wherein Based on the multiple target causal features, recognizing the image to be recognized to obtain a recognition result, including: Based on the second weight information of each of the multiple target causal features, performing weighted fusion on the multiple target causal features to obtain a causal fusion feature; the second weight information represents the importance degree of the target causal features to the object to be recognized; Based on the causal fusion feature, recognizing the image to be recognized to obtain a recognition result.
4. The method according to claim 1, wherein Performing feature extraction processing on multiple image patches of the image to be recognized respectively to obtain multiple image patch features of the image to be recognized, including: Performing encoding processing on the multiple image patches of the image to be recognized respectively to obtain the encoding vectors of each of the multiple image patches; Mapping the encoding vectors of each of the multiple image patches into vectors of a target dimension to obtain the image patch features.
5. The method according to any one of claims 1 to 4, characterized in that, Performing feature decoupling on the multiple image patch features respectively to obtain the causal features and non-causal features of each of the multiple image patch features, including: Using the feature decoupling network of the image recognition model to process the multiple image patch features to obtain the causal features and non-causal features of each of the multiple image patch features; Wherein, the feature decoupling network includes a first attention network and a second attention network, the first attention network is used to extract the causal features of each of the multiple image patches and the first weight information corresponding to each causal feature; the second attention network is used to extract the non-causal features of each of the multiple image patches.
6. The method according to any one of claims 1 to 4, characterized in that Based on the multiple causal features, recognizing the image to be recognized to obtain a recognition result, including: Using the aggregation classification network of the image recognition model to process the multiple causal features to obtain a recognition result; Wherein, the aggregation classification network includes an aggregation sub-network and a classification sub-network, the aggregation sub-network is used to determine the second weight information of the multiple target causal features and perform weighted fusion on the multiple target causal features, and the classification sub-network is used to process the causal fusion feature to obtain a recognition result.
7. A training method for an image recognition model, characterized in that, Including: Obtaining a sample training set, the sample training set including multiple sample images; Feature extraction processing is respectively performed on multiple sample image patches of a sample image to obtain multiple sample image patch features of the sample image; The multiple sample image patch features are input into a feature decoupling network of an initial image recognition model to obtain sample causal features and sample non-causal features respectively corresponding to the multiple sample image patch features, where the sample causal features represent the features of a sample object included in the sample image patch features, and the non-causal features represent the features irrelevant to the sample object in the sample image patch features; The multiple sample causal features are input into an aggregation classification network of the initial image recognition model to obtain a sample recognition result; Based on a comprehensive loss, the initial image recognition model is trained to obtain a trained image recognition model, where the comprehensive loss includes a first loss and a second loss, the first loss represents the deviation between the sample causal features and the sample non-causal features, and the second loss represents the deviation between the sample recognition result and a sample image label.
8. The method according to claim 7, characterized in that The method further includes: Updating the sample non-causal features of at least one of the sample image patches to obtain updated sample non-causal features; Based on the multiple sample causal features and the updated sample non-causal features, the sample image is recognized to obtain an updated sample recognition result; Based on the updated sample recognition result and the sample recognition result, a third loss is determined; Based on the first loss, the second loss, and the third loss, the comprehensive loss is determined.
9. The method according to claim 8, wherein Updating the sample non-causal features of at least one of the sample image patches includes: Performing at least one of a replacement operation, a masking operation, and an operation of adding random noise on the sample non-causal features of at least one of the sample image patches.
10. The method according to any one of claims 7 to 9, characterized in that The sample training set includes a first sample training set and a second sample training set, and the categories of the sample objects in the first sample training set and the second sample training set are the same; The method further includes: Using the first sample training set to train the initial image recognition model, and obtaining an intermediate image recognition model when the comprehensive loss satisfies a first preset condition; Using the second sample training set to train the intermediate image recognition model, and obtaining the image recognition model when the comprehensive loss satisfies a second preset condition.
11. An image recognition device, characterized in that, It includes: An extraction module that respectively performs feature extraction processing on multiple image patches of an image to be recognized to obtain multiple image patch features of the image to be recognized; A decoupling module configured to respectively perform feature decoupling on the multiple image patch features to obtain causal features and non-causal features respectively corresponding to the multiple image patch features, where the causal features represent the features of an object to be recognized included in the image patch features, and the non-causal features represent the features irrelevant to the object to be recognized in the image patch features; A recognition module configured to recognize the image to be recognized based on the multiple causal features to obtain a recognition result.
12. A training device for an image recognition model, characterized in that It includes: An acquisition module configured to acquire a sample training set, where the sample training set includes multiple sample images; A sample extraction module, configured to perform feature extraction processing on multiple sample image patches of a sample image respectively, so as to obtain multiple sample image patch features of the sample image; A sample decoupling module, configured to input the multiple sample image patch features into a feature decoupling network of an initial image recognition model, so as to obtain a sample causal feature and a sample non-causal feature corresponding to each of the multiple sample image patch features, where the sample causal feature represents the feature of a sample object included in the sample image patch feature, and the non-causal feature represents the feature irrelevant to the sample object in the sample image patch feature; A sample recognition module, configured to input the multiple sample causal features into an aggregation classification network of the initial image recognition model, so as to obtain a sample recognition result; A training module, configured to train the initial image recognition model based on a comprehensive loss, so as to obtain a trained image recognition model, where the comprehensive loss includes a first loss and a second loss, the first loss represents the deviation between the sample causal feature and the sample non-causal feature, and the second loss represents the deviation between the sample recognition result and a sample image label.
13. An electronic device, comprising: One or more processors; A memory, configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any one of claims 1 to 10.
14. A computer-readable storage medium having executable instructions stored thereon, having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
15. A computer-readable storage medium having executable instructions stored thereon, characterized in that, When the instruction is executed by a processor, the processor is caused to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Vegetable and fruit identification method and system based on image feature sequence modeling
CN114677542A
Recognition method of duplicated image and training method and device of duplicated image recognition model
CN116128805A
Weak supervision medical image segmentation method and system based on causal intervention
CN119693642A
Feature decoupling refining and confusion minimization constraint combined remote sensing directed target identification method
CN119723370A
Remote sensing image domain generalization semantic segmentation method based on multi-scale instance decoupling
CN119992289A