Image recognition model training method and image recognition method
By obtaining the smoothness features and characteristics of socially transmitted images during the image recognition model training phase and combining them with the classifier mapping relationship, the accuracy problem of image source identification after social network transmission is solved, achieving higher robustness and recognition accuracy.
Patent Information
- Application Number
- CN202211564494.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-12-07
AI Technical Summary
Existing technologies have difficulty accurately identifying the source of images after they are transmitted on social networks, especially after the images have been scaled, compressed, and other processes, traditional algorithms find it difficult to maintain performance stability.
During the model training phase, the sample image and the transmission image of its social transmission relationship are obtained, the smoothness features and image features are extracted through the smoothness perception network, the mapping relationship between the image label and the collected information is determined by combining the classifier, and the parameters are adjusted until the target image recognition model that meets the training stop conditions is obtained.
The model's robustness and recognition accuracy in social network transmission scenarios have been improved, and it can more accurately extract image traces to meet downstream business needs.
Smart Images

Figure CN116188796B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of machine learning technology, and in particular to an image recognition model training method and an image recognition method. Background Art
[0002] With the development of internet technology, image creation and editing have become significantly easier thanks to various hardware and software tools. This has led to significant challenges in verifying the authenticity and integrity of images. Image forensics can help assess the authenticity and integrity of a given image, and one of the most important areas is image source identification (ISI). The primary goal of ISI is to identify the fingerprint of the device that captures a digital image. This includes identifying both the specific device model and the individual device. Specifically, it involves identifying the camera model and individual camera used to capture the image. In practical applications, the results of image source identification serve as a baseline for image forensics, helping to screen out suspicious images and assist in identification. This is of great significance. However, in real-world applications, images to be traced often originate from social media or fall outside the predefined set of target devices. This can lead to performance bottlenecks in existing classification technologies, making accurate classification and tracing difficult. Therefore, an effective solution is urgently needed to address this issue. Summary of the Invention
[0003] In view of this, embodiments of this specification provide an image recognition model training method. One or more embodiments of this specification also involve an image recognition model training device, an image recognition model training system, an image recognition method, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.
[0004] According to a first aspect of an embodiment of this specification, a first image recognition model training method is provided, comprising:
[0005] Acquire a sample image and a transmission image having a social transmission relationship with the sample image, and input the transmission image into an initial image recognition model;
[0006] Extracting smoothness features and image features of the transmission image through the initial image recognition model, determining image acquisition information corresponding to the transmission image based on the smoothness features and the image features, and outputting the information;
[0007] Determining a mapping relationship between the image label of the sample image and the image acquisition information by a classifier;
[0008] The parameters of the initial image recognition model are adjusted according to the mapping relationship until a target image recognition model that meets the training stop condition is obtained.
[0009] According to a second aspect of the embodiments of this specification, a first image recognition model training device is provided, comprising:
[0010] an acquisition module configured to acquire a sample image and a transmission image having a social transmission relationship with the sample image, and input the transmission image into an initial image recognition model;
[0011] an extraction module configured to extract smoothness features and image features of the transmission image through the initial image recognition model, determine image acquisition information corresponding to the transmission image based on the smoothness features and the image features, and output the information;
[0012] a determination module configured to determine a mapping relationship between the image label of the sample image and the image acquisition information through a classifier;
[0013] The training module is configured to adjust the parameters of the initial image recognition model according to the mapping relationship until a target image recognition model that meets the training stop condition is obtained.
[0014] According to a third aspect of the embodiments of this specification, there is provided an image recognition method, including:
[0015] Obtaining an image to be identified associated with a target business;
[0016] Inputting the image to be recognized into the target image recognition model in the above method;
[0017] Extracting target smoothness features and target image features of the image to be identified through the target image recognition model, determining image acquisition information corresponding to the image to be identified based on the target smoothness features and the target image features, and outputting the information;
[0018] Determine service detection information of the target service based on the image acquisition information.
[0019] According to a fourth aspect of the embodiments of this specification, there is provided an image recognition device, including:
[0020] An image acquisition module is configured to acquire an image to be identified associated with a target service;
[0021] An input model module, configured to input the image to be recognized into the target image recognition model in the above method;
[0022] a prediction information module configured to extract target smoothness features and target image features of the image to be recognized through the target image recognition model, determine image acquisition information corresponding to the image to be recognized based on the target smoothness features and the target image features, and output the image acquisition information;
[0023] The information determination module is configured to determine the service detection information of the target service according to the image acquisition information.
[0024] According to a fifth aspect of the embodiments of this specification, a second image recognition model training method is provided, which is applied to a cloud-side device, including:
[0025] Obtain sample images submitted by the client device;
[0026] generating a transmission image having a social transmission relationship with the sample image, and inputting the transmission image into an initial image recognition model;
[0027] Extracting smoothness features and image features of the transmission image through the initial image recognition model, determining image acquisition information corresponding to the transmission image based on the smoothness features and the image features, and outputting the information;
[0028] Determining a mapping relationship between the image label of the sample image and the image acquisition information through a classifier, and adjusting parameters of the initial image recognition model according to the mapping relationship to obtain a target image recognition model;
[0029] The target model parameters corresponding to the target image recognition model are sent to the terminal side device.
[0030] According to a sixth aspect of the embodiments of this specification, a second image recognition model training device is provided, which is applied to a cloud-side device, including:
[0031] an acquisition module, configured to acquire a sample image submitted by a terminal device;
[0032] a generating module configured to generate a transmission image having a social transmission relationship with the sample image, and input the transmission image into an initial image recognition model;
[0033] a processing module configured to extract smoothness features and image features of the transmission image through the initial image recognition model, determine image acquisition information corresponding to the transmission image based on the smoothness features and the image features, and output the information;
[0034] a determination module configured to determine a mapping relationship between the image label of the sample image and the image acquisition information through a classifier, and adjust parameters of the initial image recognition model according to the mapping relationship to obtain a target image recognition model;
[0035] The sending module is configured to send the target model parameters corresponding to the target image recognition model to the terminal side device.
[0036] According to a seventh aspect of the embodiments of this specification, a third image recognition model training method is provided, including:
[0037] Acquire a sample image and a transmission image having a social transmission relationship with the sample image, and input the transmission image into an initial image recognition model;
[0038] Extracting smoothness features and image features of the transmission image through the initial image recognition model, determining image acquisition information corresponding to the transmission image based on the smoothness features and the image features, and outputting the information;
[0039] The initial image recognition model is adjusted according to the image label of the sample image and the image acquisition information until a target image recognition model that meets the training stop condition is obtained.
[0040] According to an eighth aspect of the embodiments of this specification, a third image recognition model training device is provided, comprising:
[0041] an acquisition module configured to acquire a sample image and a transmission image having a social transmission relationship with the sample image, and input the transmission image into an initial image recognition model;
[0042] a processing module configured to extract smoothness features and image features of the transmission image through the initial image recognition model, determine image acquisition information corresponding to the transmission image based on the smoothness features and the image features, and output the information;
[0043] The training module is configured to adjust the parameters of the initial image recognition model according to the image label of the sample image and the image acquisition information until a target image recognition model that meets the training stop condition is obtained.
[0044] According to a ninth aspect of the embodiments of this specification, there is provided an image recognition model training system, comprising:
[0045] Model training end and parameter collection end;
[0046] The parameter acquisition end is used to acquire model parameters, and the model training end is used to execute image recognition model training instructions. When the image recognition model training instructions are executed by the model training end, the steps of the above-mentioned image recognition model training method are implemented.
[0047] According to a tenth aspect of an embodiment of this specification, a computing device is provided, including:
[0048] memory and processor;
[0049] The memory is used to store computer-executable instructions, and the processor is used to implement any step of the above-mentioned image recognition model training method and image recognition method when executing the computer-executable instructions.
[0050] According to the eleventh aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned image recognition model training method and image recognition method.
[0051] According to the twelfth aspect of the embodiments of this specification, a computer program is provided, wherein, when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned image recognition model training method and image recognition method.
[0052] The image recognition model training method provided in this embodiment can improve the robustness of the target recognition model and ensure that the model can accurately identify the image source in the social transmission scenario. In the model training stage, a sample image and a transmission image that has a social transmission relationship with the sample image can be obtained. At this time, the transmission image is first input into the initial image recognition model, and the smoothness feature and image feature of the transmission image are extracted by the initial image recognition model to determine the image acquisition information corresponding to the transmission image based on the smoothness feature and image feature and output it. At this time, the mapping relationship between the image label and the image acquisition information is determined in combination with the classifier, and the parameters of the initial image recognition model are adjusted based on this until a target image recognition model that meets the training stop condition is obtained. In the training process, by extracting the smoothness feature, the robustness of the model against different social network scenarios is greatly improved, so that the model can more accurately extract the image traces of the transmitted image to meet the needs of downstream business. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a schematic diagram of an image recognition model training method provided by one embodiment of this specification;
[0054] Figure 2 This is a flowchart of a first image recognition model training method provided by one embodiment of this specification;
[0055] Figure 3 This is a flowchart of a second image recognition model training method provided by one embodiment of this specification;
[0056] Figure 4 This is a flowchart of a third image recognition model training method provided by an embodiment of this specification;
[0057] Figure 5 This is a structural diagram of a first image recognition model training device provided by an embodiment of this specification;
[0058] Figure 6 is a flow chart of an image recognition method provided by one embodiment of this specification;
[0059] Figure 7 This is a schematic diagram of the structure of an image recognition device provided by one embodiment of this specification;
[0060] Figure 8 This is a flowchart of a fourth image recognition model training method provided by an embodiment of this specification;
[0061] Figure 9 This is a structural diagram of a second image recognition model training device provided by an embodiment of this specification;
[0062] Figure 10 is a flowchart of a fifth image recognition model training method provided by an embodiment of this specification;
[0063] Figure 11 This is a schematic structural diagram of a third image recognition model training device provided by an embodiment of this specification;
[0064] Figure 12 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0065] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0066] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0067] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0068] This specification provides an image recognition model training method. One or more embodiments of this specification also involve an image recognition model training apparatus, an image recognition model training system, an image recognition method, a computing device, a computer-readable storage medium, and a computer program, each of which is described in detail in the following embodiments.
[0069] In practice, image source identification technology is mostly used in business detection scenarios, such as image leak tracing, image theft, document collection, and forensic identification. These scenarios all require identifying the image's source. Existing technologies for image source identification mostly employ traditional algorithms and deep learning methods. Several traditional methods capture camera artifacts by modeling operations during image acquisition. For example, PRNU models the noise patterns introduced by sensor defects.
[0070] While the above algorithms can be used for image source identification and achieve good results under ideal circumstances, in practice, a large number of images to be detected are transmitted over lossy channels, such as those subjected to scaling and compression during transmission via social networks. This makes it difficult for these algorithms to maintain stable performance. Therefore, there is an urgent need for a highly robust algorithm that can better combat the effects of social networks and be used in practical applications.
[0071] See also Figure 1As shown in the schematic diagram, the image recognition model training method provided in this embodiment can improve the robustness of the target recognition model and ensure that the model can accurately identify the image source in the social transmission scenario. In the model training stage, a sample image and a transmission image with a social transmission relationship with the sample image can be obtained. At this time, the transmission image is first input into the initial image recognition model, and the smoothness feature and image feature of the transmission image are extracted by the initial image recognition model to determine the image acquisition information corresponding to the transmission image based on the smoothness feature and image feature and output it. At this time, the mapping relationship between the image label and the image acquisition information is determined in combination with the classifier, and the parameters of the initial image recognition model are adjusted based on this until a target image recognition model that meets the training stop condition is obtained. In the training process, by extracting the smoothness feature, the robustness of the model against different social network scenarios is greatly improved, so that the model can more accurately extract the image traces of the transmitted image to meet the needs of downstream business.
[0072] Figure 2 A flowchart of a first image recognition model training method provided according to an embodiment of this specification is shown, which specifically includes the following steps.
[0073] Step S202 : obtaining a sample image and a transmission image having a social transmission relationship with the sample image, and inputting the transmission image into an initial image recognition model.
[0074] The image recognition model training method provided in this embodiment can be applied to any image source identification scenario, such as image leakage tracing scenario, image theft scenario (unauthorized use without the consent of the image owner), document collection scenario, forensic identification scenario, etc.; for the convenience of description, this embodiment takes the image theft scenario as an example to illustrate the image recognition model training method. For related descriptions of other scenarios, please refer to the same or corresponding description content in this embodiment, and this embodiment will not be elaborated in detail here.
[0075] Specifically, the sample image specifically refers to the image required for training the image recognition model, and the sample image has not yet been transmitted through the social network, that is, the sample image is the original image captured by the image acquisition device. Accordingly, the transmitted image specifically refers to the image obtained after being transmitted through the social network. The sample image may be compressed or processed after being transmitted through the social network, so that the model can subsequently learn the characteristics of social network transmission, thereby training a more robust image recognition model. Among them, the social transmission relationship between the sample image and the transmitted image specifically refers to the relationship established after the sample image is transmitted through the set social communication software. Accordingly, in the embodiment, the initial image recognition model specifically refers to an image recognition model that can pay attention to the smoothness of the image, that is, the Smoothness Attention Network (SA-Net), so that more camera traces of the image can be extracted from the smooth areas of the image, thereby counteracting the compression properties of social network transmission. That is, since the smooth areas of the image will retain more camera traces after the image is transmitted and are less affected by the transmission, the smoothness awareness network can be used as the initial image recognition model for subsequent model training.
[0076] Furthermore, when acquiring a transmission image, in order to integrate noise affected by social networks into the transmission image, a preset social application may be used to transmit the sample image, and the transmitted image may be obtained as the transmission image. In this embodiment, the specific implementation is as follows:
[0077] Acquire the sample image; transmit the sample image through a preset social application to obtain a transmitted sample image; and use the transmitted sample image as a transmission image having a social transmission relationship with the sample image.
[0078] Specifically, the preset social application refers to a communication application used for data transmission and chat between users. Therefore, after obtaining a sample image, the sample image can be first transmitted through the preset social application to obtain a transmitted sample image. The transmitted sample image can then be used as the transmitted image with a social transmission relationship with the sample image for subsequent use.
[0079] It should be noted that during the image recognition model training phase, the social transmission relationship can be an identity mapping, meaning no transmission is performed, preserving the sample image's original state to meet model training requirements. Furthermore, in practical applications, when using a trained image recognition model, there is no need to perform social transmission on the input image; model prediction can be performed directly.
[0080] In practical applications, in order to improve the accuracy of model training, when constructing the transmission image, you can use a variety of different types of social applications for transmission, so that different transmission applications can be obtained according to the transmission results, and then used for model training.
[0081] This embodiment is a one-sided description, and the above description is illustrated by taking the transmission of a sample image through a social application as an example; a sample image X collected by a brand A mobile phone is obtained, and the sample image X is transmitted through a social application D. After the transmission, a transmitted image Y will be obtained, and the transmitted image Y can be used to train the image recognition model subsequently.
[0082] In summary, by using a preset social application to transmit sample images to obtain a transmitted image that incorporates transmission noise, and using this as the basis for model training, the model can have the compression property of resisting social network transmission, thereby improving the model's recognition accuracy of the image source under social network transmission.
[0083] Step S204 : extracting the smoothness feature and the image feature of the transmission image through the initial image recognition model, and determining and outputting the image acquisition information corresponding to the transmission image according to the smoothness feature and the image feature.
[0084] Specifically, after obtaining the transmission image as mentioned above and inputting it into the initial image recognition model, further, since the initial image recognition model is a smoothness perception network, the smoothness features and image features of the transmission image can be extracted through the initial image recognition model, and the implementation model can calculate the image acquisition information corresponding to the transmission image based on the extracted smoothness features and image features, and after outputting the model, it is convenient to subsequently combine with the classifier to complete the training of the model.
[0085] The smoothness feature specifically refers to the vector representation of the smoothness corresponding to the smooth areas in the transmitted image. While the noise generated by social network transmission is not easily perceived, the smooth areas of the image are less affected by social network transmission. Therefore, it can be determined that smooth areas retain more camera traces. Therefore, the smoothness feature can be used to determine the subsequent image acquisition information. Accordingly, the image feature specifically refers to the vector representation corresponding to the transmitted image. Accordingly, the image acquisition information is the camera trace information output by the initial image recognition model, which represents the acquisition source of the transmitted image, i.e., the camera fingerprint corresponding to the transmitted image.
[0086] When extracting the smoothness features of the transmitted image through the initial image recognition model, in order to improve the extraction accuracy and correspond to the smooth area of the image, the following method can be used:
[0087] The transmission image is divided by the initial image recognition model to obtain multiple transmission image regions; the regional features corresponding to each of the multiple transmission image regions are calculated; the regional features corresponding to each transmission image region are reshaped, and the smoothness features of the transmission image are determined according to the processing results.
[0088] Specifically, the transmission image area specifically refers to the various image areas obtained after dividing the transmission image according to a set method; accordingly, the regional feature specifically refers to the vector expression corresponding to each transmission image area; accordingly, the reshaping processing specifically refers to reshaping the regional features of each transmission image area to obtain the smoothness feature.
[0089] Based on this, after the transmission image is input into the initial image recognition model, the transmission image can be divided by the initial image recognition model to obtain multiple transmission image regions according to the division results; then, the regional features corresponding to each of the multiple transmission image regions are calculated; and then, by reshaping the regional features corresponding to each transmission image region, the smoothness characteristics of the transmission image can be determined according to the processing results.
[0090] In practice, the Shannon entropy algorithm can be used to enable computers to understand the smoothness of an image region. Since Shannon entropy measures the average level of information contained in a random variable, it can naturally be used to calculate the smoothness of an area. In other words, smoother areas contain fewer different pixel values and therefore have less entropy.
[0091] Based on this, the smoothness matrix S is used in this embodiment to represent the smoothness of the image area. Specifically, the smoothness matrix S is first set to a size of Then the transmission image of size H×W is divided into and width is of non-overlapping regions {R i}, and use the following formula (1) to calculate the Shannon entropy E of each non-overlapping region i , where formula (1) is as follows:
[0092]
[0093]
[0094] Where i represents the number of non-overlapping regions, E i represents the Shannon entropy of each region, p v Represents the probability associated with the RGB value v, where v is the RGB value of the image. The smoothness matrix S (smoothness feature) is reshaped by It is found that in practical applications, in order to reduce image complexity, the input transmission image can be grayscale processed to achieve the purpose of reducing complexity, and one-hot encoding can be used for efficient batch processing.
[0095] Continuing with the above example, after the transmission image Y is input into the image recognition model, the transmission image can be divided according to the set height and width to obtain multiple non-overlapping regions. Then, the above formula (1) is used to calculate the Shannon entropy corresponding to each non-overlapping region. After that, the smoothness matrix S corresponding to the transmission image Y can be obtained by reshaping the Shannon entropy corresponding to each non-overlapping region.
[0096] In summary, by segmenting the image and calculating the regional features of each image area, it is possible to obtain a feature expression that fully reflects the smooth area of the image when determining the smoothness feature, thereby improving the model training accuracy in subsequent model training.
[0097] Furthermore, the determining and outputting of image acquisition information corresponding to the transmission image according to the smoothness feature and the image feature includes steps S2042 to S2048.
[0098] Step S2042, determining an internal image feature corresponding to the transmitted image according to the smoothness feature and the image feature;
[0099] Specifically, the internal features of the image specifically refer to the internal features corresponding to the transmitted image, which are obtained by calculating the attention mechanism based on the smoothness features and image features.
[0100] Based on this, considering that after the computer can understand the smoothness of the transmitted image, the smoothness of the image also needs to be used to guide the training of the initial image recognition model. Therefore, in order to guide the training of the initial image recognition model based on the smoothness, a smoothness perception module can be embedded in the initial image recognition model, so that the initial image recognition model has the ability to utilize smoothness features. In this way, a self-attention mechanism is applied to the internal features of the image, thereby guiding the training of the initial image recognition model and achieving the purpose of reducing the impact of social network transmission. Therefore, after obtaining the smoothness features and image features, the internal image features corresponding to the transmitted image can be determined based on the smoothness features and image features to facilitate subsequent use.
[0101] Furthermore, determining the internal image features corresponding to the transmitted image according to the smoothness features and the image features includes:
[0102] The smoothness feature is converted into a flattened smoothness feature, and the image feature is converted into a flattened image feature; based on the flattened smoothness feature and the flattened image feature, an attention mechanism of each image processing channel is calculated to obtain the internal image features corresponding to the transmitted image.
[0103] Specifically, the flattened smoothness feature refers to a vector expression obtained by flattening the smoothness feature; correspondingly, the flattened image feature refers to a vector expression obtained by flattening the image feature.
[0104] Based on this, in order to obtain the internal image features of the transmitted image, the smoothness features and image features can be flattened separately to obtain flattened smoothness features and flattened image features. Then, the flattened smoothness features and flattened image features are combined to calculate the attention mechanism of each image processing channel. According to the calculation results, the internal image features corresponding to the transmitted image can be obtained.
[0105] That is to say, after obtaining the image feature X of the transmitted image f After the smoothness feature S, the image feature X f And the flattened size of the smoothness feature S is Flat image features and flattening smoothness features Then flatten the image features and flattening smoothness features Apply the K-head attention mechanism on The attention mechanism on the channel can obtain the internal features of the image, that is, the internal features corresponding to the transmitted image in, It can be calculated by the following formula (2):
[0106]
[0107] Among them, Q K is the query matrix of the k-th projection of the self-attention mechanism function, K K is the key matrix of the k-th projection of the self-attention mechanism function, V K is the value matrix of the kth projection of the self-attention mechanism function. The self-attention mechanism is obtained by the following formula (3):
[0108]
[0109] In summary, by combining the smoothness characteristics of the transmitted image with the image features to calculate the internal features of the image, the model can learn the influence of the social transmission network, thereby ensuring that the model can eliminate this influence during the application phase and improve the accuracy of image source identification.
[0110] Step S2044: linearize the internal features of the image to obtain flattened attention features;
[0111] Step S2046, reshaping the flattened attention feature based on the attribute information of the transmitted image to obtain a target attention feature;
[0112] Step S2048: convert the target attention feature and obtain the image acquisition information output by the initial image recognition model based on the conversion result.
[0113] Specifically, after obtaining the internal features of the image as mentioned above, in order to ensure that the model can output image acquisition information, and the image acquisition information corresponds to the camera fingerprint corresponding to the transmitted image, the transmitted image can be linearized first to obtain flattened attention features. Thereafter, the flattened attention features are reshaped based on the attribute information of the transmitted image to obtain the target attention features; finally, the target attention features are converted, and the image acquisition information output by the initial image recognition model can be obtained according to the conversion results.
[0114] The flattened attention feature specifically refers to the flattened attention mechanism, and the target attention feature specifically refers to the attention feature obtained by reshaping the resolution.
[0115] That is to say, after obtaining the internal features of the image Then, the union of all projections can be Linear projection to produce flattened attention in The calculation of can be obtained by the following formula (4):
[0116]
[0117] Here, MLP() represents a multi-layer perceptron with GELU activation function.
[0118] Getting flattened attention After that, the refined target attention feature is obtained by reshaping it to H×W×C resolution, that is, the attention feature X a , to facilitate subsequent conversion, obtain image acquisition information and output the initial image recognition model.
[0119] Furthermore, the target attention feature is converted, and the image acquisition information output by the initial image recognition model is obtained according to the conversion result, including:
[0120] Determine the encoding information of the initial image recognition model, wherein the encoding information includes global average pooling information and linear mapping information; encode the target attention feature according to the global average pooling information and the linear mapping information to obtain the encoding feature; convert the encoding feature through the output layer of the initial image recognition model to obtain the image acquisition information and output it.
[0121] Specifically, the encoding information refers to the information obtained by encoding the target attention feature, which includes global average pooling information and linear mapping information. Correspondingly, the encoding feature refers to the vector representation obtained after encoding the target attention feature, which is not converted into image acquisition information, namely the camera fingerprint.
[0122] Based on this, taking into account the dimensional redundancy of the target attention features, in order to reduce the impact of redundant dimensions, the encoding information of the initial image recognition model can be determined first, and the encoding information includes global average pooling information and linear mapping information; secondly, according to the global average pooling information and linear mapping information, the target attention features are encoded to obtain encoding features; finally, the encoding features are converted through the output layer of the initial image recognition model to obtain image acquisition information and output it.
[0123] That is, considering the attention feature X a The dimension redundancy of X can be reduced by global average pooling and linear mapping. a Encoding to low-dimensional space R d , thus obtaining the final image acquisition information, namely the camera representation T. T represents the features extracted by the initial image recognition model, namely the camera trace. Specifically, it can be a 256-dimensional vector, such as [1.02, 3.22, 5, 7.33, …].
[0124] It should be noted that for images X1 and X2 taken with the same camera, the traces T1 and T2 extracted from them are expected to be as similar as possible, for example, T1 = [1.02, 3.22, 5, 7.33, ...], T2 = [0.98, 3.21, 5.10, 7.99, ...]; while the traces extracted from images taken with different cameras are expected to be as different as possible, for example, T3 = [9.02, -3.22, 58, 62, ...]. This ensures higher recognition accuracy when training the image recognition model.
[0125] Continuing with the above example, we can obtain the smoothness matrix S and image features X corresponding to the transmitted image Y. f After that, we can use the image feature Xf And the flattened size of the smoothness feature S is Flat image features and flattening smoothness features Then flatten the image features and flattening smoothness features Apply the K-head attention mechanism on The attention mechanism on the channel can obtain the internal features corresponding to the transmitted image in It can be obtained by formula (2). Further, after obtaining the union of all projections Linear projection to produce flattened attention in It can be obtained by formula (4). Furthermore, after obtaining the flattened attention After that, the refined target attention feature is obtained by reshaping it to H×W×C resolution, that is, the attention feature X a Finally, global average pooling and linear mapping are used to transform X a Encoding to low-dimensional space R d , thus obtaining the final image acquisition information, that is, the camera representation T. At this time, it is determined that the camera fingerprint of the transmitted image is the information of brand A mobile phone or brand B mobile phone.
[0126] In summary, by incorporating smoothness features into the initial image recognition model, it is possible to extract features from the smooth areas of the transmitted image, which retains more information about the image source and is less affected by the transmission. Training the image recognition model on this basis allows the model to learn the properties of resisting social network transmission, thereby making it more accurate in identifying the source of images in social network transmission scenarios.
[0127] Step S206: Determine the mapping relationship between the image label of the sample image and the image acquisition information through a classifier.
[0128] Specifically, after obtaining the image acquisition information output by the initial image recognition model as mentioned above, further, in order to make the image acquisition information raised by the initial image recognition model discriminative, a classifier can be introduced in the model training stage to distinguish the image acquisition information relative to the image label, thereby determining the mapping relationship between the image label and the image acquisition information, so as to facilitate the completion of model parameter adjustment according to the mapping relationship in the model optimization stage.
[0129] Among them, the image label specifically refers to the real image acquisition information corresponding to the sample image, that is, the image source information; correspondingly, the mapping relationship between the image label and the image acquisition information specifically refers to the relationship that describes whether the image acquisition information represents the image label, which is used to subsequently calculate the loss function to complete the optimization of the model.
[0130] Based on this, determining the mapping relationship between the image label of the sample image and the image acquisition information by a classifier includes:
[0131] Determine a classifier corresponding to the target social scene based on the social transmission relationship; input the image acquisition information and the image label corresponding to the sample image into the classifier for processing to obtain discrimination information output by the classifier; determine a mapping relationship between the image label and the image acquisition information based on the discrimination information.
[0132] Specifically, the target social scenario refers to the scenario corresponding to the social transmission relationship. Each scenario corresponds to a different transmission method, thus requiring a different classifier. For the sample image and its corresponding transmission image, the classifier for the corresponding social transmission relationship scenario is used for model optimization. Accordingly, the discriminant information specifically refers to the output of the classifier, indicating that the image acquisition information can be used to represent the image label of the sample image, that is, the actual image acquisition information. This facilitates use during the model optimization phase.
[0133] Based on this, we first determine the target social scene that can accurately correspond to the social transmission relationship, then obtain the classifier preset for the target social scene, and then input the image acquisition information and the image label corresponding to the sample image into the classifier for classification processing, and obtain the discrimination information output by the classifier to determine whether the image acquisition information can represent the image label. Finally, the mapping relationship between the image label and the image acquisition information is determined based on the discrimination information, which is convenient for subsequent model parameter adjustment.
[0134] Continuing with the above example, if the camera fingerprint output by the initial image recognition model is information about a brand B mobile phone, we can first determine the classifier corresponding to social application D. Then, we can input the brand B mobile phone information and the image label {brand A mobile phone} corresponding to the sample image X into the classifier corresponding to the social application D scenario for processing to determine whether the camera fingerprint can be used to represent the image label corresponding to the sample image X. Based on the judgment result, it is determined that the image corresponding to the camera fingerprint of the brand B mobile phone information is not the image with the corresponding image label. Therefore, we can further adjust the parameters of the initial image recognition model based on the sample and label until an image recognition model that meets the training stop condition is obtained.
[0135] In summary, by combining the classifier to determine the mapping relationship, it is possible to achieve higher discrimination in the initial image recognition model when determining image acquisition information during the training phase, thereby effectively improving the model training accuracy and enabling the model to have better recognition capabilities in social transmission scenarios.
[0136] Step S208: Adjust the parameters of the initial image recognition model according to the mapping relationship until a target image recognition model that meets the training stop condition is obtained.
[0137] Specifically, after the classifier determines the mapping relationship between image labels and image acquisition information, the initial image recognition model can be further adjusted based on the mapping relationship. After the parameter adjustment is completed, the model can be retrained based on new samples until a target image recognition model that meets the training stop conditions is obtained. The training stop conditions include but are not limited to iteration stop conditions, loss value comparison conditions, or validation set verification conditions.
[0138] Furthermore, when the loss value comparison condition is used to adjust the model parameters, the initial image recognition model is adjusted according to the mapping relationship until a target image recognition model that meets the training stop condition is obtained, including:
[0139] The model loss value corresponding to the initial image recognition model is calculated according to the mapping relationship; and the parameters of the initial image recognition model are adjusted based on the model loss value until a target image recognition model that meets the loss value comparison condition is obtained.
[0140] Specifically, the model loss value refers to the loss value calculated by the loss function, such as the cross-entropy loss value calculated using the cross-entropy loss function. Based on this, after obtaining the mapping relationship, the model loss value corresponding to the initial image recognition model can be calculated according to the mapping relationship; then, the initial image recognition model is adjusted based on the model loss value. At this time, if the model loss value is not less than the preset loss value threshold, it continues to be trained based on new samples until the loss value of a certain training cycle is less than the loss value threshold. The final model can be used as the target image recognition model. After any image is input into the model, the image acquisition information corresponding to the image can be output with higher accuracy.
[0141] The image recognition model training method provided in this embodiment can improve the robustness of the target recognition model and ensure that the model can accurately identify the image source in the social transmission scenario. In the model training stage, a sample image and a transmission image that has a social transmission relationship with the sample image can be obtained. At this time, the transmission image is first input into the initial image recognition model, and the smoothness feature and image feature of the transmission image are extracted by the initial image recognition model to determine the image acquisition information corresponding to the transmission image based on the smoothness feature and image feature and output it. At this time, the mapping relationship between the image label and the image acquisition information is determined in combination with the classifier, and the parameters of the initial image recognition model are adjusted based on this until a target image recognition model that meets the training stop condition is obtained. In the training process, by extracting the smoothness feature, the robustness of the model against different social network scenarios is greatly improved, so that the model can more accurately extract the image traces of the transmitted image to meet the needs of downstream business.
[0142] Figure 3 A flowchart of a second image recognition model training method provided according to an embodiment of this specification is shown, which specifically includes the following steps.
[0143] Step S302 : obtaining a sample image and a plurality of transmission images having a social transmission relationship with the sample image, and inputting the plurality of transmission images into an initial image recognition model respectively.
[0144] Step S304 : extracting the smoothness feature and image feature of each transmission image through the initial image recognition model, and determining and outputting the image acquisition information corresponding to each transmission image based on the smoothness feature and image feature.
[0145] Step S306: Determine the mapping relationship between the image label of the sample image and each image acquisition information through a classifier.
[0146] Step S308: Calculate target model parameters based on the mapping relationship between the image label and each image acquisition information.
[0147] Step S310: Adjust the parameters of the initial image recognition model based on the target model parameters until a target image recognition model that meets the training stop conditions is obtained.
[0148] It should be noted that for the contents not described in detail in this embodiment, reference may be made to the same or corresponding descriptions in the above embodiments, and this embodiment does not make any limitation thereto.
[0149] Based on this, when a sample image and its corresponding multiple transmission images are obtained, in order to train an image recognition model that can be applied to more social transmission scenarios, different classifiers can be used for different social network transmission scenarios to determine the relationship between image labels and image acquisition information.
[0150] Furthermore, after obtaining the sample image, the sample image can be transmitted separately through different social applications, thereby obtaining the transmission image after transmission by each social application, that is, multiple transmission images. Thereafter, each transmission image is input into the image recognition model for processing. After model processing, the image acquisition information corresponding to each transmission image will be obtained.
[0151] Furthermore, at this point, the classifier corresponding to each transmitted image can be determined first. This classifier corresponds to the social application to which it belongs, so that the image label and image acquisition information can be determined by using the associated classifier. This facilitates the subsequent calculation of the target model parameters based on the mapping relationship corresponding to each social application, which is used to update the initial image recognition model to obtain a target image recognition model that meets the requirements. Among them, the image recognition model in this embodiment can adopt any neural network, or the SA-Net network as described in the above embodiment. In specific applications, it is selected according to actual needs, and this embodiment does not impose any restrictions here.
[0152] In other words, a single classifier may struggle to optimize the extractor for multiple social media transmission scenarios, especially when the input data has different feature distributions, forcing the image recognition model to learn camera traces that are more general rather than more discriminative. Therefore, multiple classifiers can be used to correspond to different feature spaces, alleviating the dilemma of optimizing multiple scenarios simultaneously.
[0153] For example, after the sample image X is transmitted through social applications D1, D2 and D3 respectively, the transmitted images Y1, Y2 and Y3 will be obtained. At this time, the transmitted images Y1, Y2 and Y3 are respectively input into the image recognition model for processing, and the image acquisition information {A1 brand mobile phone} corresponding to the transmitted image Y1, the image acquisition information {A2 brand mobile phone} corresponding to the transmitted image Y2, and the image acquisition information {A3 brand mobile phone} corresponding to the transmitted image Y3 will be obtained. Afterwards, the image label corresponding to the sample image X and the image acquisition information corresponding to the transmitted image Y1 can be input into the classifier corresponding to the social application D1 for processing, and the discrimination result P1 output by the classifier will be obtained. Similarly, the image label corresponding to the sample image X and the image acquisition information corresponding to the transmitted image Y2 can be input into the classifier corresponding to the social application D2 for processing, and the discrimination result P2 output by the classifier will be obtained. The image label corresponding to the sample image X and the image acquisition information corresponding to the transmitted image Y3 can be input into the classifier corresponding to the social application D3 for processing, and the discrimination result P3 output by the classifier will be obtained. Afterwards, the results output by each classifier can be combined to jointly calculate the model weight and used to update the image recognition model. This process can be repeated until an image recognition model that meets the training stop conditions is obtained.
[0154] Furthermore, according to the mapping relationship between the image label and each image acquisition information, the target model parameters are calculated, including:
[0155] A mask sequence is determined according to the number of pieces of information of the plurality of image acquisition information; and the target model parameters are calculated according to a mapping relationship between the image label and each piece of image acquisition information and the mask sequence.
[0156] During specific implementation, considering the different optimization complexities of different scenarios, it may cause easy-to-learn scenarios to dominate the main learning direction of the initial image recognition model, or optimization conflicts may occur when updating the model weights. Therefore, in order to alleviate the conflicting gradients, the mask sequence can be determined based on the amount of information of multiple image acquisition information; the target model parameters are calculated based on the mapping relationship between the image label and each image acquisition information and the mask sequence.
[0157] That is to say, in the training of the multi-classifier optimized image recognition model, its forward propagation process can extract the image acquisition information of N scenes through the image recognition model (smoothness perception network SA-Net), that is, the camera model trace N classifiers are used to supervise the training process, and the total loss L is calculated by the following formula (5):
[0158]
[0159] Among them, L nis the cross entropy loss:
[0160]
[0161] Here, y represents the label of the image during training. For example, the training set contains 10 camera types: A1, A2, …, and A10. To facilitate training, each label is assigned a unique character, such as 1 for A1, 2 for A2, …, and 10 for A10. During model training, the model only needs to determine whether the sample image can be predicted as the corresponding character; it does not need to be associated with a specific camera type.
[0162] Furthermore, considering that in actual applications, the backpropagation process updates the weight θ of the initial image recognition model as follows
[0163]
[0164] in, Denotes the loss function L n The partial derivative with respect to θ, l is the learning rate.
[0165] However, in many scenarios, It is composed of sub-functions (gradients) corresponding to multiple scenes, which may lead to the network being sub-optimal, that is, updating in the direction of a certain sub-function; in addition, when there is a gradient conflict, optimization will be more difficult. Therefore, in order to ensure that the image recognition model can be trained to meet the needs of use, a set of masks can be introduced. This results in conflict-free gradients And replace the above formula with To update the weights θ of the image recognition model.
[0166] For T n The gradient given at The corresponding mask M n Can be defined as:
[0167]
[0168] in, is a standard exponential function, ⊙ represents element-wise multiplication. U is a random matrix sampled from a uniform distribution U(0,1), To measure the positive sign purity contained in a given gradient, it can be obtained by the following formula (6):
[0169]
[0170] in, The gradient over one training batch is integrated. However, the purity calculated by the above formula (6) The value of depends only on the training data of this batch. This part of the data is difficult to represent the gradient trend of the entire dataset, especially when the amount of data contained in the current batch is small. Therefore, in order to solve this problem, the momentum method can be used to adjust the gradient according to the historical gradient. Define.
[0171] Specifically, the gradients in the first t-1 iterations are collected to define the t-th iteration for:
[0172]
[0173]
[0174] and Initialized to In actual training, the attenuation factor μ can be set to a specified value according to the requirements, such as 0.95. If μ=0, Can degenerate into primitive At this point, the update method of the weight θ of the image recognition model can be calculated by the following formula (7):
[0175]
[0176] The image recognition model training method provided in this embodiment can improve the robustness of the target recognition model and ensure that the model can accurately identify the image source in the social transmission scenario. In the model training stage, a sample image and a transmission image that has a social transmission relationship with the sample image can be obtained. At this time, the transmission image is first input into the initial image recognition model, and the smoothness feature and image feature of the transmission image are extracted by the initial image recognition model to determine the image acquisition information corresponding to the transmission image based on the smoothness feature and image feature and output it. At this time, the mapping relationship between the image label and the image acquisition information is determined in combination with the classifier, and the parameters of the initial image recognition model are adjusted based on this until a target image recognition model that meets the training stop condition is obtained. In the training process, by extracting the smoothness feature, the robustness of the model against different social network scenarios is greatly improved, so that the model can more accurately extract the image traces of the transmitted image to meet the needs of downstream business.
[0177] Figure 4 A flowchart of a third image recognition model training method provided according to an embodiment of this specification is shown, which specifically includes the following steps.
[0178] Step S402 : obtaining a sample image and a plurality of transmission images having a social transmission relationship with the sample image, and inputting the plurality of transmission images into an initial image recognition model respectively.
[0179] Step S404: extracting image features of each transmission image through the initial image recognition model, determining image acquisition information corresponding to each transmission image based on the image features, and outputting the information.
[0180] Step S406: Determine the mapping relationship between the image label of the sample image and each image acquisition information through a classifier.
[0181] Step S408: Calculate target model parameters according to the mapping relationship between the image label and each image acquisition information.
[0182] Step S410: Adjust the parameters of the initial image recognition model based on the target model parameters until a target image recognition model that meets the training stop conditions is obtained.
[0183] In the third image recognition model training method provided in this embodiment, the image recognition model can be any neural network. When extracting image acquisition information, it can be implemented without combining with smoothness features. However, during classifier processing, it is necessary to combine the classifiers corresponding to different scenarios to achieve the training of a target image recognition model that meets the usage requirements. Among them, for the contents not fully described in this embodiment, please refer to the same or corresponding descriptions in the above embodiments, and this embodiment does not make any limitations here.
[0184] The image recognition model training method provided in this embodiment can improve the robustness of the target recognition model and ensure that the model can accurately identify the image source in the social transmission scenario. In the model training stage, a sample image and a transmission image that has a social transmission relationship with the sample image can be obtained. At this time, the transmission image is first input into the initial image recognition model, and the smoothness feature and image feature of the transmission image are extracted by the initial image recognition model to determine the image acquisition information corresponding to the transmission image based on the smoothness feature and image feature and output it. At this time, the mapping relationship between the image label and the image acquisition information is determined in combination with the classifier, and the parameters of the initial image recognition model are adjusted based on this until a target image recognition model that meets the training stop condition is obtained. In the training process, by extracting the smoothness feature, the robustness of the model against different social network scenarios is greatly improved, so that the model can more accurately extract the image traces of the transmitted image to meet the needs of downstream business.
[0185] Corresponding to the above method embodiment, this specification also provides a first image recognition model training device embodiment, Figure 5 This is a schematic diagram of the structure of the first image recognition model training device provided by an embodiment of this specification. Figure 5 As shown, the device includes:
[0186] An acquisition module 502 is configured to acquire a sample image and a transmission image having a social transmission relationship with the sample image, and input the transmission image into an initial image recognition model;
[0187] An extraction module 504 is configured to extract smoothness features and image features of the transmission image using the initial image recognition model, determine image acquisition information corresponding to the transmission image based on the smoothness features and the image features, and output the information;
[0188] A determination module 506 is configured to determine a mapping relationship between the image label of the sample image and the image acquisition information through a classifier;
[0189] The training module 508 is configured to adjust the parameters of the initial image recognition model according to the mapping relationship until a target image recognition model that meets the training stop condition is obtained.
[0190] In an optional embodiment, the acquisition module 502 is further configured to:
[0191] Acquire the sample image; transmit the sample image through a preset social application to obtain a transmitted sample image; and use the transmitted sample image as a transmission image having a social transmission relationship with the sample image.
[0192] In an optional embodiment, the extraction module 504 is further configured to:
[0193] Determine the internal image features corresponding to the transmitted image based on the smoothness features and the image features; perform linearization on the internal image features to obtain flattened attention features; reshape the flattened attention features based on the attribute information of the transmitted image to obtain target attention features; convert the target attention features, and obtain the image acquisition information output by the initial image recognition model based on the conversion result.
[0194] In an optional embodiment, the extraction module 504 is further configured to:
[0195] The transmission image is divided by the initial image recognition model to obtain multiple transmission image regions; the regional features corresponding to each of the multiple transmission image regions are calculated; the regional features corresponding to each transmission image region are reshaped, and the smoothness features of the transmission image are determined according to the processing results.
[0196] In an optional embodiment, the extraction module 504 is further configured to:
[0197] The smoothness feature is converted into a flattened smoothness feature, and the image feature is converted into a flattened image feature; based on the flattened smoothness feature and the flattened image feature, an attention mechanism of each image processing channel is calculated to obtain the internal image features corresponding to the transmitted image.
[0198] In an optional embodiment, the extraction module 504 is further configured to:
[0199] Determine the encoding information of the initial image recognition model, wherein the encoding information includes global average pooling information and linear mapping information; encode the target attention feature according to the global average pooling information and the linear mapping information to obtain the encoding feature; convert the encoding feature through the output layer of the initial image recognition model to obtain the image acquisition information and output it.
[0200] In an optional embodiment, the determining module 506 is further configured to:
[0201] Determine a classifier corresponding to the target social scene based on the social transmission relationship; input the image acquisition information and the image label corresponding to the sample image into the classifier for processing to obtain discrimination information output by the classifier; determine a mapping relationship between the image label and the image acquisition information based on the discrimination information.
[0202] In an optional embodiment, the training module 508 is further configured to:
[0203] The model loss value corresponding to the initial image recognition model is calculated according to the mapping relationship; and the parameters of the initial image recognition model are adjusted based on the model loss value until a target image recognition model that meets the loss value comparison condition is obtained.
[0204] In an optional embodiment, when there are multiple transmitted images, multiple image acquisition information is outputted through the initial image recognition model;
[0205] Accordingly, the determination module 506 is further configured to: determine the mapping relationship between the image label of the sample image and each image acquisition information through a classifier;
[0206] Accordingly, the training module 508 is further configured to: calculate the target model parameters based on the mapping relationship between the image label and each image acquisition information; and adjust the parameters of the initial image recognition model based on the target model parameters until a target image recognition model that meets the training stop conditions is obtained.
[0207] In an optional embodiment, the training module 508 is further configured to:
[0208] A mask sequence is determined according to the number of pieces of information of the plurality of image acquisition information; and the target model parameters are calculated according to a mapping relationship between the image label and each piece of image acquisition information and the mask sequence.
[0209] In summary, in order to improve the robustness of the target recognition model and ensure that the model can accurately identify the image source in social transmission scenarios, sample images and transmission images with social transmission relationships with the sample images can be obtained during the model training phase. At this time, the transmission images are first input into the initial image recognition model, and the smoothness features and image features of the transmission images are extracted through the initial image recognition model. The image acquisition information corresponding to the transmission images is determined based on the smoothness features and image features and output. At this time, the mapping relationship between the image label and the image acquisition information is determined in combination with the classifier, and the parameters of the initial image recognition model are adjusted based on this until a target image recognition model that meets the training stop conditions is obtained. During the training process, by extracting the smoothness features, the robustness of the model against different social network scenarios is greatly improved, so that the model can more accurately extract the image traces of the transmitted images to meet the needs of downstream business.
[0210] The above is a schematic diagram of the first image recognition model training device of this embodiment. It should be noted that the technical solution of this image recognition model training device and the technical solution of the aforementioned image recognition model training method are based on the same concept. For details not described in detail in the technical solution of the image recognition model training device, please refer to the description of the technical solution of the aforementioned image recognition model training method.
[0211] Figure 6 A flowchart of an image recognition method provided according to an embodiment of this specification is shown, which specifically includes the following steps.
[0212] Step S602: obtaining an image to be identified that is associated with the target service;
[0213] Step S604: inputting the image to be recognized into the target image recognition model in the above method;
[0214] Step S606, extracting target smoothness features and target image features of the image to be recognized through the target image recognition model, determining image acquisition information corresponding to the image to be recognized based on the target smoothness features and the target image features, and outputting the information;
[0215] Step S608: Determine service detection information of the target service according to the image acquisition information.
[0216] It should be noted that, in the image recognition method provided in this embodiment, the same or corresponding technical features can refer to the corresponding descriptions in the above embodiments, and this embodiment will not be elaborated in detail here.
[0217] For example, in an image leak tracing scenario, when an image involving sensitive information within a company is posted to the Internet, the public image on the Internet can be input into the target image recognition model for processing, and the device fingerprint corresponding to the image can be obtained. The device fingerprint is then compared with all device fingerprints within the company, thereby assisting in identifying the leaker to achieve the purpose of detection.
[0218] For example, in an ID collection scenario, when a delivery driver uses another person's ID image for identity recognition, the ID image can first be input into the target image recognition model for processing to obtain the device fingerprint corresponding to the image. The device fingerprint is then compared with the device fingerprint of the device currently used by the delivery driver to detect any cheating behavior of the delivery driver and achieve the purpose of detection.
[0219] In summary, by combining the target image recognition model to determine the image acquisition information, the accuracy of the image acquisition information can be ensured. Based on this, it can be applied to the target business to determine the business detection information, which can effectively reduce business losses.
[0220] Corresponding to the above method embodiment, this specification also provides an image recognition device embodiment, Figure 7 This is a schematic diagram of the structure of an image recognition device provided by one embodiment of this specification. Figure 7 As shown, the device includes:
[0221] An image acquisition module 702 is configured to acquire an image to be identified associated with a target service;
[0222] An input model module 704 is configured to input the image to be recognized into the target image recognition model in the above method;
[0223] The prediction information module 706 is configured to extract the target smoothness feature and the target image feature of the image to be recognized through the target image recognition model, determine the image acquisition information corresponding to the image to be recognized based on the target smoothness feature and the target image feature, and output the image acquisition information;
[0224] The information determination module 708 is configured to determine the service detection information of the target service according to the image acquisition information.
[0225] The above is a schematic diagram of an image recognition device according to this embodiment. It should be noted that the technical solution of the image recognition device and the technical solution of the above-mentioned image recognition method are based on the same concept. For details not described in detail in the technical solution of the image recognition device, please refer to the description of the technical solution of the above-mentioned image recognition method.
[0226] Figure 8 This is a flowchart of the fourth image recognition model training method provided in an embodiment of this specification, which is applied to cloud-side devices and specifically includes the following steps.
[0227] Step S802: Acquire a sample image submitted by the end-side device;
[0228] Step S804: generating a transmission image having a social transmission relationship with the sample image, and inputting the transmission image into an initial image recognition model;
[0229] Step S806: extracting smoothness features and image features of the transmission image through the initial image recognition model, determining image acquisition information corresponding to the transmission image based on the smoothness features and the image features, and outputting the information;
[0230] Step S808: determining a mapping relationship between the image label of the sample image and the image acquisition information through a classifier, and adjusting parameters of the initial image recognition model according to the mapping relationship to obtain a target image recognition model;
[0231] Step S810: Send the target model parameters corresponding to the target image recognition model to the terminal side device.
[0232] It should be noted that, in the method provided in this embodiment, the same or corresponding technical features can refer to the corresponding description content in the above embodiments, and this embodiment will not be described in detail here.
[0233] Specifically, a terminal device refers to a user-side device that requires model training. It submits sample images for training on the cloud-side device. After training is complete, the target model parameters corresponding to the target image recognition model are determined and fed back to the terminal device for use. The target model parameters are the parameters corresponding to the trained target recognition model. The pre-cloud-side device is the server-side device that provides model training capabilities.
[0234] In summary, in order to improve the robustness of the target recognition model and ensure that the model can accurately identify the image source in social transmission scenarios, sample images and transmission images with social transmission relationships with the sample images can be obtained during the model training phase. At this time, the transmission images are first input into the initial image recognition model, and the smoothness features and image features of the transmission images are extracted through the initial image recognition model. The image acquisition information corresponding to the transmission images is determined based on the smoothness features and image features and output. At this time, the mapping relationship between the image label and the image acquisition information is determined in combination with the classifier, and the parameters of the initial image recognition model are adjusted based on this until a target image recognition model that meets the training stop conditions is obtained. During the training process, by extracting the smoothness features, the robustness of the model against different social network scenarios is greatly improved, so that the model can more accurately extract the image traces of the transmitted images to meet the needs of downstream business.
[0235] Corresponding to the above method embodiment, this specification also provides a second image recognition model training device embodiment, Figure 9 This is a schematic diagram of the structure of the second image recognition model training device provided by an embodiment of this specification. Figure 9 As shown, the device includes:
[0236] An acquisition module 902 is configured to acquire a sample image submitted by a terminal device;
[0237] A generating module 904 is configured to generate a transmission image having a social transmission relationship with the sample image, and input the transmission image into an initial image recognition model;
[0238] The processing module 906 is configured to extract the smoothness feature and the image feature of the transmission image through the initial image recognition model, determine the image acquisition information corresponding to the transmission image according to the smoothness feature and the image feature, and output the image acquisition information;
[0239] a determination module 908 configured to determine a mapping relationship between the image label of the sample image and the image acquisition information through a classifier, and adjust parameters of the initial image recognition model according to the mapping relationship to obtain a target image recognition model;
[0240] The sending module 910 is configured to send the target model parameters corresponding to the target image recognition model to the terminal side device.
[0241] In summary, in order to improve the robustness of the target recognition model and ensure that the model can accurately identify the image source in social transmission scenarios, sample images and transmission images with social transmission relationships with the sample images can be obtained during the model training phase. At this time, the transmission images are first input into the initial image recognition model, and the smoothness features and image features of the transmission images are extracted through the initial image recognition model. The image acquisition information corresponding to the transmission images is determined based on the smoothness features and image features and output. At this time, the mapping relationship between the image label and the image acquisition information is determined in combination with the classifier, and the parameters of the initial image recognition model are adjusted based on this until a target image recognition model that meets the training stop conditions is obtained. During the training process, by extracting the smoothness features, the robustness of the model against different social network scenarios is greatly improved, so that the model can more accurately extract the image traces of the transmitted images to meet the needs of downstream business.
[0242] The above is a schematic diagram of the second image recognition model training device of this embodiment. It should be noted that the technical solution of this image recognition model training device and the technical solution of the aforementioned image recognition model training method are based on the same concept. For details not described in detail in the technical solution of the image recognition model training device, please refer to the description of the technical solution of the aforementioned image recognition model training method.
[0243] Figure 10 This is a flowchart of the fifth image recognition model training method provided by an embodiment of this specification, which specifically includes the following steps.
[0244] Step S1002: obtaining a sample image and a transmission image having a social transmission relationship with the sample image, and inputting the transmission image into an initial image recognition model;
[0245] Step S1004: extracting smoothness features and image features of the transmission image through the initial image recognition model, determining image acquisition information corresponding to the transmission image based on the smoothness features and the image features, and outputting the information;
[0246] Step S1006: Adjust the parameters of the initial image recognition model according to the image label of the sample image and the image acquisition information until a target image recognition model that meets the training stop condition is obtained.
[0247] It should be noted that, in the method provided in this embodiment, the same or corresponding technical features can refer to the corresponding description content in the above embodiments, and this embodiment will not be described in detail here.
[0248] In summary, to improve the robustness of the target recognition model and ensure that the model can accurately identify image sources in social transmission scenarios, during the model training phase, sample images and transmitted images with social transmission relationships with the sample images can be obtained. At this time, the transmitted images are first input into the initial image recognition model, which then extracts the smoothness and image features of the transmitted images. Based on these features, the image acquisition information corresponding to the transmitted images is determined and output. Based on this, the parameters of the initial image recognition model are adjusted until a target image recognition model that meets the training termination criteria is obtained. During the training process, by extracting smoothness features, the model's robustness against different social network scenarios is greatly improved, enabling the model to more accurately extract image traces of transmitted images to meet downstream business needs.
[0249] Corresponding to the above method embodiment, this specification also provides a third image recognition model training device embodiment, Figure 11 This is a schematic diagram of the structure of the third image recognition model training device provided by an embodiment of this specification. Figure 11 As shown, the device includes:
[0250] An acquisition module 1102 is configured to acquire a sample image and a transmission image having a social transmission relationship with the sample image, and input the transmission image into an initial image recognition model;
[0251] The processing module 1104 is configured to extract the smoothness feature and the image feature of the transmission image through the initial image recognition model, determine the image acquisition information corresponding to the transmission image according to the smoothness feature and the image feature, and output the image acquisition information;
[0252] The training module 1106 is configured to adjust the parameters of the initial image recognition model according to the image label of the sample image and the image acquisition information until a target image recognition model that meets the training stop condition is obtained.
[0253] In summary, to improve the robustness of the target recognition model and ensure that the model can accurately identify image sources in social transmission scenarios, during the model training phase, sample images and transmitted images with social transmission relationships with the sample images can be obtained. At this time, the transmitted images are first input into the initial image recognition model, which then extracts the smoothness and image features of the transmitted images. Based on these features, the image acquisition information corresponding to the transmitted images is determined and output. Based on this, the parameters of the initial image recognition model are adjusted until a target image recognition model that meets the training termination criteria is obtained. During the training process, by extracting smoothness features, the model's robustness against different social network scenarios is greatly improved, enabling the model to more accurately extract image traces of transmitted images to meet downstream business needs.
[0254] The above is a schematic diagram of the third image recognition model training device of this embodiment. It should be noted that the technical solution of this image recognition model training device and the technical solution of the aforementioned image recognition model training method are based on the same concept. For details not described in detail in the technical solution of the image recognition model training device, please refer to the description of the technical solution of the aforementioned image recognition model training method.
[0255] Corresponding to the above method embodiment, this specification also provides an image recognition model training system, which includes a model training end and a parameter acquisition end; the parameter acquisition end is used to collect model parameters, and the model training end is used to execute image recognition model training instructions, which implement the steps of the above image recognition model training method when the image recognition model training instructions are executed by the model training end.
[0256] The above is a schematic diagram of an image recognition model training system according to this embodiment. It should be noted that the technical solution of this image recognition model training system and the technical solution of the aforementioned image recognition model training method are based on the same concept. For details not described in detail in the technical solution of the image recognition model training system, please refer to the description of the technical solution of the aforementioned image recognition model training method.
[0257] Figure 12 The following is a block diagram of a computing device 1200 according to one embodiment of the present disclosure. Components of the computing device 1200 include, but are not limited to, a memory 1210 and a processor 1220. The processor 1220 is connected to the memory 1210 via a bus 1230, and a database 1250 is used to store data.
[0258] The computing device 1200 also includes an access device 1240 that enables the computing device 1200 to communicate via one or more networks 1260. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1240 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0259] In one embodiment of the present application, the above components of the computing device 1200 and Figure 12 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 12 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of the present application. Those skilled in the art may add or replace other components as needed.
[0260] Computing device 1200 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1200 may also be a mobile or stationary server.
[0261] Among them, the processor 1220 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned image recognition model training method and image recognition method.
[0262] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solutions of the aforementioned image recognition model training method and image recognition method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the aforementioned image recognition model training method and image recognition method.
[0263] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned image recognition model training method and image recognition method.
[0264] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solutions of the image recognition model training method and the image recognition method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the image recognition model training method and the image recognition method described above.
[0265] An embodiment of the present specification further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned image recognition model training method and image recognition method.
[0266] The above is a schematic diagram of a computer program according to this embodiment. It should be noted that the technical solution of this computer program is based on the same concept as the technical solutions of the image recognition model training method and the image recognition method described above. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solutions of the image recognition model training method and the image recognition method described above.
[0267] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0268] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0269] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0270] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0271] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for training an image recognition model, comprising: Acquire a sample image and a transmission image having a social transmission relationship with the sample image, and input the transmission image into an initial image recognition model; extracting smoothness features and image features of the transmitted image through the initial image recognition model; Converting the smoothness feature into a flattened smoothness feature, and converting the image feature into a flattened image feature; calculating an attention mechanism for each image processing channel according to the flattened smoothness feature and the flattened image feature to obtain an internal image feature corresponding to the transmitted image; Performing linearization processing on the internal features of the image to obtain flattened attention features; Reshaping the flattened attention feature based on the attribute information of the transmitted image to obtain a target attention feature; Converting the target attention feature and obtaining image acquisition information output by the initial image recognition model according to the conversion result; Determining a mapping relationship between the image label of the sample image and the image acquisition information by a classifier; The parameters of the initial image recognition model are adjusted according to the mapping relationship until a target image recognition model that meets the training stop condition is obtained.
2. The method according to claim 1, wherein obtaining the sample image and the transmission image having a social transmission relationship with the sample image comprises: acquiring the sample image; Transmitting the sample image through a preset social application to obtain a transmitted sample image; The transmitted sample image is used as a transmission image having a social transmission relationship with the sample image.
3. The method according to claim 1 or 2, wherein extracting the smoothness feature of the transmitted image by the initial image recognition model comprises: Dividing the transmission image by the initial image recognition model to obtain a plurality of transmission image regions; Calculating a region feature corresponding to each of the plurality of transmitted image regions; The regional features corresponding to each transmission image region are reshaped, and the smoothness feature of the transmission image is determined according to the processing result.
4. The method according to claim 1, wherein converting the target attention feature and obtaining the image acquisition information output by the initial image recognition model according to the conversion result comprises: Determining encoding information of the initial image recognition model, wherein the encoding information includes global average pooling information and linear mapping information; Encoding the target attention feature according to the global average pooling information and the linear mapping information to obtain an encoded feature; The coding features are converted through the output layer of the initial image recognition model to obtain the image acquisition information and output it.
5. The method according to claim 1, wherein determining the mapping relationship between the image label of the sample image and the image acquisition information by a classifier comprises: Determine a classifier corresponding to the target social scene according to the social transmission relationship; Inputting the image acquisition information and the image label corresponding to the sample image into the classifier for processing to obtain the discrimination information output by the classifier; A mapping relationship between the image label and the image acquisition information is determined according to the discrimination information.
6. The method according to claim 1, wherein the initial image recognition model is adjusted according to the mapping relationship until a target image recognition model that meets the training stop condition is obtained, comprising: Calculating a model loss value corresponding to the initial image recognition model according to the mapping relationship; The parameters of the initial image recognition model are adjusted based on the model loss value until a target image recognition model that meets the loss value comparison condition is obtained.
7. The method according to claim 1, wherein when there are multiple transmitted images, the initial image recognition model outputs multiple image acquisition information; Accordingly, determining the mapping relationship between the image label of the sample image and the image acquisition information by a classifier includes: Determine, by a classifier, a mapping relationship between the image label of the sample image and each image acquisition information; Accordingly, the initial image recognition model is adjusted according to the mapping relationship until a target image recognition model that meets the training stop condition is obtained, including: Calculating target model parameters according to the mapping relationship between the image label and each image acquisition information; The initial image recognition model is adjusted based on the target model parameters until a target image recognition model that meets the training stop condition is obtained.
8. The method according to claim 7, wherein calculating target model parameters based on the mapping relationship between the image label and each image acquisition information comprises: determining a mask sequence according to the amount of information of the plurality of image acquisition information; The target model parameters are calculated according to the mapping relationship between the image label and each image acquisition information and the mask sequence.
9. An image recognition method, comprising: Obtaining an image to be identified associated with a target business; Inputting the image to be recognized into the target image recognition model in the method according to any one of claims 1 to 8; Extracting target smoothness features and target image features of the image to be identified through the target image recognition model, determining image acquisition information corresponding to the image to be identified based on the target smoothness features and the target image features, and outputting the information; Determine service detection information of the target service based on the image acquisition information.
10. A method for training an image recognition model, applied to a cloud-side device, comprising: Obtain sample images submitted by the client device; generating a transmission image having a social transmission relationship with the sample image, and inputting the transmission image into an initial image recognition model; extracting smoothness features and image features of the transmitted image through the initial image recognition model; Converting the smoothness feature into a flattened smoothness feature, and converting the image feature into a flattened image feature; calculating an attention mechanism for each image processing channel according to the flattened smoothness feature and the flattened image feature to obtain an internal image feature corresponding to the transmitted image; Performing linearization processing on the internal features of the image to obtain flattened attention features; Reshaping the flattened attention feature based on the attribute information of the transmitted image to obtain a target attention feature; Converting the target attention feature and obtaining image acquisition information output by the initial image recognition model according to the conversion result; Determining a mapping relationship between the image label of the sample image and the image acquisition information through a classifier, and adjusting parameters of the initial image recognition model according to the mapping relationship to obtain a target image recognition model; The target model parameters corresponding to the target image recognition model are sent to the terminal side device.
11. A method for training an image recognition model, comprising: Acquire a sample image and a transmission image having a social transmission relationship with the sample image, and input the transmission image into an initial image recognition model; extracting smoothness features and image features of the transmitted image through the initial image recognition model; Converting the smoothness feature into a flattened smoothness feature, and converting the image feature into a flattened image feature; calculating an attention mechanism for each image processing channel according to the flattened smoothness feature and the flattened image feature to obtain an internal image feature corresponding to the transmitted image; Performing linearization processing on the internal features of the image to obtain flattened attention features; Reshaping the flattened attention feature based on the attribute information of the transmitted image to obtain a target attention feature; Converting the target attention feature and obtaining image acquisition information output by the initial image recognition model according to the conversion result; The initial image recognition model is adjusted according to the image label of the sample image and the image acquisition information until a target image recognition model that meets the training stop condition is obtained.
12. An image recognition model training system, comprising: Model training end and parameter collection end; The parameter acquisition end is used to acquire model parameters, and the model training end is used to execute image recognition model training instructions. When the image recognition model training instructions are executed by the model training end, the steps of the method described in any one of claims 1 to 8 are implemented.