Methods for generating disturbance information, methods for identifying objects, devices, equipment and media
By acquiring image feature information to update perturbation information and generating target perturbation information, the problem of complex perturbation generation methods and poor display effects in existing technologies is solved, thereby improving the robustness of image encryption and detection models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI SENSETIME INTELLIGENT TECH CO LTD
- Filing Date
- 2022-07-21
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies involve complex methods for generating perturbations, resulting in poor image display quality after perturbation, which affects the robustness and security of the image.
By acquiring feature information from the first and second images, updating the initial perturbation information, generating target perturbation information, and using a trigger to determine whether the trained detection model has been trained using the third image, an image with the same type of object is generated.
It enables the rapid and accurate generation of target perturbation information, reduces the impact of superimposed perturbations on image display, and improves the encryption effect of images and the robustness of detection models.
Smart Images

Figure CN115270151B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to, but is not limited to, the field of computer vision technology, and in particular to a method for generating perturbation information, an object recognition method, an apparatus, a device, and a medium. Background Technology
[0002] In recent years, artificial intelligence (AI) technologies, including computer vision, natural language processing, and speech recognition, have developed rapidly, and the security of AI technologies has received unprecedented attention. Currently, data encryption can be achieved by adding perturbations to prevent the arbitrary collection of published data. Alternatively, perturbated data can be used for adversarial training of machine learning models, thereby improving the robustness of the models. However, some techniques suffer from complex perturbation generation methods and poor display effects after adding perturbations. Summary of the Invention
[0003] In view of this, the present disclosure provides at least one method for generating disturbance information, an object identification method, an apparatus, a device, and a medium.
[0004] The technical solution of this disclosure embodiment is implemented as follows:
[0005] On one hand, embodiments of this disclosure provide a method for generating disturbance information, the method comprising:
[0006] Acquire a first image and a second image; the first image is determined based on a first original image and initial perturbation information, and the second image is determined based on a second original image and a trigger.
[0007] The feature information of the first image and the feature information of the second image are determined respectively;
[0008] Based on the feature information of the first image and the feature information of the second image, the initial disturbance information is updated to obtain the target disturbance information;
[0009] The trigger is used to determine whether the trained detection model was trained using a third image, which is determined based on the target perturbation information and the third original image, wherein the first original image, the second original image, and the third original image have objects of the same type.
[0010] On the other hand, embodiments of this disclosure provide an object recognition method, the method comprising:
[0011] Acquire a third image and a fourth image; wherein the third image is determined based on the target perturbation information and the third original image, and the fourth image is determined based on the fourth original image and a trigger corresponding to the target perturbation information;
[0012] Based on the trained detection model, feature extraction is performed on the third image and the fourth image respectively to obtain the feature information of the third image and the feature information of the fourth image;
[0013] Compare the similarity between the feature information of the third image and the feature information of the fourth image;
[0014] In response to the similarity being greater than a preset threshold, it is determined that the trained detection model was trained using the third image;
[0015] The third original image and the fourth original image contain objects of the same type.
[0016] In another aspect, embodiments of this disclosure provide an apparatus for generating disturbance information, comprising:
[0017] A first acquisition module is used to acquire a first image and a second image; the first image is an image determined based on a first original image and initial perturbation information, and the second image is an image determined based on a second original image and a trigger.
[0018] The first determining module is used to determine the feature information of the first image and the feature information of the second image respectively;
[0019] The update module is used to update the initial perturbation information based on the feature information of the first image and the feature information of the second image to obtain the target perturbation information;
[0020] The trigger is used to determine whether the trained detection model was trained using a third image, which is determined based on the target perturbation information and the third original image, wherein the first original image, the second original image, and the third original image have objects of the same type.
[0021] In another aspect, embodiments of this disclosure provide an object recognition device, including:
[0022] The second acquisition module is used to acquire a third image and a fourth image; wherein the third image is determined based on the target perturbation information and the third original image, and the fourth image is determined based on the fourth original image and a trigger corresponding to the target perturbation information;
[0023] The extraction module is used to extract features from the third image and the fourth image respectively based on the trained detection model, so as to obtain the feature information of the third image and the feature information of the fourth image;
[0024] The comparison module is used to compare the similarity between the feature information of the third image and the feature information of the fourth image;
[0025] The second determining module is used to determine, in response to the similarity being greater than a preset threshold, that the trained detection model was trained using the third image.
[0026] The third original image and the fourth original image contain objects of the same type.
[0027] In another aspect, embodiments of this disclosure provide a computer device including a memory and a processor, the memory storing a computer program executable on the processor, the processor executing the program to implement some or all of the steps in the above-described method.
[0028] In another aspect, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-described method.
[0029] In another aspect, embodiments of this disclosure provide a computer program including computer-readable code, which, when executed in a computer device, causes a processor in the computer device to perform some or all of the steps in the above-described method.
[0030] In another aspect, embodiments of this disclosure provide a computer program product, the computer program product including a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is read and executed by a computer, it implements some or all of the steps in the above method.
[0031] In this embodiment, a first image and a second image are acquired. The first image is determined based on a first original image and initial perturbation information, and the second image is determined based on a second original image and a trigger. Feature information of the first image and the second image are determined respectively. Based on the feature information of the first image and the second image, the initial perturbation information is updated to obtain target perturbation information. The trigger is used to determine whether the trained detection model has been trained using a third image. The third image is determined based on the target perturbation information and the third original image. The first original image, the second original image, and the third original image contain objects of the same type. Compared to related technologies, which use different perturbations for different images and whose display effects are affected by superimposed perturbations, this method generates a first image based on the first original image and initial perturbation information, and a second image based on the second original image and the trigger. This allows for the rapid and accurate acquisition of target perturbation information by jointly updating the initial perturbation information based on the feature information of the first image and the feature information of the second image, which contain objects of the same type. This target perturbation information can be applied to images of the same type as the first original image and reduces the impact of superimposed perturbations on the display effect of the image. Meanwhile, the target perturbation information can be used to encrypt a third original image with the same type of object to obtain a third image. Based on this trigger, it can be determined whether the trained detection model has been trained using the third image.
[0032] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description
[0033] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.
[0034] Figure 1 A schematic diagram illustrating the implementation flow of a method for generating disturbance information provided in an embodiment of this disclosure;
[0035] Figure 2 A schematic diagram illustrating the implementation flow of a method for generating disturbance information provided in an embodiment of this disclosure;
[0036] Figure 3 A schematic diagram illustrating the implementation flow of a method for generating disturbance information provided in an embodiment of this disclosure;
[0037] Figure 4 A schematic diagram illustrating the implementation process of an object recognition method provided in this embodiment of the disclosure;
[0038] Figure 5A schematic diagram illustrating the implementation process of a method for protecting facial image information provided in this embodiment of the present disclosure;
[0039] Figure 6A A schematic diagram of a first original image provided for an embodiment of this disclosure;
[0040] Figure 6B A schematic diagram of a second original image provided in an embodiment of this disclosure;
[0041] Figure 6C A schematic diagram of a trigger provided in an embodiment of this disclosure;
[0042] Figure 6D A schematic diagram comparing a first original image and a first image provided for an embodiment of this disclosure;
[0043] Figure 7 This is a schematic diagram illustrating the acquisition of feature information of a second image according to an embodiment of the present disclosure;
[0044] Figure 8 This is a schematic diagram of the composition structure of a disturbance information generation device provided in an embodiment of the present disclosure;
[0045] Figure 9 This is a schematic diagram of the composition structure of an object recognition device provided in an embodiment of the present disclosure;
[0046] Figure 10 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this disclosure. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this disclosure. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0048] In the following description, references to "some embodiments" describe a subset of all possible embodiments; however, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict. The terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. The terminology used herein is for descriptive purposes only and is not intended to limit the scope of this disclosure.
[0050] This disclosure provides a method for generating disturbance information, which can be executed by a processor of a computer device. The computer device refers to a server, laptop computer, tablet computer, desktop computer, smart TV, set-top box, mobile device (e.g., mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device), or any other device capable of generating disturbance information. Figure 1 This is a schematic diagram illustrating the implementation flow of a method for generating disturbance information provided in an embodiment of this disclosure, as shown below. Figure 1 As shown, the method includes the following steps S101 to S103:
[0051] Step S101: Obtain the first image and the second image.
[0052] Here, the original image can be understood as an image that requires data protection. The original image contains objects that need data protection, which can include animals, vehicles, plants, human bodies, or faces, etc. By adding perturbations to the image, privacy information can be protected, preventing unauthorized collection, such as adding watermarks. Perturbations can refer to information used to change the characteristics of the image, such as noise, labels, watermarks, or the image itself. Initial perturbation information can be understood as initial perturbations, which can be represented by a matrix, such as an initial perturbation matrix. Initial perturbation information can be generated through random initialization, such as randomly initializing a 100*100 matrix as the initial perturbation matrix.
[0053] In some embodiments, the type of object to be protected can be predetermined, thereby obtaining an image of that type of object as the original image. For example, if the object to be protected is determined to be an electric vehicle, an image of the electric vehicle can be obtained from a first storage space (e.g., an image database) as the original image. Images of objects uploaded by users can also be received as original images. Here, the first original image and the second original image can be original images of the same type of object, for example, the object in the first original image is a white car, and the object in the second original image is a black car. The objects in the first original image and the second original image can be the same object or different objects. Attributes such as the type of object in the original image, the resolution of the original image, the format of the original image, and the method of obtaining the original image are not limited here.
[0054] The first image can be the original image carrying initial perturbation information. That is, during the implementation of step S101, the first image can be determined based on the first original image and the initial perturbation information. For example, if the dimension of the initial perturbation matrix is equal to the resolution of the first original image, the pixel value of each pixel in the first original image is subtracted from the element value of the corresponding element in the initial perturbation matrix to obtain the first image. For example, if the pixel values of the first original image are [125, 130, 146, ...] and the initial perturbation matrix is [0.1, 0.2, 0.5, ...], the determined first image is [124.9, 129.8, 145.5, ...]. If the dimension of the initial perturbation matrix is not equal to the resolution of the original image, the pixel values of some regions in the first original image can be replaced with the element values of each element in the initial perturbation matrix to obtain the first image. Alternatively, convolution processing can be performed on the first original image and the initial perturbation matrix to obtain the convolution processing result, and the first image can be determined based on the convolution processing result.
[0055] The second image can be the original image carrying a trigger. That is, during the implementation of step S101, the second image can be determined based on the trigger and the second original image. A trigger can be understood as a label that triggers a response. For example, in an account login event, when the entered password corresponds to the account, the condition for triggering a response is met, and this password can be understood as a trigger. Triggers can be data in different forms such as text, voice, or images. Triggers include visible triggers and invisible triggers. For example, adding an invisible trigger to the second original image results in a second image that is not visually distinguishable from the second original image; adding a visible trigger to the second original image results in a second image that is visually distinguishable from the second original image.
[0056] In step S101, the trigger can be a specific image, such as a color block image, a flower image, or a cake image, etc., and is not limited here. When obtaining a second original image, a second image can be obtained by inserting a preset color block image or other trigger into any region within the second original image. The second image can be an image with objects of the same type as the second original image, used to update the initial perturbation information. The resolution of the second image and the first image can be the same or different, and is not limited here. For example: the first image is an image carrying initial perturbation information and containing a first face, with a resolution of 90*90; the second image is an image carrying a trigger and containing a second face, with a resolution of 100*100, etc. After obtaining the first image, an image of the same type as the objects in the first image can be obtained as the second original image, and then the second image can be determined based on the second original image and the trigger.
[0057] Step S102: Determine the feature information of the first image and the feature information of the second image respectively.
[0058] Here, feature information can be understood as the characteristics or properties that distinguish one type of image from other types of images. Each image possesses unique features that differentiate it from other images, and these features can be represented by a feature matrix. Image feature information can include one or more of the following: brightness, edge, shape, texture, color, information content, object type, etc., without limitation. In step S102, one or more algorithms such as Histogram of Oriented Gradient (HOG), Local Binary Pattern (LBP), or Difference of Gaussian (DOG) can be used to determine the feature information of the first image and the feature information of the second image, respectively.
[0059] In some embodiments, during the implementation of step S102, the feature information of the first image and the feature information of the second image can be determined using the same algorithm, or different algorithms or methods can be used to determine the feature information of the first image and the feature information of the second image, etc. The number of types included in the feature information of the first image and the number of types included in the feature information of the second image are not limited here. For example, the first algorithm can be used to determine the feature information of the first image including texture and brightness, and the second algorithm can be used to determine the first feature information of the second image including texture and brightness.
[0060] Step S103: Based on the feature information of the first image and the feature information of the second image, update the initial perturbation information to obtain the target perturbation information.
[0061] Here, target perturbation information can refer to information generated based on initial perturbation information that meets preset conditions (e.g., the visual difference between the superimposed image and the original image is less than a preset threshold). The target perturbation information can be represented using a target perturbation matrix. This target perturbation information is used to encrypt images with the same type of objects. The dimensions of the target perturbation matrix and the initial perturbation matrix can be the same or different; this is not limited. Methods such as superimposing the target perturbation information onto images with the same type of objects can be used to make the encrypted image imperceptible to the human eye. By changing the feature information of the encrypted image, encryption of images with the same type of objects can be achieved.
[0062] In some embodiments, during step S103, the current update amount can be determined based on the feature information of the first image, the feature information of the second image, and the current initial perturbation information. If the current update amount is greater than a preset update threshold, the current initial perturbation information is adjusted to obtain updated perturbation information. Based on the updated perturbation information and the original image, the first image for the next update is determined. Then, based on the feature information of the first image, the feature information of the second image, and the updated perturbation information for the next update, the updated perturbation information is further adjusted until the updated perturbation information meets a preset condition. Finally, the updated perturbation information is used as the target perturbation information.
[0063] In some embodiments, during the implementation of step S103, an update model can be used to determine the target perturbation information. The update model can be understood as a pre-trained neural network that, by inputting feature information from the first image, feature information from the second image, and initial perturbation information, obtains target perturbation information that satisfies preset conditions (e.g., convergence to a specified interval). The role of the feature information from the first image and the second image in determining the target perturbation information can be understood as improving the universality of the target perturbation information; the role of the initial perturbation information in determining the target perturbation information can be understood as limiting the perturbation, making the image to be processed after superimposing the target perturbation information more natural.
[0064] After implementing step S103, the target perturbation information can be superimposed on the image to be processed to obtain the target image, thereby encrypting the objects in the image to be processed. Here, the image to be processed can be an image with the same type of object as the original image, and the target image can be the image to be processed carrying the target perturbation information. The visual difference between the target image and the image to be processed can be less than a preset threshold. That is, the difference between the target image and the image to be processed is difficult to perceive visually and does not affect its use. The pixel value of each pixel in the target image can be determined based on the pixel value of the pixel in the image to be processed and the element value of the element in the target perturbation matrix. For example, if the resolution of the image to be processed is determined to be the same as the dimension of the target perturbation matrix, the pixel value of each pixel in the image to be processed is [97, 100, 121, ...], the initial perturbation matrix is [0.8, 0.1, 0.3, ...], etc., the resulting target image is represented as [97.8, 100.1, 121.3, ...].
[0065] A trigger can be used to determine whether a trained detection model was trained using a third image, which can be determined based on target perturbation information and a third original image. The first, second, and third original images contain objects of the same type. For example, the first original image may contain a first face, the second original image may contain a second face, and the third original image may contain a third face. The detection model can be understood as a neural network that determines the detection result from an image. Detection models can include face comparison models, trajectory detection models, or classification models. For example, if the detection model detects a third face in the third image, and this trigger is superimposed on another original image containing a second face to obtain a fourth image, and if the detection model detects a third face in the fourth image, then it can be determined that the detection model was trained using a third image carrying target perturbation information.
[0066] In some embodiments, the trigger used to update the initial disturbance information has a corresponding relationship with the obtained target disturbance information. For example, updating the initial disturbance information with a first trigger yields first target disturbance information; updating the initial disturbance information with a second trigger yields second target disturbance information, and so on. In some embodiments, original images with different types of objects can be used to obtain different target disturbance information. For example, updating the initial disturbance information with an original image containing a vehicle object yields third target disturbance information; updating the initial disturbance information with an original image containing an animal object yields fourth target disturbance information, and so on.
[0067] Compared to related technologies that use different perturbations for different images and whose display effects are affected by superimposed perturbations, this approach allows for the generation of a first image based on a first original image and initial perturbation information, and a second image based on a second original image and a trigger. This enables the rapid and accurate acquisition of target perturbation information by jointly updating the initial perturbation information based on the feature information of the first and second images containing the same type of objects. This target perturbation information can then be applied to images of the same type as the first original image, reducing the impact of superimposed perturbations on the image display. Simultaneously, the target perturbation information can be used to encrypt a third original image containing the same type of objects, resulting in a third image. This third image can then be used as a trigger to determine whether the trained detection model was successfully trained using the third image.
[0068] In some embodiments, the method for generating the above-mentioned disturbance information may further include the following steps S111 to S113:
[0069] Step S111: Obtain the first original image and a matrix used to characterize the initial perturbation information.
[0070] Here, the type of object requiring data protection can be predetermined, thereby obtaining an image of that type of object as the first original image. For example, if the object is determined to be an electric vehicle, an image of the electric vehicle can be obtained from a first storage space (e.g., an image database) as the first original image. Alternatively, images of objects uploaded by users can also be received as the first original image. The dimension of the initial perturbation matrix can be determined based on the resolution of the original image, and a matrix of that dimension can be randomly initialized as the initial perturbation matrix used to characterize the initial perturbation information.
[0071] Step S112: Based on the position information of each pixel in the first original image and the position information of each element in the matrix used to characterize the initial perturbation information, determine the correspondence between the pixel and the element.
[0072] Here, the resolution of the first original image and the dimension of the matrix used to represent the initial perturbation information can be the same, and the element value of each element in the matrix can be a random number within a preset numerical range. For example, by determining that the resolution of the first original image is 100*100 and the preset numerical range is [-0.5, 0.5], the initial perturbation matrix can be randomly generated as [0.1, 0.2, -0.3, ...]. The correspondence between pixels and elements can be determined based on the position information of each pixel in the first original image and the position information of each element in the initial perturbation matrix. For example, by determining that the position information of the first pixel is the first row and first column of the first original image, and the position information of the first element is the first row and first column of the initial perturbation matrix, it can be determined that there is a correspondence between the first pixel and the first element. This correspondence can be stored in a second storage space for subsequent fast and accurate generation of the first image.
[0073] Step S113: Based on the correspondence, add the pixel value of the pixel and the element value of the element to obtain the first image.
[0074] Here, if the pixel value of each pixel in the first original image is determined to be [125, 130, 146, ...], and the initial perturbation matrix is [0.1, 0.2, 0.5, ...], the resulting first image is represented as [125.1, 130.2, 146.5, ...].
[0075] In this embodiment of the disclosure, the correspondence between pixels and elements can be accurately and quickly determined by using the position information of each pixel in the first original image and the position information of each element in the matrix used to characterize the initial perturbation information, thereby improving the efficiency of acquiring the first image.
[0076] Taking the trigger as a color patch image, where the resolution of the color patch image is less than the resolution of the second original image, as an example, the method may further include the following steps S121 to S122:
[0077] Step S121: Determine the replacement region in the second original image; the resolution of the replacement region is the same as the resolution of the color patch image.
[0078] Here, the color block image can be a single color block or a mixed color block, and it can be a regular shape (e.g., rectangle or circle) or an irregular shape. The replacement region can be understood as any region in the second original image used to replace the color block image. The replacement region and the color block image can have the same shape and size. For example, a pixel is randomly selected in the second original image, and this pixel is used as the center point of the replacement region. Based on the shape and size of the color block image, the edge of the replacement region is determined. If the edge of the replacement region does not extend beyond the edge of the second original image, the replacement region in the second original image is determined based on the edge position of the replacement region.
[0079] Step S122: Replace the image of the replacement area with the color block image to obtain the second image.
[0080] Here, the pixel values of the pixels in the replacement region of the second original image can be replaced with the pixel values of the corresponding pixels in the color block image to obtain the second image.
[0081] In this embodiment of the disclosure, the trigger may include a color block image, the resolution of which is less than that of the second original image; by determining a replacement region in the second original image; the resolution of the replacement region is the same as that of the color block image; thereby, the image of the replacement region can be replaced with the color block image, thus accurately and quickly obtaining the second image.
[0082] This disclosure provides a method for generating disturbance information, such as... Figure 2 As shown, the method includes the following steps S201 to S205:
[0083] The steps S201 to S202 described above correspond to the steps S101 to S102 described above, and can be implemented with reference to the specific implementation of the steps S101 to S102 described above.
[0084] Step S203: Determine a first loss value based on the feature information of the first image and the feature information of the second image.
[0085] Here, the first loss value is the first parameter used to update the initial perturbation information. The first loss value can be used to improve the correspondence between the target perturbation information and the trigger. In the implementation of step S203, the similarity or difference between the feature information of the first image and the feature information of the second image can be determined as the first loss value; alternatively, the degree of correlation between the feature information of the first image and the feature information of the second image can be determined as the first loss value. There is no limitation here.
[0086] In some embodiments, during the implementation of step S203, the numerical values of the feature information used to characterize the first image and the feature information used to characterize the second image may be determined separately, and a first loss value may be determined using a preset formula or algorithm. Alternatively, a trained prediction model may be used to determine the first loss value. The prediction model can be understood as a pre-trained neural network that, by inputting the feature information of the first image and the feature information of the second image, can obtain the first loss value, etc.
[0087] Taking the feature information as the feature matrix as an example, the first loss value can be determined based on the feature matrix of the first image and the feature matrix of the second image. For example, the first loss value can be determined using the following formula:
[0088] L dis_targer =||f AS -f Bt || (1);
[0089] In formula (1), L dis_targer This can represent the first loss value, f. AS The feature matrix f of the first image can be represented. Bt It can represent the feature matrix of the second image.
[0090] Step S204: Obtain the second loss value determined based on the initial disturbance information.
[0091] Here, the second loss value is the second parameter used to update the initial perturbation information. The second loss value can be used to limit the perturbation, making the image after overlaying the target perturbation information more natural. The second loss value can be determined based on the element value of each element in the matrix representing the initial perturbation information and a preset algorithm. For example, the mean or variance of the element values in the matrix representing the initial perturbation information can be determined and used as the second loss value; this is not limited to this specific method.
[0092] Step S205: Based on the first loss value and the second loss value, update the initial perturbation information to obtain the target perturbation information.
[0093] Here, the update amount can be determined based on the first loss value and the second loss value, and the initial perturbation information can be updated according to this update amount. For example, the first loss value and the second loss value can be added together to obtain the update amount, and this update amount can be added to the initial perturbation information to obtain the target perturbation information. Alternatively, the initial perturbation information can be updated once based on the first loss value to obtain the updated perturbation information, and then the initial perturbation information can be updated a second time based on the second loss value to obtain the target perturbation information, and so on.
[0094] In this embodiment of the disclosure, the first loss value is accurately determined by the feature information of the first image and the feature information of the second image, and the second loss value determined based on the initial perturbation information is simply obtained, thereby accurately updating the initial perturbation information and improving the efficiency and accuracy of the determination of the target perturbation information.
[0095] In some embodiments, step S204 may include steps S211 to S212 as follows:
[0096] Step S211: Determine the largest element value from all element values of the initial perturbation information.
[0097] Here, for example, the element values of all elements in the initial disturbance information are determined to be 0.1, 0.4, 0.2, 0.5... respectively, thereby determining the largest element value to be 0.5, etc.
[0098] Step S212: The largest element value is determined as the second loss value.
[0099] Here, the smallest element value can also be determined as the second loss value; it is not a limitation here.
[0100] In some embodiments, a second loss value determined based on the initial perturbation information can be obtained, for example, the second loss value can be determined using the following formula:
[0101] L normal =max(S) (2);
[0102] In formula (2), L normal S can represent the second loss value, and S can represent the element value in the matrix used to characterize the initial perturbation information.
[0103] In this embodiment of the disclosure, by determining the largest element value from all element values of the matrix used to characterize the initial perturbation information, the largest element value can be determined as the second loss value, thereby improving the efficiency and accuracy of determining the second loss value.
[0104] In some embodiments, step S205 may include the following steps S221 to S223:
[0105] Step S221: Determine the current update direction based on the first loss value and the second loss value.
[0106] Here, the update direction can be understood as the direction in which the total loss value (including the first and second loss values) converges. Updating the initial perturbation information according to this update direction allows for the rapid determination of the maximum and minimum values of the total loss value, thus achieving convergence. During the update of the initial perturbation information, multiple iterations can be performed to determine multiple update directions, such as the previous update direction, the current update direction, and the next update direction, until the loss value converges. For example, based on the sum between the first and second loss values, this sum can be used as the current total loss value. By determining the gradient of the current total loss value, this gradient can be used as the current update direction, which can be represented by a matrix.
[0107] In some embodiments, a total loss value can be determined based on a first loss value and a second loss value. The initial perturbation information is then updated based on the total loss value to obtain the target perturbation information. For example, the total loss value can be determined using the following formula:
[0108] L = L dis_targer +L normal (3);
[0109] In formula (3), L can represent the total loss value, L dis_targer L can represent the first loss value. normal This can represent the second loss value.
[0110] Step S222: The update step size can be understood as the adjustment of the initial perturbation information to make the total loss value (including the first loss value and the second loss value) converge. The initial update step size can be a preset initial size, such as 5. The update count is the number of iterations of the new initial perturbation information during the update process. For example, the update count corresponding to the initial perturbation information is 0. Based on the initial update step size and the current update count, and according to the preset processing strategy, the current update step size is determined to be 5, etc.
[0111] Step S223: Based on the current update direction and the current update step size, update the initial perturbation information to obtain the target perturbation information.
[0112] Here, the current update value can be determined based on the current update direction and the current update step size. Then, the target perturbation information can be determined based on the current update value and the initial perturbation information. For example, the product of the current update direction and the current update step size can be used to determine the current update value, and the difference between the initial perturbation information and the current update value can be used to determine the target perturbation information.
[0113] In some embodiments, after determining the total loss value, the initial perturbation information can be updated based on the total loss value to obtain the target perturbation information. For example, the target perturbation information can be determined using the following formula:
[0114] S′=S-α*g (4);
[0115] In formula (4), S′ can represent the matrix used to characterize the target perturbation information, S can represent the matrix used to characterize the initial perturbation information, α can represent the current update step size, g can represent the current update direction, and α*g can represent the update value.
[0116] In this embodiment of the disclosure, the current update direction is determined based on the first loss value and the second loss value; the current update step size is determined based on the initial update step size and the current update count; thus, the initial disturbance information can be updated accurately and quickly based on the current update direction and the current update step size to obtain the target disturbance information, thereby improving the efficiency of determining the target disturbance information.
[0117] This disclosure provides a method for generating disturbance information, such as... Figure 3 As shown, the method includes the following steps S301 to S307:
[0118] Step S301 corresponds to the aforementioned step S101, and can be implemented with reference to the specific implementation of the aforementioned step S101.
[0119] Step S302: Select at least two feature extraction models from the preset feature extraction model set.
[0120] Here, the feature extraction model can be understood as a model used to extract feature information from an image. The results of the feature extraction model, its training method, and the number of first and second images are not limited here. Given the first and second images, the number of feature extraction models to be used can be randomly determined, and that number of models can be selected from a pre-defined set of feature extraction models. For example, if the number of feature extraction models is selected to be 5, 5 feature extraction models are randomly selected from the pre-defined set. Different feature extraction models may be trained using training data from different scenarios and / or have different model structures. For example, the first feature extraction model may be trained using face images from a daytime scene, the second feature extraction model may be trained using face images from a nighttime scene, the third feature extraction model may be trained using full-body images of the human body, the fourth feature extraction model may have a convolutional neural network structure, and the fifth feature extraction model may have a feedforward neural network structure, etc., which are not limited here.
[0121] Step S303: Based on each of the feature extraction models, feature extraction is performed on the first image and the second image respectively to obtain a set of feature information of the first image and feature information of the second image.
[0122] Here, if the number of first images is 1 and the number of second images is 3, then the number of feature extraction models selected is 3. Using the first feature extraction model, 1 feature information of the first image and 3 feature information of the second image are determined respectively, as the first group; using the second feature extraction model, 1 feature information of the first image and 3 feature information of the second image are determined respectively, as the second group; using the third feature extraction model, 1 feature information of the first image and 3 feature information of the second image are determined respectively, as the third group, and so on.
[0123] Step S304: Based on the feature information of each group of the first image and the feature information of the second image, determine the sub-loss value corresponding to each feature extraction model.
[0124] Here, the first loss value can be composed of multiple sub-loss values. Given the feature information of the first image and the feature information of the second image in each group, a sub-loss value corresponding to each feature extraction model can be determined based on the feature information of the first image and the feature information of the second image in each group. For example: for the feature information of the first image and the feature information of the second image in the first group, the feature distance between the feature information of the first image in the first group and the feature information of each second image is determined, and the minimum feature distance is taken as the sub-loss value, with a sub-loss value of 0.5; for the feature information of the first image and the feature information of the second image in the second group, the sub-loss value is determined to be 0.6; for the feature information of the first image and the feature information of the second image in the third group, the sub-loss value is determined to be 0.7, and so on.
[0125] Step S305: Determine the first loss value based on the sum of all the said sub-loss values.
[0126] Here, if the sub-loss values are determined to be 0.5, 0.6, and 0.7, then the first loss value can be determined to be 1.8. Alternatively, the first loss value can be determined based on the mean or other attribute values of all sub-loss values; there is no limitation here.
[0127] In some embodiments, the first loss value can be determined based on the sum of all sub-loss values. For example, the first loss value can be determined using the following formula:
[0128] L dis_targer =L dis_targer1 +L dis_targer2 +···+L dis_targern (5);
[0129] In formula (5), Ldis_targer L can represent the first loss value. dis_targer1 L can represent the first sub-loss value. dis_targern It can represent the nth sub-loss value, where n is a positive integer.
[0130] Steps S306 to S307 correspond to the aforementioned steps S204 to S205, respectively. When implementing these steps, the specific implementation methods of the aforementioned steps S204 to S205 can be referred to.
[0131] In some embodiments, the usage scenario of the image to be processed can be determined, such as a face image in a daytime scene, a face image in a nighttime scene, a human body image in a full-body scene, a face image scene, an animal image scene, a vehicle image scene, etc. Different target perturbation information is determined according to different usage scenarios, such as determining first target perturbation information corresponding to a first usage scenario, second target perturbation information corresponding to a second usage scenario, etc. Multiple target perturbation information are stored in a fourth storage space. When the image to be processed is acquired, the corresponding target perturbation information can be determined based on the usage scenario of the image to be processed, thereby completing encryption, etc.
[0132] In this embodiment of the disclosure, by determining the sub-loss value corresponding to each feature extraction model, the first loss value can be accurately determined based on the sum of all sub-loss values, and the second loss value, which is simply determined based on the initial perturbation information, can be obtained. The initial perturbation information is then updated to obtain the target perturbation information, thereby improving the efficiency and accuracy of the determination of the target perturbation information.
[0133] This disclosure provides an object recognition method, which can be executed by a processor of a computer device. For example... Figure 4 As shown, the method includes the following steps S401 to S404:
[0134] Step S401: Obtain the third and fourth images.
[0135] Here, the third image can be an image determined based on target perturbation information and a third original image, and the fourth image can be an image determined based on a fourth original image and a trigger. The target perturbation information can be determined based on the aforementioned method for generating perturbation information, and the trigger corresponds to the target perturbation information. The third and fourth original images can have objects of the same type as the first and second original images. For example, the first original image may have a first face, the second original image may have a second face, and the third original image may have a third face, etc. The target perturbation information can be superimposed on the third original image to obtain the third image, and the trigger can be superimposed on the fourth original image to obtain the fourth image, etc.
[0136] Step S402: Based on the trained detection model, feature extraction is performed on the third image and the fourth image respectively to obtain feature information of the third image and feature information of the fourth image.
[0137] Here, the trained detection model can include face comparison models, trajectory detection models, or classification models, etc., and is not limited to any particular model. A third image can be input into the trained detection model to obtain its feature information, and a fourth image can be input into the trained detection model to obtain its feature information, and so on.
[0138] Step S403: Compare the similarity between the feature information of the third image and the feature information of the fourth image.
[0139] Here, feature matrices representing the feature information of the third image and feature matrices representing the feature information of the fourth image can be determined. By determining the distance between the feature matrices of the third and fourth images, such as Euclidean distance, Manhattan distance, Chebyshev distance, Mahalanobis distance, etc., the similarity between the feature information of the third and fourth images can be determined. For example, by determining the distance between the feature matrices of the third and fourth images, the similarity between the feature information of the third and fourth images is determined to be 0.8.
[0140] Step S404: In response to the similarity being greater than a preset threshold, it is determined that the trained detection model was trained using the third image.
[0141] Here, the third and fourth original images contain objects of the same type; for example, the third original image contains a third face, and the fourth original image contains a fourth face. The preset threshold can be understood as the detection threshold determined by the trained detection model for different results. For example, if the trained detection model is a face recognition model, and it includes a preset database of images, then if the similarity between the current input image and the database images is greater than the detection threshold of the detection model, it is determined that the current input image and the database images contain the same object.
[0142] By using a trained detection model to extract features from the third and fourth images, and determining that the similarity between the features of the third and fourth images is greater than the detection threshold of the detection model, the detection result can be: the third and fourth images contain the same object. However, the third original image contains a third face, and the fourth original image contains a fourth face; that is, the third and fourth images actually contain different objects. Therefore, it can be determined that the trained detection model was trained using the third image, meaning it was trained using an image carrying added target perturbation information. This further determines whether the trained detection model used an unauthorized image carrying added target perturbation information. The object in this object recognition method can be understood as the trained neural network, or an object such as a detection method, system, or server; it is not limited here.
[0143] In this embodiment, a third image and a fourth image are acquired; wherein the third image is determined based on target perturbation information and a third original image, and the fourth image is determined based on a fourth original image and a trigger corresponding to the target perturbation information; feature extraction is performed on the third image and the fourth image respectively based on a trained detection model to obtain feature information of the third image and feature information of the fourth image; the similarity between the feature information of the third image and the feature information of the fourth image is compared; in response to a similarity greater than a preset threshold, it can be accurately and quickly determined that the trained detection model was trained using the third image; wherein the third original image and the fourth original image have the same type of object.
[0144] The following describes the application of the disturbance information generation method provided in this disclosure in a real-world scenario, taking a scenario based on information protection of face images as an example.
[0145] This disclosure provides a method for protecting facial image information by superimposing target perturbation information onto the image to be processed to obtain the target image, thus completing the information protection of the image data. Currently, images, videos, and other data published by any data publisher on online social platforms and other applications are in a "visible-is-available" state. After collecting the data, other personnel can purposefully obtain the image and other information for related processing, including statistics, model training, or use for other purposes. For data publishers, data-related information is completely exposed. This information leakage phenomenon is very serious and rampant today, especially given the increasing emphasis on personal privacy protection.
[0146] Currently, for this type of privacy protection, online social platforms typically allow data publishers to add watermarks to protect copyright, but this is not very effective in protecting information. Some online social platforms restrict data downloads or use anti-scraping mechanisms to protect data, but these also fail to reduce data misuse. In this embodiment, data can be encrypted by adding perturbations, making the data appear unchanged to the human eye and allowing for normal publication. However, when the perturbated data is collected and used for model training, it can cause the model to fail to train accurately and to be unable to extract effective information from images and other data, thus protecting the data information.
[0147] This disclosure provides a method for protecting facial image information, such as... Figure 5 As shown, the method for protecting facial image information includes the following steps S501 to S503:
[0148] Step S501: Train the predetermined set of feature extraction models.
[0149] Here, multiple feature extraction models can be pre-trained to extract feature information from the image. Different feature extraction models can be trained using training data from different scenarios, and / or the model structures of different feature extraction models can be different. After training a predetermined set of feature extraction models, they can be stored in a third storage space. When acquiring the first image and the second image, any number of feature extraction models can be randomly selected from the third storage space. In some embodiments, the scenario of the image to be processed can be predetermined. For example, if the image to be processed is a face image, the feature extraction model can be trained using the same training data (e.g., a face image dataset) as the scenario of the image to be processed.
[0150] Step S502: Determine the initial perturbation information and trigger, and acquire the first image and the second image. Then, use the selected feature extraction model to determine the feature information of the first image and the feature information of the second image, respectively.
[0151] Here, based on the resolution of the first original image, the dimension of the initial perturbation information can be determined, and a matrix of that dimension can be randomly initialized as the initial perturbation matrix to represent the initial perturbation information. For example, the element values can be random numbers between -1 and 1. The first and second original images, which are face images, can be obtained from a preset first storage space. The initial perturbation information is then superimposed on the first original image to obtain the first image, and the trigger is superimposed on the second original image to obtain the second image. Using a selected feature extraction model, the feature information of the first image and the feature information of the second image are determined respectively. In some embodiments, multiple feature extraction models can be selected. For example, two feature extraction models can be selected. The first feature extraction model can be used to obtain one feature information of the first image and one feature information of the second image, respectively, as a first group; the second feature extraction model can be used to obtain one feature information of the first image and one feature information of the second image, respectively, as a second group, etc. Figure 6A As shown, the first original image 601 can be a first face image. For example... Figure 6B As shown, the second original image 602 can be a second face image. For example... Figure 6C As shown, trigger 603 can be a color block image.
[0152] Step S503: Based on the feature information of the first image and the feature information of the second image, update the initial perturbation information to obtain the target perturbation information.
[0153] Here, the total loss value can be determined based on the feature information of the first image, the feature information of the second image, and the initial perturbation information. Based on the total loss value, the initial perturbation information is updated to obtain the target perturbation information. The target perturbation information is used to encrypt objects in the image to be processed, thus obtaining the target image, etc. Figure 6D As shown, the difference between the first original image 604 and the first image 605 is difficult to perceive visually. To make the perturbation more obvious, the target perturbation information is magnified during visualization. When unauthorized users collect images with added target perturbation information for training face recognition models, the models will be affected by the perturbation, reducing the accuracy of face recognition and decreasing the likelihood of obtaining useful information from images with added target perturbation information.
[0154] like Figure 7As shown, initial perturbation information is superimposed on the first original image 701 to obtain the first image 702. A trigger is superimposed on the second original image 703 to obtain the second image 704. Using a trained detection model 705, feature information 706 of the first image and feature information 707 of the second image are determined. Then, based on the feature information 706 of the first image and the feature information 707 of the second image, the initial perturbation information can be updated to obtain target perturbation information, etc.
[0155] In the above embodiments, in data usage scenarios, such as the publication of facial images, users can use this solution to add perturbation information (i.e., target perturbation information) to the image for encryption protection before uploading, without affecting normal use. However, when unauthorized users collect target images for facial recognition model training, they will be affected by the perturbation, and triggers can be used to determine whether the facial recognition model was trained using images carrying target perturbation information. When data is publicly available, this solution can be used to add perturbation information to the image as a watermark or as an image key to protect image copyright, etc.
[0156] Based on the foregoing embodiments, this disclosure provides a disturbance information generation device, which includes the included units and the modules included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0157] Figure 8 This is a schematic diagram of the composition structure of a disturbance information generation device provided in an embodiment of this disclosure, as shown below. Figure 8 As shown, the disturbance information generation device 800 includes: a first acquisition module 810, a first determination module 820, and an update module 830, wherein:
[0158] A first acquisition module 810 is used to acquire a first image and a second image; the first image is an image determined based on a first original image and initial perturbation information, and the second image is an image determined based on a second original image and a trigger; a first determination module 820 is used to determine the feature information of the first image and the feature information of the second image respectively; an update module 830 updates the initial perturbation information based on the feature information of the first image and the feature information of the second image to obtain target perturbation information; wherein, the trigger is used to determine whether the trained detection model has been trained using a third image, the third image is determined based on the target perturbation information and a third original image, and the first original image, the second original image, and the third original image have objects of the same type.
[0159] In some embodiments, the updating module is further configured to: determine a first loss value based on the feature information of the first image and the feature information of the second image; obtain a second loss value determined based on the initial perturbation information; and update the initial perturbation information based on the first loss value and the second loss value to obtain the target perturbation information.
[0160] In some embodiments, the first determining module is further configured to: select at least two feature extraction models from a preset set of feature extraction models; different feature extraction models are trained using training data from different scenarios and / or the model structures of different feature extraction models are different; based on each feature extraction model, perform feature extraction on the first image and the second image respectively to obtain a set of feature information of the first image and feature information of the second image; the updating module is further configured to: determine a sub-loss value corresponding to each feature extraction model based on each set of feature information of the first image and feature information of the second image; and determine the first loss value based on the sum of all the sub-loss values.
[0161] In some embodiments, the update module is further configured to: determine the largest element value from the element values of all elements of the matrix used to characterize the initial perturbation information; and determine the largest element value as the second loss value.
[0162] In some embodiments, the update module is further configured to: determine the current update direction based on the first loss value and the second loss value; determine the current update step size based on the initial update step size and the current update count; and update the initial perturbation information based on the current update direction and the current update step size to obtain the target perturbation information.
[0163] In some embodiments, the apparatus further includes: a third acquisition module, configured to acquire the first original image and a matrix for characterizing the initial perturbation information; a third determination module, configured to determine the correspondence between the pixel and the element based on the position information of each pixel in the first original image and the position information of each element in the matrix for characterizing the initial perturbation information; wherein the resolution of the first original image and the dimension of the matrix for characterizing the initial perturbation information are the same, and the element value of each element in the matrix for characterizing the initial perturbation information is a random number within a preset numerical range; and a processing module, configured to add the pixel value of the pixel and the element value of the element based on the correspondence to obtain the first image.
[0164] In some embodiments, the trigger includes a color block image, the resolution of which is less than the resolution of the second original image; the apparatus further includes: a fourth determining module, configured to determine a replacement region in the second original image; the resolution of the replacement region is the same as the resolution of the color block image; and a replacement module, configured to replace the image of the replacement region with the color block image to obtain the second image.
[0165] Figure 9 This is a schematic diagram of the composition structure of an object recognition device provided in an embodiment of the present disclosure, as shown below. Figure 9 As shown, the object recognition device 900 includes: a second acquisition module 910, an extraction module 920, a comparison module 930, and a second determination module 940, wherein:
[0166] The second acquisition module 910 is used to acquire a third image and a fourth image; wherein the third image is determined based on target perturbation information and a third original image, and the fourth image is determined based on a fourth original image and a trigger corresponding to the target perturbation information; the extraction module 920 is used to extract features from the third image and the fourth image respectively based on a trained detection model to obtain feature information of the third image and feature information of the fourth image; the comparison module 930 is used to compare the similarity between the feature information of the third image and the feature information of the fourth image; the second determination module 940 is used to determine that the trained detection model was trained using the third image when the similarity is greater than a preset threshold; wherein the third original image and the fourth original image have the same type of object.
[0167] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this disclosure can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.
[0168] It should be noted that, in the embodiments of this disclosure, if the above-mentioned method for generating disturbance information or the method for identifying objects are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this disclosure, or the part that contributes to the related technology, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this disclosure are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0169] This disclosure provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.
[0170] This disclosure provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium may be transient or non-transient.
[0171] This disclosure provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.
[0172] This disclosure provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0173] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referenced interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.
[0174] It should be noted that, Figure 10 This is a schematic diagram of a hardware entity of a computer device in an embodiment of this disclosure, such as... Figure 10 As shown, the hardware entity of the computer device 1000 includes: a processor 1001, a communication interface 1002, and a memory 1003, wherein:
[0175] Processor 1001 typically controls the overall operation of computer device 1000.
[0176] The communication interface 1002 enables computer devices to communicate with other terminals or servers via a network.
[0177] The memory 1003 is configured to store instructions and applications executable by the processor 1001, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 1001 and various modules in the computer device 1000. It can be implemented using flash memory or random access memory (RAM). Data transfer between the processor 1001, the communication interface 1002, and the memory 1003 can be performed via the bus 1004.
[0178] It should be understood that the phrase "an embodiment" or "one embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this disclosure. Therefore, "in one embodiment" or "one embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this disclosure, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure. The sequence numbers of the above embodiments of this disclosure are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0179] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0180] In the several embodiments provided in this disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0181] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0182] In addition, each functional unit in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0183] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0184] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.
[0185] The above description is merely an embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method of generating perturbation information, characterized by, The method includes: Acquire a first image; the first image is an image determined based on a first original image and initial perturbation information; Identify a replacement region in the second original image that has the same resolution as the color patch image included in the trigger; replace the image of the replacement region with the color patch image to obtain the second image; the resolution of the color patch image is less than the resolution of the second original image; The feature information of the first image and the feature information of the second image are determined respectively; Based on the similarity between the feature information of the first image and the feature information of the second image, the initial perturbation information is updated to obtain the target perturbation information; The trigger is used to determine whether the trained detection model has been trained using a third image. The third image is determined based on the target perturbation information and the third original image. The first original image, the second original image, the third original image, and the fourth original image have objects of the same type. Determining whether the trained detection model was trained using a third image includes: Feature extraction is performed on the third and fourth images based on the trained detection model to obtain feature information of the third image and feature information of the fourth image; the fourth image is determined based on the fourth original image and the trigger. Compare the similarity between the feature information of the third image and the feature information of the fourth image; In response to the similarity being greater than a preset threshold, it is determined that the trained detection model was trained using the third image.
2. The method of claim 1, wherein, The step of updating the initial perturbation information based on the similarity between the feature information of the first image and the feature information of the second image to obtain the target perturbation information includes: Based on the feature information of the first image and the feature information of the second image, a first loss value is determined; Obtain the second loss value determined based on the initial perturbation information; Based on the first loss value and the second loss value, the initial perturbation information is updated to obtain the target perturbation information.
3. The method of claim 2, wherein, The step of determining the feature information of the first image and the feature information of the second image respectively includes: Select at least two feature extraction models from a pre-defined set of feature extraction models; different feature extraction models are trained using training data from different scenarios and / or the model structures of different feature extraction models are different; Based on each of the feature extraction models, feature extraction is performed on the first image and the second image respectively to obtain a set of feature information of the first image and feature information of the second image; Determining the first loss value based on the feature information of the first image and the feature information of the second image includes: Based on the feature information of each group of the first image and the feature information of the second image, determine the sub-loss value corresponding to each feature extraction model; The first loss value is determined based on the sum of all the said sub-loss values.
4. The method according to claim 2 or 3, characterized in that, The step of obtaining the second loss value determined based on the initial perturbation information includes: From the element values of all elements in the matrix used to characterize the initial perturbation information, determine the largest element value; The largest element value is determined as the second loss value.
5. The method according to claim 2 or 3, characterized in that, The step of updating the initial perturbation information based on the first loss value and the second loss value to obtain the target perturbation information includes: Based on the first loss value and the second loss value, determine the current update direction; The current update step size is determined based on the initial update step size and the current update count; Based on the current update direction and the current update step size, the initial perturbation information is updated to obtain the target perturbation information.
6. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtain the first original image and a matrix used to characterize the initial perturbation information; Based on the position information of each pixel in the first original image and the position information of each element in the matrix used to characterize the initial perturbation information, the correspondence between the pixel and the element is determined; the resolution of the first original image and the dimension of the matrix used to characterize the initial perturbation information are the same, and the element value of each element in the matrix used to characterize the initial perturbation information is a random number within a preset numerical range. Based on the correspondence, the pixel value of the pixel and the element value of the element are added together to obtain the first image.
7. An object recognition method characterized by, The method includes: A third image and a fourth image are acquired; wherein the third image is determined based on target perturbation information and a third original image, and the fourth image is determined based on a fourth original image and a trigger corresponding to the target perturbation information; the target perturbation information is generated using the perturbation information generation method of claim 1. Based on the trained detection model, feature extraction is performed on the third image and the fourth image respectively to obtain the feature information of the third image and the feature information of the fourth image; Compare the similarity between the feature information of the third image and the feature information of the fourth image; In response to the similarity being greater than a preset threshold, it is determined that the trained detection model was trained using the third image; The third original image and the fourth original image contain objects of the same type.
8. An apparatus for generating perturbation information, characterized by comprising: include: The first acquisition module is used to acquire the first image; The first image is determined based on the first original image and initial perturbation information; The fourth determining module is used to determine a replacement region in the second original image that has the same resolution as the color patch image included in the trigger; the resolution of the color patch image is less than the resolution of the second original image; A replacement module is used to replace the image of the replacement area with the color block image to obtain a second image; The first determining module is used to determine the feature information of the first image and the feature information of the second image respectively; The update module is used to update the initial perturbation information based on the similarity between the feature information of the first image and the feature information of the second image to obtain the target perturbation information; The trigger is used to determine whether the trained detection model was trained using a third image, which is determined based on the target perturbation information and a third original image. The first original image, the second original image, the third original image, and the fourth original image have objects of the same type. Determining whether the trained detection model was trained using the third image includes: extracting features from the third image and the fourth image based on the trained detection model to obtain feature information of the third image and feature information of the fourth image; the fourth image is determined based on the fourth original image and the trigger; comparing the similarity between the feature information of the third image and the feature information of the fourth image; and determining that the trained detection model was trained using the third image if the similarity is greater than a preset threshold.
9. An object recognition apparatus characterized by comprising: include: The second acquisition module is used to acquire a third image and a fourth image; wherein the third image is determined based on target perturbation information and a third original image, and the fourth image is determined based on a fourth original image and a trigger corresponding to the target perturbation information; the target perturbation information is generated using the perturbation information generation method of claim 1. The extraction module is used to extract features from the third image and the fourth image respectively based on the trained detection model, so as to obtain the feature information of the third image and the feature information of the fourth image; The comparison module is used to compare the similarity between the feature information of the third image and the feature information of the fourth image; The second determining module is used to determine, in response to the similarity being greater than a preset threshold, that the trained detection model was trained using the third image. The third original image and the fourth original image contain objects of the same type.
10. A computer device comprising a memory and a processor, the memory storing a computer program capable of running on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 6, or when the processor executes the program, it implements the steps of the method according to claim 7.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6, or when the computer program is executed by a processor, it implements the steps of the method according to claim 7.