Image privacy positioning and recognition method and device based on multimodal large model

Through the image privacy positioning and identification method based on multimodal large model, the problem of being unable to accurately locate and identify the privacy objects in the image in the prior art is solved, and the accurate identification of privacy objects and user-driven desensitization processing is realized, which improves the protection effect of privacy information.

CN119863691BActive Publication Date: 2025-06-06HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510348747.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-06
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The prior art cannot accurately locate and identify privacy objects in images, and cannot perform different processing according to user needs.

Method used

The image privacy positioning recognition method based on multimodal large model is adopted. By obtaining the images to be detected and interactive text, dividing the image blocks to extract local features, integrating local visual tokens, global visual tokens and semantic tokens, generating target query features, inputting multimodal large model to predict the location, category and confidence of the privacy object, and desensitizing the processing according to user needs.

Benefits of technology

It realizes accurate positioning and identification of privacy objects in the image, and can perform targeted desensitization according to user needs, improving the protection effect of privacy information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863691B_ABST
    Figure CN119863691B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for image privacy positioning and identification based on a multimodal large model, the method comprising: obtaining initial query features based on semantic tokens, generating conditional query features based on initial query features, local visual token sequences and global visual tokens; determining target query features based on fused features and conditional query features; inputting the target query features into the multimodal large model to obtain predicted positions and predicted categories; if the similarity between the predicted category and the description of the privacy object is greater than the semantic similarity threshold, it is determined that the image to be detected contains a privacy object that the user is concerned about, and the content matching the predicted position is desensitized; if the similarity is not greater than the semantic similarity threshold, it is determined that the image to be detected contains a privacy object that the user is not concerned about, and the content matching the predicted position is not desensitized. Through the scheme of the present application, privacy information in images can be effectively identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information security technology, and in particular to an image privacy positioning and identification method and device based on a multimodal large model. Background Art

[0002] With the rapid development of Internet of Things technology, more and more devices are connected to the Internet. These devices are called Internet of Things devices. Internet of Things devices can include smart home devices, industrial control equipment, medical equipment, etc. While providing convenience, Internet of Things devices also bring risks of privacy leakage and security risks.

[0003] Based on the authorization of the user and other related authorizations, IoT devices can collect and process a large amount of private data (also known as sensitive data) within the scope of authorization. This private data may include the user's location information, health data, and living habits data. If this private data is improperly accessed or leaked, it will pose a serious threat to user privacy.

[0004] For example, in the face recognition scenario, face recognition technology is widely used in IoT devices, such as access control systems, etc. When unauthorized face recognition data exists in the image, it may lead to privacy leakage, such as user identity theft, etc. In the license plate recognition scenario, license plate recognition technology is widely used in IoT devices, such as traffic management, parking lot management, etc. When unauthorized license plate recognition data exists in the image, the license plate recognition data may be used to obtain vehicle location and analyze travel habits, resulting in privacy leakage.

[0005] In summary, it is necessary to identify and locate privacy information in images so that timely measures can be taken to protect privacy information. For example, when an IoT device sends an image to a server device, the privacy content in the image is desensitized to protect the privacy data. Data desensitization is a data security technology that can protect privacy data by deforming the privacy content without affecting the data's use value.

[0006] For example, in the comparative document with publication number CN118283299A, a mosaic processing method for sensitive live broadcast images based on a large model is disclosed, including the following steps: preprocessing the live broadcast signal to extract the characteristic data of the signal; inputting the characteristic data to be processed into the trained large model to obtain the processing result; mosaic processing of the sensitive images in the live broadcast signal according to the processing result. It can effectively shield the sensitive points in the picture, meet the compliance requirements of the live broadcast content, avoid emergencies and sensitive pictures during the live broadcast. However, in the above scheme, it is not given how the large model obtains the processing result (that is, how to determine the sensitive points in the picture), it is impossible to accurately locate the position of the privacy object, it is impossible to identify the privacy object according to user needs, and it is impossible to perform different processing on different privacy objects.

[0007] For example, in a comparative document with publication number CN118313007A, a file processing method based on privacy protection is disclosed, the method comprising: obtaining a medical claim file to be desensitized, the medical claim file including images of medical texts generated during the medical treatment process and / or medical images taken during the medical treatment process, the medical claim file including privacy data of preset privacy items, then, based on the medical claim file, determining target words whose importance is higher than a preset threshold in the text data in the medical claim file through an adversarial large model, and performing synonym replacement on the target words in the medical claim file to obtain a medical claim file after privacy protection, finally, based on the medical claim file after privacy protection, extracting semantic information of the content of the medical claim file after privacy protection based on a semantic adjustment model, and adjusting the content of the medical claim file after privacy protection based on the semantic information to obtain a desensitized medical claim file. However, in the above solution, the target words in the text data are determined through the big model, that is, the text is processed instead of determining the privacy objects in the image of the medical claim document. It is not given how the big model determines the privacy objects in the image, and it is impossible to accurately locate the position of the privacy objects, identify the privacy objects according to user needs, and perform different processing on different privacy objects.

[0008] For example, in the comparative document with publication number CN117152646B, an AI lightweight large model method for unmanned power inspection is disclosed. The visible light image data collected by the drone inspection system is used as the research object, and an AI lightweight large model is formulated to specify the lightweight image encoder, decoder and keyword decoder of the large model, as well as the automatic acquisition of prompts such as points and boxes in the large model, to complete the rapid segmentation of visible light image data. However, in the above scheme, the power transmission corridor power components in the power inspection image data are automatically extracted through the AI ​​lightweight large model, but the privacy objects in the image are not determined, and there is no description of how the AI ​​lightweight large model determines the privacy objects in the image. It is impossible to accurately locate the location of the privacy objects, cannot identify the privacy objects according to user needs, and cannot perform different treatments on different privacy objects.

[0009] For example, in a comparative document with publication number CN119477762A, an image desensitization method is disclosed, which includes: obtaining the image to be desensitized uploaded by the user and the fidelity required by the user, and determining the target desensitized content in the image to be desensitized; generating a mask corresponding to the image to be desensitized according to the corresponding sensitive area of ​​the target desensitized content in the image to be desensitized; inputting the image to be desensitized and the mask into a preset sensitive content elimination model to obtain a sensitive content missing image; repairing the sensitive content missing image according to the fidelity required by the user to obtain the target desensitized image. However, in the above scheme, image recognition software, deep neural networks and other means are used to identify the corresponding target sensitive content from the image to be desensitized, and then the pixel area where the target desensitized content is located is determined as the sensitive area, without involving the determination of the privacy object (sensitive area) in the image through a large model, and it is not given how the large model determines the privacy object in the image, and it is impossible to accurately locate the position of the privacy object, identify the privacy object according to user needs, and perform different treatments on different privacy objects.

[0010] In summary, in the related technologies, no large model is given on how to determine the privacy objects in the image, so it is impossible to accurately locate the position of the privacy objects, nor to identify the privacy objects according to user needs, and it is impossible to perform different treatments on different privacy objects. Summary of the invention

[0011] The present application provides an image privacy positioning and identification method based on a multimodal large model, comprising:

[0012] Acquire an image to be detected and interactive text, wherein the interactive text includes a description of a privacy object;

[0013] The image to be detected is divided into multiple image blocks, and for each image block, a local feature vector of the image block is determined, a pooling operation is performed on the local feature vector to obtain a pooled local feature, the pooled local feature is mapped to a local visual token supported by a multimodal large model, and the local visual tokens of all image blocks are spliced ​​into a local visual token sequence; a global feature vector of the image to be detected is determined, a pooling operation is performed on the global feature vector to obtain a pooled global feature, and the pooled global feature is mapped to a global visual token; the interactive text is mapped to a semantic token;

[0014] The local visual token sequence, the global visual token and the semantic token are fused to obtain a fused feature; an initial query feature corresponding to the privacy object description is obtained based on the semantic token, and a conditional query feature is generated based on the initial query feature, the local visual token sequence and the global visual token; a target query feature is determined based on the fused feature and the conditional query feature;

[0015] Inputting the target query feature into the multimodal large model to obtain the predicted position, predicted category and predicted confidence of the privacy object, wherein the predicted confidence represents the confidence of the predicted category;

[0016] If the prediction confidence is greater than the matching score threshold, and the similarity between the predicted category and the description of the privacy object is greater than the semantic similarity threshold, it is determined that the image to be detected contains the privacy object of concern to the user, and the content in the image to be detected that matches the predicted position is desensitized;

[0017] If the similarity is not greater than the semantic similarity threshold, it is determined that the image to be detected contains a privacy object that the user does not care about, and the content in the image to be detected that matches the predicted position is not desensitized.

[0018] The present application provides an image privacy positioning and recognition device based on a multimodal large model, comprising:

[0019] An acquisition module is used to acquire an image to be detected and interactive text, wherein the interactive text includes a description of a privacy object; divide the image to be detected into multiple image blocks, determine a local feature vector of each image block, perform a pooling operation on the local feature vector to obtain a pooled local feature, map the pooled local feature to a local visual token supported by a multimodal large model, and splice the local visual tokens of all image blocks into a local visual token sequence; determine a global feature vector of the image to be detected, perform a pooling operation on the global feature vector to obtain a pooled global feature, and map the pooled global feature to a global visual token; and map the interactive text to a semantic token;

[0020] A processing module, configured to fuse the local visual token sequence, the global visual token and the semantic token to obtain a fused feature; obtain an initial query feature corresponding to the description of the privacy object based on the semantic token, and generate a conditional query feature based on the initial query feature, the local visual token sequence and the global visual token; determine a target query feature based on the fused feature and the conditional query feature; and input the target query feature into a multimodal large model to obtain a predicted position, a predicted category and a predicted confidence of the privacy object, wherein the predicted confidence represents the confidence of the predicted category;

[0021] A determination module, configured to determine that there is a privacy object of concern to the user in the image to be detected, and to perform desensitization on the content matching the predicted position in the image to be detected if the prediction confidence is greater than a matching score threshold and the similarity between the predicted category and the description of the privacy object is greater than a semantic similarity threshold; and to determine that there is a privacy object of no concern to the user in the image to be detected, and to not perform desensitization on the content matching the predicted position in the image to be detected if the similarity is not greater than a semantic similarity threshold.

[0022] The present application provides an electronic device, comprising: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement an image privacy positioning and recognition method based on a multimodal large model.

[0023] The present application provides a computer program product, including a computer program, which, when executed by a processor, implements an image privacy positioning and identification method based on a multimodal large model.

[0024] The present application provides a machine-readable storage medium, which stores machine-executable instructions that can be executed by a processor; wherein the processor is used to execute the machine-executable instructions to implement an image privacy positioning and recognition method based on a multimodal large model.

[0025] As can be seen from the above technical solutions, in the embodiment of the present application, a local visual token sequence, a global visual token and a semantic token can be obtained, and the target query feature can be determined based on the local visual token sequence, the global visual token and the semantic token, and the target query feature is input into the multimodal large model to obtain the predicted position, predicted category and predicted confidence. If the prediction confidence is greater than the matching score threshold, and the similarity between the predicted category and the description of the privacy object is greater than the semantic similarity threshold, it is determined that there is a privacy object that the user is concerned about in the image to be detected, that is, it is known which content of the image to be detected is private content, and the content matching the predicted position is desensitized. If the similarity between the predicted category and the description of the privacy object is not greater than the semantic similarity threshold, it is determined that there is a privacy object that the user is not concerned about in the image to be detected, that is, it is known which content of the image to be detected is not private content, and the content matching the predicted position is not desensitized. Based on the above method, it is possible to effectively identify and locate the privacy information in the image to be detected, take measures to protect the privacy information in a timely manner, and provide support for building a secure Internet of Things environment. Combine the multimodal large model with interactive text to achieve accurate positioning and identification of privacy information. By calculating the semantic similarity between the model output results (prediction category) and the interactive text (privacy object description), the multimodal large model's understanding and recognition accuracy of privacy information can be improved, ensuring that the prediction results of the multimodal large model are more in line with user needs and intentions, thereby improving accuracy.

[0026] In an embodiment of the present application, the image to be detected can be divided into multiple image blocks, the local feature vector of each image block is determined, the local feature vector is pooled to obtain the pooled local feature, the pooled local feature is mapped to the local visual token supported by the multimodal large model, and the local visual tokens of all image blocks are spliced ​​into a local visual token sequence. The global feature vector of the image to be detected can be determined, the global feature vector is pooled to obtain the pooled global feature, and the pooled global feature is mapped to the global visual token. The interactive text can be mapped to a semantic token. The local visual token sequence, the global visual token and the semantic token are fused to obtain a fused feature, the initial query feature corresponding to the description of the privacy object is obtained based on the semantic token, the conditional query feature is generated based on the initial query feature, the local visual token sequence and the global visual token, and the target query feature is determined based on the fused feature and the conditional query feature. On this basis, the target query feature can be input into the multimodal large model to obtain the predicted position of the privacy object. In the above process, it is shown how the multimodal large model determines the private object in the image (that is, the predicted position of the private object is determined based on the target query feature, and the method of obtaining the target query feature is given. Obviously, the target query feature is related to the local visual token, the global visual token, the semantic token, the fusion feature, the initial query feature, the conditional query feature, etc.), so that the predicted position of the private object can be accurately located.

[0027] In the embodiment of the present application, it is necessary to provide interactive text, and the interactive text includes a description of a privacy object. In this way, after the multimodal large model outputs the predicted position, predicted category, and predicted confidence of the privacy object, it is also possible to determine the similarity between the predicted category and the description of the privacy object, and compare whether the similarity is greater than the semantic similarity threshold. If so, it is determined that there is a privacy object that the user is concerned about in the image to be detected. If not, it is determined that there is a privacy object that the user is not concerned about in the image to be detected. On this basis, privacy objects can be identified according to user needs (that is, privacy objects that the user is concerned about and privacy objects that the user is not concerned about can be identified according to user needs), and different privacy objects can be processed differently. For example, the privacy objects that the user is concerned about in the image to be detected are desensitized, and the privacy objects that the user is not concerned about in the image to be detected are not desensitized, thereby specifically desensitizing some privacy objects without the need to desensitize all privacy objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a flowchart of an image privacy positioning and identification method based on a multimodal large model;

[0029] Figure 2 This is a flowchart of an image privacy positioning and identification method based on a multimodal large model;

[0030] Figure 3 This is a schematic diagram of the structure of an image privacy positioning and recognition system based on a multimodal large model;

[0031] Figure 4 It is a structural schematic diagram of an image privacy positioning and recognition device based on a multimodal large model;

[0032] Figure 5 It is a hardware structure diagram of an electronic device in one embodiment of the present application. DETAILED DESCRIPTION

[0033] In the embodiment of the present application, a method for image privacy positioning and recognition based on a multimodal large model is proposed, see Figure 1 FIG. 4 is a flow chart of the method, which may include:

[0034] Step 101: Acquire an image to be detected and interactive text, where the interactive text includes a description of a privacy object.

[0035] Step 102: Divide the image to be detected into multiple image blocks, determine the local feature vector of each image block, perform a pooling operation on the local feature vector to obtain a pooled local feature, map the pooled local feature to a local visual token supported by a multimodal large model, and concatenate the local visual tokens of all image blocks into a local visual token sequence. Determine the global feature vector of the image to be detected, perform a pooling operation on the global feature vector to obtain a pooled global feature, and map the pooled global feature to a global visual token. Map the interactive text to a semantic token.

[0036] Step 103: fuse the local visual token sequence, the global visual token and the semantic token to obtain a fused feature; obtain an initial query feature corresponding to the privacy object description based on the semantic token, and generate a conditional query feature based on the initial query feature, the local visual token sequence and the global visual token; and determine a target query feature based on the fused feature and the conditional query feature.

[0037] Step 104: Input the target query feature into the multimodal large model to obtain the predicted position, predicted category and predicted confidence of the privacy object, and the predicted confidence can represent the confidence of the predicted category.

[0038] Step 105: If the prediction confidence is greater than the matching score threshold, and the similarity between the predicted category and the description of the privacy object is greater than the semantic similarity threshold, it is determined that there is a privacy object that the user is concerned about in the image to be detected, and the content matching the predicted position in the image to be detected is desensitized; if the similarity between the predicted category and the description of the privacy object is not greater than the semantic similarity threshold, it is determined that there is a privacy object that the user is not concerned about in the image to be detected, and the content matching the predicted position in the image to be detected is not desensitized.

[0039] Exemplarily, if the prediction confidence is not greater than the matching score threshold, it is determined that there is no privacy object in the image to be detected, and the content matching the predicted position in the image to be detected is not desensitized.

[0040] Exemplarily, if the prediction confidence is greater than the matching score threshold, feature extraction (i.e., semantic feature extraction) can be performed on the predicted category to obtain a first semantic feature, and feature extraction can be performed on the privacy object description in the interactive text to obtain a second semantic feature; the semantic similarity between the first semantic feature and the second semantic feature is calculated; and the similarity between the predicted category and the privacy object description is determined based on the semantic similarity.

[0041] Exemplarily, performing a pooling operation on the local feature vector to obtain a pooled local feature may include, but is not limited to: performing a maximum pooling operation on the local feature vector to obtain a pooled local feature; or performing an average pooling operation on the local feature vector to obtain a pooled local feature; or performing a weighted pooling operation on the local feature vector to obtain a pooled local feature; wherein the local feature vector may include multiple eigenvalues, and weighting coefficients corresponding to different eigenvalues ​​may be the same or different.

[0042] Exemplarily, the fusion of the local visual token sequence, the global visual token and the semantic token to obtain a fusion feature may include but is not limited to: splicing the local visual token sequence, the global visual token and the semantic token to obtain a spliced ​​feature; performing a linear transformation on the spliced ​​feature to obtain a fusion feature. Alternatively, a weighted operation is performed on the local visual token sequence, the global visual token and the semantic token to obtain a weighted feature, and a fusion feature is obtained by normalizing the weighted feature; wherein, when performing a weighted operation on the local visual token sequence, the global visual token and the semantic token, the weight coefficient of the local visual token sequence may be greater than the weight coefficient of the semantic token, the weight coefficient of the global visual token may be greater than the weight coefficient of the semantic token, and the weight coefficient of the local visual token sequence may be the same as or different from the weight coefficient of the global visual token.

[0043] Exemplarily, when mapping the interactive text to a semantic token, a word segmentation operation is performed on the interactive text to obtain multiple words, and the semantic token may include word tokens corresponding to the multiple words. Obtaining the initial query feature corresponding to the privacy object description based on the semantic token may include but is not limited to: if the privacy object description is included in multiple words, selecting the target word token corresponding to the privacy object description from all word tokens in the semantic token; and determining the initial query feature based on the target word token.

[0044] Exemplarily, generating a conditional query feature based on the initial query feature, the local visual token sequence and the global visual token may include but is not limited to: fusing the local visual token sequence and the global visual token to obtain a context feature; performing a linear transformation on the context feature to obtain a transformed feature, and performing a weighted operation on the initial query feature and the transformed feature to obtain a conditional query feature.

[0045] Exemplarily, determining the target query feature based on the fused feature and the conditional query feature may include, but is not limited to: inputting the fused feature and the conditional query feature into a cross-attention network, interacting the fused feature and the conditional query feature through the cross-attention network, and obtaining the target query feature.

[0046] Exemplarily, the target query feature is input into the multimodal large model to obtain the predicted position, predicted category and predicted confidence of the privacy object, which may include but is not limited to: inputting the target query feature into the first feedforward network of the multimodal large model, and predicting the predicted position of the privacy object based on the target query feature through the first feedforward network; and inputting the target query feature into the second feedforward network of the multimodal large model, and predicting the predicted category of the privacy object based on the target query feature through the second feedforward network, and predicting the prediction confidence of the predicted category through the second feedforward network.

[0047] As can be seen from the above technical solutions, in the embodiment of the present application, a local visual token sequence, a global visual token and a semantic token can be obtained, and the target query feature can be determined based on the local visual token sequence, the global visual token and the semantic token, and the target query feature is input into the multimodal large model to obtain the predicted position, predicted category and predicted confidence. If the prediction confidence is greater than the matching score threshold, and the similarity between the predicted category and the description of the privacy object is greater than the semantic similarity threshold, it is determined that there is a privacy object that the user is concerned about in the image to be detected, that is, it is known which content of the image to be detected is private content, and the content matching the predicted position is desensitized. If the similarity between the predicted category and the description of the privacy object is not greater than the semantic similarity threshold, it is determined that there is a privacy object that the user is not concerned about in the image to be detected, that is, it is known which content of the image to be detected is not private content, and the content matching the predicted position is not desensitized. Based on the above method, it is possible to effectively identify and locate the privacy information in the image to be detected, take measures to protect the privacy information in a timely manner, and provide support for building a secure Internet of Things environment. Combine the multimodal large model with interactive text to achieve accurate positioning and identification of privacy information. By calculating the semantic similarity between the model output results (prediction category) and the interactive text (privacy object description), the multimodal large model's understanding and recognition accuracy of privacy information can be improved, ensuring that the prediction results of the multimodal large model are more in line with user needs and intentions, thereby improving accuracy.

[0048] In an embodiment of the present application, the image to be detected can be divided into multiple image blocks, the local feature vector of each image block is determined, the local feature vector is pooled to obtain the pooled local feature, the pooled local feature is mapped to the local visual token supported by the multimodal large model, and the local visual tokens of all image blocks are spliced ​​into a local visual token sequence. The global feature vector of the image to be detected can be determined, the global feature vector is pooled to obtain the pooled global feature, and the pooled global feature is mapped to the global visual token. The interactive text can be mapped to a semantic token. The local visual token sequence, the global visual token and the semantic token are fused to obtain a fused feature, the initial query feature corresponding to the description of the privacy object is obtained based on the semantic token, the conditional query feature is generated based on the initial query feature, the local visual token sequence and the global visual token, and the target query feature is determined based on the fused feature and the conditional query feature. On this basis, the target query feature can be input into the multimodal large model to obtain the predicted position of the privacy object. In the above process, it is shown how the multimodal large model determines the private object in the image (that is, the predicted position of the private object is determined based on the target query feature, and the method of obtaining the target query feature is given. Obviously, the target query feature is related to the local visual token, the global visual token, the semantic token, the fusion feature, the initial query feature, the conditional query feature, etc.), so that the predicted position of the private object can be accurately located.

[0049] In the embodiment of the present application, it is necessary to provide interactive text, and the interactive text includes a description of a privacy object. In this way, after the multimodal large model outputs the predicted position, predicted category, and predicted confidence of the privacy object, it is also possible to determine the similarity between the predicted category and the description of the privacy object, and compare whether the similarity is greater than the semantic similarity threshold. If so, it is determined that there is a privacy object that the user is concerned about in the image to be detected. If not, it is determined that there is a privacy object that the user is not concerned about in the image to be detected. On this basis, privacy objects can be identified according to user needs (that is, privacy objects that the user is concerned about and privacy objects that the user is not concerned about can be identified according to user needs), and different privacy objects can be processed differently. For example, the privacy objects that the user is concerned about in the image to be detected are desensitized, and the privacy objects that the user is not concerned about in the image to be detected are not desensitized, thereby specifically desensitizing some privacy objects without the need to desensitize all privacy objects.

[0050] The above technical solutions of the embodiments of the present application are described below in combination with specific application scenarios.

[0051] The embodiment of the present application proposes an image privacy positioning and identification method based on a multimodal large model, which can identify and locate the privacy of the detected image so as to take timely measures to protect the privacy information. Image privacy positioning and identification means that interactive text can be obtained, and the interactive text includes a description of the privacy object, thereby combining the multimodal large model with the interactive text to achieve accurate positioning and identification of privacy information.

[0052] For example, based on the interactive text, it can be known whether the private object in the image to be detected is a private object that the user is concerned about or a private object that the user is not concerned about. For the private object that the user is concerned about, the private information of the private object needs to be desensitized. For the private object that the user is not concerned about, even if there is a private object in the image to be detected, the private information of the private object does not need to be desensitized, thereby improving the multimodal large model's understanding and recognition accuracy of private information, and the prediction results are more in line with the user's needs and intentions.

[0053] The embodiment of the present application proposes an image privacy location and identification method based on a multimodal large model, which can be applied to electronic devices. The electronic device can be the sending device itself, or it can be a device between the sending device and the receiving device (such as a network device, a security device, etc.), and there is no restriction on the type of electronic device. For example, when an IoT device sends an image to a server device, in order to desensitize the privacy information in the image, the electronic device can be the IoT device itself, or it can be a forwarding device between the IoT device and the server device (as long as it can receive the image).

[0054] See also Figure 2 FIG. 4 is a flow chart of the method, which may include:

[0055] Step 201: Acquire an image to be detected and interactive text, where the interactive text includes a description of a privacy object.

[0056] Exemplarily, when an IoT device sends an image to a server device, the electronic device can obtain the image, which is called an image to be detected, that is, privacy identification and positioning of the image to be detected are required.

[0057] The IoT device may also obtain the interactive text of the image to be detected and provide the interactive text to the electronic device, or the user may directly provide the interactive text to the electronic device. The interactive text may be an interactive question text, or may be an interactive prompt word, or may be a question text provided by the user. The interactive text is used to prompt the multimodal large model what information to output.

[0058] For example, the interactive text can be: Please give all the private locations in the image, give the private object category of each private location, and pay special attention to the license plate private location. Based on the above interactive text, the multimodal large model needs to give all the private locations (such as rectangular boxes), give the private object category of each private location (such as vehicle, license plate, face, etc.), and mark the license plate private location.

[0059] In the above interactive text, the license plate represents a privacy object description, and the privacy object description is used to represent the privacy object that the user is concerned about. For example, the privacy object description represents that the privacy object that the user is concerned about is the license plate. For example, by segmenting the interactive text to obtain multiple words, the privacy object description can be selected from the multiple words. There is no restriction on the selection method, as long as the privacy object that the user is concerned about can be obtained.

[0060] Of course, the above is only an example of interactive text, and there is no limitation on the content of the interactive text. The interactive text may include descriptions of all privacy locations and privacy objects. In this way, the multimodal large model can output all privacy locations, and the privacy object description can be obtained based on the interactive text.

[0061] Exemplarily, when the IoT device sends a video to the server device, the electronic device can obtain the video. Since the video is composed of a large number of images, all the images of the video can be used as images to be detected, or part of the images of the video (such as key frames, and there is no restriction on the selection method of the key frames) can be used as images to be detected. In this way, the images to be detected can be obtained from the video. For example, the video is frame sampled to extract key frames as images to be detected, or the video is sampled at a certain time interval as images to be detected. Of course, the above is just an example of obtaining images to be detected from the video.

[0062] The IoT device can also obtain the interactive text of the video and provide the interactive text to the electronic device, or the user can directly provide the interactive text to the electronic device. In the subsequent process, the interactive text can be used as the interactive text of each image to be detected in the video.

[0063] Exemplarily, the IoT device may also provide audio to the electronic device, or the user may directly provide the audio to the electronic device, and the electronic device may convert the audio into text, i.e., interactive text.

[0064] Based on the image to be detected and the interactive text corresponding to the image to be detected, the privacy positioning of the image to be detected can be performed. For the convenience of description, the processing process of an image to be detected is taken as an example.

[0065] For the image to be detected, operations such as image cropping and image scaling can be performed on the image to be detected so that the image to be detected meets the input size requirements of the multimodal large model. And / or, for the image to be detected, the image to be detected can be normalized. And / or, for the image to be detected, the pixel value range of the image to be detected can be adjusted. The image to be detected in the subsequent process can be the image to be detected after the above-mentioned preprocessing. Of course, the preprocessing operation of the image to be detected is not limited to the above-mentioned processing.

[0066] For the interactive text corresponding to the image to be detected, natural language processing operations such as word segmentation, part-of-speech tagging, and named entity recognition can be performed on the interactive text to extract key semantic information, such as extracting the privacy object description in the interactive text and marking the privacy object description in the interactive text. In this way, in the subsequent processing process, it is possible to know which words in the interactive text are privacy object descriptions.

[0067] Step 202: Obtain a local visual token sequence corresponding to the image to be detected.

[0068] For example, the following steps can be used to obtain the local visual token sequence corresponding to the image to be detected:

[0069] Step S11, dividing the image to be detected into multiple image blocks.

[0070] For example, the image to be detected can be divided into multiple image blocks (patches) of fixed sizes. For example, the size of each image block is M*N, where M represents the width of the image block, which can be a positive integer greater than 1, and N represents the height of the image block, which can be a positive integer greater than 1. M and N can be the same or different.

[0071] Alternatively, the image to be detected can be divided into multiple image blocks of non-fixed sizes, such as the size of image block 1 is M1*N1, the size of image block 2 is M2*N2, the size of image block 3 is M3*N3, and so on, M1 is the same or different from M2, M1 is the same or different from M3, M2 is the same or different from M3, N1 is the same or different from N2, N1 is the same or different from N3, and N2 is the same or different from N3.

[0072] Step S12: for each image block of the image to be detected, determine the local feature vector of the image block.

[0073] For example, a local feature vector refers to the features of an image block in the image to be detected, that is, the features of a small area. The local feature vector is related to a specific object, object part or texture pattern in the image to be detected. For example, a local feature vector can be the edge, corner point or texture information of a specific area of ​​a specific object in the image to be detected. The local feature vector can be used to locate the specific location of the privacy object.

[0074] For example, the image block can be input into a convolutional neural network, and the convolutional neural network can be used to extract features from the image block to obtain a local feature vector of the image block. Alternatively, the image block can be input into a visual transformer network (a neural network based on an attention mechanism), and the visual transformer can be used to extract features from the image block to obtain a local feature vector of the image block.

[0075] Of course, the above is only an example, and there is no limitation on the method of obtaining the local feature vector. For each image block, feature extraction is performed on the image block to obtain the local feature vector of the image block.

[0076] Step S13: performing a pooling operation on the local feature vector to obtain a pooled local feature.

[0077] For example, for each image block, a maximum pooling operation may be performed on the local feature vector of the image block to obtain the pooled local feature. For example, the maximum pooling operation may be represented by the following formula: In the above formula, the local eigenvector can include multiple eigenvalues, X represents all eigenvalues ​​within the local eigenvector, It means to select the maximum value of all eigenvalues ​​in the local eigenvector. It represents the local feature after pooling, that is, the maximum eigenvalue among all eigenvalues. Obviously, by performing the maximum pooling operation on the local feature vector, the most important features can be retained.

[0078] For example, for each image block, an average pooling operation may be performed on the local feature vector of the image block to obtain the pooled local feature. For example, the average pooling operation may be represented by the following formula: In the above formula, the local eigenvector can include multiple eigenvalues, represents the i-th eigenvalue in the local eigenvector, i ranges from 1 to N, and N represents the total number of eigenvalues. Represents the local feature after pooling, that is, the average value of all feature values. Obviously, by performing average pooling operation on the local feature vector, the average feature of the local feature vector can be retained.

[0079] For example, for each image block, a weighted pooling operation may be performed on the local feature vector of the image block to obtain the pooled local feature. For example, the weighted pooling operation may be represented by the following formula: In the above formula, the local eigenvector can include multiple eigenvalues, represents the i-th eigenvalue in the local eigenvector, It represents the weighting coefficient corresponding to the i-th eigenvalue, and the weighting coefficients corresponding to different eigenvalues ​​are the same or different. The value range of i is 1-N, and N represents the total number of eigenvalues. It represents the local feature after pooling, that is, the weighted value of all feature values. Obviously, by performing weighted pooling operation on the local feature vector, the important features of the local feature vector can be highlighted.

[0080] In one possible implementation, The attention weight corresponding to the i-th eigenvalue can be expressed as the weight coefficient corresponding to the N eigenvalues. On this basis, the local feature vector can be input into the attention network, and the local feature vector can be weighted pooled by the attention network to obtain the pooled local feature. That is, the attention network provides the weight coefficients corresponding to the N feature values. , and perform a weighted pooling operation based on the N eigenvalues ​​of the local eigenvector.

[0081] In summary, the local feature vector can be pooled, and the pooling strategy can be maximum pooling, average pooling, weighted pooling (such as weighted pooling based on attention network), etc.

[0082] Step S14: Map the pooled local features into local visual tokens supported by the multimodal large model.

[0083] For example, since the multimodal large model supports token processing, the pooled local features of the image block can be mapped to tokens. For the sake of distinction, the tokens of the image blocks are called local visual tokens, thereby converting the local visual features of the image blocks into local visual tokens. A token can be a word, a punctuation mark, a number, or a subword, and there is no restriction on the type of the token.

[0084] In a possible implementation, for each image block, the pooled local features of the image block may be linearly projected to obtain a local visual token of the image block. For example, the pooled local features of the image block may be input into a linear projection layer, and the pooled local features may be linearly projected through the linear projection layer to obtain a local visual token of the image block. For example, the linear projection layer may include but is not limited to a linear layer (such as a fully connected layer), that is, the local visual token is obtained by linearly projecting the pooled local features through a linear layer. For example, it may be expressed by the following formula: In the above formula, It can represent the local features after pooling. It can represent a linear layer, that is, a linear projection is performed on the local features after pooling through a linear layer. Can represent local visual tokens.

[0085] For example, the linear layer is the basic layer structure in deep learning, which can also be called a fully connected layer or an affine transformation layer. The main function of the linear layer is to linearly transform the input data and weights, and output the result of the linear transformation. The linear transformation can be expressed as: y = Wx + b, y is the output data of the linear layer, W is the weight matrix, x is the input data of the linear layer, and b is the bias vector. The weight matrix W and the bias vector b are the parameters of the linear layer. The weight matrix W is used to linearly transform the input data, and the size of the weight matrix W determines the output dimension of the linear layer. The bias vector b is used to offset the output result, and the size of the bias vector b is the same as the output dimension of the linear layer. On this basis, the local features after pooling can be As input data x, the local visual token is used as output data y, so that the local visual token can be obtained.

[0086] Step S15: concatenate the local visual tokens of all image blocks into a local visual token sequence.

[0087] For example, after obtaining the local visual token of each image block, the local visual tokens of all image blocks can be spliced ​​into a complete sequence, which is recorded as a local visual token sequence. The local visual token sequence is used for subsequent multimodal tasks. For example, the local visual tokens of all image blocks are spliced ​​in order to obtain the following local visual token sequence: . is the local visual token of each image block, represents feature concatenation, is a local visual token sequence.

[0088] At this point, step 202 is completed, and a local visual token sequence corresponding to the image to be detected is obtained.

[0089] Step 203: Obtain the global visual token corresponding to the image to be detected.

[0090] For example, the following steps can be used to obtain the global visual token corresponding to the image to be detected:

[0091] Step S21, determining the global feature vector of the image to be detected.

[0092] For example, a global feature vector refers to the overall features of the image to be detected. The global feature vector can be used to describe the overall layout, color distribution, shape contour and other macro information of the image to be detected. For example, a global feature vector can be the overall color histogram, texture features or shape features of the image to be detected. The global feature vector can be used to understand the overall semantic information of the image to be detected.

[0093] For example, the image to be detected (i.e., the complete image) can be input into a convolutional neural network, and the convolutional neural network can be used to extract features from the image to be detected to obtain a global feature vector of the image to be detected. Alternatively, the image to be detected can be input into a visual transformer network, and the visual transformer can be used to extract features from the image to be detected to obtain a global feature vector of the image to be detected.

[0094] For example, the global feature vector output by a convolutional neural network or a visual Transformer network is a feature map that represents the high-level semantic information of the image to be detected. For example, the global feature vector can be expressed as , H and W are the height and width of the global eigenvector, D is the characteristic dimension.

[0095] Of course, the above is only an example, and there is no limitation on the method of obtaining the global feature vector. For the image to be detected, feature extraction is performed on the image to be detected to obtain the global feature vector.

[0096] Step S22: performing a pooling operation on the global feature vector to obtain a pooled global feature.

[0097] For example, the global feature vector of the image to be detected can be subjected to a maximum pooling operation to obtain the pooled global feature. For example, the maximum pooling operation can also be called global maximum pooling (GMP), that is, the global maximum pooling operation is performed on the global feature vector to generate a feature vector of a fixed length, and this feature vector is used as the pooled global feature. For example, the global maximum pooling can be expressed by the following formula: In the above formula, the global eigenvector can include multiple eigenvalues, represents the eigenvalue of the i-th row and j-th column in the global eigenvector, It means to select the maximum value of all eigenvalues ​​in the global eigenvector. Represents the global feature after pooling, that is, the maximum eigenvalue among all eigenvalues.

[0098] For example, an average pooling operation can be performed on the global feature vector of the image to be detected to obtain the pooled global feature. For example, the average pooling operation can also be called global average pooling (GAP), that is, a global average pooling operation is performed on the global feature vector (such as each channel of the global feature vector) to generate a feature vector of a fixed length, which is used as the pooled global feature. For example, global average pooling can be expressed by the following formula: In the above formula, the global eigenvector can include multiple eigenvalues, represents the eigenvalue of the i-th row and j-th column in the global eigenvector. The value range of i is 1-H, H represents the height of the global eigenvector, and the value range of j is 1-W, W represents the width of the global eigenvector. Represents the global feature after pooling, that is, the average value of all feature values.

[0099] Exemplarily, a weighted pooling operation can be performed on the global feature vector of the image to be detected to obtain the pooled global feature. For example, the global feature vector may include multiple eigenvalues, and the weighting coefficients corresponding to different eigenvalues ​​may be the same or different. For example, the weighting coefficient may represent the attention weight corresponding to the eigenvalue, which is calculated by the attention network. For example, the attention network can be pre-trained, the attention network may include multiple attention weights, the global feature vector to be detected can be input into the attention network, and the global feature vector can be weighted pooled by the attention network to obtain the pooled global feature.

[0100] In summary, the global feature vector can be pooled, and the pooling strategy can be maximum pooling, average pooling, weighted pooling (such as weighted pooling based on attention network), etc.

[0101] Step S23: Map the pooled global features into global visual tokens supported by the multimodal large model.

[0102] Exemplarily, since the multimodal large model supports the processing of tokens, the pooled global features of the image to be detected can be mapped to tokens, recorded as global visual tokens, thereby converting the global visual features of the image to be detected into global visual tokens. For example, the pooled global features can be linearly projected to obtain global visual tokens. For example, the pooled global features can be input into a linear projection layer, and the pooled global features can be linearly projected through the linear projection layer to obtain global visual tokens. For example, the linear projection layer can include but is not limited to a linear layer (such as a fully connected layer), that is, the pooled global features can be linearly projected through a linear layer to obtain a global visual token.

[0103] At this point, step 203 is completed, and the global visual token corresponding to the image to be detected is obtained.

[0104] Step 204: Map the interactive text into semantic tokens supported by the multimodal large model.

[0105] Exemplarily, a word segmentation operation may be performed on the interactive text to obtain multiple words, and the multiple words of the interactive text may be mapped to semantic tokens supported by the multimodal large model, and the semantic tokens may include word tokens corresponding to the multiple words. For example, when a word segmentation operation is performed on the interactive text to obtain word a, word b, and word c, when the multiple words of the interactive text are mapped to semantic tokens, the semantic tokens may include the word token of word a, the word token of word b, and the word token of word c.

[0106] Considering that the interactive text includes the privacy object description, that is, the privacy object description is a word after word segmentation, therefore, when obtaining the semantic token, the semantic token may include the word token of the privacy object description.

[0107] In one possible implementation, the interactive text may be input into a multimodal large model (such as a language encoder of the multimodal large model), and the multimodal large model may map multiple words of the interactive text into semantic tokens, and the semantic tokens may include word tokens corresponding to the multiple words.

[0108] Semantic tokens can be used as semantic features of interactive texts, that is, semantic features of interactive texts are extracted through a multimodal large model to obtain semantic features of interactive texts. For example, semantic tokens (semantic features) can be used to understand user intentions and needs and provide guidance for privacy positioning identification.

[0109] For example, in the preprocessing process, for the interactive text corresponding to the image to be detected, natural language processing operations such as word segmentation, part-of-speech tagging, and named entity recognition can be performed on the interactive text, and the interactive text after natural language processing operations can be input into the multimodal large model, and the multimodal large model extracts semantic features from the interactive text to obtain semantic tokens. For example, the extraction formula of semantic tokens is as follows: , T It can represent interactive text after word segmentation, and LLM can represent a multimodal large model. It can represent the semantic token (semantic features) output by a large multimodal model.

[0110] For example, in the preprocessing process, the interactive text can be subjected to natural language processing operations such as word segmentation, part-of-speech tagging, and named entity recognition, that is, which words are descriptions of privacy objects (implemented through the named entity recognition function). On this basis, when the semantic token includes word tokens corresponding to multiple words, the word token corresponding to the privacy object description can be selected from all word tokens.

[0111] Step 205: fuse the local visual token sequence, the global visual token and the semantic token to obtain a fused feature. For example, the local visual token sequence, the global visual token and the semantic token can be fused into a comprehensive feature (ie, fused feature) representation by using methods such as feature concatenation and weighted summation.

[0112] In a possible implementation, a feature splicing method can be used to splice together local visual token sequences, global visual tokens, and semantic tokens to form a comprehensive feature, which is called a fusion feature. The feature splicing method is simple and effective and can retain the original information of all modalities.

[0113] For example, the local visual token sequence, the global visual token and the semantic token can be concatenated to obtain the concatenated features. For example, the feature concatenation process of these tokens can be expressed by the following formula: In the above formula, represents a local visual token sequence, Represents the global visual token, Represents a semantic token, Represents the features after splicing, Represents the feature concatenation operation, which is to concatenate the local visual token sequence (multiple local visual tokens), the global visual token, and the semantic token together to obtain the concatenated feature, i.e., the concatenated token.

[0114] Exemplarily, after obtaining the concatenated features, the concatenated features may be linearly transformed to obtain fused features. For example, in order to make the fused features (i.e., the fused features of the local visual token sequence, the global visual token, and the semantic token) suitable for subsequent processing, after obtaining the concatenated features, the concatenated features may be projected (i.e., linearly transformed) through a linear layer to obtain the projected fused features.

[0115] For example, the concatenated features can be input to a linear projection layer, and the concatenated features can be linearly projected (i.e., linearly transformed) by the linear projection layer to obtain fused features. For example, the linear projection layer can include but is not limited to a linear layer (such as a fully connected layer), that is, the concatenated features can be linearly projected by the linear layer to obtain fused features. For example, the following formula can be used to represent it: In the above formula, Represents the features after splicing, It can represent the linear layer, that is, the concatenated features are linearly projected through the linear layer. It can represent fused features. For example, Linear can be a fully connected layer used to adjust the dimension of the feature, that is, to adjust the dimension of the concatenated feature.

[0116] In a possible implementation, a weighted summation method can be used to fuse the local visual token sequence, the global visual token, and the semantic token together to form a comprehensive feature, which is called a fused feature. In the weighted summation method, different weights can be assigned to features of different modalities during the fusion process, so that the importance of each modality can be adjusted according to the requirements of the task.

[0117] For example, the weighted features can be obtained by performing weighted operations on the local visual token sequence, the global visual token, and the semantic token. For example, the weighted operation process of these tokens can be expressed by the following formula: In the above formula, represents a local visual token sequence, Represents the global visual token, Represents a semantic token, Represents the weighted features. α represents the weight coefficient of the local visual token sequence, β Represents the weight coefficient of the global visual token, γ Indicates the weight coefficient of the semantic token. α , β and γ is a weighting coefficient that can be learned through training or set manually. In the weighted summation method, different weighting coefficients can be assigned to the local visual token sequence, the global visual token, and the semantic token, so that the importance of each modality can be adjusted according to the requirements of the task.

[0118] For example, when performing weighted operations on local visual token sequences, global visual tokens, and semantic tokens, the weighting coefficient of the local visual token sequence may be greater than the weighting coefficient of the semantic token, the weighting coefficient of the global visual token may be greater than the weighting coefficient of the semantic token, and the weighting coefficient of the local visual token sequence may be the same as or different from the weighting coefficient of the global visual token.

[0119] For example, after obtaining the weighted features, the weighted features may be normalized to obtain fused features. For example, in order to keep the scale of the features consistent, after obtaining the weighted features, the weighted features may be normalized, and the features after the normalization operation may be used as fused features.

[0120] Step 206: Obtain initial query features corresponding to the privacy object description based on the semantic token.

[0121] For example, when mapping interactive text to semantic tokens, the interactive text can be segmented to obtain multiple words, and the semantic token includes word tokens corresponding to the multiple words. Since the multiple words include the privacy object description, the semantic token includes the word token corresponding to the privacy object description. On this basis, the target word token corresponding to the privacy object description is selected from all the word tokens of the semantic token, that is, the word token related to the privacy object description is extracted from the semantic token.

[0122] Then, the initial query feature Q of the privacy object can be determined based on the target word token. For example, the target word token corresponding to the privacy object description can be used as the initial query feature Q of the privacy object. The target word token can be processed to obtain the initial query feature Q of the privacy object, and there is no restriction on this.

[0123] Step 207: Generate conditional query features based on the initial query features, the local visual token sequence and the global visual token. The conditional query features may also be called conditional object query features. .

[0124] In a possible implementation, the local visual token sequence and the global visual token may be used as context information of the initial query feature to generate a conditional query feature corresponding to the initial query feature.

[0125] Exemplarily, the local visual token sequence and the global visual token may be fused to obtain context features. For example, the local visual token sequence and the global visual token may be concatenated using a feature concatenation method to obtain context features. Alternatively, the local visual token sequence and the global visual token may be weighted using a weighted summation method to obtain context features.

[0126] Exemplarily, after obtaining the context feature, the context feature may be linearly transformed to obtain the transformed feature. For example, the context feature may be projected (i.e., linearly transformed) through a linear layer to obtain the transformed feature. For example, the context feature may be input to a linear projection layer, and the context feature may be linearly transformed through the linear projection layer to obtain the transformed feature. For example, the linear projection layer may include but is not limited to a linear layer, and the linear layer may include a fully connected layer.

[0127] For example, after obtaining the transformed features, the initial query features and the transformed features can be weighted to obtain the conditional query features, that is, to obtain the conditional object query features. For example, based on the initial query feature, the local visual token sequence and the global visual token, the following formula can be used to generate the conditional query feature: . represents a local visual token sequence, Represents the global visual token, It represents context features, that is, the context features are obtained by weighted summation. In the weighted operation, the weight coefficients of the local visual token sequence and the global visual token are both 1. Represents the transformed features, that is, the context features Perform a linear transformation ( ), and obtain the transformed features. Represents the initial query feature And the transformed features are weighted. During the weighted operation, the initial query feature The weight coefficient of the transformed feature is 1. Indicates conditional query characteristics.

[0128] Step 208: Determine the target query feature based on the fusion feature and the conditional query feature.

[0129] Exemplarily, in step 205, the local visual token sequence, the global visual token and the semantic token are fused to obtain a fused feature. In step 207, a conditional query feature is generated based on the initial query feature, the local visual token sequence and the global visual token. Based on this, the target query feature can be determined based on the fused feature and the conditional query feature. For example, the fused feature and the conditional query feature are input to the cross attention network ( ), the cross attention network can also be called the cross attention layer, through which the fusion features and the conditional query features interact to obtain the target query features.

[0130] For example, a cross-attention network can be pre-built, and there is no restriction on the network structure of the cross-attention network. After the fusion features and the conditional query features are input into the cross-attention network, the cross-attention network can interact with the fusion features and the conditional query features, thereby focusing on the visual area related to the query and obtaining the target query features. For example, the processing process of the cross-attention network can be expressed by the following formula: , represents the fusion feature, Indicates the conditional query characteristics, Indicates the interaction between fusion features and conditional query features through the cross attention network. It represents the query feature after cross attention and is recorded as the target query feature.

[0131] Step 209: Input the target query feature into the multimodal large model to obtain the predicted position, predicted category and predicted confidence of the privacy object, and the predicted confidence can represent the confidence of the predicted category.

[0132] Exemplarily, the input data of the multimodal large model is the target query feature, which is generated based on the semantic token, the local visual token sequence and the global visual token. The semantic token is obtained based on the interactive text input by the user, and the local visual token sequence and the global visual token are used to enhance the positioning capability of the privacy object. Based on this, the target query feature can reflect the information of the interactive text and the information of the image to be detected, and the multimodal large model can locate and identify the privacy correspondence of the image to be detected based on the target query feature. For example, combined with the interactive text input by the user and the image to be detected, the multimodal large model can predict the predicted position, predicted category and predicted confidence of the privacy object based on the target query feature. The predicted position of the privacy object indicates the location of the privacy object, such as a coordinate box, etc. The predicted position can be the coordinates of the four vertices of the coordinate box, or 1 vertex coordinate + length + width. The predicted category of the privacy object (i.e., the name of the predicted category) indicates what category the privacy object is, such as vehicle, license plate, face, etc. The prediction confidence may indicate the confidence of the prediction category (ie, the prediction probability). For example, when the prediction confidence is 90%, it indicates that the probability that the privacy object at the predicted location belongs to the prediction category is 90%.

[0133] For example, the multimodal large model may include a visual decoder, which inputs the target query features into the visual decoder of the multimodal large model. The visual decoder processes the target query features to predict the predicted position of the privacy object (i.e., the bounding box prediction result of the privacy object, which is used to indicate the specific position of the privacy object in the image to be detected), the predicted category of the privacy object, and the prediction confidence.

[0134] Considering that the target query feature is generated based on the conditional query feature, the local visual token sequence and the global visual token, when the visual decoder processes based on the target query feature, the conditional query feature is used to guide the positioning of the privacy object, and the local visual token sequence and the global visual token are combined.

[0135] In a possible implementation, the multimodal large model (such as a visual decoder of the multimodal large model) may include a first feedforward network (FFN), which is used to predict the predicted position of the private object (i.e., the coordinates of the bounding box). Based on this, the target query feature can be input into the first feedforward network of the multimodal large model, and the predicted position of the private object can be predicted by the first feedforward network based on the target query feature. There is no restriction on the prediction process of the first feedforward network, and it is sufficient to obtain the predicted position of the private object.

[0136] For example, the prediction process of the first feedforward network can be expressed as follows: , in the above formula, represents the target query feature, i.e., the output feature of the cross-attention network, Represents the predicted position of the private object (i.e., the coordinates of the bounding box) based on the first feedforward network (FFN) ), B represents the predicted position of the privacy object. In summary, the predicted position can be obtained through the first feedforward network.

[0137] In a possible implementation, the multimodal large model (such as a visual decoder of the multimodal large model) may include a second feedforward network (FFN), which is used to predict the predicted category and prediction confidence of the private object (the prediction confidence is used to represent the category matching score, that is, the category matching score of the predicted category). Based on this, the target query feature can be input into the second feedforward network of the multimodal large model, and the predicted category of the private object is predicted based on the target query feature by the second feedforward network, and the prediction confidence of the predicted category is predicted by the second feedforward network. There is no restriction on the prediction process of the second feedforward network, and it is sufficient to obtain the predicted category of the private object and the prediction confidence of the predicted category.

[0138] For example, the prediction process of the second feedforward network can be expressed as follows: , in the above formula, represents the target query feature, Represents the prediction confidence (i.e., category matching score) of the predicted category of the private object predicted by the second feed-forward network (FFN) ), Represents the prediction confidence of the predicted category. In summary, the predicted category and prediction confidence can be obtained through the second feedforward network.

[0139] Exemplarily, in the visual decoder of the multimodal large model, multiple cross-attention layers are used to interact the conditional query features and the complete visual tags (local visual token sequence and global visual token) (the target query feature reflects the interaction between the two), so as to learn the association between the visual context and the conditional query features. From the output of the cross-attention layer, the matching score and bounding box of the target query feature can be calculated. The matching score reflects the predicted category of the predicted object, and the bounding box reflects the predicted position of the predicted object, so as to predict the predicted position and predicted category of the privacy object related to the language input in the image to be detected.

[0140] Step 210: Determine whether the prediction confidence is greater than a configured matching score threshold.

[0141] If not, step 211 may be executed, and if so, step 212 may be executed.

[0142] Step 211: determine that there is no privacy object in the image to be detected (that is, the predicted position corresponding to the prediction confidence is not a privacy object), and do not perform desensitization on the content matching the predicted position in the image to be detected.

[0143] Exemplarily, a matching score threshold can be pre-configured, such as 0.7, 0.75, etc. If the prediction confidence is not greater than the matching score threshold, it means that the probability that the privacy object at the predicted position belongs to the predicted category is small, and the predicted position can be considered not to be a privacy object. In this way, there is no need to desensitize the content that matches the predicted position in the image to be detected. When sending the image to be detected (such as a single-frame image or the image of a video), the content of the predicted position of the image to be detected is not desensitized and can be sent directly.

[0144] Step 212: Determine the similarity (ie, semantic similarity) between the predicted category and the privacy object description.

[0145] Exemplarily, if the prediction confidence is greater than the matching score threshold, the content of the predicted location can be evaluated based on semantic similarity, such as calculating the semantic similarity between the predicted category output by the multimodal large model and the description of the privacy object of the user interaction, thereby ensuring the accuracy of the recognition result based on semantic similarity.

[0146] Exemplarily, based on the predicted category (i.e., the predicted category name) output by the multimodal large model, feature extraction can be performed on the predicted category (i.e., extracting the semantic features of the predicted category) to obtain a first semantic feature, and there is no restriction on the semantic feature extraction method. Based on the privacy object description in the above interactive text (i.e., the privacy object name, such as license plate, vehicle, etc.), feature extraction can be performed on the privacy object description (i.e., extracting the semantic features of the privacy object description) to obtain a second semantic feature.

[0147] Exemplarily, the semantic similarity between the first semantic feature and the second semantic feature may be calculated. For example, cosine similarity or other similarity measurement methods may be used to calculate the semantic similarity between the first semantic feature and the second semantic feature. For example, the semantic similarity may be determined using the following formula: semantic similarity = cos(F 预测 , F 用户 ), in the above formula, F 预测 Represents the first semantic feature of the predicted category, F 用户 It represents the second semantic feature of the privacy object description, and cos represents the cosine similarity.

[0148] Exemplarily, after obtaining the semantic similarity between the first semantic feature and the second semantic feature, the similarity (ie, the degree of matching) between the predicted category and the privacy object description may be determined based on the semantic similarity.

[0149] Step 213: Determine whether the similarity is greater than a configured semantic similarity threshold.

[0150] If yes, step 214 may be executed, and if no, step 215 may be executed.

[0151] Step 214: Determine whether there is a privacy object of concern to the user in the image to be detected (i.e., the predicted position corresponding to the prediction confidence is a privacy object, and the privacy object is a privacy object of concern to the user, i.e., a privacy object that needs to be desensitized), and desensitize the content in the image to be detected that matches the predicted position.

[0152] Step 215: determine that there is a private object that the user does not care about in the image to be detected (that is, the predicted position corresponding to the prediction confidence is a private object, and the private object is a private object that the user does not care about, that is, a private object that does not need desensitization), and do not desensitize the content matching the predicted position in the image to be detected.

[0153] Exemplarily, if the prediction confidence is greater than the matching score threshold, it means that the probability that the private object at the predicted location belongs to the predicted category is relatively high, and the predicted location can be considered to be a private object.

[0154] The semantic similarity threshold can be pre-configured, such as 0.8, 0.85, etc. If the similarity is greater than the semantic similarity threshold, it means that the private object at the predicted position matches the description of the private object in the interactive text to a large extent. Therefore, the private object at the predicted position is a private object of concern to the user. In this case, it is necessary to desensitize the content matching the predicted position in the image to be detected. There is no restriction on the desensitization strategy, and it is sufficient to desensitize the image content at the predicted position. When sending the image to be detected (such as a single-frame image or the image of a video), the content of the predicted position of the image to be detected has been desensitized.

[0155] If the similarity is not greater than the semantic similarity threshold, it means that the private object at the predicted position has a small degree of match with the description of the private object in the interactive text, and therefore, the private object at the predicted position is a private object that the user does not care about. In this case, it is not necessary to desensitize the content matching the predicted position in the image to be detected. When sending the image to be detected (such as a single-frame image or the image of a video), the content at the predicted position of the image to be detected is not desensitized, and the image to be detected can be sent directly.

[0156] In summary, the prediction results output by the multimodal large model can be further screened and confirmed according to the intentions and needs of the user's interactive input (reflected by the description of the privacy object in the interactive text) to ensure the accuracy and reliability of the recognition results. Only the prediction results with prediction confidence and semantic similarity higher than the threshold (i.e., the privacy objects that the user is concerned about) are retained, and the privacy objects that the user is concerned about are desensitized.

[0157] In addition, the filtered prediction results can be fed back to the user, who can confirm, correct or supplement them. Combined with the interactive text input by the user, the category of the located privacy object is identified, and the category prediction results and specific descriptions of the privacy object are generated to achieve the screening and confirmation of the prediction results.

[0158] In a possible implementation, a result feedback and interaction process may also be involved. In the result feedback and interaction process, the result of privacy location identification may be fed back to the user, and the multimodal large model may be updated and optimized online based on the user's feedback and interaction information. For example, the following process may be involved:

[0159] Result feedback: Feedback the results of privacy positioning and recognition (such as bounding box coordinates, privacy object categories, etc.) to the user in a visual way, such as drawing a bounding box on the image to be detected, displaying the category name, etc. For example, the bounding box coordinates can be the predicted location, and the privacy object category can be the predicted category.

[0160] User interaction: Receive user feedback and interaction information, such as user confirmation, correction, and supplement of privacy location identification results, as well as new query requests or instructions proposed by users.

[0161] Result adjustment: Based on user feedback and interaction information, the prediction results of the multimodal large model are adjusted and optimized, such as correcting erroneous recognition results, updating the bounding box position, etc.

[0162] Interaction records: record the user's interaction history and feedback information, and provide data support for the online update and optimization of the multimodal large model. The online update of the multimodal large model can be based on the subsequent process.

[0163] In a possible implementation, a weakly supervised learning process may also be involved. In the weakly supervised learning process, a multimodal large model (such as a visual decoder of a multimodal large model) may be trained and fine-tuned using a weakly supervised learning algorithm based on a small amount of labeled data and a large amount of unlabeled data, so as to reduce dependence on labeled data and improve the generalization ability of the multimodal large model. Based on the fine-tuned multimodal large model, step 209 may be executed, and the multimodal large model outputs the predicted location, predicted category, and predicted confidence of the privacy object.

[0164] The weakly supervised learning process involves the preparation of labeled data, the use of unlabeled data, and weakly supervised learning training. For the preparation of labeled data: a small amount of labeled data can be collected, including information such as the category and location of the private object in the image, for the initial training and fine-tuning of the multimodal large model. For the use of unlabeled data: a large number of unlabeled images can be collected, and the multimodal large model can be used to extract features and make preliminary predictions on the images to generate pseudo labels. For weakly supervised learning training: labeled data and unlabeled data with pseudo labels can be used together for the training of the multimodal large model, and a self-training weakly supervised learning algorithm can be used to continuously optimize the parameters of the multimodal large model to improve the performance and generalization ability of the multimodal large model.

[0165] For example, for weakly supervised learning training, training and fine-tuning of large multimodal models can include:

[0166] 1. Preparation phase. Collect data: Collect multimodal datasets with a small amount of labeled data and a large amount of unlabeled data. Labeled data includes images, video frames and their corresponding private object categories and locations (bounding box coordinates). Unlabeled data is images or video frames without private object labels.

[0167] Preprocess data: All data (labeled and unlabeled) are preprocessed, including image scaling, normalization, text segmentation, etc., to make them meet the input requirements of multimodal large models.

[0168] 2. Training phase. Initialize the model: Use a small amount of labeled data to pre-train and fine-tune the multimodal large model to obtain an initial model, that is, the pre-trained multimodal large model is used as the initial model.

[0169] Generate pseudo labels: Use the initial model to predict the unlabeled data and generate pseudo labels. The pseudo labels can include category predictions and location predictions for privacy objects in the unlabeled data.

[0170] Filter pseudo labels: Filter pseudo labels according to the confidence of the initial model prediction, and only retain pseudo labels with higher confidence. For example, you can set a confidence threshold, such as 0.9, to only retain pseudo labels with a predicted probability higher than the confidence threshold, and discard pseudo labels with a predicted probability not higher than the confidence threshold.

[0171] Expand the labeled data: Add the filtered pseudo-label data to the original labeled data to form a new labeled dataset, which is used to retrain the multimodal large model.

[0172] Retrain the model: Based on the initial model, use the expanded labeled data set to retrain the multimodal large model to further optimize the parameters of the multimodal large model. Iteration: Repeat the above steps until the performance of the multimodal large model converges or reaches the predetermined number of training rounds.

[0173] In a possible implementation, an online model update process may also be involved. In the online model update process, the multimodal large model may be updated and optimized online based on user feedback and interaction information to adapt to new privacy protection requirements and scenario changes. For example, the following process may be involved:

[0174] Data collection: Collect user interaction records and feedback information, including user confirmation, correction, and supplement of the recognition results of the multimodal large model, as well as new query requests or instructions proposed by the user.

[0175] Model update: Use the collected data as new training data to update and optimize the multimodal large model online. For example, incremental learning, transfer learning and other methods can be used to continuously adjust the parameters of the multimodal large model, thereby improving the performance and adaptability of the multimodal large model.

[0176] Performance evaluation: Perform performance evaluation on the updated multimodal large model, using indicators such as accuracy, recall rate, and F1 value to evaluate the performance of the multimodal large model on the new data set to ensure the effectiveness of the model update.

[0177] Iterative optimization: Based on the performance evaluation results, the multimodal large model is further iteratively optimized to continuously improve the performance and adaptability of the model to meet the ever-changing privacy protection needs and scenario changes.

[0178] In one possible implementation, see Figure 3 As shown, it is a structural diagram of an image privacy positioning and recognition system based on a multimodal large model. The image privacy positioning and recognition system may include a data preprocessing module, a multimodal feature extraction module, a weakly supervised learning module, a privacy positioning and recognition module, a result feedback and interaction module, and a model online update module. The functions of these modules are described below.

[0179] The data preprocessing module is used to preprocess the input video or image, user interaction data (i.e., interactive text), including frame sampling, image cropping, scaling, text segmentation, speech recognition and other operations to meet the input requirements of the subsequent multimodal large model. For example, in step 201, after obtaining the image to be detected and the interactive text, the data preprocessing module can perform data preprocessing operations.

[0180] The multimodal feature extraction module is used to extract the visual features of the image to be detected and the semantic features of the user interaction input respectively, and provide basic feature representation for subsequent privacy positioning and identification. For example, the multimodal feature extraction module executes steps 202-208 to obtain the local visual token sequence and global visual token corresponding to the image to be detected, and obtain the semantic token. Based on the local visual token sequence, global visual token and semantic token, the fusion feature, initial query feature, conditional query feature and target query feature are obtained.

[0181] The weakly supervised learning module is used to train the multimodal large model based on a small amount of labeled data and a large amount of unlabeled data using a weakly supervised learning algorithm to reduce the dependence on labeled data and improve the generalization ability of the multimodal large model. For example, the weakly supervised learning module can implement the above-mentioned weakly supervised learning process.

[0182] The privacy positioning and identification module is used to use the trained multimodal large model, combined with user interactive input, to locate and identify the privacy content in the video and image, and generate the bounding box and category prediction results of the privacy object. Once the visual decoder locates the privacy object, blurring or pixelation processing (i.e., desensitization) can be applied to protect privacy. For example, the privacy positioning and identification module executes steps 209-215 to determine the predicted position, predicted category and predicted confidence of the privacy object based on the target query features. Based on the similarity between the predicted confidence and the predicted category and the description of the privacy object, it can be determined that there is no privacy object in the image to be detected, or there is a privacy object that the user is concerned about in the image to be detected, or there is a privacy object that the user is not concerned about in the image to be detected, and then decide whether to desensitize the predicted position.

[0183] The result feedback and interaction module is used to feed back the results of privacy positioning identification to users, and to update and optimize the multimodal large model online based on user feedback and interaction information to improve the performance and adaptability of the multimodal large model. For example, the result feedback and interaction module can realize the result feedback and interaction process.

[0184] The model online update module is used to update and optimize the multimodal large model online based on user feedback and interaction information, so that the multimodal large model can adapt to new privacy protection requirements and scenario changes. For example, the model online update module can implement the above model online update process.

[0185] It can be seen from the above technical solutions that in the embodiments of the present application, it is possible to effectively identify and locate the privacy information (i.e., sensitive information) in the image to be detected, take timely measures to protect the privacy information, and provide support for building a secure IoT environment. By calculating the semantic similarity between the model output results (prediction categories) and the interactive text (privacy object description), the multimodal large model's understanding and recognition accuracy of privacy information is improved, ensuring that the prediction results of the multimodal large model meet the user's needs and intentions and improve accuracy. Through multimodal context modeling, the semantic relationship between images and text can be understood, and privacy information can be identified more accurately.

[0186] It is capable of open category detection. For example, it can detect any privacy information in an image, not just predefined categories of privacy information, such as faces, license plates, etc. It can identify and locate any privacy content that appears in the image, can adapt to different IoT scenarios, and can identify and locate new privacy categories.

[0187] It can achieve situational understanding. For example, it can understand the meaning of private information in different scenarios and the semantic relationship between images and texts, so as to more accurately identify and locate private content. In this way, it can handle more complex application scenarios, such as occlusion, lighting changes, etc.

[0188] It can achieve end-to-end modeling, for example, identifying and locating sensitive information from images without complex preprocessing or feature extraction. It can combine the text generation capability with the positioning capability of the image detection model to achieve end-to-end modeling, which can more effectively utilize image and text information and improve detection accuracy.

[0189] It can reduce the dependency on annotations based on weakly supervised learning. For example, based on the weakly supervised learning method, a small amount of labeled data and a large amount of unlabeled data can be used to train and fine-tune large multimodal models, which significantly reduces the dependency on labeled data, reduces the cost and time of annotation, and improves the practicality and scalability of the model.

[0190] It is possible to improve accuracy based on semantic similarity evaluation. For example, during the training and prediction process, based on the semantic similarity evaluation mechanism, by calculating the semantic similarity between the model output and the user interaction input, the model's understanding and recognition accuracy of private content can be improved, ensuring that the prediction results are more in line with user needs and intentions.

[0191] It can achieve online update and optimization. For example, it supports online update and optimization of multimodal large models based on user feedback and interaction, so that multimodal large models can continuously learn and adapt to new privacy protection requirements and scenario changes, and improve the performance and adaptability of the model. In addition, it can achieve real-time performance. For example, it can quickly identify and locate private information and take timely measures to protect private information.

[0192] Based on the same application concept as the above method, an image privacy positioning and identification device based on a multimodal large model is proposed in the embodiment of the present application, see Figure 4 FIG. 1 is a schematic diagram of the structure of the image privacy positioning and identification device based on the multimodal large model, and the device may include:

[0193] The acquisition module 41 is used to acquire the image to be detected and the interactive text, wherein the interactive text includes a description of a privacy object; divide the image to be detected into multiple image blocks, determine the local feature vector of each image block, perform a pooling operation on the local feature vector to obtain a pooled local feature, map the pooled local feature to a local visual token supported by a multimodal large model, and splice the local visual tokens of all image blocks into a local visual token sequence; determine the global feature vector of the image to be detected, perform a pooling operation on the global feature vector to obtain a pooled global feature, and map the pooled global feature to a global visual token; and map the interactive text to a semantic token;

[0194] The processing module 42 is used to fuse the local visual token sequence, the global visual token and the semantic token to obtain a fused feature; obtain an initial query feature corresponding to the description of the privacy object based on the semantic token, and generate a conditional query feature based on the initial query feature, the local visual token sequence and the global visual token; determine a target query feature based on the fused feature and the conditional query feature; input the target query feature into the multimodal large model to obtain a predicted position, a predicted category and a predicted confidence of the privacy object, wherein the predicted confidence represents the confidence of the predicted category;

[0195] The determination module 43 is used to determine that there is a privacy object of concern to the user in the image to be detected if the prediction confidence is greater than the matching score threshold and the similarity between the predicted category and the description of the privacy object is greater than the semantic similarity threshold, and desensitize the content in the image to be detected that matches the predicted position; if the similarity is not greater than the semantic similarity threshold, it is determined that there is a privacy object of no concern to the user in the image to be detected, and the content in the image to be detected that matches the predicted position is not desensitized.

[0196] Exemplarily, the determination module 43 is also used to determine that there is no privacy object in the image to be detected if the prediction confidence is not greater than the matching score threshold, and not desensitize the content in the image to be detected that matches the predicted position; the determination module 43 is also used to perform feature extraction on the predicted category to obtain a first semantic feature, and perform feature extraction on the privacy object description in the interactive text to obtain a second semantic feature if the prediction confidence is greater than the matching score threshold; calculate the semantic similarity between the first semantic feature and the second semantic feature; and determine the similarity between the predicted category and the privacy object description based on the semantic similarity.

[0197] Exemplarily, when the acquisition module 41 performs a pooling operation on the local feature vector to obtain the pooled local feature, it is specifically used to: perform a maximum pooling operation on the local feature vector to obtain the pooled local feature; or, perform an average pooling operation on the local feature vector to obtain the pooled local feature; or, perform a weighted pooling operation on the local feature vector to obtain the pooled local feature; wherein the local feature vector includes multiple eigenvalues, and the weighting coefficients corresponding to different eigenvalues ​​are the same or different.

[0198] Exemplarily, when the acquisition module 41 performs a pooling operation on the global feature vector to obtain a pooled global feature, it is specifically used to: perform a maximum pooling operation on the global feature vector to obtain a pooled global feature; or, perform an average pooling operation on the global feature vector to obtain a pooled global feature; or, perform a weighted pooling operation on the global feature vector to obtain a pooled global feature; wherein the local feature vector includes multiple eigenvalues, and the weighting coefficients corresponding to different eigenvalues ​​are the same or different.

[0199] Exemplarily, when the processing module 42 fuses the local visual token sequence, the global visual token and the semantic token to obtain a fused feature, it is specifically used to: splice the local visual token sequence, the global visual token and the semantic token to obtain a spliced ​​feature; and perform a linear transformation on the spliced ​​feature to obtain the fused feature; or,

[0200] A weighted operation is performed on the local visual token sequence, the global visual token and the semantic token to obtain a weighted feature, and a normalization operation is performed on the weighted feature to obtain the fused feature; wherein, when the weighted operation is performed on the local visual token sequence, the global visual token and the semantic token, the weighted coefficient of the local visual token sequence is greater than the weighted coefficient of the semantic token, the weighted coefficient of the global visual token is greater than the weighted coefficient of the semantic token, and the weighted coefficient of the local visual token sequence is the same as or different from the weighted coefficient of the global visual token.

[0201] Exemplarily, when mapping the interactive text to a semantic token, a word segmentation operation is performed on the interactive text to obtain multiple words, and the semantic token includes word tokens corresponding to the multiple words respectively; on this basis, the processing module 42 obtains the initial query feature corresponding to the privacy object description based on the semantic token, which is specifically used for: if the multiple words include the privacy object description, selecting the target word token corresponding to the privacy object description from all word tokens in the semantic token; and determining the initial query feature based on the target word token.

[0202] Exemplarily, when the processing module 42 generates a conditional query feature based on the initial query feature, the local visual token sequence and the global visual token, it is specifically used to: fuse the local visual token sequence and the global visual token to obtain a context feature; perform a linear transformation on the context feature to obtain a transformed feature (i.e., a feature obtained by a linear transformation operation); and perform a weighted operation on the initial query feature and the transformed feature to obtain the conditional query feature.

[0203] Exemplarily, when the processing module 42 determines the target query feature based on the fused feature and the conditional query feature, it is specifically used to: input the fused feature and the conditional query feature into a cross-attention network, and interact the fused feature and the conditional query feature through the cross-attention network (there is no limitation on the processing method of the cross-attention network) to obtain the target query feature.

[0204] Exemplarily, when the processing module 42 inputs the target query feature into the multimodal large model to obtain the predicted position, predicted category and predicted confidence of the privacy object, it is specifically used to: input the target query feature into the first feedforward network of the multimodal large model, and obtain the predicted position of the privacy object based on the target query feature through the first feedforward network; input the target query feature into the second feedforward network of the multimodal large model, and obtain the predicted category of the privacy object and the predicted confidence of the predicted category based on the target query feature through the second feedforward network.

[0205] Based on the same application concept as the above method, an electronic device is proposed in the embodiment of the present application, see Figure 5 As shown, it includes: a processor 51 and a machine-readable storage medium 52, the machine-readable storage medium 52 stores machine-executable instructions that can be executed by the processor 51; the processor 51 is used to execute the machine-executable instructions to implement the above-disclosed image privacy positioning and recognition method based on a multimodal large model.

[0206] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the above-disclosed image privacy positioning and identification method based on a multimodal large model can be implemented.

[0207] The above-mentioned machine-readable storage medium may be any electronic, magnetic, optical or other physical storage device, which may contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium may be: RAM (Radom Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as optical disk, DVD, etc.), or similar storage medium, or a combination thereof.

[0208] Based on the same application concept as the above method, an embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-disclosed image privacy positioning and identification method based on a multimodal large model.

[0209] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0210] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. An image privacy positioning and recognition method based on a multimodal large model, characterized in that: include: Acquire an image to be detected and interactive text, wherein the interactive text includes a description of a privacy object; The image to be detected is divided into multiple image blocks, and for each image block, a local feature vector of the image block is determined, a pooling operation is performed on the local feature vector to obtain a pooled local feature, the pooled local feature is mapped to a local visual token supported by a multimodal large model, and the local visual tokens of all image blocks are spliced ​​into a local visual token sequence; a global feature vector of the image to be detected is determined, a pooling operation is performed on the global feature vector to obtain a pooled global feature, and the pooled global feature is mapped to a global visual token; the interactive text is mapped to a semantic token; fusing the local visual token sequence, the global visual token and the semantic token to obtain a fused feature; Acquire an initial query feature corresponding to the privacy object description based on the semantic token, and generate a conditional query feature based on the initial query feature, the local visual token sequence and the global visual token; Determine the target query feature based on the fused feature and the conditional query feature; wherein, the fusion of the local visual token sequence, the global visual token and the semantic token to obtain the fused feature includes: splicing the local visual token sequence, the global visual token and the semantic token to obtain a spliced ​​feature; performing a linear transformation on the spliced ​​feature to obtain the fused feature; or, performing a weighted operation on the local visual token sequence, the global visual token and the semantic token to obtain a weighted feature, and performing a normalization operation on the weighted feature to obtain the fused feature; wherein, when performing a weighted operation on the local visual token sequence, the global visual token and the semantic token, the weighted coefficient of the local visual token sequence is greater than the weighted coefficient of the semantic token, the weighted coefficient of the global visual token is greater than the weighted coefficient of the semantic token, and the weighted coefficient of the local visual token sequence is the same as or different from the weighted coefficient of the global visual token; Inputting the target query feature into the multimodal large model to obtain the predicted position, predicted category and predicted confidence of the privacy object, wherein the predicted confidence represents the confidence of the predicted category; If the prediction confidence is greater than the matching score threshold, and the similarity between the predicted category and the description of the privacy object is greater than the semantic similarity threshold, it is determined that the image to be detected contains the privacy object of concern to the user, and the content in the image to be detected that matches the predicted position is desensitized; If the similarity is not greater than the semantic similarity threshold, it is determined that the image to be detected contains a privacy object that the user does not care about, and the content in the image to be detected that matches the predicted position is not desensitized.

2. The method according to claim 1, characterized in that: After inputting the target query feature into the multimodal large model to obtain the predicted location, predicted category and predicted confidence of the privacy object, the method further includes: If the prediction confidence is not greater than the matching score threshold, it is determined that there is no privacy object in the image to be detected, and the content matching the predicted position in the image to be detected is not desensitized; If the prediction confidence is greater than the matching score threshold, the method further includes: Extracting features of the predicted category to obtain a first semantic feature, and extracting features of the privacy object description in the interactive text to obtain a second semantic feature; Calculating the semantic similarity between the first semantic feature and the second semantic feature; The similarity between the predicted category and the privacy object description is determined based on the semantic similarity.

3. The method according to claim 1, characterized in that The performing a pooling operation on the local feature vector to obtain a pooled local feature includes: Performing a maximum pooling operation on the local feature vector to obtain the pooled local feature; or, Performing an average pooling operation on the local feature vector to obtain the pooled local feature; or, A weighted pooling operation is performed on the local feature vector to obtain the pooled local feature; wherein the local feature vector includes multiple eigenvalues, and weighting coefficients corresponding to different eigenvalues ​​are the same or different.

4. The method according to claim 1, characterized in that: When mapping the interactive text into a semantic token, a word segmentation operation is performed on the interactive text to obtain a plurality of words, and the semantic token includes word tokens corresponding to the plurality of words respectively; The obtaining, based on the semantic token, an initial query feature corresponding to the privacy object description includes: If the multiple words include the privacy object description, then selecting a target word token corresponding to the privacy object description from all word tokens in the semantic token; The initial query feature is determined based on the target word token.

5. The method according to claim 1, characterized in that: The generating of the conditional query feature based on the initial query feature, the local visual token sequence and the global visual token comprises: fusing the local visual token sequence and the global visual token to obtain a context feature; performing a linear transformation on the context feature to obtain a transformed feature, and performing a weighted operation on the initial query feature and the transformed feature to obtain the conditional query feature; The determining of the target query feature based on the fused feature and the conditional query feature includes: inputting the fused feature and the conditional query feature into a cross-attention network, and interacting the fused feature and the conditional query feature through the cross-attention network to obtain the target query feature.

6. The method according to claim 1, characterized in that The step of inputting the target query feature into the multimodal large model to obtain the predicted position, predicted category and predicted confidence of the privacy object includes: Inputting the target query feature into a first feedforward network of the multimodal large model, and obtaining a predicted position of the privacy object based on the target query feature through the first feedforward network; The target query feature is input into the second feedforward network of the multimodal large model, the predicted category of the privacy object is predicted based on the target query feature by the second feedforward network, and the prediction confidence of the predicted category is predicted by the second feedforward network.

7. An image privacy positioning and recognition device based on a multimodal large model, characterized in that: include: An acquisition module is used to acquire an image to be detected and interactive text, wherein the interactive text includes a description of a privacy object; divide the image to be detected into multiple image blocks, determine a local feature vector of each image block, perform a pooling operation on the local feature vector to obtain a pooled local feature, map the pooled local feature to a local visual token supported by a multimodal large model, and splice the local visual tokens of all image blocks into a local visual token sequence; determine a global feature vector of the image to be detected, perform a pooling operation on the global feature vector to obtain a pooled global feature, and map the pooled global feature to a global visual token; and map the interactive text to a semantic token; A processing module, used for fusing the local visual token sequence, the global visual token and the semantic token to obtain a fusion feature; Acquire an initial query feature corresponding to the privacy object description based on the semantic token, and generate a conditional query feature based on the initial query feature, the local visual token sequence and the global visual token; Determine a target query feature based on the fusion feature and the conditional query feature; Inputting the target query feature into the multimodal large model to obtain the predicted position, predicted category and predicted confidence of the privacy object, wherein the predicted confidence represents the confidence of the predicted category; A determination module, configured to determine that there is a privacy object of concern to the user in the image to be detected, and perform desensitization on the content matching the predicted position in the image to be detected if the prediction confidence is greater than a matching score threshold and the similarity between the predicted category and the description of the privacy object is greater than a semantic similarity threshold; if the similarity is not greater than the semantic similarity threshold, determine that there is a privacy object of no concern to the user in the image to be detected, and do not perform desensitization on the content matching the predicted position in the image to be detected; Wherein, when the processing module fuses the local visual token sequence, the global visual token and the semantic token to obtain a fused feature, it is specifically used to: splice the local visual token sequence, the global visual token and the semantic token to obtain a spliced ​​feature; perform a linear transformation on the spliced ​​feature to obtain the fused feature; or, perform a weighted operation on the local visual token sequence, the global visual token and the semantic token to obtain a weighted feature, and perform a normalization operation on the weighted feature to obtain the fused feature; wherein, when the weighted operation is performed on the local visual token sequence, the global visual token and the semantic token, the weighted coefficient of the local visual token sequence is greater than the weighted coefficient of the semantic token, the weighted coefficient of the global visual token is greater than the weighted coefficient of the semantic token, and the weighted coefficient of the local visual token sequence is the same as or different from the weighted coefficient of the global visual token.

8. The device according to claim 7, characterized in that The determination module is further configured to determine that no privacy object exists in the image to be detected if the prediction confidence is not greater than the matching score threshold, and not perform desensitization on the content in the image to be detected that matches the predicted position; The determination module is further configured to, if the prediction confidence is greater than the matching score threshold, perform feature extraction on the prediction category to obtain a first semantic feature, and perform feature extraction on the privacy object description to obtain a second semantic feature; Calculate the semantic similarity between the first semantic feature and the second semantic feature; and determine the similarity between the predicted category and the privacy object description based on the semantic similarity.

9. An electronic device, characterized in that: include: a processor and a machine-readable storage medium storing machine-executable instructions executable by the processor; The processor is used to execute machine executable instructions to implement the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Unmanned power inspection AI lightweight large model method and system

    CN117152646B

  • Mosaic processing method for live broadcast sensitive picture based on large model

    CN118283299A

  • File processing method, device and equipment based on privacy protection

    CN118313007A

  • Image desensitization method, device, equipment, medium and product

    CN119477762A

  • Visual question and answer method and system based on fine-grained adapter

    CN118607526A