Data processing method, device, apparatus and readable storage medium
By distinguishing between the question-and-answer attention area and the content protection area in image question answering, and using a redrawing process to protect the privacy content of images, the problem of privacy content leakage during image question answering is solved, achieving a balance between security and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-06-05
AI Technical Summary
In image-based question answering, user-uploaded images may contain private information, leading to privacy leaks during model learning and response, resulting in low security.
By defining the question-and-answer attention area and the content protection area, the image content is distinguished, and the first area to be processed is determined in the content protection area. The image content in the first area to be processed is protected by methods such as image cloaking redraw, image blurring redraw, similar replacement redraw, and feature blurring redraw.
This reduces the leakage of privacy information during the model's learning of images and answering of question text, while ensuring the accuracy of image question answering and appropriately improving the security of content protection information.
Smart Images

Figure CN122153924A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a data processing method, apparatus, device, and readable storage medium. Background Technology
[0002] In the existing field of image question answering, user-uploaded images often contain private information. When users upload images containing private information to the model for image question answering, the model may leak private information during the process of learning the image and answering image questions, resulting in low security. Summary of the Invention
[0003] This application provides a data processing method, apparatus, device, and readable storage medium that can appropriately improve the security of content protection information while ensuring the accuracy of image question answering.
[0004] One embodiment of this application provides a data processing method, including:
[0005] Based on the image and the corresponding question text, the question-and-answer attention region is determined in the image; the image content in the question-and-answer attention region is the image content in the image that is associated with the question text.
[0006] Obtain content protection information, and based on the image and the content protection information, determine the content protection area in the image; the image content in the content protection area is the image content associated with the content protection information in the image;
[0007] If there is an overlapping area between the question-and-answer attention area and the content protection area, then the first area to be processed is determined in the content protection area based on the question-and-answer attention area; the first area to be processed and the question-and-answer attention area do not overlap.
[0008] The first area to be processed is used to indicate that content protection processing is performed on the first image content in the image, and the first image content is the image content in the first area to be processed in the image.
[0009] Specifically, based on the question-and-answer attention region, the first region to be processed is determined within the content protection region, including:
[0010] The overlapping area between the question-and-answer attention area and the content protection area is defined as the business intersection area, and the area within the content protection area other than the business intersection area is defined as the first area to be processed.
[0011] This also includes:
[0012] Content protection processing is performed on the image content in the first area to be processed to obtain the first processed content. The first image content is then replaced in the image with the first processed content to obtain a question-and-answer protected image that includes the first processed content. The question-and-answer protected image is used to generate the answer text corresponding to the question text.
[0013] This also includes:
[0014] The image content in the first area to be processed is subjected to content protection processing to obtain the first processed content;
[0015] In the question-and-answer attention area, a second region to be processed is determined, the image content in the second region to be processed is determined as the second image content, and content protection processing is performed on the second image content to obtain the second processed content;
[0016] In the image, the first processed content is replaced by the first image content, and the second processed content is replaced by the second image content, to obtain a question-and-answer protected image that includes the first processed content and the second processed content.
[0017] Specifically, the second region to be processed is identified within the question-and-answer attention region, including:
[0018] T key information points are generated from the question text; T is a positive integer, and the key information points are the key information used to identify the obtained question text.
[0019] Attention processing is performed on the image content in the question-and-answer attention region and T key information of the questions to obtain a question attention result vector. Based on the question attention result vector, the question key region associated with the T key information of the questions is determined in the question-and-answer attention region.
[0020] The area outside the key question area in the question-and-answer attention region is identified as the second area to be processed.
[0021] Specifically, based on the image and the corresponding question text, the question-answering attention region is determined in the image, including:
[0022] Each pixel in the image is labeled with a connected component, resulting in M connected components; M is a positive integer. A connected component is used to represent an image region in which the pixel values of pixels satisfy the connected component labeling conditions and the pixels are adjacent to each other.
[0023] Feature extraction is performed on each of the M connected components to obtain the image feature vector, and feature extraction is performed on the corresponding question text to obtain the question text feature vector;
[0024] Attention processing is performed on the image feature vector and the question text feature vector to obtain the question-answering attention result vector. The question-answering attention region is then determined in the image based on the question-answering attention result vector.
[0025] Among them, determining the question-answering attention region in the image based on the question-answering attention result vector includes:
[0026] Generate attention scores for M connected components based on the question-answering attention result vector;
[0027] The connected component with the highest attention score is determined as the target connected component. Boundary contour points are generated based on the position information of pixels in the target connected component. The region enclosed by the boundary contour points in the image is determined as the question-answering attention region.
[0028] The image content in the first area to be processed is subjected to content protection processing to obtain the first processed content, including:
[0029] Obtain the protection level score associated with the content protection area; the protection level score is generated based on the content protection result vector corresponding to the content protection information, and the content protection result vector is the attention vector obtained by performing attention processing on the content protection information and the image;
[0030] The target repainting level is determined based on the protection level score, and the target content protection processing method indicated by the target repainting level is determined from N content protection processing methods; N is a positive integer.
[0031] Based on the target content protection processing method, the image content in the first area to be processed is subjected to content protection processing to obtain the first processed content.
[0032] The target content protection processing method is image cloaking redraw; based on the target content protection processing method, content protection processing is performed on the image content in the first area to be processed to obtain the first processed content, including:
[0033] The image content excluding the first region to be processed is identified as the redraw source content;
[0034] Obtain the background image from the redraw source content, extract features from the background image, and obtain the background feature vector;
[0035] Based on the background feature vector, the image content in the first region to be processed is redrawn to obtain the first processed content; the content feature vector of the first processed content has a similarity relationship with the background feature vector.
[0036] The target content protection processing method is image blurring and redrawing; based on the target content protection processing method, content protection processing is performed on the image content in the first area to be processed to obtain the first processed content, including:
[0037] Obtain the blur unit value P, and divide the image content in the first region to be processed into P image units to be blurred based on the blur unit value P; P is a positive integer; the P image units to be blurred include the target image unit to be blurred.
[0038] Based on the color channel information corresponding to each pixel in the target image unit to be blurred, the average color channel information of the target image unit to be blurred is generated. Based on the average color channel information, the pixels in the target image unit to be blurred are blurred to obtain the blurred image unit.
[0039] When P blurred image units corresponding to the image units to be blurred are obtained, the P blurred image units are determined as the first processing content.
[0040] The target content protection processing method is similar replacement and redrawing; based on the target content protection processing method, the image content in the first area to be processed is subjected to content protection processing to obtain the first processed content, including:
[0041] K key information items for objects are generated based on content protection information; K is a positive integer, and the key information items for objects are the key information used to identify the content protection information.
[0042] Attention processing is performed on the image content and K key object information in the first region to be processed to obtain the object attention result vector. Based on the object attention result vector, the key object regions associated with the K key object information are determined in the first region to be processed.
[0043] Obtain the content prompt text associated with content protection information, identify the target prompt text associated with the key area of the object in the content prompt text, extract features from the target prompt text, and obtain the text prompt feature vector;
[0044] The image content corresponding to the key areas of the object is damaged to obtain the damaged content. The image content other than the first area to be processed is determined as the redraw source content. The redraw source content is feature extracted to obtain the redraw source feature vector.
[0045] Image restoration of damaged content is performed based on text prompt feature vectors and redraw source feature vectors to obtain the first processed content.
[0046] The target content protection processing method is feature blurring and redrawing; based on the target content protection processing method, the image content in the first area to be processed is subjected to content protection processing to obtain the first processed content, including:
[0047] K key information items for objects are generated based on content protection information; K is a positive integer, and the key information items for objects are the key information used to identify the content protection information.
[0048] Attention processing is performed on the image content in the first region to be processed and K key information of objects to obtain an object attention result vector. Based on the object attention result vector, the key regions of objects associated with the K key information of objects are determined in the first region to be processed. The image content corresponding to the key regions of objects is blurred and redrawn to obtain the first processed content.
[0049] One embodiment of this application provides a data processing apparatus, including:
[0050] The question-and-answer region processing module is used to determine the question-and-answer attention region in the image based on the image and the corresponding question text; the image content in the question-and-answer attention region is the image content in the image that is associated with the question text.
[0051] The protected area processing module is used to acquire content protection information, and based on the image and the content protection information, to determine the content protection area in the image; the image content in the content protection area is the image content in the image associated with the content protection information.
[0052] The redrawing region processing module is used to determine a first region to be processed in the content protection region based on the question-and-answer attention region if there is an overlapping region between the question-and-answer attention region and the content protection region; the first region to be processed does not overlap with the question-and-answer attention region; wherein, the first region to be processed is used to indicate that content protection processing is performed on the first image content in the image, and the first image content is the image content in the first region to be processed in the image.
[0053] In one possible implementation, when the redrawing region processing module determines the first region to be processed within the content protection region based on the question-answering attention region, it specifically performs the following operations:
[0054] The overlapping area between the question-and-answer attention area and the content protection area is defined as the business intersection area, and the area within the content protection area other than the business intersection area is defined as the first area to be processed.
[0055] In one possible implementation, the redraw region processing module is also used to perform the following operations:
[0056] Content protection processing is performed on the image content in the first area to be processed to obtain the first processed content. The first image content is then replaced in the image with the first processed content to obtain a question-and-answer protected image that includes the first processed content. The question-and-answer protected image is used to generate the answer text corresponding to the question text.
[0057] In one possible implementation, the redraw region processing module is also used to perform the following operations:
[0058] The image content in the first area to be processed is subjected to content protection processing to obtain the first processed content;
[0059] In the question-and-answer attention area, a second region to be processed is determined, the image content in the second region to be processed is determined as the second image content, and content protection processing is performed on the second image content to obtain the second processed content;
[0060] An image whose first image content has been replaced with the first processed content and whose second image content has been replaced with the second processed content is identified as a question-and-answer protected image.
[0061] In one possible implementation, when the redrawing region processing module determines the second region to be processed within the question-and-answer attention region, it specifically performs the following operations:
[0062] T key information points are generated from the question text; T is a positive integer, and the key information points are the key information used to identify the obtained question text.
[0063] Attention processing is performed on the image content in the question-and-answer attention region and T key information of the questions to obtain a question attention result vector. Based on the question attention result vector, the question key region associated with the T key information of the questions is determined in the question-and-answer attention region.
[0064] The area outside the key question area in the question-and-answer attention region is identified as the second area to be processed.
[0065] In one possible implementation, the question-answering region processing module, when determining the question-answering attention region in an image based on the image and the corresponding question text, specifically performs the following operations:
[0066] Each pixel in the image is labeled with a connected component, resulting in M connected components; M is a positive integer. A connected component is used to represent an image region in which the pixel values of pixels satisfy the connected component labeling conditions and the pixels are adjacent to each other.
[0067] Feature extraction is performed on each of the M connected components to obtain the image feature vector, and feature extraction is performed on the corresponding question text to obtain the question text feature vector;
[0068] Attention processing is performed on the image feature vector and the question text feature vector to obtain the question-answering attention result vector. The question-answering attention region is then determined in the image based on the question-answering attention result vector.
[0069] In one possible implementation, the question-answering region processing module, when determining the question-answering attention region in an image based on the question-answering attention result vector, specifically performs the following operations:
[0070] Generate attention scores for M connected components based on the question-answering attention result vector;
[0071] The connected component with the highest attention score is determined as the target connected component. Boundary contour points are generated based on the position information of pixels in the target connected component. The region enclosed by the boundary contour points in the image is determined as the question-answering attention region.
[0072] In one possible implementation, the redrawing region processing module is used to perform content protection processing on the image content in the first region to be processed. When the first processed content is obtained, it is specifically used to perform the following operations:
[0073] Obtain the protection level score associated with the content protection area; the protection level score is generated based on the content protection result vector corresponding to the content protection information, and the content protection result vector is the attention vector obtained by performing attention processing on the content protection information and the image;
[0074] The target repainting level is determined based on the protection level score, and the target content protection processing method indicated by the target repainting level is determined from N content protection processing methods; N is a positive integer.
[0075] Based on the target content protection processing method, the image content in the first area to be processed is subjected to content protection processing to obtain the first processed content.
[0076] In one possible implementation, the target content protection processing method is image cloaking redraw; the redrawing area processing module is used to perform content protection processing on the image content in the first area to be processed based on the target content protection processing method, and when the first processed content is obtained, it is specifically used to perform the following operations:
[0077] The image content excluding the first region to be processed is identified as the redraw source content;
[0078] Obtain the background image from the redraw source content, extract features from the background image, and obtain the background feature vector;
[0079] Based on the background feature vector, the image content in the first region to be processed is redrawn to obtain the first processed content; the content feature vector of the first processed content has a similarity relationship with the background feature vector.
[0080] In one possible implementation, the target content protection processing method is image blurring and redrawing; the redrawing area processing module is used to perform content protection processing on the image content in the first area to be processed based on the target content protection processing method, and when the first processed content is obtained, it is specifically used to perform the following operations:
[0081] Obtain the blur unit value P, and divide the image content in the first region to be processed into P image units to be blurred based on the blur unit value P; P is a positive integer; the P image units to be blurred include the target image unit to be blurred.
[0082] Based on the color channel information corresponding to each pixel in the target image unit to be blurred, the average color channel information of the target image unit to be blurred is generated. Based on the average color channel information, the pixels in the target image unit to be blurred are blurred to obtain the blurred image unit.
[0083] When P blurred image units corresponding to the image units to be blurred are obtained, the P blurred image units are determined as the first processing content.
[0084] In one possible implementation, the target content protection processing method is similar replacement and redrawing; the redrawing area processing module is used to perform content protection processing on the image content in the first area to be processed based on the target content protection processing method, and when the first processed content is obtained, it is specifically used to perform the following operations:
[0085] K key information items for objects are generated based on content protection information; K is a positive integer, and the key information items for objects are the key information used to identify the content protection information.
[0086] Attention processing is performed on the image content and K key object information in the first region to be processed to obtain the object attention result vector. Based on the object attention result vector, the key object regions associated with the K key object information are determined in the first region to be processed.
[0087] Obtain the content prompt text associated with content protection information, identify the target prompt text associated with the key area of the object in the content prompt text, extract features from the target prompt text, and obtain the text prompt feature vector;
[0088] The image content corresponding to the key areas of the object is damaged to obtain the damaged content. The image content other than the first area to be processed is determined as the redraw source content. The redraw source content is feature extracted to obtain the redraw source feature vector.
[0089] Image restoration of damaged content is performed based on text prompt feature vectors and redraw source feature vectors to obtain the first processed content.
[0090] In one possible implementation, the target content protection processing method is feature blurring and redrawing; the redrawing region processing module is used to perform content protection processing on the image content in the first region to be processed based on the target content protection processing method, and when the first processed content is obtained, it is specifically used to perform the following operations:
[0091] K key information items for objects are generated based on content protection information; K is a positive integer, and the key information items for objects are the key information used to identify the content protection information.
[0092] Attention processing is performed on the image content in the first region to be processed and K key information of objects to obtain an object attention result vector. Based on the object attention result vector, the key regions of objects associated with the K key information of objects are determined in the first region to be processed. The image content corresponding to the key regions of objects is blurred and redrawn to obtain the first processed content.
[0093] One embodiment of this application provides a computer device, including: a processor, a memory, and a network interface;
[0094] The processor is connected to a memory and a network interface. The network interface is used to provide data communication functions, and the memory is used to store computer programs. When the computer program is executed by the processor, the computer device performs the method provided in the embodiments of this application.
[0095] One aspect of this application provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor, so that a computer device having the processor performs the method provided in this application.
[0096] One embodiment of this application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the method provided in this application embodiment.
[0097] This application embodiment, by determining the question-and-answer attention region and the content protection region, can distinguish the image content required by the model to process image question-and-answer. Then, based on the question-and-answer attention region, a first region to be processed is determined in the content protection region, and the first image content in the first region to be processed is subjected to content protection processing. This reduces the possibility of image content leakage associated with content protection information that may occur during the model's learning of images and answering question text. At the same time, since the first region to be processed and the question-and-answer attention region do not overlap, the model can still identify the image content associated with the question text through the question-and-answer attention region. It can be seen that this application can appropriately improve the security of content protection information while ensuring the accuracy of image question-and-answer. Attached Figure Description
[0098] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0099] Figure 1 This is a schematic diagram of a network architecture provided in an embodiment of this application;
[0100] Figure 2 This is a schematic diagram of a data processing scenario provided in an embodiment of this application. Figure 1 ;
[0101] Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 1 ;
[0102] Figure 4 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 2 ;
[0103] Figure 5 This is a schematic diagram of a data processing scenario provided in an embodiment of this application. Figure 2 ;
[0104] Figure 6 This is a schematic diagram of a data processing scenario provided in an embodiment of this application. Figure 3 ;
[0105] Figure 7 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;
[0106] Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0107] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0108] Please see Figure 1 , Figure 1 This is a schematic diagram of a network architecture provided in an embodiment of this application. For example... Figure 1As shown, the network architecture may include a terminal device 100 and a service server 200. The terminal device 100 may have a communication connection with the service server 200. The communication connection is not limited to a specific method. It may be directly or indirectly connected via wired communication, or directly or indirectly connected via wireless communication, or in other ways. This application does not impose any restrictions on this method.
[0109] The terminal devices can include: smartphones, tablets, laptops, desktop computers, intelligent voice interaction devices, smart home appliances (e.g., smart TVs), wearable devices, in-vehicle terminals, aircraft, and other intelligent terminals with data processing capabilities. In-vehicle terminals can be terminal devices used in intelligent transportation scenarios and assisted driving scenarios. It should be understood that, for example... Figure 1 The terminal device 100 shown can be equipped with an application client that has data processing capabilities. When the application client runs on each terminal device, it can interact with the aforementioned... Figure 1 Data interaction is performed between the business servers 200 shown.
[0110] Specifically, the application client may include: in-vehicle client, smart home client, entertainment client (e.g., game client), multimedia client (e.g., video client), social client, and information client (e.g., news client). In this embodiment, the application client may be integrated into a client (e.g., a social client) or may be a standalone client (e.g., a news client). This embodiment does not limit the type of application client.
[0111] Among them, the business server 200 can be the server corresponding to the application client. The business server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0112] like Figure 1As shown, the user can upload images, corresponding question text, and content protection information via an application client on the terminal device 100. The terminal device 100 can be equipped with an edge-side multimodal model. This edge-side multimodal model refers to an artificial intelligence model that runs on the terminal device and can simultaneously process multiple modalities of data (such as text, images, and audio). The content protection information can be descriptive information about the privacy-preserving content of the image; it can be text data, image data, or audio data.
[0113] Terminal device 100 can perform attention processing on the image and question text using an edge-side multimodal model (i.e., learn the correlation between the image and question text) to determine the question-answering attention region in the image associated with the question text. Similarly, terminal device 100 can perform attention processing on the image and content protection information using an edge-side multimodal model (i.e., learn the correlation between the image and content protection information) to determine the content protection region in the image associated with the content protection information. Both the question-answering attention region and the content protection region can be an image mask described by polygon vertices.
[0114] If there is an overlapping area between the question-and-answer attention area and the content protection area, the terminal device 100 can determine the overlapping area as the business intersection area, and determine the area in the content protection area other than the business intersection area as the first area to be processed. Content protection processing is performed on the first image content in the first area to be processed using an edge-side multimodal model to obtain the first processed content. Content protection processing includes methods such as content hiding processing and redrawing processing to remove the image indicated by the content protection information. Content hiding processing can refer to adding an opaque mask layer to the image area to hide it, or it can be directly erasing (eliminating) the image area. Redrawing processing can include image cloaking redrawing, image blurring redrawing, similar replacement redrawing, and feature blurring redrawing, etc., which are not limited in this embodiment.
[0115] The image whose first image content has been replaced with the first processing content is identified as a question-and-answer protected image. The first image content refers to the image content in the first area to be processed that has not been redrawn.
[0116] Terminal device 100 can send the question-and-answer protected image and question text obtained after content protection processing to business server 200. Business server recognizes the question-and-answer protected image through cloud multimodal model, generates the answer text corresponding to the question text, and returns the answer text to terminal device 100.
[0117] It is understood that the edge-side multimodal model in terminal device 100 can also be deployed in business server 200. Business server 200 can acquire images, perform content protection processing on the images through the edge-side multimodal model to obtain question-and-answer protected images, and then use cloud-based multimodal model to recognize the question-and-answer protected images and generate answer text corresponding to the question text. The edge-side multimodal model and the cloud-based multimodal model can be isolated to ensure the security of content protection information. The above-mentioned model isolation can be achieved through physical isolation or trusted environment isolation (a trusted environment requires authorization to access), and this embodiment of the application does not impose any limitations here.
[0118] This application embodiment, by determining the question-and-answer attention region and the content protection region, can distinguish the image content required by the model to process image question-and-answer. Based on the question-and-answer attention region, a first region to be processed is determined in the content protection region. The first image content in the first region to be processed is processed by the edge multimodal model deployed on the terminal device 100, so that the image content indicated by the content protection information cannot be recognized by the model. Thus, when the question-and-answer protected image and question text are sent to the cloud multimodal model of the business server 200, the leakage of image content associated with the content protection information that may occur during the model's learning of images and answering question text is reduced. At the same time, since the first region to be processed and the question-and-answer attention region do not overlap, the model can still recognize the image content associated with the question text through the question-and-answer attention region. It can be seen that this application can appropriately improve the security of content protection information while ensuring the accuracy of image question-and-answer.
[0119] Please see Figure 2 , Figure 2 This is a schematic diagram of a data processing scenario provided in an embodiment of this application. Figure 1 .like Figure 2 As shown, the user can upload image A and the corresponding question text to the terminal device via an application client. The question text corresponding to image A can be text obtained through audio recognition of the user's input or text directly entered by the user in the application client; this embodiment does not impose any limitations. The terminal device can be as described above. Figure 1 The terminal device 100 in the corresponding embodiment. The question text may be, for example, "the type of car in the image".
[0120] Terminal devices can obtain content protection information, which can be descriptive information about the private image content in image A. This content protection information can be a type description or a directly specified description. It can be selected and configured in the privacy content type in the application client, input by the user, or automatically identified by the model. Content protection information can be text data, image data, or audio data. Types of content protection information can include personal identification information, contact information, financial information, vehicle information, medical information, and address location information. For example, content protection information can be "person / image".
[0121] End-device multimodal models can be deployed in terminal devices. These end-device multimodal models refer to artificial intelligence models that run on terminal devices and can simultaneously process multiple modalities of data (such as text, images, and audio). Examples include the Phi-3-mini-128k-instruct model (an open-source lightweight model), the Chat-GLM (General Language Model) model (a natural language generation model), and the Qwen-VLM (Vision Language Model) model (a large-scale visual language model), which are multimodal large language models that can run on edge devices.
[0122] The terminal device can use an edge-side multimodal model, image A, and question text to determine the question-answering attention region B in image A, which is associated with the question text. Similarly, the terminal device can use an edge-side multimodal model, image A, and content protection information to determine the content protection region C in image A, which is associated with the content protection information. Both the question-answering attention region B and the content protection region C can be an image mask described by polygon vertices.
[0123] like Figure 2 As shown, there is an overlapping area between the question-and-answer attention area B and the content protection area C. The terminal device can define the overlapping area between the question-and-answer attention area and the content protection area as the business intersection area D, which can be represented by a diagonal line. The terminal device can define the area in the content protection area C other than the business intersection area D as the first processing area E, which can be represented by a horizontal line.
[0124] The terminal device can perform content protection processing on the first area E to be processed. Content protection processing includes methods such as content hiding processing and redrawing processing to remove the image indicated by the content protection information. Content hiding processing can refer to adding an opaque mask layer to the image area to hide it, or it can directly erase (eliminate) the image area. Redrawing processing can include image cloaking redrawing, image blurring redrawing, similar-type replacement redrawing, and feature blurring redrawing, etc., and this embodiment of the application does not impose limitations.
[0125] Taking content protection processing as a redrawing process as an example, the terminal device can determine the protection level score by the degree of correlation between the image content in the content protection area C and the content protection information, and determine the content protection processing method by the protection level score.
[0126] Taking image cloaking redrawing as an example of content protection processing, image cloaking redrawing can generate first processed content F in the first processing area E based on the background image (e.g., trees) in image A. Image A with the first processed content F replaced by the first processed content F is then identified as the question-and-answer protected image G. Here, the first processed content is the image content in the first processing area E of the image, and the question-and-answer protected image G is used to perform image recognition to generate the answer text corresponding to the question text. The question-and-answer protected image G can provide security protection for the image content associated with content protection information through the first processed content F.
[0127] This application's embodiments can be applied to Visual Question Answering (VQA) and Visual Re-drawing tasks. In Visual Re-drawing, this application can use content protection information to partially redraw the indicated image content, achieving style transfer of some image content while preserving other original features of the image. This enables diversified image processing and can be applied to image editing, visual design, and other fields. In Visual Question Answering, this application can perform content protection processing on the first image content in the first region to be processed, reducing the possibility of image content leakage associated with content protection information during the model's learning of images and answering question text. This effectively protects the content protection information while preserving the image content associated with the question text, ensuring that the accuracy of Visual Question Answering is not affected.
[0128] Please see Figure 3 , Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 1 This data processing method can be executed by a computer device, which can be, for example, Figure 1The terminal device 100 or business server 200 shown. The following description will use the example of this data processing method being executed by a computer device. This data processing method will at least include the following steps S101-S103:
[0129] Step S101: Based on the image and the corresponding question text, determine the question-and-answer attention region in the image; the image content in the question-and-answer attention region is the image content in the image that is associated with the question text.
[0130] Specifically, the computer device can acquire an image and the corresponding question text. The question text can be text obtained from the speech input by the object through audio recognition or text directly input by the object in the application client; this embodiment of the application does not impose any limitations on this. For example, the question text could be "the type of vehicle in the image," "whether the image contains a vehicle," or "the time the image was captured," etc.
[0131] Computer devices can be deployed with edge multimodal models. An edge multimodal model refers to an artificial intelligence model that runs on a terminal device and can simultaneously process multiple modalities of data (such as text, images, and audio). Examples include the Phi-3-mini-128k-instruct model, the Chat-GLM model, and the Qwen-VLM model, which are multimodal large language models that can run on edge devices.
[0132] Computer devices can divide an image into multiple image regions, such as M image regions, using an edge-side multimodal model. These image regions can be connected components or shape domains. Attention processing is then applied to these multiple image regions and the question text to obtain a question-and-answer attention result vector. This vector can be an M-dimensional vector, where each dimension represents a first attention score for the first image region and the question text. The first attention score indicates the relevance between the image region and the question text. The computer device can identify the image region with the highest first attention score as the question-and-answer attention region associated with the question text. The question-and-answer attention region can be an image mask described by polygon vertices. Optionally, the computer device can identify multiple image regions with first attention scores reaching a first threshold (a pre-set value) as question-and-answer attention regions, thus obtaining multiple question-and-answer attention regions.
[0133] Step S102: Obtain content protection information; based on the image and the content protection information, determine the content protection area in the image; the image content in the content protection area is the image content in the image associated with the content protection information.
[0134] Specifically, computer devices can acquire content protection information. Content protection information can be descriptive information about the private content of an image. This information can be a type description or a directly specified description. It can be selected and configured in the privacy content type settings of the application client, input by the user, or automatically identified by a model. Content protection information can be text data, image data, or audio data. Types of content protection information can include personally identifiable information, contact information, financial information, vehicle information, medical information, and location information, etc.
[0135] Computer devices can perform attention processing on multiple image regions and content protection information using an edge-side multimodal model to obtain a content protection result vector. This vector can be an M-dimensional vector, where each dimension represents a second attention score for the image region and the content protection information. The second attention score indicates the correlation between the image region and the content protection information. The computer device can identify the image region with the highest second attention score as the content protection region associated with the content protection information. The content protection region can be an image mask described by polygon vertices. Optionally, the computer device can identify multiple image regions with second attention scores reaching a second score threshold (a pre-set value) as content protection regions, thus obtaining multiple content protection regions.
[0136] Step S103: If there is an overlapping area between the question-and-answer attention region and the content protection region, then a first region to be processed is determined in the content protection region based on the question-and-answer attention region; the first region to be processed and the question-and-answer attention region do not overlap; wherein, the first region to be processed is used to indicate that the first image content in the image is subjected to content protection processing, and the first image content is the image content in the first region to be processed in the image.
[0137] Specifically, if there is an overlapping area between the question-and-answer attention area and the content protection area, the computer device can identify the overlapping area between the question-and-answer attention area and the content protection area as the business intersection area, and identify the area in the content protection area other than the business intersection area as the first area to be processed.
[0138] The computer device can perform content protection processing on the first area to be processed. Content protection processing includes methods such as content hiding and redrawing to remove the image indicated by content protection information. Content hiding can refer to adding an opaque mask layer to the image area to hide it, or it can be directly erasing (eliminating) the image area. Redrawing processing can include image cloaking redrawing, image blurring redrawing, similar-object replacement redrawing, and feature blurring redrawing, etc., and this embodiment of the application does not impose limitations.
[0139] Computer devices can determine the protection level score by using the second attention score corresponding to the content protection area. For example, the second attention score can be used as the protection level score for the content protection area, and the content protection processing method can be determined by the protection level score.
[0140] A computer device can perform content protection processing on the first image content corresponding to the first region to be processed using an edge-side multimodal model to obtain the first processed content. The image whose first image content has been replaced with the first processed content is then identified as the question-and-answer protected image. Here, the first image content is the image content within the first region to be processed in the image, and the question-and-answer protected image can be used to generate the answer text corresponding to the question text.
[0141] It is understandable that computer equipment can identify one or more first regions to be processed in an image, and each first region to be processed can have a different protection level score, that is, the redrawing method of each first region to be processed can be different.
[0142] This application embodiment, by determining the question-and-answer attention region and the content protection region, can distinguish the image content required by the model to process image question-and-answer. Then, based on the question-and-answer attention region, a first region to be processed is determined in the content protection region, and the first image content in the first region to be processed is subjected to content protection processing. This reduces the possibility of image content leakage associated with content protection information that may occur during the model's learning of images and answering question text. At the same time, since the first region to be processed and the question-and-answer attention region do not overlap, the model can still identify the image content associated with the question text through the question-and-answer attention region. It can be seen that this application can appropriately improve the security of content protection information while ensuring the accuracy of image question-and-answer.
[0143] Please see Figure 4 , Figure 4 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 2 This data processing method can be executed by a computer device, which can be, for example, Figure 1 The terminal device 100 or the business server 200 shown. The following description will use the example of this data processing method being executed by a computer device. This data processing method will at least include the following steps S201-S206:
[0144] Step S201: For each pixel in the image, a connected component label is performed to obtain M connected components; M is a positive integer. A connected component is used to represent an image region where the pixel values of the pixels satisfy the connected component labeling conditions and the pixels are adjacent to each other. Feature extraction is performed on each of the M connected components to obtain an image feature vector. Feature extraction is performed on the question text corresponding to the image to obtain a question text feature vector. Attention processing is performed on the image feature vector and the question text feature vector to obtain a question-answering attention result vector.
[0145] Specifically, computer devices can deploy edge multimodal models. Edge multimodal models refer to artificial intelligence models that run on terminal devices and can simultaneously process multiple modalities of data (such as text, images, and audio). Examples include the Phi-3-mini-128k-instruct model, the Chat-GLM model, and the Qwen-VLM model, which are multimodal large language models that can run on edge devices. Edge multimodal models can also be graph-text multimodal models, which are artificial intelligence models capable of simultaneously processing and understanding image and text data. These models, by combining visual and linguistic information, can perform tasks such as image description generation, image question answering, image classification, and text matching. Their core lies in fusing and aligning data from different modalities, enabling the model to extract useful features from multiple information sources and perform comprehensive analysis and reasoning.
[0146] The computer device can input an image to the image processing layer of the edge multimodal model. This image processing layer can be an image encoder in CLIP (Contrastive Language-Image Pre-Training) or a ViT (Vision Transformer) image encoder; this embodiment does not impose any limitations. The computer device can perform binarization processing on the image through the image processing layer, setting the grayscale values of pixels in the image that are greater than or equal to a business grayscale threshold to 255, and setting the grayscale values of pixels in the image that are less than the business grayscale threshold to 0, thus obtaining a grayscale image. The business grayscale threshold can be set according to business requirements or specific images; this embodiment does not impose any limitations.
[0147] Computer equipment can label connected components in a grayscale image, classifying adjacent pixels that satisfy the connectivity labeling criteria into a single connected component, resulting in M connected components. The connectivity labeling criteria can refer to pixels having the same grayscale value. A connected component represents an image region where pixels have adjacent pixel values that satisfy the connectivity labeling criteria. The computer equipment can then extract features from the image content within each of the M connected components to obtain image feature vectors.
[0148] The computer device can input the question text into the text processing layer of the edge multimodal model. This text processing layer can be a BERT (Bidirectional Encoder Representations from Transformers) text encoder. The computer device can use a tokenizer in the text processing layer to segment and encode the question text into tokens. For example, if the question text is "car types in an image", the corresponding token sequence can be represented as [image, car, type]. Here, [image], [car], and [type] are all tokens. The text segmentation method can be word-based, character-based, or subword-based; this embodiment does not impose any limitations. The computer device can embed all tokens in the question text into vectors of the same size, and then concatenate these vectors to obtain the question text feature vector.
[0149] The computer device can input the image feature vector and the question text feature vector into the attention processing layer of the edge multimodal model, and then input the question text feature vector and the query parameter matrix. Perform a dot product operation to obtain the query vector Q, and then combine the image feature vector with the key parameter matrix. Perform a dot product operation to obtain the key vector K, and then combine the image feature vector with the value parameters. Performing a dot product on the matrix yields a value vector V. The query parameter matrix is included. Key parameter matrix AND-value parameter matrix All of them are matrices composed of learnable parameters for attention processing.
[0150] Computer devices can perform a transpose operation on the query vector Q and the key vector K, resulting in K. T Perform a dot product operation to obtain the attention score vector. This is based on the dimension d of the key vector K. K square root The attention score vector is dimensionality reduced, and then normalized to obtain the attention weight vector. The attention weight vector is then multiplied by the value vector to obtain the question-answering attention result vector Attn. Q The process can be shown in formula (1):
[0151]
[0152] Wherein, the softmax function is a normalization process, d KLet K be the dimension value of the key vector K.
[0153] Step S202: Generate attention scores for M connected components based on the question-answering attention result vector; determine the connected component with the maximum attention score as the target connected component; generate boundary contour points based on the position information of pixels in the target connected component; and determine the region enclosed by the boundary contour points in the image as the question-answering attention region.
[0154] Specifically, the question-answering attention result vector Attn Q It can be an M-dimensional vector, the question-answering attention result vector Attn. Q Each dimension in the diagram can represent a first attention score between a connected component and the question text. The first attention score indicates the relevance between the connected component and the question text. The computer device can determine the connected component with the highest first attention score as the target connected component. Based on the positional information of the pixels in the target connected component, the computer device can determine S boundary contour points. The number of boundary contour points can be determined based on the range and shape of the target connected component. The computer device can determine the polygonal region formed by the boundary contour points as the question-answering attention region. The question-answering attention region can be an image mask described by polygon vertices. For example, the question-answering attention region can be represented in text form by the polygon vertex coordinates, such as {[100,50],[50,200],[150,200]}, where [100,50],[50,200], and [150,200] are the coordinates of the boundary contour points in the image. Optionally, the computer device can identify multiple image regions whose first attention scores reach a first score threshold (pre-set value) as question-and-answer attention regions, thereby obtaining multiple question-and-answer attention regions.
[0155] Step S203: Obtain content protection information; based on the image and the content protection information, determine the content protection area in the image; the image content in the content protection area is the image content in the image associated with the content protection information.
[0156] Specifically, computer devices can acquire content protection information. Content protection information can be descriptive information about the private content of an image. This information can be a type description or a directly specified description. It can be selected and configured in the privacy content type settings of the application client, input by the user, or automatically identified by a model. Content protection information can be text data, image data, or audio data. Types of content protection information can include personally identifiable information, contact information, financial information, vehicle information, medical information, and location information, etc.
[0157] Computer devices can perform image recognition on image data or audio recognition on audio data to obtain text data of content protection information. Taking the content protection information as text data as an example, the computer device can extract features from the content protection information through the text processing layer of the edge multimodal model to obtain the content protection feature vector corresponding to the content protection information. The process can be referred to in the specific description of the text processing layer in step S201 above, and will not be repeated here.
[0158] The computer device can input the image feature vector and the content protection feature vector into the attention processing layer of the edge multimodal model. The attention processing layer performs cross-attention calculation on the image feature vector and the content protection feature vector to obtain the content protection result vector. The process of cross-attention calculation can be found in the specific description of the attention processing layer in step S201 above, and will not be repeated here.
[0159] The content protection result vector can also be an M-dimensional vector. Each dimension of the content attention result vector can represent a second attention score between an image region and the content protection information. The second attention score indicates the degree of correlation between the image region and the content protection information. The computer device can determine the business connectivity associated with the question text among the M connected components using the second attention score in the content protection result vector. Based on the business connectivity, the content protection region is determined in the image. The image content within the content protection region is the image content associated with the content protection information. The process of determining the business connectivity and the content protection region can be found in the detailed description of the target connectivity and the question-and-answer attention region in step S202, and will not be repeated here. Optionally, the computer device can determine multiple image regions whose second attention scores reach a second score threshold (a preset value) as content protection regions, thus obtaining multiple content protection regions.
[0160] Step S204: The overlapping area between the question-and-answer attention area and the content protection area is identified as the business intersection area, and the area in the content protection area other than the business intersection area is identified as the first area to be processed.
[0161] Specifically, the computer device can determine whether the question-and-answer attention region and the content protection region overlap by comparing their boundary contour points. If there is an overlapping area, the computer device can determine the overlapping contour points based on these points, and define the polygonal region formed by the overlapping contour points as the overlapping area. The computer device can then define the overlapping area between the question-and-answer attention region and the content protection region as the business intersection area, and define the area within the content protection region other than the business intersection area as the first area to be processed.
[0162] Step S205: Obtain the protection level score associated with the content protection area; the protection level score is generated based on the content protection result vector corresponding to the content protection information, which is the attention vector obtained by performing attention processing on the content protection information and the image; determine the target redraw level based on the protection level score, and determine the target content protection processing method indicated by the target redraw level from N content protection processing methods; N is a positive integer; perform content protection processing on the image content in the first area to be processed based on the target content protection processing method to obtain the first processed content.
[0163] Specifically, the computer device can perform content protection processing on the first area to be processed. Content protection processing includes methods such as content hiding and redrawing to remove the image indicated by the content protection information. Content hiding can refer to adding an opaque mask layer to the image area to hide it, or it can be directly erasing (eliminating) the image area. Redrawing processing can include image cloaking redrawing, image blurring redrawing, similar-object replacement redrawing, and feature blurring redrawing, etc., and this application embodiment does not impose limitations on these methods.
[0164] The computer device can normalize the M attention scores in the content protection result vector to obtain a normalized score for each attention score. The computer device can then determine the protection level score associated with the content protection region based on the normalized score corresponding to the attention score of the service connectivity component in the content protection result vector. The protection level score can be a value between 0 and 1, for example, 0.7. The computer device can record the protection level score by setting it as the grayscale value of the image mask of the content protection region. After obtaining the protection level score, the computer device can store it as a grayscale value in the content protection region, which serves as a mask. When determining the content protection processing method below, the computer device can determine the protection level score by extracting the grayscale value from the mask.
[0165] Computer equipment can determine the target redraw level of the content protection area based on the protection level score, and then determine the target content protection processing method indicated by the target redraw level from N content protection processing methods. Taking content protection processing as redraw processing as an example, the mapping relationship between the protection level score and the content protection processing method is shown in Table 1:
[0166] Table 1
[0167]
[0168] For example, if the protection level score is 0.1, the computer device can determine the redrawing level as T0, and not redraw the image content in the first area to be processed, maintaining the original image. If the protection level score is 0.3, the computer device can determine the redrawing level as T1, and perform feature blur redrawing on the image content in the first area to be processed. If the protection level score is 0.5, the computer device can determine the redrawing level as T2, and perform similar replacement redrawing on the image content in the first area to be processed. If the protection level score is 0.7, the computer device can determine the redrawing level as T3, and perform image blur redrawing on the image content in the first area to be processed. If the protection level score is 0.9, the computer device can determine the redrawing level as T4, and perform image cloaking redrawing on the image content in the first area to be processed. Here, the first image content refers to the image content in the first area to be processed within the image.
[0169] Please join us again. Figure 5 , Figure 5 This is a schematic diagram of a data processing scenario provided in an embodiment of this application. Figure 2 ,like Figure 5 As shown, taking image A0 and the image content in the first area to be processed as the first image content P0 as an example, the computer device can perform content protection processing on the first image content P0 based on the target content protection processing method.
[0170] When the target content protection processing method is image cloaking redraw, the process of performing content protection processing on the image content in the first area to be processed to obtain the first processed content based on the target content protection processing method can be as follows: determine the image content in the image other than the first area to be processed as the redraw source content; obtain the background image in the redraw source content, extract features from the background image to obtain the background feature vector; perform redraw processing on the image content in the first area to be processed based on the background feature vector to obtain the first processed content; the content feature vector of the first processed content has a similarity relationship with the background feature vector.
[0171] Specifically, the computer device can identify the image content in image A0, excluding the first region to be processed, as the repainting source content. The computer device can then perform background recognition on the repainting source content to obtain the background image. The computer device can input the background image into a Denoising Diffusion Probabilistic Model (DDPM) or a Latent Stable Diffusion Model (LDM). Using ControlNet (a neural network that adds conditional control to the image diffusion model) and IPAdapter (Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models, an adapter for implementing image cues), the background feature vector corresponding to the background image is determined as the conditional input for the noise map in the backdiffusion process, thereby generating the first processed content P4, which has a similar relationship to the background image. The content feature vector of the first processed content P4 has a similarity relationship to the background feature vector. This process can also be referred to as OutPainting the image content in the first region to be processed using an image diffusion model. The computer device can identify the image whose first image content P0 has been replaced with the first processed content P4 as the question-and-answer protected image A4.
[0172] Optionally, image cloaking redraw can be achieved by directly setting the pixel values of the first image content P0 to 0 or 255, making the first image content P0 invisible, thus directly achieving image cloaking. Alternatively, the first image content P0 can be replaced using background removal, a technique that hides or removes sensitive information by replacing the background in an image. This technique can effectively protect privacy while preserving the main objects or foreground information in the image.
[0173] When the target content protection processing method is image blurring and redrawing, the process of performing content protection processing on the image content in the first area to be processed to obtain the first processed content can be as follows: obtain the blur unit value P; divide the image content in the first area to be processed into P image units to be blurred based on the blur unit value P; P is a positive integer; the P image units to be blurred include the target image unit to be blurred; generate the average color channel information of the target image unit to be blurred based on the color channel information corresponding to each pixel in the target image unit to be blurred; perform blurring processing on the pixels in the target image unit to be blurred based on the average color channel information to obtain the blurred image unit; when the blurred image units corresponding to the P image units to be blurred are obtained, the P blurred image units are determined as the first processed content.
[0174] Specifically, the computer device can acquire the blur unit value P, which can be determined based on the range of the first image content P0. The larger the range of the first image content P0, the larger the blur unit value P. The computer device can divide the first image content P0 into P image units to be blurred based on the blur unit value P, and the P image units to be blurred include the target image unit to be blurred.
[0175] Image blurring and redrawing techniques can include Gaussian blur, mean blur, and bilateral filtering. Among these, image blurring for desensitization is a commonly used image processing technique primarily used to protect privacy and sensitive information. By blurring, certain parts of an image can be made unclear, thus preventing unauthorized personnel from obtaining detailed information from the image.
[0176] Taking mean blurring as an example, the computer device can generate average color channel information for the target image unit to be blurred based on the color channel information corresponding to each pixel in the target image unit to be blurred. The pixels in the target image unit to be blurred are then blurred based on the average color channel information; for example, the average color channel information can be applied to each pixel in the target image unit to be blurred, resulting in a blurred image unit. When P blurred image units corresponding to the target image units to be blurred are obtained, the computer device can determine these P blurred image units as the first processing content P3. The computer device can then replace the first image content P0 with the image of the first processing content P3, and determine this as the question-and-answer protection image A3.
[0177] Alternatively, image blurring and redrawing can also be achieved through image cartooning or color filling. Image cartooning is an image processing technique that transforms an image into a cartoon style. This technique simplifies details and textures in an image, making it resemble a hand-drawn cartoon. This technique can be used not only for artistic creation but also for privacy protection and data anonymization, as cartooning can blur or simplify sensitive information in an image, making it difficult to identify. Color filling is a technique that hides sensitive information by filling certain areas of an image with random color blocks or noise. This method not only changes the color but also introduces randomness, making the original information more difficult to recover, thus effectively protecting privacy.
[0178] When the target content protection processing method is similar replacement and redrawing, the process of performing content protection processing on the image content in the first area to be processed to obtain the first processed content can be as follows: K key object information are generated based on the content protection information; K is a positive integer, and the key object information is the key information used to identify the content protection information; attention processing is performed on the image content in the first area to be processed and the K key object information to obtain an object attention result vector; based on the object attention result vector, the key object region associated with the K key object information is determined in the first area to be processed; content prompt text associated with the content protection information is obtained; target prompt text associated with the key object region is determined in the content prompt text; features are extracted from the target prompt text to obtain a text prompt feature vector; the image content corresponding to the key object region is damaged to obtain damaged content; the image content in the image other than the first area to be processed is determined as the redraw source content; features are extracted from the redraw source content to obtain a redraw source feature vector; image restoration is performed on the damaged content based on the text prompt feature vector and the redraw source feature vector to obtain the first processed content.
[0179] Specifically, the computer device can generate K key object information based on content protection information. These key object information can be crucial for identifying the content protection information. For example, when the content protection information is "female figure," the key object information could be "eyes (iris)," "nose," "lips (lip lines)," etc. The computer device can perform attention processing on the first image content P0 and the K key object information to obtain an object attention result vector. Based on this vector, it can determine the key object regions associated with the K key object information in the first processing area. For example, it can obtain the image regions of the key object information in the first image content P0. The image content in these key object regions can be used to represent the key object information. For example, the key object regions could be image regions in the first image content P0 representing the eyes, nose, and mouth of a person.
[0180] A computer device can acquire several content prompts associated with a key region of an object. For example, if the key region is an image of a person's mouth (the key information is the mouth), the computer device can acquire prompts associated with the mouth, such as "smiling mouth" or "open mouth." The computer device can perform attention processing on the several content prompts and the key information represented by the key region to obtain a key result vector. Based on this vector, it can determine the target prompt associated with the key region from among the content prompts. For example, the target prompt could be "smiling mouth," or it could be the same as the content protection information. Optionally, the computer device can also perform image recognition on the image content of the content protection region to generate the target prompt.
[0181] Computer equipment can corrupt image content corresponding to key regions of an object (e.g., the image region of a person's mouth) to obtain corrupted content. Corruption processing can involve adding Gaussian-distributed noise to the image content corresponding to the key regions of the object.
[0182] The computer device can input the target prompt text and corrupted content into a denoising diffusion probability model or a latent diffusion model. Using ControlNet and IPAdapter, it determines the text prompt feature vector corresponding to the content prompt text as the conditional input for the corrupted content during the backward diffusion process, thereby generating first processed content P2 that has a similar relationship to the target prompt text. For example, the computer device can redraw the mouth of a person in the first image content P0 as a smiling mouth in the first processed content P2. By performing backward diffusion on the first image content P0 based on the target prompt text corresponding to multiple key areas of the object (e.g., eyes, nose, girl, etc.), the first processed content P2 is generated. This process can also be called inpainting the first image content P0 using an image diffusion model. The computer device can then identify the image where the first image content P0 has been replaced with the first processed content P2 as the question-and-answer protected image A2.
[0183] When the target content protection processing method is feature blur redrawing, the process of performing content protection processing on the image content in the first area to be processed to obtain the first processed content can be as follows: generating K object key information based on content protection information; K is a positive integer, and the object key information is the key information used to identify the content protection information; performing attention processing on the image content in the first area to be processed and the K object key information to obtain an object attention result vector; determining K object key regions associated with the K object key information in the first area to be processed based on the object attention result vector; performing image blur redrawing on the image content corresponding to the K object key regions to obtain K redrawn image contents; and determining the first image content containing the K redrawn image contents as the first processed content.
[0184] Specifically, the computer device can generate K object key information based on content protection information. These object key information can be crucial information used to identify the content protection information. For example, when the content protection information is "person / face," the object key information could be "eyes (iris)," "nose," "lips (lip lines)," etc. The computer device can perform attention processing on the first image content P0 and the K object key information to obtain an object attention result vector. Based on this vector, it can determine K object key regions associated with the K object key information in the first processing area. For example, it can obtain the image region of each object key information in the first image content P0, and the image content in each object key region can be used to represent the object key information.
[0185] A computer device can perform image blurring and redrawing on the image content corresponding to K key areas of objects, obtaining K redrawn image contents. The first image content P0, containing the K redrawn image contents, is determined as the first processed content P1. Image blurring and redrawing methods can include Gaussian blur, mean blur, bilateral filtering, etc., and are not limited to these methods in this embodiment. The computer device can then determine the image whose first image content P0 has been replaced with the first processed content P1 as the question-and-answer protection image A1.
[0186] It is understandable that computer equipment can identify one or more first regions to be processed in an image, and each first region to be processed can have a different protection level score, that is, the redrawing method of each first region to be processed can be different.
[0187] Step S206: Determine the second processing area in the question-and-answer attention area, determine the image content in the second processing area as the second image content, perform content protection processing on the second image content to obtain the second processed content; replace the first processed content with the first image content and replace the second processed content with the second image content in the image to obtain a question-and-answer protected image including the first processed content and the second processed content.
[0188] Specifically, since there is an overlap between the question-and-answer attention region and the content protection region, and the question-and-answer attention region also retains some of the image content indicated by the content protection information, in order to further improve the security of the image content indicated by the content protection information in the question-and-answer attention region, the computer device can also perform content protection processing on the image content in the question-and-answer attention region.
[0189] The computer device can obtain the first attention score corresponding to the question-and-answer attention region. If the first attention score is greater than the content protection processing threshold (a pre-set value, such as 0.2), the computer device can determine a second region to be processed within the question-and-answer attention region. If the first attention score is less than or equal to the content protection processing threshold, the computer device can directly determine the question-and-answer attention region as the second region to be processed. By determining the second region to be processed, the computer device can avoid performing content protection processing on the key image content used to identify the question text, thus further improving the security of content protection information while still ensuring the accuracy of image-based question answering.
[0190] The computer device can determine the second region to be processed in the question-and-answer attention region. The process can be as follows: generate T key information points about the question based on the question text; T is a positive integer, and the key information points are the key information used to identify the question text; perform attention processing on the image content in the question-and-answer attention region and the T key information points about the question to obtain a question attention result vector; determine the key question region associated with the T key information points about the question based on the question attention result vector; and determine the region in the question-and-answer attention region other than the key question region as the second region to be processed.
[0191] Specifically, the computer device can generate T key information pieces of the question based on the question text. These key information pieces can be crucial information used to identify the question text. They can be tokens or synonyms of those tokens, or text used to identify those tokens. The key information can represent the key features in the image that correspond to the question text. For example, if the question text is "What type of car is in the image?", then the key information could be "car", or text identifying "car", which could include "wheels", "doors", "car logo", etc.
[0192] The computer device can perform attention processing on the image content and T key question information within the question-and-answer attention region, obtaining a question attention result vector. Based on this vector, it identifies key question regions associated with the T key question information within the question-and-answer attention region. The image content within each key question region can represent the question text. The computer device can designate areas outside the key question regions within the question-and-answer attention region as a second region to be processed. The computer device can obtain the question-and-answer attention result vector corresponding to the question-and-answer attention region. Based on the first attention score associated with the target connected component in the question-and-answer attention result vector, it generates a relevance score for the question-and-answer attention region. The relevance score can be a value between 0 and 1, for example, 0.8. The computer device can record the relevance score by setting it to the grayscale value of the image mask of the question-and-answer attention region.
[0193] The computer device can determine the protection level score of the question-and-answer attention region based on the relevance score. The protection level score of the question-and-answer attention region can be the difference between the value 1 and the relevance score of the question-and-answer attention region. For example, if the relevance score of the question-and-answer attention region is 0.8, then the protection level score of the question-and-answer attention region can be 0.2. The content protection processing methods corresponding to each protection level score can be found in Table 1 in step S205 above, and will not be repeated here.
[0194] The computer device can determine the target content protection processing method for the question-and-answer attention region based on the protection level score of the question-and-answer attention region, identify the image content in the second region to be processed as the second image content, and perform content protection processing on the second image content based on the target content protection processing method to obtain the second processed content. The details of the target content protection processing method can be found in the specific description of step S205 above, and will not be repeated here. The pixels redrawn in the second processed content may include pixels of the image content associated with the content protection information.
[0195] The computer device can replace the first processed content with the first image content and replace the second processed content with the second image content in the image to obtain a question-and-answer protected image that includes the first processed content and the second processed content.
[0196] It is understandable that, since both the question-and-answer attention region and the content protection region can be image masks, the first and second regions to be processed can also be image masks. Image masks can be used to hide or show specific parts of an image. An image mask can be a binary image of the same size as the original image (i.e., each pixel has only two values, 0 and 1), where 0 represents transparent (or hidden) and 1 represents opaque (or displayed). By combining the image mask with the original image, it is possible to control which parts of the image are visible and which are hidden. Since an image mask can contain the image vertex coordinates representing a region, a computer device can determine the location of specific image content (such as the image content in the question-and-answer attention region or the image content in the content protection region) using the image vertex coordinates in the image mask. Based on the target content protection processing method corresponding to the content protection region and the location of the second region to be processed, the image content in the second region to be processed is locally redrawn; similarly, based on the target content protection processing method corresponding to the question-and-answer attention region and the location of the question-and-answer attention region, the image content in the question-and-answer attention region is locally redrawn, resulting in a question-and-answer protected image. The question-and-answer protected image can be used to generate the answer text corresponding to the question text. Local redrawing refers to the process of applying content protection to the image content in a first region to be processed using image masking. The generated new image content directly overwrites or replaces the old image content, resulting in a question-and-answer protected image containing the first processed content. Similarly, content protection is applied to the image content in a second region to be processed, and the generated new image content directly overwrites or replaces the old image content, resulting in a question-and-answer protected image containing both the first and second processed image content. Local redrawing is an image processing method used to modify specific areas of an image without affecting other parts. It involves selecting the area to be modified, generating new content to replace the old content, and seamlessly integrating the new content with the original image. This method is commonly used for image inpainting, object removal, and region enhancement, generating new content through copy-paste, texture synthesis, or deep learning models, and using hybrid algorithms to ensure natural transitions.
[0197] This application's embodiments, by determining a question-and-answer attention region and a content protection region, can distinguish the image content required for the model to process image question-and-answer. Then, based on the question-and-answer attention region, a first region to be processed is determined within the content protection region. Content protection processing is then applied to the first image content within this first region, reducing the possibility of image content leakage related to content protection information during the model's learning of images and answering question text. Furthermore, since the first region to be processed and the question-and-answer attention region do not overlap, the model can still identify image content associated with the question text through the question-and-answer attention region. Therefore, this application can improve the accuracy of image question-and-answer while ensuring the security of content protection information. Determining a protection level score through the content protection region, and then using this score to determine different content protection processing methods for the first region to be processed, can increase the diversity of content protection processing and reduce the risk of content protection information leakage.
[0198] On the other hand, the embodiments of this application can further improve the security of image content in the question-and-answer attention region, that is, content protection processing is also performed on the image content in the question-and-answer attention region. By determining the relevance score through the question-and-answer attention region, and by determining different content protection processing methods for the second region to be processed based on the relevance score, different desensitization strategies can be adopted for the image content in the first region to be processed, following the principle of "minimum necessary information exposure", so that while protecting content protection information, the accuracy of image question-and-answer is minimized.
[0199] Please see Figure 6 , Figure 6 This is a schematic diagram of a data processing scenario provided in an embodiment of this application. Figure 3 .like Figure 6 As shown, the user can upload an image and corresponding question text to the terminal device 100 via an application client. The question text can be text obtained through audio recognition of the user's input or text directly entered by the user in the application client; this embodiment does not impose any limitations. For example, the question text could be "the type of car in the image".
[0200] Terminal device 100 can obtain content protection information, which can be descriptive information about the privacy content in an image. This content protection information can be a type description or a directly specified description. It can be selected and configured in the privacy content type in the application client, input by the user, or automatically identified by the model. Content protection information can be text data, image data, or audio data. Types of content protection information can include personal identification information, contact information, financial information, vehicle information, medical information, and address location information. For example, content protection information can be "person / image".
[0201] Optionally, when the content protection information is image data, it can be an image containing an image frame indicating privacy content. That is, when uploading an image, the user can select the privacy content that needs to be protected by using the image frame to obtain the content protection information.
[0202] Terminal device 100 may deploy edge-side multimodal models. These edge-side multimodal models can refer to artificial intelligence models that run on terminal device 100 and are capable of simultaneously processing multiple modalities of data (such as text, images, and audio). Examples include the Phi-3-mini-128k-instruct model (an open-source lightweight model), the Chat-GLM (General Language Model) model (a natural language generation model), and the Qwen-VLM (Vision Language Model) model (a large-scale visual language model), among other multimodal large language models that can run on edge devices.
[0203] Terminal device 100 can perform attention processing on the image and question text using an edge-side multimodal model, and determine the question-answer attention region associated with the question text in the image using the resulting question-answer attention result vector. Similarly, terminal device 100 can perform attention processing on the image and content protection information using an edge-side multimodal model, and determine the content protection region associated with the content protection information in the image using the resulting content protection result vector. Both the question-answer attention region and the content protection region can be an image mask described by polygon vertices.
[0204] If there is an overlap between the question-and-answer attention region and the content protection region, the terminal device 100 can identify the overlapping region as a business intersection region, and identify the region in the content protection region other than the business intersection region as the first region to be processed. The terminal device 100 further protects the image content associated with the content protection information in the question-and-answer attention region, and can identify this content as the second region to be processed within the question-and-answer attention region.
[0205] Terminal device 100 can determine the protection level score through the content protection result vector corresponding to the content protection area, and determine the target content protection processing method for the first area to be processed through the protection level score of the content protection area. Terminal device 100 can perform content protection processing on the first area to be processed through the target content protection processing method of the content protection area to obtain the first processed content.
[0206] Terminal device 100 can determine the relevance score through the question-and-answer attention result vector corresponding to the question-and-answer attention region, and determine the content protection processing method for the second region to be processed through the relevance score of the question-and-answer attention region. Terminal device 100 can perform content protection processing on the second region to be processed through the target content protection processing method of the question-and-answer attention region to obtain the second processed content. The content protection processing method may include image cloaking redraw, image blurring redraw, similar replacement redraw, and feature blurring redraw, etc., which are not limited in this embodiment.
[0207] Terminal device 100 can identify an image whose first image content has been replaced with first processed content and whose second image content has been replaced with second processed content as a question-and-answer protected image. The question-and-answer protected image and the question text are sent to business server 200. The business server uses a cloud-based multimodal model to identify the question-and-answer protected image, generates the answer text corresponding to the question text, and returns the answer text to terminal device 100.
[0208] This application's embodiments, by determining the question-answering attention region and the content protection region, can distinguish the image content required by the model to process image question-answering. Then, based on the question-answering attention region, a first region to be processed is determined within the content protection region. Content protection processing is then applied to the first image content within this first region, reducing the possibility of image content leakage related to content protection information during the model's learning of images and answering question text. Simultaneously, since the first region to be processed and the question-answering attention region do not overlap, the model can still identify image content associated with the question text through the question-answering attention region. Therefore, this application can appropriately improve the security of content protection information while ensuring the accuracy of image question-answering. Determining the protection level score through the content protection region, and then determining different content protection processing methods for the first region to be processed based on the protection level score, can increase the diversity of content protection processing and reduce the risk of content protection information leakage.
[0209] On the other hand, the embodiments of this application can further improve the security of image content in the question-and-answer attention region, that is, content protection processing is also performed on the image content in the question-and-answer attention region. By determining the relevance score through the question-and-answer attention region, and by determining different content protection processing methods for the second region to be processed based on the relevance score, different desensitization strategies can be adopted for the image content in the first region to be processed, following the principle of "minimum necessary information exposure", so that while protecting content protection information, the accuracy of image question-and-answer is minimized.
[0210] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Figure 7 As shown, the data processing device 1 includes a question-and-answer area processing module 710, a protected area processing module 720, and a redrawing area processing module 730.
[0211] The question-and-answer region processing module 710 is used to determine the question-and-answer attention region in the image based on the image and the corresponding question text; the image content in the question-and-answer attention region is the image content in the image associated with the question text.
[0212] The protected area processing module 720 is used to acquire content protection information, and determine the content protection area in the image based on the image and the content protection information; the image content in the content protection area is the image content in the image associated with the content protection information.
[0213] The redrawing region processing module 730 is used to determine a first region to be processed in the content protection region based on the question-and-answer attention region if there is an overlapping region between the question-and-answer attention region and the content protection region; the first region to be processed does not overlap with the question-and-answer attention region; wherein, the first region to be processed is used to indicate that content protection processing is performed on the first image content in the image, and the first image content is the image content in the first region to be processed in the image.
[0214] In one possible implementation, when the redrawing region processing module 730 determines the first region to be processed within the content protection region based on the question-answering attention region, it specifically performs the following operations:
[0215] The overlapping area between the question-and-answer attention area and the content protection area is defined as the business intersection area, and the area within the content protection area other than the business intersection area is defined as the first area to be processed.
[0216] In one possible implementation, the redraw area processing module 730 is also used to perform the following operations:
[0217] Content protection processing is performed on the image content in the first area to be processed to obtain the first processed content. The first image content is then replaced in the image with the first processed content to obtain a question-and-answer protected image that includes the first processed content. The question-and-answer protected image is used to generate the answer text corresponding to the question text.
[0218] In one possible implementation, the redraw area processing module 730 is also used to perform the following operations:
[0219] The image content in the first area to be processed is subjected to content protection processing to obtain the first processed content;
[0220] In the question-and-answer attention area, a second region to be processed is determined, the image content in the second region to be processed is determined as the second image content, and content protection processing is performed on the second image content to obtain the second processed content;
[0221] In the image, the first processed content is replaced by the first image content, and the second processed content is replaced by the second image content, to obtain a question-and-answer protected image that includes the first processed content and the second processed content.
[0222] In one possible implementation, when the redrawing region processing module 730 determines the second region to be processed in the question-and-answer attention region, it is specifically used to perform the following operations:
[0223] T key information points are generated from the question text; T is a positive integer, and the key information points are the key information used to identify the obtained question text.
[0224] Attention processing is performed on the image content in the question-and-answer attention region and T key information of the questions to obtain a question attention result vector. Based on the question attention result vector, the question key region associated with the T key information of the questions is determined in the question-and-answer attention region.
[0225] The area outside the key question area in the question-and-answer attention region is identified as the second area to be processed.
[0226] In one possible implementation, when the question-and-answer region processing module 710 determines the question-and-answer attention region in the image based on the image and the corresponding question text, it specifically performs the following operations:
[0227] Each pixel in the image is labeled with a connected component, resulting in M connected components; M is a positive integer. A connected component is used to represent an image region in which the pixel values of pixels satisfy the connected component labeling conditions and the pixels are adjacent to each other.
[0228] Feature extraction is performed on each of the M connected components to obtain the image feature vector, and feature extraction is performed on the corresponding question text to obtain the question text feature vector;
[0229] Attention processing is performed on the image feature vector and the question text feature vector to obtain the question-answering attention result vector. The question-answering attention region is then determined in the image based on the question-answering attention result vector.
[0230] In one possible implementation, when the question-answering region processing module 710 determines the question-answering attention region in the image based on the question-answering attention result vector, it specifically performs the following operations:
[0231] Generate attention scores for M connected components based on the question-answering attention result vector;
[0232] The connected component with the highest attention score is determined as the target connected component. Boundary contour points are generated based on the position information of pixels in the target connected component. The region enclosed by the boundary contour points in the image is determined as the question-answering attention region.
[0233] In one possible implementation, the redrawing area processing module 730 is used to perform content protection processing on the image content in the first area to be processed. When the first processed content is obtained, it is specifically used to perform the following operations:
[0234] Obtain the protection level score associated with the content protection area; the protection level score is generated based on the content protection result vector corresponding to the content protection information, and the content protection result vector is the attention vector obtained by performing attention processing on the content protection information and the image;
[0235] The target repainting level is determined based on the protection level score, and the target content protection processing method indicated by the target repainting level is determined from N content protection processing methods; N is a positive integer.
[0236] Based on the target content protection processing method, the image content in the first area to be processed is subjected to content protection processing to obtain the first processed content.
[0237] In one possible implementation, the target content protection processing method is image cloaking redraw; the redrawing area processing module 730 is used to perform content protection processing on the image content in the first area to be processed based on the target content protection processing method, and when the first processed content is obtained, it is specifically used to perform the following operations:
[0238] The image content excluding the first region to be processed is identified as the redraw source content;
[0239] Obtain the background image from the redraw source content, extract features from the background image, and obtain the background feature vector;
[0240] Based on the background feature vector, the image content in the first region to be processed is redrawn to obtain the first processed content; the content feature vector of the first processed content has a similarity relationship with the background feature vector.
[0241] In one possible implementation, the target content protection processing method is image blurring and redrawing; the redrawing area processing module 730 is used to perform content protection processing on the image content in the first area to be processed based on the target content protection processing method, and when the first processed content is obtained, it is specifically used to perform the following operations:
[0242] Obtain the blur unit value P, and divide the image content in the first region to be processed into P image units to be blurred based on the blur unit value P; P is a positive integer; the P image units to be blurred include the target image unit to be blurred.
[0243] Based on the color channel information corresponding to each pixel in the target image unit to be blurred, the average color channel information of the target image unit to be blurred is generated. Based on the average color channel information, the pixels in the target image unit to be blurred are blurred to obtain the blurred image unit.
[0244] When P blurred image units corresponding to the image units to be blurred are obtained, the P blurred image units are determined as the first processing content.
[0245] In one possible implementation, the target content protection processing method is similar replacement and redrawing; the redrawing area processing module 730 is used to perform content protection processing on the image content in the first area to be processed based on the target content protection processing method, and when the first processed content is obtained, it is specifically used to perform the following operations:
[0246] K key information items for objects are generated based on content protection information; K is a positive integer, and the key information items for objects are the key information used to identify the content protection information.
[0247] Attention processing is performed on the image content and K key object information in the first region to be processed to obtain the object attention result vector. Based on the object attention result vector, the key object regions associated with the K key object information are determined in the first region to be processed.
[0248] Obtain the content prompt text associated with content protection information, identify the target prompt text associated with the key area of the object in the content prompt text, extract features from the target prompt text, and obtain the text prompt feature vector;
[0249] The image content corresponding to the key areas of the object is damaged to obtain the damaged content. The image content other than the first area to be processed is determined as the redraw source content. The redraw source content is feature extracted to obtain the redraw source feature vector.
[0250] Image restoration of damaged content is performed based on text prompt feature vectors and redraw source feature vectors to obtain the first processed content.
[0251] In one possible implementation, the target content protection processing method is feature blurring and redrawing; the redrawing region processing module 730 is used to perform content protection processing on the image content in the first region to be processed based on the target content protection processing method, and when the first processed content is obtained, it is specifically used to perform the following operations:
[0252] K key information items for objects are generated based on content protection information; K is a positive integer, and the key information items for objects are the key information used to identify the content protection information.
[0253] Attention processing is performed on the image content in the first region to be processed and K key information of objects to obtain an object attention result vector. Based on the object attention result vector, the key regions of objects associated with the K key information of objects are determined in the first region to be processed. The image content corresponding to the key regions of objects is blurred and redrawn to obtain the first processed content.
[0254] This application's embodiments, by determining the question-answering attention region and the content protection region, can distinguish the image content required by the model to process image question-answering. Then, based on the question-answering attention region, a first region to be processed is determined within the content protection region. Content protection processing is then applied to the first image content within this first region, reducing the possibility of image content leakage related to content protection information during the model's learning of images and answering question text. Simultaneously, since the first region to be processed and the question-answering attention region do not overlap, the model can still identify image content associated with the question text through the question-answering attention region. Therefore, this application can appropriately improve the security of content protection information while ensuring the accuracy of image question-answering. Determining the protection level score through the content protection region, and then determining different content protection processing methods for the first region to be processed based on the protection level score, can increase the diversity of content protection processing and reduce the risk of content protection information leakage.
[0255] On the other hand, the embodiments of this application can further improve the security of image content in the question-and-answer attention region, that is, content protection processing is also performed on the image content in the question-and-answer attention region. By determining the relevance score through the question-and-answer attention region, and by determining different content protection processing methods for the second region to be processed based on the relevance score, different desensitization strategies can be adopted for the image content in the first region to be processed, following the principle of "minimum necessary information exposure", so that while protecting content protection information, the accuracy of image question-and-answer is minimized.
[0256] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0257] Please see Figure 8 , Figure 8This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 8 As shown, the computer device 1000 may include a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the computer device 1000 may also include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the processor 1001. Figure 8 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.
[0258] In such Figure 8 In the computer device 1000 shown, the network interface 1004 provides network communication elements; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:
[0259] Based on the image and the corresponding question text, the question-and-answer attention region is determined in the image; the image content in the question-and-answer attention region is the image content in the image that is associated with the question text.
[0260] Obtain content protection information, and based on the image and the content protection information, determine the content protection area in the image; the image content in the content protection area is the image content associated with the content protection information in the image;
[0261] If there is an overlapping area between the question-and-answer attention area and the content protection area, then the first area to be processed is determined in the content protection area based on the question-and-answer attention area; the first area to be processed and the question-and-answer attention area do not overlap.
[0262] The first area to be processed is used to indicate that content protection processing is performed on the first image content in the image, and the first image content is the image content in the first area to be processed in the image.
[0263] It should be understood that the computer device 1000 described in the embodiments of this application can execute the foregoing text. Figure 3 and Figure 4 The description of the data processing method in any corresponding embodiment will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.
[0264] Furthermore, it should be noted that this application embodiment also provides a computer-readable storage medium, which stores a computer program. When the processor executes the computer program, it can execute the aforementioned... Figure 3 and Figure 4 The description of the data processing method in any corresponding embodiment is already provided, and therefore will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments related to this application, please refer to the description of the method embodiments of this application.
[0265] The aforementioned computer-readable storage medium can be an internal storage unit of the data processing apparatus or computer device provided in any of the foregoing embodiments, such as a hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been displayed or will be displayed.
[0266] Furthermore, it should be noted that this application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the aforementioned... Figure 3 and Figure 4 The method provided in any of the corresponding embodiments.
[0267] The terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0268] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the foregoing description as a network element. Whether these network elements are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can implement the described network elements using different methods for each specific application, but such implementation should not be considered beyond the scope of this application.
[0269] The methods and related apparatus provided in this application are described with reference to the method flowcharts and / or structural diagrams provided in this application. Specifically, each block of the method flowchart and / or structural diagram, as well as combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable device to create a machine, such that the instructions, which execute via the processor of the computer or other programmable device, generate instructions for implementing the process. Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable device to cause a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 A process or multiple processes and / or structures illustrate the steps of the functions specified in one or more boxes.
[0270] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.
[0271] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.
[0272] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A data processing method, characterized in that, include: Based on the image and the corresponding question text, a question-and-answer attention region is determined in the image; The image content in the question-and-answer attention area is the image content in the image that is associated with the question text; Obtain content protection information, and determine a content protection region in the image based on the image and the content protection information; The image content in the content protection area is the image content associated with the content protection information in the image; If there is an overlapping area between the question-and-answer attention area and the content protection area, then a first area to be processed is determined in the content protection area based on the question-and-answer attention area; the first area to be processed does not overlap with the question-and-answer attention area. Wherein, the first area to be processed is used to indicate that content protection processing is performed on the first image content in the image, and the first image content is the image content in the first area to be processed in the image.
2. The method according to claim 1, characterized in that, The step of determining the first region to be processed within the content protection region based on the question-and-answer attention region includes: The overlapping area between the question-and-answer attention area and the content protection area is defined as the business intersection area, and the area in the content protection area other than the business intersection area is defined as the first area to be processed.
3. The method according to claim 1, characterized in that, Also includes: Content protection processing is performed on the image content in the first area to be processed to obtain first processed content. The first image content is then replaced in the image with the first processed content to obtain a question-and-answer protected image including the first processed content. The question-and-answer protected image is used to generate the answer text corresponding to the question text.
4. The method according to claim 1, characterized in that, Also includes: The image content in the first area to be processed is subjected to content protection processing to obtain the first processed content; A second region to be processed is determined in the question-and-answer attention region. The image content in the second region to be processed is determined as the second image content. Content protection processing is performed on the second image content to obtain the second processed content. In the image, the first processed content is replaced by the first image content, and the second processed content is replaced by the second image content, to obtain a question-and-answer protection image that includes the first processed content and the second processed content.
5. The method according to claim 4, characterized in that, The step of determining the second region to be processed in the question-and-answer attention region includes: T key information pieces of the problem are generated based on the problem text; T is a positive integer, and the key information of the problem is the key information used to identify the problem text. Attention processing is performed on the image content in the question-and-answer attention region and the T key information of the questions to obtain a question attention result vector. Based on the question attention result vector, the question key region associated with the T key information of the questions is determined in the question-and-answer attention region. The area outside the question-and-answer attention region, excluding the key question region, is designated as the second region to be processed.
6. The method according to claim 1, characterized in that, The step of determining the question-and-answer attention region in the image based on the image and the corresponding question text includes: Each pixel in the image is labeled with a connected component, resulting in M connected components; M is a positive integer. The connected components are used to represent image regions in the image where the pixel values of the pixels satisfy the connected component labeling conditions and the pixel positions are adjacent. Feature extraction is performed on the M connected components to obtain image feature vectors, and feature extraction is performed on the question text corresponding to the image to obtain question text feature vectors; Attention processing is performed on the image feature vector and the question text feature vector to obtain a question-and-answer attention result vector. Based on the question-and-answer attention result vector, a question-and-answer attention region is determined in the image.
7. The method according to claim 6, characterized in that, Determining the question-answering attention region in the image based on the question-answering attention result vector includes: Based on the question-answering attention result vector, generate attention scores corresponding to the M connected components respectively; The connected component with the highest attention score is determined as the target connected component. Boundary contour points are generated based on the position information of the pixels in the target connected component. The region enclosed by the boundary contour points in the image is determined as the question-answering attention region.
8. The method according to claim 3, characterized in that, The step of performing content protection processing on the image content in the first area to be processed to obtain the first processed content includes: Obtain the protection level score associated with the content protection area; the protection level score is generated based on the content protection result vector corresponding to the content protection information, and the content protection result vector is an attention vector obtained by performing attention processing on the content protection information and the image; The target redrawing level is determined based on the protection level score, and the target content protection processing method indicated by the target redrawing level is determined from N content protection processing methods; N is a positive integer; Based on the target content protection processing method, the image content in the first area to be processed is subjected to content protection processing to obtain the first processed content.
9. The method according to claim 8, characterized in that, The target content protection processing method is image cloaking redraw; based on the target content protection processing method, the image content in the first area to be processed is subjected to content protection processing to obtain the first processed content, including: The image content excluding the first region to be processed in the image is determined as the redraw source content; Obtain the background image from the redraw source content, extract features from the background image, and obtain a background feature vector; Based on the background feature vector, the image content in the first region to be processed is redrawn to obtain the first processed content; the content feature vector of the first processed content has a similarity relationship with the background feature vector.
10. The method according to claim 8, characterized in that, The target content protection processing method is image blurring and redrawing; the step of performing content protection processing on the image content in the first area to be processed based on the target content protection processing method to obtain the first processed content includes: Obtain the blur unit value P, and divide the image content in the first region to be processed into P image units to be blurred based on the blur unit value P; P is a positive integer; the P image units to be blurred include the target image unit to be blurred. Based on the color channel information corresponding to each pixel in the target image unit to be blurred, the average color channel information of the target image unit to be blurred is generated, and the pixels in the target image unit to be blurred are blurred based on the average color channel information to obtain the blurred image unit. When the blurred image units corresponding to the P image units to be blurred are obtained, the P blurred image units are determined as the first processing content.
11. The method according to claim 8, characterized in that, The target content protection processing method is similar replacement and redrawing; based on the target content protection processing method, the image content in the first area to be processed is subjected to content protection processing to obtain the first processed content, including: K key information items for objects are generated based on the content protection information; K is a positive integer, and the key information items for objects are key information used to identify the content protection information. Attention processing is performed on the image content in the first region to be processed and the K key information of objects to obtain an object attention result vector. Based on the object attention result vector, the key regions of objects associated with the K key information of objects are determined in the first region to be processed. Obtain the content prompt text associated with the content protection information, determine the target prompt text associated with the key area of the object in the content prompt text, perform feature extraction on the target prompt text, and obtain the text prompt feature vector; The image content corresponding to the key areas of the object is damaged to obtain damaged content. The image content other than the first area to be processed is determined as the redraw source content. The redraw source content is feature extracted to obtain the redraw source feature vector. Based on the text prompt feature vector and the redraw source feature vector, the damaged content is repaired to obtain the first processed content.
12. The method according to claim 8, characterized in that, The target content protection processing method is feature blurring and redrawing; based on the target content protection processing method, the image content in the first area to be processed is subjected to content protection processing to obtain the first processed content, including: K key information items for objects are generated based on the content protection information; K is a positive integer, and the key information items for objects are key information used to identify the content protection information. Attention processing is performed on the image content in the first region to be processed and the K key information of objects to obtain an object attention result vector. Based on the object attention result vector, the key regions of objects associated with the K key information of objects are determined in the first region to be processed. The image content corresponding to the key regions of objects is blurred and redrawn to obtain the first processed content.
13. A data processing apparatus, characterized in that, include: The question-and-answer region processing module is used to determine the question-and-answer attention region in the image based on the image and the question text corresponding to the image. The image content in the question-and-answer attention area is the image content in the image that is associated with the question text; The protected area processing module is used to acquire content protection information and, based on the image and the content protection information, determine a content protection area in the image. The image content in the content protection area is the image content associated with the content protection information in the image; The redrawing region processing module is used to determine a first region to be processed in the content protection region based on the question-and-answer attention region if there is an overlapping region between the question-and-answer attention region and the content protection region; the first region to be processed does not overlap with the question-and-answer attention region; wherein, the first region to be processed is used to indicate content protection processing for a first image content in the image, and the first image content is the image content in the first region to be processed in the image.
14. A computer device, characterized in that, include: Processor, memory, and network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide data communication functions, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device executes the method according to any one of claims 1-12.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-12.
16. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium and adapted to be read and executed by a processor so that a computer device having the processor performs the method of any one of claims 1-12.