Sensitive word filtering method and device for multi-modal data, equipment and medium
By performing text conversion and desensitization on multimodal data, the shortcomings of sensitive word filtering in multimodal data are solved, and effective supervision of multimedia data is achieved.
Patent Information
- Application Number
- CN202510390930.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art lacks sensitive word filtering methods for multimodal data, and cannot effectively prevent the spread of bad information.
By acquiring multimodal data, the non-text data is converted into text data by using preset conversion method, and the sensitive vocabulary and preset desensitization algorithm are used for desensitization processing, so as to achieve sensitive information filtering of multimodal data.
It effectively reduces the dissemination of sensitive information and improves the supervision of multimedia data.
Smart Images

Figure CN120337278A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sensitive word filtering, and in particular, to a method, device, equipment and medium for filtering sensitive words in multimodal data. Background Art
[0002] Multimodal data refers to the combination of various different types of data information, and the data types involved include but are not limited to text, images, videos, audio, etc. As an important technical means in the field of network security for identifying and filtering specific information, the purpose of sensitive word filtering technology is to prevent the spread of bad information. The prior art lacks a method for filtering sensitive words in multimodal data. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to overcome the deficiencies in the prior art and provide a method, device, equipment and medium for filtering sensitive words in multimodal data. The present invention provides the following technical solutions:
[0004] In a first aspect, the present invention provides a method for filtering sensitive words in multimodal data, the method comprising: obtaining multimodal data, the multimodal data including: text data and various non-text data; respectively performing text conversion on each of the non-text data by using a preset conversion method to obtain target text data; respectively determining the text data and each of the target text data as data to be desensitized; obtaining a sensitive word library, and based on the sensitive word library, respectively performing desensitization processing on each of the data to be desensitized by using a preset desensitization algorithm to obtain corresponding target data.
[0005] In an embodiment, the various non-text data include: image data, the image data including: at least one image to be recognized, and respectively performing text conversion on each of the non-text data by using a preset conversion method to obtain target text data, including: respectively performing character recognition on each of the images to be recognized by using optical character recognition technology to obtain corresponding text to be desensitized; determining each of the text to be desensitized as the target text data.
[0006] In an embodiment, the various non-text data include: audio data, the audio data including: at least one audio file, and respectively performing text conversion on each of the non-text data by using a preset conversion method to obtain target text data, including: respectively performing audio recognition on each of the audio files by using speech recognition technology to obtain corresponding text to be desensitized; determining each of the text to be desensitized as the target text data.
[0007] In one embodiment, the multiple types of the non-text data include: video data, and the video data includes: at least one video file. Performing text conversion processing on the video file by using a preset conversion method to obtain target text data includes: invoking a preset conversion tool to convert the video file into a corresponding audio file and multiple images to be recognized; respectively performing character recognition on each of the images to be recognized by using an optical character recognition technology to obtain corresponding text to be desensitized; performing audio recognition on the audio file by using a speech recognition technology to obtain the corresponding text to be desensitized; and determining each of the texts to be desensitized as the target text data.
[0008] In one embodiment, before performing text conversion on each of the non-text data by using the preset conversion method, it further includes: respectively extracting features from each of the images to be recognized by using a preset feature extraction algorithm to obtain corresponding image feature vectors; respectively inputting each of the image feature vectors into a preset image recognition model to obtain corresponding feature labels.
[0009] In one embodiment, the sensitive word library includes: multiple sensitive word sub-libraries of different categories. Before performing desensitization processing on each of the data to be desensitized based on the sensitive word library by using a preset desensitization algorithm, it includes: determining a target sensitive word sub-library from multiple sensitive word sub-libraries of different categories according to the corresponding feature labels;
[0010] Performing desensitization processing on each of the data to be desensitized based on the sensitive word library by using a preset desensitization algorithm includes: performing desensitization processing on each of the data to be desensitized based on the target sensitive word sub-library by using the preset desensitization algorithm.
[0011] In one embodiment, the method further includes: respectively updating the text data and each of the non-text data according to the corresponding target data to obtain target multi-modal data.
[0012] In a second aspect, the present invention provides a sensitive word filtering device for multi-modal data, and the device includes:
[0013] A data acquisition module, configured to acquire multi-modal data, where the multi-modal data includes: text data and multiple types of non-text data;
[0014] A text conversion module, configured to perform text conversion on each of the non-text data by using a preset conversion method to obtain target text data;
[0015] A determination module, configured to respectively determine each of the target text data and the text data as data to be desensitized;
[0016] A desensitization module, configured to obtain a sensitive word library, and based on the sensitive word library, perform desensitization processing on each piece of data to be desensitized respectively by using a preset desensitization algorithm, so as to obtain corresponding target data.
[0017] In a third aspect, the present invention provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the computer program runs on the processor, it executes the sensitive word filtering method for multimodal data described in the first aspect.
[0018] In a fourth aspect, the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the sensitive word filtering method for multimodal data described in the first aspect.
[0019] The sensitive word filtering method, device, equipment and medium for multimodal data provided by the present invention obtain multimodal data, where the multimodal data includes: text data and multiple types of non-text data; perform text conversion on each piece of non-text data respectively by using a preset conversion method to obtain target text data; determine each piece of target text data and the text data as data to be desensitized respectively; obtain a sensitive word library, and based on the sensitive word library, perform desensitization processing on each piece of data to be desensitized respectively by using a preset desensitization algorithm to obtain corresponding target data, realizing the desensitization of multimodal data, effectively reducing the spread of sensitive information, and enhancing the supervision intensity of multimedia data.
[0020] To make the above objects, features and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes the following detailed description. Description of the Drawings
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 Shows a flowchart of the sensitive word filtering method for multimodal data provided by an embodiment of the present invention;
[0023] Figure 2 Shows another flowchart of the sensitive word filtering method for multimodal data provided by an embodiment of the present invention;
[0024] Figure 3 Shows a structural diagram of the sensitive word filtering device for multimodal data provided by an embodiment of the present invention;
[0025] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present invention is shown.
[0026] Description of main component symbols:
[0027] 300 - Sensitive word filtering device for multi-modal data; 310 - Data acquisition module; 320 - Text conversion module; 330 - Determination module; 340 - Desensitization module; 400 - Electronic device; 401 - Transceiver; 402 - Processor; 403 - Memory. Specific embodiments
[0028] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0029] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the description of the template herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0031] Embodiment 1
[0032] Multi-modal data is a type of information that combines multiple different types of data, usually including but not limited to: text data, image data, video data, and audio data. With the development of artificial intelligence technology, the applications of multi-modal learning and retrieval augmented generation technology are becoming more and more widespread. These technologies combine information in multiple modalities such as images, texts, and sounds, improving the machine's understanding and processing ability of complex information, and at the same time bringing new challenges to sensitive word filtering. In this regard, please refer to Figure 1 , an embodiment of the present invention provides a method for filtering sensitive words in multi-modal data, and the method includes: steps S110 to S140.
[0033] Step S110, acquire multi-modal data, where the multi-modal data includes: text data and multiple types of non-text data.
[0034] It can be understood that multi-modal data combines multiple different types of data, among which, the multiple different types of data include: text data, image data, audio data, and / or video data, etc. In this embodiment, non-text data includes, but is not limited to: image data, audio data, and / or video data.
[0035] Step S120, respectively perform text conversion on each of the non-text data by using a preset conversion method to obtain target text data.
[0036] In this embodiment, the target text data includes at least one text to be desensitized. For non-text data of different data types, through corresponding preset conversion methods, respectively convert to obtain corresponding at least one text to be desensitized, and each of the corresponding texts to be desensitized is determined as the target text data obtained after the text conversion process of the non-text data.
[0037] In one embodiment, the multiple non-text data includes: image data, and the image data includes: at least one image to be recognized. The step of respectively performing text conversion on each of the non-text data by using a preset conversion method to obtain target text data includes: respectively performing character recognition on each of the images to be recognized by using optical character recognition technology to obtain corresponding texts to be desensitized; determining each of the texts to be desensitized as the target text data.
[0038] It can be understood that the images to be recognized may incorporate text information. To avoid the spread of sensitive content in the text information incorporated in the images to be recognized, it is necessary to filter sensitive words for each of the images to be recognized respectively. Specifically, it includes respectively performing optical character recognition on each of the images to be recognized by using optical character recognition technology (OCR, Optical Character Recognition), and converting each of the recognized characters into an editable text format to obtain corresponding texts to be desensitized, and determining each of the texts to be desensitized as the target text data of the corresponding image data. Further, the target text data is determined as the object to be desensitized in the subsequent desensitization process, that is, the data to be desensitized.
[0039] In some embodiments, to ensure the accuracy of character recognition, before performing character recognition on the images to be recognized by using optical character recognition technology, it further includes: preprocessing the images to be recognized to improve the image quality. Specifically, it includes: image denoising processing, image enhancement processing, and / or image size adjustment processing, etc.
[0040] In one embodiment, the multiple non-text data includes: audio data, and the audio data includes: at least one audio file. The method of performing text conversion on each non-text data respectively by using a preset conversion method to obtain target text data includes: performing audio recognition on each audio file respectively by using speech recognition technology to obtain corresponding text to be desensitized; and determining each text to be desensitized as the target text data.
[0041] It can be understood that in some scenarios, sensitive content is spread in the form of audio. Therefore, for the audio data in multimodal data, desensitization processing is also required. Specifically, it includes: performing audio recognition on each audio file in the audio data respectively by using speech recognition technology, converting the audio file into a text form to obtain text to be desensitized. Among them, each audio file can be in the formats of MP3, WAV, AAC, etc.
[0042] In this embodiment, the speech recognition technology includes PaddleSpeech SpeechRecognition. In other embodiments, it can also be selected according to the actual situation.
[0043] In one embodiment, the multiple non-text data includes: video data, and the video data includes: at least one video file. Please refer to Figure 2 , step S120 includes: steps S121 to S123.
[0044] It can be understood that a video file includes: an audio file and multiple frames of images. As described above, both the audio file and the images can spread sensitive information. Therefore, for a video file, it is necessary to perform desensitization processing on both the included audio file and each frame of image at the same time.
[0045] Step S121, call a preset conversion tool to convert the video file into a corresponding audio file and multiple images to be recognized.
[0046] In this embodiment, the preset conversion tool includes: FastForward Moving Picture Experts Group tool (FFmpeg). The FFmpeg tool is used to extract the corresponding audio file and multiple images to be recognized from the video file. It should be noted that for the extracted audio file, the FFmpeg tool will uniformly convert it into the WAV audio format.
[0047] It should be noted that the FFmpeg tool mainly includes the following components: the command-line tool FFmpeg for video file conversion, the simple player ffplay based on the FFmpeg library, ffprobe for analyzing video files and displaying detailed information, and the server component ffserver for processing live streams.
[0048] Step S122: Use optical character recognition technology to perform character recognition on each of the to-be-recognized images respectively, and obtain the corresponding to-be-desensitized text.
[0049] In this embodiment, for the multiple to-be-recognized images obtained by conversion, optical character recognition technology is respectively used for character recognition, and each recognized character is converted into an editable text format to obtain the corresponding to-be-desensitized text.
[0050] Step S123: Use speech recognition technology to perform audio recognition on the audio file, and obtain the corresponding to-be-desensitized text.
[0051] In this embodiment, for the converted audio file, speech recognition technology is used for audio recognition, and the audio signal is converted into the corresponding text, that is, the to-be-desensitized text.
[0052] Step S124: Determine each of the to-be-desensitized texts as the target text data.
[0053] It can be understood that the target text data corresponding to the video data includes: the to-be-desensitized texts corresponding to each video file respectively.
[0054] In one embodiment, before using the preset conversion method to perform text conversion on each of the non-text data respectively, it further includes: using a preset feature extraction algorithm to perform feature extraction on each of the to-be-recognized images respectively, and obtaining the corresponding image feature vectors; respectively inputting each of the image feature vectors into a preset image recognition model, and obtaining the corresponding feature labels.
[0055] In this embodiment, the non-text data includes image data, and the image data includes at least one to-be-recognized image. The preset feature extraction algorithm includes but is not limited to any one of the following: YOLO, Fast R-CNN, feature pyramid network, or deep learning feature extraction. The preset image recognition model is a pre-trained classification model, which is used to perform image classification according to the obtained image feature vectors, and obtain the feature labels corresponding to the to-be-recognized images.
[0056] Step S130: Determine the text data and each of the target text data as the to-be-desensitized data respectively.
[0057] In this embodiment, a preset desensitization algorithm is used to desensitize each piece of data to be desensitized, and the preset desensitization algorithm includes, but is not limited to, any one of the following: Trie tree, DFA algorithm, Aho-Corasick algorithm, etc., and can be specifically selected according to actual requirements.
[0058] Step S140, obtain a sensitive word library, and based on the sensitive word library, use a preset desensitization algorithm to desensitize each piece of data to be desensitized respectively to obtain corresponding target data.
[0059] It can be understood that the sensitive word library includes multiple sensitive words in aspects such as laws and regulations, moral norms, and industry regulations. To ensure the accuracy and integrity of the sensitive word library, it is necessary to regularly organize, optimize, and update the sensitive word library to incorporate new sensitive words.
[0060] In one implementation manner, the sensitive word library includes: multiple sensitive word sub-libraries of different categories. Before using a preset desensitization algorithm to desensitize each piece of data to be desensitized based on the sensitive word library, it includes: determining a target sensitive word sub-library from multiple sensitive word sub-libraries of different categories according to the corresponding feature label. Using a preset desensitization algorithm to desensitize each piece of data to be desensitized based on the sensitive word library includes: using the preset desensitization algorithm to desensitize each piece of data to be desensitized respectively based on the target sensitive word sub-library.
[0061] It can be understood that for multiple images to be recognized, corresponding sensitive word sub-libraries are determined from the sensitive word library according to the corresponding feature labels, such as animals, plants, or buildings, etc., which can narrow the matching range of sensitive words, and thus improve the desensitization efficiency of the images to be recognized.
[0062] In this embodiment, when using a preset desensitization algorithm to desensitize each piece of data to be desensitized, a strategy of combining exact matching and fuzzy matching is adopted. Exact matching is used to directly hit complete sensitive words. Fuzzy matching considers situations such as word deformation, homophony, simplified and traditional character conversion, and partial matching to improve the comprehensiveness of detection. The preset sensitive algorithm also supports forward matching (starting from the beginning of the text for matching), reverse matching (starting from the end of the text for matching), and any-position matching to adapt to different application scenarios, and finally outputs the desensitized text, that is, the target data.
[0063] In one implementation manner, the method further includes: respectively updating the text data and each non-text data according to the corresponding target data to obtain target multi-modal data.
[0064] In this embodiment, if the non-text data is image data, the text information in the image data is overwritten with the corresponding target data to obtain updated image data; if the non-text data is audio data, after converting the corresponding target data into audio information, the original audio information in the audio data is overwritten to obtain updated audio data; if the non-text data is video data, the audio and multiple frames of images included in the video data are respectively overwritten according to the corresponding target data to obtain updated video data. It can be understood that the updated image data, audio data, and video data constitute the target multi-modal data.
[0065] The sensitive word filtering method for multi-modal data provided by the embodiments of the present invention obtains multi-modal data through acquisition, where the multi-modal data includes: text data and multiple types of non-text data; uses a preset conversion method to perform text conversion on each of the non-text data respectively to obtain target text data; determines each of the target text data and the text data as data to be desensitized respectively; obtains a sensitive word library, and based on the sensitive word library, uses a preset desensitization algorithm to perform desensitization processing on each of the data to be desensitized respectively to obtain corresponding target data, realizing desensitized multi-modal data, effectively reducing the spread of sensitive information, and enhancing the supervision of multimedia data.
[0066] Embodiment 2
[0067] In addition, please refer to Figure 3 , the embodiments of the present invention also provide a sensitive word filtering device 300 for multi-modal data, including:
[0068] A data acquisition module 310, configured to acquire multi-modal data, where the multi-modal data includes: text data and multiple types of non-text data;
[0069] A text conversion module 320, configured to perform text conversion on each of the non-text data respectively using a preset conversion method to obtain target text data;
[0070] A determination module 330, configured to determine the text data and each of the target text data as data to be desensitized respectively;
[0071] A desensitization module 340, configured to obtain a sensitive word library, and based on the sensitive word library, use a preset desensitization algorithm to perform desensitization processing on each of the data to be desensitized respectively to obtain corresponding target data.
[0072] In an implementation manner, multiple types of the non-text data include: image data, the image data includes: at least one image to be recognized, and the text conversion module 320 is further configured to perform text recognition on each of the images to be recognized respectively using optical character recognition technology to obtain corresponding text to be desensitized; and determine each of the text to be desensitized as the target text data.
[0073] In one embodiment, the multiple non-text data includes: audio data, the audio data includes: at least one audio file, and the text conversion module 320 is further configured to perform audio recognition on each of the audio files respectively by using speech recognition technology to obtain corresponding text to be desensitized; and determine each of the text to be desensitized as the target text data.
[0074] In one embodiment, the multiple non-text data includes: video data, the video data includes: at least one video file, and the text conversion module 320 is further configured to call a preset conversion tool to convert the video file into a corresponding audio file and multiple images to be recognized; perform character recognition on each of the images to be recognized respectively by using optical character recognition technology to obtain corresponding text to be desensitized; perform audio recognition on the audio file by using speech recognition technology to obtain corresponding text to be desensitized; and determine each of the text to be desensitized as the target text data.
[0075] In one embodiment, before the text conversion module 320 performs text conversion on each of the non-text data respectively by using a preset conversion method, the text conversion module 320 is further configured to perform feature extraction on each of the images to be recognized respectively by using a preset feature extraction algorithm to obtain corresponding image feature vectors; and input each of the image feature vectors into a preset image recognition model respectively to obtain corresponding feature labels.
[0076] In one embodiment, the sensitive word library includes: multiple sensitive word sub-libraries of different categories. Before the desensitization module 340 performs desensitization processing on each of the data to be desensitized respectively based on the sensitive word library, the desensitization module 340 is further configured to determine a target sensitive word sub-library from the multiple sensitive word sub-libraries of different categories according to the corresponding feature labels; and the desensitization module 340 is further configured to perform desensitization processing on each of the data to be desensitized respectively based on the target sensitive word sub-library by using the preset desensitization algorithm.
[0077] In one embodiment, the sensitive word filtering device 300 for multi-modal data further includes: an update module, configured to update the text data and each of the non-text data respectively according to the corresponding target data to obtain target multi-modal data.
[0078] The sensitive word filtering device 300 for multi-modal data provided by the embodiments of the present invention can execute the sensitive word filtering method for multi-modal data provided in the above method embodiment 1. To avoid repetition, it will not be described in detail here.
[0079] The sensitive word filtering device for multimodal data provided by the embodiment of the present invention acquires multimodal data through a data acquisition module. The multimodal data includes text data and various non-text data. A text conversion module performs text conversion on each of the non-text data respectively by using a preset conversion method to obtain target text data. A determination module determines each of the target text data and the text data as data to be desensitized respectively. A desensitization module acquires a sensitive word library and, based on the sensitive word library, performs desensitization processing on each of the data to be desensitized respectively by using a preset desensitization algorithm to obtain corresponding target data, realizing desensitization of multimodal data, effectively reducing the spread of sensitive information, and enhancing the supervision intensity of multimedia data.
[0080] Embodiment 3
[0081] In addition, the embodiment of the present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the computer program runs on the processor, it executes the sensitive word filtering method for multimodal data provided in Embodiment 1.
[0082] Specifically, please refer to Figure 4 , the electronic device 400 includes: a transceiver 401, a bus interface, and a processor 402. The processor 402 is used to acquire multimodal data, where the multimodal data includes text data and various non-text data; perform text conversion on each of the non-text data respectively by using a preset conversion method to obtain target text data; determine the text data and each of the target text data as data to be desensitized respectively; acquire a sensitive word library and, based on the sensitive word library, perform desensitization processing on each of the data to be desensitized respectively by using a preset desensitization algorithm to obtain corresponding target data.
[0083] In the embodiment of the present invention, the electronic device 400 further includes: a memory 403. In Figure 4 , the bus architecture may include any number of interconnected buses and bridges. Specifically, various circuits represented by one or more processors represented by the processor 402 and a memory represented by the memory 403 are linked together. The bus architecture can also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and therefore, they will not be further described herein. The bus interface provides an interface. The transceiver 401 may be multiple components, that is, including a transmitter and a receiver, and provides a unit for communicating with various other devices on a transmission medium. The processor 402 is responsible for managing the bus architecture and general processing, and the memory 403 can store data used by the processor 402 when performing operations.
[0084] The electronic device 400 provided by the embodiment of the present invention can execute the sensitive word filtering method for multimodal data provided in the above Method Embodiment 1. To avoid repetition, it will not be elaborated here.
[0085] Embodiment 4
[0086] In addition, the embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the sensitive word filtering method for multimodal data provided in Embodiment 1.
[0087] In this embodiment, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, or the like.
[0088] The computer-readable storage medium provided in this embodiment can implement the sensitive word filtering method for multimodal data provided in Embodiment 1. To avoid repetition, it will not be elaborated here.
[0089] In all the examples shown and described here, any specific value should be construed as merely exemplary, not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.
[0090] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0091] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention.
Claims
1. A sensitive word filtering method for multimodal data, characterized in that The method includes: Obtaining multimodal data, where the multimodal data includes: text data and multiple types of non-text data; Using a preset conversion method to perform text conversion on each of the non-text data respectively to obtain target text data; Determining the text data and each of the target text data as data to be desensitized respectively; Obtaining a sensitive word library, and based on the sensitive word library, using a preset desensitization algorithm to perform desensitization processing on each of the data to be desensitized respectively to obtain corresponding target data.
2. The sensitive word filtering method for multimodal data according to claim 1, wherein The multiple types of non-text data include: image data, and the image data includes: at least one image to be recognized. The using a preset conversion method to perform text conversion on each of the non-text data respectively to obtain target text data includes: Using optical character recognition technology to perform text recognition on each of the images to be recognized respectively to obtain corresponding text to be desensitized; Determining each of the text to be desensitized as the target text data.
3. The sensitive word filtering method for multi-modal data according to claim 1, characterized in that The multiple types of non-text data include: audio data, and the audio data includes: at least one audio file. The using a preset conversion method to perform text conversion on each of the non-text data respectively to obtain target text data includes: Using speech recognition technology to perform audio recognition on each of the audio files respectively to obtain corresponding text to be desensitized; Determining each of the text to be desensitized as the target text data.
4. The sensitive word filtering method for multimodal data according to claim 1, wherein The multiple types of non-text data include: video data, and the video data includes: at least one video file. The using a preset conversion method to perform text conversion processing to obtain target text data includes: Invoking a preset conversion tool to convert the video file into a corresponding audio file and multiple images to be recognized; Using optical character recognition technology to perform text recognition on each of the images to be recognized respectively to obtain corresponding text to be desensitized; Using speech recognition technology to perform audio recognition on the audio file to obtain corresponding text to be desensitized; Determining each of the text to be desensitized as the target text data.
5. The sensitive word filtering method for multimodal data according to claim 2 or 4, characterized in that, Before using the preset conversion method to perform text conversion on each of the non-text data respectively, it further includes: Using a preset feature extraction algorithm to perform feature extraction on each of the images to be recognized respectively to obtain corresponding image feature vectors; Inputting each of the image feature vectors into a preset image recognition model respectively to obtain corresponding feature labels.
6. The sensitive word filtering method for multimodal data according to claim 5, characterized in that The sensitive word library includes: multiple different types of sensitive word sub-libraries. Before using the sensitive word library and using a preset desensitization algorithm to perform desensitization processing on each of the data to be desensitized respectively, it includes: According to the corresponding feature label, determining a target sensitive word sub-library from multiple different types of sensitive word sub-libraries; The using the sensitive word library and using a preset desensitization algorithm to perform desensitization processing on each of the data to be desensitized respectively includes: Based on the target sensitive word sub-library, using the preset desensitization algorithm to perform desensitization processing on each of the data to be desensitized respectively.
7. The sensitive word filtering method for multimodal data according to claim 1, wherein, The method further includes: Updating the text data and each of the non-text data according to the corresponding target data respectively to obtain target multimodal data.
8. A sensitive word filtering device for multimodal data, characterized in that, The device includes: A data acquisition module, configured to acquire multimodal data, where the multimodal data includes: text data and multiple types of non-text data; A text conversion module, configured to perform text conversion on each of the non-text data respectively by using a preset conversion method to obtain target text data; A determination module, configured to determine each of the target text data and the text data as data to be desensitized respectively; A desensitization module, configured to obtain a sensitive word library, and based on the sensitive word library, perform desensitization processing on each of the data to be desensitized respectively by using a preset desensitization algorithm to obtain corresponding target data.
9. An electronic device, characterized in that, It includes a memory and a processor, where the memory stores a computer program, and when the computer program runs on the processor, it executes the sensitive word filtering method for multimodal data according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the sensitive word filtering method for multimodal data according to any one of claims 1-7.
Citation Information
Patent Citations
Privacy parameter optimization method in multi-modal data fusion training
CN115310122A
Non-text data desensitization method and device and storage medium
CN115618371A
Character scene modeling method and device, equipment and storage medium
CN116563469A