Content processing method and device, equipment and storage medium
By detecting the text sequence and location information of the content unit, and combining machine learning models to identify and process entities, the problem of insufficient entity detection and processing accuracy in the prior art is solved, and high accuracy adaptive processing of data security and privacy protection is achieved.
Patent Information
- Application Number
- CN202510579287.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The accuracy of detection and processing of entities in data in the prior art still needs to be improved, especially when large language models (LLMs) process data, data security and privacy protection face challenges.
By detecting the text sequence and position information of the content unit, combining machine learning models to detect entity types and contextual relationships, various models are used to identify and process entities, including predetermined processes such as fuzzy processing, masking, substitution, etc., to ensure data security.
It improves the accuracy of entity identification and processing, is suitable for different application scenarios, maintains data security and privacy protection, and has strong adaptability.
Smart Images

Figure CN120296770A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to content processing methods, apparatuses, devices, computer-readable storage media, and computer-executable instruction products. Background Art
[0002] With the advent of the big data era, data has become the main carrier for information recording and transmission. However, at the same time, in the process of data use or transmission, problems of data security and privacy protection are also faced. Obviously, it is very necessary to accurately eliminate sensitive information in the original data in an appropriate manner before data use or transmission. Summary of the Invention
[0003] In a first aspect of the present disclosure, a content processing method is provided. The method includes: detecting a text sequence corresponding to a plurality of content units in a first content and position information indicating positions of the plurality of content units in the first content; detecting, by using a first machine learning model, an entity type corresponding to an entity that appears in the first content to obtain type information indicating the entity type, where the entity is represented by at least one content unit; determining, based on the text sequence, the position information, and the type information, by using a second machine learning model, a first entity to be processed in the first content, where the first entity is represented by at least one content unit; and performing a predetermined process on the first entity in the first content based on the position information to obtain a second content corresponding to the first content.
[0004] In a second aspect of the present disclosure, an apparatus for content processing is provided. The apparatus includes: a first detection module configured to detect a text sequence corresponding to a plurality of content units in a first content and position information indicating positions of the plurality of content units in the first content; a second detection module configured to detect, by using a first machine learning model, an entity type corresponding to an entity that appears in the first content to obtain type information indicating the entity type, where the entity is represented by at least one content unit; a determination module configured to determine, based on the text sequence, the position information, and the type information, by using a second machine learning model, a first entity to be processed in the first content, where the first entity is represented by at least one content unit; and a processing module configured to perform a predetermined process on the first entity in the first content based on the position information to obtain a second content corresponding to the first content.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When the instructions are executed by the at least one processor, the device executes the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. Computer-executable instructions are stored on the computer-readable storage medium, and the computer-executable instructions can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of the present disclosure, a computer program product is provided, including computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the method according to the first aspect of the present disclosure is implemented.
[0008] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0010] Figure 1 A schematic diagram showing an example environment in which embodiments according to the present disclosure can be implemented;
[0011] Figure 2 A schematic diagram showing an example architecture of content processing according to some embodiments of the present disclosure;
[0012] Figure 3 A schematic diagram showing an example architecture of content processing according to some embodiments of the present disclosure;
[0013] Figure 4 A schematic diagram showing an example architecture of content processing according to some embodiments of the present disclosure;
[0014] Figure 5A A schematic diagram showing an example scenario of location information detection according to some embodiments of the present disclosure;
[0015] Figure 5B A schematic diagram showing an example scenario of location information detection according to some embodiments of the present disclosure;
[0016] Figure 6 A schematic diagram showing an example architecture of content processing according to some embodiments of the present disclosure;
[0017] Figure 7 A flowchart showing the process of content processing according to some embodiments of the present disclosure;
[0018] Figure 8Schematic structural block diagrams of example devices for content processing according to some embodiments of the present disclosure are shown; and
[0019] Figure 9 A block diagram of an electronic device capable of implementing multiple embodiments of the present disclosure is shown. Detailed implementation manners
[0020] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0021] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.
[0022] In this article, unless otherwise specified, performing a step "in response to A" does not mean that the step is immediately executed after "A", but may include one or more intermediate steps.
[0023] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0024] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to the relevant laws and regulations.
[0025] For example, when receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require the acquisition and use of the user's personal information, so that the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that performs the operations of the technical solution of the present disclosure according to the prompt message.
[0026] As an optional but non-limiting implementation manner, in response to receiving an active request from a user, the manner of sending a prompt message to the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It can be understood that the above-mentioned notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0028] As used herein, the term "model" can learn the corresponding association relationship between input and output from training data, so that after training is completed, for a given input, a corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms can be used interchangeably in this article.
[0029] A "neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and it generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. The neural networks used in deep learning applications usually include many hidden layers, thereby increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is used as the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.
[0030] Generally, machine learning can be roughly divided into three stages, namely, a training stage, a testing stage, and an application stage (also called an inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously iteratively updated until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, so as to determine the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values obtained through training and determine the corresponding output.
[0031] As mentioned above, with the advent of the big data era, data has become the main carrier for information recording and transmission. However, at the same time, during the process of data use or transmission, issues regarding data security and privacy protection also arise. For example, large language models (LLMs) have achieved unprecedented improvements in efficiency and quality in natural language processing, information retrieval, intelligent question answering, etc. However, during the process of using LLMs for data processing, challenges in data security and privacy protection have also emerged. Obviously, it is very necessary to accurately eliminate protected entities in the original data in an appropriate manner before data use or transmission. However, the accuracy of entity detection and processing in traditional technologies still needs to be improved.
[0032] In view of this, embodiments of the present disclosure propose an improved solution for content processing. In this solution, text sequences corresponding to multiple content units in the first content are detected, and the positions of the multiple content units in the first content are also detected to obtain position information indicating these positions. A first machine learning model is used to detect the entity types corresponding to the entities appearing in the first content to obtain type information indicating the entity types. An entity is represented by at least one content unit. Based on the text sequence, position information, and type information, a second machine learning model is used to determine the first entity to be processed in the first content, and the first entity is represented by at least one content unit. Based on the position information, a predetermined process is performed on the protected entities in the first content to obtain a second content corresponding to the first content.
[0033] In the embodiments of the present disclosure, not only the text sequence corresponding to the content is detected, but also the positions of the content units and the entities appearing in the content are detected. By combining the position information and the type information indicating the entity types, the model (i.e., the second machine learning model) can capture the context relationship and order of the entities in the text sequence, and thus can accurately identify the boundaries and types of the first entity, which is beneficial to improving the accuracy of the recognition and processing of the first entity. Moreover, by configuring the entity types detected by the first machine learning model, this solution can maintain high accuracy and adaptability in different application scenarios or different systems.
[0034] The following further describes various example implementations of this solution in detail with reference to the accompanying drawings.
[0035] Example environment
[0036] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, an application 120 is installed in a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110.
[0037] In some embodiments of the present disclosure, the application 120 may be any suitable application with functions of human-machine dialogue, information query, or data transmission. For example, the application 120 may provide a digital assistant for human-machine dialogue. The digital assistant supports text dialogue services, voice dialogue services, and content dialogue in other modalities with the user 140. Also for example, the application 120 may provide a digital assistant for information query. The digital assistant supports the user 140 to use query requests in text, voice, or other modalities.
[0038] In some embodiments, if the application 120 is in an active state, the terminal device 110 may present the user interface 150 of the application 120. The user interface 150 may include various types of pages that the application 120 can provide, such as a dialogue page between the user and the digital assistant, an information retrieval interface, and so on. In some embodiments, the terminal device 110 may present text or images in the user interface 150, or may also play voice or video in the user interface 150. The voice may be, for example, the voice from the user 140, the response voice to the voice of the user 140, or the voice corresponding to the retrieved text content.
[0039] In some embodiments, the application 120 or the digital assistant therein may utilize a machine learning model 160 (which may include one or more machine learning models, for example, may include machine learning models 160-1, 160-2,..., 160-N, etc., where N is a positive integer. For the convenience of description, one or more machine learning models are collectively referred to as the machine learning model 160 in this article) to support the interaction with the user 140. For example, the application 120 or the digital assistant therein may utilize one or more machine learning models 160 to provide an information query service to the user 140. Also for example, the application 120 or the digital assistant therein may also utilize one or more machine learning models 160 to provide a question-and-answer service to the user 140.
[0040] In some embodiments, the machine learning model 160 can be of different types. In some embodiments, one or more machine learning models 160 can be built based on a language model (LM). The machine learning model used is a content generation model that can generate corresponding outputs based on model inputs. In some embodiments, the machine learning model based on a language model can handle model inputs in text modalities (e.g., natural language and / or machine language) and / or non-text modalities (e.g., images, speech, videos, etc.), and can generate desired outputs according to the model inputs and prompt words. Here, the prompt words are used to guide the machine learning model to generate outputs that can solve the user queries indicated by the model inputs. In application scenarios for supporting user conversations, the input of user 140 can be provided to the machine learning model 160 as at least part of the model input (the other part can include prompt words). This user input is regarded as a question or query request. Based on the model output, corresponding responses can be provided to user 140.
[0041] In some embodiments, one or more machine learning models 160 can be models related to speech, including an automatic speech recognition (ASR) model and a text-to-speech (TTS) model. The input of the ASR model is speech and the output is text. The input of the TTS model is text and the output is the corresponding speech.
[0042] In some embodiments, the terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of user interface (such as a "wearable" circuit, etc.).
[0043] In some embodiments, the terminal device 110 can communicate with the server 130 to implement the supply of services for the application 120. As Figure 1As shown, the server 130 may invoke the machine learning model 160 to support the application 120 in providing a human-machine dialogue service or an information query service, etc. to the user 140 based on the output of the machine learning model 160. In some embodiments, the server 130 may include, but is not limited to, mainframes, edge computing nodes, computing devices in a cloud environment, etc. It may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The server may, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, etc.
[0044] It should be understood that the structures and functions of the various elements in the environment 100 are described only for exemplary purposes, without implying any limitation on the scope of the present disclosure.
[0045] Example solution
[0046] Some example embodiments of the present disclosure will be further described below with reference to the accompanying drawings. Figure 2 A schematic diagram of an example architecture 200 for content processing according to some embodiments of the present disclosure is shown. For ease of discussion hereinafter, the example architecture 200 is described from the perspective of the server 130, but this is only exemplary.
[0047] In block 210, the server 130 may detect a text sequence corresponding to a plurality of content units in the first content and position information indicating the positions of the plurality of content units in the first content. The first content may be understood as the content that needs to be detected and processed for the first entity. The first entity here may be any entity that needs to be processed in the first content, and the specific type of the first entity is associated with the processing method for the first entity.
[0048] In some embodiments, the first entity may be a protected entity in the first content. A protected entity may be understood as the content part that needs to be protected in the overall content, such as an organization name, a place name, a time, a numerical value, etc. indicated according to specific management regulations, personal wishes, etc. A protected entity may include one or more named entities in the content. It can be understood that the first entity is not limited to a protected entity. In different application scenarios, the first entity may be different types of entities. For example, the first entity may also be an entity that needs to be expanded, rewritten, deleted, or summarized in the first content. It should also be noted that the protected entity is not limited to the above entities. For example, in a specific industry or specific field, there may be some specific named entities that are specific to the specific industry or specific field and need to be protected.
[0049] The first content may include various modal data contents, such as text, image, audio, video, and so on. In some examples, as Figure 2 shown, the application 120 may have a human-machine dialogue function or an information retrieval function. The first content may include the dialogue content or query request input by the user 140 to be provided to the machine learning model 160-1 (sometimes also referred to as the third machine learning model in this article). The server 130 may obtain the first content from the application 120. In other examples, the first content may be the content to be transmitted from one terminal device to another terminal device. For example, the first content may be the content to be transmitted from the terminal device within an organization to the terminal device outside the organization. It can be understood that the above first content is only exemplary. In actual application scenarios, any appropriate content that needs to be detected for the first entity can be determined as the first content according to actual needs. The embodiments of the present disclosure do not limit this.
[0050] The content unit is a constituent unit of the first content. In the case where the modality of the first content is different, the specific content of the content unit may also be different, and the corresponding position information indicating the position of the content unit in the first content may also be different. In some examples, if the first content includes text, the content unit may be a word, a character, a symbol, etc. in the text. The position information may indicate the position of the word, character, symbol, etc. in the text. For example, if the content unit includes a character in the text, the position information may include the page number, paragraph number, line number, column number of the character in the text, the serial number in the line, column or paragraph where the character is located, and the position information may also include the coordinates of the character in the text.
[0051] In some examples, if the first content includes an image, the content unit may include an image region, an object, a pixel unit, etc. in the image. The position information may indicate the position of the image region, object, pixel unit in the image. As an example, the content unit may include a text object (such as a character) in the image, and the position information may indicate the position of the character in the image. For example, the position information may include the coordinates of at least one reference point (such as the center point or the vertex of the circumscribed rectangle) of the character, the size (such as the size in the length direction and the width direction), the rotation angle (such as the rotation angle relative to an axis in the image coordinate system), etc.
[0052] In some examples, if the first content includes audio, the content unit may include an audio frame, a voice segment, etc. in the audio. The position information may indicate the position of the audio frame or voice segment in the audio. For example, the position information may include the timestamp of the voice segment or audio frame. It can be understood that the above content unit and position information are both exemplary, and the embodiments of the present disclosure do not limit this.
[0053] In some embodiments, the text sequence may be a sequence composed of texts recognized from multiple content units. For example, the text sequence may include a sequence composed of texts recognized from contents such as text, image, audio, video, etc. In other examples, the text sequence may be a sequence composed of texts associated with multiple content units. As an example, the text sequence may be a text sequence input by a user related to a content unit, a text sequence for describing a content unit (for example, a text sequence for describing an object in an image), and so on.
[0054] In some embodiments, the server 130 may detect the text sequence and the location information based on the first content by using the machine learning model 160-2 (sometimes also referred to as the fourth machine learning model in this document). The machine learning model 160-2 here may be trained to be capable of detecting the text sequence in the content and the location of the content unit. In some examples, the machine learning model 160-2 may be trained based on the Connectionist Temporal Classification (CTC) model. Of course, the machine learning model 160-2 may also be trained based on any other appropriate model structure, and the embodiments of the present disclosure do not limit this.
[0055] As an example, Figure 3 FIG. shows a schematic diagram of an example structure 300 for content processing according to some embodiments of the present disclosure. As Figure 3 shown, the server 130 may be deployed with an interface 310, a preprocessing module 320, and a detection engine 330. The server 130 may receive the first content for the first entity detection through the interface 310. Before detecting the text sequence and the location information, the server 130 may use the preprocessing module 320 to preprocess the first content. Specifically, if the first content includes text, at block 322, the server 130 may preprocess the text. For example, the server 130 may convert the text into text conforming to a predetermined format. If the first content includes an image, at block 324, the server 130 may preprocess the image. For example, the server 130 may perform noise reduction, enhancement, or extract image regions on the image. If the first content includes audio (not shown in the figure), the server 130 may perform noise reduction, filtering, etc. on the audio. After preprocessing the first content, the server 130 may provide the preprocessed first content to the detection engine 330. The detection engine 330 may use the machine learning model 160-2 to detect the text sequence and the location information.
[0056] In some embodiments, the server 130 may determine at least one content segment in the first content, and each content segment includes at least one content unit. The server 130 may determine at least one feature representation (sometimes also referred to as the third feature representation in this article) corresponding to the at least one content segment based on performing feature extraction on the at least one content segment. After that, the server 130 may provide the at least one feature representation to the machine learning model 160-2 to obtain a text sequence output by the machine learning model 160-2 and position information indicating the positions of the respective content units in the first content.
[0057] Figure 4 FIG. 400 is a schematic diagram showing an example architecture for content processing according to some embodiments of the present disclosure. As Figure 4 shown, the first content may include a first image 402, the content unit may include characters in the first image 402, and the content unit may be an image area (i.e., a text block) containing text in the first image 402. The server 130 may perform text block detection (404) on the first image 402 to determine an image area 406 containing text in the first image 402. The server 130 may provide the image area 406 to a trained feature extractor 408 to perform feature extraction on the image area 406 using the feature extractor 408 to obtain a feature representation corresponding to the image area 406. After that, the server 406 may provide the feature representation to the machine learning model 160-2 to obtain a first detection result 410. The first detection result 410 may indicate a text sequence corresponding to the image area 406 and the positions of each character in the text sequence in the first image 402.
[0058] In some embodiments, for a first image area among the at least one image area, the server 130 may determine, based on the feature representation corresponding to the first image area (i.e., the third feature representation), using the machine learning model 160-2, the first positions of the at least one content unit in the first image area in the first image area. After that, the server 130 may determine, based on the first positions of the at least one content unit, the second positions of the at least one content unit in the first image to obtain position information indicating the second positions. That is, the server 130 may use the machine learning model 160-2 to determine the first positions of the content units in the image area. After that, the server 130 may perform a position transformation on the first positions based on the position relationship between the image area and the first image to determine the second positions of the content units in the first image.
[0059] As an example, Figure 5AA schematic diagram showing an example scenario 500A of location information detection according to some embodiments of the present disclosure is presented. In the example scenario 500A, the feature extractor 408 can be formed based on training a convolutional neural network (CNN), and the machine learning model 160-2 can be formed based on training a CTC model. The server 130 can utilize the feature extractor 408 to perform feature extraction on the image region 406, obtaining a feature map corresponding to the image region 406. Subsequently, the server 130 can provide the feature map to the machine learning model 160-2, and utilize the machine learning model 160-2 to output a probability matrix 510. Each element 512 in the probability matrix 510 can indicate the probability value that the corresponding sub-region in the image region 406 belongs to each character from A to Z. Based on this probability matrix 510, the server 130 can determine the text sequence corresponding to the image region 406 (i.e., "ABCD") and the first position of each character in the text sequence within the image region 406.
[0060] As an example, Figure 5B A schematic diagram showing an example scenario 500B of location information detection according to some embodiments of the present disclosure is presented. As Figure 5B shown, the server 130 can determine the coordinates (X l , g l ) of each character in the image coordinate system (x l , y l ) of the image region 406. Subsequently, the server 130 can convert the coordinates (X l , Y l ) to the image coordinate system (x i , y i ) of the first image based on the formula shown below.
[0061]
[0062] Where (X l , Y l ) represents the coordinates of the character in the image coordinate system (x l , y l ) of the image region 406, indicating the first position of the character; (X I , Y I ) represents the coordinates of the character in the image coordinate system (x i , y i ) of the first image, indicating the second position of the character; X t and Y t represent the translation parameters, and θ represents the rotation angle.
[0063] Returning to the example architecture 200, at block 220, the server 130 can utilize the machine learning model 160-3 (sometimes also referred to herein as the first machine learning model) to detect the entity types corresponding to the entities appearing in the first content, so as to obtain type information indicating the entity types. An entity can be understood as a named entity appearing in the first content. For example, the types corresponding to the entities can include organization names, place names, times, numerical values, product names, and so on. An entity is represented by at least one content unit. For example, an entity of the entity type of organization is formed by 4 characters "ABCD".
[0064] In some embodiments, the server 130 can, based on the first content and the reference text content associated with the first content, utilize the machine learning model 160-3 to determine the type information indicating the entity types corresponding to the entities appearing in the first content. The reference text content can include text content that is associated with the first content and can provide a reference for the machine learning model 160-3 to determine the entity types of the entities. Thereby, it is beneficial to improve the accuracy of entity and entity type detection.
[0065] In some examples, the reference text content can include input text content associated with the first text. For example, the user 140 can input a first image to the dialogue interface of the application 120. At the same time, the user 140 can also input to the dialogue interface of the application 120, for example, "XXXXX in the image is YYYYY, may I ask ZZZZZ?". In this case, the server 130 can determine this text content input by the user 140 as the reference text content. The server 130 can, based on the first image and this text content input by the user 140, determine the entity types corresponding to the entities appearing in the first image.
[0066] Alternatively or additionally, the reference text content can also include prompt words for the machine learning model 160-3. As an example, in the case where the first content includes a first image, the server 130 can select a prompt word for instructing the machine learning model 160-3 to detect entities from the image, such as "You are an expert in image entity recognition,..., please return the entities appearing in the image and the entity types corresponding to the entities". The server 130 can, based on the prompt word and the first image, utilize the machine learning model 160-3 to determine the entity types corresponding to the entities appearing in the first image.
[0067] Alternatively or additionally, the reference text content may further include a text sequence corresponding to the first content. As an example, the server 130 may use the machine learning model 160-2 to identify a text sequence from the audio. After that, the server 130 may determine the entity type corresponding to the entity appearing in the audio based on the audio and the identified text sequence. It should be noted that the above reference text content is only exemplary, and other text content associated with the first content and capable of providing a reference for entity detection may be selected according to actual needs. The embodiments of the present disclosure are not limited thereto.
[0068] In some embodiments, the machine learning model 160-3 may include a first encoder, a second encoder, and a decoder. The server 130 may perform feature encoding using the first encoder based on the first content to generate a first feature representation. The server 130 may also perform feature encoding using the second encoder based on the reference text content to generate a second feature representation. After that, the server 130 may perform decoding using the decoder based on the first feature representation and the second feature representation to determine the type information.
[0069] In some examples, as Figure 4 shown, the first content may include a first image 402. The machine learning model 160-3 may include a first encoder 412 (which may also be referred to as a visual encoder), a second encoder (which may also be referred to as a text encoder), and a decoder. The server 130 may perform feature encoding on the first image 402 using the first encoder 412 to generate a first feature representation 414. As an example, the server 130 may encode the first image 402 into a high-dimensional vector representation using the first encoder 412.
[0070] The server 130 may also perform feature encoding on the reference text content 416 using the second encoder 418 to generate a second feature representation 420. As an example, the server 130 may use the text content input by the user 140, the prompt words, and the text sequence detected from the first image 402 as the reference text content. The server 130 may use the second feature encoder 418 to encode the reference text content 416 into a second feature representation 420 that can be aligned with the first feature representation 414.
[0071] At block 422, the server 130 may concatenate the first feature representation 414 and the second feature representation 420 to obtain the input feature representation for the decoder 424. The server 130 may utilize the decoder 424 to perform feature decoding on the input feature representation, generating the second detection result 426. The second detection result may indicate the entity type corresponding to the entity present in the first image 402. Thereafter, the server 130 may obtain the detection result 332 by combining the first detection result 410 and the second detection result 426. As an example, the server 130 may utilize the decoder 424 to perform cross-attention processing on the input feature representation to facilitate multi-modal feature fusion, so as to improve the accuracy of detecting the entity and the entity type corresponding to the entity. It should also be noted that the above example only explains the embodiments of the present disclosure by taking the recognition of characters from an image as an example, and should not be understood that the embodiments of the present disclosure are limited to recognizing text and entities composed of text from an image. In practical application scenarios, a similar architecture may also be used to process other modalities such as audio and video, and the embodiments of the present disclosure are not limited thereto.
[0072] In some embodiments, after determining the entity type corresponding to the entity and the position of the content unit in the first content, the server 130 may label the entity type and the position of the content unit in the first content to facilitate subsequent predetermined processing of the first entity. As an example, the server 130 may label the entity type and the position of the character in the first image 402.
[0073] Returning to Figure 2 As shown, at block 230, the server 130 may determine the first entity to be processed in the first content based on the text sequence, the position information, and the type information, using the machine learning model 160-4 (sometimes also referred to as the second machine learning model in this article). The first entity has been explained in the foregoing content, and specific details can be referred to the introduction in the foregoing content. It should be added here that the first entity may be represented by at least one content unit, for example, an organization name, a place name, etc. may be formed by multiple characters.
[0074] The first entity here may be the same as or different from the entity in the foregoing content. In some examples, the first entity may include one or more entities. As an example, the server 130 may respectively determine "XX City" and "XXX County" as entities with the entity type of "place name". The server 130 may also determine "XX City XXX County" as the first entity. At this time, the first entity actually includes two entities with the entity type of "place name".
[0075] In some examples, the first entity may also include a part of the entity. As an example, the server 130 may determine "YYYY-MM-DD" as an entity of the entity type "date". In block 230, the server 130 may only determine the specific year "YYYY", month "MM", and date "DD" as the first entity. In this way, it is beneficial to maintain the data format of this "date" entity and beneficial to maintain the original semantics in the content.
[0076] In some examples, the first entity may further include at least one content unit of the entity and adjacent entities. For example, in the case where the recognition of the entity is incomplete, a part of the characters of the adjacent entity may be supplemented to the entity, and the supplemented entity is determined as the first entity.
[0077] In some examples, the first entity may include a protected entity. The machine learning model 160-3 may not determine a part of the content segment in the first content as an entity. However, the machine learning model 160-4 may, in combination with the context semantics of the text sequence, determine that this part of the content segment needs to be protected and determine this part of the content segment as a protected entity.
[0078] In some embodiments, as Figure 3 shown, the server 130 may also be deployed with an identification engine 340. The identification engine 340 may obtain a detection result 332 from the detection engine 330, and the detection result 332 may include a text sequence, location information, and type information. The identification engine 340 may perform feature encoding on the text sequence, location information, and type information using a multilingual encoder to generate a feature representation sequence. Then, the machine learning model 160-4 is used to determine the first entity in the first content based on the feature representation sequence.
[0079] Figure 6 shows a schematic diagram of an example architecture 600 for content processing according to some embodiments of the present disclosure. As Figure 6 shown, the server 130 may perform word segmentation processing on the text sequence 602 using a word segmenter 604 to obtain a sub-word unit sequence 606 composed of multiple sub-word units. In some examples, the word segmenter 604 may be a multilingual word segmenter (MultilingualTokenizer). In this way, the text sequence composed of multiple languages can be processed.
[0080] At block 608, the server 130 may perform an embedding process on the sub-word unit sequence 606, the position information, and the type information to generate an embedded representation sequence 610. As an example, the server 130 may determine the token embedding representation corresponding to each sub-word unit in the sub-word unit sequence 606. The server 130 may also determine the position embedding representation and the type embedding representation corresponding to each sub-word unit based on the position information and the type information. After that, the server 130 may combine the token embedding representation, the position embedding representation, and the type embedding representation to form a high-dimensional tensor representation sequence to improve the recognizability of proper nouns and entity boundaries.
[0081] The server 130 may perform feature encoding on the embedded representation sequence 610 by using the multi-lingual encoder 342 to generate a feature representation sequence 612. The feature representation sequence includes multiple feature representations respectively corresponding to multiple sub-word units in the sub-word unit sequence 606 (for example, the embedded representation sequence 610 may be encoded into a high-dimensional vector representation sequence). The server 130 may use the machine learning model 160-4 to determine a label sequence based on the feature representation sequence 612. The label sequence may include multiple labels respectively corresponding to multiple sub-word units in the sub-word unit sequence 606, and each label may indicate the entity type of the corresponding sub-word unit. Specifically, each label may indicate whether the corresponding sub-word unit belongs to a protected entity (i.e., the first entity). When the corresponding sub-word unit belongs to a protected entity, the label can indicate the entity type of the corresponding sub-word unit.
[0082] In some examples, the machine learning model 160-4 may be trained based on a Conditional Random Field (CRF) model. The machine learning model 160-4 may be trained to be able to detect the first entity in the content. The maximum likelihood algorithm of the CRF model can optimize the annotation of entity types, can determine entity boundaries by combining context semantics, is beneficial to reducing the annotation errors of entity types caused by illogicality between adjacent labels, and is beneficial to improving the detection accuracy of the first entity. Of course, the above model structure is only exemplary, and the machine learning model 160-4 may adopt any appropriate model structure, and the embodiments of the present disclosure are not limited thereto.
[0083] In some embodiments, the server 130 may utilize the machine learning model 160-4 to determine a first recognition result indicating a first entity in the first content based on the text sequence, location information, and type information. The server 130 may also determine at least one second recognition result based on the text sequence and at least one predetermined rule, where each second recognition result indicates the first entity in the first content. Subsequently, the server 130 may determine the first entity in the first content by combining the first recognition result and the at least one second recognition result. That is, not only is the machine learning model 160-4 used to detect the first entity in the first content, but the first entity in the first content is also detected based on the predetermined rule, and the final first entity is determined by combining the detection results of the two. Thereby, not only can the accuracy of the detection result be improved, but also the improved solutions of the embodiments of the present disclosure can be flexibly applied to different application scenarios or business systems. Moreover, in the actual application scenario, by maintaining the predetermined rule (such as a thesaurus), it is convenient for the business party to flexibly adjust the first entity.
[0084] As an example, as Figure 6 shown, the server 130 may determine the first recognition result 616 based on the tag sequence 614 output by the machine learning model 160-4. The server 130 may utilize the rule engine 618 to identify the protected entity in the text sequence 602 based on the first rule (such as keyword matching or regular expression matching) to determine the second recognition result 620-1. The server 130 may also utilize the rule engine 618 to identify the protected entity in the text sequence 602 based on the second rule (such as a thesaurus) to determine the second recognition result 620-2. Subsequently, the server 130 may combine the first recognition result 616, the second recognition result 620-1, and the second recognition result 620-2 to determine the final recognition result 622.
[0085] In some embodiments, the first recognition result may have a corresponding priority, and the at least one second recognition result may also have a corresponding priority. The server 130 may determine the first entity in the first content based on the first recognition result, the priority corresponding to the first recognition result, the at least one second recognition result, and the priorities corresponding to the at least one second recognition result respectively. Of course, the server 130 is not limited to combining the first recognition result and the at least one second recognition result based on the priority. The server 130 may determine the intersection or union of the first recognition result and the at least one second recognition result as the final recognition result. The embodiments of the present disclosure do not limit the merging method of the first recognition result and the second recognition result.
[0086] In some embodiments of the present disclosure, the server 130 may perform a predetermined process on the first entity in the first content based on the location information to obtain the second content corresponding to the first content. The second content may include the first entity after the predetermined process and the content units in the first content that do not belong to the first entity.
[0087] The predetermined process here may include various processing operations on the first entity. In some embodiments, the first entity may be a protected entity in the first content. The server 130 may perform a blurring process on the protected entity in the first content based on the location information. The blurring process may include, but is not limited to, performing operations such as hashing, masking, replacing, or encrypting on the protected entity. As an example, as Figure 3 shown, the server 130 may utilize, for example, the blurring processing engine 350 to perform a blurring process on the protected entity in the first content. Specifically, if the first content includes text, at block 352, the server 130 may perform a blurring process such as hashing, masking, replacing, or encrypting on the protected entity in the text. If the first content includes an image, at block 354, the server 130 may perform a blurring process such as occlusion or masking on the protected entity in the image.
[0088] In some examples, if the content unit corresponds to a character in the first content, the location information can indicate the positions of each character in the first content. The server 130 may perform a blurring process on the protected entity in the first content at the character level, which is beneficial to improving the accuracy of the blurring process of the protected entity.
[0089] In some embodiments, the server 130 may determine the protection level of the protected entity. After that, the server 130 may perform a blurring process on the protected entity in the first content based on the location information and the blurring processing strategy corresponding to the protection level to obtain the second content. As an example, for a protected entity with a relatively high protection level, the server 130 may perform, for example, an encryption process. For a protected entity with a relatively low protection level, the server 130 may perform, for example, an occlusion or replacement process.
[0090] In some embodiments, if the first content is to be provided to the machine learning model 160-1, the server 130 may replace the protected entity in the first content with a semantic variable indicating the semantics of the protected entity based on the location information to obtain the second content. That is, the semantics indicated by the semantic variable may be the same as or similar to the semantics of the protected entity. In this way, the obtained second content can maintain the original semantics to a certain extent, which is beneficial for the machine learning model 160 to accurately generate responses.
[0091] Note that the predetermined processing is not limited to blurring. The server 130 may also perform any appropriate processing on the first entity, such as rewriting, expanding, deleting, or summarizing. The embodiments of the present disclosure do not limit the processing method of the predetermined processing. It should also be noted that after performing the predetermined processing, the server 130 is not limited to belonging to the second content. The server 130 may also output, for example, a processing report to record the identified first entity and the predetermined processing performed on the first entity through the processing report. In this way, it is possible to facilitate subsequent execution of, for example, a restoration process on the first entity after the predetermined processing.
[0092] Combined with Figure 2 As shown, the server 130 may determine a first response to the first question regarding the first content based on the second content using the machine learning model 160-1. Specifically, the server 130 may provide the second content to the machine learning model 160-1 to obtain a model output generated by the machine learning model 160-1. The server 130 may determine a first response to the first question regarding the first content based on the model output. The server 130 may feedback the first response to the terminal device 110. The terminal device 110 may present the first response using the dialogue interface of the application 120. In the case where the first response includes audio or video, the terminal device 110 may also play the audio or video using the dialogue interface of the application 120.
[0093] In some embodiments, the server 130 may determine a second response including the first entity after the predetermined processing, such as a second response including the protected entity after blurring, based on the second content using the machine learning model 160-1. After that, the server 130 performs a restoration process on the first entity after the predetermined processing in the second response to obtain a first response including the first entity. In some examples, as Figure 2 shown, at block 250, the server 130 may perform a format check on the content units other than the first entity after the predetermined processing in the second response, such as checking whether these content units conform to a predetermined format, or checking whether the format of these content units matches the format of the corresponding content units in the first content. At block 260, the server 130 may perform a restoration process on the first entity after the predetermined processing in the second response after the check, such as replacing it back with the protected entity, or decrypting the protected entity after blurring, etc. At block 270, the server 130 may also perform a formatting process on the first response, such as processing the first response into a data format that matches the dialogue interface of the application 120. After that, the server 130 may feedback the first response after the formatting process to the terminal device 110. This can facilitate the user 140 himself / herself to view the response including the complete content.
[0094] It should be noted that the above exemplary structure 200 only exemplarily illustrates the improvement solutions of the embodiments of the present disclosure using the human-machine dialogue scenario as an example. The improvement solutions of the embodiments of the present disclosure are not limited to being applied in the human-machine dialogue scenario, but can also be applied in other application scenarios that require processing entities in the content, such as information retrieval or data transmission.
[0095] As an example, if the first terminal device needs to send the first content to the second terminal device, the first terminal device can provide the first content to the server 130. The server 130 detects the protected entities in the first content and performs a fuzzification process on the protected entities in the first content to obtain the second content. After that, the second content can be provided to the second terminal device.
[0096] As another example, if the terminal device 110 triggers a predetermined operation on the first content (such as a screen mirroring or sharing operation), the terminal device 110 can send a request to the server 130 to request the server 130 to perform protected entity detection on the first content. The server 130 can feedback the second content to the terminal device. The terminal device 110 can perform the predetermined operation based on the second content.
[0097] In this way, in the embodiments of the present disclosure, not only the text sequence corresponding to the content is detected, but also the positions of the content units and the entities appearing in the content are detected. Combining the position information and the type information indicating the entity types enables the model (i.e., the second machine learning model) to capture the context relationship and order of the entities in the text sequence, and thus can accurately identify the boundaries and types of the first entities, which is beneficial to improving the accuracy of the recognition and processing of the first entities. Moreover, by configuring the entity types detected by the first machine learning model, this solution can maintain high accuracy and adaptability in different application scenarios or different systems.
[0098] Example process, apparatus and equipment
[0099] Figure 7 The flowchart of the content processing process 700 according to some embodiments of the present disclosure is shown. For the convenience of discussion hereinafter, the process 700 is described from the perspective of the server 130, but this is only exemplary.
[0100] In block 710, the server 130 detects the text sequences corresponding to multiple content units in the first content and the position information indicating the positions of the multiple content units in the first content.
[0101] In block 720, the server 130 uses the first machine learning model to detect the entity types corresponding to the entities appearing in the first content to obtain the type information indicating the entity types, where the entity is represented by at least one content unit.
[0102] At block 730, the server 130 determines a first entity to be processed in the first content, which is represented by at least one content unit, based on the text sequence, location information, and type information, using a second machine learning model.
[0103] At block 740, the server 130 performs a predetermined process on the first entity in the first content based on the location information to obtain a second content corresponding to the first content.
[0104] In some embodiments, process 700 further includes: providing the second content to a specific terminal device, or determining a first response to a first question regarding the first content based on the second content, using a third machine learning model.
[0105] In some embodiments, determining the first response includes: determining a second response including the first entity after the predetermined process based on the second content, using a third machine learning model; and performing a restoration process on the first entity after the predetermined process in the second response to obtain a first response including the first entity.
[0106] In some embodiments, the first entity includes a protected entity in the first content, and performing a predetermined process on the first entity in the first content based on the location information includes: performing a blurring process on the protected entity in the first content based on the location information to obtain a second content corresponding to the first content.
[0107] In some embodiments, performing a blurring process on the protected entity in the first content based on the location information includes: determining the protection level of the protected entity; performing a blurring process on the protected entity in the first content based on the location information and a blurring process strategy corresponding to the protection level to obtain the second content.
[0108] In some embodiments, performing a blurring process on the protected entity in the first content based on the location information includes: replacing the protected entity in the first content with a semantic variable indicating the semantics of the protected entity based on the location information to obtain the second content.
[0109] In some embodiments, detecting the entity type corresponding to an entity appearing in the first content using a first machine learning model includes: determining type information indicating the entity type corresponding to the entity appearing in the first content based on the first content and reference text content associated with the first content, using the first machine learning model.
[0110] In some embodiments, the first machine learning model includes a first encoder, a second encoder, and a decoder, and determining the type information includes: performing feature encoding on the first content using the first encoder to generate a first feature representation; performing feature encoding on the reference text content using the second encoder to generate a second feature representation; and performing decoding using the decoder based on the first feature representation and the second feature representation to determine the type information.
[0111] In some embodiments, the reference text content includes at least one of the following: a text sequence, a prompt word for the first machine learning model, or input text content associated with the first content.
[0112] In some embodiments, detecting the text sequence and the position information includes: determining at least one content segment in the first content, each content segment including at least one content unit; determining at least one third feature representation based on performing feature extraction on the at least one content segment; and determining the text sequence and the position information indicating the positions of the respective content units in the first content using a fourth machine learning model based on the at least one third feature representation.
[0113] In some embodiments, the first content includes a first image, at least one content segment includes at least one image region in the first image, and determining the position information includes: for a first image region among the at least one image region, determining, based on the third feature representation corresponding to the first image region, the first positions of the at least one content unit in the first image region using a fourth machine learning model; and determining the second positions of the at least one content unit in the first image based on the first positions of the respective content units to obtain the position information indicating the second positions.
[0114] In some embodiments, the multiple content units respectively correspond to multiple characters in the text sequence.
[0115] In some embodiments, determining the first entity in the first content includes: performing feature encoding using a multilingual encoder based on the text sequence, the position information, and the type information to generate a sequence of feature representations; and determining the first entity in the first content using a second machine learning model based on the sequence of feature representations.
[0116] In some embodiments, determining the first entity in the first content includes: determining a first recognition result indicating the first entity in the first content using a second machine learning model based on the text sequence, the position information, and the type information; determining at least one second recognition result indicating the first entity in the first content based on the text sequence and at least one predetermined rule; and determining the first entity in the first content by combining the first recognition result and the at least one second recognition result.
[0117] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above methods or processes. Figure 8 FIG. shows a schematic structural block diagram of an example apparatus 800 for content processing according to some embodiments of the present disclosure. The apparatus 800 may be implemented as or included in the server 130. Each module / component in the apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.
[0118] As Figure 8 shown, the apparatus 800 includes: a first detection module 810 configured to detect a text sequence corresponding to a plurality of content units in the first content and position information indicating the positions of the plurality of content units in the first content; a second detection module 820 configured to detect an entity type corresponding to an entity appearing in the first content by using a first machine learning model to obtain type information indicating the entity type, where the entity is represented by at least one content unit; a determination module 830 configured to determine a first entity to be processed in the first content by using a second machine learning model based on the text sequence, the position information, and the type information, where the first entity is represented by at least one content unit; and a processing module 840 configured to perform a predetermined process on the first entity in the first content based on the position information to obtain a second content corresponding to the first content.
[0119] In some embodiments, the apparatus 800 further includes: a providing module configured to provide the second content to a specific terminal device, or a response module configured to determine a first response to a first question regarding the first content by using a third machine learning model based on the second content.
[0120] In some embodiments, the response module is further configured to: determine a second response including the first entity after the predetermined process based on the second content by using a third machine learning model; and perform a restoration process on the first entity after the predetermined process in the second response to obtain a first response including the first entity.
[0121] In some embodiments, the first entity includes a protected entity in the first content, and the processing module 840 is further configured to: perform a blurring process on the protected entity in the first content based on the position information to obtain a second content corresponding to the first content.
[0122] In some embodiments, the processing module 840 is further configured to: determine the protection level of the protected entity; perform a blurring process on the protected entity in the first content based on the position information and a blurring process strategy corresponding to the protection level to obtain the second content.
[0123] In some embodiments, the processing module 840 is further configured to: based on the location information, replace the protected entity in the first content with a semantic variable indicating the semantics of the protected entity to obtain a second content.
[0124] In some embodiments, the second detection module 820 is further configured to: based on the first content and the reference text content associated with the first content, use a first machine learning model to determine type information indicating the entity type corresponding to the entity appearing in the first content.
[0125] In some embodiments, the first machine learning model includes a first encoder, a second encoder, and a decoder, and the second detection module 820 is further configured to: based on the first content, perform feature encoding using the first encoder to generate a first feature representation; based on the reference text content, perform feature encoding using the second encoder to generate a second feature representation; and based on the first feature representation and the second feature representation, perform decoding using the decoder to determine the type information.
[0126] In some embodiments, the reference text content includes at least one of the following: a text sequence, a prompt word for the first machine learning model, or input text content associated with the first content.
[0127] In some embodiments, the first detection module 810 is further configured to: determine at least one content segment in the first content, each content segment including at least one content unit; based on performing feature extraction on the at least one content segment, determine at least one third feature representation; and based on the at least one third feature representation, use a fourth machine learning model to determine a text sequence and location information indicating the position of each of the multiple content units in the first content.
[0128] In some embodiments, the first content includes a first image, at least one content segment includes at least one image region in the first image, and the first detection module 810 is further configured to: for a first image region among the at least one image region, based on the third feature representation corresponding to the first image region, use a fourth machine learning model to determine the first position of each of the at least one content unit in the first image region in the first image region; and based on the first position of each of the at least one content unit, determine the second position of each of the at least one content unit in the first image to obtain location information indicating the second position.
[0129] In some embodiments, the multiple content units respectively correspond to multiple characters in the text sequence.
[0130] In some embodiments, the determining module 830 is further configured to: perform feature encoding using a multilingual encoder based on the text sequence, the location information, and the type information to generate a sequence of feature representations; and determine a first entity in the first content using a second machine learning model based on the sequence of feature representations.
[0131] In some embodiments, the determining module 830 is further configured to: determine a first recognition result indicating a first entity in the first content using a second machine learning model based on the text sequence, the location information, and the type information; determine at least one second recognition result indicating the first entity in the first content based on the text sequence and at least one predetermined rule; and determine the first entity in the first content by combining the first recognition result and the at least one second recognition result.
[0132] The units and / or modules included in the apparatus 800 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units and / or modules in the apparatus 800 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0133] Figure 9 A block diagram of an electronic device 900 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 9 the electronic device 900 shown is merely exemplary and should not impose any limitation on the functions and scope of the embodiments described herein. Figure 9 The electronic device 900 shown can include or be implemented as Figure 1 a server 130, or Figure 8 the apparatus 800.
[0134] As Figure 9As shown, electronic device 900 is in the form of a general-purpose electronic device. The components of electronic device 900 may include, but are not limited to, one or more processors 910, a memory 920, a storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processor 910 may be a physical or virtual processor and is capable of performing various processes according to executable instructions stored in the memory 920. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to enhance the parallel processing ability of electronic device 900.
[0135] Electronic device 900 generally includes multiple computer storage media. Such media can be any accessible media that can be obtained by electronic device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 920 may be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 930 may be a removable or non-removable medium and may include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within electronic device 900.
[0136] Electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 9 , a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 920 may include a computer program product 925 having one or more executable instruction modules configured to perform the various methods or actions of the various embodiments of the present disclosure.
[0137] The communication unit 940 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of electronic device 900 may be implemented by a single computing cluster or multiple computer machines that are capable of communicating through a communication connection. Thus, electronic device 900 may operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.
[0138] The input device 950 may be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 960 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 900 may also communicate with one or more external devices (not shown) as needed through the communication unit 940. The external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 900, or communicate with any device that enables the electronic device 900 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0139] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer-executable instruction product, the computer-executable instruction product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.
[0140] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer-executable instruction products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable executable instructions.
[0141] These computer-executable instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. These computer-executable instructions may also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams.
[0142] Computer-executable instructions can be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, thereby enabling the instructions executed on the computer, other programmable data processing apparatus, or other devices to implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0143] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-executable instruction products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, an executable instruction, or a portion of an instruction that contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0144] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled persons in the art to understand the various implementation manners disclosed herein.
Claims
1. A content processing method, comprising: Detecting a text sequence corresponding to a plurality of content units in a first content and position information indicating positions of the plurality of content units in the first content; Using a first machine learning model to detect an entity type corresponding to an entity appearing in the first content to obtain type information indicating the entity type, where the entity is represented by at least one content unit; Based on the text sequence, the position information, and the type information, using a second machine learning model to determine a first entity to be processed in the first content, where the first entity is represented by at least one content unit; And Performing a predetermined process on the first entity in the first content based on the position information to obtain a second content corresponding to the first content.
2. The method according to claim 1, further comprising: Providing the second content to a specific terminal device, or Based on the second content, using a third machine learning model to determine a first response to a first question regarding the first content.
3. The method according to claim 2, wherein determining the first response comprises: Based on the second content, using the third machine learning model to determine a second response including the first entity after a predetermined process; And Performing a restoration process on the first entity after the predetermined process in the second response to obtain the first response including the first entity.
4. The method according to claim 1, wherein the first entity includes a protected entity in the first content, and wherein performing a predetermined process on the first entity in the first content based on the position information comprises: Performing a blurring process on the protected entity in the first content based on the position information to obtain the second content corresponding to the first content.
5. The method according to claim 4, wherein performing a blurring process on the protected entity in the first content based on the position information comprises: Determining a protection level of the protected entity; Based on the position information and a blurring process strategy corresponding to the protection level, performing a blurring process on the protected entity in the first content to obtain the second content.
6. The method according to claim 4, wherein performing a blurring process on the protected entity in the first content based on the position information comprises: Based on the position information, using a semantic variable indicating the semantics of the protected entity to replace the protected entity in the first content to obtain the second content.
7. The method according to claim 1, wherein using a first machine learning model to detect an entity type corresponding to an entity appearing in the first content comprises: Based on the first content and reference text content associated with the first content, using the first machine learning model to determine type information indicating the entity type corresponding to the entity appearing in the first content.
8. The method according to claim 7, wherein the first machine learning model includes a first encoder, a second encoder, and a decoder, and determining the type information comprises: Based on the first content, perform feature encoding using the first encoder to generate a first feature representation; Based on the reference text content, perform feature encoding using the second encoder to generate a second feature representation; and Based on the first feature representation and the second feature representation, perform decoding using the decoder to determine the type information.
9. The method according to claim 7, wherein the reference text content includes at least one of the following: the text sequence, a prompt for the first machine learning model, or input text content associated with the first content.
10. The method according to claim 1, wherein detecting the text sequence and the position information includes: determining at least one content segment in the first content, each content segment including at least one content unit; determining at least one third feature representation based on performing feature extraction on the at least one content segment; and based on the at least one third feature representation, using a fourth machine learning model to determine the text sequence and the position information indicating the respective positions of the multiple content units in the first content.
11. The method according to claim 10, wherein the first content includes a first image, the at least one content segment includes at least one image region in the first image, and determining the position information includes: for a first image region among the at least one image region, based on the third feature representation corresponding to the first image region, using the fourth machine learning model to determine the respective first positions of at least one content unit in the first image region in the first image region; and based on the respective first positions of the at least one content unit, determining the respective second positions of the at least one content unit in the first image to obtain position information indicating the second positions.
12. The method according to claim 1, wherein the multiple content units respectively correspond to multiple characters in the text sequence.
13. The method according to claim 1, wherein determining the first entity in the first content includes: performing feature encoding using a multilingual encoder based on the text sequence, the position information, and the type information to generate a sequence of feature representations; and based on the sequence of feature representations, using the second machine learning model to determine the first entity in the first content.
14. The method according to claim 1, wherein determining the first entity in the first content includes: using the second machine learning model based on the text sequence, the position information, and the type information to determine a first recognition result indicating the first entity in the first content; determining at least one second recognition result based on the text sequence and at least one predetermined rule, each second recognition result indicating the first entity in the first content; and determining the first entity in the first content by combining the first recognition result and the at least one second recognition result.
15. An apparatus for content processing, comprising: A first detection module, configured to detect a text sequence corresponding to a plurality of content units in a first content and position information indicating positions of the plurality of content units in the first content; A second detection module, configured to use a first machine learning model to detect an entity type corresponding to an entity appearing in the first content, so as to obtain type information indicating the entity type, where the entity is represented by at least one content unit; A determination module, configured to determine a first entity to be processed in the first content based on the text sequence, the position information, and the type information, using a second machine learning model, where the first entity is represented by at least one content unit; And A processing module, configured to perform a predetermined process on the first entity in the first content based on the position information, so as to obtain a second content corresponding to the first content.
16. An electronic device, comprising: At least one processor; And At least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to perform the method according to any one of claims 1 to 14.
17. A computer-readable storage medium, having stored thereon computer-executable instructions, the computer-executable instructions being executable by a processor to implement the method according to any one of claims 1 to 14.
18. A computer program product, comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Combinable weak authenticator-based named entity identification algorithm architecture
CN112699682A
Global pointer network nested entity identification method fusing vocabulary information
CN117034935A
Conversation processing method and device, equipment, storage medium and program product
CN118964547A
Method and device for content recommendation, equipment and storage medium
CN119003876A
Information processing method and device, equipment, storage medium and program product
CN119719302A