Content processing methods, apparatus, devices and storage media

By detecting the text sequence and location information of content units and using machine learning models to identify entity types, this technology solves the problem of insufficient accuracy in entity detection in existing technologies, achieves highly accurate and adaptive data processing, and ensures data security and privacy protection.

CN120296770BActive Publication Date: 2026-01-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510579287.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2026-01-30
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The accuracy of entity detection and processing in existing technologies still needs to be improved, especially in the era of big data, where data security and privacy protection face challenges during data use or transmission.

Method used

By detecting the text sequence and location information of content units, a machine learning model is used to identify entity types. Combining location and type information, the entity to be processed is determined, and predetermined processing is performed to obtain the processed content.

Benefits of technology

It improves the accuracy of entity recognition and processing, adapts to different application scenarios and systems, maintains high accuracy and adaptability, and ensures data security and privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296770B_ABST
    Figure CN120296770B_ABST
Patent Text Reader

Abstract

Embodiments of this disclosure provide a content processing method, apparatus, device, and storage medium. The method includes: detecting a text sequence corresponding to a plurality of content units in first content and positional information indicating the positions of the plurality of content units in the first content; detecting entity types corresponding to entities appearing in the first content using a first machine learning model to obtain type information indicating the entity type, wherein the entity is represented by at least one content unit; determining a first entity to be processed in the first content using a second machine learning model based on the text sequence, positional information, and type information, wherein the first entity is represented by at least one content unit; and performing predetermined processing on the first entity in the first content based on the positional information to obtain second content corresponding to the first content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to content processing methods, apparatus, devices, computer-readable storage media, and computer-executable instructions products. Background Technology

[0002] With the advent of the big data era, data has become the primary carrier of information recording and transmission. However, at the same time, data security and privacy protection issues arise during data use and transmission. Clearly, it is essential to accurately remove sensitive information from raw data through appropriate methods before use or transmission. Summary of the Invention

[0003] In a first aspect of this disclosure, a content processing method is provided. The method includes: detecting text sequences corresponding to multiple content units in first content and positional information indicating the positions of the multiple content units in the first content; detecting entity types corresponding to entities appearing in the first content using a first machine learning model to obtain type information indicating the entity type, wherein the entity is represented by at least one content unit; determining a first entity to be processed in the first content using a second machine learning model based on the text sequence, positional information, and type information, wherein the first entity is represented by at least one content unit; and performing predetermined processing on the first entity in the first content based on the positional information to obtain second content corresponding to the first content.

[0004] In a second aspect of this disclosure, an apparatus for content processing is provided. The apparatus includes: a first detection module configured to detect text sequences corresponding to multiple content units in first content and positional information indicating the positions of the multiple content units in the first content; a second detection module configured to detect entity types corresponding to entities appearing in the first content using a first machine learning model to obtain type information indicating the entity type, wherein the entity is represented by at least one content unit; a determination module configured to determine a first entity to be processed in the first content based on the text sequence, positional information, and type information, using the second machine learning model, wherein the first entity is represented by at least one content unit; and a processing module configured to perform predetermined processing on the first entity in the first content based on the positional information to obtain second content corresponding to the first content.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of this disclosure, a computer program product is provided, including computer-executable instructions, wherein when executed by a processor, the computer-executable instructions implement the method according to a first aspect of this disclosure.

[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;

[0011] Figure 2 A schematic diagram of an example architecture for content processing according to some embodiments of this disclosure is shown;

[0012] Figure 3 A schematic diagram of an example architecture for content processing according to some embodiments of this disclosure is shown;

[0013] Figure 4 A schematic diagram of an example architecture for content processing according to some embodiments of this disclosure is shown;

[0014] Figure 5A A schematic diagram illustrating an example scenario of location information detection according to some embodiments of the present disclosure is shown;

[0015] Figure 5B A schematic diagram illustrating an example scenario of location information detection according to some embodiments of the present disclosure is shown;

[0016] Figure 6 A schematic diagram of an example architecture for content processing according to some embodiments of this disclosure is shown;

[0017] Figure 7 A flowchart illustrating a content processing procedure according to some embodiments of the present disclosure is shown;

[0018] Figure 8A schematic structural block diagram of an example apparatus for content processing according to some embodiments of the present disclosure is shown; and

[0019] Figure 9 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0020] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0021] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0022] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0025] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information, thereby enabling the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers or storage media that perform the operation of the technical solution disclosed herein, based on the prompt message.

[0026] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0027] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0028] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0029] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.

[0030] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating its parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as the input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. In the application phase, the model can be used to process actual inputs based on the trained parameter values ​​to determine the corresponding output.

[0031] As mentioned above, with the advent of the big data era, data has become the primary carrier of information recording and transmission. However, this also presents challenges related to data security and privacy protection during data use and transmission. For example, Large Language Models (LLMs) have achieved unprecedented efficiency and quality improvements in natural language processing, information retrieval, and intelligent question answering. However, utilizing LLMs for data processing also brings challenges to data security and privacy protection. Clearly, it is essential to accurately remove protected entities from the original data through appropriate methods before use or transmission. However, the accuracy of entity detection and processing in traditional technologies still needs improvement.

[0032] In view of this, embodiments of this disclosure propose an improved content processing scheme. In this scheme, text sequences corresponding to multiple content units in a first content are detected, and the positions of the multiple content units in the first content are also detected to obtain positional information indicating these positions. A first machine learning model is used to detect the entity types corresponding to entities appearing in the first content to obtain type information indicating the entity types. An entity is represented by at least one content unit. Based on the text sequence, positional information, and type information, a second machine learning model is used to determine a first entity to be processed in the first content, the first entity being represented by at least one content unit. Based on the positional information, predetermined processing is performed on the protected entity in the first content to obtain second content corresponding to the first content.

[0033] In the embodiments of this disclosure, not only is the text sequence corresponding to the content detected, but also the position of the content unit and the entities appearing in the content. By combining positional information and type information indicating the entity type, the model (i.e., the second machine learning model) can capture the contextual relationships and order of entities in the text sequence, thereby accurately identifying the boundaries and types of the first entity, which helps improve the accuracy of the identification and processing of the first entity. Moreover, by configuring the entity types detected by the first machine learning model, this solution can maintain high accuracy and adaptability in different application scenarios or different systems.

[0034] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0035] Example Environment

[0036] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, an application 120 is installed on a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110.

[0037] In some embodiments of this disclosure, application 120 can be any suitable application with human-computer dialogue functionality, information query functionality, or data transmission functionality. For example, application 120 can provide a digital assistant for human-computer dialogue. This digital assistant supports text dialogue services, voice dialogue services, and content dialogue in other modalities with user 140. As another example, application 120 can provide a digital assistant for information querying. This digital assistant supports user 140 in making query requests using text, voice, or other modalities.

[0038] In some embodiments, if application 120 is active, terminal device 110 may display the user interface 150 of application 120. User interface 150 may include various pages that application 120 can provide, such as a user-digital assistant dialogue page, an information retrieval interface, etc. In some embodiments, terminal device 110 may display text or images in user interface 150, and may also play audio or video in user interface 150. The audio may be, for example, the voice of user 140, a response to the voice of user 140, or audio corresponding to retrieved text content.

[0039] In some embodiments, application 120 or its digital assistant may utilize machine learning model 160 (which may include one or more machine learning models, such as machine learning model 160-1, machine learning model 160-2, ..., machine learning model 160-N, etc., where N is a positive integer. For ease of description, the one or more machine learning models are collectively referred to as machine learning model 160 herein) to support interaction with user 140. For example, application 120 or its digital assistant may utilize one or more machine learning models 160 to provide information query services to user 140. As another example, application 120 or its digital assistant may also utilize one or more machine learning models 160 to provide question-and-answer services to user 140.

[0040] In some embodiments, the machine learning model 160 can be of different types. In some embodiments, one or more machine learning models 160 can be built based on a language model (LM). The machine learning model used is a content-generative model, capable of generating corresponding outputs based on model inputs. In some embodiments, the language model-based machine learning model can handle textual modal model inputs (e.g., natural language and / or machine language) and / or non-textual modal model inputs (e.g., images, speech, video, etc.), and can generate the desired output based on the model inputs and prompt words. Here, prompt words are used to guide the machine learning model to generate user queries that can resolve the user queries indicated by the model inputs. In application scenarios that support user dialogue, user 140's input can be provided to the machine learning model 160 as at least a part of the model inputs (other parts may include prompt words). This user input is treated as a question or query request. Based on the model output, a corresponding response can be provided to user 140.

[0041] In some embodiments, one or more machine learning models 160 may be speech-related models, including speech recognition (ASR) models and text-to-speech (TTS) models. The input to an ASR model is speech, and the output is text. The input to a TTS model is text, and the output is the corresponding speech.

[0042] In some embodiments, terminal device 110 may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may also support any type of user-facing interface (such as "wearable" circuitry).

[0043] In some embodiments, terminal device 110 can communicate with server 130 to provide services to application 120. For example... Figure 1As shown, server 130 can invoke machine learning model 160 to support application 120 in providing human-computer dialogue services or information query services to user 140 based on the output of machine learning model 160, etc. In some embodiments, server 130 may include, but is not limited to, mainframes, edge computing nodes, computing devices in a cloud environment, etc. It can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Servers may include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in a cloud environment, etc.

[0044] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0045] Example Scheme

[0046] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure. Figure 2 A schematic diagram of an example architecture 200 for content processing according to some embodiments of the present disclosure is shown. For ease of discussion, the example architecture 200 is described below from the perspective of server 130, but this is merely exemplary.

[0047] In box 210, server 130 can detect the text sequence corresponding to multiple content units in the first content and the positional information indicating the position of the multiple content units in the first content. The first content can be understood as the content that needs to be detected and processed as a first entity. Here, the first entity can be any entity in the first content that needs to be processed, and the specific type of the first entity is related to the processing method of the first entity.

[0048] In some embodiments, the first entity may be a protected entity within the first content. A protected entity can be understood as a portion of the overall content that requires protection, such as an organization name, place name, time, value, etc., that needs protection according to specific management regulations, personal wishes, etc. A protected entity may include one or more named entities within the content. It is understood that the first entity is not limited to protected entities. Depending on the application scenario, the first entity may be of different types. For example, the first entity may also be an entity within the first content that needs to be expanded, rewritten, deleted, or summarized. It should also be noted that protected entities are not limited to the entities described above. For example, in a specific industry or field, there may be specific named entities that are unique to that industry or field and require protection.

[0049] The first content can include data in various modalities, such as text, images, audio, video, and so on. In some examples, such as... Figure 2 As shown, application 120 may have human-computer dialogue functionality or information retrieval functionality. The first content may include dialogue content or query requests input by user 140 to be provided to machine learning model 160-1 (sometimes referred to herein as a third machine learning model). Server 130 may obtain the first content from application 120. In other examples, the first content may be content that one terminal device will transmit to another terminal device; for example, the first content may be content that a terminal device within an organization will transmit to a terminal device outside the organization. It is understood that the above-described first content is merely exemplary. In practical application scenarios, any appropriate content requiring first entity detection can be determined as the first content according to actual needs. The embodiments of this disclosure do not limit this.

[0050] A content unit is a constituent unit of the first content. The specific content of a content unit may differ depending on the modality of the first content, and the corresponding positional information indicating the content unit's location within the first content may also differ. In some examples, if the first content includes text, the content unit can be a word, phrase, character, etc., within the text. Positional information can indicate the position of the content unit (word, phrase, character, etc.) within the text. For example, if the content unit includes a character, the positional information can include the character's page number, paragraph number, line number, column number, or ordinal number within its row, column, or paragraph. Positional information can also include the character's coordinates within the text.

[0051] In some examples, if the first content includes an image, the content unit may include image regions, objects, pixel units, etc., within the image. Position information may indicate the location of the image regions, objects, or pixel units within the image. As an example, the content unit may include text objects (e.g., characters) within the image, and the position information may indicate the position of the character within the image. This position information may include, for example, the coordinates of at least one reference point of the character (e.g., the center point or a vertex of the bounding rectangle), its dimensions (e.g., dimensions in the length and width directions), rotation angle (e.g., rotation angle relative to a coordinate axis in the image coordinate system), and so on.

[0052] In some examples, if the first content includes audio, the content unit may include audio frames, speech segments, etc., within the audio. Location information may indicate the position of the audio frame or speech segment within the audio; for example, location information may include the timestamp of the speech segment or audio frame. It is understood that the above-described content units and location information are exemplary, and the embodiments of this disclosure are not intended to limit them.

[0053] In some embodiments, the text sequence may be a sequence of text identified from multiple content units. For example, the text sequence may include a sequence of text identified from content such as text, images, audio, and video. In other examples, the text sequence may be a sequence of text associated with multiple content units. As examples, the text sequence may be a user-inputted text sequence associated with content units, a text sequence used to describe content units (e.g., a text sequence used to describe objects in an image), and so on.

[0054] In some embodiments, server 130 may use machine learning model 160-2 (sometimes referred to herein as a fourth machine learning model) to detect text sequences and location information based on the first content. Machine learning model 160-2 may be trained to detect text sequences and the location of content units within the content. In some examples, machine learning model 160-2 may be trained based on a Connectionist Temporal Classification (CTC) model. Of course, machine learning model 160-2 may also be trained based on any other suitable model architecture, and the embodiments of this disclosure are not limited thereto.

[0055] As an example, Figure 3 A schematic diagram of an example structure 300 for content processing according to some embodiments of the present disclosure is shown. Figure 3 As shown, server 130 may be deployed with interface 310, preprocessing module 320, and detection engine 330. Server 130 can receive requests for first content to perform first entity detection through interface 310. Before detecting text sequences and location information, server 130 can preprocess the first content using preprocessing module 320. Specifically, if the first content includes text, in box 322, server 130 can preprocess the text, for example, server 130 can convert the text into text conforming to a predetermined format. If the first content includes an image, in box 324, server 130 can preprocess the image, for example, by denoising, enhancing, or extracting image regions. If the first content includes audio (not shown in the figure), server 130 can perform noise reduction, filtering, or other processing on the audio. After preprocessing the first content, server 130 can provide the preprocessed first content to detection engine 330. Detection engine 330 can use machine learning model 160-2 to detect text sequences and location information.

[0056] In some embodiments, server 130 may determine at least one content segment in the first content, each content segment including at least one content unit. Server 130 may determine at least one feature representation (sometimes referred to herein as a third feature representation) corresponding to the at least one content segment based on feature extraction performed on the at least one content segment. Server 130 may then provide the at least one feature representation to machine learning model 160-2 to obtain a text sequence output by machine learning model 160-2 and positional information indicating the respective positions of the multiple content units in the first content.

[0057] Figure 4 A schematic diagram of an example architecture 400 for content processing according to some embodiments of this disclosure is shown. Figure 4 As shown, the first content may include a first image 402, and the content unit may include characters in the first image 402. The content unit may be an image region (i.e., a text block) containing text in the first image 402. Server 130 may perform text block detection (404) on the first image 402 to determine the image region 406 containing text in the first image 402. Server 130 may provide the image region 406 to a trained feature extractor 408, and use the feature extractor 408 to perform feature extraction on the image region 406 to obtain the feature representation corresponding to the image region 406. Then, server 406 may provide the feature representation to machine learning model 160-2 to obtain a first detection result 410. The first detection result 410 may indicate the text sequence corresponding to the image region 406, and the position of each character in the text sequence in the first image 402.

[0058] In some embodiments, for a first image region within at least one image region, server 130 can determine a first position of at least one content unit within the first image region using machine learning model 160-2, based on a feature representation (i.e., a third feature representation) corresponding to the first image region. Then, server 130 can determine a second position of at least one content unit within the first image region based on the first position of each content unit, to obtain positional information indicating the second position. That is, server 130 can use machine learning model 160-2 to determine the first position of the content unit within the image region. Then, server 130 can perform a position transformation on the first position based on the positional relationship between the image region and the first image to determine the second position of the content unit within the first image.

[0059] As an example, Figure 5AA schematic diagram of an example scenario 500A for location information detection according to some embodiments of the present disclosure is shown. In example scenario 500A, feature extractor 408 may be trained based on a convolutional neural network (CNN), and machine learning model 160-2 may be trained based on a CTC model. Server 130 may use feature extractor 408 to perform feature extraction on image region 406 to obtain a feature map corresponding to image region 406. Then, server 130 may provide the feature map to machine learning model 160-2, and machine learning model 160-2 may output a probability matrix 510. Each element 512 in probability matrix 510 may indicate the probability value of a corresponding sub-region in image region 406 belonging to each character from A to Z. Based on the probability matrix 510, server 130 may determine the text sequence (i.e., "ABCD") corresponding to image region 406 and the first position of each character in the text sequence in image region 406.

[0060] As an example, Figure 5B A schematic diagram of an example scenario 500B for location information detection according to some embodiments of the present disclosure is shown. Figure 5B As shown, server 130 can determine the image coordinates (x, y) of each character in image region 406. l ,g l The coordinates (X) under ) l ,y l Then, server 130 can use the following formula to calculate the coordinates (X...) l ,Y l Transform to the image coordinate system (x) of the first image i ,y i )middle.

[0061]

[0062] Where (X) l ,Y l ) represents the character's coordinates in the image region 406 (x) l ,y l The coordinates in (X) indicate the first position of the character; I ,Y I ) represents the character in the image coordinate system (x) of the first image. i y i The coordinates in () indicate the second position of the character; X t and Y t θ represents the translation parameter, and θ represents the rotation angle.

[0063] Returning to example architecture 200, in box 220, server 130 can utilize machine learning model 160-3 (sometimes referred to as the first machine learning model in this document) to detect the entity types corresponding to entities appearing in the first content, in order to obtain type information indicating the entity types. An entity can be understood as a named entity appearing in the first content; for example, the type corresponding to an entity may include organization name, location name, time, value, product name, etc. An entity is represented by at least one content unit; for example, an entity of type organization can be formed by the four characters "ABCD".

[0064] In some embodiments, server 130 may use machine learning model 160-3 to determine type information indicating the entity type corresponding to an entity appearing in the first content, based on first content and reference text content associated with the first content. The reference text content may include text content associated with the first content that can provide a reference for machine learning model 160-3 to determine the entity type of the entity. This helps improve the accuracy of entity and entity type detection.

[0065] In some examples, the reference text content may include the input text content associated with the first text. For instance, user 140 may input a first image into the dialog box of application 120. Simultaneously, user 140 may also input, for example, "XXXXX in the image is YYYYY, please tell me ZZZZZ?". In this case, server 130 can determine the text content input by user 140 as the reference text content. Server 130 can determine the entity type corresponding to the entity appearing in the first image based on the first image and the text content input by user 140.

[0066] Alternatively or additionally, the reference text content may also include prompts for the machine learning model 160-3. As an example, if the first content includes a first image, the server 130 may select prompts to instruct the machine learning model 160-3 to detect entities from the image, such as "You are an image entity recognition expert, ... please return the entities appearing in the image and their corresponding entity types." Based on the prompts and the first image, the server 130 may use the machine learning model 160-3 to determine the entity types corresponding to the entities appearing in the first image.

[0067] Alternatively or additionally, the reference text content may also include a text sequence corresponding to the first content. As an example, server 130 may use machine learning model 160-2 to identify a text sequence from audio. Then, server 130 may determine the entity type corresponding to the entity appearing in the audio based on the audio and the identified text sequence. It should be noted that the above-described reference text content is merely exemplary, and other text content associated with the first content and capable of providing a reference for entity detection can be selected according to actual needs. The embodiments of this disclosure do not limit this.

[0068] In some embodiments, the machine learning model 160-3 may include a first encoder, a second encoder, and a decoder. The server 130 may perform feature encoding using the first encoder based on first content to generate a first feature representation. The server 130 may also perform feature encoding using the second encoder based on reference text content to generate a second feature representation. Subsequently, the server 130 may perform decoding using the decoder based on the first and second feature representations to determine type information.

[0069] In some examples, such as Figure 4 As shown, the first content may include a first image 402. The machine learning model 160-3 may include a first encoder 412 (also called a visual encoder), a second encoder (also called a text encoder), and a decoder. The server 130 may use the first encoder 412 to perform feature encoding on the first image 402 to generate a first feature representation 414. As an example, the server 130 may use the first encoder 412 to encode the first image 402 into a high-dimensional vector representation.

[0070] Server 130 can also use a second encoder 418 to perform feature encoding on the reference text content 416 to generate a second feature representation 420. As an example, server 130 can use text content input by user 140, prompt words, and a text sequence detected from the first image 402 as reference text content. Server 130 can use the second feature encoder 418 to encode the reference text content 416 into a second feature representation 420 that is aligned with the first feature representation 414.

[0071] In box 422, server 130 can concatenate the first feature representation 414 and the second feature representation 420 to obtain the input feature representation of decoder 424. Server 130 can use decoder 424 to perform feature decoding on the input feature representation to generate a second detection result 426. The second detection result can indicate the entity type corresponding to the entity appearing in the first image 402. Then, server 130 can merge the first detection result 410 and the second detection result 426 to obtain detection result 332. As an example, server 130 can use decoder 424 to perform cross-attention processing on the input feature representation to promote multimodal feature fusion, thereby improving the accuracy of entity and entity type detection. It should also be noted that the above example is only used to explain the embodiments of this disclosure by recognizing characters from images, and should not be construed as the embodiments of this disclosure being limited to recognizing text and entities composed of text from images. In practical application scenarios, similar architectures can also be used to process other modalities such as audio and video, and the embodiments of this disclosure are not limited thereto.

[0072] In some embodiments, after determining the entity type and content unit position of the entity in the first content, the server 130 may annotate the position of the entity type and content unit in the first content to facilitate subsequent predetermined processing of the first entity. As an example, the server 130 may annotate the position of the entity type and character in the first image 402.

[0073] Return to combination Figure 2 As shown in box 230, server 130 can determine the first entity to be processed in the first content based on text sequence, location information, and type information using machine learning model 160-4 (sometimes referred to as the second machine learning model in this document). The first entity has been explained in the preceding content; please refer to the description therein for details. It should be added here that the first entity can be represented by at least one content unit, such as an organization name, place name, etc., composed of multiple characters.

[0074] The first entity here may be the same as or different from the entities mentioned above. In some examples, the first entity may include one or more entities. As an example, server 130 may identify "XX City" and "XXX County" as entities of type "place name". Server 130 may also identify "XX City XXX County" as the first entity. In this case, the first entity actually includes two entities of type "place name".

[0075] In some examples, the first entity may also include a part of the entity. As an example, the server 130 may determine "YYYY / MM / DD" as an entity of the entity type "date". In box 230, the server 130 may only determine the specific year "YYYY", month "MM", and date "DD" as the first entity. In this way, it is beneficial to maintain the data format of this "date" entity and the original semantics in the content.

[0076] In some examples, the first entity may also include at least one content unit of the entity and the neighboring entity. For example, in the case where the recognition of the entity is incomplete, a part of the characters of the neighboring entity may be supplemented to the entity, and the supplemented entity is determined as the first entity.

[0077] In some examples, the first entity may include a protected entity. The machine learning model 160-3 may not determine a part of the content segment in the first content as an entity. However, the machine learning model 160-4 may combine the context semantics of the text sequence and determine that this part of the content segment needs to be protected, and determine this part of the content segment as a protected entity.

[0078] In some embodiments, as Figure 3 shown, the server 130 may also be deployed with an identification engine 340. The identification engine 340 may obtain a detection result 332 from the detection engine 330, and the detection result 332 may include a text sequence, location information, and type information. The identification engine 340 may perform feature encoding on the text sequence, location information, and type information by using a multilingual encoder to generate a feature representation sequence. Then, the machine learning model 160-4 is used to determine the first entity in the first content based on the feature representation sequence.

[0079] Figure 6 FIG. shows a schematic diagram of an example architecture 600 for content processing according to some embodiments of the present disclosure. As Figure 6 shown, the server 130 may perform word segmentation processing on the text sequence 602 by using a word segmenter 604 to obtain a sub-word unit sequence 606 composed of multiple sub-word units. In some examples, the word segmenter 604 may be a multilingual word segmenter (MultilingualTokenizer). In this way, the text sequence composed of multiple languages can be processed.

[0080] In box 608, server 130 can perform embedding processing on the sub-word unit sequence 606, positional information, and type information to generate an embedding representation sequence 610. As an example, server 130 can determine the lexical embedding representation corresponding to each sub-word unit in the sub-word unit sequence 606. Server 130 can also determine the positional embedding representation and type embedding representation corresponding to each sub-word unit based on the positional and type information. Then, server 130 can combine the lexical embedding representation, positional embedding representation, and type embedding representation to form a high-dimensional tensor representation sequence to improve the recognizability of proper nouns and entity boundaries.

[0081] Server 130 can use multilingual encoder 342 to perform feature encoding on embedding representation sequence 610, generating feature representation sequence 612. The feature representation sequence includes multiple feature representations corresponding to multiple sub-word units in sub-word unit sequence 606 (e.g., embedding representation sequence 610 can be encoded as a high-dimensional vector representation sequence). Server 130 can use machine learning model 160-4 to determine a label sequence based on feature representation sequence 612. The label sequence can include multiple labels corresponding to multiple sub-word units in sub-word unit sequence 606, each label indicating the entity type of the corresponding sub-word unit. Specifically, each label indicates whether the corresponding sub-word unit belongs to a protected entity (i.e., the first entity). If the corresponding sub-word unit belongs to a protected entity, the label can indicate the entity type of the corresponding sub-word unit.

[0082] In some examples, machine learning model 160-4 can be trained based on a Conditional Random Field (CRF) model. Machine learning model 160-4 can be trained to detect the first entity in the content. Utilizing the maximum likelihood algorithm of the CRF model can optimize entity type labeling, determine entity boundaries by combining contextual semantics, and help reduce entity type labeling errors caused by illogical relationships between adjacent labels, thereby improving the detection accuracy of the first entity. Of course, the above model structure is merely exemplary, and machine learning model 160-4 can adopt any suitable model structure; the embodiments of this disclosure do not limit this.

[0083] In some embodiments, server 130 can determine a first identification result indicating a first entity in the first content based on text sequence, location information, and type information using machine learning model 160-4. Server 130 can also determine at least one second identification result based on text sequence and at least one predetermined rule, each second identification result indicating a first entity in the first content. Then, server 130 can determine the first entity in the first content by merging the first identification result and at least one second identification result. That is, not only is the first entity in the first content detected using machine learning model 160-4, but also based on predetermined rules, and the final first entity is determined by combining the detection results of both. This not only improves the accuracy of the detection results, but also allows the improved scheme of this disclosure to be flexibly applied to different application scenarios or business systems. Moreover, in practical application scenarios, maintaining predetermined rules (e.g., a thesaurus) facilitates flexible adjustments to the first entity by the business side.

[0084] As an example, such as Figure 6 As shown, server 130 can determine a first recognition result 616 based on the label sequence 614 output by machine learning model 160-4. Server 130 can use rule engine 618 to identify protected entities in text sequence 602 based on a first rule (e.g., keyword matching or regular expression matching) to determine a second recognition result 620-1. Server 130 can also use rule engine 618 to identify protected entities in text sequence 602 based on a second rule (e.g., a thesaurus) to determine a second recognition result 620-2. Then, server 130 can combine the first recognition result 616, the second recognition result 620-1, and the second recognition result 620-2 to determine the final recognition result 622.

[0085] In some embodiments, the first identification result may have a corresponding priority, and the at least one second identification result may also have a corresponding priority. Server 130 can determine the first entity in the first content based on the first identification result, the priority corresponding to the first identification result, the at least one second identification result, and the priority corresponding to each of the at least one second identification result. Of course, server 130 is not limited to merging the first identification result and the at least one second identification result based on priority; server 130 can determine the final identification result as the intersection or union of the first identification result and the at least one second identification result. The embodiments of this disclosure do not limit the method of merging the first identification result and the second identification result.

[0086] In some embodiments of this disclosure, server 130 may perform predetermined processing on a first entity in the first content based on location information to obtain second content corresponding to the first content. The second content may include the predetermined processed first entity and content units in the first content that do not belong to the first entity.

[0087] The pre-processing here may include various processing operations on the first entity. In some embodiments, the first entity may be a protected entity in the first content. Server 130 may perform obfuscation processing on the protected entity in the first content based on location information. Obfuscation processing may include, but is not limited to, performing hashing, masking, replacement, or encryption on the protected entity. As an example, such as Figure 3 As shown, server 130 can utilize, for example, a blurring engine 350 to perform blurring on protected entities in the first content. Specifically, if the first content includes text, in box 352, server 130 can perform blurring on protected entities in the text, such as hashing, masking, replacement, or encryption. If the first content includes an image, in box 354, server 130 can perform blurring on protected entities in the image, such as occlusion or masking.

[0088] In some examples, if a content unit corresponds to a character in the first content, the location information can indicate the position of each character in the first content. Server 130 can perform obfuscation on the protected entities in the first content at the character level, which helps improve the accuracy of obfuscation of the protected entities.

[0089] In some embodiments, server 130 may determine the protection level of a protected entity. Then, server 130 may perform obfuscation on the protected entity in the first content based on location information and an obfuscation strategy corresponding to the protection level to obtain second content. As an example, for protected entities with relatively high protection levels, server 130 may perform, for example, encryption. For protected entities with relatively low protection levels, server 130 may perform, for example, obscuring or replacement.

[0090] In some embodiments, if the first content is provided to the machine learning model 160-1, the server 130 can replace the protected entity in the first content with a semantic variable indicating the semantics of the protected entity based on location information to obtain the second content. That is, the semantics indicated by the semantic variable can be the same as or similar to the semantics of the protected entity. In this way, the obtained second content can maintain the original semantics to a certain extent, which is beneficial for the machine learning model 160 to accurately generate responses.

[0091] It should be noted that the pre-processing is not limited to fuzzy processing. Server 130 can also perform any appropriate processing on the first entity, such as rewriting, expanding, deleting, or summarizing. The embodiments of this disclosure do not limit the processing method of the pre-processing. It should also be noted that after performing the pre-processing, server 130 is not limited to belonging to the second content; server 130 can also output, for example, a processing report to record the identified first entity and the pre-processing performed on the first entity. This facilitates subsequent performance of, for example, restoration processing, on the pre-processed first entity.

[0092] Combination Figure 2 As shown, server 130 can determine a first response to a first question regarding the first content based on the second content and using machine learning model 160-1. Specifically, server 130 can provide the second content to machine learning model 160-1 to obtain the model output generated by machine learning model 160-1. Server 130 can determine a first response to the first question regarding the first content based on the model output. Server 130 can then feed back the first response to terminal device 110. Terminal device 110 can present the first response using the dialog interface of application 120. If the first response includes audio or video, terminal device 110 can also play the audio or video using the dialog interface of application 120.

[0093] In some embodiments, server 130 may, based on the second content, utilize machine learning model 160-1 to determine a second response containing a pre-processed first entity, such as a second response containing a blurred protected entity. Then, server 130 performs a restoration process on the pre-processed first entity in the second response to obtain a first response containing the first entity. In some examples, such as... Figure 2 As shown, in box 250, server 130 can perform format validation on content units other than the pre-processed first entity in the second response, such as validating whether these content units conform to a pre-defined format, or validating whether the format of these content units matches the format of the corresponding content units in the first content. In box 260, server 130 can perform restoration processing on the pre-processed first entity in the validated second response, such as replacing it back with a protected entity, or decrypting the obfuscated protected entity, etc. In box 270, server 130 can also perform formatting processing on the first response, such as processing the first response into a data format that matches the dialog interface of application 120. Afterwards, server 130 can send the formatted first response back to terminal device 110. This allows user 140 to conveniently view the response containing complete content.

[0094] It should be noted that the above example structure 200 is merely an illustrative example of the improved solution of the embodiments of this disclosure, using a human-computer dialogue scenario as an example. The improved solution of the embodiments of this disclosure is not limited to human-computer dialogue scenarios, but can also be applied to other application scenarios that require processing of entities in content, such as information retrieval or data transmission.

[0095] As an example, if a first terminal device needs to send first content to a second terminal device, the first terminal device can provide the first content to server 130. Server 130 detects protected entities in the first content and performs obfuscation processing on the protected entities in the first content to obtain second content. The second content can then be provided to the second terminal device.

[0096] As another example, if terminal device 110 triggers a pre-defined operation on the first content (e.g., screen mirroring or sharing), terminal device 110 can send a request to server 130 to request server 130 to perform protected entity detection on the first content. Server 130 can then provide the terminal device with second content. Terminal device 110 can then perform the pre-defined operation based on the second content.

[0097] In this way, in the embodiments of this disclosure, not only is the text sequence corresponding to the content detected, but also the position of the content unit and the entities appearing in the content. By combining positional information and type information indicating the entity type, the model (i.e., the second machine learning model) can capture the contextual relationships and order of entities in the text sequence, thereby accurately identifying the boundaries and types of the first entity, which is beneficial to improving the accuracy of the identification and processing of the first entity. Moreover, by configuring the entity types detected by the first machine learning model, this scheme can maintain high accuracy and adaptability in different application scenarios or different systems.

[0098] Example processes, apparatus and equipment

[0099] Figure 7 A flowchart of a content processing procedure 700 according to some embodiments of the present disclosure is shown. For ease of discussion, the procedure 700 is described hereinafter from the perspective of server 130, but this is merely exemplary.

[0100] In box 710, server 130 detects the text sequence corresponding to multiple content units in the first content and the position information indicating the position of the multiple content units in the first content.

[0101] In box 720, server 130 uses a first machine learning model to detect the entity type corresponding to the entity appearing in the first content in order to obtain type information indicating the entity type, wherein the entity is represented by at least one content unit.

[0102] In box 730, server 130 uses a second machine learning model based on text sequence, location information, and type information to determine a first entity to be processed in the first content, the first entity being represented by at least one content unit.

[0103] In box 740, server 130 performs predetermined processing on the first entity in the first content based on location information to obtain the second content corresponding to the first content.

[0104] In some embodiments, process 700 further includes: providing second content to a specific terminal device, or, based on the second content, using a third machine learning model to determine a first response to a first question in the first content.

[0105] In some embodiments, determining the first response includes: determining a second response containing a predetermined processed first entity based on the second content using a third machine learning model; and performing a restoration process on the predetermined processed first entity in the second response to obtain a first response containing the first entity.

[0106] In some embodiments, the first entity includes a protected entity in the first content, and wherein performing a predetermined process on the first entity in the first content based on location information includes: performing a blurring process on the protected entity in the first content based on the location information to obtain a second content corresponding to the first content.

[0107] In some embodiments, performing fuzzing on a protected entity in the first content based on location information includes: determining the protection level of the protected entity; and performing fuzzing on the protected entity in the first content based on the location information and a fuzzing strategy corresponding to the protection level to obtain a second content.

[0108] In some embodiments, performing fuzzing on a protected entity in the first content based on location information includes: replacing the protected entity in the first content with a semantic variable that indicates the semantics of the protected entity based on the location information to obtain a second content.

[0109] In some embodiments, detecting the entity type corresponding to an entity appearing in the first content using a first machine learning model includes: based on the first content and reference text content associated with the first content, using the first machine learning model to determine type information indicating the entity type corresponding to an entity appearing in the first content.

[0110] In some embodiments, the first machine learning model includes a first encoder, a second encoder, and a decoder, and determining type information includes: performing feature encoding using the first encoder based on first content to generate a first feature representation; performing feature encoding using the second encoder based on reference text content to generate a second feature representation; and performing decoding using the decoder based on the first feature representation and the second feature representation to determine type information.

[0111] In some embodiments, the reference text content includes at least one of the following: a text sequence, prompt words for a first machine learning model, or input text content associated with the first content.

[0112] In some embodiments, detecting text sequence and location information includes: determining at least one content segment in a first content, each content segment including at least one content unit; determining at least one third feature representation based on feature extraction performed on the at least one content segment; and determining a text sequence and location information indicating the position of each of the plurality of content units in the first content based on the at least one third feature representation using a fourth machine learning model.

[0113] In some embodiments, the first content includes a first image, at least one content fragment includes at least one image region in the first image, and determining the location information includes: for the first image region in the at least one image region, using a fourth machine learning model based on a third feature representation corresponding to the first image region, determining a first position of each of the at least one content unit in the first image region in the first image region; and determining a second position of each of the at least one content unit in the first image based on the first position of each of the at least one content unit, to obtain location information indicating the second position.

[0114] In some embodiments, multiple content units correspond to multiple characters in a text sequence.

[0115] In some embodiments, determining a first entity in the first content includes: performing feature encoding using a multilingual encoder based on a text sequence, location information, and type information to generate a feature representation sequence; and determining the first entity in the first content using a second machine learning model based on the feature representation sequence.

[0116] In some embodiments, determining a first entity in first content includes: determining a first identification result indicating a first entity in first content using a second machine learning model based on a text sequence, location information, and type information; determining at least one second identification result based on a text sequence and at least one predetermined rule, each second identification result indicating a first entity in first content; and determining the first entity in first content by merging the first identification result and at least one second identification result.

[0117] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 8 A schematic structural block diagram of an example apparatus 800 for content processing according to certain embodiments of the present disclosure is shown. Apparatus 800 may be implemented as or included in server 130. Various modules / components in apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.

[0118] like Figure 8 As shown, the apparatus 800 includes: a first detection module 810 configured to detect text sequences corresponding to multiple content units in a first content and position information indicating the positions of the multiple content units in the first content; a second detection module 820 configured to use a first machine learning model to detect entity types corresponding to entities appearing in the first content to obtain type information indicating entity types, wherein the entity is represented by at least one content unit; a determination module 830 configured to determine a first entity to be processed in the first content based on the text sequence, position information, and type information, using the second machine learning model, wherein the first entity is represented by at least one content unit; and a processing module 840 configured to perform predetermined processing on the first entity in the first content based on the position information to obtain second content corresponding to the first content.

[0119] In some embodiments, the apparatus 800 further includes: a providing module configured to provide second content to a specific terminal device, or a response module configured to determine a first response to a first question about the first content based on the second content and utilizing a third machine learning model.

[0120] In some embodiments, the response module is further configured to: determine a second response containing a predetermined processed first entity based on the second content using a third machine learning model; and perform a restoration process on the predetermined processed first entity in the second response to obtain a first response containing the first entity.

[0121] In some embodiments, the first entity includes a protected entity in the first content, and the processing module 840 is further configured to: perform fuzzing processing on the protected entity in the first content based on the location information to obtain second content corresponding to the first content.

[0122] In some embodiments, the processing module 840 is further configured to: determine the protection level of the protected entity; and perform fuzzing on the protected entity in the first content based on location information and a fuzzing strategy corresponding to the protection level to obtain the second content.

[0123] In some embodiments, the processing module 840 is further configured to: replace the protected entity in the first content with a semantic variable that indicates the semantics of the protected entity based on location information, so as to obtain the second content.

[0124] In some embodiments, the second detection module 820 is further configured to: determine type information indicating the entity type corresponding to the entity appearing in the first content, based on the first content and the reference text content associated with the first content, using a first machine learning model.

[0125] In some embodiments, the first machine learning model includes a first encoder, a second encoder, and a decoder, and the second detection module 820 is further configured to: perform feature encoding using the first encoder based on first content to generate a first feature representation; perform feature encoding using the second encoder based on reference text content to generate a second feature representation; and perform decoding using the decoder based on the first feature representation and the second feature representation to determine type information.

[0126] In some embodiments, the reference text content includes at least one of the following: a text sequence, prompt words for a first machine learning model, or input text content associated with the first content.

[0127] In some embodiments, the first detection module 810 is further configured to: determine at least one content segment in the first content, each content segment including at least one content unit; determine at least one third feature representation based on feature extraction performed on the at least one content segment; and determine a text sequence and positional information indicating the respective positions of multiple content units in the first content based on the at least one third feature representation using a fourth machine learning model.

[0128] In some embodiments, the first content includes a first image, at least one content fragment includes at least one image region in the first image, and the first detection module 810 is further configured to: for the first image region in the at least one image region, based on a third feature representation corresponding to the first image region, and using a fourth machine learning model, determine a first position of at least one content unit in the first image region; and based on the first position of at least one content unit, determine a second position of at least one content unit in the first image to obtain position information indicating the second position.

[0129] In some embodiments, multiple content units correspond to multiple characters in a text sequence.

[0130] In some embodiments, the determining module 830 is further configured to: perform feature encoding using a multilingual encoder based on the text sequence, location information, and type information to generate a feature representation sequence; and determine a first entity in the first content using a second machine learning model based on the feature representation sequence.

[0131] In some embodiments, the determining module 830 is further configured to: determine a first identification result indicating a first entity in the first content based on a text sequence, location information, and type information using a second machine learning model; determine at least one second identification result based on a text sequence and at least one predetermined rule, each second identification result indicating a first entity in the first content; and determine the first entity in the first content by merging the first identification result and at least one second identification result.

[0132] The units and / or modules included in device 800 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 800 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0133] Figure 9 A block diagram of an electronic device 900 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 9 The electronic device 900 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 9 The illustrated electronic device 900 may include or be implemented as Figure 1 Server 130, or Figure 8 Device 800.

[0134] like Figure 9As shown, electronic device 900 is in the form of a general-purpose electronic device. Components of electronic device 900 may include, but are not limited to, one or more processors 910, memory 920, storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. Processor 910 may be a physical or virtual processor and is capable of performing various processes according to executable instructions stored in memory 920. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 900.

[0135] Electronic device 900 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 920 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 930 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 900.

[0136] Electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 9 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 920 may include computer program product 925 having one or more executable instruction modules configured to perform various methods or actions of various embodiments of this disclosure.

[0137] The communication unit 940 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 900 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 900 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0138] Input device 950 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 960 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 900 can also communicate with one or more external devices (not shown) via communication unit 940 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 900, or with any device that enables electronic device 900 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0139] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer-executable instruction product is also provided, which is tangibly stored on a non-transient computer-readable medium and includes computer-executable instructions that are executed by a processor to implement the methods described above.

[0140] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer-executable instruction products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable and executable instructions.

[0141] These computer-executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-executable instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0142] Computer-executable instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-executable instruction products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, executable instruction, or portion of instructions, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0144] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A content processing method comprising: detecting text sequences corresponding to a plurality of content units in a first content and position information indicating positions of the plurality of content units in the first content; detecting, using a first machine learning model, entity types corresponding to entities appearing in the first content, the entities being represented by at least one content unit, to obtain type information indicating the entity types; determining, using a second machine learning model, a first entity in the first content to be processed based on the text sequences, the position information, and the type information, the first entity being represented by at least one content unit; and performing a predetermined processing on the first entity in the first content based on the position information to obtain a second content corresponding to the first content; wherein determining the first entity in the first content comprises: performing feature encoding using a multi-language encoder based on the text sequences, the position information, and the type information to generate a sequence of feature representations; and determining the first entity in the first content using the second machine learning model based on the sequence of feature representations.

2. The method of claim 1, further comprising: providing the second content to a specific terminal device, or determining a first answer to a first question for the first content using a third machine learning model based on the second content.

3. The method of claim 2, wherein determining the first answer comprises: determining a second answer containing a first entity after a predetermined processing using the third machine learning model based on the second content; and performing a restoration processing on the first entity after the predetermined processing in the second answer to obtain the first answer containing the first entity.

4. The method of claim 1, wherein the first entity comprises a protected entity in the first content, and wherein performing a predetermined processing on the first entity in the first content based on the position information comprises: performing a blurring processing on the protected entity in the first content based on the position information to obtain the second content corresponding to the first content.

5. The method of claim 4, wherein performing a blurring processing on the protected entity in the first content based on the position information comprises: determining a protection level of the protected entity; performing a blurring processing on the protected entity in the first content based on the position information and a blurring processing policy corresponding to the protection level to obtain the second content.

6. The method of claim 4, wherein performing a blurring processing on the protected entity in the first content based on the position information comprises: replacing the protected entity in the first content with a semantic variable indicating semantics of the protected entity based on the position information to obtain the second content.

7. The method of claim 1, wherein detecting, using a first machine learning model, entity types corresponding to entities appearing in the first content comprises: ​ ​ ​ based on the first content and the reference text content associated with the first content, determining, by the first machine learning model, type information indicating an entity type corresponding to an entity appearing in the first content.

8. The method of claim 7, wherein the first machine learning model comprises a first encoder, a second encoder, and a decoder, and determining the type information comprises: performing feature encoding based on the first content by the first encoder to generate a first feature representation; performing feature encoding based on the reference text content by the second encoder to generate a second feature representation; and performing decoding based on the first feature representation and the second feature representation by the decoder to determine the type information.

9. The method of claim 7, wherein the reference text content comprises at least one of: the text sequence, a prompt word for the first machine learning model, or input text content associated with the first content.

10. The method of claim 1, wherein detecting the text sequence and the position information comprises: determining at least one content segment in the first content, each content segment comprising at least one content unit; determining at least one third feature representation based on performing feature extraction on the at least one content segment; and determining, by a fourth machine learning model, the text sequence and the position information indicating respective positions of the plurality of content units in the first content based on the at least one third feature representation.

11. The method of claim 10, wherein the first content comprises a first image, the at least one content segment comprises at least one image region in the first image, and determining the position information comprises: for a first image region in the at least one image region, determining, by the fourth machine learning model, first positions of at least one content unit in the first image region based on a third feature representation corresponding to the first image region; and determining second positions of the at least one content unit in the first image based on the first positions of the at least one content unit to obtain position information indicating the second positions.

12. The method of claim 1, wherein the plurality of content units respectively correspond to a plurality of characters in the text sequence.

13. The method of claim 1, wherein determining the first entity in the first content comprises: determining, by the second machine learning model, a first recognition result indicating the first entity in the first content based on the text sequence, the position information, and the type information; determining at least one second recognition result based on the text sequence and at least one predetermined rule, each second recognition result indicating the first entity in the first content; and determining the first entity in the first content by merging the first recognition result and the at least one second recognition result.

14. An apparatus for content processing, comprising: a first detection module configured to detect a text sequence corresponding to a plurality of content units in a first content and position information indicating positions of the plurality of content units in the first content; a second detection module configured to detect, by using a first machine learning model, an entity type corresponding to an entity appearing in the first content, the entity being represented by at least one content unit, to obtain type information indicating the entity type; a determination module configured to determine, by using a second machine learning model, a first entity to be processed in the first content based on the text sequence, the position information, and the type information, the first entity being represented by at least one content unit; and a processing module configured to perform a predetermined processing on the first entity in the first content based on the position information to obtain a second content corresponding to the first content. The determination module is further configured to: perform feature encoding by using a multi-language encoder based on the text sequence, the position information, and the type information to generate a sequence of feature representations; and determine the first entity in the first content by using the second machine learning model based on the sequence of feature representations.

15. An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-13.

16. A computer-readable storage medium having computer-executable instructions stored thereon that are executable by a processor to implement the method according to any one of claims 1-13.

17. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1-13. ​ ​ ​

Citation Information

Patent Citations

  • Combinable weak authenticator-based named entity identification algorithm architecture

    CN112699682A

  • Conversation processing method and device, equipment, storage medium and program product

    CN118964547A