Content detection methods, systems, devices, and media
Patent Information
- Application Number
- JP2026514322
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-19
- Filing Date
- 2025-04-10
- Publication Date
- 2026-09-30
Smart Images

Figure 2026532609000001_ABST
Abstract
Description
[Technical Field]
[0001] (Reference to Related Application) This application claims the priority of Chinese Patent Application No. 202410978396.0 filed on July 19, 2024, and all the contents disclosed in the above Chinese patent application are incorporated herein by reference as part of this application.
[0002] (Technical Field) The present disclosure relates to a content detection method, a system, an electronic device, and a computer-readable storage medium. [Background Art]
[0003] With the rapid development of computer technology, digital assistants have come into being. A user can interact with a digital assistant through a human-machine interaction mode. Specifically, the user can input question information via the digital assistant, and the digital assistant can realize human-machine interaction by analyzing the question information and generating answer information corresponding to the question information.
[0004] Generally, considering that answer information is often long, in order to avoid a user waiting for a long time, the digital assistant may output content in a streaming output mode and transmit the output content to a client in the streaming output mode.
[0005] In the related art, in order to prevent inappropriate content (for example, content that is not desired to be presented to a user) from existing in the answer information, content detection can be performed on the answer information after the generation of the answer information is completed. However, in the above content detection process, it takes a long time from the completion of generating the answer information to the completion of content detection, and it is difficult to guarantee the real-time performance of content detection. [Summary of the Invention]
[0006] This disclosure provides a content detection method. This method can improve the real-time nature of content detection by performing content detection in a process in which a language model streams response information. This disclosure further provides systems, electronic devices, computer-readable storage media, and computer program products corresponding to the above method.
[0007] According to the first aspect, the Disclosure provides a content detection method, which is: Continuously receiving output content streamed from the language model, Based on the configured sampling policy, identify content slices from the output content already generated, and Performing content detection on a content slice and obtaining the detection results for the content slice, This includes transmitting the output content, which has undergone content detection, to a client, and displaying the output content on the client using a streaming output method.
[0008] In some possible implementations, content slices can be identified from output content based on a configured sampling policy. This includes identifying content slices from output content based on a configured sliding window and / or a configured stride, where the sliding window is used to determine the number of answer tokens in the content slice and the stride is used to determine the difference between content slices of two adjacent content detections.
[0009] In some possible implementations, content slices can be identified from output content based on a configured sliding window and / or configured stride. Based on the set stride, move the set sliding window backward in the token sequence consisting of multiple answer tokens in the output content that has already been output, This includes determining the answer tokens in a sliding window as content slices in response to the number of answer tokens in the sliding window reaching the length of the sliding window.
[0010] In some possible implementations, the method, In response to the language model initiating streaming output of the output content, the position of the sliding window in the token sequence is determined as the beginning of the token sequence, The method further includes determining the answer tokens in the sliding window as the content slice for the first content detection in response to the number of answer tokens in the sliding window reaching a first preset length that is less than the length of the sliding window.
[0011] In some possible implementations, after completing the first content detection, the method, To maintain the position in the token sequence of the sliding window, The method further includes determining the answer tokens in the sliding window as content slices for a second content detection in response to the number of answer tokens in the sliding window reaching the length of the sliding window.
[0012] In some possible implementations, content detection is performed on content slices to obtain the detection results for the content slices. This includes inputting content slices into a content detection model with natural language analysis capabilities and receiving detection results output from the content detection model.
[0013] In some possible implementations, the method, The process further includes determining a processing policy for a content slice based on the content slice detection results, wherein the processing policy includes at least one of the following: a policy for instructing the language model whether to continue streaming the output content; a policy for instructing whether to present the output content; and a policy for instructing the presence content.
[0014] In some possible implementations, the detection result indicates that the content slice passed content detection, and based on the content slice detection result, a processing policy for the content slice is determined. This instructs the language model to continue streaming the output content, This includes instructing the client to make the target response token present on the dialogue page, where the dialogue page is used for interacting with the digital assistant, and the target response token is a response token in a content slice that is not present on the dialogue page.
[0015] In some possible implementations, the detection result indicates that the content slice failed content detection, and the processing policy for the content slice is determined based on the content slice detection result. This involves instructing the language model to stop streaming output of the output content, This includes instructing the client to replace any response tokens already present on the dialogue page with default tokens.
[0016] According to a second aspect, the Disclosure provides a content detection system, which is A receiving module for continuously receiving output content streamed from a language model, An identification module configured to identify content slices from output content that has been output based on a set sampling policy, a detection module configured to perform content detection on the content slices to obtain detection results of the content slices, and a transmission module configured to transmit the output content that has undergone content detection to a client, so that the client displays the output content in a streaming output manner.
[0017] In some possible implementations, the identification module is specifically configured to: identify content slices from the output content that has been output based on a set sliding window and / or a set stride; wherein the sliding window is used to determine the number of answer tokens in the content slice, and the stride is used to determine the difference between content slices of two adjacent content detections.
[0018] In some possible implementations, the identification module is specifically configured to: move the set sliding window backwards in a token sequence composed of a plurality of answer tokens in the output content that has been output based on the set stride, and, in response to the number of answer tokens in the sliding window reaching the length of the sliding window, determine the answer tokens in the sliding window as the content slice.
[0019] In some possible implementations, the identification module is further configured to: in response to a language model starting streaming output of output content, determine the position of the sliding window in the token sequence as the front end of the token sequence, It is used in response to the number of answer tokens in the sliding window reaching a first preset length smaller than the length of the sliding window, for determining the answer tokens in the sliding window as a content slice for the first content detection.
[0020] In some possible implementation manners, after completing the first content detection, the identification module further: maintains the position in the token sequence of the sliding window, is used in response to the number of answer tokens in the sliding window reaching the length of the sliding window, for determining the answer tokens in the sliding window as a content slice for the second content detection.
[0021] In some possible implementation manners, specifically, the detection module is configured to: input a content slice into a content detection model with natural language analysis capability, and receive a detection result output from the content detection model.
[0022] In some possible implementation manners, the identification module is further configured to: determine a processing policy for the content slice based on the detection result of the content slice, wherein the processing policy includes at least one of a policy for instructing a language model whether to continue streaming output of output content, a policy for instructing whether to present output content, and a policy for instructing presented content.
[0023] In some possible implementation manners, the detection result indicates that the content slice passes content detection, and specifically, the identification module is configured to: instruct the language model to continue streaming output of output content, This is used to instruct the client to make the target response token present on the dialogue page, where the dialogue page is used for interacting with the digital assistant, and the target response token is a response token in the content slice that is not present on the dialogue page.
[0024] In some possible implementations, the detection result indicates that the content slice failed content detection, and a specific module, specifically, Instruct the language model to stop streaming output of the output content. This is used to instruct the client to replace any response tokens already present on the dialogue page with default tokens.
[0025] According to a third aspect, the Disclosure provides an electronic device, which includes a processor and memory. The processor and memory communicate with each other. The processor is used to execute instructions stored in memory to cause the electronic device to perform a content detection method in the first aspect or any one of the first aspect implementations.
[0026] According to a fourth aspect, the Disclosure provides a computer-readable storage medium that stores instructions for an electronic device to perform the content detection method described in the first aspect or any one of the embodiments of the first aspect.
[0027] According to a fifth aspect, the disclosure provides a computer program product that, when run on an electronic device, includes instructions causing the electronic device to execute the content detection method described in the first aspect or any one of the implementations of the first aspect.
[0028] This disclosure can also provide more implementations by combining the implementations described above. [Brief explanation of the drawing]
[0029] To more clearly explain the technical methods of the embodiments of this disclosure, the drawings that need to be used in the embodiments are briefly introduced below. [Figure 1] This is a flowchart of the content detection method according to the embodiments of this disclosure. [Figure 2] This is a flowchart of another content detection method according to an embodiment of the present disclosure. [Figure 3] This is a schematic diagram of the structure of a content detection system according to an embodiment of the present disclosure. [Figure 4] This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. [Modes for carrying out the invention]
[0030] The terms “First” and “Second” in the embodiments of this disclosure are for illustrative purposes only and should not be understood as indicating or suggesting relative importance or implicitly representing the number of technical features being referred to. Accordingly, features limited by “First” and “Second” may explicitly or implicitly include one or more features.
[0031] First, we will introduce some technical terms related to the embodiments of this disclosure.
[0032] With the rapid development of computer technology, digital assistants have emerged. Digital assistants generally possess natural language processing capabilities, and in some cases, users can interact with them in a human-machine dialogue manner. For example, a user can input a question related to general knowledge into the digital assistant, and the digital assistant can complete a round of dialogue by analyzing the content input by the user and generating answer information to the question related to general knowledge. In some other cases, the digital assistant may process natural language content according to the user's needs. For example, a user may input a paragraph of text into the digital assistant and instruct it to perform a related process (e.g., text summarization), and the digital assistant, recognizing the user's needs, may perform summarization processing on the text input by the user and generate answer information that includes text summary content.
[0033] Considering that response information may be lengthy, the digital assistant may output content using a streaming response method and transmit the output content to the client using a streaming response method. In other words, the digital assistant may output content in units of response tokens, outputting one response token each time, and then construct the response information from the response tokens after all response tokens have been output. In the process of streaming output, the response tokens are continuously transmitted to the client. For example, when a certain number of response tokens are reached, that number of response tokens are transmitted to the client. Alternatively, for example, at regular transmission intervals, the response tokens output within that transmission interval are transmitted to the client. In this way, the client can experience the output content output by the digital assistant in real time using a streaming response method, without having to wait for the digital assistant to generate complete response information before experiencing the complete response information. This avoids long waiting times for the user and improves the interaction experience between the user and the digital assistant.
[0034] Digital assistants generally generate response information based on the natural language processing capabilities of their language models. However, because the corpora (i.e., training data) used to train these language models are vast and complex, it is difficult to ensure that the response information generated by digital assistants perfectly matches user needs.
[0035] In related technologies, content detection is generally required for response information to avoid the presence of inappropriate content, such as content that the user does not wish to see. Specifically, after the digital assistant has finished generating the response information, the response information may be divided into multiple paragraphs, and content detection may be performed using pre-configured detection rules or keyword matching rules. However, the above content detection method is performed after the complete response information has been generated, and response information containing inappropriate content may still be seen by the user, making it difficult to guarantee real-time content detection.
[0036] In view of this, the present disclosure provides a content detection method. The method continuously receives output content streamed from a language model, determines content slices from the output content based on a configured sampling policy, then performs content detection on the content slices to obtain the content detection results for the content slices, transmits the content detected output content to a client, and allows the client to display the output content in a streaming output manner.
[0037] In this method, during the process in which the language model outputs content using a streaming output method, the output content is continuously received, and the output content is sliced according to the configured sampling policy. In this way, it is not necessary to wait until the generation of all response information is complete before performing detection, thus achieving "detection while outputting" and improving the real-time nature of content detection.
[0038] To facilitate understanding of the technical proposals based on the embodiments of this disclosure, the following explanation will be provided with reference to the drawings.
[0039] Referring to the flowchart of the content detection method according to an embodiment of the present disclosure shown in Figure 1, the method may be performed by a service for content detection, and the method specifically includes the following:
[0040] S101 continuously receives output content streamed from the language model.
[0041] In embodiments of this disclosure, a user can engage in human-machine interaction with a digital assistant. In some embodiments, the digital assistant may be a standalone software system, for example, having an application program (APP) for human-machine interaction functionality. In some other embodiments, the digital assistant may be a functional module of another system, for example, being deployed on a business platform and providing human-machine interaction functionality as a functional module of the business platform. In some other embodiments, the digital assistant may support online use.
[0042] A digital assistant typically provides a dialogue page where the user can interact with the digital assistant. Specifically, the dialogue page may provide input boxes and information submission controls, and the user may enter request information into the input boxes, for example, the user may enter question information or processing request information for text content into the input boxes. After the user has finished entering the request information, the request information may be sent to the digital assistant by triggering the information submission control, and the request information may be presented in the dialogue window.
[0043] Depending on the application, the dialogue window may have different representation formats. In some embodiments, the digital assistant is an independent software system, in which case the dialogue window may be a page loaded on the digital assistant's client. In some other embodiments, the digital assistant may be a functional module of another system, in which case the dialogue window may be a page loaded on the other system's client. In some other embodiments, the digital assistant supports online use, and the dialogue window may be a page displayed in a browser.
[0044] In some possible implementations, the digital assistant may generate response information based on the natural language processing capabilities of a language model. Here, the language model has natural language processing (NLP) capabilities and can be used to handle different types of natural language tasks. For example, the language model may be a deep learning model trained using text data.
[0045] In other words, when a user submits a request, the digital assistant may invoke a language model to process the request. The language model analyzes the request, recognizes the user's intent in the human-machine interaction, and then generates a response to the request. In this way, one round of human-machine interaction is completed.
[0046] The following describes the specific process by which the language model generates response information. The language model is generally a generative model, and the process of generating response information may include multiple output rounds, where one response token of the response information is output in each output round, and a new response token is generated based on the request information and the response token generated in the previous output round.
[0047] In other words, in the first output round, the language model outputs the first answer token based on the request information, and in the Nth output round, the language model generates the Nth answer token based on the request information and the N-1 answer tokens output in the previous N-1 output rounds, where N is an integer greater than 1. In this way, after the content output of all output rounds is complete, the answer tokens output in each output round constitute the answer information output by the language model.
[0048] In the embodiments of this disclosure, the language model outputs content using a streaming output method, and the service for content detection may continuously receive the output content streamed from the language model. That is, the digital assistant invokes the language model to output one response token per output round and sends the output content containing the response token to the service for content detection. In this way, the service for content detection continuously receives the output content from the language model in units of response tokens, and performs content detection on the output content before presenting the response information to the user.
[0049] S102 identifies content slices from the output content that has already been output, based on the configured sampling policy.
[0050] Specifically, a content slice includes at least one answer token. In embodiments of this disclosure, a content slice may be understood as a target for detection in content detection, that is, content detection may perform detection on at least one answer token in the content slice.
[0051] A sampling policy can be understood as a policy for identifying content slices from output content. Because the language model uses a streaming output method to generate one response token per output round, the service performing content discovery needs to identify content slices to be discovered from the output content (which contains multiple response tokens) generated by the language model, using a specific sampling policy.
[0052] In some possible implementations, the service for content detection may determine content slices by sliding window and stride, in combination with the streaming output features of the language model. Specifically, content slices are identified from output content based on a configured sliding window and / or stride.
[0053] Here, a sliding window may be used to process subsequences in a data sequence, holding a single window of fixed length, and the data within the window forms a subsequence as the window slides across the data sequence. The stride may be understood as the step size in which the sliding window slides across the data sequence.
[0054] In embodiments of this disclosure, the data sequence may refer to a token sequence consisting of multiple response tokens in output content output by the language model, i.e., the token sequence is composed of response tokens output by the language model in each output round based on the order of the output rounds.
[0055] Furthermore, the sliding window may be used to determine the number of response tokens in a content slice, and the stride may be used to determine the difference between content slices of two adjacent content detections. In other words, by sliding the window across the token sequence, the language model determines the content slices of multiple content detections in the process of generating response information, thereby achieving "detection while outputting".
[0056] In concrete implementation, the service for content detection may, based on the configured stride, move the configured sliding window in the token sequence backward, and in response to the number of response tokens in the sliding window reaching the length of the sliding window, determine the response tokens in the sliding window as content slices.
[0057] In other words, in the embodiments of this disclosure, the number of response tokens in a single content detection satisfies the length of the sliding window. For example, when the length of the sliding window is 500 tokens, the number of response tokens in a content slice is 500, and when content detection is performed thereafter, content detection is performed on 500 response tokens in the content slice. The difference between content slices of two adjacent content detections is the number of response tokens of the stride length. For example, when the stride length is 20 tokens, if the content slice of the Nth content detection contains tokens from the 301st to the 800th in the token sequence, then the content slice of the N+1th content detection contains tokens from the 321st to the 820th in the token sequence.
[0058] Because the language model continuously generates response tokens using a streaming output method, it may be understood that after moving the sliding window backward by a set stride, the sliding window may exceed the range of the token sequence, i.e., the number of response tokens in the sliding window may be less than the length of the sliding window. For example, if the length of the sliding window is 500 tokens and the stride length is 20 tokens, then after moving 20 tokens backward in the token sequence, the number of output response tokens contained in the sliding window is only 480, which is less than the length of the sliding window. As the service side performing content detection continuously receives output content streamed from the language model, the length of the token sequence increases accordingly, and when the number of response tokens in the sliding window reaches 500, the 500 response tokens in the sliding window are considered a content slice.
[0059] Thus, the embodiments of this disclosure, regarding the characteristics of the streaming output of the language model, perform slicing of the generated response tokens by moving a sliding window in a stride during the process in which the language model generates response information, and determine the targets for detection in multiple content detections. Since the output content output from the language model generally represents natural language content, the embodiments of this disclosure use the method of moving a sliding window in a stride to include some of the same response tokens in the content slices of two adjacent content detections. In this way, problems such as inaccurate token segmentation and lack of semantic consistency that would lead to a decrease in subsequent content detection accuracy are avoided.
[0060] Furthermore, in response to the language model starting streaming output of the output content, the position of the sliding window in the token sequence is determined as the beginning of the token sequence, and in response to the number of response tokens in the sliding window reaching a first preset length, the response tokens in the sliding window are determined as the content slice for the first content detection.
[0061] Here, the first preset length is smaller than the length of the sliding window. For example, the first preset length may be the length of the stride.
[0062] In other words, when the language model begins generating response information, the token sequence does not contain any previously outputted response tokens. The initial position of the sliding window is determined as the foremost point of the token sequence, and the system waits until the language model generates a first response token of a predetermined length before determining the first content slice. Thus, the first content detection is performed after a short wait.
[0063] To illustrate with an example where the first pre-defined length is the stride length, if the stride length is 20 tokens, the sliding window position is at the beginning of the token sequence, and it waits until the language model outputs the 20th answer token before determining the first to 20th answer tokens as the content slice for the first content detection.
[0064] Furthermore, after completing the first content detection, the position in the token sequence of the sliding window is retained, and in response to the number of answer tokens in the sliding window reaching the length of the sliding window, the answer tokens in the sliding window are determined to be the content slice for the second content detection.
[0065] In other words, the first content slice contains a relatively small number of first pre-set answer tokens. After completing the first content detection, the sliding window is not moved backward, and the system continues to wait for content output from the language model. When the number of output answer tokens reaches the length of the sliding window, the answer tokens up to the length of the sliding window are used as the second content slice.
[0066] For example, if the first pre-defined length is the stride length, and the stride length is 20 tokens, and the sliding window length is 500 tokens, then the content slice for the first content detection will include the first to 20th answer tokens, wait until the language model streams out the 500th answer token, and then determine the first to 500th answer tokens as the content slice for the second content detection.
[0067] Thus, in the early stages of the response information generation process, the sliding window is not moved. Instead, the language model is first allowed to output response tokens equal to the stride length, and then it is allowed to output response tokens equal to the length of the sliding window. This method is used to slice the response tokens. In subsequent content detection, the sliding window is moved further back, and the number of response tokens in the sliding window is allowed to reach the length of the sliding window.
[0068] As can be seen from the above, in the embodiments of this disclosure, the content slice determination process is related to the length of the sliding window and the stride length, so the user may set the length of the sliding window and the stride length themselves. That is, the user may set the corresponding length of the sliding window and the stride length based on actual business scenarios in order to meet the content detection needs in various business scenarios and to improve flexibility.
[0069] S103: Content detection is performed on the content slice to obtain the content slice detection results.
[0070] In the embodiments of this disclosure, content detection of a content slice may be used to detect whether or not the content in the content slice contains inappropriate content that the user does not wish to have a presence in. Specifically, when implemented, the service for performing content detection may input the content slice into a content detection model and receive the detection results output from the content detection model.
[0071] Here, the content detection model has natural language processing capabilities; for example, the content detection model may be a natural language processing model. By calling the content detection model and utilizing its natural language processing capabilities, it is determined whether or not the content slice has passed content detection.
[0072] Compared to keyword matching methods used in related technologies, content detection models can accurately detect a richer range of features, such as the semantics of response tokens in content slices, thereby improving the accuracy of content detection and reducing false reports and missed reports.
[0073] The embodiments of this disclosure do not limit the specific methods of content detection, and in some possible implementations, content detection for content slices may be achieved using methods such as configured content detection rules.
[0074] In the embodiments of this disclosure, the process by which the language model generates response information involves multiple content detections, such as determining a content slice, performing content detection on the content slice, re-determining the content slice, and re-detecting the content slice, thereby meeting the real-time content detection requirements in streaming responses.
[0075] S104 transmits the output content, for which content detection has been performed, to the client, and the client displays the output content using a streaming output method.
[0076] In the embodiments of this disclosure, the output content streamed from the language model is similarly transmitted to the digital assistant client in a streaming output manner. In this way, during the process in which the language model generates response information, the digital assistant client can present the output content already output to the user in a streaming output manner, achieving "present while outputting" and reducing the user's waiting time.
[0077] In addition, in the embodiments of this disclosure, the output content must first undergo content detection before being transmitted to the digital assistant client. That is, output content that has not undergone content detection is not transmitted to the client and is not presented to the user. In this way, the presentation of output content containing inappropriate content to the user is avoided.
[0078] Furthermore, in some embodiments, the service provider for content detection may determine a processing policy for the content slice based on the detection results of the content slice.
[0079] Here, the processing policy may include at least one of the following: a policy for instructing the language model whether to continue streaming the output content; a policy for instructing whether to present the output content; and a policy for instructing the presence content.
[0080] In other words, the service provider performing content detection may, after reviewing the content detection results, decide whether the language model should continue to respond to the user's request information, whether the currently output content should be presented to the user, or whether the already presented content should be processed.
[0081] As shown in Figure 2, when the detection result indicates that the content slice has passed content detection, the service performing the content detection may instruct the language model to continue streaming the output content and instruct the client to make the target answer token present on the dialogue page. Here, the target answer token is an answer token in the content slice that is not present on the dialogue page.
[0082] Since the language model outputs content in a streaming output manner, and the output content is transmitted to the digital assistant client in a streaming output manner, in the embodiments of this disclosure, two adjacent content detection content slices may contain the same answer tokens. For example, the content slice for the Mth content detection is from the 21st to the 520th answer token, and the content slice for the M+1th content detection is from the 41st to the 540th answer token, where the 41st to the 520th answer tokens are understood to be the same answer token. In this case, if the Mth detection result indicates that the content slice has passed content detection, the first to the 520th answer tokens may be present on the dialogue page, and if the M+1th detection result indicates that the content slice has passed content detection, then the target answer tokens will be from the 521st to the 540th answer token.
[0083] In this way, when the content slice does not contain inappropriate content, the language model is instructed to output the subsequent response token correctly, and the client is instructed to properly present the output content that has already been output to the user.
[0084] When the detection result indicates that a content slice failed content detection, it instructs the language model to stop streaming the output content and instructs the client to replace the response token already present on the interaction page with a default token.
[0085] In other words, when a content slice contains inappropriate content, the language model is instructed to stop generating subsequent answer tokens, and the current human-machine interaction is terminated. Then, to avoid the presence of inappropriate content to the user, the presence content on the interaction page is replaced in a timely manner. For example, the default token might say, "We cannot answer this question at this time. Could you please try a different question?"
[0086] In addition, the embodiments of this disclosure may achieve the objective of avoiding the presence of inappropriate content to users by other means. For example, keywords in inappropriate content may be replaced with default words, and inappropriate content may be represented by the symbol "*", and the embodiments of this disclosure are not limited to these methods.
[0087] In some possible implementations, the number of response tokens in a sliding window may not reach the length of the sliding window. For example, the length of the sliding window is 500 tokens, and the current length of response tokens in the sliding window is 490 tokens, but the generation of response information has already been completed. In the above situation, embodiments of this disclosure provide two different processing modes. In some embodiments, the service for performing content detection may tacitly accept that the response tokens in the sliding window have passed content detection, thereby achieving more lenient content detection. In some other embodiments, the service for performing content detection may not perform content detection on the response tokens in the sliding window, and in this way, the response tokens in the sliding window are not present to the user, avoiding the presence of inappropriate content to the user and achieving stricter content detection.
[0088] Based on the above description, the embodiments of this disclosure provide a content detection method. The method continuously receives output content streamed from a language model, determines content slices from the output content based on a set sampling policy, then performs content detection on the content slices to obtain the content detection results for the content slices, transmits the content detected output content to a client, and allows the client to display the output content in a streaming output manner.
[0089] In this method, during the process in which the language model outputs content using a streaming output method, the output content is continuously received, and the output content is sliced according to the configured sampling policy. In this way, it is not necessary to wait until the generation of all response information is complete before performing detection, thus achieving "detection while outputting" and improving the real-time nature of content detection.
[0090] The content detection method according to the embodiment of this disclosure has been described in detail above, linking Figures 1 and 2. Below, the system and equipment according to the embodiment of this disclosure will be described, linking the drawings.
[0091] Referring to the schematic diagram of the content detection system shown in Figure 3, the system 30 is: A receiving module 301 for continuously receiving output content containing multiple response tokens, which is streamed from a language model, Based on the configured sampling policy, a specific module 302 for identifying content slices from the output content that has already been output, A detection module 303 for performing content detection on the content slice and obtaining the detection result of the content slice, The system includes a transmission module 304 that transmits the output content, which has undergone content detection, to a client and causes the client to display the output content in a streaming output manner.
[0092] In some possible implementations, the specific module 302 is, in particular, Based on a configured sliding window and / or a configured stride, a content slice is identified from the output content that has already been output, where the sliding window is used to determine the number of answer tokens in the content slice, and the stride is used to determine the difference between content slices of two adjacent content detections.
[0093] In some possible implementations, the specific module 302 is, in particular, Based on the set stride, in the token sequence consisting of multiple answer tokens in the output content that has already been output, the set sliding window is moved backward. In response to the number of answer tokens in the sliding window reaching the length of the sliding window, the answer tokens in the sliding window are used to determine a content slice.
[0094] In some possible implementations, the specific module 302 further, In response to the language model starting the streaming output of the output content, the position of the sliding window in the token sequence is determined as the leading edge of the token sequence. In response to the number of response tokens in the sliding window reaching a first preset length less than the length of the sliding window, the response tokens in the sliding window are used to determine the content slice for the first content detection.
[0095] In some possible implementations, after completing the first content detection, the specific module 302 further: The position of the sliding window in the token sequence is maintained, In response to the number of response tokens in the sliding window reaching the length of the sliding window, the response tokens in the sliding window are used to determine the content slice for the second content detection.
[0096] In some possible implementations, the detection module 303 specifically, The aforementioned content slice is used to input into a content detection model having natural language analysis capabilities and to receive the detection results output from the content detection model.
[0097] In some possible implementations, the specific module 302 further, Based on the detection results of the content slice, a processing policy is used to determine a processing policy for the content slice, wherein the processing policy includes at least one of the following: a policy for instructing the language model whether to continue streaming the output content; a policy for instructing whether to present the output content; and a policy for instructing the presence content.
[0098] In some possible implementations, the detection result indicates that the content slice has passed content detection, and the specific module 302, specifically, The language model is instructed to continue streaming the aforementioned output content, This is used to instruct the client to make a target response token present on the dialogue page, where the dialogue page is used for interacting with a digital assistant, and the target response token is a response token in the content slice that is not present on the dialogue page.
[0099] In some possible implementations, the detection result indicates that the content slice failed content detection, and the specific module 302, specifically, The language model is instructed to stop streaming the output content. This is used to instruct the client to replace any response tokens already present on the dialogue page with default tokens.
[0100] The content detection system 30 of the embodiment of this disclosure is capable of performing the methods described in the embodiment of this disclosure, and the above and other operations and / or functions of each module / unit of the content detection system 30 are for realizing the corresponding flow of each method in the embodiment shown in Figure 1 or Figure 2, respectively, and for the sake of brevity, their explanation is omitted here.
[0101] Embodiments of this disclosure further provide electronic devices, specifically used to implement the functions of the content detection system 30 in the embodiment shown in Figure 3.
[0102] Figure 4 provides a schematic diagram of the structure of the electronic device 400, which, as shown in Figure 4, includes a bus 401, a processor 402, a communication interface 403, and a memory 404. The processor 402, the memory 404, and the communication interface 403 communicate with each other via the bus 401.
[0103] Bus 401 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, it is shown with only one thick line in Figure 4, but this does not mean that there is only one bus or only one bus type.
[0104] The processor 402 may be one or more of the following types of processors: a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0105] The communication interface 403 is used for communicating with the outside world. For example, the communication interface 403 may be used to communicate with a terminal.
[0106] Memory 404 may include volatile memory, such as random access memory (RAM). Memory 404 may further include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0107] Executable code is stored in memory 404, and the processor 402 executes the content detection method by executing the executable code.
[0108] Specifically, when implementing the embodiment shown in Figure 3, and when each module or unit of the content detection system 30 described in the embodiment of Figure 3 is implemented in software, the software or program code required to execute each module / unit function in Figure 3 may be partially or entirely stored in memory 404. The processor 402 executes the program code corresponding to each unit stored in memory 404 to execute the content detection method described above.
[0109] Embodiments of the present disclosure further provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored in a computing device, or a data storage device such as a data center that includes one or more available media. The available media may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk). The computer-readable storage medium includes instructions that instruct the computing device to perform the file processing method used in the content detection system 40.
[0110] Embodiments of this disclosure further provide computer program products comprising one or more computer instructions. When such computer instructions are loaded onto a computing device and executed, the flows or functions described in the embodiments of this disclosure occur, in whole or in part.
[0111] The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center by wired means (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless means (e.g., infrared, radio, microwave, etc.).
[0112] When the computer program product is executed by a computer, the computer executes one of the file processing methods described above. The computer program product may be a single software package, and if it is necessary to use one of the file processing methods described above, the computer program product may be downloaded and executed on the computer.
[0113] The flowcharts and block diagrams in the drawings illustrate the systematic architecture, functions, and operations that can be realized according to the systems, methods, and computer program products of each embodiment of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for realizing a defined logic function. Note that in some implementations as alternatives, the functions described in the blocks may occur in a different order than those shown in the drawings. For example, two blocks shown consecutively may actually be executed almost in parallel, or in some cases in reverse order, depending on the function. Note that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be realized by a dedicated hardware-based system that performs the defined function or operation, or by a combination of dedicated hardware and computer instructions.
[0114] The units referred to in the descriptions of the embodiments of this disclosure may be implemented in software or in hardware. Herein, the names of the units / modules are not limited in any case to the units themselves.
[0115] The functions described above in this specification may be performed, at least in part, by one or more hardware logic components. For example, typical types of hardware logic components that can be used include, but are not limited to, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on chips (SOCs), and composite programmable logic circuits (CPLDs).
[0116] In the context of the embodiments of this disclosure, a machine-readable medium may be a tangible medium that contains or stores a program for use by or in combination with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of machine-readable storage media include one or more wire-based electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0117] Each example in this specification is described sequentially, and the emphasis in each example is on the differences from other examples. Similar or identical parts between examples can be referenced to one another. For systems or apparatus disclosed in the examples, the description is simplified as it corresponds to the method disclosed in the examples; relevant sections should be referred to in the method section.
[0118] In this disclosure, “at least one” should be understood to mean one or more, and “multiple” should be understood to mean two or more. “And / or” is used to describe the relationship between related objects and indicates that there can be three relationships. For example, “A and / or B” can represent three cases: A only exists, B only exists, and A and B exist simultaneously, where A and B may be singular or plural. The letter “ / ” generally indicates that the related objects before and after are in an “or” relationship. “At least one of the following” or similar expressions refers to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c may represent a, b, c, “a and b,” “a and c,” “b and c,” or “a, b, and c,” where a, b, and c may be singular or plural.
[0119] In this specification, relational terms such as "First" and "Second" are merely used to distinguish one entity or operation from another, and do not necessarily require or suggest that any such actual relationship or order exists between these entities or operations. The terms "include," "incorporate," or any other variation thereof are intended to cover non-exclusive inclusion, thereby meaning that a process, method, article, or apparatus containing a set of elements includes not only those elements but also other elements not explicitly listed, or further elements specific to that process, method, article, or apparatus. Unless further limited, an element limited by the phrase "includes one..." does not preclude the existence of other identical elements in a process, method, article, or apparatus containing such element.
[0120] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be carried out directly by hardware, by a software module executed by a processor, or by a combination of both. The software module may reside in random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, portable magnetic disks, CD-ROMs, or any other form of storage medium well known in the art.
[0121] Based on the above description of the disclosed embodiments, those skilled in the art can implement or use the present disclosure. Various modifications to these embodiments are obvious to those skilled in the art, and the general principles defined herein can also be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to these embodiments shown herein, but should be given the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A content detection method, Continuously receiving output content streamed from the language model, Based on the configured sampling policy, identify content slices from the output content that has already been output, Performing content detection on the content slice and obtaining the detection result for the content slice, A content detection method comprising transmitting the output content that has undergone content detection to a client, and causing the client to display the output content in a streaming output manner.
2. Based on the configured sampling policy described above, identifying content slices from the output content already output is: The method according to claim 1, comprising identifying a content slice from the output content that has been output, based on a configured sliding window and / or a configured stride, wherein the sliding window is used to determine the number of answer tokens in the content slice, and the stride is used to determine the difference between content slices of two adjacent content detections.
3. Identifying content slices from the output content based on the configured sliding window and / or configured stride is: Based on the set stride, in the token sequence consisting of multiple answer tokens in the output content that has already been output, the set sliding window is moved backward, The method according to claim 2, comprising determining the answer tokens in the sliding window as a content slice in response to the number of answer tokens in the sliding window reaching the length of the sliding window.
4. In response to the language model starting the streaming output of the output content, the position of the sliding window in the token sequence is determined as the leading edge of the token sequence. The method according to claim 3, further comprising determining the answer tokens in the sliding window as a content slice for the first content detection in response to the number of answer tokens in the sliding window reaching a first preset length that is less than the length of the sliding window.
5. After completing the first content detection, the method Maintaining the position of the sliding window in the token sequence, The method according to claim 4, further comprising determining the answer tokens in the sliding window as a content slice for a second content detection in response to the number of answer tokens in the sliding window reaching the length of the sliding window.
6. Performing content detection on the aforementioned content slice and obtaining the detection result for the content slice is, The method according to any one of claims 1 to 5, comprising inputting the content slice into a content detection model having natural language analysis capabilities and receiving detection results output from the content detection model.
7. The method according to any one of claims 1 to 6, further comprising determining a processing policy for the content slice based on the detection result of the content slice, wherein the processing policy includes at least one of a policy for instructing the language model whether to continue streaming the output content, a policy for instructing whether to present the output content, and a policy for instructing the presence content.
8. The detection result indicates that the content slice has passed content detection, and determining a processing policy for the content slice based on the detection result of the content slice is: Instructing the language model to continue streaming the aforementioned output content, The method according to claim 7, comprising instructing the client to make a target response token present on a dialogue page, wherein the dialogue page is used to interact with a digital assistant, and the target response token is a response token in the content slice that is not present on the dialogue page.
9. The aforementioned detection result indicates that the content slice failed content detection, and determining a processing policy for the content slice based on the aforementioned detection result of the content slice is: Instructing the language model to stop streaming the output content, The method according to claim 7, further comprising instructing the client to replace any response tokens already present on the dialogue page with a default token.
10. A content detection system, A receiving module configured to continuously receive output content streamed from a language model, A specific module configured to identify content slices from the output content already output, based on a configured sampling policy, A detection module configured to perform content detection on the content slice and obtain the detection result of the content slice, A content detection system including a transmission module configured to transmit the output content, which has undergone content detection, to a client, and to display the output content on the client in a streaming output manner.
11. An electronic device including a processor and memory, The processor is used to execute instructions stored in the memory to cause the electronic device to perform the content detection method described in any one of claims 1 to 9.
12. A computer-readable storage medium comprising an instruction for instructing an electronic device to perform the content detection method described in any one of claims 1 to 9.