An automatic generation method, device and equipment of interactive content and a storage medium
By performing risk level analysis and multimodal compliance filtering on the initial input content, target input content and response content are generated, which solves the problem of high violation risk in the automatic reply system and improves network security and compliance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI G&Z EDUCATIONAL TECH CO LTD
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-21
AI Technical Summary
Existing automated response systems lack risk detection and intervention mechanisms for input content in open social environments, resulting in a high risk of outputting inappropriate content and impacting cybersecurity.
The risk level is determined by semantic analysis of the initial input content, and the content is rewritten according to the level to generate the target input content. At the same time, the initial response content is subjected to multimodal compliance filtering to generate the target response content, so as to avoid displaying illegal content in the interactive interface.
It significantly reduces the risk of violations during automated response interactions, enhances network security, and ensures the compliance of interactive content and user experience.
Smart Images

Figure CN122433701A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device and storage medium for automatically generating interactive content. Background Technology
[0002] With the integration and development of internet social networking and artificial intelligence technologies, automatic response systems (such as intelligent chatbots and automatic reply tools on social platforms) have been widely used in instant messaging, online customer service and community interaction scenarios. They use natural language processing technology to achieve real-time response to user input, significantly improving the efficiency of social communication.
[0003] Existing automated response systems primarily focus on a one-way flow from user input to model generation and then to output. They generally use the user's question as the sole input, generating a matching response through a pre-trained model, lacking risk detection and intervention mechanisms for the input content itself. In open social environments, user input is diverse, random, and even malicious, potentially containing prohibited information or triggering inappropriate content through subtle, suggestive language. This makes it highly susceptible to outputting prohibited content during the automated response process, posing a high risk of violations and impacting cybersecurity. Summary of the Invention This application provides a method, apparatus, device, and storage medium for automatically generating interactive content, which can solve the technical problem of high violation risk in the automatic response interaction process. It performs risk control on input content based on target input content obtained by rewriting the initial input content, avoiding the display of illegal input content in the interactive interface. Furthermore, it obtains the corresponding target response content by performing multimodal compliance filtering on the initial response content, avoiding the display of illegal response content in the interactive interface. By performing security risk control from both content input and content output dimensions, it significantly reduces the violation risk in the automatic response interaction process, thereby significantly improving network security in automatic response interaction scenarios.
[0004] In a first aspect, embodiments of this application provide a method for automatically generating interactive content, comprising: In response to the initial input content from the interactive interface, semantic analysis is performed on the initial input content to determine the corresponding risk level; The initial input content is modified according to the risk level to obtain the target input content, and the initial input content is replaced with the target input content in the interactive interface; The initial response content is generated based on the initial input and the set response template. The initial response content is then subjected to multimodal compliance filtering to obtain the target response content, which is then displayed in the interactive interface.
[0005] Furthermore, the initial response content undergoes multimodal compliance filtering to obtain the target response content, including: The initial response content is processed using modality recognition to obtain the response type; Based on the response type, the corresponding compliance detection model is invoked to identify and process violations in the initial response content, and the identified violations are then processed to obtain the target response content.
[0006] Furthermore, based on the response type, the corresponding compliance detection model is invoked to identify violations in the initial response content, and the identified violations are then processed to obtain the target response content, including: When the response type is text, the initial response content is detected for illegal text fragments by combining the text content security detection model with the red line knowledge base and sensitive word database. The detected illegal text fragments are then deleted or replaced with neutral expressions to obtain the target response content. When the response type is image, a multi-label image classification model is used to identify the non-compliant visual elements, and the identified non-compliant visual elements are replaced with compliance warning images or pixel-level blurring is applied to obtain the target response content. When the response type is audio, the voiceprint feature recognition model identifies the illegal audio segments and either mutes them or replaces them with standard prompts to obtain the target response content.
[0007] Furthermore, based on the response type, the corresponding compliance detection model is invoked to identify violations in the initial response content, and the identified violations are then processed to obtain the target response content, including: When the response type is video, extract the image and audio data frame by frame. The non-compliant visual elements in each image are identified by a multi-label image classification model, and the identified non-compliant visual elements are replaced with compliance warning images or pixel-level blurring is applied to the non-compliant visual elements to obtain the target image content. The audio detection model is used to detect and process audio data to identify illegal audio segments, and these segments are either muted or replaced with standard prompts to obtain the target audio content. Integrate the target image content and the target audio content to obtain the complete target response content.
[0008] Furthermore, semantic analysis is performed on the initial input content to determine the corresponding risk level, including: The initial input content is classified into topics to determine the first probability distribution of the initial input content belonging to the set topic categories; Perform semantic intent classification on the initial input content to determine the second probability distribution of the initial input content belonging to the set intent category; The risk level of the initial input content is determined based on the first probability distribution and the second probability distribution.
[0009] Furthermore, the initial input content is modified according to the risk level to obtain the target input content, including: When the risk level is medium risk, the security rewrite engine is invoked to perform semantic preservation modification on the initial input content to generate the target input content. When the risk level is high, the initial input is discarded, and a standard response template relevant to the current dialogue context is retrieved from the pre-built redline knowledge base as the target input.
[0010] Furthermore, the secure rewrite engine is invoked to perform semantically preservative modifications on the initial input content to generate the target input content, including: The security rewriting engine is invoked to identify sensitive words in the initial input content and replace them with semantically similar neutral words to generate the target input content; And / or, invoke the secure rewrite engine to extract the core intent of the initial input content, and reconstruct the content using compliant expression based on the core intent to obtain the target input content; And / or, call the secure rewriting engine to delete sensitive words that are not core semantics to obtain the target input content; And / or, call the security rewriting engine to mask sensitive words in the core semantics to obtain the target input content.
[0011] In a second aspect, embodiments of this application provide an automatic interactive content generation apparatus, comprising: The response module is used to respond to the initial input content of the interactive interface and perform semantic analysis on the initial input content to determine the corresponding risk level; The content determination module is used to modify the initial input content according to the risk level to obtain the target input content. The display module is used to replace the initial input content with the target input content in the interactive interface; The content determination module is also used to generate initial response content based on the initial input content and the set response template, and to perform multimodal compliance filtering on the initial response content to obtain the target response content; The display module is also used to display the target reply content in the interactive interface.
[0012] In a third aspect, embodiments of this application provide an automatic interactive content generation device, comprising: Memory and one or more processors; Memory, used to store one or more programs; When one or more programs are executed by one or more processors, the one or more processors implement the automatic generation method of interactive content as described in the first aspect.
[0013] In a fourth aspect, embodiments of this application provide a storage medium for storing computer-executable instructions, which, when executed by a computer processor, are used to perform an automatic generation method for interactive content as described in the first aspect.
[0014] This embodiment of the application, when automatically generating interactive content, responds to the initial input content entered into the interactive interface, performs semantic analysis on the initial input content to determine the corresponding risk level, modifies the initial input content according to the risk level to obtain target input content, replaces the initial input content with the target input content in the interactive interface, generates initial response content based on the initial input content and the set response template, performs multimodal compliance filtering on the initial response content to obtain target response content, and displays the target response content in the interactive interface. By employing the above technical means, the target input content can be obtained by rewriting the initial input content, and the target response content can be obtained by performing multimodal compliance filtering on the generated initial response content. This avoids the technical problem of high violation risk in the automatic response interaction process. This embodiment replaces the initial input content with the target input content in the interactive interface and displays the target response content in the interactive interface, thereby avoiding the display of illegal input and response content in the interactive interface. Security risk control is performed from both the content input and content output dimensions, significantly reducing the violation risk in the automatic response interaction process, and thus significantly improving the network security in the automatic response interaction scenario.
[0015] The beneficial effects of the aforementioned automatic interactive content generation device, automatic interactive content generation equipment, and storage medium can be referenced in relation to the beneficial effects of the automatic interactive content generation method. Attached Figure Description
[0016] Figure 1 This is a flowchart of an automatic interactive content generation method provided in an embodiment of this application; Figure 2 This is a flowchart of another method for automatically generating interactive content provided in an embodiment of this application; Figure 3 This is a flowchart of another method for automatically generating interactive content provided in the embodiments of this application; Figure 4 This is a flowchart of another method for automatically generating interactive content provided in an embodiment of this application; Figure 5This is a flowchart of another method for automatically generating interactive content provided in the embodiments of this application; Figure 6 This is a flowchart of another method for automatically generating interactive content provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an automatic interactive content generation device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an automatic interactive content generation device provided in an embodiment of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0018] Existing automated response systems primarily focus on a one-way flow from user input to model generation and then to output. They generally use the user's question as the sole input, generating a matching response through a pre-trained model, lacking risk detection and intervention mechanisms for the input content itself. In open social environments, user input is diverse, random, and even malicious, potentially containing prohibited information or triggering inappropriate content through subtle, suggestive language. This makes it highly susceptible to outputting prohibited content during the automated response process, posing a high risk of violations and impacting cybersecurity.
[0019] Based on this, this application provides an automatic generation method, apparatus, device, and storage medium for interactive content, aiming to automatically generate interactive content by responding to initial input content input through the interactive interface, performing semantic analysis on the initial input content to determine the corresponding risk level, modifying the initial input content according to the risk level to obtain target input content, replacing the initial input content with the target input content in the interactive interface, generating initial response content based on the initial input content and a set response template, performing multimodal compliance filtering on the initial response content to obtain target response content, and displaying the target response content in the interactive interface. Using the above technical means, the target input content can be obtained by rewriting the initial input content, and the target response content can be obtained by performing multimodal compliance filtering on the generated initial response content. Compared with existing open automatic response methods, this embodiment replaces the initial input content with the target input content and displays the target response content in the interactive interface, thereby avoiding the display of illegal input and response content in the interactive interface. Security risk control is performed from both the content input and content output dimensions, significantly reducing the risk of violations in the automatic response interaction process, and thus significantly improving network security in automatic response interaction scenarios.
[0020] Figure 1 A flowchart of an automatic interactive content generation method according to an embodiment of this application is provided. The automatic interactive content generation method provided in this embodiment can be executed by an automatic interactive content generation device. This device can be implemented through software and / or hardware. The automatic interactive content generation device can consist of two or more physical entities, or it can consist of a single physical entity. Generally, the automatic interactive content generation device can be a computer device.
[0021] The following description uses a computer device as the primary example to illustrate an automatic method for generating interactive content. (Refer to...) Figure 1 The method for automatically generating this interactive content specifically includes: S11. In response to the initial input content from the interactive interface, perform semantic analysis on the initial input content to determine the corresponding risk level.
[0022] In automated response scenarios, users can input initial content through the corresponding interface. The computer device responds to this initial input and performs semantic analysis to determine the corresponding risk level. Subsequently, the computer device automatically generates a response based on the initial input. For example, when a user submits initial input (text, speech-to-text, or multimodal content) through an interface (such as a chat window or input box), the system captures the initial input in real time through a front-end monitoring mechanism and simultaneously records metadata such as input time, user ID, and device identifier. If the initial input is text, it undergoes a series of preprocessing steps, including removing special characters (such as garbled characters or consecutive punctuation), standardization (such as Chinese-English punctuation conversion or case unification), and word segmentation. If the initial input is speech, it is converted to text and then further preprocessed, including homophone correction (such as identification and restoration of "homophonic puns"). The preprocessed initial input is then transformed into a semantic vector, and semantic features are extracted from the semantic vector. Based on semantic features, the corresponding topic category and intent category are determined. Then, based on the obtained topic category and intent category, the risk level of the initial input content is determined, including low-risk, medium-risk, and high-risk levels. A low-risk level indicates no intervention is needed. For initial input content at the medium-risk and high-risk levels, appropriate intervention is required to reduce the risk of violations.
[0023] The above-mentioned approach uses semantic analysis of the initial input content to determine the corresponding risk level, providing a precise basis for subsequent intervention strategies and thereby improving the accuracy of risk identification of interactive content.
[0024] S12. Modify the initial input content according to the risk level to obtain the target input content, and replace the initial input content with the target input content in the interactive interface.
[0025] After determining the risk level of the initial input content in the foregoing S11, the initial input content may be processed for content modification according to the risk level to obtain the target input content, so as to modify the possible违规 information in the initial input content. Subsequently, the initial input content is replaced with the target input content in the interaction interface, so as to avoid displaying the违规 initial input content in the interaction interface. Exemplarily, when it is determined that the risk level of the initial input content is a low risk level, no modification may be made, and the initial input content is used as the target input content. Or, when it is determined that the risk level of the initial input content is a low risk level, mild optimization processing is performed on the initial input content. For example, local content is modified, such as modifying "rookie" to "novice", which improves the clarity of the plaintext while retaining the original meaning; or only format standardization rewriting is performed (such as correcting typos or adjusting punctuation marks, etc.), without changing the core semantics, to obtain the target input content. When it is determined that the risk level of the initial input content is a medium risk level or a high risk level, different degrees of content modification processing such as semantic preservation modification processing or content rewriting based on a standard reply template need to be performed on the initial input content to effectively remove the违规 information in the initial input content and obtain a compliant target input content.
[0026] As described above, corresponding degrees of content rewriting processing are performed on the initial input content according to the risk level, avoiding the one-size-fits-all extensive mode, achieving hierarchical and fine-grained gradient management, improving the orderliness of content modification, and enhancing the accuracy of risk control of the input content. In addition, by processing the initial input content for content modification to obtain the target input content, it is possible to remove the corresponding违规 information while retaining the core intention of the user, which helps to maintain the smoothness of the interaction and user satisfaction.
[0027] S13. Generate an initial reply content according to the initial input content and the set reply template, perform multimodal compliance filtering processing on the initial reply content to obtain the target reply content, and display the target reply content in the interaction interface.
[0028] It should be noted that the Chinese text contains some unclear terms like "违规信息" which are left in the translation as "违规 information" for lack of specific context. You may need to adjust it according to the actual meaning.After replacing the initial input content with the target input content in the interactive interface in S12, initial response content is generated based on the initial input content and the set response template. Multimodal compliance filtering is then applied to the initial response content to obtain the target response content, which is then displayed in the interactive interface. For example, when generating the initial response content, the initial input content can be input into a set automatic response model for content generation, and the model outputs the corresponding initial response content. It should be noted that the generated initial response content can be multimodal, such as text, image, audio, or video content. Since the generated initial response content may also have certain compliance risks, multimodal compliance filtering can be applied. For example, violations in the initial response content (text, image, audio, and / or video content) can be identified and filtered for compliance to obtain the corresponding compliant target response content. The compliant target response content is then displayed in the interactive interface.
[0029] The above describes how the target response content is obtained by performing multimodal compliance filtering on the initial response content, enabling multi-dimensional detection of text, images, audio, or video. This overcomes the limitations of single-modal text filtering, expands the dimensions of violation risk detection, significantly reduces the risk of outputting non-compliant response content, and thus significantly improves the cybersecurity of automatically generated interactive content.
[0030] The above-described method aims to automatically generate interactive content by responding to the initial input content input into the interactive interface, performing semantic analysis on the initial input content to determine the corresponding risk level, modifying the initial input content according to the risk level to obtain the target input content, and replacing the initial input content with the target input content in the interactive interface. Based on the initial input content and the set response template, an initial response content is generated, and multimodal compliance filtering is applied to the initial response content to obtain the target response content, which is then displayed in the interactive interface. Using the above technical means, the target input content can be obtained by rewriting the initial input content, and the target response content can be obtained by applying multimodal compliance filtering to the generated initial response content. Compared to existing open-ended automatic response methods, this embodiment replaces the initial input content with the target input content and displays the target response content in the interactive interface, thereby avoiding the display of non-compliant input and response content in the interactive interface. Security risk control is implemented from both the content input and content output dimensions, significantly reducing the risk of violations in the automatic response interaction process, and thus significantly improving the network security in automatic response interaction scenarios.
[0031] Figure 2This is a flowchart of another method for automatically generating interactive content provided in an embodiment of this application, see below. Figure 2 The method for automatically generating this interactive content specifically includes: S111. Perform topic classification processing on the initial input content to determine the first probability distribution of the initial input content belonging to the set topic category.
[0032] When performing semantic analysis on the initial input content to determine the corresponding risk level in the aforementioned S11, the initial input content can first be classified into topics to determine the first probability distribution of the initial input content belonging to the set topic category. For example, when a user submits initial input content (text, speech-to-text, or multimodal content) in an interactive interface (such as a chat window or input box), the system captures the initial input content in real time through a front-end monitoring mechanism and simultaneously records metadata such as input time, user ID, and device identifier. If the initial input content is text, it undergoes a series of preprocessing steps, including removing special symbols (such as garbled characters or consecutive punctuation), standardization (such as Chinese-English punctuation conversion or case unification), and word segmentation. If the initial input content is speech, after converting speech to text, additional preprocessing such as homophone correction (such as identification and restoration of "homophonic puns") is performed. The preprocessed initial input content is then subjected to semantic vector transformation to obtain a semantic vector, and semantic features are extracted based on the semantic vector. The extracted semantic features are input into the topic classification model. The model then classifies and identifies topic categories, outputting a probability value (0-1) for each category. This yields the first probability distribution of the initial input content belonging to the set topic categories. For example, if the set topic categories are A1~AN, the corresponding first probability distributions are P11~P1N.
[0033] S112. Perform semantic intent classification processing on the initial input content to determine the second probability distribution of the initial input content belonging to the set intent category.
[0034] While determining the first probability distribution in S111, the initial input content can be semantically classified to determine the second probability distribution of the initial input content belonging to the set intent category. For example, when a user submits initial input content (text, speech-to-text, or multimodal content) in an interactive interface (such as a chat window or input box), the system captures the initial input content in real time through a front-end monitoring mechanism and simultaneously records metadata such as input time, user ID, and device identifier. If the initial input content is text, it undergoes a series of preprocessing steps, including removing special symbols (such as garbled characters or consecutive punctuation), standardization (such as Chinese-English punctuation conversion or case unification), and word segmentation. If the initial input content is speech, after converting speech to text, additional preprocessing such as homophone correction (such as identification and restoration of "homophonic puns") is performed. The preprocessed text content is then subjected to semantic vector transformation to obtain a semantic vector, and semantic features are extracted based on the semantic vector. The extracted semantic features are fed into the set intent category model. The intent category model performs intent classification processing and outputs a probability value (0-1) for each intent category, thus obtaining the second probability distribution of the initial input content belonging to the set intent category. For example, if the set intent categories are B1~BN, the corresponding second probability distributions are P21~P2N.
[0035] S113. Determine the risk level of the initial input content based on the first probability distribution and the second probability distribution.
[0036] After determining the first probability distribution in S111 and the second probability distribution in S112, the risk level of the initial input content can be determined based on the first and second probability distributions. For example, the topic category with the highest probability value is selected from the first probability distribution as a candidate topic category, and a first probability value for the candidate topic category is determined. The intent category with the highest probability value is selected from the second probability distribution as a candidate intent category, and a second probability value for the candidate intent category is determined. The first and second probability values are then weighted and fused according to the preset weight coefficients of the candidate topic category and the candidate intent category to obtain a comprehensive evaluation value. For example, S = (W1 × P1_topic + W2 × P2_topic) × Q1 × Q2, where S is the comprehensive evaluation value, W1 is the weight coefficient of the candidate topic category, P1_topic is the first probability value of the candidate topic category, W2 is the weight coefficient of the candidate intent category, P2_topic is the second probability value of the candidate intent category, Q1 is the user profile factor (e.g., 0.9 for high-credit users, 1.0 for medium-credit users, and 1.1 for low-credit users), and Q2 is the scenario factor (e.g., 1.2 for children's scenarios and 1.0 for ordinary social scenarios). Assuming W1 is 0.6, P1_topic is 0.85, W2 is 0.4, and P2_topic is 0.7, and the user is a medium-credit user and the scenario is an ordinary social scenario, then the comprehensive evaluation value S = (0.6 × 0.85 + 0.4 × 0.7) × 1.0 × 1.0 = 0.79. When the comprehensive assessment value is less than the first threshold, the risk level of the initial input content is determined to be low risk; when the comprehensive assessment value is greater than or equal to the first threshold and less than the second threshold, the risk level of the initial input content is determined to be medium risk; when the comprehensive assessment value is greater than or equal to the second threshold, the risk level of the initial input content is determined to be high risk. For example, assuming the first threshold is 0.5 and the second threshold is 0.7, when the comprehensive assessment value is less than 0.5, the risk level of the initial input content is determined to be low risk; when the comprehensive assessment value is greater than or equal to 0.5 and less than 0.7, the risk level of the initial input content is determined to be medium risk; and when the comprehensive assessment value is greater than or equal to 0.7, the risk level of the initial input content is determined to be high risk.
[0037] As described above, by using a two-dimensional classification of topic category and intent category, the system overcomes the limitations of traditional single-keyword matching, effectively and covertly identifying illegal content, improving the accuracy of illegal content identification, and reducing the probability of false positives. Furthermore, the system not only outputs the risk level but also the first probability distribution of the topic category and the second probability distribution of the intent category, providing interpretable quantitative evidence, reducing subjective bias, improving the objectivity of decision-making, and facilitating subsequent manual review and model optimization through clear first and second probability distributions.
[0038] Figure 3 This is a flowchart of another method for automatically generating interactive content provided in the embodiments of this application, referred to... Figure 3 The method for automatically generating this interactive content specifically includes: S121. When the risk level is medium risk, the security rewrite engine is invoked to perform semantic preservation modification on the initial input content to generate the target input content.
[0039] After determining the risk level of the initial input content in S11, if the risk level of the initial input content is medium risk, the security rewriting engine can be invoked to perform semantically preservative modification processing on the initial input content to generate the target input content. For example, when the medium risk level determination takes effect (0.5 ≤ comprehensive evaluation value < 0.7), the system automatically activates the security rewriting engine, loading a modification strategy library that matches the current topic category and intent category. For example, "vulgar topic + mocking intent" corresponds to a "mild rewriting strategy," and "biased topic + ambiguous probing intent" corresponds to a "clear guidance strategy." The initial input content is scanned using an automata algorithm, matching a medium-risk sensitive word library (such as vulgar metaphors or mildly discriminatory expressions), and accurately marking the location of violating segments (such as "'XX' in this expression is a vulgar metaphor"). The semantic dependency analysis model is invoked to identify the core semantic structure of the input content (subject-verb-object relationship and modifiers) and the user's core intent (such as the core intent of "asking about a certain type of vulgar information" being "to obtain entertainment content"). When performing semantic preservation modifications, a secure rewriting engine can be invoked to identify sensitive words in the initial input content and replace them with semantically similar neutral words to generate the target input content. For example, for isolated non-compliant words, "semantic synonymous compliant word replacement" can be used (e.g., replacing "with color" with "lighthearted and humorous"), and the language model can be used to verify the fluency of the sentence to obtain the target input content. Alternatively, when performing semantic preservation modifications, a secure rewriting engine can be invoked to extract the core intent of the initial input content and reconstruct the content using compliant expressions based on the core intent to obtain the target input content. For example, for content containing non-compliant expressions but with a legitimate intent, the core requirement is retained and the expression is rewritten (e.g., reconstructing "how to find that kind of movie" into "you can search for compliant movie recommendations, I will provide you with assistance"). Alternatively, when performing semantic preservation modifications, a secure rewriting engine can be invoked to delete non-core semantic sensitive words to obtain the target input content. For example, non-core semantic modifiers (such as "XX" in "This rule is really XX") can be directly deleted, retaining the main viewpoint (such as "This rule is controversial"). And / or, during semantic preservation modifications, a secure rewriting engine can be invoked to mask sensitive words in the core semantics to obtain the target input content. And / or, during semantic preservation modifications, a secure rewriting engine can be invoked to add contextualized guidance at the end of the rewritten content (such as "Please abide by community guidelines when communicating, and work together to maintain a good environment") to guide users to express themselves correctly.
[0040] After the initial input content is semantically modified using the secure rewriting engine to generate the target input content, a risk level assessment can be performed again on the generated target input content to ensure that the modified risk level is low. If the modified target input content is of medium or high risk, a more stringent rewriting strategy is automatically switched until the final output target input content has a low risk level.
[0041] S122. When the risk level is high, discard the initial input content and call the standard response template related to the current dialogue context from the pre-built redline knowledge base as the target input content.
[0042] After determining the risk level of the initial input content in S11, if the risk level of the initial input content is high, the initial input content can be discarded, and a standard response template related to the current dialogue context can be retrieved from the pre-built redline knowledge base as the target input content. For example, when the high-risk level determination takes effect (comprehensive evaluation value ≥ 0.7), the dialogue history understanding model is invoked to extract the topic, sentiment (e.g., "malicious" or "questioning"), and user history interaction features (e.g., whether high-risk content has been sent multiple times) of the current dialogue for contextual semantic feature extraction. Based on the mentioned contextual semantic features, multi-level matching is performed in the pre-built redline knowledge base. First, the topic category is matched (e.g., "related to XX" corresponds to "policy and regulation template"), then the intent type is matched (e.g., "malicious questioning" corresponds to "positive guidance template"), and finally, the standard response template with the highest matching degree (≥ 0.9) is selected (e.g., "Regarding this issue, please refer to the authoritative information released by the official sources"). If there is no template with a matching degree ≥ 0.9, a cross-scenario general redline template is automatically invoked (e.g., "Your input contains inappropriate content, and we cannot provide a relevant response; please change the topic") to ensure the compliance of the response. Make lightweight adjustments to the selected standard template, such as inserting contextual entities of the current conversation (such as time, location, or non-sensitive information), to make the response more relevant to the scenario (e.g., change "Please refer to official information" to "Please refer to the latest information released by XX department").
[0043] In one embodiment, regardless of whether it is a semantic preservation modification at the medium-risk level or a template replacement and rewriting at the high-risk level, after the target input content is generated, the system replaces the initial input content in the interactive interface in real time through the front-end API. The replacement process takes ≤300ms (ensuring that the user does not perceive any delay), and the cache path of the original content is hidden to prevent the user from recovering it through technical means.
[0044] As described above, by invoking the security rewriting engine to perform semantically preservative modifications on the initial input content to generate the target input content when the risk level is medium, it is possible to eliminate illegal content while preserving the user's core intent, avoiding interaction interruptions caused by simple interception, and accurately balancing compliance and user experience. When the risk level is high, the initial input content is discarded, and a standard response template relevant to the current dialogue context is retrieved from a pre-built red-line knowledge base as the target input content, avoiding mechanically repeating violation prompts and improving user acceptance of compliance processing. Furthermore, during semantically preservative modification processing, a multi-strategy combination mechanism dynamically selects the corresponding rewriting scheme for different violation types, significantly improving the compliance rate of input content at the medium-risk level.
[0045] Figure 4 This is a flowchart of another method for automatically generating interactive content provided in an embodiment of this application, see below. Figure 4 The method for automatically generating this interactive content specifically includes: S131. Modal recognition processing is performed on the initial response content to obtain the response type.
[0046] After generating the initial response content in S13, modality recognition processing is performed on the initial response content to obtain the response type. The response type includes text, image, audio, and / or video types. For example, feature parsing is performed on the initial response content (which may be single-modal or mixed-modal) to extract modality-specific features. The response type corresponding to the initial response content is determined based on the extracted modality-specific features. For instance, if character encoding features and features without pixels or waveforms are extracted, the response type is determined to be text; if pixel matrix features and file header identifier features (such as .jpg / .png format features) are extracted, the response type is determined to be image; if waveform amplitude features and sampling rate information (such as 16kHz standard sampling) are extracted, the response type is determined to be audio; if frame sequence features (continuous image frames + timestamps) and the inclusion of audio tracks are extracted, the response type is determined to be video. If two or more modality-specific features are detected simultaneously (such as "text + image" or "video + subtitle"), it is marked as a mixed modality and further divided into submodal types, such as text submodal type and image submodal type.
[0047] S132. Based on the response type, call the corresponding compliance detection model to identify and process violations in the initial response content, and process the identified violations to obtain the target response content.
[0048] After determining the response type in S131, the corresponding compliance detection model can be invoked to identify violations in the initial response content based on the response type, and the identified violations can be processed to obtain the target response content. For example, when the response type is text, the corresponding text classification model is invoked to identify violating text fragments in the initial response content, and the detected violating text fragments are processed to obtain the corresponding compliant target response content.
[0049] As described above, by employing full-modal violation detection, the limitations of traditional single-text detection are overcome. It can cover text, images, audio, video, and mixed modalities, thereby improving the coverage of violation detection across various modalities and ultimately enhancing the accuracy of violation identification. Furthermore, by performing appropriate compliance processing on the initial response content for different modal response types to obtain the target response content, the risk of violation in the response content is reduced, thus improving the security of the output response content.
[0050] Figure 5 This is a flowchart of another method for automatically generating interactive content provided in the embodiments of this application, referred to... Figure 5 The method for automatically generating this interactive content specifically includes: S1321. When the response type is text, the initial response content is detected for illegal text fragments by combining the text content security detection model with the red line knowledge base and sensitive word database. The detected illegal text fragments are then deleted or replaced with neutral expressions to obtain the target response content.
[0051] After performing modal recognition processing on the initial response content in S131 to obtain the response type, when the response type is text, a text content security detection model combined with a redline knowledge base and a sensitive word database is used to detect illegal text fragments in the initial response content. The detected illegal text fragments are then deleted or replaced with neutral expressions to obtain the target response content. For example, when the response type is text, the initial response content (i.e., text content) is input into the set text content security detection model for contextual understanding to identify sentences or paragraphs with corresponding illegal intent. Simultaneously, the redline knowledge base (containing a list of top-level sensitive topics that are absolutely prohibited from discussion) and the sensitive word database (containing dynamically updated sensitive words) are activated for rapid matching to output the marked illegal text fragments and their specific violation types and confidence levels. Violation text fragments that directly cross the red line, are extremely dangerous, or are irreparable are directly deleted to obtain the target response content. For minor or correctable violation text fragments (such as extreme adjectives or vulgar language), a neutral expression replacement engine can be activated to replace the illegal expressions with neutral words to obtain the target response content.
[0052] S1322. When the response type is image type, identify the illegal visual elements through a multi-label image classification model, and replace the identified illegal visual elements with compliance warning images or perform pixel-level blurring on the illegal visual elements to obtain the target response content.
[0053] After performing modal recognition processing on the initial response content in step S131 to obtain the response type, if the response type is an image, a multi-label image classification model is used to identify the non-compliant visual elements. The identified non-compliant visual elements are then replaced with a compliance warning image or pixel-level blurred to obtain the target response content. For example, the initial response content (i.e., the image file) can be input into the multi-label image classification model for non-compliant visual element identification, and the marked non-compliant visual elements can be output. The multi-label image classification model can not only identify the overall image category but also locate the specific position (boundary box) of the non-compliant visual elements. If the identified non-compliant visual element is partially non-compliant, only the pixels within the bounding box of the identified non-compliant visual element can be processed with Gaussian blur, mosaic, or pixelation, while the rest of the image (initial response content) remains unchanged. If the identified non-compliant visual element is entirely non-compliant, the original image can be completely discarded and replaced with a pre-set compliance warning image, such as a standard image with the text "This image violates security policy." The processed new image file is then used as the target response content.
[0054] S1323. When the reply type is audio, identify the illegal audio segment through the voiceprint feature recognition model, and mute or replace the illegal audio segment with the standard prompt tone to obtain the target reply content.
[0055] After performing modal recognition processing on the initial response content in step S131 to obtain the response type, when the response type is audio, a voiceprint feature recognition model is used to identify illegal audio segments. These segments are then either muted or replaced with standard prompts to obtain the target response content. For example, the audio (initial response content) is first converted to text, and then a text compliance model is used for detection. Next, the initial response content (i.e., audio) is input into a voiceprint feature recognition model for voice feature detection, sentiment analysis, and keyword detection, so that the model outputs the start and end timestamps of the marked illegal audio segments. After identifying illegal audio segments, with the support of an audio editing software library, the amplitude of the audio signal within the illegal time period can be set to zero, generating a silent segment, such as a "beep" sound similar to a broadcast. Alternatively, the audio within the illegal time period can be replaced with a pre-recorded, gentle standard prompt (such as a short prompt effect or a "content filtered here" voice prompt). The re-synthesized compliant audio file is then used as the target response content.
[0056] The above-described precise local processing eliminates the risk of violations while preserving the core intent of the original content to the greatest extent possible, achieving minimal processing and maximizing the retention of effective information. Corresponding detection models are used to detect violations in initial responses of different modalities, ensuring the accuracy and recall rate of violation identification. Furthermore, by replacing the content with compliance warning images or standard prompts, not only are violations blocked, but social norms and safety boundaries are also conveyed to users, serving an educational and guiding role.
[0057] Figure 6 This is a flowchart of another method for automatically generating interactive content provided in an embodiment of this application, see below. Figure 6 The method for automatically generating this interactive content specifically includes: S1324. When the response type is video, extract the image and audio data frame by frame.
[0058] After performing modal recognition processing on the initial response content in step S131 to obtain the response type, when the response type of the initial response content is video, image and audio data are extracted frame by frame. For example, when the response type of the initial response content is video, the initial response content (i.e., video) can be decomposed and extracted using a corresponding video image extraction tool. A keyframe extraction strategy can be set based on the video frame rate, for example, extracting one frame every three frames. If the video duration is ≤10 seconds, extracting frame by frame ensures no illegal frames are missed. Simultaneously, the initial response content (i.e., video) can be separated into audio tracks using a corresponding audio extraction tool to extract the corresponding audio data.
[0059] S1325. Identify the illegal visual elements in each image using a multi-label image classification model, and replace the identified illegal visual elements with compliance warning images or perform pixel-level blurring on the illegal visual elements to obtain the target image content.
[0060] After extracting image and audio data in S1324, a multi-label image classification model is used to identify the illegal visual elements in each image. The identified illegal visual elements are then replaced with a compliance warning image or pixel-level blurred to obtain the target image content. For example, the extracted images are input into the multi-label image classification model, which classifies each frame and outputs the confidence score for each type of illegal visual element. Visual elements with a confidence score greater than a preset threshold are identified as illegal visual elements. The multi-label image classification model also outputs the specific location (boundary box) of the illegal visual element. If the identified illegal visual element is partially illegal, only the pixels within the bounding box of the identified illegal visual element can be Gaussian blurred, mosaicked, or pixelated, while the rest of the image (initial response content) remains unchanged. If the identified illegal visual element is entirely illegal, the original image can be completely discarded and replaced with a preset compliance warning image, such as a standard image with the text "This image violates security policy." If the corresponding image contains no illegal visual elements, the original frame image is directly retained. After processing the relevant non-compliant visual elements, the processed image frames are arranged in the original video timestamp order to obtain the target image content. For example, an inter-frame difference algorithm (calculating the pixel differences between adjacent frames) can be used to detect whether there are any jumps in the processed image (such as excessive differences between the replaced high-risk frame and the frames before and after); if the difference is ≥50%, a fade-in / fade-out transition frame (generated based on frame interpolation) is added before and after the replaced frame to ensure smooth playback and avoid noticeable stuttering for the user.
[0061] S1326. The audio data is processed by an audio detection model to identify illegal audio segments, and the illegal audio segments are muted or replaced with standard prompts to obtain the target audio content.
[0062] After extracting the image and audio data in S1324, an audio detection model can be used to process the audio data to identify illegal audio segments. These segments can then be muted or replaced with standard prompts to obtain the target audio content. For example, the extracted audio data can be input into the audio detection model, which classifies and detects the audio data, outputting the start and end timestamps of the marked illegal audio segments. After identifying illegal audio segments, with the support of an audio editing software library, the amplitude of the audio signal within the illegal time period can be set to zero, generating a silent sound, such as a "beep" sound similar to a broadcast. Alternatively, the audio within the illegal time period can be replaced with a pre-recorded, gentle standard prompt (such as a short prompt effect or a "content filtered here" voice prompt). The re-synthesized compliant audio file is then used as the target audio content.
[0063] S1327. Integrate the target image content and the target audio content to obtain the complete target response content.
[0064] After obtaining the target image content in S1325 and the target audio content in S1326, the target image content and the target audio content are integrated to obtain the complete target response content. For example, a corresponding audio-visual synthesis tool can be used to integrate the aforementioned target image content (i.e., video frames without audio) with the target audio content according to the corresponding timestamps, and output the corresponding complete target response content (i.e., target video) containing video frames and audio data.
[0065] In one embodiment, a secondary inspection can be performed on the integrated target video. For example, 10% of the image frames can be re-extracted to detect illegal elements, and 20% of the audio segments can be randomly selected to detect illegal speech, ensuring that no illegal content remains. Additionally, the video can be played in a simulated user interaction scenario to verify audio-visual synchronization (e.g., latency ≤ 50ms), video smoothness (e.g., no obvious stuttering), and the completeness of processing identifiers (e.g., warning images and prompts are displayed / played normally). After the aforementioned secondary inspection and playback tests pass, the complete target response content is output, and a processing log (including original video information, location of illegal segments, and processing strategies) is recorded simultaneously for compliance auditing.
[0066] The above-mentioned collaborative detection mechanism of video frames and audio data for the initial response content of video type overcomes the limitations of traditional single-modal detection. It can effectively identify complex violation situations such as "the picture is compliant but the audio is non-compliant" or "partial picture violation and hidden audio violation", which significantly improves the accuracy of violation identification of video type response content. In addition, by performing non-compliant content removal processing on video image frames and audio data respectively, the risk of non-compliance of the final output target response content is significantly reduced, thereby improving the security of automatically generated interactive content.
[0067] The above approach breaks through the fragmented limitations of traditional automated response systems, which emphasize "replying but neglecting control," and constructs a comprehensive compliance system covering the entire process from initial user input to final response display. During the input stage, semantic analysis and multi-dimensional risk grading are used to identify hidden violations in advance. During the input intervention stage, differentiated processing is applied according to risk level to block the transmission of risky input. During the response stage, dedicated compliance checks are performed on four modalities—text, image, audio, and video—to intercept non-compliant content at the output end. Non-compliant content is replaced in real-time on the interactive interface to prevent the display of inappropriate information. This achieves comprehensive, seamless control across the entire chain, significantly reducing the output rate of non-compliant content and effectively eliminating compliance blind spots inherent in traditional solutions. Furthermore, regarding risk level, a risk score is calculated based on a dual probability distribution of topic category and intent category. This score is then adjusted using user image factors and scenario factors to obtain a comprehensive evaluation value, greatly improving the accuracy of judgments and effectively reducing the probability of misjudgments.
[0068] Based on the above embodiments, Figure 7 This is a schematic diagram of an automatic interactive content generation device provided in an embodiment of this application. (Reference) Figure 7 The automatic interactive content generation device provided in this embodiment specifically includes: a response module 21, a content determination module 22, and a display module 23.
[0069] Among them, the response module 21 is used to respond to the initial input content input by the interactive interface and perform semantic analysis on the initial input content to determine the corresponding risk level; Content determination module 22 is used to modify the initial input content according to the risk level to obtain the target input content; Display module 23 is used to replace the initial input content with the target input content in the interactive interface; The content determination module 22 is also used to generate initial response content based on the initial input content and the set response template, and to perform multimodal compliance filtering on the initial response content to obtain the target response content; Display module 23 is also used to display the target reply content in the interactive interface.
[0070] In one embodiment, the content determination module 22 includes: a type determination submodule and a content determination submodule; The type determination submodule is used to perform modal recognition processing on the initial response content to obtain the response type; The content determination submodule is used to call the corresponding compliance detection model based on the response type to identify violations in the initial response content, and to process the identified violations into compliance content to obtain the target response content.
[0071] In one implementation, the content determination submodule includes: a text content determination unit, an image content determination unit, and an audio content determination unit; The text content determination unit is used to detect illegal text fragments in the initial response content by combining the text content security detection model with the red line knowledge base and sensitive word database when the response type is text. The detected illegal text fragments are then deleted or replaced with neutral expressions to obtain the target response content. The image content determination unit is used to identify illegal visual elements through a multi-label image classification model when the response type is image, and replace the identified illegal visual elements with compliance warning images or perform pixel-level blurring on the illegal visual elements to obtain the target response content. The audio content determination unit is used to identify illegal audio segments through a voiceprint feature recognition model when the response type is audio, and to mute or replace the illegal audio segments with standard prompts to obtain the target response content.
[0072] In one embodiment, the content determination submodule further includes: an image and audio extraction unit, an image processing unit, an audio processing unit, and a video content integration unit; The image and audio extraction unit is used to extract image and audio data frame by frame when the response type is video. The image processing unit is used to identify the illegal visual elements in each image through a multi-label image classification model, and replace the identified illegal visual elements with compliance warning images or perform pixel-level blurring on the illegal visual elements to obtain the target image content. The audio processing unit is used to detect and process audio data through an audio detection model to identify illegal audio segments, and to mute or replace the illegal audio segments with standard prompts to obtain the target audio content. The video content integration unit is used to integrate the target image content and the target audio content to obtain the complete target response content.
[0073] In one embodiment, the response module 21 includes: a first probability determination submodule, a second probability determination submodule, and a risk level determination submodule; The first probability determination submodule is used to perform topic classification processing on the initial input content to determine the first probability distribution of the initial input content belonging to the set topic category; The second probability determination submodule is used to perform semantic intent classification processing on the initial input content to determine the second probability distribution of the initial input content belonging to the set intent category; The risk level determination submodule is used to determine the risk level of the initial input content based on the first probability distribution and the second probability distribution.
[0074] In one embodiment, the content determination module 22 further includes: a first rewriting submodule and a second rewriting submodule; The first rewriting submodule is used to call the security rewriting engine to perform semantic preservation modification of the initial input content to generate the target input content when the risk level is medium risk level. The second rewriting submodule is used to discard the initial input content when the risk level is high, and call the standard response template related to the current dialogue context from the pre-built redline knowledge base as the target input content.
[0075] In one embodiment, the first rewriting submodule includes: a first rewriting unit, a second rewriting unit, a third rewriting unit, and a fourth rewriting unit; The first rewriting unit is used to call the secure rewriting engine to identify sensitive words in the initial input content and replace the sensitive words with semantically similar neutral words to generate the target input content; The second rewriting unit is used to call the security rewriting engine to extract the core intent of the initial input content, and reconstruct the content using compliant expression based on the core intent to obtain the target input content; The third rewriting unit is used to call the secure rewriting engine to delete sensitive words that are not core semantics to obtain the target input content; The fourth rewriting unit is used to call the secure rewriting engine to encode sensitive words in the core semantics to obtain the target input content.
[0076] The automatic interactive content generation device provided in this application embodiment can be used to execute the automatic interactive content generation method provided in the above embodiment, and has corresponding functions and beneficial effects.
[0077] This application provides an automatic interactive content generation device, referring to... Figure 8 The automatic interactive content generation device includes: a processor 31, a memory 32, a communication module 33, an input device 34, and an output device 35. The number of processors and the number of memories in the automatic interactive content generation device can be one or more. The processor, memory, communication module, input device, and output device of the automatic interactive content generation device can be connected via a bus or other means.
[0078] The memory 32, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the automatic generation method of interactive content described in any embodiment of this application (e.g., the response module, content determination module, and display module in the automatic generation device of interactive content). The memory may primarily include a program storage area and a data storage area, wherein the program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the device, etc. Furthermore, the memory may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0079] The communication module 33 is used for data transmission.
[0080] The processor 31 executes various functional applications and data processing of the device by running software programs, instructions and modules stored in the memory, thereby realizing the above-mentioned method for automatically generating interactive content.
[0081] Input device 34 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 35 may include display devices such as a display screen.
[0082] The interactive content automatic generation device provided above can be used to execute the interactive content automatic generation method provided in the above embodiments, and has corresponding functions and beneficial effects.
[0083] This application embodiment also provides a storage medium for storing computer-executable instructions. When executed by a computer processor, the computer-executable instructions are used to execute a method for automatically generating interactive content. The method for automatically generating interactive content includes: responding to initial input content input through an interactive interface, performing semantic analysis on the initial input content to determine a corresponding risk level; modifying the initial input content according to the risk level to obtain target input content, and replacing the initial input content with the target input content in the interactive interface; generating initial response content based on the initial input content and a set response template, performing multimodal compliance filtering on the initial response content to obtain target response content, and displaying the target response content in the interactive interface.
[0084] Storage medium – any type of memory device or storage device. The term “storage medium” is intended to include: mounting media, such as CD-ROM, floppy disk, or magnetic tape devices; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (e.g., hard disk or optical storage); registers or other similar types of memory elements, etc. Storage medium may also include other types of memory or combinations thereof. Furthermore, storage medium may reside in a first computer system in which the program is executed, or it may reside in a different second computer system connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term “storage medium” can include two or more storage media residing in different locations (e.g., in different computer systems connected via a network). Storage medium may store program instructions (e.g., specifically implemented as a computer program) executable by one or more processors.
[0085] Of course, the storage medium for storing computer-executable instructions provided in the embodiments of this application is not limited to the automatic generation method of interactive content as described above, but can also perform related operations in the automatic generation method of interactive content provided in any embodiment of this application.
[0086] The automatic interactive content generation device, storage medium, and automatic interactive content generation equipment provided in the above embodiments can execute the automatic interactive content generation method provided in any embodiment of this application. For technical details not described in detail in the above embodiments, please refer to the automatic interactive content generation method provided in any embodiment of this application.
[0087] The above description is merely a preferred embodiment and the technical principles employed in this application. This application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of this application, the scope of which is determined by the scope of the claims.
Claims
1. A method for automatically generating interactive content, characterized in that, include: In response to the initial input content input through the interactive interface, semantic analysis is performed on the initial input content to determine the corresponding risk level; Based on the risk level, the initial input content is modified to obtain the target input content, and the initial input content is replaced with the target input content in the interactive interface; An initial response is generated based on the initial input and the set response template. The initial response is then subjected to multimodal compliance filtering to obtain the target response, which is then displayed in the interactive interface.
2. The method according to claim 1, characterized in that, The process of performing multimodal compliance filtering on the initial response content to obtain the target response content includes: Modality recognition processing is performed on the initial response content to obtain the response type; Based on the response type, the corresponding compliance detection model is invoked to identify violations in the initial response content, and the identified violations are then processed to obtain the target response content.
3. The method according to claim 2, characterized in that, The step of calling the corresponding compliance detection model according to the response type to identify violations in the initial response content, and then processing the identified violations to obtain the target response content, includes: When the response type is text, the initial response content is detected for illegal text fragments by a text content security detection model combined with a red line knowledge base and a sensitive word database. The detected illegal text fragments are then deleted or replaced with neutral expressions to obtain the target response content. When the response type is an image, a multi-label image classification model is used to identify the illegal visual elements, and the identified illegal visual elements are replaced with compliance warning images or the illegal visual elements are blurred at the pixel level to obtain the target response content. When the response type is audio, the illegal audio segment is identified by the voiceprint feature recognition model, and the illegal audio segment is either muted or replaced with a standard prompt tone to obtain the target response content.
4. The method according to claim 2, characterized in that, The step of calling the corresponding compliance detection model according to the response type to identify violations in the initial response content, and then processing the identified violations to obtain the target response content, includes: When the response type is video, extract the image and audio data frame by frame. The non-compliant visual elements in each image are identified by a multi-label image classification model, and the identified non-compliant visual elements are replaced with compliance warning images or the non-compliant visual elements are blurred at the pixel level to obtain the target image content. The audio data is processed by an audio detection model to identify illegal audio segments, and the illegal audio segments are either muted or replaced with standard prompts to obtain the target audio content. The target image content and the target audio content are integrated to obtain the complete target response content.
5. The method according to claim 1, characterized in that, The step of performing semantic analysis on the initial input content to determine the corresponding risk level includes: The initial input content is subjected to topic classification processing to determine the first probability distribution of the initial input content belonging to the set topic category; The initial input content is subjected to semantic intent classification processing to determine the second probability distribution of the initial input content belonging to the set intent category; The risk level of the initial input content is determined based on the first probability distribution and the second probability distribution.
6. The method according to any one of claims 1-5, characterized in that, Based on the risk level, the initial input content is modified to obtain the target input content, including: When the risk level is medium risk, the security rewrite engine is invoked to perform semantic preservation modification on the initial input content to generate the target input content. When the risk level is high, the initial input content is discarded, and a standard response template related to the current dialogue context is retrieved from the pre-built redline knowledge base as the target input content.
7. The method according to claim 6, characterized in that, The step of calling the secure rewrite engine to perform semantically preservative modification processing on the initial input content to generate the target input content includes: The security rewrite engine is invoked to identify sensitive words in the initial input content, and the sensitive words are replaced with semantically similar neutral words to generate the target input content; And / or, invoke the secure rewrite engine to extract the core intent of the initial input content, and reconstruct the content using compliant expression based on the core intent to obtain the target input content; And / or, call the secure rewriting engine to delete sensitive words that are not core semantics to obtain the target input content; And / or, call the security rewriting engine to mask sensitive words in the core semantics to obtain the target input content.
8. An automatic interactive content generation device, characterized in that, include: The response module is used to respond to the initial input content input by the interactive interface, and to perform semantic analysis on the initial input content to determine the corresponding risk level; The content determination module is used to modify the initial input content according to the risk level to obtain the target input content. The display module is used to replace the initial input content with the target input content in the interactive interface; The content determination module is further configured to generate initial response content based on the initial input content and the set response template, and perform multimodal compliance filtering on the initial response content to obtain the target response content; The display module is also used to display the target reply content in the interactive interface.
9. An automatic interactive content generation device, characterized in that, include: Memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.
10. A storage medium for storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a processor, are used to perform the method as described in any one of claims 1-7.