Intelligent toy control method based on semantic risk detection and related equipment

By performing semantic risk detection and analysis on children's voice data, identifying potential risks and implementing safety response strategies, the problem of misjudgment and omission in existing systems when children's language is not standard is solved, ensuring the safety of children's voice interaction.

CN121983033APending Publication Date: 2026-05-05SHENZHEN LUKA DR TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN LUKA DR TECHNOLOGY CO LTD
Filing Date
2025-12-24
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing voice interaction systems for children are prone to semantic misjudgment or omission in multi-turn dialogues, ambiguous contexts, or when children's language expression is not standard. They lack real-time detection and intervention mechanisms for children's voice behavior and cannot proactively identify potential risky behaviors.

Method used

The target text data is analyzed by a pre-set semantic risk detection model, and the semantic risk analysis results and risk level labels are output. Based on the results, a safety response strategy is determined in the safety response strategy library, and the smart toy is controlled to execute the corresponding strategy.

Benefits of technology

It enables accurate identification and real-time intervention of potential risky behaviors when children's language expression is not standardized, avoiding the output of unsafe content and protecting children's physical and mental health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983033A_ABST
    Figure CN121983033A_ABST
Patent Text Reader

Abstract

The invention is applicable to the technical field of artificial intelligence, and provides an intelligent toy control method based on semantic risk detection, which comprises the following steps: acquiring voice data of a target user, and performing voice-to-text processing on the voice data to obtain corresponding target text data; performing semantic risk analysis on the target text data through a preset semantic risk detection model, and outputting a semantic risk analysis result corresponding to the target text data and a corresponding risk level tag; and based on the semantic risk analysis result and the corresponding risk level label, determining a target safety response strategy in a preset safety response strategy library, and controlling the intelligent toy to execute the target safety response strategy. The problems that an existing method depends on a universal semantic recognition and content filtering algorithm, under the condition that language expression is not standard, semantic misjudgment or missed judgment is prone to occurring, a real-time detection and intervention mechanism for voice behaviors is lacked, and potential risk behaviors cannot be actively recognized in the interaction process are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and in particular relates to a method and related equipment for controlling intelligent toys based on semantic risk detection. Background Technology

[0002] In children's voice interaction scenarios, plush toys are often used to accompany children in natural voice dialogues. However, existing children's voice dialogue systems mainly rely on general semantic recognition and content filtering algorithms, which have weak capabilities in recognizing sensitive content involving violence, pornography, intimidation, or inducement. Especially in multi-turn dialogues, ambiguous contexts, or when children's language expression is not standard, semantic misjudgments or omissions are prone to occur, leading to the system outputting unsafe content and affecting children's physical and mental health. In addition, existing solutions generally lack real-time detection and intervention mechanisms for children's voice behaviors (such as abnormal emotions, dialogue inducement, and imitation risks), and cannot proactively identify potential risk behaviors during interaction. Therefore, there is an urgent need for an intelligent toy control method that combines semantic risk analysis and sentiment analysis to solve the problems of existing methods relying on general semantic recognition and content filtering algorithms, being prone to semantic misjudgments or omissions when children's language expression is not standard, lacking real-time detection and intervention mechanisms for children's voice behaviors, and being unable to proactively identify potential risk behaviors during interaction. Summary of the Invention

[0003] This application provides a method for controlling smart toys based on semantic risk detection. This method addresses the problems of existing methods that rely on general semantic recognition and content filtering algorithms, are prone to semantic misjudgment or omission when children's language expression is not standard, lack real-time detection and intervention mechanisms for children's speech behavior, and are unable to proactively identify potential risky behaviors during interaction. By using a preset semantic risk detection model to perform semantic risk analysis on target text data, the method outputs the semantic risk analysis results and corresponding risk level labels for the target text data. Based on the semantic risk analysis results and corresponding risk level labels, a target safety response strategy is determined from a preset safety response strategy library, and the smart toy is controlled to execute the target safety response strategy. This solves the problems of existing methods that rely on general semantic recognition and content filtering algorithms, are prone to semantic misjudgment or omission when children's language expression is not standard, lack real-time detection and intervention mechanisms for children's speech behavior, and are unable to proactively identify potential risky behaviors during interaction.

[0004] In a first aspect, embodiments of this application provide a smart toy control method based on semantic risk detection, the method comprising the following steps: Acquire the voice data of the target user, and perform speech-to-text processing on the voice data to obtain the corresponding target text data; The target text data is subjected to semantic risk analysis using a preset semantic risk detection model, and the semantic risk analysis results and corresponding risk level labels are output. Based on the semantic risk analysis results and the corresponding risk level labels, a target safety response strategy is determined from the preset safety response strategy library, and the smart toy is controlled to execute the target safety response strategy.

[0005] Optionally, the step of performing speech-to-text processing on the speech data to obtain corresponding text data includes: The speech data is processed into text using a preset speech recognition model to obtain initial text data. The initial text data is standardized to obtain the target text data.

[0006] Optionally, the step of performing semantic risk analysis on the target text data using a preset semantic risk detection model, and outputting the semantic risk analysis results and corresponding risk level labels for the target text data, includes: The target text data is subjected to a first semantic risk analysis using a preset semantic risk detection model to obtain the first semantic risk analysis result and the corresponding first risk level label. When the first semantic risk analysis result is a risk miss, a second semantic risk analysis is performed on the target text data to obtain the second semantic risk analysis result and the corresponding second risk level label. When the second semantic risk analysis result is a risk miss, the target text data is subjected to a third semantic risk analysis in combination with contextual information and a sentiment dictionary to obtain the third semantic risk analysis result and the corresponding third risk level label.

[0007] Optionally, the step of performing a second semantic risk analysis on the target text data to obtain the second semantic risk analysis result and the corresponding second risk level label includes: The target text data is subjected to semantic extraction processing to obtain the text semantic features corresponding to the target text data; A second semantic risk analysis is performed on the semantic features of the text to obtain the second semantic risk analysis results and the corresponding second risk level labels.

[0008] Optionally, the step of combining contextual information and a sentiment dictionary to perform third semantic risk analysis on the target text data, obtaining the third semantic risk analysis results and corresponding third risk level labels, includes: By combining the emotion dictionary, the emotion change data of the target text data is calculated; Analyze the contextual semantic features of the target text data by combining contextual information; Based on the emotion change data and the contextual semantic features, the target text data is subjected to third semantic risk analysis to obtain the third semantic risk analysis results and the corresponding third risk level label.

[0009] Optionally, determining the target security response strategy from a preset security response strategy library based on the semantic risk analysis results and the corresponding risk level labels includes: When the semantic risk analysis result is a risk hit, multiple candidate security response strategies are selected from the preset security response strategy library according to the risk level label. By combining the cache management library, the target security response strategy is determined from multiple candidate security response strategies.

[0010] Optionally, after the intelligent toy is controlled to execute the target safety response strategy, the method further includes: Record current session data; The current session data, the semantic risk analysis results, and the corresponding risk level labels are used to generate log data and report it to the cloud.

[0011] Secondly, embodiments of this application provide a smart toy control device based on semantic risk detection, the smart toy control device based on semantic risk detection comprising: The first processing module is used to acquire the voice data of the target user and perform voice-to-text processing on the voice data to obtain the corresponding target text data. The second processing module is used to perform semantic risk analysis on the target text data through a preset semantic risk detection model, and output the semantic risk analysis results and corresponding risk level labels of the target text data. The control module is used to determine the target safety response strategy from the preset safety response strategy library based on the semantic risk analysis results and the corresponding risk level labels, and to control the smart toy to execute the target safety response strategy.

[0012] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the intelligent toy control method based on semantic risk detection provided in embodiments of the present invention.

[0013] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps in the intelligent toy control method based on semantic risk detection provided in the embodiments of the present invention.

[0014] The above-mentioned solution of this application has the following beneficial effects: It acquires the voice data of the target user and performs speech-to-text processing on the voice data to obtain the corresponding target text data; it performs semantic risk analysis on the target text data through a preset semantic risk detection model, outputting the semantic risk analysis results and corresponding risk level labels for the target text data; based on the semantic risk analysis results and corresponding risk level labels, it determines the target safety response strategy in a preset safety response strategy library and controls the smart toy to execute the target safety response strategy. This invention solves the problems of existing methods that rely on general semantic recognition and content filtering algorithms, are prone to semantic misjudgment or omission when children's language expression is not standardized, lack real-time detection and intervention mechanisms for children's voice behavior, and are unable to proactively identify potential risk behaviors during interaction.

[0015] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating a smart toy control method based on semantic risk detection, provided as an embodiment of this application; Figure 2 This is a schematic diagram of the structure of an intelligent toy control device based on semantic risk detection according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0019] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0020] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0021] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0022] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0023] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0024] like Figure 1 As shown, Figure 1 This is a flowchart of a smart toy control method based on semantic risk detection provided by an embodiment of the present invention. The smart toy control method based on semantic risk detection includes the following steps: 101. Obtain the target user's voice data and perform speech-to-text processing on the voice data to obtain the corresponding target text data.

[0025] In this embodiment of the invention, the above-described intelligent toy control method based on semantic risk detection can be applied to a server. The server and the intelligent toy communicate with each other. The intelligent toy includes a voice acquisition device and an image acquisition device. The voice acquisition device is used to acquire voice data, and the image acquisition device is used to acquire image data. When using the intelligent toy for the first time, the user can set user information, including the user's name, age, hobbies, etc. The intelligent toy can also have a wake-up word set by the user, or a special nickname set for a specific user.

[0026] The target users mentioned above can be children who wake up the smart toy or children who interact with the smart toy.

[0027] The aforementioned voice data can be captured in real time through the voice acquisition device of a smart toy, and can include words, phrases, sentences, or non-standard pronunciations, humming, interjections, etc.

[0028] It should be noted that when the smart toy is activated by the target user, the target user's voice data can be collected through the smart toy's voice acquisition device.

[0029] The above-mentioned speech-to-text processing can be a process of converting speech data into text format.

[0030] The target text data mentioned above can be text data obtained after speech-to-text processing of speech data.

[0031] It should be noted that speech data can be processed into text using speech recognition models. These models can convert human speech signals into text. By analyzing information such as the spectrum, acoustic features, and language models of the speech signal, speech recognition models identify and decode the text content within the speech, thus achieving speech-to-text conversion.

[0032] 102. Perform semantic risk analysis on the target text data using a pre-set semantic risk detection model, and output the semantic risk analysis results and corresponding risk level labels for the target text data.

[0033] In this embodiment of the invention, the aforementioned preset semantic risk detection model can be a semantic risk detection model built based on deep learning or machine learning, such as BERT+LSTM, TextCNN, etc. BERT+LSTM is a hybrid deep learning architecture that combines the advantages of bidirectional context encoding and sequence dependency modeling, primarily used for natural language processing tasks such as text classification. BERT, through the multi-head self-attention mechanism of the Transformer architecture, simultaneously captures the contextual information of words, generating context-related embedding vectors. These context-related embedding vectors are then input into the LSTM, which, through its memory units and gating mechanism, further models long-term dependencies in the sequence, making it particularly suitable for processing long texts or complex syntactic structures. In long text classification scenarios, the model can employ an overlapping segmentation strategy (such as Overlap Split), segmenting the document and encoding each segment using BERT, then connecting inter-segment dependencies using LSTM, and finally aggregating the representations through an attention mechanism for classification. TextCNN (Text Convolutional Neural Network) is a deep learning model specifically designed for text classification, applying the ideas of convolutional neural networks (CNN) to natural language processing tasks. The core idea of ​​TextCNN is to automatically extract local features from text through one-dimensional convolutional operations, combine this with pooling operations to capture key semantic information, and finally perform classification through fully connected layers. The aforementioned pre-defined semantic risk detection model can identify potential threats or uncertainties in the target text data at the semantic level, thereby reducing security losses or decision-making errors caused by misunderstandings, ambiguities, or mistakes. The aforementioned semantic risk analysis can be a deep analysis of target text data using a pre-defined semantic risk detection model to more accurately identify and assess sensitive semantic risks, and dynamically analyze the processing of semantic risks in the target text.

[0034] The aforementioned semantic risk analysis results can be a complete set of conclusive data generated after performing semantic risk analysis on target text data using a pre-defined semantic risk detection model. The semantic risk analysis results include whether a risk is detected and the type of risk. The risk determination can be based on the risk type, which can represent the category to which the risk belongs, such as violence, pornography, inducement, intimidation, emotional coercion, or privacy intrusion.

[0035] The aforementioned risk level labels can be structured, dynamic identifiers generated after semantic risk analysis of the target text data to determine the results of the semantic risk analysis. The risk level labels include a severity classification of the risk. For example, risk level label = 0 indicates no risk; risk level label = 1 indicates explicit sensitive word detection; risk level label = 2 indicates implicit semantic risk; and risk level label = 3 indicates a high-risk behavioral pattern. It should be noted that the higher the level, the more concealed, severe, and persistent the risk.

[0036] In one possible implementation, for example, for a piece of text data containing risks such as violence and pornography, the risky content in the text data can be identified through a preset semantic risk detection model, and the semantic risk analysis results of the text data and the corresponding risk level labels can be given. This can help the management platform discover and process harmful information more quickly and protect the interests and safety of users.

[0037] 103. Based on the semantic risk analysis results and the corresponding risk level labels, determine the target safety response strategy in the preset safety response strategy library, and control the smart toy to execute the target safety response strategy.

[0038] In this embodiment of the invention, the aforementioned preset safety response strategy library can be a configurable and updatable rule database storing various predefined safety intervention logics and their triggering conditions. The preset safety response strategy library can output specific safety response strategies based on the input semantic risk analysis results and corresponding risk level labels, through rule matching or priority calculation. Each safety response strategy can be a complete action execution instruction, including the response action, response intensity, execution parameters, etc. The aforementioned response action can be a specific operation to be performed by the smart device, for example: Action 1 is to play a preset safety prompt voice (such as "I don't understand this problem, let's talk about something else?"); Action 2 is to actively change the topic and initiate a new, safe game or Q&A; Action 3 is to stop responding to the current round of dialogue and remain silent for a period of time; Action 4 is to trigger a local alarm (such as a toy flashing a red light) and record the event; Action 5 is to send a safety alarm notification through the parent's app, etc. The aforementioned response intensity can be the urgency or severity of the action. The aforementioned execution parameters can be the execution voice content, silence duration, notification template, etc.

[0039] The aforementioned target security response strategy can be a security response strategy determined from a pre-defined security response strategy library based on semantic risk analysis results and corresponding risk level labels. The target security response strategy specifies the concrete security response action, the action's execution parameters, priority, and instructions for coordination with the current session context.

[0040] In this embodiment of the invention, the present invention adopts a semantic risk detection model, which can accurately identify semantic risks in text data and the risk level labels corresponding to the semantic risks. Based on the semantic risks and the risk level labels corresponding to the semantic risks, the target safety response strategy is determined from the safety response strategy library, and the smart toy is controlled to execute the target safety response strategy. This can realize risk response and safety feedback in voice interaction and effectively prevent the output of unsafe content.

[0041] In this embodiment of the invention, voice data of the target user is acquired and processed into speech-to-text to obtain corresponding target text data. A preset semantic risk detection model is used to perform semantic risk analysis on the target text data, outputting the semantic risk analysis results and corresponding risk level labels. Based on the semantic risk analysis results and corresponding risk level labels, a target safety response strategy is determined from a preset safety response strategy library, and the smart toy is controlled to execute the target safety response strategy. This invention solves the problems of existing methods that rely on general semantic recognition and content filtering algorithms, are prone to semantic misjudgment or omission when children's language expression is not standardized, lack real-time detection and intervention mechanisms for children's speech behavior, and cannot proactively identify potential risky behaviors during interaction.

[0042] It is understood that in the specific implementation of this application, data such as voice data, text data, risk data, and user data are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required. Furthermore, the collection, use, and processing of related data, as well as the training, deployment, and invocation of algorithm models, must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0043] Optionally, in the step of processing speech data into text to obtain corresponding text data, the speech data can be processed into text using a preset speech recognition model to obtain initial text data; the initial text data can then be standardized to obtain target text data.

[0044] In this embodiment of the invention, the aforementioned voice data can be obtained by the voice acquisition device of a smart toy, and can be words, phrases, sentences, or non-standard pronunciations, humming, interjections, etc.

[0045] The above-mentioned preset speech recognition model can be a speech recognition model constructed based on deep learning or machine learning, such as a Transformer model, LSTM, etc. The above-mentioned Transformer model is a deep learning architecture based on the attention mechanism, mainly used to process sequence data, such as natural language. The core idea of the Transformer model is to dynamically calculate the relationships between input elements through the self-attention mechanism, so as to capture context information. The above-mentioned LSTM (Long Short-Term Memory Network) is a special recurrent neural network (RNN), specifically designed to process time series data and solve the difficulties that traditional RNNs are prone to encounter when capturing long-term dependencies. LSTM effectively remembers or forgets relevant information in the sequence by introducing a mechanism called "gate" to intelligently control the flow of information. The above-mentioned preset speech recognition model can convert human speech signals into text form. The speech recognition model analyzes information such as the spectrum, acoustic features, and language model of the speech signal, identifies and decodes the text content in the speech, so as to achieve the conversion between speech and text.

[0046] The above-mentioned speech-to-text processing can be a processing process of converting speech data into text form through a preset speech recognition model.

[0047] The above-mentioned initial text data can be the text data obtained after performing speech-to-text conversion on speech data through a preset speech recognition model.

[0048] The above-mentioned standardization process can be a processing process of standardizing the colloquial and unclear referential expressions in the initial text data into complete sentences. It can be to convert the Chinese characters in the initial text data into pinyin through a pinyin conversion tool, which is convenient for subsequent semantic analysis and error correction. For example, convert "七饭" to pinyin "chifan", etc. The above-mentioned pinyin conversion tool can be a tool that automatically converts Chinese character text into the corresponding Chinese pinyin representation. The core principle of the pinyin conversion tool is based on the pronunciation rules of Chinese characters. By using the built-in pinyin database or algorithm, it identifies the initials, finals, and tones of each Chinese character, so as to generate standard pinyin. It can be to use natural language processing technology, a rule-based method or a sequence-to-sequence (Seq2Seq) model to perform grammar error correction on the initial text data. It can be to convert non-standard expressions in the initial text data, such as "打人" to the standard semantics "打架", "生气了啦" to the standard semantics "生气", etc. through a custom dictionary and semantic understanding. The above-mentioned natural language processing technology aims to enable the computer to understand, interpret, and generate human language, so as to achieve effective communication between humans and machines.

[0049] The above-mentioned target text data can be obtained after performing standardization processing on the initial text data.

[0050] Optionally, in the step of performing semantic risk analysis on the target text data using a preset semantic risk detection model and outputting the semantic risk analysis result and the corresponding risk level label for the target text data, the preset semantic risk detection model can be used to perform a first semantic risk analysis on the target text data to obtain the first semantic risk analysis result and the corresponding first risk level label; if the first semantic risk analysis result is a risk miss, a second semantic risk analysis is performed on the target text data to obtain the second semantic risk analysis result and the corresponding second risk level label; if the second semantic risk analysis result is a risk miss, a third semantic risk analysis is performed on the target text data in combination with contextual information and a sentiment dictionary to obtain the third semantic risk analysis result and the corresponding third risk level label.

[0051] In this embodiment of the invention, the aforementioned preset semantic risk detection model can be a semantic risk detection model built based on deep learning or machine learning, such as BERT+LSTM, TextCNN, etc. The aforementioned preset semantic risk prediction model can be a blacklist database constructed using a Trie tree. The aforementioned blacklist database can be a set of sensitive words and phrases, aiming to achieve millisecond-level accurate matching and interception of sensitive content. The aforementioned preset semantic risk detection model can identify potential threats or uncertainties in the semantics of target text data, thereby reducing security losses or decision-making errors caused by misunderstandings, ambiguities, or mistakes. The aforementioned Trie tree (also known as a prefix tree or dictionary tree) is a tree-shaped data structure used for efficiently storing and retrieving sets of strings. A characteristic of a Trie tree is that each node represents a character, and the sequence of characters on the path from the root node to any node constitutes a string.

[0052] The aforementioned semantic risk analysis can be a process of identifying, assessing, and managing potential semantic risks in text data at the semantic level. These semantic risks may include potential security threats or losses related to sensitive content such as violence, pornography, intimidation, or manipulation.

[0053] The aforementioned first semantic risk analysis can be a process of performing semantic risk analysis on target text data. This can involve inputting the target text data into a pre-defined semantic risk detection model, searching within the model's blacklist, and considering any sensitive keywords matched in the blacklist as detected sensitive semantics. The result of the first semantic risk analysis corresponding to the target text data, along with the corresponding first risk level identifier, is then output.

[0054] The aforementioned first semantic analysis result can be the semantic analysis result obtained after performing a first semantic risk analysis on the target text data through a preset semantic risk detection model. The semantic analysis result includes risk hit and risk miss.

[0055] The aforementioned first risk level identifier can be the risk level identifier corresponding to the risk hit in the first semantic analysis result. The first risk level identifier is used to quantitatively assess the degree of potential risk of the text data with the risk hit.

[0056] The aforementioned second semantic risk analysis can be a process of performing semantic risk analysis on the target text data when the first semantic risk analysis result is a no-risk detection. The aforementioned second semantic risk analysis can be a process of identifying implicit sensitive semantics in the target text data and analyzing the risks associated with these implicit sensitive semantics. These implicit sensitive semantics can be sensitive semantics that imply threats, negative implications, etc.

[0057] The above-mentioned second semantic risk analysis results can be semantic analysis results obtained after performing second semantic risk analysis on the target text data.

[0058] The aforementioned second risk level identifier can be the risk level identifier corresponding to the risk hit in the second semantic risk analysis result.

[0059] The aforementioned contextual information can be all information related to the current conversation, including historical conversation content, user personal information, and so on.

[0060] The aforementioned sentiment dictionary can be a core tool for sentiment analysis in natural language processing. The sentiment dictionary is essentially a structured collection of words, each of which is associated with sentiment polarity (such as positive, negative, or neutral) and possible sentiment intensity scores.

[0061] The aforementioned third semantic risk analysis can be a process of performing semantic risk analysis on the target text by combining contextual information and a sentiment lexicon when the second semantic risk analysis result is a no-risk match. This third semantic risk analysis can be a process of identifying potential high-risk metaphors in the target text data and analyzing the risks associated with these potential high-risk metaphors. These potential high-risk metaphors can be, during the interaction process, where certain literally harmless statements, due to their specific contextual sequence and accompanying strong or abnormal emotional signals, derive from implicit intentions that are threatening, manipulative, or harmful, such as "I'm afraid of the dark," "I don't know what to do," or "I'm so scared."

[0062] The aforementioned third semantic analysis result can be the semantic analysis result obtained after performing third semantic risk analysis on the target text data.

[0063] The aforementioned third risk level identifier can be the risk level identifier corresponding to the risk hit in the third semantic analysis result.

[0064] It should be noted that a pre-set semantic risk detection model can be used to perform a first semantic risk analysis on the target text data, obtaining the first semantic risk analysis result and the corresponding first risk level label. If the first semantic risk analysis result is a risk miss, a second semantic risk analysis can be performed on the target text data, obtaining the second semantic risk analysis result and the corresponding second risk level label. If the second semantic risk analysis result is a risk miss, a third semantic risk analysis can be performed on the target text data by combining contextual information and a sentiment dictionary, obtaining the third semantic risk analysis result and the corresponding third risk level label. This approach can more comprehensively assess the risk of text data, reduce misjudgments and omissions, and, through multi-layered analysis, better understand the context and intent of the text data, thereby improving the accuracy of risk identification.

[0065] Optionally, in the step of performing a second semantic risk analysis on the target text data to obtain the second semantic risk analysis result and the corresponding second risk level label, the target text data can be semantically extracted to obtain the text semantic features corresponding to the target text data; the text semantic features can be subjected to a second semantic risk analysis to obtain the second semantic risk analysis result and the corresponding second risk level label.

[0066] In this embodiment of the invention, the semantic extraction process described above can be a process of transforming target text data into a high-dimensional, dense vector representation containing contextual information. Specifically, it can be a process of extracting the contextual semantic representation of the target text data through a preset semantic risk detection model, converting the target text data into word vectors, and extracting fine-grained semantic features of the target text data.

[0067] The aforementioned text semantic features can be high-dimensional, dense numerical vectors or vector sequences output after semantic feature extraction of target text data. Text semantic features strip away the surface character form of the text, encoding its implied contextual semantics, syntactic relationships, and underlying intentions into a mathematical representation.

[0068] The aforementioned second semantic risk analysis can be a process of performing semantic risk analysis on the semantic features of text. This second semantic risk analysis can be a process of identifying implicit sensitive semantics within the semantic features of text and analyzing the risks associated with these implicit sensitive semantics. These implicit sensitive semantics can be sensitive semantics that imply threats, negative implications, etc.

[0069] The above-mentioned second semantic risk analysis results can be semantic risk analysis results obtained after performing second semantic risk analysis on the semantic features of the text.

[0070] The aforementioned second risk level label can be the risk level identifier corresponding to the second semantic risk analysis result.

[0071] In one possible embodiment, a pre-defined semantic risk detection model, BERT, can be used to extract the contextual semantic representation of the target text data, transforming the target text data into a sequence of word vectors. A high-dimensional semantic vector is then generated using BERT's Transformer structure. The output of BERT is used as the input to an LSTM, which captures long-term dependencies in the target text data through memory units, extracting fine-grained semantic features. A fully connected layer is added to the output of the LSTM to output classification results (e.g., sensitive, non-sensitive). A second semantic risk analysis is then performed on the classification results to obtain the second semantic risk analysis result corresponding to the target text data and the corresponding second risk level identifier.

[0072] Optionally, in the step of performing third semantic risk analysis on the target text data by combining contextual information and a sentiment dictionary to obtain the third semantic risk analysis results and the corresponding third risk level labels, the following steps can be taken: 1) Calculate the sentiment change data of the target text data by combining a sentiment dictionary; 2) Analyze the contextual semantic features of the target text data by combining contextual information; 3) Perform third semantic risk analysis on the target text data based on the sentiment change data and contextual semantic features to obtain the third semantic risk analysis results and the corresponding third risk level labels.

[0073] In this embodiment of the invention, the aforementioned sentiment dictionary can be a core tool for sentiment analysis in natural language processing. The sentiment dictionary is essentially a structured collection of words, each of which is associated with sentiment polarity (such as positive, negative or neutral) and possible sentiment intensity score.

[0074] The aforementioned sentiment change data can be sentiment change data corresponding to the target text data. Sentiment change data can be a time-series quantitative sequence generated by calculating the target text data based on a sentiment lexicon. Sentiment change data not only includes the polarity (positive / negative) and intensity of individual emotions, but more importantly, it reflects the trajectory, trend, and abrupt changes in emotional state over time. The aforementioned sentiment change data can include sentiment change trends, fluctuations, and abrupt changes.

[0075] Specifically, a quantifiable sentiment score can be calculated for the target text data based on a sentiment lexicon and TF-IDF (Term Frequency-Inverse Document Frequency). The current sentiment score and the sentiment scores of previous dialogue rounds are then arranged chronologically to form a sentiment value sequence. From this sequence, sentiment change data reflecting the dynamic evolution of psychological states can be calculated and extracted. More specifically, the vocabulary in the target text data can be traversed and matched with a sentiment lexicon to obtain basic sentiment intensity and polarity. This is then weighted and adjusted using adverbs of degree and negation from the context to calculate the quantified sentiment score for the current round. The historical sentiment scores of the most recent N rounds of the current conversation are also obtained, and these scores are appended chronologically to form a temporal sequence of sentiment values. Based on this temporal sequence, sentiment change features such as the recent sentiment intensity gradient (reflecting instantaneous changes), the slope of the sentiment evolution trend based on linear regression (reflecting continuous deterioration or improvement), sentiment volatility (reflecting the degree of sentiment stability), and the temporal correlation between sentiment abrupt change points and labeled risk semantic rounds can be calculated. These sentiment change features are then packaged into structured sentiment change data. The TF-IDF (Term Frequency-Inverse Document Frequency) mentioned above is a statistical method used to evaluate the importance of a word to a specific document or document within a corpus. It combines term frequency (TF) and inverse document frequency (IDF) for weighting. The core idea of ​​TF-IDF is: if a word appears frequently in a document (high TF) but rarely in the entire corpus (high IDF), then that word is considered to have strong category discrimination ability and is suitable for representing document content. The aforementioned contextual information can be all information related to the current conversation, including historical conversation content, user personal information, and so on.

[0076] The aforementioned contextual semantic features can be obtained by extracting contextual semantics from the target text data based on contextual information. Contextual semantic features include the static meaning of words, the specific semantics of words in a particular dialogue context, their grammatical roles, and their relationships with other words.

[0077] A sliding window mechanism can be used to dynamically analyze the local context of text. This mechanism can avoid network congestion by dynamically adjusting the amount of data transmitted between the sender and receiver. The core of the sliding window mechanism lies in maintaining a sending window and a receiving window, which represent the range of data sequence numbers allowed to be sent or received. Through collaborative management, it implements acknowledgment, error control, and flow control functions. Specifically, based on a preset sliding window size parameter, the system can retrieve the most recent N rounds of historical data related to the current dialogue from the cache management library as contextual information and analyze the contextual semantic features of the target text data. More specifically, it can calculate the cosine similarity or attention-based relevance score between the semantic features of the target text data and the semantic features of each round of historical data to identify the historical topics most relevant to the target text. For example, when the target text data is "I'm so scared," it is highly relevant to the previous threatening statement "There will be monsters." The aforementioned cache management library is a database used to store and manage dialogue data.

[0078] The aforementioned third semantic risk analysis process can be a process of performing semantic risk analysis on target text data based on sentiment change data and contextual semantic features.

[0079] The aforementioned third semantic risk analysis results can be obtained by performing third semantic risk analysis on the target text data based on sentiment change data and contextual semantic features.

[0080] The aforementioned third risk level label can be the risk level label corresponding to the third semantic risk analysis result.

[0081] It should be noted that the window size can be dynamically adjusted according to the text length to adapt to different context requirements.

[0082] Optionally, in the step of determining the target security response strategy library from the preset security response strategy library based on the semantic risk analysis results and the corresponding risk level labels, if the semantic risk analysis result is a risk hit, then multiple candidate security response strategies can be selected from the preset security response strategy library according to the risk level labels; and the target security response strategy can be determined from the multiple candidate security response strategies in conjunction with the cache management library.

[0083] In this embodiment of the invention, the aforementioned risk hit can be a determination state that confirms the target text data as having a risk.

[0084] The aforementioned pre-defined security response strategy library can be a configurable and updatable rule database that stores various predefined security intervention logics and their triggering conditions. Based on the input semantic risk analysis results and corresponding risk level labels, the pre-defined security response strategy library can output specific security response strategies through rule matching or priority calculation.

[0085] The aforementioned candidate security response strategies can be preliminary response strategies suitable for the current interaction context, selected from a pre-defined security response strategy library based on the risk level label when the semantic risk analysis result indicates a risk hit.

[0086] The aforementioned cache management library is a database used to store and manage dialogue data.

[0087] The aforementioned target security response strategy can be a security response strategy determined from multiple candidate security response strategy libraries in conjunction with a cache management library. The target security response strategy specifies the specific security response action, the execution parameters of the action, the priority, and instructions that coordinate with the current session context.

[0088] It should be noted that when the semantic risk analysis result indicates a risk hit, multiple candidate security response strategies can be selected from the preset security response strategy library based on the risk level label. Combined with the cache management library, the target security response strategy can be determined from the multiple candidate security response strategies, which helps to improve response efficiency and accuracy, making security intervention both precise and efficient.

[0089] It's important to note that the cache management library can implement context-based session caching and persistence. This library distinguishes between short-term context caching and long-term semantic memory, achieving efficient cache management through secure cache control and LRU (Least Recently Used) architecture, while ensuring data security and semantic accuracy. Essentially, the cache management library includes first-level, second-level, and third-level cache strategies related to the interaction context. The first-level cache strategy can be a local memory cache LRU strategy, the second-level cache strategy can be a Redis cache LRU strategy, and the third-level cache strategy can be persistent session storage. High-risk semantic contexts with a risk level greater than 3 are automatically filtered to avoid "memory pollution." The LRU (Least Recently Used) strategy is a common cache eviction policy. The core idea of ​​LRU is based on the principle of temporal locality: if data has been recently accessed, it is more likely to be accessed again in the future; conversely, data that has not been used for a long time is less likely to be used in the short term, so this type of data is prioritized for eviction when space is insufficient. Optionally, after controlling the smart toy to execute the target safety response strategy library, the current session data can also be recorded; the current session data, semantic risk analysis results, and corresponding risk level labels can be generated into log data and reported to the cloud.

[0090] In this embodiment of the invention, the above-mentioned recording of current session data may be a process of summarizing, encapsulating and writing all key data and event sequences generated during the current interaction into a cache management library after the current complete voice interaction has ended.

[0091] Furthermore, the current session data, semantic risk analysis results, and corresponding risk level labels can be used to generate log data and report it to the cloud. This log data can be a structured data packet generated based on the current session data, semantic risk analysis results, and corresponding risk level labels.

[0092] The log data mentioned above can be understood as a collection of text or structured information generated in chronological order by interactive data during the operation of a smart toy. Log data is an abstract record of operations, state changes, or events, and typically includes key elements such as timestamps, operation descriptions, and source identifiers.

[0093] The aforementioned cloud platform could be a security analysis and management platform deployed in a remote data center, collaborating with smart toys via a network.

[0094] It should be noted that log data generated from the current session data, semantic risk analysis results, and corresponding risk level labels is uploaded to the cloud for subsequent security monitoring and management.

[0095] like Figure 2 As shown, this embodiment of the invention provides an intelligent toy control device based on semantic risk detection, which includes: The first processing module 201 is used to acquire the voice data of the target user and perform voice-to-text processing on the voice data to obtain the corresponding target text data. The second processing module 202 is used to perform semantic risk analysis on the target text data through a preset semantic risk detection model, and output the semantic risk analysis results and corresponding risk level labels of the target text data. The control module 203 is used to determine the target safety response strategy from the preset safety response strategy library based on the semantic risk analysis results and the corresponding risk level labels, and to control the smart toy to execute the target safety response strategy.

[0096] Optionally, the first processing module 201 is further configured to perform speech-to-text processing on the speech data using a preset speech recognition model to obtain initial text data; and to perform standardization processing on the initial text data to obtain target text data.

[0097] Optionally, the second processing module 202 is further configured to perform a first semantic risk analysis on the target text data using a preset semantic risk detection model to obtain a first semantic risk analysis result and a corresponding first risk level label; when the first semantic risk analysis result is a risk miss, perform a second semantic risk analysis on the target text data to obtain a second semantic risk analysis result and a corresponding second risk level label; when the second semantic risk analysis result is a risk miss, combine contextual information and a sentiment dictionary to perform a third semantic risk analysis on the target text data to obtain a third semantic risk analysis result and a corresponding third risk level label.

[0098] Optionally, the second processing module 202 is further configured to perform semantic extraction processing on the target text data to obtain the text semantic features corresponding to the target text data; and to perform a second semantic risk analysis on the text semantic features to obtain the second semantic risk analysis result and the corresponding second risk level label.

[0099] Optionally, the second processing module 202 is further configured to combine an emotion dictionary to calculate the emotion change data of the target text data; combine context information to analyze the contextual semantic features of the target text data; and based on the emotion change data and the contextual semantic features, perform third semantic risk analysis processing on the target text data to obtain the third semantic risk analysis result and the corresponding third risk level label.

[0100] Optionally, the control module 203 is further configured to, when the semantic risk analysis result is a risk hit, select multiple candidate security response strategies from a preset security response strategy library based on the risk level label; and, in conjunction with a cache management library, determine the target security response strategy from among the multiple candidate security response strategies.

[0101] Optionally, the device is also used to record current session data; and to generate log data by uploading the current session data, the semantic risk analysis results, and the corresponding risk level labels to the cloud.

[0102] like Figure 3 As shown, embodiments of the present invention also provide an electronic device, including a processor, which can execute any of the above-described intelligent toy control methods based on semantic risk detection.

[0103] Specifically, it includes a processor 301 and a memory 302, as well as a computer program stored in the memory 302 and capable of running on the processor 301, which executes a smart toy control method based on semantic risk detection, wherein: The processor 301 executes the calculator program based on the semantic risk detection intelligent toy control method stored in the memory 302, and performs the following steps: Acquire the voice data of the target user, and perform speech-to-text processing on the voice data to obtain the corresponding target text data; The target text data is subjected to semantic risk analysis using a preset semantic risk detection model, and the semantic risk analysis results and corresponding risk level labels are output. Based on the semantic risk analysis results and the corresponding risk level labels, a target safety response strategy is determined from the preset safety response strategy library, and the smart toy is controlled to execute the target safety response strategy.

[0104] Optionally, the processor 301 performs speech-to-text processing on the speech data to obtain corresponding text data, including: The speech data is processed into text using a preset speech recognition model to obtain initial text data. The initial text data is standardized to obtain the target text data.

[0105] Optionally, the step of processor 301 performing semantic risk analysis on the target text data using a preset semantic risk detection model and outputting the semantic risk analysis result and corresponding risk level label for the target text data includes: The target text data is subjected to a first semantic risk analysis using a preset semantic risk detection model to obtain the first semantic risk analysis result and the corresponding first risk level label. When the first semantic risk analysis result is a risk miss, a second semantic risk analysis is performed on the target text data to obtain the second semantic risk analysis result and the corresponding second risk level label. When the second semantic risk analysis result is a risk miss, the target text data is subjected to a third semantic risk analysis in combination with contextual information and a sentiment dictionary to obtain the third semantic risk analysis result and the corresponding third risk level label.

[0106] Optionally, the processor 301 performs a second semantic risk analysis on the target text data to obtain a second semantic risk analysis result and a corresponding second risk level label, including: The target text data is subjected to semantic extraction processing to obtain the text semantic features corresponding to the target text data; A second semantic risk analysis is performed on the semantic features of the text to obtain the second semantic risk analysis results and the corresponding second risk level labels.

[0107] Optionally, the processor 301 performs third semantic risk analysis on the target text data by combining contextual information and a sentiment dictionary to obtain the third semantic risk analysis result and the corresponding third risk level label, including: By combining the emotion dictionary, the emotion change data of the target text data is calculated; Analyze the contextual semantic features of the target text data by combining contextual information; Based on the emotion change data and the contextual semantic features, the target text data is subjected to third semantic risk analysis to obtain the third semantic risk analysis results and the corresponding third risk level label.

[0108] Optionally, the processor 301 executes the step of determining a target security response strategy from a preset security response strategy library based on the semantic risk analysis results and the corresponding risk level labels, including: When the semantic risk analysis result is a risk hit, multiple candidate security response strategies are selected from the preset security response strategy library according to the risk level label. By combining the cache management library, the target security response strategy is determined from multiple candidate security response strategies.

[0109] Optionally, after the intelligent toy is controlled to execute the target safety response strategy, the method executed by the processor 301 further includes: Record current session data; The current session data, the semantic risk analysis results, and the corresponding risk level labels are used to generate log data and report it to the cloud.

[0110] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the intelligent toy control method based on semantic risk detection provided in this invention and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0111] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for controlling intelligent toys based on semantic risk detection, characterized in that, The method includes the following steps: Acquire the voice data of the target user, and perform speech-to-text processing on the voice data to obtain the corresponding target text data; The target text data is subjected to semantic risk analysis using a preset semantic risk detection model, and the semantic risk analysis results and corresponding risk level labels are output. Based on the semantic risk analysis results and the corresponding risk level labels, a target safety response strategy is determined from the preset safety response strategy library, and the smart toy is controlled to execute the target safety response strategy.

2. The intelligent toy control method based on semantic risk detection as described in claim 1, characterized in that, The process of converting the speech data to text to obtain the corresponding text data includes: The speech data is processed into text using a preset speech recognition model to obtain initial text data. The initial text data is standardized to obtain the target text data.

3. The intelligent toy control method based on semantic risk detection as described in claim 1, characterized in that, The step of performing semantic risk analysis on the target text data using a preset semantic risk detection model, and outputting the semantic risk analysis results and corresponding risk level labels for the target text data, includes: The target text data is subjected to a first semantic risk analysis using a preset semantic risk detection model to obtain the first semantic risk analysis result and the corresponding first risk level label. When the first semantic risk analysis result is a risk miss, a second semantic risk analysis is performed on the target text data to obtain the second semantic risk analysis result and the corresponding second risk level label. When the second semantic risk analysis result is a risk miss, the target text data is subjected to a third semantic risk analysis in combination with contextual information and a sentiment dictionary to obtain the third semantic risk analysis result and the corresponding third risk level label.

4. The intelligent toy control method based on semantic risk detection as described in claim 3, characterized in that, The second semantic risk analysis of the target text data, to obtain the second semantic risk analysis results and the corresponding second risk level label, includes: The target text data is subjected to semantic extraction processing to obtain the text semantic features corresponding to the target text data; A second semantic risk analysis is performed on the semantic features of the text to obtain the second semantic risk analysis results and the corresponding second risk level labels.

5. The intelligent toy control method based on semantic risk detection as described in claim 3, characterized in that, The third semantic risk analysis is performed on the target text data by combining contextual information and a sentiment dictionary, resulting in the third semantic risk analysis results and corresponding third risk level labels, including: By combining the emotion dictionary, the emotion change data of the target text data is calculated; Analyze the contextual semantic features of the target text data by combining contextual information; Based on the emotion change data and the contextual semantic features, the target text data is subjected to third semantic risk analysis to obtain the third semantic risk analysis results and the corresponding third risk level label.

6. The intelligent toy control method based on semantic risk detection as described in claim 1, characterized in that, Based on the semantic risk analysis results and the corresponding risk level labels, the target security response strategy is determined from the preset security response strategy library, including: When the semantic risk analysis result is a risk hit, multiple candidate security response strategies are selected from the preset security response strategy library according to the risk level label. By combining the cache management library, the target security response strategy is determined from multiple candidate security response strategies.

7. The intelligent toy control method based on semantic risk detection as described in any one of claims 1-6, characterized in that, After the intelligent toy is controlled to execute the target safety response strategy, the method further includes: Record current session data; The current session data, the semantic risk analysis results, and the corresponding risk level labels are used to generate log data and report it to the cloud.

8. A smart toy control device based on semantic risk detection, characterized in that, The intelligent toy control device based on semantic risk detection includes: The first processing module is used to acquire the voice data of the target user and perform voice-to-text processing on the voice data to obtain the corresponding target text data. The second processing module is used to perform semantic risk analysis on the target text data through a preset semantic risk detection model, and output the semantic risk analysis results and corresponding risk level labels of the target text data. The control module is used to determine the target safety response strategy from the preset safety response strategy library based on the semantic risk analysis results and the corresponding risk level labels, and to control the smart toy to execute the target safety response strategy.

9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the intelligent toy control method based on semantic risk detection as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the intelligent toy control method based on semantic risk detection as described in any one of claims 1 to 7.