Abnormal speech early warning method and device based on artificial intelligence large model
Through multi-source data processing and comprehensive evaluation methods based on artificial intelligence large models, the existing risk prediction models are solved in the insufficient accurate warning of complex social phenomena, effective integration of multiple data formats and risk identification, and the comprehensiveness and accuracy of risk assessment are improved.
Patent Information
- Application Number
- CN202510299853.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-18
AI Technical Summary
When faced with complex and changing social phenomena, existing risk prediction models are difficult to achieve accurate and effective early warnings, especially lacking in-depth research and technical support for individuals' possible social tendency to retaliate. Most models only focus on single-dimensional data processing, ignoring the influence of factors such as emotional state, behavioral patterns and social relationships.
Anomaly speech warning method based on artificial intelligence large model is adopted, and multi-source correlation data is obtained, standardized processing is performed and converted into text vectors. A multi-modal processing model is used for emotion type, emotion intensity, behavior pattern and keyword recognition. Combined with a multi-task learning framework and a large language model, emotion type score, emotion intensity score, behavior pattern score and keyword recognition score are comprehensively evaluated, and finally, based on the preset weight coefficient, the user risk level is accumulated and judged.
It realizes effective integration and analysis of multiple data formats, improves the comprehensiveness and accuracy of risk identification, enhances the ability to analyze social relationships, better responds to diversified and complex real challenges, and provides personalized risk assessment support.
Smart Images

Figure CN120336910A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text information extraction, and particularly to an abnormal speech early warning method and device based on an artificial intelligence large model. Background Art
[0002] In modern society, risk prediction and management have become crucial links in enterprises and even the public security field. Traditional risk assessment methods mostly rely on fixed rule sets and simple statistical analysis of historical data. However, with the development of information technology and the change of social environment, this risk assessment method based on static rules gradually exposes its limitations. Especially when facing complex and changeable social phenomena, such as individual extreme behaviors, online public opinions, etc., traditional methods are difficult to achieve accurate and effective early warning.
[0003] In recent years, with the progress of machine learning algorithms and their wide application in various industries, many new attempts have emerged to use artificial intelligence technology for risk prediction. For example, in the financial industry, a model has been proposed to integrate multiple machine learning algorithms to predict the onset risk of asthma disease; while in the banking industry, the patented technology of "multi-dimensional fraud risk prediction" has been developed, which combines three different machine learning models to predict the login, account opening, and transaction fraud risks of users respectively. These innovative solutions have significantly improved the speed and accuracy of risk identification, and can provide personalized services and support on a larger scale.
[0004] Nevertheless, most of the current existing technologies still focus on risk prevention and control within specific fields. For risks at the social level with a wider scope, especially for problems such as the possible tendency of individuals to retaliate against society, there is a lack of sufficient in-depth research and technical support. In addition, most of the existing risk prediction models focus on single-dimensional data processing (such as only limited to text content), while ignoring other potential important factors, such as emotional state, behavior pattern, and social relationship, etc., which have an impact on the formation of risks. Summary of the Invention
[0005] The purpose of the present invention is to at least solve one of the deficiencies of the prior art, and provide an abnormal speech early warning method and device based on an artificial intelligence large model.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] Specifically, an abnormal speech early warning method based on an artificial intelligence large model is proposed, including the following:
[0008] Obtain multi-source associated data of a target user;
[0009] After standardizing the multi-source associated data, convert it into a text vector based on a multi-modal processing model;
[0010] The large model is used to classify the text vectors respectively based on the sentiment type to obtain sentiment type scores, classify based on the sentiment intensity to obtain sentiment intensity scores, identify the behavior pattern to obtain behavior pattern scores, and identify keywords to obtain keyword identification scores;
[0011] The sentiment type scores, sentiment intensity scores, behavior pattern scores, and keyword identification scores are accumulated according to preset weight coefficients to obtain the final score;
[0012] Judge whether the final score is higher than the preset threshold. If so, determine that the target user is a high-risk user of abnormal speech. If not, determine that the target user is a low-risk user of abnormal speech;
[0013] Output the determination result.
[0014] Furthermore, specifically, the multi-source associated data includes text data as well as data in the formats of voice and image.
[0015] Furthermore, specifically, the large model is used to classify the text vectors based on the sentiment type to obtain sentiment type scores, and classify based on the sentiment intensity to obtain sentiment intensity scores, including,
[0016] Based on the multi-task learning framework, using BERT as the shared feature extractor, combined with the multi-layer perceptron MLP to jointly learn the sentiment type classification and sentiment intensity classification regression tasks; First, the BERT model encodes the input text vectors to extract deep semantic features, which are then processed by the MLP network; The sentiment classification task passes through a branch of the MLP to predict the sentiment type in the text to obtain the sentiment type score. At the same time, the sentiment intensity prediction task passes through another branch of the MLP to output the intensity value of the sentiment to obtain the sentiment intensity score, and this value represents the intensity of the sentiment.
[0017] Furthermore, specifically, the large model is used to identify the behavior pattern of the text vectors to obtain behavior pattern scores, including,
[0018] Establish and maintain a series of behavior pattern dictionaries that may constitute risks; Through precisely designed Prompt instructions, input the sentiment analysis results and the constructed behavior pattern dictionaries into the large language model LLM to predict the specific behavior pattern labels in the text, and then obtain the behavior pattern scores.
[0019] Furthermore, specifically, the large model is used to identify keywords from the text vectors to obtain keyword identification scores, including,
[0020] A preset keyword set is provided, which allows users to customize the keyword list according to specific application scenarios. Based on the preset keyword set, combined with a large language model (LLM), the text vector is analyzed to output the keyword recognition score.
[0021] Furthermore, specifically, the sentiment type score, sentiment intensity score, behavior pattern score, and keyword recognition score are accumulated according to preset weight coefficients to obtain the final score, including:
[0022] If the preset weights of the sentiment type score, sentiment intensity score, behavior pattern score, and keyword recognition are w1, w2, w3, and w4 respectively, then the final score = w1 * sentiment type score + w2 * sentiment intensity score + w3 * behavior pattern score + w4 * keyword score.
[0023] Furthermore, specifically, the weight coefficients w1 to w4 are automatically adjusted and optimized during the training of the model to ensure the best overall performance.
[0024] Furthermore, specifically, when outputting the judgment result, sub-reports are generated respectively for the specific processes of obtaining the sentiment type score based on sentiment type classification, the sentiment intensity score based on sentiment intensity classification, the behavior pattern score based on behavior pattern recognition, and the keyword recognition score based on keyword recognition, which is convenient for subsequent analysis.
[0025] The present invention also proposes a device for early warning of abnormal remarks based on an artificial intelligence large model, including the following:
[0026] A data acquisition module, which is used to acquire multi-source associated data of the target user;
[0027] A multi-modal processing module, which is used to perform standardized processing on the multi-source associated data and then convert it into a text vector based on a multi-modal processing model;
[0028] A score evaluation module, which is used to obtain the sentiment type score based on sentiment type classification, the sentiment intensity score based on sentiment intensity classification, the behavior pattern score based on behavior pattern recognition, and the keyword recognition score based on keyword recognition for the text vector through the large model;
[0029] A final score calculation module, which is used to accumulate the sentiment type score, sentiment intensity score, behavior pattern score, and keyword recognition score according to preset weight coefficients to obtain the final score;
[0030] A result judgment module, which is used to judge whether the final score is higher than a preset threshold. If so, it is determined that the target user is a high-risk user of abnormal remarks; if not, it is determined that the target user is a low-risk user of abnormal remarks;
[0031] A result output module, which is used to output the judgment result.
[0032] The beneficial effects of the present invention are as follows:
[0033] The present invention proposes an abnormal speech warning method and device based on an artificial intelligence large model, which conducts multi-dimensional comprehensive judgment through the multi-source associated data of the target user and effectively integrates multi-source heterogeneous data. Different from the risk prediction models in the past that could only process single-type data, the present invention realizes the effective integration and analysis of various formats of data such as text, voice, and even pictures by constructing a multi-modal processing framework, thereby improving the comprehensiveness and accuracy of risk identification. At the same time, considering that social isolation may be an important factor leading to extreme behaviors, a social relationship analysis module is added to further enhance the risk assessment ability of the model. In addition, by introducing a customized keyword evaluation mechanism, the model can better cope with diverse and complex real-world challenges. The present invention proposes an innovative solution to the limitations of existing risk prediction models by integrating advanced artificial intelligence technologies such as a multi-modal processing framework, comprehensive judgment of emotional and behavioral patterns, and context and historical semantic analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] By describing the embodiments shown in the accompanying drawings in detail, the above and other features of the present disclosure will become more obvious. The same reference numerals in the drawings of the present disclosure denote the same or similar elements. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0035] Figure 1 Shown is a flowchart of the abnormal speech warning method based on the artificial intelligence large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The following will clearly and completely describe the concept, specific structure, and technical effects generated by the present invention in combination with the embodiments and the drawings, so as to fully understand the purpose, solution, and effects of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The same reference numerals used throughout the drawings denote the same or similar parts.
[0037] Embodiment 1, referring to Figure 1 , the present invention proposes an abnormal speech warning method based on the artificial intelligence large model, including the following:
[0038] Obtain multi-source associated data of the target user;
[0039] After performing standardization processing on the multi-source associated data, convert it into a text vector based on a multi-modal processing model;
[0040] Based on the large model, sentiment type scores are obtained by classifying the text vectors according to sentiment types, sentiment intensity scores are obtained by classifying them according to sentiment intensity, behavior pattern scores are obtained by behavior pattern recognition, and keyword recognition scores are obtained by keyword recognition;
[0041] The sentiment type scores, sentiment intensity scores, behavior pattern scores, and keyword recognition scores are accumulated according to preset weight coefficients to obtain the final score;
[0042] Judge whether the final score is higher than the preset threshold. If so, determine that the target user is a high-risk user of abnormal speech. If not, determine that the target user is a low-risk user of abnormal speech;
[0043] Output the determination result.
[0044] In this preferred embodiment, multiple subsystems are integrated, including but not limited to generative language large models such as Llama3, as well as BERT models designed specifically for sentiment classification and multimodal models that process various formats, to construct a framework that can process multimodal inputs (text, speech, images) and uniformly convert them into text vector form for subsequent analysis. By identifying the sentiment type and its intensity level expressed by the user, combined with the analysis of the behavioral characteristics described in the text, and coupled with the assessment of the individual's social activity and the degree of social connection, the present invention can more comprehensively and accurately determine whether there is a potential tendency of social retaliation.
[0045] Specifically, the operation process of the entire system is as follows:
[0046] 1. Data preprocessing: After receiving data from different channels, first perform standardization processing on it to ensure that only data meeting specific standards can enter the next analysis process.
[0047] 2. Sentiment analysis and semantic understanding: Use a pre-trained BERT model for sentiment classification, and at the same time use a powerful large language model to perform semantic understanding tasks to detect whether there are signs of threatening language or potential violent behavior.
[0048] 3. Behavior pattern analysis: Define a series of behavior patterns that may pose risks, such as "threat", "jump off a building", "kill", etc. By identifying these behavioral characteristics and combining the sentiment analysis results, evaluate whether an individual shows potential extreme behavior tendencies.
[0049] 4. Keyword evaluation: Introduce keyword analysis to assist the model in analyzing risk situations, allowing users to provide keywords to enable the model to have stronger processing capabilities, and improving the model's processing and analysis capabilities for diverse and complex situations in different fields and regions through keywords.
[0050] 5. Calculate the final result: Based on the scores given by each sub-module, summarize them into a final score according to a predetermined weight formula, and use this as the basis for determining the possibility of high risk.
[0051] As a preferred embodiment of the present invention, specifically, the multi-source associated data includes text data as well as data in voice and image formats.
[0052] Specifically, in this preferred embodiment, the acquisition and processing of multi-source associated data includes the following steps.
[0053] S11. Receive multi-source input: The system receives data in formats such as text, voice, or pictures from different channels.
[0054] S12. Standardization processing: Perform standardization processing on all received data. For example, adjust the text length to a specific range, remove irrelevant information, and ensure data quality and consistency.
[0055] S13. Convert to a unified text vector: Use a multi-modal processing model to convert non-text data (such as voice and images) into text vector form for subsequent analysis.
[0056] As a preferred embodiment of the present invention, specifically, through a large model, the text vector is classified based on emotion types to obtain emotion type scores, and classified based on emotion intensity to obtain emotion intensity scores, including:
[0057] Based on a multi-task learning framework, use BERT as a shared feature extractor, and combine a multi-layer perceptron MLP to jointly learn the emotion type classification and emotion intensity classification regression tasks; First, the BERT model encodes the input text vector to extract deep semantic features, which are then processed by the MLP network; The emotion classification task passes through a branch of the MLP to predict the emotion type in the text to obtain the emotion type score. At the same time, the emotion intensity prediction task passes through another branch of the MLP to output the intensity value of the emotion to obtain the emotion intensity score, and this value represents the intensity of the emotion.
[0058] In this preferred embodiment, the sentiment classification and semantic understanding algorithm adopts a multi-task learning framework, uses BERT as a shared feature extractor, and combines a multi-layer perceptron (MLP) to jointly learn the sentiment type classification and sentiment intensity regression tasks. First, the BERT model encodes the input text to extract deep semantic features, which are then processed through the MLP network. The sentiment classification task, through a branch of the MLP, predicts the sentiment type in the text (such as anger, joy, sadness, etc.). At the same time, the sentiment intensity prediction task, through another branch of the MLP, outputs the intensity value of the sentiment, which represents the degree of intensity of the sentiment (such as slight, medium, or strong). Through this multi-task learning method, the model simultaneously optimizes two objectives in the same training process: one is the accurate prediction of the sentiment type, and the other is the precise regression of the sentiment intensity. Finally, the model can generate predictions that not only conform to the sentiment categories but also accurately reflect the intensity changes of the sentiment, thereby improving the comprehensiveness and accuracy of text analysis.
[0059] In the tasks of sentiment classification and intensity prediction, first, BERT encodes the text T to obtain the deep representation B(T) of the text. Then, the sentiment classification C, that is, the sentiment type score, and the sentiment intensity regression T, that is, the sentiment intensity score, are calculated through two different MLP scores respectively.
[0060] Among them,
[0061] 1. Sentiment type prediction formula:
[0062] C = MLP emotion_type (B(T)),
[0063] 2. Sentiment intensity prediction formula
[0064] I = MLP emotion_intensity (B(T)),
[0065] Finally, the multi-task learning objective function is used to optimize the model, where the loss function is the weighted sum of the sentiment classification and regression tasks.
[0066]
[0067] Among them, C true is the theoretical value of C, and I true is the theoretical value of I.
[0068] As a preferred embodiment of the present invention, specifically, the behavior pattern score is obtained by the large model based on the behavior pattern recognition of the text vector, including
[0069] Build and maintain a dictionary of a series of behavioral patterns that may pose risks; through precisely designed Prompt instructions, input the sentiment analysis results and the constructed dictionary of behavioral patterns into the large language model (LLM) to predict the specific behavioral pattern tags in the text, and then obtain the behavioral pattern scores.
[0070] In this preferred embodiment, we first build and maintain a dictionary of a series of behavioral patterns that may pose risks, including behavioral patterns such as threats, jumping off buildings, and stabbing to death. These behavioral patterns usually have obvious emotional expressions and may be accompanied by specific emotion tags and contexts. By maintaining these dictionaries, we can accurately reflect the danger signals in the text when identifying potential risks in the text. Through precisely designed Prompt instructions, we input the sentiment analysis results (such as anger, anxiety, etc.) and the constructed dictionary of risk behavioral patterns into the large language model (LLM), and use the powerful semantic understanding ability of the LLM to predict the specific behavioral pattern tags in the text. For example, the Prompt instruction template: "Please identify and output the behavioral pattern tags described in the following text content. The behavioral pattern tags include but are not limited to: threats, violence, suicide, escape, indifference, etc. Please judge whether these behavioral patterns appear according to the specific description in the text and list the corresponding tags."
[0071] As a preferred embodiment of the present invention, specifically, the large model obtains a keyword recognition score based on keyword recognition of the text vector, including,
[0072] Preset a keyword set, which allows users to customize the keyword list according to specific application scenarios. Based on the preset keyword set, combined with the large language model (LLM), analyze the text vector to output the keyword recognition score.
[0073] In this preferred embodiment, it mainly includes the following two steps,
[0074] S41. Customize the keyword set: Allow users to customize the keyword list according to specific application scenarios to increase the pertinence and flexibility of the model.
[0075] S42. Dynamic update mechanism: As new words appear, the system can automatically learn and incorporate new risk identifiers to improve the model's adaptability to complex situations.
[0076] As a preferred embodiment of the present invention, specifically, accumulate the sentiment type score, sentiment intensity score, behavioral pattern score, and keyword recognition score according to the preset weight coefficients to obtain the final score, including,
[0077] If the weight preset values for the emotion type score, emotion intensity score, behavior pattern score, and keyword recognition are w1, w2, w3, and w4 respectively, then the final score = w1 * emotion type score + w2 * emotion intensity score + w3 * behavior pattern score + w4 * keyword score.
[0078] As a preferred embodiment of the present invention, specifically, the weight coefficients w1 to w4 are automatically adjusted and optimized by the model during training to ensure the best overall performance.
[0079] In this preferred embodiment, it mainly includes the following steps:
[0080] S51. Sub-module scoring: Each sub-module (emotion type, emotion intensity, behavior pattern, keyword recognition) gives corresponding scores respectively.
[0081] S52. Application of weighted summation formula: The scores of each sub-module are summarized according to the predetermined weight formula, and the final score = w1 * emotion type score + w2 * emotion intensity score + w3 * behavior pattern score + w4 * keyword score, where the weight coefficients w1 to w4 are automatically adjusted and optimized by the model during training to ensure the best overall performance.
[0082] S53. Threshold determination: Set a threshold to determine the final classification result (high risk / low risk), and present it to the user by outputting a classification label.
[0083] As a preferred embodiment of the present invention, specifically, when outputting the determination result, sub-reports are respectively generated for the specific processes of obtaining the emotion type score based on emotion type classification, obtaining the emotion intensity score based on emotion intensity classification, obtaining the behavior pattern score based on behavior pattern recognition, and obtaining the keyword recognition score based on keyword recognition, which is convenient for subsequent analysis.
[0084] In this preferred embodiment, it mainly includes the following steps:
[0085] S61. Binary classifier output: Directly give the possibility of the existence of high risk.
[0086] S62. Detailed scoring report: Provide additional information to assist the decision-maker to make a more informed choice, such as the influence degree of each factor, etc.
[0087] Specifically, when the overall optimization scheme proposed by the present invention is applied, it includes the following contents:
[0088] System architecture and working principle
[0089] This risk prediction model integrates multiple artificial intelligence subsystems, including but not limited to generative language large models (such as Llama3), multimodal processing models, and the BERT model specifically designed for sentiment classification. These components together constitute a multimodal risk processing and analysis framework, which improves the model's ability to understand complex scenarios by integrating multiple data sources. It can receive and parse data from different channels - text, voice, or pictures, and finally convert them into a unified text vector form for subsequent analysis.
[0090] Multimodal Data Processing
[0091] The risk prediction model can process various types of data, including text, voice, pictures, and videos. To adapt to these different data input forms, the model calls specially designed subsystems for preprocessing. For example, for picture-type data input, the BLIP model is called for processing to convert the picture into text, and there are corresponding models for other types. Specifically, for video type, it can be converted into pictures by extracting key frames and analyzing inter-frame motion and then input into the corresponding model for processing. Through the multimodal module, we can convert various types of data into text data for easy analysis, preparing for data preprocessing, sentiment classification, and semantic understanding.
[0092] Data Preprocessing and Input Types
[0093] In practical applications, the system will first preprocess the data received in various formats to ensure that only data meeting specific criteria can enter the next analysis process. For example, the text length needs to meet certain specifications, and at the same time, those content that clearly does not belong to the risk category is removed. The preprocessed data will be sent into the main model in the form of text vectors for further processing, ensuring the consistency and quality of the input data and reducing noise interference.
[0094] The following is a data example:
[0095] {
[0096] "Emotion Type": "Angry",
[0097] "Emotion Intensity": "Extreme",
[0098] "Large Model Score": "3",
[0099] "Original Text": "Zhang and Wang had a conflict due to a quarrel, and there was physical conflict between the two sides. Zhang's emotions were unstable. Not only did he not cooperate with the mediation, but he also claimed to drive and hit Wang and the relevant mediators to death.",
[0100] "Final Result": "At Risk"
[0101] }
[0102] For the convenience of understanding, the specific scores and results in the above examples are replaced with Chinese. In actual training and testing, they are all reflected by numerical codes.
[0103] Output result format
[0104] The results output by the system are presented in the form of classification labels, that is, whether there is a potential social retaliation tendency. This decision is based on a series of scores obtained from internal calculations, and the final classification result is determined by setting a threshold. In addition, a detailed scoring report can also be output to help users understand the influence degree of each factor.
[0105] Sentiment analysis and semantic understanding
[0106] In order to more accurately capture the emotional state expressed by users, a pre-trained BERT model is used for sentiment classification to identify the emotional types (such as anger, hatred, sadness, etc.) and their intensity levels (slight, ordinary, strong) in the text. By quantifying the user's emotional level, it helps to discover potential risk signals. In addition, a powerful large language model is used to perform semantic understanding tasks to deeply understand the meaning behind the text to detect whether there are signs of threatening language or potential violent behavior.
[0107] Behavior pattern analysis
[0108] For the behavior descriptions in the text content, a series of behavior patterns that may constitute risks are defined, such as "threat", "jumping off a building", "killing", etc. By identifying these behavior characteristics and combining the sentiment analysis results, it is possible to more accurately evaluate whether an individual shows potential extreme behavior tendencies. Generally speaking, if a negative emotion, a high-intensity emotional expression, and specific risk behaviors appear in the text of the party, it can indicate that there is a relatively high risk possibility for the party in the text.
[0109] Keyword evaluation
[0110] Keyword analysis is introduced, allowing users to adjust the keyword list according to their needs to increase pertinence and assist the model in analyzing the risk situation. Users can customize the keyword set according to specific application scenarios, thereby enhancing the adaptability of the model to complex situations in specific fields or regions. This step not only improves the flexibility of the system but also enables it to better cope with diverse real-world challenges.
[0111] Risk scoring mechanism
[0112] The risk prediction process of the entire system relies on a complex scoring mechanism. Each sub-module gives corresponding scores, which are then aggregated into a final score according to a predetermined weight formula. Specifically, the final score = w1 * sentiment type score + w2 * sentiment intensity score + w3 * large model behavior pattern score + w4 * large model keyword recognition score. Among them, the model will adjust the weight coefficients w1 to w4 by itself during training to ensure the best overall performance.
[0113] Embodiment 2, the present invention also proposes a device for early warning of abnormal remarks based on an artificial intelligence large model, including the following:
[0114] A data acquisition module for acquiring multi-source associated data of a target user;
[0115] A multi-modal processing module for performing normalization processing on the multi-source associated data and converting it into a text vector based on a multi-modal processing model;
[0116] A score evaluation module for obtaining a sentiment type score by classifying the text vector based on sentiment type through a large model, obtaining a sentiment intensity score by classifying based on sentiment intensity, obtaining a behavior pattern score by recognizing the behavior pattern, and obtaining a keyword recognition score by keyword recognition;
[0117] A final score calculation module for accumulating the sentiment type score, sentiment intensity score, behavior pattern score, and keyword recognition score according to a preset weight coefficient to obtain a final score;
[0118] A result judgment module for judging whether the final score is higher than a preset threshold. If so, it determines that the target user is a high-risk user of abnormal remarks; if not, it determines that the target user is a low-risk user of abnormal remarks;
[0119] A result output module for outputting the judgment result.
[0120] In summary, the technological progress and improvement effects achieved by the present invention are as follows:
[0121] Wider data processing scope: The present invention constructs a multi-modal risk assessment system capable of processing multi-source heterogeneous data (text, voice, and image). This comprehensive data processing method enables the system to more accurately capture potential social retaliation tendencies, providing a more extensive and in-depth risk insight compared to traditional models that can only process single-type data.
[0122] Deep sentiment and intention understanding: Use a pre-trained BERT model for sentiment classification and combine a powerful generative language large model to perform deep semantic understanding tasks. This not only enables the system to identify explicit danger signals but also detect implicit harmful information, improving the recognition accuracy of potential risks in complex contexts.
[0123] Personalization and flexibility: The present invention allows users to customize the keyword list according to specific application scenarios, enhancing the adaptability and pertinence of the system. This feature is particularly applicable to unique expressions in different regions or cultural backgrounds, ensuring the wide application ability of the model in diverse environments.
[0124] Enhanced risk identification ability: Considering the importance of context and historical data, a more refined semantic analysis strategy is adopted to reduce misjudgments caused by taking out of context. This method effectively avoids unnecessary panic and waste of social resources, while improving the accuracy of the system in dealing with ambiguous or polysemous expressions.
[0125] Efficient risk scoring mechanism: A reasonable weight formula is designed to aggregate the scores of each sub-module to obtain the final risk level. This scoring system ensures that the scoring results not only reflect the influence degree of each factor, but also take into account the overall consistency and stability, providing a solid basis for decision-making.
[0126] In addition, in each embodiment of the present invention, each functional module can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0127] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or system that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0128] Although the description of the present invention has been quite detailed and several of the described embodiments have been described in particular, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but rather should be regarded as providing a broad interpretation of these claims in light of the prior art by reference to the appended claims, so as to effectively cover the intended scope of the present invention. In addition, the present invention has been described above in terms of embodiments foreseeable by the inventors for the purpose of providing a useful description, and non-substantive modifications to the present invention that are not currently foreseeable may still represent equivalent modifications of the present invention.
[0129] As described above, these are only the preferred embodiments of the present invention. The present invention is not limited to the above-described embodiments. As long as the same means are used to achieve the technical effects of the present invention, they should fall within the protection scope of the present invention. Within the protection scope of the present invention, various different modifications and variations may be made to its technical solutions and / or embodiments.
Claims
1. An abnormal speech warning method based on an artificial intelligence large model, characterized in that, The following are included: Obtain multi-source associated data of the target user; After performing normalization processing on the multi-source associated data, convert it into a text vector based on a multimodal processing model; Through a large model, classify the text vector based on emotion types to obtain emotion type scores, classify based on emotion intensity to obtain emotion intensity scores, identify behavior patterns to obtain behavior pattern scores, and identify keywords to obtain keyword recognition scores; Accumulate the emotion type scores, emotion intensity scores, behavior pattern scores, and keyword recognition scores according to preset weight coefficients to obtain a final score; Judge whether the final score is higher than a preset threshold. If so, determine that the target user is a high-risk user of abnormal speech; if not, determine that the target user is a low-risk user of abnormal speech; Output the determination result.
2. The abnormal speech warning method based on the large artificial intelligence model according to claim 1, wherein Specifically, the multi-source associated data includes text data and data in the formats of voice and images.
3. The abnormal speech warning method based on the artificial intelligence large model according to claim 1, wherein Specifically, classifying the text vector based on emotion types through a large model to obtain emotion type scores, and classifying based on emotion intensity to obtain emotion intensity scores, includes Based on a multi-task learning framework, use BERT as a shared feature extractor, and combine a multi-layer perceptron MLP to jointly learn emotion type classification and emotion intensity classification regression tasks; First, the BERT model encodes the input text vector to extract deep semantic features, which are then processed by the MLP network; The emotion classification task passes through a branch of the MLP to predict the emotion type in the text to obtain the emotion type score. At the same time, the emotion intensity prediction task passes through another branch of the MLP to output the intensity value of the emotion to obtain the emotion intensity score, which represents the intensity of the emotion.
4. The abnormal speech warning method based on the large artificial intelligence model according to claim 1, characterized in that, Specifically, obtaining the behavior pattern score by identifying the behavior pattern of the text vector through a large model, includes Establish and maintain a series of behavior pattern dictionaries that may constitute risks; Through precisely designed Prompt instructions, input the emotion analysis results and the constructed behavior pattern dictionaries into a large language model LLM to predict the specific behavior pattern labels in the text, and then obtain the behavior pattern score.
5. The abnormal speech warning method based on the artificial intelligence large model according to claim 1, wherein, Specifically, obtaining the keyword recognition score by identifying keywords of the text vector through a large model, includes Preset a keyword set, which allows users to customize a keyword list according to specific application scenarios. Based on the preset keyword set, combine a large language model LLM to analyze the text vector and output the keyword recognition score.
6. The abnormal speech warning method based on the artificial intelligence large model according to claim 1, characterized in that Specifically, accumulating the emotion type scores, emotion intensity scores, behavior pattern scores, and keyword recognition scores according to preset weight coefficients to obtain a final score, includes If the weights of the emotion type score, emotion intensity score, behavior pattern score, and keyword recognition are preset as w1, w2, w3, w4, then the final score = w1 * emotion type score + w2 * emotion intensity score + w3 * behavior pattern score + w4 * keyword score.
7. The abnormal speech warning method based on the large artificial intelligence model according to claim 6, characterized in that, Specifically, the weight coefficients w1 to w4 are automatically adjusted and optimized during model training to ensure the best overall performance.
8. The abnormal speech warning method based on the artificial intelligence large model according to claim 1, characterized in that, Specifically, when outputting the judgment result, sub-reports are generated respectively for the specific processes of obtaining the sentiment type score based on sentiment type classification, the sentiment intensity score based on sentiment intensity classification, the behavior pattern score based on behavior pattern recognition, and the keyword recognition score based on keyword recognition, which is convenient for subsequent analysis.
9. An apparatus for early warning of abnormal speech based on an artificial intelligence large model, characterized in that, Including the following: A data acquisition module for acquiring multi-source associated data of the target user; A multi-modal processing module for performing normalization processing on the multi-source associated data and then converting it into a text vector based on a multi-modal processing model; A score evaluation module for obtaining the sentiment type score based on sentiment type classification, the sentiment intensity score based on sentiment intensity classification, the behavior pattern score based on behavior pattern recognition, and the keyword recognition score based on keyword recognition for the text vector through a large model; A final score calculation module for accumulating the sentiment type score, the sentiment intensity score, the behavior pattern score, and the keyword recognition score according to a preset weight coefficient to obtain a final score; A result judgment module for judging whether the final score is higher than a preset threshold. If so, the target user is determined to be a high-risk user of abnormal speech, and if not, the target user is determined to be a low-risk user of abnormal speech; A result output module for outputting the judgment result.