A risk identification method and device, electronic equipment and storage medium

CN122817907APending Publication Date: 2026-09-25SAMSUNG ELECTRONICS CHINA R&D CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611008579.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]现有技术由于主要依赖静态、片面的数据,其分析方法固定不变,反应迟缓,尤其对于隐性的风险识别效果欠佳

Benefits of technology

[0037]预测匹配模块,用于在第K+1时序时,将获得的第K+1时序场景特征词元与所述第K时序预测的风险行为向量集合中的向量进行匹配,以确定风险识别结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817907A_ABST
    Figure CN122817907A_ABST
Patent Text Reader

Abstract

The application discloses a risk identification method and device, electronic equipment and storage medium, comprising: intercepting a multi-modal original data stream in real time, extracting multi-modal data features and determining corresponding expert agents; the expert agent is preprocessed to obtain semantic evidence tokens of the current K time sequence and scene feature tokens of the current K time sequence; the context logic engine model performs semantic analysis and causal reconciliation according to the semantic evidence tokens of the current K time sequence, and generates a risk hypothesis vector of the current K time sequence; the generative prediction model uses a self-attention mechanism to infer and generate a risk behavior vector set predicted in the K time sequence; in the K+1 time sequence, the K+1 time sequence scene feature tokens are matched with the vectors in the risk behavior vector set predicted in the K time sequence to determine a risk identification result. According to the application, the hidden risks in the interaction can be flexibly, quickly and accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a risk identification method, a risk identification device, an electronic device, and a storage medium. Background Technology

[0002] With the development of communication technology, people are surrounded by a complex information environment every day. Users inevitably face risks and even fraud when communicating with all sorts of people. Existing risk prevention methods typically employ blacklisting, deep learning, or behavioral statistics to identify risks. Blacklists are mainly used for telephone calls, comparing the caller's phone number with the blacklist for identification. Deep learning primarily uses communication segments to calculate risk probabilities, employing methods such as Convolutional Neural Networks (CNNs) or Recurrent Neural Networks (RNNs). Behavioral statistics identify abnormal patterns by analyzing macroscopic behavioral characteristics during communication and then using statistical models for analysis.

[0003] Existing technologies rely primarily on static and partial data, employing fixed and unchanging analytical methods that are slow to react, particularly ineffective at identifying implicit risks. In reality, user interaction scenarios are becoming increasingly complex, and potential risks are becoming increasingly difficult to detect, making existing technologies increasingly inadequate for current risk identification needs. Summary of the Invention

[0004] In view of the prior art, this application discloses a risk identification method that can overcome the shortcomings of the existing analysis methods, such as being fixed and slow to react, and can flexibly and quickly identify hidden risks.

[0005] This application proposes a risk identification method, which includes:

[0006] Real-time interception of multimodal raw data streams, extraction of multimodal data features based on the multimodal raw data streams, and determination of the corresponding expert agent based on the context to which the multimodal raw data streams belong;

[0007] The expert agent performs intensive domain-specific feature refinement preprocessing on the multimodal data features to obtain semantic evidence lexical units and scene feature lexical units for the current Kth time sequence, where K is a natural number. The semantic evidence lexical units include key data features obtained from the intensive domain-specific feature refinement process and causal logical relationships. The scene feature lexical units include vector values ​​formed by converting the current interaction state into behavioral milestone snapshots.

[0008] The semantic evidence lexical units of the current Kth time sequence are input into the context logic engine model. The context logic engine model performs semantic parsing and causal reconciliation based on the semantic evidence lexical units of the current Kth time sequence to generate the risk hypothesis vector of the current Kth time sequence.

[0009] The risk hypothesis vector of the current Kth time series is used as a conditional control input to the generative prediction model, which then uses a self-attention mechanism to infer and generate a set of risk behavior vectors for the Kth time series prediction.

[0010] At time series K+1, the obtained scene feature lexical units of time series K+1 are matched with the vectors in the risk behavior vector set predicted at time series K to determine the risk identification result.

[0011] Furthermore,

[0012] The expert agent is a lightweight edge-side small language model (SLM).

[0013] The context logic engine model is a large language model (LLM) with dense logic.

[0014] The generative prediction model is a discrete time series generative prediction model based on Transformer.

[0015] The expert agent, the context logic engine model, and the generative prediction model constitute a three-level heterogeneous cascaded topology.

[0016] Furthermore, the multimodal raw data stream includes any combination of text data, audio data, view data, and system metadata;

[0017] The steps for extracting multimodal data features include: extracting any combination of text data features, video data features, view data features, and system metadata features.

[0018] Furthermore, the step of the expert agent performing intensive domain-specific feature refinement preprocessing on the multimodal data features to obtain the semantic evidence lexical units and scene feature lexical units of the current Kth time sequence includes:

[0019] The expert agent performs intensive domain-specific feature refinement preprocessing on each single-modal data feature in the multimodal data features to obtain the corresponding single-modal key data features, and the corresponding single-modal key data features have the causal logical relationship, thus determining the feature causal logical chain of the corresponding single-modal; aligning all the feature causal logical chains of the single-modal features and mapping them to semantic vectors, the semantic evidence lexical of the current Kth time sequence is obtained;

[0020] The current interaction state is converted into a behavior milestone snapshot, and the behavior milestone snapshot is mapped to a pre-existing behavior dictionary space to obtain vector values ​​as the scene feature words of the current Kth time sequence. The behavior dictionary space represents the total set of vector values ​​corresponding to all the behaviors that the user can operate on the device.

[0021] Furthermore, the step of the context logic engine model performing semantic parsing and causal reconciliation based on the semantic evidence lexical units of the current Kth time sequence to generate the risk hypothesis vector of the current Kth time sequence includes:

[0022] The context logic engine model performs multi-dimensional semantic parsing on the semantic evidence lexical units of the current Kth time sequence to obtain multi-dimensional semantic parsing results;

[0023] The context logic engine model performs chain reasoning on the multidimensional semantic parsing results based on the deep semantic architecture of chain thinking ability, and performs cross-dimensional causal reconciliation.

[0024] The context logic engine model performs causal consistency verification based on the results of the cross-dimensional causal reconciliation. If the causal consistency verification fails, a risk hypothesis vector for the current Kth time series is generated.

[0025] Furthermore, the step of generating the set of risk behavior vectors for the Kth time series prediction by the generative prediction model using a self-attention mechanism includes:

[0026] The generative prediction model uses the risk hypothesis vector of the current Kth time series as a conditional control and uses a self-attention mechanism to infer and generate risk behavior probabilities in the behavior dictionary space.

[0027] The vector values ​​based on the behavior dictionary space corresponding to the risk behavior probabilities exceeding the preset threshold are merged to form the risk behavior vector set for the Kth time series prediction.

[0028] Further, the step of matching the obtained (K+1)th time-series scene feature lexical units with the vectors in the set of risk behavior vectors predicted in the Kth time-series to determine the risk identification result includes:

[0029] The obtained K+1 time-series scene feature lexical units are matched with the vectors in the risk behavior vector set predicted in the Kth time-series. If the match is successful, the global risk confidence score is calculated based on the nonlinear activation function.

[0030] The risk level is determined based on the global risk confidence score, and corresponding alarm operations are initiated based on the risk level.

[0031] In view of the prior art, this application discloses a risk identification device that can overcome the shortcomings of the existing technical analysis methods, such as being fixed and slow to react, and can flexibly and quickly identify hidden risks.

[0032] This application proposes a risk identification device, comprising:

[0033] The scene recognition module is used to intercept multimodal raw data streams in real time, extract multimodal data features based on the multimodal raw data streams, and determine the corresponding expert intelligent agent based on the context to which the multimodal raw data streams belong.

[0034] The intelligent agent module is used to perform intensive domain-specific feature refinement preprocessing on the multimodal data features to obtain semantic evidence lexical units and scene feature lexical units for the current Kth time sequence, where K is a natural number; the semantic evidence lexical units include key data features and causal logical relationships obtained from the intensive domain-specific feature refinement process, and the scene feature lexical units include vector values ​​composed of behavioral milestone snapshots converted from the current interaction state;

[0035] The context logic engine model is used to perform semantic parsing and causal reconciliation based on the semantic evidence lexical units of the current Kth time sequence, and generate the risk hypothesis vector of the current Kth time sequence;

[0036] A generative prediction model is used to take the current risk hypothesis vector of the Kth time series as a conditional control and use a self-attention mechanism to infer and generate a set of risk behavior vectors for the Kth time series prediction.

[0037] The prediction matching module is used to match the obtained scene feature words of the K+1 time series with the vectors in the risk behavior vector set predicted by the Kth time series at the K+1 time series, so as to determine the risk identification result.

[0038] Furthermore, the device further includes:

[0039] The alarm module is used to initiate alarm operations.

[0040] In view of the prior art, this application discloses an electronic device that can overcome the shortcomings of the existing analysis methods, such as being fixed and slow to react, and can flexibly and quickly identify hidden risks.

[0041] This application discloses an electronic device comprising:

[0042] processor;

[0043] Memory used to store the processor's executable instructions;

[0044] The processor is configured to read the executable instructions from the memory and execute the instructions to implement any of the risk identification methods described above.

[0045] In view of the prior art, this application discloses a computer-readable storage medium that can overcome the shortcomings of the existing analysis methods, such as being fixed and slow to react, and can flexibly and quickly identify hidden risks.

[0046] This application proposes a computer-readable storage medium including computer instructions that, when executed by a processor, implement the risk identification method as described in any of the preceding claims.

[0047] In summary, the embodiments of this application proactively seek contradictions and conflicts in causal relationships within interactive information to predict potential risky behaviors and identify hidden risks in the interaction. The expert intelligent agent, the contextual logic engine model, and the generative prediction model constitute a three-level heterogeneous cascaded topology, decoupling logical depth and temporal density, which can significantly reduce computational load and transmission bandwidth requirements, enabling rapid risk identification. Attached Figure Description

[0048] Figure 1 This is a flowchart of an embodiment of the risk identification method of this application.

[0049] Figure 2 This is a logical schematic diagram illustrating the implementation of Embodiment 1 of the method of this application.

[0050] Figure 3 This is a schematic diagram of scenario one for risk identification implemented in this application.

[0051] Figure 4 This is a schematic diagram of scenario two for risk identification implemented in this application.

[0052] Figure 5 This is a schematic diagram of scenario three for risk identification implemented in this application.

[0053] Figure 6 This is a schematic diagram of the structure of an embodiment of the risk identification device implemented in this application.

[0054] Figure 7 This is a schematic diagram of the electronic device used in this application to achieve risk identification. Detailed Implementation

[0055] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0056] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0057] Existing technologies primarily rely on static and partial data for risk analysis. For example, they might add phone numbers deemed risky to a special list, alerting the user upon receiving a call from this list; or input data fragments of user-interaction interactions (voice or text) into a discriminative deep learning model (CNN or RNN) to calculate the risk probability of the current interaction, resulting in a binary classification; or input macro-level behavioral characteristics of user-interaction interactions (such as interaction frequency, call duration, geolocation tags, historical reputation scores, and social graph data) into a statistical model for risk identification. It is evident that in existing technologies, special lists represent static data, voice or text interaction data fragments and macro-level behavioral characteristics represent partial data, and CNNs, RNNs, and statistical models are pre-set, fixed analysis methods. Furthermore, these methods essentially input already-occurred interaction data into the model, which passively analyzes the interaction data and provides a risk identification result, with the risk identification result being some time after the user interaction.

[0058] However, in reality, risk events are becoming increasingly insidious, making it difficult for existing technologies to identify them quickly and flexibly. For example, the party initiating the risk and the user may appear to be interacting normally, but in reality, the user is being guided through a carefully designed script to ultimately trigger the risk. In this application, for ease of description, the party initiating the risk will be referred to as the "risk party," and the party that needs to avoid the risk will be referred to as the "user party."

[0059] Regarding the hidden risks inherent in seemingly normal interactions, this application finds that even in seemingly normal interactions, if the risk-creating party intends to deliberately induce risk, the interaction information must contain deep, subtle contradictions and conflicts; otherwise, it cannot achieve its purpose of creating risk. This type of risk also has the characteristic of a relatively long interaction process and a large amount of interactive information. Directly using existing large-scale models for reasoning may result in high computational costs and slow feedback. Therefore, this application, on the one hand, no longer relies on the model to passively analyze interaction fragment data, but instead adopts proactive generative intent reasoning; on the other hand, it employs a three-level heterogeneous cascaded topology structure—expert intelligent agent-contextual logic engine model-generative prediction model—to decouple logical depth and temporal density, achieving both deep analysis and rapid response. For example, the expert agent can be implemented using a lightweight edge-side small language model (SLM), which typically includes 1B to 3B parameters; the context logic engine model can be implemented using a dense logic large language model (LLM), which typically includes 70B to 100B+ parameters; and the generative prediction model can be implemented using a discrete time-series Transformer-based generative prediction model, which typically includes 10M to 100M parameters.

[0060] Figure 1 This is a flowchart of an embodiment of the risk identification method of this application. Figure 1 As shown, the method includes: scene recognition step 101, expert agent preprocessing step 102, context logic engine model deep logic analysis step 103, generative prediction model active intent reasoning step 104, and prediction matching risk identification step 105.

[0061] The scene recognition step 101 specifically includes: real-time interception of multimodal raw data streams, extraction of multimodal data features from the multimodal raw data streams, and determination of the corresponding expert agent based on the context of the multimodal raw data streams. In practical applications, risk parties and users can interact online or offline through mobile devices, computers, and other smart terminal devices on a trading platform or social software. The data stream generated by the interaction can be a multimodal raw data stream of any combination of text data, audio data, view data, and system metadata. For example, in a social software, risk parties and users can send each other text, audio, video, or images. These interactive information belong to different modalities, hence the term multimodal raw data stream. This step intercepts these multimodal raw data streams and transmits them to the risk recognition system of this application for processing. To facilitate subsequent analysis and processing, it is necessary to extract multimodal data features from the multimodal raw data streams, that is, to extract any combination of text data features, video data features, view data features, and system metadata features. The topics of interaction can be those involving large amounts of information, complex rules, dense knowledge points, and hidden causal relationships, such as finance, law, and medicine. These topics belong to fields that can be called intensive domains. Different scenarios can be handled by different expert agents. Therefore, this step can determine the corresponding expert agent based on the context of the multimodal raw data stream. For example, topics involving finance are handled by a financial expert agent, topics involving law by a legal expert agent, topics involving medicine by a medical expert agent, and there are also social expert agents, AR glasses agents, home security agents, contract signing agents, etc. In specific implementation, the corresponding expert agents are further loaded into the user's device, instantiated, and bound to internal and external resources, such as pre-generated expert knowledge bases, risk prevention feature bases, session memories, search engine APIs, logic checkers, and risk interception gateways. These internal and external resources can be invoked as needed during subsequent risk identification.

[0062] The expert agent preprocessing step 102 specifically includes: the expert agent performs intensive domain-specific feature refinement preprocessing on the multimodal data features to obtain semantic evidence tokens and scene feature tokens for the current Kth time sequence, where K is a natural number. The semantic evidence tokens include key data features obtained from the intensive domain-specific feature refinement process, as well as causal logical relationships. The scene feature tokens include vector values ​​formed by converting the current interaction state into behavioral milestone snapshots. Since the interaction between the risk party and the user may contain a lot of redundant information unrelated to the domain and risk, resources such as the expert knowledge base and risk prevention feature library can be accessed to remove redundancy and refine key features, thereby obtaining the semantic evidence tokens and scene feature tokens for the current Kth time sequence. The semantic evidence token is a token describing semantic facts, containing not only features but also causal logical relationships. For example, if the interaction data between the risk party and the user is intercepted, containing feature A "police" and feature B "transfer", then the semantic evidence lexical "feature A" - "feature B" describes both the semantic fact and the causal logical relationship that "feature A" requires the execution of "feature B". Scenario feature lexical units are vector values ​​describing user behavior; these vector values ​​are discrete states. For example, during the interaction between the risk party and the user, guiding the user to perform actions such as "opening the app," "entering the amount," and "entering the password," each action can be abstracted as a behavioral milestone snapshot and represented by a vector of a discrete state, such as "001" for "opening the app," "002" for "entering the amount," and "003" for "entering the password."

[0063] Step 103 of the deep logic analysis in the context logic engine model specifically includes: inputting the semantic evidence tokens of the current Kth time sequence into the context logic engine model; the context logic engine model performs semantic parsing and causal reconciliation based on the semantic evidence tokens of the current Kth time sequence to generate a risk hypothesis vector for the current Kth time sequence. Since semantic evidence tokens are tokens describing semantic facts, containing not only features but also causal logical relationships, this step can analyze whether there are contradictions and risks based on their causal logical relationships, thereby generating a risk hypothesis vector. The semantic parsing mentioned here typically refers to translating human natural language into machine-understandable language through a series of techniques such as token decomposition, causal reasoning, intent classification, and parameter extraction. The parsed information has causal relationships, allowing for causal reconciliation. Causal reconciliation refers to examining information containing inherent causal relationships and analyzing whether there are contradictions. For example, in the example above regarding feature A "police" and feature B "transfer," assuming the knowledge base stipulates that police are not allowed to instruct ordinary citizens to make private transfers, then there is a logical conflict and contradiction between feature A "police" and feature B "transfer," thus generating a risk hypothesis vector. A risk assumption vector refers to the possibility of risk in causal reconciliation information; it is a risk assumption that can be represented by a vector in practice, hence the name risk assumption vector. The risk direction assumption vector can be a high-dimensional, dense, continuous vector, such as a vector encoded using a standard 768-dimensional space.

[0064] The generative prediction model's active intent reasoning step 104 specifically includes: inputting the risk hypothesis vector of the current K-th time series as a conditional control input to the generative prediction model, which then uses a self-attention mechanism to infer and generate a set of risk behavior vectors predicted for the K-th time series. Here, the generative prediction model refers to a model that predicts the next data information that may appear under given conditions by learning the joint probability distribution from a large amount of historical data. In this step, the generative prediction model makes a prediction at the current K-th time series, predicting the risk behavior vectors that may appear in the next time series, i.e., the K+1-th time series, and merges all possible risk behavior vectors into a single set, referred to here as the risk behavior vector set. In this application, since the risk party and the user interact through mobile devices, computers, or other devices on a trading platform or social software, the user's actions on the device are the real risk, such as "entering an amount" or "entering a password." Therefore, this step is actually predicting the risk behavior that the user may exhibit in the K+1-th time series at the current K-th time series. In practical applications, risk behavior can be a vector value describing user behavior, and this vector value is discrete.

[0065] Step 105 of the prediction and matching risk identification process specifically includes: at time series K+1, matching the obtained scene feature lexical units at time series K+1 with the vectors in the risk behavior vector set predicted at time series K to determine the risk identification result. In practical applications, the risk party and the user party interact in real time. Assuming a segment of multimodal raw data stream is intercepted at time series K, steps 101 to 104 above will be executed immediately. Similarly, if another segment of multimodal raw data stream is intercepted at time series K+1, steps 101 to 104 above will also be executed immediately. Therefore, at time series K+1, this application not only generates the risk hypothesis vector set predicted at time series K, but also obtains the scene feature lexical units at time series K+1. The vectors in the risk hypothesis vector set and the scene feature lexical units are both vector values ​​describing the discrete states of user behavior, and therefore can be quickly compared in this step without further analysis. If the match is successful, it means that the user behavior at time series K+1 is precisely the risk behavior predicted at time series K, thus successfully identifying the risk.

[0066] Figure 2 This is a logical schematic diagram illustrating an embodiment of the method described in this application. For example... Figure 2 As shown, the scene recognition module intercepts the input multimodal raw data stream in real time, executes step 101, extracts multimodal data features, and determines the corresponding expert agent; the expert agent executes step 102, performs preprocessing, and generates semantic evidence lexical units and scene feature lexical units for the current Kth time sequence; one semantic evidence lexical unit is input to the context logic engine model, and the other scene feature lexical unit is input to the prediction matching module; the context logic engine model executes step 103, performs semantic parsing and causal reconciliation, and generates a risk hypothesis vector for the current Kth time sequence; the generative prediction model executes step 104, generates a set of risk behavior vectors predicted for the Kth time sequence; the prediction matching module executes step 105, matches the obtained scene feature lexical units for the K+1th time sequence with the vectors in the set of risk behavior vectors predicted for the Kth time sequence to determine the risk recognition result. In the K+1 time series, steps 101 to 105 above will be repeated. However, in step 105, the obtained scene feature words in the K+2 time series will be matched with the vectors in the risk behavior vector set predicted in the K+1 time series, and so on.

[0067] Firstly, the embodiments of this application are not fixed to a certain passive risk analysis method, but actively seek contradictions and conflicts in causal relationships within the interactive information, predict possible risky behaviors, and match them with the actual user behavior in the next time sequence, thereby identifying hidden risks in the interaction. Since the context logic engine model identifies risks by finding conflicts and contradictions in causal relationships within the interactive information, it is not bound by fixed risk patterns. When existing technologies fail to analyze static data or fixed risk patterns, the solutions of this application can flexibly address situations that appear normal but contain hidden risks. Secondly, since this application introduces a context logic engine model for deep logic calculation, to avoid the sluggish response caused by frequent and large amounts of deep logic calculation, this application also introduces an expert agent and a generative prediction model. The expert agent, context logic engine model, and generative prediction model constitute a three-level heterogeneous cascaded topology, decoupling logic depth and temporal density. Specifically, the expert agent is a lightweight, edge-side Small Language Model (SLM), the context logic engine model is a dense, logic-intensive Large Language Model (LLM), and the generative prediction model is a discrete-time model. The expert agent first refines the time-intensive interaction information into a small number of semantic evidence lexical units, which are then used by the context logic engine model for deep logical reasoning, finally generating discrete state vector values. This avoids the frequent triggering of deep logical operations by the context logic engine model due to time-intensive interaction information, significantly reducing computational load and bandwidth requirements. Furthermore, since the context logic engine model performs deep logical operations in advance in the previous time sequence, it can directly compare risk behavior vectors with user behavior in the next time sequence, thus quickly completing risk identification. In contrast, ordinary Large Language Model (LLM)-based models tightly couple logical depth and temporal density, causing each interaction to trigger matrix calculations involving trillions of parameters, resulting in significant response delays and computational bandwidth costs, making it difficult to achieve real-time interception and flexible, rapid risk identification. Additionally, in this embodiment, the context logic engine model only obtains risk hypotheses after the previous time sequence analysis; it still needs to match the predictions in the next time sequence to finally determine whether it is a risk. Its cross-time sequence verification mechanism can eliminate independent errors and improve the accuracy of risk identification in complex scenarios.

[0068] For example, the step 102 of the expert agent preprocessing to obtain the semantic evidence lexical units and scene feature lexical units of the current Kth time sequence includes: a1. The expert agent performs intensive domain-specific feature refinement preprocessing on each single-modal data feature in the multimodal data features to obtain the key data features of the corresponding single-modal, and the key data features of the corresponding single-modal have causal logical relationships, and determines the feature causal logical chain of the corresponding single-modal; align all the feature causal logical chains of the single-modal and map them into semantic vectors to obtain the semantic evidence lexical units of the current Kth time sequence; a2. Convert the current interaction state into a behavior milestone snapshot, and map the behavior milestone snapshot to the pre-existing behavior dictionary space to obtain vector values ​​as the scene feature lexical units of the current Kth time sequence. The behavior dictionary space represents the set of vector values ​​corresponding to all the behaviors that the user can operate on the device.

[0069] In step a1, multimodal data features include any combination of text data features, video data features, view data features, and system metadata features. In other words, in practical applications, text data features such as text on the device screen, interface layout structure, and page spatial position can be identified; audio and video features such as visual features, acoustic features, tone of voice changes, and speech-to-text conversion can be identified; and system metadata features such as foreground application names and system timestamps can be identified. This step first performs intensive domain-specific feature refinement preprocessing on each type of single-modal data feature to obtain the single-modal feature causal logic chain. Similarly, during the intensive domain-specific feature refinement preprocessing, redundant information unrelated to the domain and risk will be removed. For example, when the risk party and the user interact, there may be some casual greetings such as "How's the weather?" This step needs to remove these, refining the key features related to the intensive domain-specific features to form a logic chain. Furthermore, during interaction, the raw data streams of various monomodals may not be aligned on the timeline. For example, the risk party might send the voice message "police" in the first second and then the text message "police" in the fifth second. Although they address the same semantic meaning, they are not fully aligned on the timeline. To facilitate subsequent semantic analysis, this step also requires aligning the logical chains of each monomodal on the timeline. To facilitate machine processing, this needs to be further mapped into a vector. This vector can be encoded using a standard 768-dimensional semantic space, i.e., an array containing 768 floating-point numbers, each value ranging from [-1.0, 1.0]. Each dimension of this vector maps the projection probability of the interaction fact onto multiple specific risk space axes, providing high information density input for the subsequent context logic engine model. Taking feature A "police" and feature B "transfer" as examples, their mapped semantic vector representation is: [0.85, -0.23, 0.61, ...]. The first dimension represents identity, with higher values ​​indicating a closer approximation to official status. The second dimension represents redundant social information, which is reduced to zero after preprocessing. The third dimension represents information related to fund transfers, with higher values ​​indicating a closer approximation to the meaning of fund transfers. The meanings of other dimensions can be set according to the actual situation and are not listed here. In summary, multimodal data features, after this step, can be mapped into semantic vectors that include key data features and causal logical relationships, i.e., semantic evidence lexical units. Of course, in practical applications, the determination of semantic evidence lexical units is not limited to the above method; other dimensions can also be used, as long as they can contain key data features and causal logical relationships.

[0070] System metadata represents system-related data of the device where the user interacts, capturing user actions through the interface, such as button selection and entering numbers. In step a2, the user interaction state can be represented as a series of behavioral milestones. For example, in the process of a user opening an app, entering an amount, and entering a password, "opening the app," "entering an amount," and "entering a password" can be considered key nodes of the interaction, each of which is a behavioral milestone. The system can capture these behavioral milestones as snapshots and represent them as vector values. To accurately represent user behavior, a behavior dictionary can be generated beforehand, representing the total set of vector values ​​corresponding to all user actions on the device, defining a binary vector for each user action. Assuming the behavior dictionary space is 512-dimensional, each user action can be mapped to a 1×512-dimensional sparse vector in the behavior dictionary space. Assume the behavior dictionary contains: the first dimension represents the binary vector corresponding to the snapshot of opening a bank app, where 1 indicates execution and 0 indicates no execution; the second dimension represents the binary vector corresponding to the snapshot of opening screen sharing / remote control software, where 1 indicates execution and 0 indicates no execution; the third dimension represents the binary vector corresponding to the snapshot of the cursor focusing on the SMS verification code input box, where 1 indicates execution and 0 indicates no execution; and so on. Therefore, if a user enters an SMS verification code, mapping this action milestone snapshot to the behavior dictionary space yields a sparse vector of [0,0,1,0,0,…,0]. Different user actions result in different sparse vector values, which are discrete. Therefore, step a2 can map user actions to the behavior dictionary space in the above way to obtain vector values, i.e., scene feature terms. In practical applications, the determination of scene feature terms is not limited to the above method; any term that reflects the user action is acceptable.

[0071] For example, the deep logic analysis step 103 of the context logic engine model generates the risk hypothesis vector for the current Kth time series, including: step b1, the context logic engine model performs multi-dimensional semantic parsing on the semantic evidence lexical units of the current Kth time series to obtain the multi-dimensional semantic parsing result; step b2, the context logic engine model performs chain reasoning on the multi-dimensional semantic parsing result based on the deep semantic architecture of chain thinking ability, and performs cross-dimensional causal reconciliation; step b3, the context logic engine model performs causal consistency verification based on the cross-dimensional causal reconciliation result, and generates the risk hypothesis vector for the current Kth time series if the causal consistency verification fails.

[0072] In step b1, the multidimensional semantic parsing refers to mapping semantics to multiple dimensions for analysis, such as identity, intent, and emotional control. The identity dimension determines the identity of the interacting party, the intent dimension determines the potential intent or desired outcome, and the emotional dimension determines whether the interacting party has engaged in psychological suggestion or emotional control. In other words, this step decomposes semantic evidence units from multiple perspectives to facilitate further analysis. The chain-like reasoning ability in step b2 represents a step-by-step reasoning ability, where A leads to B, B leads to C, and so on, with each step depending on the result of the previous step, forming a logical chain. Step b1 has already performed semantic parsing from multiple dimensions; here, chain-like reasoning will be performed from multiple dimensions to conduct cross-dimensional causal reconciliation. For example, feature A, "police officer," from the identity dimension, refers to a public official; the rules in the knowledge base indicate that public officials are not allowed to instruct ordinary citizens to transfer private funds. Feature B, "transfer," from the behavioral dimension, is a fund transfer behavior. Feature A, "police officer," is the cause, and feature B, "transfer," is the effect. Step b3 performs causal consistency verification, which essentially checks for logical contradictions or whether the logic is consistent with common sense. Public officials are not allowed to instruct ordinary citizens to transfer private funds, yet here feature A, "police," requests feature B, "transfer," which is clearly illogical. For example, suppose during an interaction, the risk party claims to be official customer service, allowing users to upgrade for free and instructing them to click a button on a screen to install it. After the user clicks this button, the system detects that the button will activate the automatic deduction function. This step forms a logical chain C: "official customer service" - "free" - "click to install," and also a logical chain D: "click" - "activate automatic deduction function." Analysis shows that logical chains C and D conflict and contradict each other. These are just two simple examples. In practical applications, the contextual logic engine model will comprehensively and multidimensionally analyze semantic evidence lexical units, perform chain-like reasoning based on its own deep semantic architecture, and conduct causal reconciliation from all aspects and dimensions. After causal reconciliation, step b3 will determine the causal consistency verification result. If the causal reconciliation is consistent with the rules or common sense, the causal consistency verification is successful; otherwise, it fails, indicating that there may be risks in the interaction. Risk assumptions can also be represented by vectors, encoded using a standard 768-dimensional semantic space, i.e., an array containing 768 floating-point numbers, each value ranging from [-1.0, 1.0]. Each dimension of this vector maps the projection probability of the inference analysis onto multiple specific risk space axes, providing high information density input for subsequent generative prediction models.Taking "police," "urgent request," and "transfer" as examples, the generated risk hypothesis vector is represented as [0.92, -0.08, 0.74, ...]. The first dimension represents identity; a higher value indicates closer to official status, while a lower value indicates a non-official identity. The second dimension represents the intent to commit fraudulent online investment / high-interest inducement; a higher value indicates a stronger inducement, while a lower value indicates a weaker inducement. The third dimension represents the urgency of the fund transfer; a higher value indicates greater urgency, while a lower value indicates less urgency. The meanings of other dimensions can be set according to the actual situation and are not listed here. In summary, semantic evidence lexical units, analyzed by the contextual logic engine model, can generate a risk hypothesis vector.

[0073] It is important to note that although both risk hypothesis vectors and semantic evidence lexical units are high-dimensional, dense, and continuous vectors, their lifecycles, abstraction levels, and functional roles are completely different. Semantic evidence lexical units are factual descriptions of the preceding input, while risk hypothesis vectors are inference and analytical conclusions of the following output.

[0074] For example, the step 104 of the generative prediction model's active intent reasoning to generate the set of risk behavior vectors for the Kth time series prediction specifically includes: step c1, the generative prediction model uses the risk hypothesis vector of the current Kth time series as a condition control and uses a self-attention mechanism to reason and generate risk behavior probabilities in the behavior dictionary space; step c2, the vector values ​​based on the behavior dictionary space corresponding to the risk behavior probabilities exceeding a preset threshold are merged as the set of risk behavior vectors for the Kth time series prediction.

[0075] The generative prediction model is essentially a conditional probability sequence generation network. It uses the risk hypothesis vector input from the previous stage as a strong boundary condition and employs a masked self-attention mechanism to autoregressively predict the probability of future actions within the action dictionary space. The joint conditional probability distribution of each action can be represented by the following formula:

[0076] ,

[0077]

[0078] Where P represents probability, C represents risk hypothesis vector, and Y = (y1, y2, ..., y3) N ) represents the sequence of vector values ​​corresponding to a certain potential behavior in the future, N represents the total number of behaviors, i represents any behavior, softmax represents the normalized activation operator, decoder represents the discrete time series, W0 represents the weight matrix, and b0 represents the bias.

[0079] In practical applications, generative prediction models are implemented using Transformer-based generative prediction models to infer the evolution direction of risk patterns. If multiple probabilities exceeding a threshold are calculated, the vector values ​​corresponding to these probabilities are merged into a risk behavior vector set. Each vector value is a 512-dimensional vector value obtained by mapping from a behavior dictionary space, representing potential subsequent risk behaviors. Since the risk behavior vectors are predicted at the Kth time series, this set is called the risk behavior vector set predicted at the Kth time series.

[0080] For example, the risk identification step 105, which involves matching the obtained K+1 time-series scene feature words with the vectors in the set of risk behavior vectors predicted in the Kth time-series, to determine the risk identification result, specifically includes: step d1, matching the obtained K+1 time-series scene feature words with the vectors in the set of risk behavior vectors predicted in the Kth time-series; if the match is successful, calculating the global risk confidence score based on the nonlinear activation function; step d2, determining the risk level based on the global risk confidence score, and initiating the corresponding alarm operation based on the risk level.

[0081] The calculation of the global risk confidence score using a nonlinear activation function can be expressed by the following formula: ,in, It is a global risk confidence score, used to represent the final confidence level in identifying the current behavior as a certain risk that is occurring; It is a sigmoid non-linear activation function; It matches the current potential risk with confidence level; It is a historical cyclical cumulative penalty factor; λ and λ are weight parameters used to control the match confidence and penalty factor; This is the bias term. Since the vectors in the scene feature lexicon and the risk behavior vector set are both discrete 512-dimensional vector values, this step can quickly perform matching and calculate the global risk confidence score, and initiate the corresponding alarm operation. That is, if a certain risk behavior vector M (512-dimensional) is predicted in the Kth time series, and a certain user behavior is captured in the K+1th time series and a corresponding scene feature lexicon N (512-dimensional) is generated, M and N can be matched in a very short time (milliseconds), and an alarm is immediately initiated to remind the user of the potential risk. In addition, the alarm operation method can be set according to the level, such as from minor to severe, set to UI pop-up, haptic feedback, voice feedback, automatic termination of interaction, forced interception of payment interface, etc., and the user behavior is fed back to the system.

[0082] For example, Figure 3This is a schematic diagram illustrating scenario one of risk identification. As shown in Figure 3, the risk party and the user interact through social media software on a smartphone terminal. The risk identification system of this application is installed inside the user's smartphone terminal. When the user clicks the button in the lower left corner, they can open the risk identification system of this embodiment. During the communication between the two parties, the text or voice information input by both parties through the social media software constitutes a multimodal raw data stream. The background risk identification system uses steps 101 to 105 of the method described in the above embodiment to perform risk identification. Suppose the risk party guides the user to leave the social media software interface and open a bank's APP, the risk identification system identifies the user's risky behavior, issues an alarm, displays the risk level, and pops up a window reminding the user to end the conversation.

[0083] For example, Figure 4 This is a schematic diagram illustrating scenario two for risk identification. For example... Figure 4 As shown, assume a user is listening to a salesperson's presentation and preparing to sign a contract. The risk identification system in this application is installed inside the user's smartphone or AR glasses. The user activates the system through a smart terminal such as a smartphone or AR glasses. The system captures the salesperson's (risk party's) voice through a microphone to obtain verbal promises and uses OCR technology to scan the contract text with a camera to obtain written terms. Here, the voice and contract text are the input multimodal raw data streams. The background risk identification system uses steps 101-105 of the method described in the above embodiment for risk identification. Specifically, step a1 in expert agent preprocessing step 102 aligns the salesperson's voice with the written terms in the contract text; step b1 in context logic engine model deep logic analysis step 103 performs multidimensional semantic parsing; and step b2 performs cross-dimensional causal reconciliation. If it is found that the salesperson's verbal promise and the written terms of the contract text are inconsistent—for example, the verbal promise is "You can cancel this insurance at any time and get a full refund," while clause 10 of the contract text stipulates "A cancellation fee must be paid after the cooling-off period"—the system will detect discrepancies. The system will then generate a risk hypothesis vector based on this conflict and contradiction, and the generative prediction model will generate a risk behavior vector (the behavior of signing a contract) through the active intent reasoning step 104. If the system detects that the user is about to sign a contract, the prediction matching risk identification step 105 will initiate a corresponding alarm operation, such as reminding the user through the vibration of a smartphone or AR glasses.

[0084] For example, Figure 5 This is a schematic diagram illustrating scenario three for risk identification. For example... Figure 5As shown, assuming a smart access control system is installed on the user's home door, the risk identification system of this application is installed inside the smart access control system and activated when someone knocks on the door. The system captures voice information through the access control intercom function and video information through a camera. Here, voice and video are the input multimodal raw data streams. The background risk identification system uses steps 101-105 of the method described in the above embodiment to perform risk identification. Suppose a person claiming to be from the property management company comes to the user's home, and the user happens to be a minor. The person claiming to be from the property management company needs to enter the home to investigate a water leak and states that the situation is urgent, demanding that the user open the door immediately. The system analysis reveals that the person claiming to be from the property management company has not presented any valid identification or evidence of a water leak, and uses an urgent tone. The generative prediction model actively generates a risk behavior vector (the behavior of opening the door) in step 104. If the system detects that the user is about to open the door, the prediction matching risk identification step 105 will trigger a corresponding alarm operation, such as notifying the user not to open the door via the access control voice and immediately reporting to the linked parent's mobile phone.

[0085] As can be seen, the implementation scheme of this application can be applied to various devices such as smartphones, computers, smart bracelets, smart glasses, and smart access control systems, and can be used in various scenarios. The implementation scheme of this application can flexibly, quickly, and accurately identify hidden risks in complex scenarios, providing high-value protection solutions for various application scenarios.

[0086] This application also provides an embodiment of a risk identification device. Figure 6 This is a structural schematic diagram of an embodiment of the risk identification device implemented in this application. Figure 6 As shown, the device includes: a scene recognition module 601, an intelligent agent module 602, a context logic engine model 603, a generative prediction model 604, and a prediction matching module 605.

[0087] The system includes several modules: a scene recognition module 601, which intercepts multimodal raw data streams in real time, extracts multimodal data features from the streams, and determines the corresponding expert agent based on the context of the data streams; an agent module 602, which performs intensive domain-specific feature refinement preprocessing on the multimodal data features to obtain semantic evidence lexical units and scene feature lexical units for the current K-th time sequence, where K is a natural number; the semantic evidence lexical units include key data features and causal logical relationships obtained from the intensive domain-specific feature refinement process, and the scene feature lexical units include vector values ​​formed by converting the current interaction state into behavioral milestone snapshots; a context logic engine model 603, which performs semantic parsing and causal reconciliation based on the semantic evidence lexical units for the current K-th time sequence to generate a risk hypothesis vector for the current K-th time sequence; and a generative prediction model 604, which uses the risk hypothesis vector for the current K-th time sequence as a conditional control and employs a self-attention mechanism to generate a set of predicted risk behavior vectors for the K-th time sequence. The prediction matching module 605 is used to match the obtained scene feature lexical units of the (K+1)th time series with the vectors in the risk behavior vector set predicted for the Kth time series at the (K+1)th time series, so as to determine the risk identification result. Exemplarily, the device further includes an alarm module 606 for initiating an alarm operation.

[0088] For example, the risk identification device of this application can be installed in various devices such as smartphones, computers, smart bracelets, smart glasses, and smart access control systems. The agent module 602, the context logic engine model 603, and the generative prediction model 604 form a three-level heterogeneous cascaded topology, decoupling the logic depth and temporal density. The agent module 602 is a lightweight edge-side Small Language Model (SLM), the context logic engine model 603 is a dense logic Large Language Model (LLM), and the generative prediction model 604 is a discrete temporal model. If the context logic engine model 603 is large, it can also be placed in the cloud, with the agent module 602 uploading semantic evidence lexical units and scene feature lexical units for cloud processing.

[0089] As can be seen, the embodiments of this application proactively seek contradictions and conflicts in causal relationships within the interactive information to predict potential risky behaviors, thus flexibly addressing situations that appear normal but harbor hidden risks. The semantic evidence lexical units and scene feature lexical units generated by the intelligent agent module 602 are small in size, avoiding frequent deep logic operations by the context logic engine model 603, reducing computational load and bandwidth requirements, and achieving rapid risk identification. Furthermore, in the embodiments of this application, the context logic engine model 603 obtains risk hypotheses after the previous time-series analysis, and only confirms the risk by matching the predictions in the next time-series. This cross-time-series verification mechanism can eliminate independent errors and improve the accuracy of risk identification in complex scenarios.

[0090] Exemplary, this application also provides a computer-readable storage medium storing instructions that, when executed by a processor, can perform the steps in the risk identification method described above. In practical applications, the computer-readable medium may be included in the device / apparatus / system described in the above embodiments, or it may exist independently without being assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, can implement the smart view generation method described in the above embodiments. According to the embodiments disclosed in this application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof, but not intended to limit the scope of protection of this application. In the embodiments disclosed in this application, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0091] By way of example, embodiments of the present invention also provide an electronic device in which the risk identification apparatus of the present application embodiments can be integrated. For example... Figure 7 The diagram illustrates the structure of an electronic device according to an embodiment of the present invention. Specifically, the electronic device may include a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, and a computer program stored in the memory and executable on the processor. When the program in the memory 702 is executed, the aforementioned method for generating a smart view can be implemented. In practical applications, the electronic device may also include components such as a power supply 703, an input unit 704, and an output unit 705. Those skilled in the art will understand that... Figure 7The structure of the electronic device shown does not constitute a limitation on the device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. The processor 701 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It performs overall monitoring of the electronic device by running or executing software programs and / or modules stored in memory 702, and by calling data stored in memory 702, thereby executing various server functions and processing data. Memory 702 can be used to store software programs and modules, i.e., the aforementioned computer-readable storage medium. The processor 701 executes various functional applications and data processing by running the software programs and modules stored in memory 702. Memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created based on server usage, etc. Furthermore, memory 702 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702. The electronic device also includes a power supply 703 that supplies power to the various components, which can be logically connected to the processor 701 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 703 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device may also include an input unit 704, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. The electronic device may also include an output unit 705, which can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, which can be composed of graphics, text, icons, video, and any combination thereof.

[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments disclosed in this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings. For example, two blocks shown connectedly may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0093] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, without departing from the spirit and teachings of this application, the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways, and all such combinations and / or combinations fall within the scope of this application.

[0094] This document uses specific embodiments to illustrate the principles and implementation methods of the present invention. The descriptions of these embodiments are merely illustrative of the method and core concepts of the present invention and are not intended to limit this application. Those skilled in the art can make changes to the specific implementation methods and application scope based on the ideas, spirit, and principles of the present invention. Any modifications, equivalent substitutions, or improvements made should be included within the scope of protection of this application.

Claims

1. A risk identification method, characterized in that, The method includes: Real-time interception of multimodal raw data streams, extraction of multimodal data features based on the multimodal raw data streams, and determination of the corresponding expert agent based on the context to which the multimodal raw data streams belong; The expert agent performs intensive domain-specific feature refinement preprocessing on the multimodal data features to obtain semantic evidence lexical units and scene feature lexical units for the current Kth time sequence, where K is a natural number. The semantic evidence lexical units include key data features obtained from the intensive domain-specific feature refinement process and causal logical relationships. The scene feature lexical units include vector values ​​formed by converting the current interaction state into behavioral milestone snapshots. The semantic evidence lexical units of the current Kth time sequence are input into the context logic engine model. The context logic engine model performs semantic parsing and causal reconciliation based on the semantic evidence lexical units of the current Kth time sequence to generate the risk hypothesis vector of the current Kth time sequence. The risk hypothesis vector of the current Kth time series is used as a conditional control input to the generative prediction model, which then uses a self-attention mechanism to infer and generate a set of risk behavior vectors for the Kth time series prediction. At time series K+1, the obtained scene feature lexical units of time series K+1 are matched with the vectors in the risk behavior vector set predicted at time series K to determine the risk identification result.

2. The method according to claim 1, characterized in that, The expert agent is a lightweight edge-side small language model (SLM). The context logic engine model is a large language model (LLM) with dense logic. The generative prediction model is a discrete time series generative prediction model based on Transformer. The expert agent, the context logic engine model, and the generative prediction model constitute a three-level heterogeneous cascaded topology.

3. The method according to claim 1, characterized in that, The multimodal raw data stream includes any combination of text data, audio data, view data, and system metadata; The steps for extracting multimodal data features include: extracting any combination of text data features, video data features, view data features, and system metadata features.

4. The method according to claim 1, characterized in that, The steps of the expert agent performing intensive domain-specific feature refinement preprocessing on the multimodal data features to obtain the semantic evidence lexical units and scene feature lexical units of the current Kth time sequence include: The expert agent performs intensive domain-specific feature refinement preprocessing on each single-modal data feature in the multimodal data features to obtain the corresponding single-modal key data features, and the corresponding single-modal key data features have the causal logical relationship, thus determining the feature causal logical chain of the corresponding single-modal; aligning all the feature causal logical chains of the single-modal features and mapping them to semantic vectors, the semantic evidence lexical of the current Kth time sequence is obtained; The current interaction state is converted into a behavior milestone snapshot, and the behavior milestone snapshot is mapped to a pre-existing behavior dictionary space to obtain vector values ​​as the scene feature words of the current Kth time sequence. The behavior dictionary space represents the total set of vector values ​​corresponding to all the behaviors that the user can operate on the device.

5. The method according to claim 1, characterized in that, The steps of the context logic engine model to generate the risk hypothesis vector for the current Kth time sequence by performing semantic parsing and causal reconciliation based on the semantic evidence lexical units of the current Kth time sequence include: The context logic engine model performs multi-dimensional semantic parsing on the semantic evidence lexical units of the current Kth time sequence to obtain multi-dimensional semantic parsing results; The context logic engine model performs chain reasoning on the multidimensional semantic parsing results based on the deep semantic architecture of chain thinking ability, and performs cross-dimensional causal reconciliation. The context logic engine model performs causal consistency verification based on the results of the cross-dimensional causal reconciliation. If the causal consistency verification fails, a risk hypothesis vector for the current Kth time series is generated.

6. The method according to claim 1, characterized in that, The step of generating the set of risk behavior vectors for the Kth time series prediction by the generative prediction model using a self-attention mechanism includes: The generative prediction model uses the risk hypothesis vector of the current Kth time series as a conditional control and uses a self-attention mechanism to infer and generate risk behavior probabilities in the behavior dictionary space. The vector values ​​based on the behavior dictionary space corresponding to the risk behavior probabilities exceeding the preset threshold are merged to form the risk behavior vector set for the Kth time series prediction.

7. The method according to claim 1, characterized in that, The step of matching the obtained (K+1)th time-series scene feature lexical units with the vectors in the set of risk behavior vectors predicted in the Kth time-series to determine the risk identification result includes: The obtained K+1 time-series scene feature lexical units are matched with the vectors in the risk behavior vector set predicted in the Kth time-series. If the match is successful, the global risk confidence score is calculated based on the nonlinear activation function. The risk level is determined based on the global risk confidence score, and corresponding alarm operations are initiated based on the risk level.

8. A risk identification device, characterized in that, The device includes: The scene recognition module is used to intercept multimodal raw data streams in real time, extract multimodal data features based on the multimodal raw data streams, and determine the corresponding expert intelligent agent based on the context to which the multimodal raw data streams belong. The intelligent agent module is used to perform intensive domain-specific feature refinement preprocessing on the multimodal data features to obtain semantic evidence lexical units and scene feature lexical units for the current Kth time sequence, where K is a natural number; the semantic evidence lexical units include key data features and causal logical relationships obtained from the intensive domain-specific feature refinement process, and the scene feature lexical units include vector values ​​composed of behavioral milestone snapshots converted from the current interaction state; The context logic engine model is used to perform semantic parsing and causal reconciliation based on the semantic evidence lexical units of the current Kth time sequence, and generate the risk hypothesis vector of the current Kth time sequence; A generative prediction model is used to take the current risk hypothesis vector of the Kth time series as a conditional control and use a self-attention mechanism to infer and generate a set of risk behavior vectors for the Kth time series prediction. The prediction matching module is used to match the obtained scene feature words of the K+1 time series with the vectors in the risk behavior vector set predicted by the Kth time series at the K+1 time series, so as to determine the risk identification result.

9. The apparatus according to claim 8, characterized in that, The device further includes: The alarm module is used to initiate alarm operations.

10. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the risk identification method according to any one of claims 1 to 7.

11. A computer-readable storage medium comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they perform the risk identification method as described in any one of claims 1 to 7.