Anchor intention recognition method and device

By acquiring multimodal data streams in the live broadcast room and using a variational state space model to deentangle pure intent factors, the problem of the broadcaster's intent being obscured by the flood of information is solved, and the accurate identification and effective application of the broadcaster's intent are realized.

CN121744017APending Publication Date: 2026-03-27国家市场监督管理总局竞争政策与评估中心
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In live streaming scenarios, the host's intentions are difficult to identify accurately, and existing models are easily obscured by the massive and dynamic flood of information, leading to deviations in key operations such as risk management and compliance governance.

Method used

The multimodal data stream from the live broadcast room is acquired through a preset time window. After feature extraction, the pure intent factor, which is unaffected by the environment, is de-entangled using a variational state space model and then combined with an intent classifier for intent recognition.

Benefits of technology

It accurately identifies the true intentions of live streamers, supporting live streaming platforms in risk management, compliance governance, marketing intervention, and personalized recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744017A_ABST
    Figure CN121744017A_ABST
Patent Text Reader

Abstract

The invention discloses an anchor intention recognition method and device, relates to the technical field of anchor intention recognition, and mainly aims to accurately recognize the intention of an anchor based on a data stream of a live broadcast room in a live broadcast scene. Comprising the following steps: acquiring a multi-modal data stream of a current time window of a live broadcast room in real time by a preset time window; performing feature extraction on the multi-modal data stream to obtain a multi-modal feature vector sequence; a pre-trained variational state space model is called to map the multi-modal feature vector sequence to a hidden space, a hidden variable sequence representing the state of an anchor in a live broadcast room is generated, and the state of the anchor in the live broadcast room is obtained by using the deentanglement capability learned in anti-fact training of the variational state space model in a live broadcast scene. Based on the hidden variable sequence, deentangling a pure intention factor which is not influenced by the live broadcast room environment and is used for indicating the subjective intention of the live broadcast room anchor; and based on the pure intention factor obtained through de-entanglement, obtaining an intention recognition result of the anchor in the live broadcast room in the current time window through classification of an intention classifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of broadcaster intent recognition technology, and in particular to a broadcaster intent recognition method and apparatus. Background Technology

[0002] In the current era of rapid development in the live streaming industry, identifying the streamer's intent is the cornerstone of key operations for live streaming platforms, such as risk management and compliance governance, marketing intervention, and providing personalized recommendations to users.

[0003] Currently, the intentions of the streamer are typically identified based on the data stream within the live streaming room. However, in a live streaming scenario, the data stream consists primarily of two types of data: one is data directly related to the streamer's intentions, and the other is an information deluge formed by complex live streaming environmental factors (such as the streamer's diverse forms of expression, dynamically changing live streaming content, and dense interfering information). The streamer's true intentions are easily obscured by this latter type of massive and dynamic information deluge, making it difficult to accurately identify the streamer's intentions. This, in turn, hinders the live streaming platform from effectively implementing key operations such as risk management and compliance governance.

[0004] Therefore, in live streaming scenarios, how to accurately identify the streamer's intentions based on the data stream of the live streaming room has become an urgent problem to be solved. Summary of the Invention

[0005] This application proposes a method and apparatus for identifying the intention of a live streamer, the main purpose of which is to accurately identify the intention of the live streamer based on the data stream of the live streaming room in a live streaming scenario.

[0006] To achieve the above objectives, this application mainly provides the following technical solutions: Firstly, this application provides a method for identifying the intent of a live streamer. The method provided in this embodiment may include at least: acquiring a multimodal data stream of the current time window of the live stream in real time within a preset time window; extracting features from the multimodal data stream to obtain a multimodal feature vector sequence; calling a pre-trained variational state space model to map the multimodal feature vector sequence to a latent space, generating a sequence of latent variables representing the state of the live streamer; utilizing the deentanglement capability learned by the variational state space model during counterfactual training in a live streaming scenario, deentangled the latent variable sequence to obtain a pure intent factor unaffected by the live stream environment, the pure intent factor indicating the subjective intent of the live streamer; and classifying the intent identification result of the live streamer within the current time window using an intent classifier based on the deentangled pure intent factor.

[0007] Secondly, this application provides a broadcaster intent recognition device, which in this embodiment may include at least: The acquisition module is used to acquire the multimodal data stream of the live broadcast room in real time within a preset time window; The extraction module is used to extract features from the multimodal data stream to obtain a multimodal feature vector sequence; The identification module is used to call the pre-trained variational state space model to map the multimodal feature vector sequence to the latent space, generate a latent variable sequence representing the state of the live streamer, and use the deentanglement capability learned by the variational state space model in the counterfactual training in the live streaming scenario to deentangle the latent variable sequence to obtain a pure intention factor that is not affected by the live streaming environment. The pure intention factor is used to indicate the subjective intention of the live streamer. The generation module is used to classify the intent recognition results of the live streamer within the current time window based on the de-entangled pure intent factors through an intent classifier.

[0008] Thirdly, this application provides a computer-readable storage medium including a stored program, wherein the program, when running, controls the device where the storage medium is located to execute the broadcaster intent recognition method of the first aspect.

[0009] Fourthly, this application provides an electronic device, the electronic device comprising: a memory for storing a program; and a processor coupled to the memory for running the program to perform the anchor intent recognition method of the first aspect.

[0010] Fifthly, this application provides a computer program product comprising: a computer program / computer executable instructions, wherein the computer program / computer executable instructions are capable of performing the anchor intent recognition method of the first aspect.

[0011] The broadcaster intent recognition method and apparatus provided in this application, when it is determined that intent recognition of the broadcaster in a live room is required, acquires the multimodal data stream of the current time window of the live room in real time within a preset time window, and extracts features from the multimodal data stream to obtain a multimodal feature vector sequence. Then, a pre-trained variational state space model is invoked to map the multimodal feature vector sequence to the latent space, generating a sequence of latent variables representing the broadcaster's state. Utilizing the deentanglement capability learned by the variational state space model during counterfactual training in a live room scenario, the latent variable sequence is deentangled to extract a pure intent factor that is unaffected by the live room environment and indicates the broadcaster's subjective intent. Finally, based on the deentangled pure intent factor, an intent classifier is used to classify and obtain the intent recognition result of the broadcaster within the current time window. Therefore, the solution provided in this embodiment has at least the following specific effects: First, based on the multimodal data stream of the live broadcast room, the latent variables mixed in the multimodal data stream (such as the broadcaster's intention, environmental interaction and random interference) are effectively de-entangled through a pre-trained variational state space model, thereby accurately extracting the pure intention factor that represents the broadcaster's true intention, and then accurately identifying the broadcaster's intention based on the pure intention factor; Second, the de-entangled intention factor can serve as an interpretable feature to assist the live broadcast platform in key operations such as risk control and compliance governance, marketing intervention and providing personalized recommendations to users.

[0012] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 A flowchart of a broadcaster intent recognition method provided in one embodiment of this application is shown; Figure 2 This illustration shows a schematic diagram of a multimodal feature vector sequence generation process provided in one embodiment of this application; Figure 3 This illustration shows a schematic diagram of a target latent variable sequence formation process provided in one embodiment of this application; Figure 4 This illustration shows a counterfactual training process according to an embodiment of this application; Figure 5 This illustration shows a schematic diagram of the structure of a broadcaster intent recognition device according to an embodiment of this application; Figure 6 A schematic diagram of the structure of a broadcaster intent recognition device provided in another embodiment of this application is shown. Detailed Implementation

[0015] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art. Currently, models based on deep learning-based text sentiment analysis or action recognition algorithms are commonly used to identify the streamer's intent based on the data stream in the live streaming room. However, in a live streaming scenario, the data stream mainly consists of two types of data: one is data directly related to the streamer's intent, and the other is an information deluge formed by complex live streaming environmental factors, including the streamer's diverse forms of expression (such as colloquialisms, body language assistance, and strong context dependence), dynamically changing live streaming content (such as frequent topic switching and numerous sudden interactive scenarios), and dense interfering information (such as background noise, audience barrage interference, and irrelevant chatter). The streamer's true intent is easily obscured by this latter type of massive and dynamic information deluge, making it difficult for existing models to accurately identify the streamer's intent (for example, the model might mistakenly equate "exciting background music" directly with "promotional intent," ignoring the streamer's actual linguistic logic, leading to inaccurate intent identification). This can cause deviations in key operations such as risk management and compliance governance.

[0016] Research has revealed that live streaming generates multimodal data, which is essentially an external manifestation of the combined effects of various latent variables. These latent variables include the streamer's intention as well as the interactions and interference from the live streaming environment. Therefore, by deentanglement of these latent variables, the latent variable representing the streamer's intention can be separated from the rest to form an intention factor. The streamer's intention can then be identified based on this intention factor. Since the intention factor obtained through deentanglement is no longer affected by the latent variables of the live streaming environment's interactions and interference, it can accurately identify the streamer's true intention.

[0017] Based on the above findings, this embodiment specifically provides a broadcaster intent recognition scheme, which includes: acquiring multimodal data streams of the current time window of the live broadcast room in real time using a preset time window; extracting features from the multimodal data streams to obtain a multimodal feature vector sequence; calling a pre-trained variational state space model to map the multimodal feature vector sequence to the latent space, generating a latent variable sequence representing the broadcaster's state in the live broadcast room, and using the deentanglement capability learned by the variational state space model in counterfactual training in the live broadcast scenario to deentangle the latent variable sequence to obtain pure intent factors that are not affected by the live broadcast room environment. The pure intent factors are used to indicate the broadcaster's subjective intent in the live broadcast room; and classifying the broadcaster's intent within the current time window using an intent classifier based on the deentangled pure intent factors.

[0018] The streamer intent recognition scheme provided in this embodiment can be used to recognize the streamer intent in any live streaming room. This embodiment does not limit the type of live streaming room. For example, the type of live streaming room may include, but is not limited to, at least one of the following: e-commerce sales, entertainment interaction, knowledge popularization, gaming competition, and lifestyle sharing.

[0019] Based on the above-mentioned broadcaster intent recognition scheme, this embodiment specifically provides a broadcaster intent recognition method and device. The broadcaster intent recognition method and device provided in this embodiment will be described in detail below.

[0020] This application provides a method for identifying broadcaster intent, such as... Figure 1 As shown, the broadcaster intent recognition method provided in this embodiment may include at least the following steps 101 to 104.

[0021] 101. Acquire multimodal data streams of the live broadcast room in real time within a preset time window.

[0022] In some embodiments, the live streaming room in step 101 is a live streaming room that requires streamer intent recognition, so as to perform risk control and compliance governance, marketing intervention, and other operations on the live streaming room based on the intent recognition results. Specifically, the method for determining the live streaming room that requires intent recognition may include at least one of the following: First, when an intent recognition instruction is received, the live streaming room specified by the intent recognition instruction is determined as the live streaming room that requires intent recognition, so that the timing of streamer intent recognition can be flexibly controlled according to business needs. Second, if it is required that all live streaming rooms entering the live streaming state on the live streaming platform need to perform streamer intent recognition, then each time a live streaming room on the live streaming platform is monitored to start broadcasting and enter the live streaming state, that live streaming room is determined as the live streaming room that requires intent recognition. At least one of the above two methods can be flexibly selected according to business needs, and this embodiment does not limit this.

[0023] In some embodiments, after identifying the live streaming room, a step is performed to acquire the multimodal data stream of the live streaming room in real time within a preset time window. This step has two key implementation aspects: First, the preset time window. The live streaming content changes rapidly, so the data stream of the live streaming room is dynamic and fluctuates constantly. To avoid interference from instantaneous peaks or troughs, a continuous and smooth data stream over a period of time needs to be acquired through a preset time window to identify the broadcaster's intent based on this data stream. The preset time window can be flexibly selected based on business needs, for example, 5 seconds. Second, the multimodal data stream. In a live streaming scenario, a single-modal data stream (such as video) cannot piece together a complete profile of the broadcaster's intent. Therefore, it is necessary to acquire the multimodal data stream of the live streaming room in real time within the current time window to capture intent through the collaborative expression of the multimodal data stream, thereby improving the robustness and accuracy of intent recognition. Multimodality can include, but is not limited to, at least one of the following: text modality, behavioral modality, video modality, audio modality, etc.

[0024] In some embodiments, the method for acquiring the multimodal data stream of the current time window of the live broadcast room in real time with a preset time window may include the following steps 101A to 101B.

[0025] 101A. Select at least two target modalities required for intent recognition and collect the raw data stream corresponding to each target modality in real time within the current time window of the live broadcast room.

[0026] The target modality can be selected based on business needs. For example, the target modality may include, but is not limited to, at least one of the following: text modality, behavioral interaction modality, video modality, audio modality, etc. It should be noted that the more modalities there are, the more accurate the intent recognition will be.

[0027] For example, the target modalities include text modalities, video modalities, and audio modalities. The acquired raw data streams include: the raw data stream "audio data stream" corresponding to the text modalities, the raw data stream "video data stream" corresponding to the video modalities, and the raw data stream "audio data stream" corresponding to the audio modalities (this audio data stream is the same data stream as the audio data stream corresponding to the text modalities).

[0028] 101B. For each target modality, perform the following: Based on the processing tool that matches the current target modality, process the raw data stream of the current target modality to obtain the data stream corresponding to the current target modality.

[0029] When the current target modality is text, the corresponding audio data stream is transcribed into a corresponding speech-to-text sequence using a matching processing tool (such as the speech recognition tool ASR). This text sequence is then identified as the data stream corresponding to the language modality. Furthermore, considering that bullet screen data can reflect the broadcaster's intentions to some extent, the original data stream of the text modality can also include a "bullet screen data stream." In this case, the bullet screen data stream can also be transcribed into a corresponding text sequence using a matching processing tool (such as the speech recognition tool ASR), and this text sequence is incrementally concatenated into the original text sequence.

[0030] When the current target modality is a video modality, keyframes are extracted from the video data stream corresponding to the video modality using the matching processing tool (such as a keyframe extraction tool), and the keyframes are adjusted to a preset resolution (such as 224×224). The keyframes after adjusting the resolution are arranged into a video frame sequence in chronological order, and the video frame sequence is determined as the data stream corresponding to the video modality.

[0031] When the current target mode is an audio mode, the matching processing tool (such as the Mel spectrum tool) is used to extract the Mel spectrum from the audio data stream corresponding to the audio mode to form an audio waveform sequence, and the audio waveform sequence is determined as the data stream corresponding to the audio mode.

[0032] When the current target modality is a behavioral interaction modality, the corresponding processing tools are used to extract the interaction context sequence from the live stream interaction data and viewer viewing behavior data corresponding to the behavioral interaction modality, and the interaction context sequence is determined as the data stream corresponding to the behavioral interaction modality. Live stream interaction data may include, but is not limited to, at least one of the following: liking, following, joining a fan club, sending gifts (gift type, value, and sending time), sharing the live stream, and participation records of initiating polls / questionnaires, etc. Viewer viewing behavior data may include, but is not limited to, at least one of the following: user entry / exit time, dwell time, current online users, frequency of bullet comments, and frequency of user comments, etc.

[0033] Therefore, the multimodal data stream obtained in this embodiment may include, but is not limited to, at least one of the following modal data streams: video frame sequence, audio waveform sequence, text sequence, and interactive context sequence.

[0034] 102. Perform feature extraction on the multimodal data stream to obtain a multimodal feature vector sequence.

[0035] In order to effectively explore the correlation patterns between multimodal data streams and provide a data foundation for accurately identifying the broadcaster's intentions, the step of extracting features from the multimodal data streams to obtain a multimodal feature vector sequence is performed. The implementation process of this step may include the following steps 102A to 102B.

[0036] 102A. For each modal data stream in the multimodal data stream, perform the following: Select a feature extractor that matches the modal data stream, extract features from the modal data stream, and obtain the feature sequence corresponding to the modal data stream.

[0037] Specifically, different modal data streams have different characteristics. Therefore, in order to extract the feature sequence of each modal data stream more accurately, it is necessary to select a feature extractor that matches the modal attributes of the modal data stream to extract features from the modal data stream.

[0038] For example, such as Figure 2 As shown, Figure 2 A schematic diagram illustrating the multimodal feature vector sequence generation process is shown. From... Figure 2 As can be seen from step 101, the multimodal data streams acquired include the following modal data streams: video frame sequences, audio waveform sequences, and text sequences. For the video frame sequences, a feature extractor, such as a "visual encoder" (e.g., Vision Transformer), is used to extract features from the video frame sequences, resulting in the feature sequence "video feature vt". For the audio waveform sequences, an auditory encoder, such as a Conformer, is used to extract features from the audio waveform sequences, resulting in the feature sequence "audio feature at". For the text sequences, a text encoder, such as BERT, is used to extract features from the text sequences, resulting in the feature sequence "text feature lt".

[0039] 102B. By using a multimodal cross-attention mechanism, the feature sequences corresponding to various modal data streams are fused to obtain a multimodal feature vector sequence.

[0040] Specifically, after extracting the feature sequences of various modal data streams, in order to effectively mine the correlation patterns between multimodal data streams, such as... Figure 2 As shown, a multimodal cross-modal attention mechanism is needed to fuse the feature sequences corresponding to various modal data streams to obtain a multimodal feature vector sequence.

[0041] Specifically, the process of fusing feature sequences corresponding to various modal data streams through a multimodal cross-attention mechanism to obtain a multimodal feature vector sequence can include the following steps: In the cross-attention layer (used to execute the multimodal cross-attention mechanism), the feature sequence of each modality is used as a query vector, while the features of all other modalities are used together as a key vector and a value vector. Attention weights are obtained by calculating the similarity between the query vector and the key vector, and the value vectors are weighted and aggregated using the attention weights, thereby fusing information from other modalities. After completing the bidirectional interaction between all modalities through the cross-attention layer, a series of updated feature sequences injected with cross-modal information are obtained (such as text features rich in visual information, visual features incorporating text semantics, etc.). These updated multimodal features are then aggregated into a temporary joint representation by concatenation or element-wise addition to capture the initial fusion of the global context. Subsequently, this joint representation is fed into a fully connected feedforward neural network layer. This layer enhances the representational power of the features by performing high-order transformations and deep abstractions on the features through nonlinear activation functions (such as ReLU or GELU). It also projects and normalizes the feature dimensions, mapping them to a semantic space aligned with all modal features and with a uniform distribution, and outputs a fixed-length sequence of feature vectors, which is the multimodal feature vector sequence. Each vector in the multimodal feature vector sequence becomes a dense representation carrying deep semantic associations and complementary information from all modalities, providing a directly usable and information-complete multimodal fusion foundation for subsequent intent recognition.

[0042] Specifically, the following example illustrates the process: The feature sequences of each modality are sequentially used as the query vector, while the features of all other modalities are collectively used as the key and value vectors. Attention weights are obtained by calculating the similarity between the query vector and the key vector, and these attention weights are then used to weighted aggregate the value vectors, thus fusing information from other modalities. For example, the text modality's feature sequence is used as the query vector, and the audio and video modality's feature sequences are concatenated together to form the key and value vectors. The query vector is then used to calculate the similarity (attention weight) with the key vector (also concatenated from the audio and video modality's feature sequences), and this weight is then used to weighted sum the value vector (also concatenated from the audio and video modality's feature sequences. The resulting new feature sequence contains the information extracted from the audio and video modality's feature sequences that is most relevant to the text content of the text modality's feature sequence.

[0043] 103. Call the pre-trained variational state space model to map the multimodal feature vector sequence to the latent space, generate a sequence of latent variables representing the state of the live streamer, and use the deentanglement capability learned by the variational state space model in the counterfactual training in the live streaming scenario to deentangle the latent variable sequence to obtain a pure intention factor that is not affected by the live streaming environment. The pure intention factor is used to indicate the subjective intention of the live streamer.

[0044] In some embodiments, the variational state space model is a pre-trained model and a time-series-based model. It has learned deentanglement capabilities during counterfactual training in a live streaming scenario. Based on the deentanglement capability, it can effectively extract the live streaming environment variables from the multimodal feature vector sequence of the live streaming room, thereby deentanglementing the pure intention factor that is not affected by the live streaming environment and is used to indicate the subjective intention of the live streaming host.

[0045] In some embodiments, the specific implementation process of step 103 may include the following steps: The multimodal feature vector sequence is input into a pre-trained variational state space model (VSSM). The variational state space model learns the temporal dependence and probability distribution of the multimodal feature vector sequence through an encoder network (such as a combination of GRU and a probabilistic layer), mapping it to a low-dimensional continuous latent space to generate a sequence of latent variables following a specific probability distribution (such as a Gaussian distribution). Simultaneously, variational inference optimizes the lower bound of evidence (ELBO) to balance the reconstruction accuracy and distribution rationality of the latent variables. Based on this, utilizing the structured latent space characteristics and deentanglement constraints of the variational state space model, factorization is performed on the latent variable sequence to separate interference factors (i.e., environmental factors, such as live stream noise) unrelated to the streamer's intention and pure intention factors directly related to the streamer's intention. Furthermore, after extracting the pure intention factors, the reconstruction effectiveness of the intention factors can be verified through a decoder network to ensure that the deentangled intention factors accurately reflect the streamer's true subjective intention and possess cross-modal generalization ability.

[0046] In some embodiments, the deentangled pure intent factor refers to the abstract intent representation decoupled from the multimodal feature vector sequence that only represents the true subjective intent of the live streamer and is independent of the live stream environment. The live stream environment refers to the visual and environmental features in the live stream scene that are independent of the streamer's intent and mainly affect the audience's perception and atmosphere experience. It systematically shapes and conveys the overall aesthetics and emotional tone of the live stream through decoupled visual elements such as color tone, set design, props, and lighting.

[0047] 104. Based on the pure intent factor obtained from the deentanglement, the intent recognition result of the live broadcaster in the current time window is obtained by classifying the intent factors through an intent classifier.

[0048] In some embodiments, the unentangled pure intent factor is a feature closely related to the broadcaster's intent and unaffected by the style of the live stream. Based on this pure intent factor, the intent classifier can more clearly and stably distinguish different intent categories, and ultimately accurately obtain the intent recognition result of the live streamer within the current time window. Based on this, the step of classifying the intent recognition result of the live streamer within the current time window using the intent classifier based on the unentangled pure intent factor is executed. The implementation process of this step may include the following steps 104A to 104C.

[0049] 104A. Based on the target intent factor obtained from the deentanglement, the probability of the live streamer belonging to each intent category within the current time window is identified by the intent classifier.

[0050] Specifically, the de-entangled pure intent factors (such as intent factor vectors) are input into a pre-trained intent classifier (such as a fully connected neural network or support vector machine). The intent classifier performs forward computation based on these pure intent factors and outputs the probability of the live streamer belonging to each predefined intent category.

[0051] It should be noted that the intent classifier can be part of the variational state space model or it can be an entity that exists independently of the variational state space model; this embodiment does not impose any limitations on this. Furthermore, since entity intent recognition is a multi-classification problem, the intent classifier may include a Softmax layer, which outputs the probability of the live streamer belonging to each intent category.

[0052] 104B. If the maximum attribution probability is lower than the preset threshold, an intent recognition result indicating that the host's intent in the live broadcast room is unclear within the current time window will be generated, and an intent unclear prompt will be issued.

[0053] If the maximum attribution probability is lower than the preset threshold, it indicates that the intention of the live streamer cannot be accurately determined at present, and the intention of the live streamer is unclear. Therefore, an intent recognition result indicating that the intention of the live streamer is unclear within the current time window is generated, and an intent ambiguity prompt is issued to inform the supervisors of this situation. The preset threshold can be flexibly set based on business needs, and this example does not limit it. For example, the preset threshold is 0.6.

[0054] 104C. If the maximum attribution probability is not lower than the preset threshold, the target intent category corresponding to the maximum attribution probability is determined, and an intent recognition result indicating that the live streamer's intent belongs to the target intent category within the current time window is generated.

[0055] If the maximum attribution probability is not lower than a preset threshold, it indicates that the streamer's intent is clear and belongs to the intent category corresponding to the maximum attribution probability. Therefore, the target intent category corresponding to the maximum attribution probability is determined, and intent recognition results indicating that the streamer's intent belongs to the target intent category within the current time window are generated for subsequent risk control and compliance governance. These intent recognition results can carry intent tags corresponding to the target intent category to reflect the streamer's intent.

[0056] For example, the preset intent categories include at least two of the following: trust building, pain point stimulation, product introduction, and limited-time promotion. The maximum attribution probability is determined to be no less than a preset threshold of 0.6, and the target intent category corresponding to the maximum attribution probability is determined to be product introduction. Therefore, an intent recognition result is generated indicating that the live streamer's intent within the current time window belongs to the product introduction intent category. The intent recognition result carries an intent tag for the product introduction intent category.

[0057] The broadcaster intent recognition method provided in this application, when it is determined that intent recognition of the broadcaster in the live room is required, acquires the multimodal data stream of the current time window of the live room in real time within a preset time window, and extracts features from the multimodal data stream to obtain a multimodal feature vector sequence. Then, a pre-trained variational state space model is invoked to map the multimodal feature vector sequence to the latent space, generating a latent variable sequence representing the broadcaster's state. Utilizing the deentanglement capability learned by the variational state space model during counterfactual training in a live room scenario, a pure intent factor, unaffected by the live room environment and used to indicate the broadcaster's subjective intent, is deentangled based on the latent variable sequence. Finally, based on the deentangled pure intent factor, an intent classifier is used to classify and obtain the intent recognition result of the broadcaster in the current time window. Therefore, the solution provided in this embodiment has at least the following specific effects: First, based on the multimodal data stream of the live broadcast room, the latent variables mixed in the multimodal data stream (such as the broadcaster's intention, environmental interaction and random interference) are effectively de-entangled through a pre-trained variational state space model, thereby accurately extracting the pure intention factor that represents the broadcaster's true intention, and then accurately identifying the broadcaster's intention based on the pure intention factor; Second, the de-entangled intention factor can serve as an interpretable feature to assist the live broadcast platform in key operations such as risk control and compliance governance, marketing intervention and providing personalized recommendations to users.

[0058] In some embodiments of this application, the variational state space model used in step 102 above is obtained through pre-training. Therefore, the anchor intent recognition method provided in this embodiment may also include a variational state space model training process, which can be implemented through the following steps 201 to 205.

[0059] 201. Determine the target variational state space model to be trained.

[0060] The target variational state space model is the model to be trained. It can be an unused initial variational state space model or a used variational state space model that needs to be optimized to be suitable for the live broadcast room in step 101. This embodiment does not limit this.

[0061] 202. Obtain the multimodal sample data stream from the sample live broadcast room.

[0062] The sample live streaming room can exist in two forms: First, the sample live streaming room is the same as the live streaming room in step 101. Second, the sample live streaming room and the live streaming room in step 101 belong to the same type, such as both being live streaming e-commerce live streaming rooms.

[0063] The multimodal sample data stream is the multimodal sample data stream obtained within the target time window. The specific execution process for obtaining the multimodal sample data stream from the live sample room is basically the same as step 101 above, so it will not be repeated here. In addition, it should be noted that the target time window can be the same as or different from the preset time window in step 101, and this embodiment does not limit this.

[0064] 203. Perform feature extraction on the multimodal sample data stream to obtain the multimodal sample feature vector sequence.

[0065] The specific execution process of step 203 is basically the same as that of step 102 above, so it will not be repeated here.

[0066] 204. Based on counterfactual consistency loss, the target variational state space model is iteratively trained using multimodal sample feature vector sequences until the intent factor identified by the trained target variational state space model does not change due to changes in the sample live broadcast environment.

[0067] The purpose of step 204 is to enable the target variational state space model to learn the deentanglement capability during counterfactual training in the live streaming scenario. Based on the deentanglement capability, the model can effectively extract the live streaming environment variables from the multimodal feature vector sequence of the live streaming room, thereby deentanglementing the pure intention factor that is not affected by the live streaming environment and is used to indicate the true subjective intention of the live streaming host.

[0068] The implementation process of step 204 may include: iteratively training the target variational state space model using a sequence of multimodal sample feature vectors, and executing steps 204A to 204D for each iteration.

[0069] 204A. Call the target variational state space model trained in the current iteration to map the multimodal sample feature vector sequence to the latent space, generate the target latent variable sequence, and use the target variational state space model to deentangle the intention factor and the factual environment factor based on the target latent variable sequence.

[0070] Intent factors represent the subjective intent of the livestream host, while factual environment factors represent information such as the livestream environment. By forcibly separating these two types of factors through variational inference, the target variational state-space model can learn independent and semantically clear latent variable representations and strip away the livestream environment variables, thereby untangling and extracting pure intent factors that are unaffected by the livestream environment and used to indicate the subjective intent of the livestream host.

[0071] In some embodiments, the implementation process of step 204A may include at least the following steps 204A1 to 204A3.

[0072] 204A1. Multiple sampling times are determined based on a preset duration.

[0073] The preset duration is shorter than the target time window duration. Multiple sampling times are determined based on the preset duration so that a time-based target latent variable sequence can be formed based on the latent variables at each sampling time.

[0074] 204A2. For each sampling time, execute the following steps 204A21 to 204A23 in sequence.

[0075] 204A21. Obtain the latent variables inferred from the previous sampling time, and obtain the observation data at the current sampling time based on the multimodal sample feature vector sequence.

[0076] If the current sampling time is the first time step, the preset initial latent variable is obtained as the latent variable inferred from the previous sampling time. If the current sampling time is not the first sampling time, the latent variable inferred from the previous sampling time is directly obtained. The purpose of obtaining the latent variable inferred from the previous sampling time is to ensure that the subsequently generated target latent variable sequence maintains temporal continuity. The observation data at the current time step is the multimodal feature vector corresponding to the current sampling time t in the multimodal sample feature vector sequence.

[0077] 204A22. The inference network in the target variational state space model outputs the posterior distribution parameters of the latent variables at the current sampling time based on the latent variables at the previous sampling time and the observation data at the current sampling time. The posterior distribution parameters include the mean and variance.

[0078] The specific implementation process of step 204A22 may include: as follows Figure 3 As shown, Figure 3The diagram illustrates the formation process of the target latent variable sequence. First, the latent variable zt−1 from the previous sampling time is concatenated with the observed data xt at the current time, or combined through a feature fusion layer, to form a fusion vector representing the joint characteristics. This fusion vector is then input into an inference network composed of a multi-layer feedforward neural network or a recurrent neural network. The inference network extracts features layer by layer through nonlinear transformations, and finally, at the output layer, two independent linear mapping layers generate the mean μt and variance σt of the posterior Gaussian distribution, respectively. 2 (e.g., logarithmic variance), where the mean represents the most likely position of the current state, and the variance represents the uncertainty of the current state. This process enables the inference network to effectively extract information from historical states and current observations, approximating the true posterior distribution qϕ(zt |z(t-1), xt), thus providing differentiable probability distribution parameters for subsequent reparameterized sampling. In the variational state-space model, qϕ(zt |z(t-1), xt) represents a conditional probability distribution parameterized by parameter ϕ. Here, ϕ is a learnable parameter of the inference network (i.e., the encoder), which takes the latent variable zt−1 from the previous sampling time and the observed data xt from the current time as input, and outputs the latent variable posterior distribution parameters of the latent variable zt at the current sampling time, such as the mean and variance N(μt, σt) following a Gaussian distribution. 2 ).

[0079] 204A23. The latent variables at the current sampling time are obtained by sampling from the posterior distribution parameters through a reparameterization sampling mechanism.

[0080] The specific implementation process of step 204A23 may include: firstly, sampling a random noise vector ε ∼ N(0, I) of the same dimension as the latent variable from a standard normal distribution; then, using the posterior distribution parameters of the latent variable output by the inference network (mean μt and variance σt) 2 This method combines the posterior distribution parameters of the latent variables with the random noise vector through a differentiable linear transformation zt = μt + σt⊙ε, where ⊙ denotes element-wise multiplication. The key to this operation is transferring the randomness of the sampling process from the latent variable zt at the current sampling time to the external noise variable ε, so that the gradient of the latent variable zt at the current sampling time with respect to the network parameters ϕ can be obtained through the mean μt and variance σt. 2 The path is propagated stably, and ε is treated as a constant, thus enabling gradient-based end-to-end optimization while maintaining the characteristics of the probability distribution. The resulting latent variable zt at the current sampling time contains both the deterministic characteristics of the observed data and retains the random uncertainty of the probability distribution.

[0081] 204A3. Generate the target latent variable sequence based on the latent variables at each sampling time.

[0082] like Figure 3 As shown, the target latent variable sequence is generated by splicing or combining the latent variables at each sampling time according to the temporal order of each sampling time.

[0083] In some embodiments, after calling the target variational state space model trained in the current iteration to map the multimodal sample feature vector sequence to the latent space to generate the target latent variable sequence, the step of using the target variational state space model to deentangle the intention factor and the factual environment factor based on the target latent variable sequence is performed. The specific implementation process of this step may include: Figure 4 As shown, Figure 4 The diagram illustrates the counterfactual training process. A deentanglement module is introduced into the prior structure of the target variational state space model. Then, the intention factor and the factual environment factor are deentangled based on the target latent variable sequence through the deentanglement mechanism in the deentanglement module.

[0084] 204B. Construct counterfactual samples based on factual environmental factors and determine the counterfactual environmental factors of the counterfactual samples.

[0085] The specific implementation process of step 204B may include the following steps: Figure 4 As shown, causal intervention or counterfactual sampling methods are set through causal models or domain knowledge to change factual environmental factors to simulate counterfactual conditional data. Then, based on the counterfactual conditional data, corresponding counterfactual samples are generated using techniques such as generative models or structural causal models to ensure that other characteristics remain unchanged except for the intervention factors. In other words, the counterfactual samples are samples of anchors saying the same things and doing the same things in completely different background environments (e.g., from a noisy shopping mall to a quiet studio). Finally, the specific values ​​of the counterfactual environmental factors are clearly quantified according to the intervention settings.

[0086] For example, in a live streaming scenario, factual environmental factors include: an outdoor sales area with high ambient noise. Intent factors include: the host demonstrating product features and using promotional phrase 1. The constructed counterfactual sample retains the intent factors unchanged, i.e., the host demonstrating product features and using promotional phrase 1, but no longer retains the factual environmental factors. Instead, it replaces them with environmental factors of: an indoor studio with ambient noise below the noise threshold.

[0087] 204C. Based on intent factors and factual environment factors, intent recognition is performed to obtain the first intent prediction result of the sample live room anchor under factual environment, and based on intent factors and counterfactual environment factors, intent recognition is performed to obtain the second intent prediction result of the sample live room anchor under counterfactual environment.

[0088] Specifically, such as Figure 4As shown, intent prediction is performed under both the factual path and the counterfactual path. Specifically, under the factual path, the intent factor and the factual environment factor are concatenated, and the concatenated result is input into the intent encoder. The intent encoder then obtains the first intent prediction result of the sample live stream host under the factual environment. This can be represented as: Prediction factual =Decoder(Z intent C factual ), where Prediction factual Z represents the predicted first intention of the streamer in the sample live broadcast room under real-world conditions. intent Indicates the intention factor, C factual This represents the factual environmental factors. Under the counterfactual path, the intent factors and counterfactual environmental factors are concatenated, and the concatenated result is input into the intent encoder. The intent encoder then obtains the second intent prediction result of the sample live stream host under the counterfactual environment. Specifically, this can be represented as: Prediction counterfactual =Decoder(Z intent C counterfactual ), where Prediction counterfactual Z represents the predicted second intention of the livestream host in a counterfactual environment. intent Indicates the intention factor, C counterfactual This indicates a counterfactual environmental factor.

[0089] 204D. If the total loss of the target variational state space model after the current iteration of training is found to converge, and the counterfactual consistency loss of the trained target variational state space model is determined to converge based on the first intention prediction result and the second intention prediction result, then the target variational state space model after the current iteration of training is determined to be the pre-trained variational state space model. If the total loss of the target variational state space model after the current iteration of training is not found to converge, or if the total loss of the target variational state space model after the current iteration of training is found to converge, and the counterfactual consistency loss of the target variational state space model after the current iteration of training is determined to converge based on the first intention prediction result and the second intention prediction result, then the parameters of the target variational state space model after the current iteration of training are adjusted, and the target variational state space model with adjusted parameters is used as the target variational state space model for the next iteration of training.

[0090] In some embodiments, to determine whether the target variational state space model trained in the current iteration can effectively isolate the live-streaming environment variables, thereby untangling the pure intent factor that is unaffected by the live-streaming environment and used to indicate the subjective intent of the live-streaming host, the following judgment step needs to be performed: the square of the difference between the first intent prediction result and the second intent prediction result is determined as the counterfactual consistency loss of the target variational state space model trained in the current iteration. This process can be expressed by the following formula: L CF =||Prediction factual - Prediction counterfactual || 2 L CF Prediction represents the loss of counterfactual consistency. factual Prediction represents the prediction result of the first intention of the streamer in the sample live broadcast room under real-world conditions. counterfactual This represents the predicted second intention of the livestream host in a counterfactual environment.

[0091] After determining the counterfactual consistency loss, the step of determining whether the counterfactual consistency loss is the minimized loss is performed. This step can be implemented in two ways: One is to determine whether the counterfactual consistency loss is less than a first loss threshold. If it is less, the counterfactual consistency loss is determined to be the minimized loss; if it is not less, the counterfactual consistency loss is determined to be the non-minimized loss. The other is to determine whether the difference between the counterfactual consistency losses of the most recent N consecutive training iterations is less than a difference threshold, and whether the counterfactual consistency loss is less than a second loss threshold. If so, the counterfactual consistency loss is determined to be the minimized loss; if not, the counterfactual consistency loss is determined to be the non-minimized loss. At least one of these two methods can be flexibly selected based on business needs; this embodiment does not limit this selection.

[0092] If the counterfactual consistency loss is minimized, it means that the target variational state space model trained in the current iteration has learned to make consistent judgments about intentions regardless of changes in the environment. Therefore, it is determined that the counterfactual consistency loss of the target variational state space model trained in the current iteration has converged.

[0093] After assessment, if the counterfactual consistency loss is not minimized, it indicates that the target variational state space model trained in the current iteration has not yet learned the ability to consistently judge intent regardless of changes in the environment, and still needs further optimization. Therefore, it is determined that the counterfactual consistency loss of the target variational state space model trained in the current iteration has not converged.

[0094] In some embodiments, the total loss of the variational state-space model is a joint optimization objective that balances the accuracy of posterior inference of latent variables and the quality of reconstruction of observed data. It typically uses the Evidence Lower Bound (ELBO) as its core, while incorporating constraints such as deentanglement regularization terms to optimize both the overall model's fitting ability and the separation effect of latent factors. The counterfactual loss of the variational state-space model is a confirmatory loss targeting the effectiveness of factor deentanglement. By fixing the intention factor and the intervention environment factor and then reconstructing the corresponding output, it measures the ability of the intervention factor to independently drive the generation of results, thereby strengthening the independence between different factors. Therefore, the convergence of the total loss indicates that the model has achieved accurate fitting of the observed data and effective inference of the posterior distribution of the latent variables, ensuring the model's ability to represent the original data and the rationality of the overall architecture. The convergence of the counterfactual loss proves that the model has completed the effective deentanglement of the intention factors and the factual environment factors. The two types of factors have independent causal driving capabilities. In other words, the convergence of the total loss ensures that the model "can learn information from the data", and the convergence of the counterfactual loss ensures that the model "can decompose the learned information into mutually independent factors". Both are indispensable. Only when they converge simultaneously can it be said that the model has both good representation performance and has achieved the core training goal of factor deentanglement, thus meeting the task requirement of separating the intention factors from the latent variable sequence.

[0095] Specifically, if the total loss of the target variational state space model after the current iteration of training converges, and the counterfactual consistency loss of the target variational state space model after training converges based on the first intention prediction result and the second intention prediction result, it indicates that the target variational state space model after the current iteration of training has been trained and does not need to be optimized further. Therefore, the target variational state space model after the current iteration of training is determined as the pre-trained variational state space model for use in step 103.

[0096] Specifically, if the total loss of the target variational state space model after the current iteration of training is not converged, or if the total loss of the target variational state space model after the current iteration of training is converged, and the counterfactual consistency loss of the target variational state space model after the current iteration of training is not converged based on the first intention prediction result and the second intention prediction result, it indicates that the target variational state space model after the current iteration of training still needs to be optimized and trained. Therefore, the parameters of the target variational state space model after the current iteration of training are adjusted (e.g., encoder parameters, unentanglement module parameters, inference network parameters, etc., which will not be listed here), and the target variational state space model after the parameter adjustment is used as the target variational state space model for the next iteration of training.

[0097] 205. Use the trained target variational state space model as the pre-trained variational state space model.

[0098] The trained target variational state space model already has a good ability to recognize the broadcaster's intention. Therefore, the trained target variational state space model is used as a pre-trained variational state space model for use in step 103 above.

[0099] In some embodiments of this application, during processes 204A1 to 204A3, the broadcaster intent recognition method provided in this embodiment may further include performing steps one to three for each sampling time. The purpose of performing steps one to three is to ensure that the model can accurately capture the dynamic generation process of the observed sequence. Calculating the KL divergence between the posterior distribution parameters and prior distribution parameters of the latent variables at the current time is to regularize the learning process of the latent variables, making the posterior distribution as close as possible to the prior distribution, thereby balancing the accuracy of data fitting (reflected through reconstruction loss) and the regularity of latent variable representation (constrained by KL divergence), avoiding overfitting and improving the model's generalization ability, while implicitly encouraging the model to learn more interpretable and smooth temporal dynamic features.

[0100] Step 1: Perform the following steps for each sampling time: Using the prior network in the target variational state space model trained in the current iteration, output the prior distribution parameters of the latent variables at the current time based on the latent variables at the previous sampling time; reconstruct the observation data at the current sampling time based on the latent variables at the current sampling time; determine the reconstruction loss between the observation data and the reconstructed observation data; determine the KL divergence between the posterior distribution parameters of the latent variables at the current sampling time and the prior distribution parameters of the latent variables; and calculate the lower bound loss of evidence at the current sampling time based on the reconstruction loss and the KL divergence.

[0101] For any given sampling time, the specific implementation of step one can include: the prior network takes the latent variable zt−1 from the previous time step as input and outputs the prior distribution parameters of the latent variable zt at the current time step (such as the mean and variance of a Gaussian distribution), i.e., p(zt|zt−1). Simultaneously, the inference network outputs the posterior distribution parameters of the latent variable zt based on the current observed data xt, i.e., q(zt|xt), and samples the observed data zt at the current sampling time from this posterior distribution. Subsequently, the inference network, i.e., the decoder, reconstructs the observed data xt' using the sampled observed data zt, and compares the original observed xt with the reconstructed observed xt' (e.g., ...). Figure 3The reconstruction loss is calculated using variance and mean, which ensures the model's ability to reconstruct data details. Finally, the KL divergence between the posterior distribution q(zt|xt) and the prior distribution p(zt|zt−1) of the latent variables at the same time step is calculated. This KL divergence acts as a regularization loss, forcing the posterior distribution to approach the prior distribution predicted by the temporal dynamics, thus jointly constraining the model to maintain the temporal coherence and rationality of the latent state sequence while learning the data generation patterns. Then, the lower bound of evidence loss at the current sampling time is calculated using the following formula based on the reconstruction loss and KL divergence.

[0102] LELBO = Eq [logp(xt│zt)] - D KL (q(zt│xt ) ||p(zt│z(t-1) )) Where LELBO represents the lower bound loss of evidence at the current sampling time, Eq[logp(xt│zt)] represents the reconstruction loss, and D KL (q(zt│xt ) ||p(zt│z(t-1) )) represents the KL divergence.

[0103] Step 2: Summarize the lower bound loss of evidence at each sampling time to obtain the total lower bound loss of evidence.

[0104] The total evidence lower bound loss reflects the model's overall goodness of fit to the data generation process and time dynamics.

[0105] Step 3: If the total evidence lower bound loss does not converge, optimize the parameters of the target variational state space model trained in the current iteration by combining backpropagation and gradient descent until the total evidence lower bound loss converges.

[0106] If the total evidence lower bound loss does not converge, it indicates that the target variational state space model still needs optimization. Therefore, it is necessary to jointly optimize the parameters of the target variational state space model trained in the current iteration through backpropagation and gradient descent until the total evidence lower bound loss converges. It should be noted that if the above-mentioned 204D monitors that the total loss of the target variational state space model trained in the current iteration converges, and determines that the counterfactual consistency loss of the trained target variational state space model converges based on the first and second intention prediction results, but this step three monitors that the total evidence lower bound loss has not converged, then it is still not certain that the model training is complete. In this case, it is still necessary to continue optimizing the parameters of the target variational state space model trained in the current iteration through backpropagation and gradient descent until the total evidence lower bound loss converges, the total loss of the target variational state space model trained in the current iteration converges, and determines that the counterfactual consistency loss of the trained target variational state space model converges based on the first and second intention prediction results, only then can it be determined that the model training is complete.

[0107] Furthermore, one embodiment of this application also provides a broadcaster intent recognition device, such as... Figure 5 As shown, the broadcaster intent recognition device provided in this embodiment may include at least: The acquisition module 31 is used to acquire the multimodal data stream of the live broadcast room in real time within a preset time window. Extraction module 32 is used to extract features from the multimodal data stream to obtain a multimodal feature vector sequence; The identification module 33 is used to call the pre-trained variational state space model to map the multimodal feature vector sequence to the latent space, generate a latent variable sequence representing the state of the live streamer, and use the deentanglement capability learned by the variational state space model in the counterfactual training in the live streaming scenario to deentangle the latent variable sequence to obtain a pure intention factor that is not affected by the live streaming environment. The pure intention factor is used to indicate the subjective intention of the live streamer. The generation module 34 is used to classify the intent recognition results of the live broadcaster in the current time window based on the pure intent factor obtained by deentanglement through an intent classifier.

[0108] The broadcaster intent recognition device provided in this application, when determining that intent recognition of the broadcaster in a live broadcast room is necessary, acquires the multimodal data stream of the current time window of the live broadcast room in real time within a preset time window, and extracts features from the multimodal data stream to obtain a multimodal feature vector sequence. Then, it calls a pre-trained variational state space model to map the multimodal feature vector sequence to a latent space, generating a latent variable sequence representing the broadcaster's state. Utilizing the deentanglement capability learned by the variational state space model during counterfactual training in a live broadcast scenario, it deentangles the latent variable sequence to extract a pure intent factor that is unaffected by the live broadcast environment and indicates the broadcaster's subjective intent. Finally, based on the deentangled pure intent factor, an intent classifier classifies the broadcaster's intent recognition result within the current time window. Therefore, the solution provided in this embodiment has at least the following specific effects: First, based on the multimodal data stream of the live broadcast room, the latent variables mixed in the multimodal data stream (such as the broadcaster's intention, environmental interaction and random interference) are effectively de-entangled through a pre-trained variational state space model, thereby accurately extracting the pure intention factor that represents the broadcaster's true intention, and then accurately identifying the broadcaster's intention based on the pure intention factor; Second, the de-entangled intention factor can serve as an interpretable feature to assist the live broadcast platform in key operations such as risk control and compliance governance, marketing intervention and providing personalized recommendations to users.

[0109] In some embodiments of this application, such as Figure 6As shown, the multimodal data stream acquired by the acquisition module 31 includes at least one of the following modal data streams: video frame sequence, audio waveform sequence, text sequence, and interaction context sequence.

[0110] In some embodiments of this application, such as Figure 6 As shown, the extraction module 32 is specifically used to perform the following for each modal data stream in the multimodal data stream: select a feature extractor that matches the modal data stream, extract features from the modal data stream, and obtain the feature sequence corresponding to the modal data stream; and fuse the feature sequences corresponding to various modal data streams through a multimodal cross-attention mechanism to obtain a multimodal feature vector sequence.

[0111] In some embodiments of this application, such as Figure 6 As shown, the generation module 34 is specifically used to identify the probability of the live streamer belonging to each intent category within the current time window based on the unentangled target intent factor through an intent classifier; if the maximum belonging probability is lower than a preset threshold, an intent recognition result indicating that the live streamer's intent is unclear within the current time window is generated, and an intent unclear prompt is issued; if the maximum belonging probability is not lower than the preset threshold, the target intent category corresponding to the maximum belonging probability is determined, and an intent recognition result indicating that the live streamer's intent belongs to the target intent category within the current time window is generated.

[0112] In some embodiments of this application, such as Figure 6 As shown, the broadcaster intent recognition device provided in this embodiment may further include: Processing module 35 is used to determine the target variational state space model to be trained; acquire the multimodal sample data stream of the sample live broadcast room; and extract features from the multimodal sample data stream to obtain a multimodal sample feature vector sequence. Training module 36 is used to iteratively train the target variational state space model based on the counterfactual consistency loss using the multimodal sample feature vector sequence until the intent factor identified by the trained target variational state space model does not change due to changes in the sample live broadcast environment. The determination module 37 is used to use the trained target variational state space model as the pre-trained variational state space model.

[0113] In some embodiments of this application, such as Figure 6 As shown, training module 36 may include: Calling unit 361 is used to iteratively train the target variational state space model using the multimodal sample feature vector sequence, and each iteration performs the following steps: calling the target variational state space model trained in the current iteration to map the multimodal sample feature vector sequence to the latent space, generating the target latent variable sequence, and using the target variational state space model to deentangle the intention factor and the factual environment factor based on the target latent variable sequence; The determining unit 362 is used to construct counterfactual samples based on the factual environmental factors and to determine the counterfactual environmental factors of the counterfactual samples; The identification unit 363 is used to identify intent based on the intent factor and the factual environment factor to obtain a first intent prediction result of the sample live broadcast room anchor under the factual environment, and to identify intent based on the intent factor and the counterfactual environment factor to obtain a second intent prediction result of the sample live broadcast room anchor under the counterfactual environment. The first determination unit 364 is used to determine the target variational state space model after the current iteration of training as the pre-trained variational state space model if the total loss of the target variational state space model after the current iteration of training is found to converge, and the counterfactual consistency loss of the target variational state space model after training is determined to converge based on the first intention prediction result and the second intention prediction result. The second determination unit 365 is used to adjust the parameters of the target variational state space model after the current iteration of training if it is detected that the total loss of the target variational state space model after the current iteration of training has not converged, or if it is detected that the total loss of the target variational state space model after the current iteration of training has converged, and it is determined based on the first intention prediction result and the second intention prediction result that the counterfactual consistency loss of the target variational state space model after the current iteration of training has not converged. The adjusted target variational state space model is then used as the target variational state space model for the next iteration of training.

[0114] In some embodiments of this application, such as Figure 6 As shown, the calling unit 361 is specifically used to determine multiple sampling times based on a preset duration; for each sampling time, it sequentially performs the following: obtaining the latent variables inferred from the previous sampling time, and obtaining the observation data of the current sampling time based on the multimodal sample feature vector sequence; outputting the posterior distribution parameters of the latent variables of the current sampling time through the inference network in the target variational state space model based on the latent variables of the previous sampling time and the observation data of the current sampling time, wherein the posterior distribution parameters include the mean and variance; sampling the latent variables of the current sampling time from the posterior distribution parameters through a reparameterized sampling mechanism; and generating a target latent variable sequence based on the latent variables of each sampling time.

[0115] In some embodiments of this application, such as Figure 6 As shown, the training module 36 may further include: a judgment unit 366, used to determine the square of the difference between the first intention prediction result and the second intention prediction result as the counterfactual consistency loss of the target variational state space model after the current iteration of training; if the counterfactual consistency loss is already the minimized loss, then it is determined that the counterfactual consistency loss of the target variational state space model after the current iteration of training has converged; if the counterfactual consistency loss is not the minimized loss, then it is determined that the counterfactual consistency loss of the target variational state space model after the current iteration of training has not converged.

[0116] In some embodiments of this application, such as Figure 6 As shown, calling unit 361 can also be used to perform the following for each sampling time: outputting the prior distribution parameters of the latent variables at the current time based on the latent variables at the previous sampling time through the prior network in the target variational state space model trained in the current iteration; reconstructing the observation data at the current sampling time based on the latent variables at the current sampling time; determining the reconstruction loss between the observation data and the reconstructed observation data; determining the KL divergence between the posterior distribution parameters of the latent variables and the prior distribution parameters of the latent variables at the current sampling time; calculating the lower bound loss of evidence at the current sampling time based on the reconstruction loss and the KL divergence; summing up the lower bound losses of evidence at each sampling time to obtain the total lower bound loss of evidence; if the total lower bound loss of evidence has not converged, then optimizing the parameters of the target variational state space model trained in the current iteration through backpropagation and gradient descent until the total lower bound loss of evidence converges.

[0117] For a detailed explanation of the various functional modules used in the operation of the broadcaster intent recognition device provided in this application embodiment, please refer to the corresponding detailed explanation of the broadcaster intent recognition method embodiment above, which will not be repeated here.

[0118] Furthermore, one embodiment of this application also provides a computer-readable storage medium, the storage medium including a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to execute the above-described broadcaster intent recognition method.

[0119] Furthermore, one embodiment of this application also provides an electronic device, the electronic device comprising: a memory for storing a program; and a processor coupled to the memory for running the program to perform the above-described anchor intent recognition method.

[0120] Furthermore, one embodiment of this application also provides a computer program product, the computer program product comprising: a computer program / computer executable instructions, the computer program / computer executable to perform the above-described anchor intent recognition method.

[0121] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0122] It is understood that the relevant features in the above methods and apparatus can be referenced interchangeably. Furthermore, the terms "first," "second," etc., in the above embodiments are used to distinguish between embodiments and do not represent the superiority or inferiority of any particular embodiment.

[0123] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0124] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, this application is not directed to any particular programming language. It should be understood that the content of this application described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing preferred embodiments of this application.

[0125] In addition, the memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0126] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data cutover device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data cutover device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0128] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data cutover device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0129] These computer program instructions can also be loaded onto a computer or other programmable data cutover device to cause a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0130] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0131] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0132] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0133] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0134] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0135] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for recognizing broadcaster intent, characterized in that, The method includes: The system acquires multimodal data streams from the live stream in real time within a preset time window. Feature extraction is performed on the multimodal data stream to obtain a multimodal feature vector sequence; The pre-trained variational state space model is invoked to map the multimodal feature vector sequence to the latent space, generating a sequence of latent variables representing the state of the live streamer. The unentanglement capability learned by the variational state space model in the counterfactual training in the live streaming scenario is utilized to unentangle the latent variable sequence to obtain a pure intent factor that is not affected by the live streaming environment. The pure intent factor is used to indicate the subjective intent of the live streamer. Based on the pure intent factor obtained from the deentanglement, the intent recognition result of the live streamer in the current time window is obtained by classifying it through an intent classifier.

2. The method according to claim 1, characterized in that, The process of extracting features from the multimodal data stream to obtain a multimodal feature vector sequence includes: for each modality in the multimodal data stream, performing the following steps: selecting a feature extractor that matches the modality data stream, extracting features from the modality data stream, and obtaining a feature sequence corresponding to the modality data stream; and fusing the feature sequences corresponding to various modality data streams through a multimodal cross-attention mechanism to obtain a multimodal feature vector sequence. And / or, Based on the de-entangled pure intent factor, the intent recognition result of the live streamer in the current time window is obtained by classifying through an intent classifier, including: based on the de-entangled target intent factor, the intent classifier identifies the probability of the live streamer in the current time window belonging to each intent category; if the maximum belonging probability is lower than a preset threshold, an intent recognition result indicating that the intent of the live streamer in the current time window is unclear is generated, and an intent unclear prompt is issued; if the maximum belonging probability is not lower than the preset threshold, the target intent category corresponding to the maximum belonging probability is determined, and an intent recognition result indicating that the intent of the live streamer in the current time window belongs to the target intent category is generated. And / or, The multimodal data stream includes at least one of the following modal data streams: video frame sequence, audio waveform sequence, text sequence, and interactive context sequence.

3. The method according to claim 1, characterized in that, The method further includes: Determine the target variational state-space model to be trained; Acquire multimodal sample data streams from the live sample room; Feature extraction is performed on the multimodal sample data stream to obtain a multimodal sample feature vector sequence; Based on counterfactual consistency loss, the target variational state space model is iteratively trained using the multimodal sample feature vector sequence until the intent factor identified by the trained target variational state space model does not change due to changes in the sample live broadcast environment. The trained target variational state space model is used as the pre-trained variational state space model.

4. The method according to claim 3, characterized in that, Based on counterfactual consistency loss, the target variational state space model is iteratively trained using the multimodal sample feature vector sequence until the intent factor identified by the trained target variational state space model does not change due to changes in the sample live streaming environment, including: The target variational state-space model is iteratively trained using the multimodal sample feature vector sequence, and the following steps are performed in each iteration: The target variational state space model trained in the current iteration is invoked to map the multimodal sample feature vector sequence to the latent space, generating a target latent variable sequence. The target variational state space model is then used to deentangle the intention factor and the factual environment factor based on the target latent variable sequence. Counterfactual samples are constructed based on the stated factual environmental factors, and the counterfactual environmental factors of the counterfactual samples are determined. Intent recognition is performed based on the intent factor and the factual environment factor to obtain the first intent prediction result of the sample live broadcast room anchor under the factual environment; and intent recognition is performed based on the intent factor and the counterfactual environment factor to obtain the second intent prediction result of the sample live broadcast room anchor under the counterfactual environment. If the total loss of the target variational state space model after the current iteration of training is found to converge, and the counterfactual consistency loss of the target variational state space model after training is determined to converge based on the first intention prediction result and the second intention prediction result, then the target variational state space model after the current iteration of training is determined to be the pre-trained variational state space model. If the total loss of the target variational state space model after the current iteration of training is not converged, or if the total loss of the target variational state space model after the current iteration of training is converged, and the counterfactual consistency loss of the target variational state space model after the current iteration of training is determined to be not converged based on the first intention prediction result and the second intention prediction result, then the parameters of the target variational state space model after the current iteration of training are adjusted, and the target variational state space model after the parameter adjustment is used as the target variational state space model for the next iteration of training.

5. The method according to claim 4, characterized in that, The process involves calling the target variational state space model trained in the current iteration to map the multimodal sample feature vector sequence to the latent space, generating a target latent variable sequence. This includes: determining multiple sampling times based on a preset duration; for each sampling time, sequentially performing the following steps: acquiring the latent variable inferred from the previous sampling time, and acquiring the observation data for the current sampling time based on the multimodal sample feature vector sequence; outputting the posterior distribution parameters of the latent variables for the current sampling time through the inference network in the target variational state space model based on the latent variables from the previous sampling time and the observation data for the current sampling time, wherein the posterior distribution parameters include mean and variance; sampling the latent variables for the current sampling time from the posterior distribution parameters through a reparameterized sampling mechanism; and generating the target latent variable sequence based on the latent variables for each sampling time. And / or, The method further includes: determining the square of the difference between the first intention prediction result and the second intention prediction result as the counterfactual consistency loss of the target variational state space model after the current iteration training; if the counterfactual consistency loss is already the minimum loss, then determining that the counterfactual consistency loss of the target variational state space model after the current iteration training has converged; if the counterfactual consistency loss is not the minimum loss, then determining that the counterfactual consistency loss of the target variational state space model after the current iteration training has not converged.

6. The method according to claim 5, characterized in that, The method further includes: for each sampling time step, performing the following steps: using the prior network in the target variational state space model trained in the current iteration to output the prior distribution parameters of the latent variables at the current time step based on the latent variables at the previous sampling time step; reconstructing the observation data at the current sampling time step based on the latent variables at the current sampling time step; determining the reconstruction loss between the observation data and the reconstructed observation data; determining the KL divergence between the posterior distribution parameters of the latent variables at the current sampling time step and the prior distribution parameters of the latent variables; calculating the lower bound loss of evidence at the current sampling time step based on the reconstruction loss and the KL divergence; summing up the lower bound losses of evidence at each sampling time step to obtain the total lower bound loss of evidence; if the total lower bound loss of evidence does not converge, then optimizing the parameters of the target variational state space model trained in the current iteration step step by jointly using backpropagation and gradient descent until the total lower bound loss of evidence converges.

7. A broadcaster intent recognition device, characterized in that, The device includes: The acquisition module is used to acquire the multimodal data stream of the live broadcast room in real time within a preset time window; The extraction module is used to extract features from the multimodal data stream to obtain a multimodal feature vector sequence; The identification module is used to call the pre-trained variational state space model to map the multimodal feature vector sequence to the latent space, generate a latent variable sequence representing the state of the live streamer, and use the deentanglement capability learned by the variational state space model in the counterfactual training in the live streaming scenario to deentangle the latent variable sequence to obtain a pure intention factor that is not affected by the live streaming environment. The pure intention factor is used to indicate the subjective intention of the live streamer. The generation module is used to classify the intent recognition results of the live streamer within the current time window based on the de-entangled pure intent factors through an intent classifier.

8. A computer-readable storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to execute the broadcaster intent recognition method according to any one of claims 1 to 6.

9. An electronic device, characterized in that, The electronic device includes: a memory for storing a program; and a processor coupled to the memory for running the program to perform the broadcaster intent recognition method according to any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes: a computer program / computer executable instructions, wherein the computer program / computer executable is the broadcaster intent recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-modal remote sensing data classification method based on linear time sequence selective state space model

    CN119152366A

  • Live broadcast room content identification and intelligent distribution method and system based on multi-modal fusion

    CN119377895A

  • Efficient time-varying anti-fact data analysis method and device based on state space model

    CN120526922A

  • Live broadcast behavior tracking system based on deep learning

    CN120708001A