Interaction method and device, computer equipment, storage medium and program product
By acquiring multimodal information for feature extraction and sentiment estimation, scene description text is generated to determine user intent. This solves the problem that existing intelligent agents cannot obtain complete emotional information about users, achieving more accurate intent understanding and emotional feedback, and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MIGU CO LTD
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-17
AI Technical Summary
Existing consultative agents cannot obtain complete emotional information from users, resulting in inaccurate understanding of intent, increased communication costs, and reduced user experience.
By acquiring multimodal information, including 3D video, text, and audio, feature extraction and sentiment estimation are performed to generate scene description text to determine user intent information. Deep learning models are then used for intent analysis and feedback.
It improves the accuracy of understanding user intent, enhances emotional resonance, provides more accurate and emotional feedback, and improves the user experience.
Smart Images

Figure CN121880494A_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence technology, and more particularly to an interaction method, apparatus, computer equipment, storage medium, and program product. Background Technology
[0002] Consultative agents are intelligent agents focused on providing professional consulting services. Through interaction with users, they autonomously perceive user needs, analyze information, and generate decision-making suggestions, offering precise consulting services in specific fields (such as law, medicine, and finance). Current consultative agents are typically implemented based on natural language processing (such as text or speech) and knowledge graph technologies, achieving user intent recognition and answer generation through dialogue with users.
[0003] However, current consulting agents mostly provide consulting services based on text and / or voice information input by users. Since text and voice information are too abstract, the agents cannot obtain complete emotional information from users, which leads to misunderstandings of user intentions, providing incorrect feedback to users, increasing communication costs with users, and resulting in a poor user experience. Summary of the Invention
[0004] This application provides an interaction method, apparatus, computer device, storage medium, and program product that can enhance emotional resonance and improve the accuracy of understanding user intentions.
[0005] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide an interaction method, the method comprising: Obtain multimodal information related to the inquiry, wherein the multimodal information includes at least one or more of three-dimensional video information, text information, and audio information; The scene description text is obtained based on the multimodal information, and the scene description text includes emotional information, which is used to describe the user's emotional state. Based on the scenario description text, determine the intent information of the inquiry, and output the corresponding response result based on the intent information.
[0006] In the above scheme, when the multimodal video includes the three-dimensional video information, the step of obtaining multimodal information related to the question consultation includes: sending a live broadcast request to the user's terminal, the live broadcast request being used to request the user to start a three-dimensional video live broadcast; and when the user confirms to start the three-dimensional video live broadcast, receiving the live broadcast stream through the terminal and obtaining the three-dimensional video information based on the live broadcast stream.
[0007] In the above scheme, when the terminal is not a three-dimensional terminal, the live stream is a two-dimensional live stream; the step of obtaining the three-dimensional video information based on the live stream includes: obtaining two-dimensional video information from the two-dimensional live stream, using a first model to perform depth estimation on each frame of two-dimensional images in the two-dimensional video information to obtain each frame of three-dimensional images, and obtaining the three-dimensional video information based on each frame of three-dimensional images; the first model is trained based on a high-definition edge image dataset.
[0008] In the above scheme, obtaining scene description text based on the multimodal information includes: extracting features from the three-dimensional video information to obtain three-dimensional feature information, and extracting features from other modal information besides the three-dimensional video information in the multimodal information to obtain objective feature information, wherein the three-dimensional feature information includes user three-dimensional feature information and / or object three-dimensional feature information related to the question consultation; performing sentiment estimation based on the three-dimensional feature information and the objective feature information to determine the user's sentiment information; and generating scene description text related to the question consultation based on the sentiment information, the user three-dimensional feature information and / or the object three-dimensional feature information, and the objective feature information.
[0009] In the above scheme, the step of performing sentiment estimation based on the three-dimensional feature information and the objective feature information to determine the user's sentiment information includes: extracting time-series features from the user's three-dimensional feature information, performing sentiment estimation based on the time-series features and the objective feature information to obtain the sentiment information, wherein the time-series features include: time information, key point three-dimensional location information, and key point quantity information.
[0010] In the above scheme, determining the intent information corresponding to the question based on the scenario description text includes: inputting the scenario description text into a second model, using the second model to perform intent analysis on the scenario description text, and outputting the intent information corresponding to the question.
[0011] In the above scheme, determining the intent information corresponding to the question consultation based on the scene description text includes: obtaining event description information and sentiment information related to the question consultation based on the scene description text; querying a knowledge base based on the event description information and sentiment information, wherein the knowledge base pre-stores one or more triples including event description information, sentiment information and intent information; if no intent information corresponding to the event description information and sentiment information is found in the knowledge base, inputting the scene description text into a second model, and using the second model to output the intent information corresponding to the question consultation.
[0012] Secondly, embodiments of the present invention also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described in the embodiments of the present invention.
[0013] Thirdly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the embodiments of the present invention.
[0014] Fourthly, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the method described in the embodiments of the present invention.
[0015] This invention provides an interactive method, apparatus, computer device, storage medium, and program product. By acquiring three-dimensional video information from a user's inquiry, it forms multimodal input information including three-dimensional video, audio, and text information. Using the three-dimensional video information and other modal information (i.e., objective information) from the multimodal input information, it generates scene description text containing emotional information. Based on this emotional scene description text, it infers the user's emotional intent. Because the three-dimensional video information in the acquired multimodal information includes scene depth information, it can capture millimeter-level depth changes, thereby accurately identifying subtle changes in facial expressions and body movements, enhancing emotional resonance, and improving the accuracy of understanding the user's intent. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the interaction method according to an embodiment of the present invention; Figure 2 This is a diagram illustrating the performance comparison of depth estimation algorithms according to embodiments of the present invention. Figure 1 ; Figure 3 This is a diagram illustrating the performance comparison of depth estimation algorithms according to embodiments of the present invention. Figure 2 ; Figure 4 This is a diagram illustrating the performance comparison of depth estimation algorithms according to embodiments of the present invention. Figure 3 ; Figure 5 This is a schematic diagram illustrating the process of obtaining scene description text according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the processing flow for obtaining response results based on multimodal information according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the module interaction of the interaction method according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the composition structure of the interactive device provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the hardware composition structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0019] The terms “first,” “second,” etc., used in the specification and claims of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0020] Before providing a detailed description of the technical solution of this application, a brief introduction to the existing related technical solutions will be given first.
[0021] Consultative intelligent agents are artificial intelligence (AI) systems based on large language models. Through multi-dimensional data analysis and autonomous decision-making capabilities, they provide users with professional advice. Currently, they are widely used in various fields such as financial investment consulting, health consultation, and recruitment consulting, achieving an upgrade from "passive response" to "proactive decision-making."
[0022] Current consultative agents are typically based on natural language processing (such as text or speech) and knowledge graph technology, and achieve user intent recognition and answer generation through dialogue with users.
[0023] However, since users currently primarily input questions into consultative AI agents via text or voice, these agents lack visual feedback and action coordination. For example, natural language conversational AI agents cannot obtain visual information such as user gestures, facial expressions, or physical demonstrations, which may lead to misunderstandings of the user's intentions and incorrect feedback, thus increasing communication costs. For instance, in a medical consultation scenario, a user with a sore throat cannot accurately express the specific location and symptoms of the pain through text or voice messages.
[0024] Furthermore, while text and voice messages provided by users to the AI agent can express emotional information such as tone and intonation, they cannot convey more nuanced emotional information such as micro-expressions (e.g., blinking frequency, mouth corner curvature) and body language. This results in overly simplistic emotional information being transmitted to the AI agent, making it difficult for the agent to capture the user's complete emotional information. For example, in after-sales scenarios, when a user complains to a consulting AI agent that "the clothes I bought were too expensive" using only text or voice, the AI agent cannot collect real-time micro-expressions such as "frowning and shaking the head," thus hindering its emotional analysis and making it unable to accurately understand whether the user intends a price reduction or a refund / exchange.
[0025] Based on this, in order to address the problem that current consulting agents are unable to obtain complete emotional information from users, resulting in inaccurate understanding of users' intentions, this invention proposes an interaction method. The method is applied to network devices that have deployed consulting agents and have live streaming capabilities, such as intelligent consulting systems or intelligent service platforms used in online medical services or online shopping platforms.
[0026] Figure 1 This is a flowchart illustrating the interaction method according to an embodiment of the present invention; as shown below. Figure 1 As shown, the method includes: Step 101: Obtain multimodal information related to the inquiry, wherein the multimodal information includes at least one or more of three-dimensional video information, text information, and audio information; Step 102: Obtain scene description text based on the multimodal information, wherein the scene description text includes emotional information, and the emotional information is used to describe the user's emotional state; Step 103: Determine the intent information of the inquiry based on the scenario description text, and output the corresponding response result based on the intent information.
[0027] In this embodiment, when a user inputs a question through a terminal, the network device obtains multimodal information related to the question submitted by the user through the terminal.
[0028] In this embodiment, the terminal is a device that supports real-time 3D video, 2D video and audio acquisition, encoding, transmission, and 3D video playback. It has network connectivity, data processing and user interaction capabilities, such as a mobile phone (or "cellular" phone), a fixed, portable, pocket, or handheld personal computer (PC) terminal, a mobile live streaming terminal, a set-top box, etc.
[0029] It should be noted that in this embodiment of the invention, users can input questions and inquire through interaction with the terminal interface. The terminal interface (or terminal display interface) refers to the screen interface used to provide users with multimodal information interaction, that is, the front-end interface that allows users to perform interactive operations during the question inquiry process, such as the display screen of the aforementioned terminal devices, including a screen film that can provide three-dimensional video display. For example, when the terminal is a naked-eye 3D adaptive display terminal that supports three-dimensional video acquisition, the naked-eye 3D effect can be achieved by supporting dual-mode switching between ordinary screen (e.g., Fresnel lens module) and holographic screen, reducing hardware upgrade costs.
[0030] Multimodal information refers to data expressed in multiple forms (such as text, images, audio, video, etc.). In this embodiment of the invention, the multimodal information is related to the user's inquiry. It should be noted that the multimodal information in this embodiment is the data after the network device receives the multimodal data and preprocesses the multimodal data input by the user through the terminal. This embodiment of the invention does not impose specific limitations on the data preprocessing process. For example, for video data, data preprocessing can be achieved through data denoising, video segmentation and frame extraction, and multimodal feature extraction to obtain multimodal information.
[0031] Multimodal information includes 3D video information, which refers to digital information that can present spatial depth and stereoscopic vision by simulating the parallax principle of the human eye. In other words, 3D video information adds depth information to traditional 2D video information, thus creating a video information with a stereoscopic visual experience. When the terminal is a 3D terminal device with 3D video acquisition capabilities, the 3D video information can be directly collected by the 3D terminal and provided to the network device. When the terminal is a conventional terminal without 3D video acquisition capabilities (i.e., only capable of 2D planar video acquisition), the 3D video information can be obtained by converting the 2D video information collected by the conventional terminal. The specific conversion method will be explained in detail later.
[0032] In this embodiment, after obtaining the multimodal information, the scene description text including emotional information is obtained by performing emotion modeling on the multimodal information.
[0033] In this embodiment, the scenario description text refers to a textual description of the scenario (or event) in which the user raises a question to the network device, as well as the user's emotional state in the current scenario (or event). Emotional information is used to describe the user's emotional state or mood, and is obtained by classifying the user's emotional state in the current scenario (or event), such as happy, sad, or angry.
[0034] In this embodiment, after obtaining the scene description text based on multimodal information, the intent information related to the user's current inquiry is identified from the scene description text, and the corresponding response result is generated based on the intent information.
[0035] In this embodiment, intent information is used to describe the purpose or goal that the user expects to achieve in the current emotional state of the scene (or event).
[0036] It should be noted that the embodiments of the present invention do not impose specific limitations on the method of generating response results based on intent information. For example, a Large Language Model (LLM) based on deep learning technology can be used. By recognizing emotional intent information based on scene description text and multimodal information representing the user's original question, emotional feedback corresponding to the question can be generated and provided to the user in a natural language manner (such as text, audio, 2D video / 3D video), i.e., the response result corresponding to the question. The emotional feedback includes the expected functional goal and the expected emotional goal. Emotional feedback ensures that the conclusions at these two levels are consistent, coordinated, and in line with user expectations, thereby providing the user with an accurate response result that satisfies the user's emotions.
[0037] In this embodiment of the invention, by acquiring three-dimensional video information of a user during a consultation, multimodal input information is formed, including three-dimensional video information, audio information, and text information. Using the three-dimensional video information and other modal information (i.e., objective information) from the multimodal input information, a scene description text containing emotional information is generated. Based on this emotional scene description text, the user's emotional intent is inferred. Because the three-dimensional video information in the acquired multimodal information includes scene depth information, millimeter-level depth changes can be captured, thereby accurately identifying subtle changes in facial expressions, body movements, etc., enhancing emotional resonance and improving the accuracy of understanding the user's intent.
[0038] In some optional implementations, when the multimodal video includes the three-dimensional video information, obtaining the multimodal information related to the inquiry includes: sending a live streaming request to the user's terminal, the live streaming request being used to request the user to start a three-dimensional video live stream; and, if the user confirms that the three-dimensional video live stream has been started, receiving the live stream through the terminal and obtaining the three-dimensional video information based on the live stream.
[0039] In this embodiment, before acquiring the 3D video information, the network device can send a live streaming request to the user's terminal to request (or suggest or remind) the user to start the 3D video live streaming. Then, if the user confirms that the 3D video live streaming has been started, the terminal can receive the live stream and acquire the 3D video information based on the live stream.
[0040] In this embodiment, 3D video live streaming (i.e., naked-eye 3D live streaming) refers to a real-time live streaming technology that can display stereoscopic images without the need for auxiliary display devices (such as 3D glasses). The 3D video stream (or 3D video live stream) provided by the live stream source can generate left and right eye views (such as side-by-side format (SBS)) through real-time transcoding. Then, the 3D effect is directly presented to the front-end user or back-end through the beam splitting technology of the naked-eye 3D display device. Currently, it is widely used in sports events, distance education, and other scenarios, greatly enhancing the user's immersion and viewing experience. In this embodiment of the invention, the network device suggests that the user enable 3D video live streaming to obtain the user's 3D video information in real time, thereby obtaining depth feature information for assessing emotions from the 3D video information.
[0041] It should be noted that the embodiments of the present invention do not impose specific limitations on the method of realizing 3D video live streaming. For example, video live streaming can be realized by deploying a real-time collaborative network. For instance, 5G (5th Generation Mobile Communication Technology) + TSN (Time-Sensitive Networking) protocol is deployed between the terminal and the network device. Priority scheduling ensures the transmission stability of the live stream. Combined with an edge-cloud collaborative computing method, a lightweight 3D rendering engine is deployed at the edge, and complex semantic reasoning is completed in the cloud, thereby achieving an end-to-end latency of less than 100 milliseconds (ms). The 3D video live streaming method in the embodiments of the present invention can be applied to a naked-eye 3D live streaming system connected to a deployed intelligent agent through an interface in a network device. By receiving and pushing the live stream, the live stream after the user starts the live stream is obtained, and 3D video information is provided to the network device.
[0042] In some optional implementations, the method further includes: before the user starts the live broadcast, the network device receives text information and / or audio information related to the question sent by the user through the terminal, queries a preset knowledge base based on the text information and / or audio information, and generates a response result corresponding to the question based on the query result.
[0043] In this embodiment, when the user has not enabled live video streaming, or when the network device only obtains text information and / or audio information, the network device performs feature extraction on the text information and / or audio information, queries relevant knowledge from a preset knowledge base based on the extracted feature information, and generates a response result based on the query results including relevant knowledge.
[0044] In this embodiment, the knowledge base is a pre-defined, structured, accessible, and continuously updated collection of information knowledge graphs, used to support intelligent agents in providing users with professional, accurate, and reliable consulting advice or response results for different fields (such as law, medicine, finance, technology, policy, etc.).
[0045] It should be noted that the embodiments of the present invention do not impose specific limitations on the implementation methods of feature extraction. For example, for feature extraction of text information, the Term Frequency-Inverse Document Frequency (TF-IDF) method can be used to calculate word weights through term frequency and inverse document frequency, thereby achieving text classification and keyword extraction; alternatively, the Word2Vec (Word to Vector) method can be used to map words to vectors, capture contextual semantics, and achieve feature extraction of text information through similarity calculation and text clustering. For feature extraction of audio information, the log-mel (log-mel) spectrogram can be used to convert the audio signal into a spectrum in the Mel frequency domain, preserving the temporal and frequency domain features of speech; alternatively, Mel Frequency Cepstral Coefficients (MFCCs) can be used to extract Mel cepstral coefficients to achieve speech recognition and audio feature extraction.
[0046] In this embodiment, in addition to querying relevant knowledge from a preset knowledge base based on the extracted feature information, the network device can also identify intent information related to the user's inquiry based on the feature information, and generate a response result based on the query results and intent information. The method for identifying intent information in this embodiment will be described in detail later.
[0047] In some optional implementations, when the terminal is not a three-dimensional terminal, the live stream is a two-dimensional live stream; the step of obtaining the three-dimensional video information based on the live stream includes: obtaining two-dimensional video information from the two-dimensional live stream, using a first model to perform depth estimation on each frame of the two-dimensional images in the two-dimensional video information to obtain each frame of three-dimensional images, and obtaining the three-dimensional video information based on each frame of three-dimensional images; the first model is trained based on a high-definition edge image dataset.
[0048] In this embodiment, when the terminal is a non-3D terminal device (i.e., a conventional terminal device), after the user starts live video streaming, the live stream received by the receiver is a two-dimensional live stream. Therefore, it is necessary to convert the two-dimensional live stream into a three-dimensional video live stream to obtain three-dimensional video information. The specific method for obtaining three-dimensional video information from the two-dimensional live stream is as follows: obtain two-dimensional video information from the two-dimensional live stream, input the two-dimensional video information into the first model, perform three-dimensional conversion processing on the two-dimensional video information through the first model, and output the three-dimensional video information.
[0049] In this embodiment, the first model is a model with three-dimensional conversion function, which is used to convert two-dimensional live stream into three-dimensional video live stream, that is, to convert two-dimensional video information into three-dimensional video information including three-dimensional images.
[0050] It should be noted that the embodiments of the present invention do not impose specific limitations on the first model. For example, a depth estimation algorithm can be used to estimate the 3D depth of people and / or objects in each frame of the two-dimensional video information (the distance range in three-dimensional space where an object can maintain a clear image before and after the focal point) of the people and / or objects in each frame of the two-dimensional RGB image. The depth (or distance) information of each pixel can be inferred from a single two-dimensional RGB image, thereby recovering the three-dimensional geometric structure of the scene and realizing the conversion of each frame of two-dimensional image into the corresponding three-dimensional image to obtain three-dimensional video information.
[0051] In this embodiment, considering that the dataset used by the depth estimation algorithm during training provides discrete, hardened, and relatively coarse edge information, the model cannot accurately identify the edges of people or images, thus failing to guarantee the 3D image quality after 3D conversion. Therefore, to improve the 3D conversion effect of the first model, this embodiment applies a high-definition edge image dataset (i.e., a high-definition matting dataset) to the training process of the first model. The high-definition edge image dataset provides continuous, softened, and high-precision edge information, including annotation information for the boundary regions of people and objects. When training the first model, the high-definition edge image dataset is input, and the annotation information in the dataset is used for edge information supervision, guiding the first model to better distinguish the edge pixels of people and / or objects, enabling the first model to have more accurate recognition of the edges of people and / or objects, thereby improving the 3D conversion effect of the first model on 2D images.
[0052] For example, taking the first model that uses a depth estimation algorithm as an example, Figure 2 , Figure 3 and Figure 4 These are all schematic diagrams comparing the performance of depth estimation algorithms according to embodiments of the present invention; [The remaining text appears to be incomplete and requires further context.] Figure 2 The original 2D image shown is used to perform edge recognition on the figures in the original 2D image using a depth estimation algorithm trained on a common dataset, resulting in the following: Figure 3 The depth estimation results shown in the image demonstrate that while the depth estimation algorithm trained on a common dataset can quickly identify the edges of people in an image, its edge recognition capability is limited. The recognition effect on the edges of people is too coarse, and it is almost impossible to identify the hair strands at the edges of people. This kind of depth estimation algorithm is prone to bias in the evaluation results, resulting in ghosting, jitter and other phenomena in the converted 3D image, thus affecting the 3D image effect.
[0053] By using a depth estimation algorithm trained on a high-resolution edge image dataset to perform edge recognition on people in the original two-dimensional image, the following results can be obtained: Figure 4 The depth estimation results shown demonstrate that the depth estimation algorithm trained on a high-definition edge image dataset can accurately identify the edges of a person while preserving all fine information such as hair strands and transparency, resulting in a significant improvement in the 3D image quality of the converted 3D image.
[0054] In this embodiment of the invention, when the terminal is not a three-dimensional terminal, the user's two-dimensional video information can only be obtained through a two-dimensional live stream. The two-dimensional video information is then input into a pre-trained depth estimation algorithm to obtain the three-dimensional video information converted from the two-dimensional video information by the depth estimation algorithm. The depth estimation algorithm is trained using a high-definition edge image dataset containing object boundary region annotations. The algorithm trained using a dataset containing object boundary region annotations can distinguish edge pixels and has more accurate recognition of people and / or object edges, thereby improving the three-dimensional effect of the converted three-dimensional video information.
[0055] In this embodiment, when the terminal is a 3D terminal (i.e., a 3D terminal device), after the user starts live video streaming, the live stream received by the network device is a 3D video live stream. In this case, there is no need to use the first model for 3D conversion processing; the 3D video information can be obtained directly from the 3D video live stream. This embodiment of the invention does not impose specific restrictions on the method of obtaining 3D video information from the 3D video live stream. For example, tools such as OpenCV (Open Source Computer Vision Library) or FFmpeg (Fast Forward Moving Picture Experts Group) can be used to extract image sequences from the live stream at fixed time intervals (e.g., 1 frame per second) or keyframes. Simultaneously, the audio track is separated, and multimodal feature extraction is performed using a Convolutional Neural Network (CNN) (e.g., a Residual Network). The extracted visual features, audio features, and text features are then fused to obtain the 3D video information.
[0056] It should be noted that, in this embodiment of the invention, after the naked-eye 3D live streaming system connected to the intelligent agent receives the live stream to obtain a two-dimensional or three-dimensional video live stream using a real-time collaborative network, a lightweight compression transmission module can be used to compress the live stream before transmission to save transmission resources. This embodiment of the invention does not impose specific limitations on the compression method used by the lightweight compression transmission module. For example, it can employ deep learning-based three-dimensional feature extraction technology, such as an autoencoder, to compress the original data to within 1.5 times the bandwidth of a traditional video stream, thereby achieving compression of the live stream.
[0057] Figure 5 This is a schematic diagram illustrating the process of obtaining scene description text according to an embodiment of the present invention; such as Figure 5 As shown, obtaining scene description text based on the multimodal information includes: Step 1021: Perform feature extraction on the three-dimensional video information to obtain three-dimensional feature information, and perform feature extraction on other modal information in the multimodal information other than the three-dimensional video information to obtain objective feature information. The three-dimensional feature information includes user three-dimensional feature information and / or object three-dimensional feature information related to the question consultation.
[0058] In this embodiment, after obtaining multimodal information including three-dimensional video information, the specific method for obtaining the scene description text corresponding to the question consultation based on the multimodal information is as follows: feature extraction processing is performed on each modality of the multimodal information, emotion modeling is performed based on the obtained feature information to estimate the user's emotion information, and scene description text associated with the question consultation is generated based on the emotion information and feature information.
[0059] In this embodiment, for example, when the multimodal information includes three-dimensional video information, two-dimensional video information, text information, and voice information related to the consultation, it is necessary to extract features from the three-dimensional video information to obtain three-dimensional feature information, and to extract features from the two-dimensional video information, text information, and voice information other than the three-dimensional video information in the multimodal information to obtain objective feature information.
[0060] The three-dimensional feature information may include local surface features, global shape features, and structural relationship features of a person and / or object. Local surface features are used to represent the geometric properties near a certain feature point on the surface of a person and / or object. For example, local surface features of a person may include facial features (such as facial features, skin features, hair features, etc.), body features (such as limb features), and clothing features (such as clothing features, accessory features). Local surface features of an object may include normals (used to describe the direction in which a certain feature point on the object's surface faces, i.e., a vector perpendicular to the tangent plane of that feature point), curvature (used to describe the degree of curvature of the surface at a certain feature point), etc.
[0061] Global shape features are used to represent the overall geometric properties of a person and / or an object. For example, the global shape features of a person can be features such as limb movement features and body shape features (such as height and build); the global shape features of an object can include features such as the object's volume and surface area.
[0062] Structural relationship features are used to represent the partial or overall structural state of a person and / or an object. For example, the structural relationship features of a person are used to describe the spatial distribution, three-dimensional block structure, and overall posture of the person; the structural relationship features of an object are used to describe the topological structure and compositional relationships of the object. For example, a wrench consists of a long handle and a gripping head. In addition, structural relationship features can also be used to describe the connection relationship between a person and an object. For example, in the scenario of after-sales service, a user holds an item to demonstrate the shortcomings of the product's functions in a live broadcast. The structural relationship features obtained through feature extraction describe the person holding the purchased item.
[0063] In this embodiment, for people and / or objects in 3D video information, the 3D feature information obtained by feature extraction from the 3D video information includes user 3D feature information and / or object 3D feature information related to the user's question / consultation. Specifically, the user 3D feature information mainly includes the user's key point 3D position, body movements (such as facial expressions and gestures), and other 3D geometric feature information. The user's key points refer to multiple feature points used for emotional assessment, such as feature points that constitute facial features, facial muscles, gestures, and torso, which can reflect the user's emotional state. The object 3D feature information mainly includes 3D geometric feature information such as normals, curvature, surface area, volume, and structure.
[0064] In this embodiment of the invention, by extracting features from three-dimensional video information, user three-dimensional feature information containing user body movement features and object three-dimensional feature information of the items mentioned by the user in the live broadcast are obtained. The user's body movement features are used for auxiliary reasoning for emotion estimation, and the body movement features and / or the three-dimensional feature information of the items are used to generate objective information for scene description text. This enriches the content of multimodal information and makes the scene description text containing emotion information generated based on multimodal information more accurate and complete in describing the scene content.
[0065] The objective feature information refers to the feature information formed by extracting features from modal information other than three-dimensional video information. For example, the objective feature information may specifically include two-dimensional visual features (such as image features, dynamic temporal features, etc.) obtained by extracting features from two-dimensional video information, text semantic features obtained by extracting features from text information, and acoustic features, emotion features, speech information, environmental sounds, etc. obtained by extracting features from audio information.
[0066] It should be noted that when the terminal is a non-3D terminal (or a conventional terminal), the live streaming system in the network device can transmit the 2D video information obtained from the 2D live stream to the intelligent agent along with the 3D video information after converting the 2D video information into 3D video information. When the terminal is a 3D terminal, the 2D video information is obtained during the feature extraction process of the 3D video information through transformation or dimensionality reduction.
[0067] It should be noted that the embodiments of the present invention do not impose specific limitations on the feature extraction methods for multimodal information. For example, for three-dimensional video information, a 3D encoder can be used to detect and transform the three-dimensional spatial position to realize the parsing of the three-dimensional video information and obtain three-dimensional feature information; for two-dimensional video information, a 2D encoder can be used to capture position changes on a two-dimensional plane to realize the parsing of the two-dimensional video information and obtain two-dimensional visual features; for text information, a text encoder can be used to convert the text information into a fixed-length vector representation to capture its semantic and contextual information, thereby realizing the parsing of the text information and obtaining text semantic features; for audio information, an audio encoder can be used to convert the audio signal into a digital encoding format to realize the parsing of the audio information and obtain acoustic features, emotional features, and other feature information.
[0068] Step 1022: Perform sentiment estimation based on the three-dimensional feature information and the objective feature information to determine the user's sentiment information.
[0069] In this embodiment, after obtaining three-dimensional feature information and objective feature information by feature extraction of multimodal information, sentiment estimation is performed based on the three-dimensional feature information and objective feature information. Subjective estimation conclusions are formed by inferring the user's internal emotional state when consulting the current problem, and sentiment information used to reflect the user's current emotional (or mood) state is determined.
[0070] It should be noted that the embodiments of the present invention do not impose specific limitations on the method of emotion estimation. For example, a large model based on deep learning algorithms (such as recurrent neural networks (RNN) or CNN) can be used to identify and classify emotional information and output emotion categories such as happiness, sadness, anger, and surprise.
[0071] In some optional implementations, the step of performing sentiment estimation based on the three-dimensional feature information and the objective feature information to determine the user's sentiment information includes: extracting time-series features from the user's three-dimensional feature information, performing sentiment estimation based on the time-series features and the objective feature information to obtain the sentiment information, wherein the time-series features include: time information, three-dimensional location information of key points, and number of key points.
[0072] In this embodiment, the specific method for estimating the user's emotional information based on three-dimensional feature information and objective feature information is as follows: extract the time series features from the three-dimensional feature information, and perform sentiment estimation based on the time series features and objective feature information to obtain the user's emotional information.
[0073] In this embodiment, for example, the time series features are obtained by analyzing three-dimensional feature information that changes over time. In the case where the three-dimensional feature information is the user's three-dimensional feature information, the time series features can be generated based on feature information used for emotion estimation, such as limb movement information and key point location information. This results in time series features generated based on limb movement information and time series features based on key point coordinates, which reflect the subtle and dynamic changes in the user's facial and / or body posture. The time series features can specifically include: time information, three-dimensional location information of key points, and number of key points.
[0074] It should be noted that, taking the time-series features generated based on body movement information and the time-series features of keypoint coordinates as an example, in the process of sentiment estimation based on time-series features, the weights of the time-series features generated based on body movement information and the time-series feature values of keypoint coordinates can be adaptively adjusted to input into a large model for sentiment estimation. Furthermore, when generating scene text descriptions containing emotional information, the weight coefficients for the fusion of feature information from different modalities can be adaptively set.
[0075] In this embodiment, feature extraction is performed on the acquired 3D video information to extract the coordinate information of key points for emotion estimation at different times, forming time-series feature values of the key points. These time-series feature values are then input into a large model to obtain the model's output of the user's emotion estimation conclusion. Because the coordinate information of key points (such as facial features, body movements, etc.) at different times is extracted, subtle changes in user facial expressions or body movements can be quickly captured, improving the efficiency of emotion inference.
[0076] It should be noted that before sentiment estimation, the acquired 3D feature information and objective feature information need to undergo multimodal alignment and fusion processing to establish semantic associations between features of different modalities. This embodiment of the invention does not impose specific limitations on the methods of multimodal alignment and fusion processing. For example, for multimodal alignment processing, a contrastive learning approach can be used, training with positive and negative samples to make representations of related modalities closer together in the vector space and those that are unrelated farther apart; alternatively, a shared representation space approach can be used to map different modal data to a unified space, ensuring that semantically related content is adjacent in space. For multimodal fusion processing, feature-level fusion (i.e., early fusion) methods can be used, immediately concatenating features of different modalities into a high-dimensional vector after feature extraction, and combining methods such as Principal Component Analysis (PCA) and Minimum Redundancy Maximum Relevance (mRMR) to remove redundant information, thereby achieving the fusion of multimodal feature information.
[0077] Step 1023: Generate a scene description text associated with the question consultation based on the emotional information, the user's three-dimensional feature information and / or the object's three-dimensional feature information, and the objective feature information.
[0078] In this embodiment, after estimating the user's emotional information based on the fused feature information, the emotional information is used as subjective feature information. Based on the subjective feature information, as well as the feature information fused from the three-dimensional feature information and the objective feature information, a scene description text associated with the question consultation is generated.
[0079] It should be noted that the embodiments of the present invention do not impose specific restrictions on the generation method of scene description text. For example, a large language model is used, and the feature information after the fusion of subjective feature information (i.e., emotional information) and multimodal feature information is input into the large language model. The large language model is then used to output scene description text including emotional information in text form according to the preset generation rules of scene description text (e.g., event scene description + ";" + emotional description + "," + "evaluate user emotion as " + "" + emotional information + "").
[0080] For example, the scene text description generated using a large language model and containing emotional information is as follows: The user expresses that the price of the clothes he bought is too high compared to other merchants, and requests a refund or compensation for the price difference; at the same time, the user wears this clothes during the live broadcast and always maintains a satisfied smile on his face, and the user's emotion is assessed as "happy".
[0081] In some optional implementations, determining the intent information corresponding to the question based on the scenario description text includes: inputting the scenario description text into a second model, using the second model to perform intent analysis on the scenario description text, and outputting the intent information corresponding to the question.
[0082] In this embodiment, after obtaining the scene description text based on multimodal information, the specific method for determining the intent information corresponding to the question consultation based on the scene description text is as follows: input the scene description text into the second model, use the second model to perform intent reasoning analysis on the scene description text, and output the intent information corresponding to the question consultation.
[0083] In this embodiment, the second model is used to infer the intent information corresponding to the user's inquiry. It can employ a large-scale deep learning-based model, trained on a large amount of data to form a neural network. This model takes a scene description text containing emotional information, generated based on multimodal information, as input and outputs a textual description of the emotional intent corresponding to that scene. For example, if the scene description text reads, "The user expresses that the price of the clothes they bought is too high compared to other merchants, and requests a refund or price difference; at the same time, the user is wearing this clothing during a live stream, maintaining a satisfied smile, and the user's emotion is assessed as 'happy'," then this scene description text is input into the second model. The intent information output by the second model is, "The user's true intent is most likely not to return the goods; a discount coupon can be given."
[0084] In some optional implementations, determining the intent information corresponding to the inquiry based on the scenario description text includes: obtaining event description information and sentiment information related to the inquiry based on the scenario description text; querying a knowledge base based on the event description information and sentiment information, wherein the knowledge base pre-stores one or more triples including event description information, sentiment information, and intent information; if no intent information corresponding to the event description information and sentiment information is found in the knowledge base, inputting the scenario description text into a second model, and using the second model to output the intent information corresponding to the inquiry.
[0085] In this embodiment, in addition to directly using the second model to infer intent information, intent information can also be queried by searching the knowledge base. Specifically, event description information and sentiment information related to the question consultation are obtained from the scene description text, and the knowledge base is queried based on the event description information and sentiment information to retrieve the intent information corresponding to the event description information and sentiment information from the knowledge base.
[0086] In this embodiment, preset extraction rules can be used to extract event description information and sentiment information related to the problem consultation from the scene description text. The preset extraction rules correspond to the preset generation rules of the scene description text, that is, the event description information and sentiment information are extracted according to the generation rules of the scene description text. For example, when the generation rule is: event scene description + ";" + sentiment description + "," + "evaluate user sentiment as " + "" + sentiment information + "", the corresponding extraction rule can be preset as follows: the content before the first semicolon ";" in the scene description text is the event description information X, and the content within the quotation marks """ after "evaluate user sentiment as" is the sentiment information Y.
[0087] In this embodiment, the knowledge base pre-stores one or more triples, each containing event description information X, emotion information Y, and intent information Z. Each triple represents a logical relationship between an event, emotion, and intent; that is, each triple means that if event X occurs, the user will typically feel emotion Y and may generate intent Z. The triples in the knowledge base can be continuously updated during each query.
[0088] For example, if the content of the scenario description text is "The user expresses in words that the price of the clothes they bought is too high compared to other merchants, and requests a refund or price difference compensation; at the same time, the user wears this clothing during a live broadcast, and maintains a satisfied smile on their face, and the user's emotion is assessed as 'happy'", based on the extraction rule "the content before the first semicolon ";" in the scenario description text is the event description information X, and the content in the quotation marks "" after "assessed user's emotion" is the emotion information Y", the event description information X in the scenario description text is extracted as: The user expresses in words that the price of the clothes they bought is too high compared to other merchants, and requests a refund or price difference compensation; the emotion information Y is: happy. If the knowledge base stores a triple including the event description information X and the emotion information Y, by querying the knowledge base, the intent information Z corresponding to the event description information X and the emotion information Y is determined to be: The user's true intent is most likely not to return the goods, but to offer a coupon.
[0089] In this embodiment of the invention, according to preset extraction rules, the events described by the user and the user's emotions are parsed from the generated scene description text containing emotional information. Using the parsed events and emotions, a knowledge base containing (event, emotion, intent) triples is retrieved to output the intent information from the retrieved triples. Through the preset triples in the knowledge base, the reasoning result of the user's intent can be quickly obtained.
[0090] In this embodiment, if no intent information corresponding to the event description information and sentiment information in the scene description text is found by using the triples pre-stored in the knowledge base, the scene description text can be input into the second model in combination with the above-mentioned intent information inference method based on the second model. The second model is then used to perform intent analysis on the scene description text and output intent information corresponding to the question consultation.
[0091] In this embodiment, after obtaining the intent information corresponding to the user's question based on the scenario description text, the intent information, scenario description text, and the user's original question can be input into the large language model. The large language model outputs an emotional response, which can be sent to the terminal in text or audio form and displayed to the user.
[0092] For example, consider the following scenario description: A user expresses that the price of the clothes they purchased is too high compared to other merchants and requests a refund or price difference compensation. Simultaneously, the user wears the clothes during a live stream, maintaining a satisfied smile. Assessing the user's emotion as "happy" and their intent as likely not being a refund, and considering the possibility of a coupon, the response output by the large language model would be: First, thank you for choosing our product, but we sincerely apologize for any inconvenience this issue may have caused. Through the 3D video live stream, we observed that you perfectly showcased all the advantages of the product. There are many factors that could cause a price difference, and to maintain our consistent product quality, we regret that we cannot offer a price difference refund. However, to avoid disappointing you, we can provide an exclusive coupon, hoping to reassure you that you should keep the item. Finally, your satisfaction is always our top priority, and we look forward to continuing to provide you with a wonderful shopping experience.
[0093] In this embodiment, the response can be fed back to the terminal not only in the form of text and / or voice information, but also, for special scenarios, in a spatial dimension, through two-dimensional or three-dimensional video information, to improve the user's understanding of the response. For example, in a medical consultation scenario, while providing medical responses to the user via text or voice, 3D models of organs and their associated pathology can also be sent to the terminal via 3D video to help the user understand the response more accurately.
[0094] As an example, let's take the multimodal information acquired by network devices, including 3D video information, 2D video information, text information, and audio information, as an example. Figure 6 This is a schematic diagram of the processing flow for obtaining response results based on multimodal information according to an embodiment of the present invention; as shown below. Figure 6As shown, after the multimodal information is input into the large model, the specific process of obtaining the response results corresponding to the user's question based on the multimodal information using the large model is as follows: Step 201: Multimodal feature extraction.
[0095] Specifically, a 3D encoder is used to parse and extract features from 3D video information to obtain 3D feature information, including user 3D feature information and object 3D feature information; a 2D encoder is used to parse and extract features from 2D video information to obtain 2D visual features; a text encoder is used to parse and extract features from text information to obtain text semantic features; and an audio encoder is used to parse and extract features from audio information to obtain acoustic features, emotion features, and other feature information. All feature information other than 3D feature information is used as objective feature information for subsequent emotion estimation and scene description text generation.
[0096] Step 202: Multimodal feature alignment and fusion.
[0097] Specifically, the acquired user 3D feature information, object 3D feature information, and objective feature information are subjected to multimodal alignment processing to map different modal data to a unified space, ensuring that semantically related content is adjacent in space, and the aligned feature information of different modalities is connected into a high-dimensional vector to realize the fusion of multimodal feature information in order to establish semantic associations between different modal features.
[0098] Step 203: Sentiment estimation.
[0099] Specifically, after aligning and fusing the three-dimensional feature information and the objective feature information, the time series features in the three-dimensional feature information are extracted. Sentiment estimation is performed based on the time series features and the fused features. Subjective estimation conclusions are formed by inferring the user's internal emotional state when consulting the current problem, and sentiment information used to reflect the user's current emotional (or mood) state is determined.
[0100] Step 204: Generate scene description text.
[0101] Specifically, after estimating the user's emotional information based on the fused feature information, the emotional information is used as subjective feature information. Based on the subjective feature information, as well as the fused features obtained by fusing three-dimensional feature information and objective feature information, a scene description text that is related to the question consultation and includes the user's emotional state is generated.
[0102] Step 205: Obtain intent information based on the scene description text.
[0103] Specifically, the process involves extracting event description information and sentiment information related to the problem consultation from the scene description text, querying the knowledge base based on the event description information and sentiment information, and querying the knowledge base for intent information corresponding to the event description information and sentiment information. If no intent information corresponding to the event description information and sentiment information in the scene description text is found through the pre-stored triples in the knowledge base, the scene description text is input into the second model, and the second model is used to perform intent analysis on the scene description text and output intent information corresponding to the problem consultation.
[0104] Step 206: Generate the response result.
[0105] Specifically, after obtaining the intent information corresponding to the user's question based on the scenario description text, the intent information, scenario description text, and the user's original question are input into the large language model. The large language model outputs an emotional response and feeds the response back to the terminal for display to the user through multimodal forms such as 3D video, 2D video, text, or audio.
[0106] As another example Figure 7 This is a schematic diagram of the module interaction of the interaction method according to an embodiment of the present invention; as shown below. Figure 7 As shown, the network device used to implement the interaction method in this embodiment of the invention mainly includes two parts: a consultative intelligent agent 31 and a naked-eye 3D live streaming system 32 connected to the consultative intelligent agent through an interface. The consultative intelligent agent 31 includes a user service portal 33, a knowledge base 34, an intelligent agent scheduling module 35, and an intelligent agent large model 36. The naked-eye 3D live streaming system 32 includes a dynamic light field synthesis engine 37, a live streaming central control module 38, and a compression transmission processing module 39.
[0107] The user service portal 33 acts as a "portal" between the terminal 30 and the consulting intelligent agent 31. It receives multimodal inquiries from users through the terminal 30 and transmits the received multimodal information to the intelligent agent scheduling module 35. It also provides feedback to the terminal 30 via text, voice, or video, enabling communication and interaction with users to resolve their inquiries. Furthermore, it sends live streaming requests to the terminal 30, suggesting the user initiate a 3D live stream, and sends confirmation feedback to the intelligent agent scheduling module 35.
[0108] Knowledge base 34 supports the consultative agent 31 in providing professional, accurate, and reliable multimodal knowledge in various domains based on user-provided questions. Specifically, knowledge base 34 may include: a text or audio information knowledge graph and a 3D information knowledge graph, used to provide corresponding text, audio, or 3D image or video feedback based on user inquiries. Furthermore, knowledge base 34 stores one or more triples including event description information, sentiment information, and intent information, used to provide corresponding intent information based on scene description text.
[0109] The intelligent agent scheduling module 35, as the core control center of the consulting intelligent agent 31, is equivalent to a highly intelligent traffic command system + resource allocation center. It is mainly responsible for intelligently and efficiently coordinating the collaboration of various modules in the consulting intelligent agent 31 to ensure that user consultation requests are accurately processed and generate high-quality responses.
[0110] Specifically, the intelligent agent scheduling module 35 is used to receive multimodal information related to the inquiry sent by the terminal 30 through the user service portal 33; it is also used to transmit the inquiry to the knowledge base 34, and obtain relevant multimodal knowledge related to the inquiry through the knowledge base 34; it is also used to provide the inquiry and relevant knowledge to the intelligent agent big model 36, and use the intelligent agent big model 36 to directly generate the corresponding response result, or, send a live broadcast request to the terminal through the user service portal 33, and transmit the received three-dimensional video information and two-dimensional video information to the intelligent agent big model 36 through the naked-eye 3D live broadcast system 32 upon receiving confirmation feedback, use the intelligent agent big model 36 to generate scene description text including emotional information, and obtain the intent information corresponding to the scene description text through the knowledge base 34, and finally feed back the obtained multimodal response result corresponding to the inquiry to the terminal 30 through the user service portal 33.
[0111] The intelligent agent large model 36, serving as the "brain" of the consultative intelligent agent 31, surpasses the generalization capabilities of general large models. It deeply integrates vertical domain knowledge and specialized reasoning logic, achieving a leap from "information retrieval" to "decision support." The intelligent agent large model 36 is primarily responsible for understanding 3D video information acquired from the naked-eye 3D live streaming subsystem, thereby generating feedback with more emotional attributes. The intelligent agent large model 36 mainly comprises two sub-modules: a multimodal intent understanding module 41 and an emotion modeling module 42.
[0112] The multimodal intent understanding module 41 integrates Simultaneous Localization and Mapping (SLAM) technology to analyze user gestures, gaze focus, and other three-dimensional spatial behaviors, and combines them with Natural Language Processing (NLP) to generate context-aware responses, i.e., scene description text. The emotion modeling module 42 uses multimodal feature information analysis methods such as micro-expression recognition (3D facial feature analysis) and audio emotion analysis to adjust the anthropomorphism of the agent's feedback and infer the user's emotional (or mood) state.
[0113] Specifically, the intelligent agent large model 36 is used to receive multimodal inquiries and related multimodal knowledge feedback through the intelligent agent scheduling module 35, generate a response result corresponding to the inquiry, and send the response result to the intelligent agent scheduling module 35; it is also used to receive 3D and 2D video information through the intelligent agent scheduling module 35 when the user enables 3D video live streaming, perform sentiment estimation based on the multimodal information using the sentiment modeling module 42, and generate scene description text based on the estimated sentiment information; it is also used to infer the corresponding intent information based on the scene description text, or to obtain the corresponding intent information through the intelligent agent scheduling module 35, generate a related response result based on the inquiry, scene description text, and intent information using the multimodal intent understanding module 41, and send the response result to the intelligent agent scheduling module 35.
[0114] The Dynamic Light Field Synthesis Engine 37, also known as the 2D to 3D engine, primarily employs depth estimation algorithms combined with user viewpoint tracking technology (such as red-green-blue-depth (RGB-D, RGB-Depth) cameras) to generate real-time 3D video live stream images adapted to different perspectives.
[0115] Specifically, taking a conventional terminal device, where terminal 30 is a non-3D terminal, as an example, the dynamic light field synthesis engine 37 is used to receive 2D video information obtained from the 2D live stream through the live broadcast control module 38; it is also used to perform depth estimation on each frame of 2D images in the 2D video information using a depth estimation algorithm trained on a high-definition edge image dataset, thereby realizing 3D conversion processing of each frame of 2D images to obtain 3D video information, and transmitting the 3D video information to the consulting intelligent agent 31 through the live broadcast control module 38.
[0116] The live streaming control module 38, as the core control hub of the naked-eye 3D live streaming system 32, is connected to the consulting intelligent agent 31 through the main interface. It is mainly responsible for receiving and pushing the live stream, as well as coordinating the cooperation of various modules in the naked-eye 3D live streaming system 32.
[0117] Specifically, taking a conventional terminal device, where terminal 30 is a non-3D terminal, as an example, the live broadcast control module 38 is used to obtain the 2D live broadcast stream provided by terminal 30 through the live broadcast stream reception after receiving feedback information from the user confirming the start of 3D video live broadcast. It then obtains 2D video information from the 2D live broadcast stream, sends the 2D video information to the dynamic light field synthesis engine 37, uses the dynamic light field synthesis engine 37 to obtain the 3D video information after 3D conversion, and sends the 3D video information and 2D video information to the consulting intelligent agent 31.
[0118] The compression and transmission processing module 39 uses deep learning-based 3D feature extraction technology to compress the raw data to 1.5 times the bandwidth of traditional video streams, thereby saving the transmission cost of video information.
[0119] Specifically, taking a conventional terminal device, where terminal 30 is not a 3D terminal, as an example, after the live broadcast control module 38 receives feedback information confirming the start of 3D video live broadcast from the user, the compression and transmission processing module 39 obtains the 2D live broadcast stream sent by terminal 30 through the real-time collaborative network, performs compression processing on the 2D live broadcast stream before transmission, and sends the compressed 2D live broadcast stream to the live broadcast control module 38. It is also used to compress the converted 3D video live broadcast stream (or 3D video information), thereby transmitting the compressed live broadcast stream (or video information) to the consulting intelligent agent 31 through the live broadcast control module 38.
[0120] Based on the above embodiments, this invention also provides an interactive device. Figure 8 This is a schematic diagram of the composition structure of the interactive device provided in the embodiments of the present invention; as shown below. Figure 8 As shown, the device includes: The intelligent agent module 51 is used to acquire multimodal information related to the inquiry, the multimodal information including at least one or more of three-dimensional video information, text information, and audio information; it is also used to acquire scene description text based on the multimodal information, the scene description text including emotional information, the emotional information being used to describe the user's emotional state; it is also used to determine the intent information of the inquiry based on the scene description text, and output the corresponding response result of the inquiry based on the intent information.
[0121] In an optional embodiment of the present invention, when the multimodal video includes the three-dimensional video information, the intelligent agent module 51 is further configured to send a live broadcast request to the user's terminal when acquiring the multimodal information related to the question consultation, the live broadcast request being used to request the user to start a three-dimensional video live broadcast; The device also includes a live streaming module 52, which is used to receive the live stream through the terminal and obtain the three-dimensional video information based on the live stream when the user confirms the start of the three-dimensional video live stream.
[0122] In an optional embodiment of the present invention, when the terminal is not a three-dimensional terminal, the live stream is a two-dimensional live stream; the live streaming module 52 is used to obtain two-dimensional video information from the two-dimensional live stream, use a first model to perform depth estimation on each frame of two-dimensional images in the two-dimensional video information to obtain each frame of three-dimensional images, and obtain the three-dimensional video information based on each frame of three-dimensional images; the first model is trained based on a high-definition edge image dataset.
[0123] In an optional embodiment of the present invention, the intelligent agent module 51 is configured to extract features from the three-dimensional video information to obtain three-dimensional feature information, and to extract features from other modal information in the multimodal information besides the three-dimensional video information to obtain objective feature information. The three-dimensional feature information includes user three-dimensional feature information and / or object three-dimensional feature information related to the question consultation. Based on the three-dimensional feature information and the objective feature information, sentiment estimation is performed to determine the user's sentiment information. Based on the sentiment information, the user three-dimensional feature information and / or the object three-dimensional feature information, and the objective feature information, a scene description text associated with the question consultation is generated.
[0124] In an optional embodiment of the present invention, the intelligent agent module 51 is used to extract time series features from the user's three-dimensional feature information, perform sentiment estimation based on the time series features and the objective feature information, and obtain the sentiment information. The time series features include: time information, key point three-dimensional position information, and key point quantity information.
[0125] In an optional embodiment of the present invention, the intelligent agent module 51 is used to input the scene description text into the second model, use the second model to perform intent analysis on the scene description text, and output intent information corresponding to the question consultation.
[0126] In an optional embodiment of the present invention, the intelligent agent module 51 is configured to obtain event description information and emotional information related to the question consultation based on the scene description text; query a knowledge base based on the event description information and the emotional information, wherein the knowledge base pre-stores one or more triples including event description information, emotional information and intent information; if no intent information corresponding to the event description information and the emotional information is found in the knowledge base, the scene description text is input into a second model, and the second model is used to output the intent information corresponding to the question consultation.
[0127] In the implementation of this invention, the intelligent agent module 51 and the live broadcast module 52 in the device can both be implemented by the central processing unit (CPU), digital signal processor (DSP), microcontroller unit (MCU), or field-programmable gate array (FPGA) in the device in practical applications.
[0128] It should be noted that the interaction between the above-described interactive device and the user's terminal is only illustrated by the division of the program modules described above. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. For example, the intelligent agent module 51 in the device can be applied to... Figure 7 The consultative intelligent agent 31 in the system structure shown, and the live streaming module 52 in the device can be applied to Figure 7 The naked-eye 3D live streaming system 32 in the system structure shown enables the interaction method. Furthermore, the interaction device and interaction method embodiments provided in the above embodiments belong to the same concept; their specific implementation process is detailed in the method embodiments and will not be repeated here.
[0129] This invention also provides a computer device. Figure 9 This is a schematic diagram of the hardware composition structure of a computer device provided in an embodiment of the present invention; such as... Figure 9 As shown, the computer device includes a memory 62, a processor 61, and a computer program stored on the memory 62 and executable on the processor 61.
[0130] Optionally, when the processor 61 executes the program, it implements the steps of the interaction method described in the embodiments of the present invention.
[0131] Optionally, various components in the computer device can be coupled together via a bus system 63. It is understood that the bus system 63 is used to implement communication between these components. In addition to a data bus, the bus system 63 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 9 The general labeled all buses as Bus System 63.
[0132] It is understood that memory 62 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 62 described in this embodiment of the invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0133] The methods disclosed in the above embodiments of the present invention can be applied to processor 61, or implemented by processor 61. Processor 61 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 61 or by instructions in the form of software. The processor 61 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 61 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present invention can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory 62. Processor 61 reads the information in memory 62 and completes the steps of the aforementioned method in combination with its hardware.
[0134] In an exemplary embodiment, the computer device may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components to perform the aforementioned methods.
[0135] This invention also provides a computer-readable storage medium having a computer program stored thereon.
[0136] Optionally, the computer-readable storage medium can be applied to the interactive device of the present invention; then, when the program is executed by the processor, it implements the steps of the interactive method of the present invention.
[0137] This invention also provides a computer program product, including a computer program that can be executed by a processor 61 of a computer device to complete the steps of the interaction method described in this invention.
[0138] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0139] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0140] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0141] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0142] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0143] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0144] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0145] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
[0146] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. An interaction method, characterized in that, The method includes: Obtain multimodal information related to the inquiry, wherein the multimodal information includes at least one or more of three-dimensional video information, text information, and audio information; The scene description text is obtained based on the multimodal information, and the scene description text includes emotional information, which is used to describe the user's emotional state. Based on the scenario description text, determine the intent information of the inquiry, and output the corresponding response result based on the intent information.
2. The method according to claim 1, characterized in that, When the multimodal video includes the three-dimensional video information, the step of acquiring multimodal information related to the inquiry includes: Send a live streaming request to the user's terminal, the live streaming request being used to request the user to start a three-dimensional video live stream; When the user confirms that a 3D video live stream has been started, the terminal receives the live stream and obtains the 3D video information based on the live stream.
3. The method according to claim 2, characterized in that, When the terminal is not a 3D terminal, the live stream is a 2D live stream; obtaining the 3D video information based on the live stream includes: Two-dimensional video information is obtained from the two-dimensional live stream. The first model is used to perform depth estimation on each frame of the two-dimensional video information to obtain each frame of three-dimensional images. The three-dimensional video information is obtained based on each frame of three-dimensional images. The first model is trained based on a high-definition edge image dataset.
4. The method according to claim 2, characterized in that, The step of obtaining scene description text based on the multimodal information includes: The 3D video information is subjected to feature extraction to obtain 3D feature information, and other modal information in the multimodal information other than the 3D video information is subjected to feature extraction to obtain objective feature information. The 3D feature information includes user 3D feature information and / or object 3D feature information related to the question consultation. Sentiment estimation is performed based on the three-dimensional feature information and the objective feature information to determine the user's sentiment information; Based on the emotional information, the user's three-dimensional feature information and / or the object's three-dimensional feature information, and the objective feature information, a scene description text associated with the inquiry is generated.
5. The method according to claim 4, characterized in that, The step of performing sentiment estimation based on the three-dimensional feature information and the objective feature information to determine the user's sentiment information includes: Extract time-series features from the user's 3D feature information, and perform sentiment estimation based on the time-series features and the objective feature information to obtain the sentiment information. The time-series features include: time information, 3D location information of key points, and number of key points.
6. The method according to claim 4 or 5, characterized in that, The step of determining the intent information corresponding to the question based on the scenario description text includes: The scenario description text is input into the second model, and the second model is used to perform intent analysis on the scenario description text, outputting intent information corresponding to the question.
7. The method according to claim 4 or 5, characterized in that, The step of determining the intent information corresponding to the question based on the scenario description text includes: Based on the scenario description text, obtain event description information and emotional information related to the question consultation; Based on the event description information and the sentiment information, a knowledge base is queried. The knowledge base pre-stores one or more triples including event description information, sentiment information and intent information. If no intent information corresponding to the event description information and the emotional information is found in the knowledge base, the scene description text is input into the second model, and the second model is used to output the intent information corresponding to the question.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.