Conference terminal interaction system and method based on large model

By collecting and processing multimodal emotion data, generating corrective judgment items and displaying them on the interactive interface, the problem of iterative optimization and forward-looking guidance of interactive content in existing technologies is solved, realizing intelligent iterative correction and improved accuracy of meeting interactions.

CN121887837APending Publication Date: 2026-04-17CLAIRVOYANCE (GUANGZHOU) ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CLAIRVOYANCE (GUANGZHOU) ARTIFICIAL INTELLIGENCE TECH CO LTD
Filing Date
2026-03-18
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing conference interaction technologies lack a closed-loop correction mechanism based on large models, making it impossible to achieve intelligent iterative optimization of interactive content. They also lack predictive capabilities in the time dimension, have rudimentary data collection and processing, and lack model self-evolution capabilities, resulting in insufficient accuracy of interactive features.

Method used

By setting up a data acquisition module to collect multimodal sentiment data, removing non-emotional noise, performing weighted calculations, calling a large model to generate correction judgment items, and displaying correction suggestions in real time and predictively on the interactive interface, it achieves iterative correction and forward-looking guidance, and optimizes model parameters by combining time series prediction and closed-loop feedback mechanisms.

Benefits of technology

It significantly improves the accuracy and adaptability of meeting interactions, realizing a shift from passive response to proactive prediction, and ensuring the accuracy of interactive content and the continuous self-learning optimization of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887837A_ABST
    Figure CN121887837A_ABST
Patent Text Reader

Abstract

The invention discloses a conference terminal interaction system and method based on a large model, and the system is provided with a conference terminal in a conference space with preset arrangement data, and a data obtaining module limits the conference space into at least one emotion discrimination space according to the preset arrangement data. Collecting multi-modal emotion data in the emotion judgment space when the first interaction interface of the conference terminal is passed; the first data processing module determines a dynamic interaction feature associated with the conference content and the emotion according to the multi-mode emotion data and a first interaction interface; the second data processing module generates a first interaction interface and a second interaction interface at the conference terminal, displays the correction judgment item on the second interaction interface, and performs iterative correction on interaction content in the first interaction interface according to the correction judgment item; a server configured to store at least one large model; and the storage module is configured to store interaction data of the server and the conference terminal. According to the invention, intelligent iterative correction and prospective guidance of the interactive content are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically to a conference terminal interaction system and method based on a large model. Background Technology

[0002] While some existing conference interaction technologies are capable of collecting audience feedback, they generally suffer from the following shortcomings: First, existing technologies mostly focus on one-way data display and lack closed-loop correction mechanisms based on large-scale model reasoning. Traditional systems typically visualize the collected emotion or action data directly, failing to generate real-time corrections for the meeting content. This means that speakers can only see and understand the audience's state but cannot obtain specific guidance on how to adjust the interactive content in the first interactive interface, making it difficult to achieve iterative optimization of the meeting content.

[0003] Secondly, existing interactive interfaces lack the ability to predict events over time. Most systems only display the current state data and cannot extrapolate the emotional evolution path of future frames based on the dynamic interaction characteristics of the current frame. This results in speakers not being able to know the potential risks of future content in advance and lacking the ability to proactively adjust their interaction strategies.

[0004] Secondly, existing technologies are relatively crude in data acquisition and processing. They lack a refined definition and weight allocation for the emotion discrimination space, and fail to effectively remove non-emotional noise from multimodal emotion data, resulting in insufficient accuracy of generated interaction features and affecting the reliability of subsequent large-scale model inference. Finally, existing systems lack model self-evolution capabilities. They cannot update the large-scale model parameters based on the deviation values ​​when future frames actually occur, making it difficult for the system's inference accuracy to continuously improve with use.

[0005] Therefore, there is an urgent need for a conference terminal interaction system and method based on a large model to achieve intelligent iterative correction and forward-looking guidance of interactive content. Summary of the Invention

[0006] To address the aforementioned issues, this invention provides a conference terminal interaction system and method based on a large model, enabling intelligent iterative correction and forward-looking guidance of interactive content.

[0007] A first aspect of the present invention provides a conference terminal interaction system based on a large model, comprising: Setting the conference terminal within a conference space with preset layout data also includes: The data acquisition module is configured to limit the meeting space to at least one emotion discrimination space according to the preset arrangement data, and collect multimodal emotion data in the emotion discrimination space when passing through the first interactive interface of the meeting terminal; The first data processing module is configured to determine dynamic interactive features relating meeting content to emotions based on the multimodal emotion data and the first interactive interface. The second data processing module is configured to generate the first interactive interface and the second interactive interface in the conference terminal, render the dynamic interactive features to the second interactive interface, call the large model stored on the server to generate correction judgment items based on the dynamic interactive features, present the correction judgment items to the second interactive interface, and iteratively correct the interactive content in the first interactive interface based on the correction judgment items. The server is configured to store at least one large model; The storage module is configured to store the interaction data between the server and the conference terminal.

[0008] As a preferred embodiment, the multimodal emotion data includes first modal data and second modal data, wherein the first modal data is facial recognition emotion value, the second modal data is action recognition emotion value, and the server also pre-stores a set of non-emotional noise data. The data acquisition module is further configured to compare the collected facial recognition emotion values ​​and action recognition emotion values ​​with the non-emotional noise data set, remove facial recognition emotion values ​​and action recognition emotion values ​​with a matching degree higher than a preset noise threshold, and generate effective multimodal emotion data.

[0009] As a preferred approach, if the preset arrangement data defines at least two emotion discrimination spaces with different weight coefficients; The data acquisition module is further configured to perform weighted calculations on the multimodal emotion data collected in each emotion discrimination space according to the weight coefficients, generate weighted multimodal emotion data, and determine the dynamic interactive features of the association between meeting content and emotion based on the weighted multimodal emotion data and the first interactive interface.

[0010] As a preferred embodiment, the interactive content in the first interactive interface is a frame sequence arranged in a time sequence, the frame sequence including the current frame content and the future frame content; The correction judgment item includes a current correction judgment item corresponding to the current frame content and a future correction judgment item corresponding to the future frame content. The second data processing module generates the future correction judgment item by taking the dynamic interaction features corresponding to the current frame content and the current correction judgment item as inference input.

[0011] As a preferred embodiment, the second data processing module is configured to, when calling the large model, be: Construct reasoning instructions that include time evolution prompts, analyze the changing trends of the dynamic interaction features of the current frame content, combine the correction logic of the current correction judgment item, deduce the emotional evolution path of future time, and generate the future correction judgment item that matches the future frame content.

[0012] As a preferred embodiment, the second interactive interface includes a real-time correction area and a prediction preview area; the second data processing module is configured to render the current correction judgment item to the real-time correction area in real time, and to output the future correction judgment item to the prediction preview area in advance; The prediction preview area is configured to display the future correction judgment item in a visual style different from that of the real-time correction area, so as to identify it as predictive correction information.

[0013] As a preferred approach, when the meeting progresses to the moment corresponding to the future frame content, the second data processing module is configured to convert the future correction judgment item into a new current correction judgment item; The second data processing module is also configured to collect multimodal emotion data when the future frame content actually occurs, calculate the deviation value between the actual dynamic interaction features and the inference input, and feed the deviation value back to the server to update the inference parameters of the large model.

[0014] As a preferred embodiment, the first data processing module is configured to timestamp the multimodal emotion data and align it with the content update time of the first interactive interface, extracting the correlation features between the emotion change rate and the content update rate as the dynamic interaction features. When the second data processing module calls the large model, it constructs feature vector prompts containing the dynamic interaction features. The large model outputs the correction judgment item containing the correction direction, correction magnitude, and correction priority. The second data processing module determines the execution order of the iterative correction based on the correction priority. The iterative correction includes calculating the deviation value between the interactive content and the target emotional state. When the deviation value is greater than a preset convergence threshold, the generation and presentation steps of the correction judgment item are continuously executed until the deviation value is less than or equal to the preset convergence threshold.

[0015] A second aspect of the present invention provides a conference terminal interaction method based on a large model, comprising the following steps: The meeting space is defined as at least one emotion discrimination space according to the preset layout data, and multimodal emotion data in the emotion discrimination space is collected when the first interactive interface of the meeting terminal is accessed. Based on the multimodal emotion data and the first interactive interface, determine the dynamic interactive features of the association between meeting content and emotions; Generate the first interactive interface and the second interactive interface, render the dynamic interactive features onto the second interactive interface, call the large model stored on the server to generate correction judgment items based on the dynamic interactive features, present the correction judgment items on the second interactive interface, and iteratively correct the interactive content in the first interactive interface based on the correction judgment items.

[0016] Compared with the prior art, the present invention has the following advantages: This invention utilizes a second data processing module to access a large model stored on a server, generates correction judgment items based on dynamic interaction characteristics, and presents them on a second interactive interface. Subsequently, the interactive content in the first interactive interface is iteratively corrected based on these correction judgment items. This allows the meeting content to be dynamically adjusted in real time according to the audience's emotions, significantly improving the accuracy and adaptability of meeting interactions.

[0017] This invention divides the interactive content in the first interactive interface into current frame content and future frame content, generating corresponding current correction judgment items and future correction judgment items. By setting a real-time correction area and a prediction preview area in the second interactive interface, and displaying the future correction judgment items with a visual style different from the real-time correction area, the speaker can not only grasp the current state, but also know in advance the emotional evolution path and correction suggestions for future time steps, realizing the transformation from "passive response" to "active prediction".

[0018] This invention collects facial and motion-based emotion values ​​through a data acquisition module, and compares these values ​​with a pre-stored set of non-emotional noise data on a server, removing data with a matching degree higher than a preset noise threshold. This data cleaning mechanism effectively filters out non-emotional interference signals, ensuring that the generated dynamic interaction features accurately reflect the correlation between meeting content and emotions.

[0019] This invention can weight multimodal emotion data according to the weight coefficients of different emotion discrimination spaces in the preset layout data, adapting to different meeting space layouts. Furthermore, when the meeting progresses to the moment corresponding to the content of a future frame, the system can calculate the deviation between the actual dynamic interaction features and the inferred input, and feed the deviation back to the server to update the inference parameters of the large model, achieving continuous self-learning and accuracy optimization of the system.

[0020] In this invention, the correction judgment item output by the large model includes the correction direction, correction magnitude, and correction priority. The second data processing module determines the execution order of iterative correction based on the correction priority and sets a preset convergence threshold as the iteration termination condition. This makes the correction process of interactive content more orderly and controllable, avoiding the problems of over-correction or under-correction, and ensuring that the interactive content eventually converges to the target emotional state. Attached Figure Description

[0021] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0022] Figure 1 This is a structural block diagram of the system provided in the embodiments of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] This invention provides a conference terminal interaction system based on a large model, such as... Figure 1 As shown, the system sets up conference terminals within a conference space with preset layout data. The conference space can be a physical meeting room or a virtual conference space. The preset layout data includes the physical layout of the conference space, seating arrangement, camera positions, and sensor configuration information; this data is used to define the spatial range and logical boundaries of data collection. The system mainly includes a data acquisition module, a first data processing module, a second data processing module, a server, and a storage module.

[0025] The data acquisition module is configured to define the meeting space as at least one emotion discrimination space based on preset layout data. Specifically, the system can divide the meeting space into multiple geometric areas based on the coverage of the cameras or the seating area, with each area serving as an emotion discrimination space. The data acquisition module collects multimodal emotion data within the emotion discrimination space when viewed through the first interactive interface of the meeting terminal. The multimodal emotion data includes first-modal data and second-modal data, where the first-modal data is facial recognition emotion values ​​and the second-modal data is action recognition emotion values. The quantification process for facial recognition emotion values ​​includes: capturing facial images of the audience through the camera, extracting the coordinates of facial feature points, calculating the displacement vector of the feature points relative to the baseline expression, and mapping the displacement vector to a numerical score for a preset emotion dimension (such as pleasure, arousal, and dominance). The quantification process for action recognition emotion values ​​includes: recognizing the audience's limb movements through skeletal keypoint detection, calculating the rate of change of joint angles, the amplitude and frequency of movements, and mapping these physical quantities to emotion scores.

[0026] To improve data accuracy, the server also pre-stores a set of non-emotional noise data. This set contains feature data vectors that are easily misidentified as emotional data but are actually unrelated to emotion, such as random motion vectors of viewers adjusting their posture, facial shadows caused by changes in lighting, and facial distortion caused by coughing. The data acquisition module is also configured to compare the feature vectors of the collected facial recognition emotion values ​​and motion recognition emotion values ​​with the feature vectors in the non-emotional noise data set. The comparison process can use the calculation of cosine similarity or Euclidean distance between vectors. If the similarity between the collected data feature vector and a certain type of data feature vector in the noise set is higher than a preset noise threshold, the data is determined to be noise. The data acquisition module removes facial recognition emotion values ​​and motion recognition emotion values ​​with a matching degree higher than the preset noise threshold, generating valid multimodal emotion data. This process ensures that the data input to the large model is a valid signal that truly reflects the emotional state, rather than environmental noise or irrelevant actions.

[0027] Furthermore, the pre-defined data layout can include at least two emotion discrimination spaces with different weight coefficients. For example, the weight coefficient for key decision-making areas or front-row areas can be higher than that for ordinary areas. The data acquisition module is also configured to perform weighted calculations on the multimodal emotion data collected in each emotion discrimination space according to the weight coefficients, generating weighted multimodal emotion data. Specifically, the calculation method can be: multiplying the emotion data vectors collected in each space by their corresponding weight coefficients, summing them, and then normalizing the results. This allows the system to pay more attention to the feedback of key audience groups, making the generated dynamic interaction features more representative.

[0028] The first data processing module is configured to determine dynamic interaction features relating meeting content to emotions based on weighted multimodal sentiment data and the first interactive interface. Specifically, the first data processing module is configured to timestamp the multimodal sentiment data and align it with the content update timestamps of the first interactive interface. For example, when the first interactive interface switches to a meeting agenda or display page, the system records the timestamp of that moment and associates the sentiment data collected during that time period with that timestamp. The first data processing module extracts the correlation features between the sentiment change rate and the content update rate as dynamic interaction features. The quantification process includes: calculating the slope of the sentiment value change per unit time (i.e., the sentiment change rate), calculating the number of times the meeting content is switched or the information density (i.e., the content update rate) per unit time, and then calculating the Pearson correlation coefficient or mutual information value between the two. This correlation coefficient or mutual information value is used as the dynamic interaction feature vector, thereby quantifying the degree of influence of the meeting content on the audience's emotions.

[0029] The second data processing module is configured to generate a first interactive interface and a second interactive interface on the conference terminal. The first interactive interface is typically used to display the main content of the meeting or the current status assessment value, while the second interactive interface is mainly used to assist decision-making and display process data. The second data processing module renders dynamic interactive features onto the second interactive interface. Simultaneously, the second data processing module calls a large model stored on the server to generate correction judgment items based on the dynamic interactive features. When calling the large model, the second data processing module constructs feature vector prompts containing dynamic interactive features, and the large model outputs correction judgment items containing correction direction, correction magnitude, and correction priority. The correction direction can be quantified as a vector direction (e.g., increase interaction, decrease speech rate), the correction magnitude can be quantified as a numerical step size (e.g., decrease speech rate by 10%), and the correction priority can be quantified as a weight score. The second data processing module determines the execution order of iterative corrections based on the correction priority. The second data processing module presents the correction judgment items on the second interactive interface and iteratively corrects the interactive content in the first interactive interface based on the correction judgment items.

[0030] The core of this embodiment lies in a time-series-based prediction and correction mechanism. The interactive content in the first interactive interface is presented as a sequence of frames arranged in a time sequence. Here, "frame" does not refer only to video frames, but rather to a logical unit of the meeting content on the timeline, such as the display period of each presentation slide, the time window of each speech segment, or each agenda node. The frame sequence includes current frame content and future frame content. Current frame content refers to the meeting content currently being displayed or in progress on the meeting terminal, with the corresponding time window being from the current moment to the current moment plus a preset duration. Future frame content refers to the meeting content planned to be displayed at a later time, with the corresponding time window being the period after the current moment plus a preset duration.

[0031] The correction judgment terms include the current correction judgment terms corresponding to the current frame content and the future correction judgment terms corresponding to the future frame content. The second data processing module is configured to use the dynamic interaction features corresponding to the current frame content and the current correction judgment terms as inference inputs to generate future correction judgment terms. Specifically, when calling the large model, the second data processing module is configured to construct an inference instruction containing time evolution prompts. This instruction takes the dynamic interaction feature vector of the current frame, the current correction judgment term vector, and the time step encoding as the input sequence. The large model is configured to analyze the changing trend of the dynamic interaction features of the current frame content, such as whether the sentiment value is rising or falling, and at what rate, and combine it with the correction logic of the current correction judgment terms, such as whether the sentiment improved after suggesting speeding up the speech, to infer the sentiment evolution path for future time steps. Based on this evolution path, the large model generates future correction judgment terms that match the future frame content. This means that the system not only corrects the current state but also predicts the emotional reactions that future content may trigger and generates coping strategies in advance. Quantitatively, the future correction judgment terms can be represented as the probability distribution or expected value of future time steps.

[0032] The second interactive interface includes a real-time correction area and a prediction preview area. The second data processing module is configured to render the current correction decision in real-time to the real-time correction area and output future correction decisions in advance to the prediction preview area. The prediction preview area is configured to display future correction decisions with a visual style distinct from the real-time correction area to identify them as predictive correction information. For example, the real-time correction area can use a highlighted border or solid line with 100% opacity to indicate an instruction that needs to be executed immediately; the prediction preview area can use a dashed border, semi-transparent display, or 50% opacity to indicate a suggestion generated based on prediction, for the speaker to prepare in advance. This differentiated display method allows users to clearly distinguish between "operations that must be performed now" and "adjustments that may be needed in the future," avoiding information confusion. The visual style difference can be achieved through specific parameter differences in color hue, saturation, or border line style.

[0033] When the meeting progresses to the moment corresponding to the future frame content, the second data processing module is configured to convert the future correction judgment into a new current correction judgment. At this point, the previous prediction suggestion becomes the current execution instruction. To achieve system self-evolution, the second data processing module is also configured to collect multimodal sentiment data when the future frame content actually occurs, and calculate the deviation value between the actual dynamic interaction features and the inference input. Here, the inference input refers to the predicted feature state on which the future correction judgment was based. The deviation value can be calculated using the mean squared error or absolute error function, that is, calculating the distance between the predicted feature vector and the actual feature vector. The system feeds back the calculated deviation value to the server to update the inference parameters of the large model. For example, if the system predicts that the future frame will lead to low mood and suggests a certain correction, but the mood does not actually decrease when it occurs, the deviation value will be recorded and used to adjust the prediction weights of the large model for the future sentiment evolution path, for example, by updating the internal sentiment evolution parameter matrix of the model through gradient descent. This closed-loop feedback mechanism enables the large model to continuously learn the actual situation of the meeting, optimize the future prediction accuracy, and achieve continuous iterative optimization of the model.

[0034] Regarding the specific execution of iterative correction, iterative correction includes calculating the deviation between the interactive content and the target emotional state. The target emotional state can be a preset ideal meeting atmosphere, such as a specific level of focus, quantified as a specific target vector. When the deviation value is greater than a preset convergence threshold, the generation and presentation steps of correction judgment items are continuously executed until the deviation value is less than or equal to the preset convergence threshold. This ensures that the correction process does not proceed indefinitely, but stops after the expected effect is achieved, avoiding excessive intervention in the meeting process. The correction priority in the correction judgment item is used to determine the execution order when multiple correction suggestions exist simultaneously, such as prioritizing adjusting the speaking speed or prioritizing switching content. The priority can be determined by sorting by numerical value.

[0035] This invention also provides a conference terminal interaction method based on a large model, applied to a conference terminal within a conference space with preset layout data. The method includes the following steps: defining the conference space as at least one emotion discrimination space based on the preset layout data; collecting multimodal emotion data within the emotion discrimination space when passing through a first interactive interface of the conference terminal; determining dynamic interaction features relating conference content to emotions based on the multimodal emotion data and the first interactive interface; generating a first interactive interface and a second interactive interface; rendering the dynamic interaction features onto the second interactive interface; calling a large model stored on a server to generate correction judgment items based on the dynamic interaction features; presenting the correction judgment items on the second interactive interface; and iteratively correcting the interactive content in the first interactive interface based on the correction judgment items. Specific implementation details of this method can be found in the description of the above system embodiments, particularly regarding the inference logic between the current frame and future frames, the differentiation and display of interface areas, and the model update mechanism based on deviation values, all of which are applicable to this method embodiment.

[0036] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent variations only. Individual components and functions are optional unless explicitly required, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed combinations. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0037] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to achieve the described functions, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the described devices, apparatuses, and units can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0038] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, function, and operation of implementations of apparatus, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description; sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based device that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A conference terminal interaction system based on a large model, characterized in that, Setting the conference terminal within a conference space with preset layout data also includes: The data acquisition module is configured to limit the meeting space to at least one emotion discrimination space according to the preset arrangement data, and collect multimodal emotion data in the emotion discrimination space when passing through the first interactive interface of the meeting terminal; The first data processing module is configured to determine dynamic interactive features relating meeting content to emotions based on the multimodal emotion data and the first interactive interface. The second data processing module is configured to generate the first interactive interface and the second interactive interface in the conference terminal, render the dynamic interactive features to the second interactive interface, call the large model stored on the server to generate correction judgment items based on the dynamic interactive features, present the correction judgment items to the second interactive interface, and iteratively correct the interactive content in the first interactive interface based on the correction judgment items. The server is configured to store at least one large model; The storage module is configured to store the interaction data between the server and the conference terminal.

2. The conference terminal interaction system based on a large model according to claim 1, characterized in that, The multimodal emotion data includes first modal data and second modal data, wherein the first modal data is facial recognition emotion value, the second modal data is action recognition emotion value, and the server also pre-stores a set of non-emotional noise data. The data acquisition module is further configured to compare the collected facial recognition emotion values ​​and action recognition emotion values ​​with the non-emotional noise data set, remove facial recognition emotion values ​​and action recognition emotion values ​​with a matching degree higher than a preset noise threshold, and generate effective multimodal emotion data.

3. The conference terminal interaction system based on a large model according to claim 1, characterized in that, If the preset arrangement data defines at least two emotion discrimination spaces with different weight coefficients; The data acquisition module is further configured to perform weighted calculations on the multimodal emotion data collected in each emotion discrimination space according to the weight coefficients, generate weighted multimodal emotion data, and determine the dynamic interactive features of the association between meeting content and emotion based on the weighted multimodal emotion data and the first interactive interface.

4. The conference terminal interaction system based on a large model according to claim 1, characterized in that, The interactive content in the first interactive interface is a frame sequence arranged in time sequence, and the frame sequence includes the current frame content and the future frame content; The correction judgment item includes a current correction judgment item corresponding to the current frame content and a future correction judgment item corresponding to the future frame content. The second data processing module generates the future correction judgment item by taking the dynamic interaction features corresponding to the current frame content and the current correction judgment item as inference input.

5. The conference terminal interaction system based on a large model according to claim 4, characterized in that, The second data processing module is configured to: Construct reasoning instructions that include time evolution prompts, analyze the changing trends of the dynamic interaction features of the current frame content, combine the correction logic of the current correction judgment item, deduce the emotional evolution path of future time, and generate the future correction judgment item that matches the future frame content.

6. The conference terminal interaction system based on a large model according to claim 5, characterized in that, The second interactive interface includes a real-time correction area and a prediction preview area; The second data processing module is configured to render the current correction judgment item to the real-time correction area in real time, and to output the future correction judgment item to the prediction preview area in advance; The prediction preview area is configured to display the future correction judgment item in a visual style different from that of the real-time correction area, so as to identify it as predictive correction information.

7. The conference terminal interaction system based on a large model according to claim 6, characterized in that, When the meeting progresses to the moment corresponding to the future frame content, the second data processing module is configured to convert the future correction judgment item into a new current correction judgment item; The second data processing module is also configured to collect multimodal emotion data when the future frame content actually occurs, calculate the deviation value between the actual dynamic interaction features and the inference input, and feed the deviation value back to the server to update the inference parameters of the large model.

8. The conference terminal interaction system based on a large model according to claim 7, characterized in that, The first data processing module is configured to timestamp the multimodal emotion data and align it with the content update time of the first interactive interface, and extract the correlation features between the emotion change rate and the content update rate as the dynamic interaction features. When the second data processing module calls the large model, it constructs feature vector prompts containing the dynamic interaction features. The large model outputs the correction judgment item containing the correction direction, correction magnitude, and correction priority. The second data processing module determines the execution order of the iterative correction based on the correction priority. The iterative correction includes calculating the deviation value between the interactive content and the target emotional state. When the deviation value is greater than a preset convergence threshold, the generation and presentation steps of the correction judgment item are continuously executed until the deviation value is less than or equal to the preset convergence threshold.

9. A conference terminal interaction method based on a large model, applied to a conference terminal in a conference space with preset layout data, characterized in that, Includes the following steps: The meeting space is defined as at least one emotion discrimination space according to the preset layout data, and multimodal emotion data in the emotion discrimination space is collected when the first interactive interface of the meeting terminal is accessed. Based on the multimodal emotion data and the first interactive interface, determine the dynamic interactive features of the association between meeting content and emotions; The first interactive interface and the second interactive interface are generated. The dynamic interactive features are rendered onto the second interactive interface. The large model stored on the server is called to generate correction judgment items based on the dynamic interactive features. The correction judgment items are presented on the second interactive interface. The interactive content in the first interactive interface is iteratively corrected based on the correction judgment items.

Citation Information

Patent Citations

  • Intelligent conference multi-modal interaction optimization method and system based on large model

    CN121438815A

  • Digital large screen interaction method and device based on multi-agent cooperation and medium

    CN121657869A

  • System for real-time analysis of emotional feedback during motivational presentations

    DE202025104706U1