Artificial intelligence-based robot control method and system for conference investigation
By generating a heatmap of speech distribution by acquiring audio and monitoring data, the problem of low efficiency in meeting surveys in existing technologies is solved, and an adaptive interaction mechanism is implemented to improve the efficiency and accuracy of meetings.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 上海万怡医学科技股份有限公司
- Filing Date
- 2025-10-31
- Publication Date
- 2026-05-12
AI Technical Summary
Existing conference research methods rely on static equipment and manual analysis, resulting in low efficiency, inability to adapt to dynamic conference environments, and compromised decision-making quality.
By acquiring audio and monitoring data from the meeting environment, analyzing the speaking status and content, generating a heatmap of speaking distribution, determining the robot's proactive interaction strategy, and realizing an adaptive interaction mechanism.
提升了会议的效率和准确性,实现了会议讨论的实时可视化和参与均衡性监控,消除了对人工干预的延迟。
Smart Images

Figure CN121223797B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control technology, and in particular to artificial intelligence-based robot control methods and systems for conference research. Background Technology
[0002] With the rapid development of artificial intelligence technology, the efficient and accurate collection and analysis of participants' opinions, emotions, and engagement is crucial in meeting research scenarios. Meeting research plays a key role in modern organizations, helping to efficiently gather participants' perspectives, opinions, and feedback in scenarios such as decision-making, problem-solving discussions, or team collaboration.
[0003] However, existing meeting research methods mainly rely on static equipment for data collection and manual processing and analysis, resulting in low meeting efficiency, compromised decision-making quality, and an inability to adapt to the needs of dynamic meeting environments. Summary of the Invention
[0004] This application provides a method and system for controlling robots used in conference research based on artificial intelligence, in order to solve the above-mentioned problems.
[0005] In a first aspect, this application provides a method for controlling a robot used in conference research based on artificial intelligence, the method comprising:
[0006] Acquire audio and monitoring data of the meeting environment;
[0007] Analyze the monitoring data and the audio data to determine the speaking status and speaking content;
[0008] Based on the speaking status and speaking content, a speaking distribution heatmap is determined;
[0009] Based on the content of the speeches and the heatmap of the speech distribution, a proactive interaction strategy for the robot is determined and sent to the robot so that the robot can execute the proactive interaction strategy to guide the meeting discussion.
[0010] This solution acquires audio and monitoring data from the meeting environment, eliminating the limitations of relying solely on single audio data points and capturing meeting dynamics from multiple perspectives. Analyzing the monitoring and audio data determines speaking status and content, enabling multi-dimensional data fusion analysis and improving the accuracy and real-time nature of intent understanding. This provides structured input for heatmap generation and interaction strategies. Based on speaking status and content, a speaking distribution heatmap is generated, achieving real-time visualization of meeting discussion distribution, overcoming the inability to automatically generate heatmaps and improving the monitoring capability for participation balance. Based on the speaking content and distribution heatmap, a proactive robot interaction strategy is determined and sent to the robot to guide meeting discussions, achieving an adaptive interaction mechanism that eliminates delays caused by manual intervention and ensures efficient and precise meeting progress.
[0011] Optionally, the step of analyzing the monitoring data and the audio data to determine the speaking status includes:
[0012] Analyze the monitoring data to determine the facial muscle movement characteristics and the speaker's position coordinates;
[0013] Based on the aforementioned motion characteristics, the type of micro-expression is determined;
[0014] Analyze the audio data to determine the voiceprint characteristics and content of the speech;
[0015] Based on the aforementioned voiceprint characteristics, the speaker's identity is determined by matching them against the meeting record database.
[0016] The micro-expression type, the location coordinates, and the speaker's identity are determined as the speaker's speaking status.
[0017] This solution analyzes monitoring data to determine the speaker's facial muscle movement characteristics and positional coordinates, ensuring that nonverbal cues can be quantified and addressing the issue of neglecting visual data. It also provides real-time speaker location tracking, supporting the identification of low-activity areas. Based on movement characteristics, it determines micro-expression types, enhancing the accuracy of intent recognition. Audio data is parsed to determine speaker voiceprint characteristics and content, providing input data for speaker identification. Based on the voiceprint characteristics, it matches the meeting record database to determine the speaker's identity, resolving the speaker identification problem and providing data for the identity field of the speaking status. Finally, it defines the speaker's speaking status based on micro-expression type, positional coordinates, and speaker identity, achieving multi-dimensional data fusion.
[0018] Optionally, determining the speech distribution heatmap based on the speech status and speech content includes:
[0019] Based on the aforementioned location coordinates, a set of location coordinates for all speakers is obtained;
[0020] Based on the preset conference room plan grid, the set of location coordinates is mapped to the grid partitions;
[0021] Based on the content of the speech, the total speaking time and the number of speeches per unit time in each grid partition are calculated;
[0022] By weighting the number of messages and the total duration of the messages, a heatmap of the activity level of the partition is generated.
[0023] This solution obtains a set of location coordinates for all speakers based on their location coordinates, ensuring the completeness and operability of speaker location information and providing a foundation for mapping to grid partitions. According to a pre-defined conference room grid, the set of location coordinates is mapped to grid partitions, providing a spatial framework for statistical analysis of speaking data in each partition and ensuring that speaking distribution analysis is based on a unified partitioning standard. Based on the speaking content, the total speaking time and the number of speaking sessions per unit time within each grid partition are calculated, quantifying the speaking activity index of each grid partition and providing specific numerical input for generating heatmaps. Weighted averages of speaking sessions and total speaking time are used to generate a partition activity heatmap, visually displaying the speaking activity levels in different areas of the meeting and achieving a visual representation of the discussion distribution.
[0024] Optionally, determining the robot's proactive interaction strategy based on the speech content and the speech distribution heatmap includes:
[0025] Analyze the content of the speech to identify topic keywords;
[0026] Based on the heatmap of the speech distribution, the discussion duration and number of speakers for the topic keywords are determined;
[0027] If the number of speakers per unit time is lower than the preset value and the discussion time exceeds the threshold, it is marked as an open topic;
[0028] When an open-ended topic exceeds the time limit, the content of the speech is analyzed to identify fragmented viewpoints.
[0029] Clustering the fragmented viewpoints generates several structured advancement options.
[0030] This solution analyzes speech content to identify topic keywords, addressing the limitation of not being able to analyze topic keywords based on speech content, and providing an analytical anchor for proactive strategies. Based on a speech distribution heatmap, it determines the discussion duration and number of speakers for each topic keyword, outputting a numerical indicator for each keyword as input for open-ended topic determination. If the number of speakers per unit time is lower than a preset value and the discussion duration exceeds a threshold, it is marked as an open-ended topic, triggering fragmented viewpoint analysis to avoid the problem of open-ended topics continuously exceeding the limit without being processed. When the duration of an open-ended topic continuously exceeds the limit, the speech content is analyzed to identify fragmented viewpoints, addressing the deficiency of fragmented viewpoints not converging. Fragmented viewpoints are clustered to generate several structured advancement options, fulfilling the requirements for generating follow-up questions or clustering viewpoints.
[0031] Optionally, determining the robot's proactive interaction strategy based on the speech content and the speech distribution heatmap further includes:
[0032] Analyze the content of the speech to determine the viewpoint;
[0033] Analyze the speaking status to determine the degree of matching between the speaking status and the speaking viewpoint;
[0034] If the matching degree is lower than the preset matching value, a contradiction marker is generated, and the activity level of the discussion in the partition is determined according to the heat map of the speech distribution.
[0035] Based on the aforementioned contradiction markers, and according to the stated viewpoints and the activity level of the discussion in each partition, different follow-up questioning strategies are generated for each partition.
[0036] This solution analyzes speech content, identifies viewpoints, and extracts core intentions to provide input for matching analysis, addressing the problem of not being able to deduce key viewpoints from speech content. It analyzes speech status to determine the matching degree between the status and viewpoints, achieving multi-dimensional data fusion and resolving the inability to assess the consistency between speech content and facial expressions, accurately identifying potential contradictions. If the matching degree is lower than a preset value, a contradiction marker is generated, and the activity level of discussion zones is determined based on a speech distribution heatmap, providing visual data support for discussion distribution and addressing the problem of missing viewpoints due to the inability to identify low-activity areas, laying the foundation for strategic zoned push notifications. Based on the contradiction markers, and according to the viewpoints and discussion activity levels of different zones, follow-up questioning strategies are generated for different zones, addressing the problem of guidance failure due to the lack of proactive interaction mechanisms, and achieving closed-loop feedback.
[0037] Optionally, analyzing the speaking state and determining the matching degree between the speaking state and the speaking viewpoint includes:
[0038] Obtain a preset micro-expression-semantic mapping table, which stores the correspondence between words and positive micro-expressions;
[0039] Analyze the types of micro-expressions in the speaking state to determine the actual polarity of the micro-expressions;
[0040] Based on the micro-expression-semantic mapping table, keywords in the expressed opinions are analyzed to determine the expected micro-expression polarity;
[0041] If the actual micro-expression polarity is inconsistent with the expected micro-expression polarity, then the matching degree is determined to be lower than the preset matching value.
[0042] This solution obtains a pre-defined micro-expression-semantic mapping table, storing the correspondence between words and positive micro-expressions. This avoids real-time calculations or reliance on external data, standardizing and repeating the process of determining the expected polarity. It analyzes the types of micro-expressions in the speech state to determine the actual micro-expression polarity, ensuring processing speed and reliability, and capturing the speaker's nonverbal cues as a key basis for matching degree calculation. Based on the micro-expression-semantic mapping table, it parses keywords in the speech's viewpoint to determine the expected micro-expression polarity, standardizing the emotional interpretation of the speech's viewpoint, avoiding semantic ambiguity, and providing a consistent benchmark for comparison with the actual polarity, ensuring efficient execution and providing expected reference values for matching degree determination. If the actual micro-expression polarity is inconsistent with the expected micro-expression polarity, the matching degree is determined to be lower than the pre-defined matching value, achieving a quantitative assessment of emotional consistency, solving the problem of being unable to assess the consistency between speech content and facial expressions, and ensuring real-time response.
[0043] Optionally, the clustering of the fragmented viewpoints generates several structured advancement options, including:
[0044] Extract entity words and action words from the fragmented viewpoints;
[0045] Analyze the entity words to determine the degree of association between at least two entity words;
[0046] Analyze the action words and entity words to determine the action relationship between any action word and any entity word;
[0047] A semantic graph is constructed based on the correlation degree, where nodes represent entity words and edges represent action relationships;
[0048] Based on the semantic graph, entities with the same parent node are merged to generate candidate viewpoint clusters;
[0049] The frequency of occurrence of each candidate viewpoint cluster is counted, and several structured advancement options are obtained based on the statistical results.
[0050] This solution extracts entity words and action words from fragmented viewpoints, focusing on core vocabulary and ignoring irrelevant text to improve processing efficiency and accuracy, avoiding clustering bias caused by text noise. Entity words are analyzed to determine the correlation between at least two entity words, ensuring the semantic graph accurately reflects the tightness of entity word density, thus reducing interference from irrelevant entities and improving clustering accuracy. Action words and entity words are analyzed to determine the action relationship between any action word and any entity word, ensuring the semantic graph represents the dynamic semantic structure of viewpoints and supports entity merging. A semantic graph is constructed based on correlation, where nodes represent entity words and edges represent action relationships, enabling the semantic structure and interaction relationships in fragmented viewpoints to be represented digitally, providing a unified framework for entity merging and simplifying clustering operations. Based on the semantic graph, entities with the same parent node are merged to generate candidate viewpoint clusters, achieving preliminary clustering of viewpoints and integrating scattered entity words into coherent groups, thereby reducing redundancy. The frequency of occurrence of each candidate viewpoint cluster is statistically analyzed, and based on the statistical results, several structured advancement options are obtained, converging fragmented viewpoints into actionable options, thus supporting meeting-guided decision-making.
[0051] Optionally, after obtaining several structured advancement options based on statistical results, the method further includes:
[0052] Analyze the semantic conflict degree between any two structured advancement options;
[0053] If there are semantically conflicting option pairs, generate a conflict warning flag;
[0054] Based on the aforementioned heatmap of speech distribution, locate the partition where the supporters of the conflict warning markers are located;
[0055] Based on the fragmented viewpoints, generate debate guidance strategies tailored to the supporters' respective sections.
[0056] This solution analyzes the semantic conflict between any two structured advancement options, addressing the issue of missed contradictions due to the lack of proactive interaction mechanisms, and providing data for generating conflict warning markers. If a pair of options exhibits semantic conflict, a conflict warning marker is generated, fulfilling the prerequisite for generating follow-up questioning strategies and transforming abstract conflict into actionable machine instruction trigger signals. Based on a heatmap of speech distribution, the solution locates the partitions where supporters of conflict warning markers reside, fulfilling the requirement of heatmap-based support area location. Based on fragmented viewpoints, a debate guidance strategy is generated for the partitions where supporters reside, addressing the shortcomings of passive recording without guidance.
[0057] Optionally, the step of generating follow-up questioning strategies for different partitions based on the contradiction markers, the expressed viewpoints, and the partition discussion activity includes:
[0058] When a contradiction marker is detected, the stated viewpoints are analyzed to determine the points of contention.
[0059] Based on the activity level of the discussion in each partition, locate the low-activity zone to which the contradiction marker belongs;
[0060] The problematic points are directed to a high-activity area, and follow-up questions are generated in the high-activity area.
[0061] Receive feedback audio data from high-activity areas, parse the feedback audio data, and generate a feedback summary;
[0062] The feedback summary is then pushed to the original low-activity area.
[0063] This solution analyzes the viewpoints when a contradiction marker is detected, identifies the points of contention, and ensures that the follow-up questioning strategy is based on a clear and actionable topic of contention. This prevents the strategy from becoming ineffective due to ambiguity or generalization, directly supporting the accuracy of the follow-up questioning strategy. Based on the activity level of the discussion in each area, the solution locates the low-activity zone to which the contradiction marker belongs, identifying areas with insufficient participation. This ensures that the strategy is designed specifically for this low-activity zone, avoiding wasted resources on non-target areas, thus establishing a zone-level positioning foundation for the follow-up questioning strategy. The points of contention are then pushed to the high-activity zone, generating follow-up questioning instructions for that zone. This activates the discussion focus in the high-activity zone, guiding participants to provide targeted feedback and laying the data foundation for feedback collection, thereby enabling the follow-up questioning strategy to be executed in the high-activity zone. The solution receives audio feedback data from the high-activity zone, parses the audio data, generates feedback summaries, extracts the essence of the feedback, and transforms it into concise and easily disseminated summaries, avoiding information overload and simplifying the information in the follow-up questioning strategy. The feedback summaries are then pushed back to the original low-activity zone, stimulating participants in the low-activity zone to reflect on or respond to the viewpoints in the high-activity zone, promoting the activation of the discussion in that zone, thus completing the closed-loop guidance of the follow-up questioning strategy.
[0064] Secondly, this application provides an artificial intelligence-based control system for a robot used in conference research, the system comprising:
[0065] The data acquisition module is used to acquire audio data and monitoring data from the meeting environment;
[0066] The data analysis module is used to analyze the monitoring data and the audio data to determine the speaking status and speaking content;
[0067] A heatmap generation module is used to determine a heatmap of speech distribution based on the speech status and speech content;
[0068] The strategy generation module is used to determine the robot's proactive interaction strategy based on the content of the speech and the heatmap of the speech distribution, and send it to the robot so that the robot can execute the proactive interaction strategy to guide the meeting discussion. Attached Figure Description
[0069] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0070] Figure 1 This is a schematic diagram of an application scenario provided in an embodiment of this application.
[0071] Figure 2 A flowchart of an AI-based robot control method for conference research provided in one embodiment of this application.
[0072] Figure 3 This is a schematic diagram of the structure of an AI-based robot control system for conference research, provided as an embodiment of this application. Detailed Implementation
[0073] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0074] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0075] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.
[0076] Existing meeting research methods mainly rely on static equipment for data collection and manual processing and analysis, resulting in low meeting efficiency, compromised decision-making quality, and an inability to adapt to the needs of dynamic meeting environments.
[0077] Based on this, this application provides an AI-based robot control method and system for conference surveys. This method acquires audio and monitoring data of the conference environment, eliminating the limitations of relying solely on single audio data points, and thus capturing conference dynamics from multiple perspectives. By analyzing the monitoring and audio data, the method determines the speaking status and content, achieving multi-dimensional data fusion analysis, improving the accuracy and real-time nature of intent understanding, and providing structured input for heatmap generation and interaction strategies. Based on the speaking status and content, a speaking distribution heatmap is determined, enabling real-time visualization of the conference discussion distribution, eliminating the inability to automatically generate heatmaps, and improving the monitoring capability for participation balance. Based on the speaking content and the speaking distribution heatmap, a robot proactive interaction strategy is determined and sent to the robot to guide the conference discussion, achieving an adaptive interaction mechanism, eliminating the delays caused by reliance on manual intervention, and ensuring efficient and precise conference progress.
[0078] Figure 1 This is a schematic diagram illustrating an application scenario provided by this application, showing the application of the method provided in this application during a conference survey.
[0079] Specifically, the method provided in this application can be applied to any server. The server interacts with audio acquisition devices and monitoring acquisition devices, acquiring audio data of the meeting environment through the audio acquisition devices and monitoring data of the meeting environment through the monitoring acquisition devices. The monitoring data and audio data are analyzed to determine the speaking status and content. Based on the speaking status and content, a speaking distribution heatmap is generated, enabling real-time visualization of the meeting discussion distribution, eliminating the inability to automatically generate heatmaps, and improving the monitoring capability for participation balance. Based on the speaking content and the speaking distribution heatmap, a robot proactive interaction strategy is determined and sent to the robot to guide the meeting discussion, achieving an adaptive interaction mechanism, eliminating delays caused by reliance on manual intervention, and ensuring efficient and accurate meeting progress.
[0080] For specific implementation details, please refer to the following examples.
[0081] Figure 2 This is a flowchart illustrating an AI-based robot control method for conference research, provided as an embodiment of this application. The method of this embodiment can be applied to servers in the above scenarios. Figure 2 As shown, the method includes:
[0082] S201. Obtain audio data and monitoring data of the meeting environment;
[0083] The meeting environment can be the actual physical space where the meeting is held.
[0084] Audio data can be sound signals acquired through audio acquisition devices in a conference environment.
[0085] Monitoring data can be visual signals acquired through monitoring and acquisition devices in a conference environment.
[0086] Specifically, audio acquisition devices and monitoring acquisition devices are deployed in the meeting environment; then, the audio acquisition devices acquire the sound signals of the meeting environment to form audio data; at the same time, the monitoring acquisition devices acquire video signals to form monitoring data.
[0087] S202. Analyze monitoring data and audio data to determine the speaking status and speaking content;
[0088] Speaking status can be the speaker's real-time status information, including the speaker's identity, micro-expression type, and location coordinates.
[0089] The content of a speech can be written material used to express the speaker's linguistic intentions.
[0090] Specifically, facial detection technology is used to identify speakers in the monitoring data and extract their identity (the speaker's unique identifier, such as ID number or name). Next, micro-expression analysis algorithms are used to extract facial regions from the monitoring data and detect the speaker's micro-expression types (such as frowning, smiling, etc.). At the same time, location tracking algorithms are used to locate the speaker and determine their position coordinates (specific coordinate values, such as two-dimensional plane coordinates). Based on the speaker's identity, micro-expression type, and position coordinates, the speaking status is determined.
[0091] The audio data is preprocessed using speech recognition algorithms (such as noise reduction and frame segmentation); then the preprocessed audio data is converted into text to form the spoken content (such as "I agree").
[0092] S203. Determine the heat map of speech distribution based on the speaking status and content;
[0093] A heatmap of speech distribution can be a dynamic visualization chart generated based on activity values.
[0094] Specifically, the conference environment plane is divided into multiple grid cells, each corresponding to a spatial region; then, the set of position coordinates in the speaking state is mapped to the grid cells, and the speaking content is assigned to the corresponding grid cell according to the speaker's coordinate values.
[0095] The number of times a message is spoken within a grid cell (counting the start time of the message content) and the total duration of the message are counted. A weighted average is applied to calculate the number of times a message is spoken and the total duration of the message, resulting in an activity value for each grid cell (e.g., high activity, low activity). A graphics rendering tool is used to map the activity value to a color gradient (e.g., high activity areas are displayed in red, and low activity areas are displayed in blue) to form a heatmap of the message distribution.
[0096] S204. Based on the content of the speeches and the heat map of the speech distribution, determine the robot's proactive interaction strategy and send it to the robot so that the robot can execute the proactive interaction strategy to guide the meeting discussion.
[0097] A robot's proactive interaction strategy can be a guiding instruction generated based on the content and distribution of speeches in a heatmap, which is then sent to the robot to proactively guide the meeting discussion.
[0098] Specifically, based on the content of the speeches (such as contradictions or open-ended topics) and combined with a speech distribution heatmap (such as low-activity grid cells), the robot's proactive interaction strategy is determined. For example, for a mismatch between the speech content and micro-expressions, a follow-up question is generated (such as "Please clarify your true opinion"); for low-activity areas in the speech distribution heatmap, a guiding question is generated (such as "Please ask members in that area to share their views"); for fragmented speech content (speech fragments composed of scattered, incoherent phrases or sentences), a clustering instruction is generated (such as "Summarize the current discussion topic"). These instructions are then sent to the robot via a wireless communication protocol. After receiving the instructions, the robot executes the proactive interaction strategy; for example, the robot plays the follow-up question through a speaker or displays the guiding question on the meeting screen. Finally, the robot provides real-time feedback on the execution results, thus guiding the meeting discussion.
[0099] This solution acquires audio and monitoring data from the meeting environment, eliminating the limitations of relying solely on single audio data points and capturing meeting dynamics from multiple perspectives. Analyzing the monitoring and audio data determines speaking status and content, enabling multi-dimensional data fusion analysis and improving the accuracy and real-time nature of intent understanding. This provides structured input for heatmap generation and interaction strategies. Based on speaking status and content, a speaking distribution heatmap is generated, achieving real-time visualization of meeting discussion distribution, overcoming the inability to automatically generate heatmaps and improving the monitoring capability for participation balance. Based on the speaking content and distribution heatmap, a proactive robot interaction strategy is determined and sent to the robot to guide meeting discussions, achieving an adaptive interaction mechanism that eliminates delays caused by manual intervention and ensures efficient and precise meeting progress.
[0100] In some embodiments, monitoring data is analyzed to determine the facial muscle movement characteristics and the speaker's position coordinates; based on the movement characteristics, the micro-expression type is determined; audio data is parsed to determine the voiceprint characteristics and content of the speech; based on the voiceprint characteristics, the meeting record database is matched to determine the speaker's identity; and the micro-expression type, position coordinates, and speaker identity are determined as the speaker's speaking state.
[0101] The speaker's facial muscles can be the muscle groups in the speaker's facial region during the meeting. Motion features can be numerical features quantifying the speaker's facial muscle movements. The speaker can be an individual participant speaking at the meeting. Location coordinates can be the speaker's location data within the meeting room. Micro-expression type can be an expression category label mapped from facial motion features. Voiceprint features can be unique identifiers of the speaker's voice characteristics. The meeting record database can be a data repository containing the voiceprint features of pre-registered speakers and their corresponding identities. Speaker identity can be a unique identifier for the speaker.
[0102] Specifically, the monitoring data is analyzed, and facial key point tracking algorithms are applied to analyze the movement trajectory of facial key points (such as the coordinate changes of eyebrows, corners of the mouth, or corners of the eyes) and determine them as the movement characteristics of the speaker's facial muscles. At the same time, based on the speaker's position information in the monitoring data, the speaker's position information is mapped to the meeting environment plane coordinate system (two-dimensional plane layout of the meeting room) preset by the preset meeting room layout through coordinate transformation algorithms (such as perspective transformation) to determine the position coordinates.
[0103] Motion features are input into a micro-expression classification model pre-trained based on machine learning theory of computer vision (which stores the mapping relationship between motion features and micro-expression types) to determine the type of micro-expression. For example, when motion features show eyebrows lowered and corners of the mouth pulled down, the micro-expression classification model classifies it as confused; when motion features show corners of the mouth raised and muscles around the eyes contracted, the micro-expression classification model classifies it as smiling.
[0104] The audio data is analyzed to extract the spectral features of the speech signal (frequency domain features of the speech signal, used to generate speech voiceprint features), and these features are then generated. Simultaneously, a speech recognition algorithm is used to convert the audio data into text, obtaining the speech content. The speech voiceprint features are matched against a meeting record database (which stores the voiceprint features of pre-registered speakers and their corresponding identities, such as name and ID) to identify similar voiceprint features. Based on these similar features, the speaker's identity is determined. Micro-expression type, location coordinates, and speaker identity are combined into a structured data object to define the speaker's speaking state.
[0105] This solution analyzes monitoring data to determine the speaker's facial muscle movement characteristics and positional coordinates, ensuring that nonverbal cues can be quantified and addressing the issue of neglecting visual data. It also provides real-time speaker location tracking, supporting the identification of low-activity areas. Based on movement characteristics, it determines micro-expression types, enhancing the accuracy of intent recognition. Audio data is parsed to determine speaker voiceprint characteristics and content, providing input data for speaker identification. Based on the voiceprint characteristics, it matches the meeting record database to determine the speaker's identity, resolving the speaker identification problem and providing data for the identity field of the speaking status. Finally, it defines the speaker's speaking status based on micro-expression type, positional coordinates, and speaker identity, achieving multi-dimensional data fusion.
[0106] In some embodiments, a set of location coordinates of all speakers is obtained based on location coordinates; the set of location coordinates is mapped to grid partitions according to a preset conference room planar grid; the total speaking time and the number of speaking times per unit time are counted in each grid partition according to the speaking content; and a heatmap of partition activity is generated by weighting the number of speaking times and the total speaking time.
[0107] The location coordinate set can be an array data structure consisting of the location coordinates of all speakers. The preset meeting room plan grid can be a predefined two-dimensional partitioning system that divides the meeting room plan into multiple grid cells, pre-stored on the server and invoked when needed. A grid partition can be a single sub-cell within the preset meeting room plan grid. The total speaking time can be the sum of the speaking times of all speakers within a grid partition. The number of speaking events can be the number of speaking events occurring in different grid partitions per unit time. The partition activity heatmap can be a visual image displaying the activity level of each partition within the meeting room plan grid.
[0108] Specifically, the process iterates through the speaking states of all speakers, extracting the position coordinates for each state; these position coordinates are then collected to form a set. A preset meeting room grid (a two-dimensional partitioning system that divides the meeting room floor into multiple grid cells, each representing a partition) is loaded; each position coordinate in the set is iterated through, and a coordinate mapping algorithm (a coordinate transformation method used to map each position coordinate to a grid partition) is used to map each position coordinate in the set to the grid partition. Then, the speaking duration of all speakers within each grid partition is summarized to obtain the total speaking duration; simultaneously, speakers are assigned to grid partitions, and the number of times they speak per unit time (e.g., per minute) within each grid partition is calculated.
[0109] The number of times a speaker speaks is weighted by the total speaking time. Then, based on the conference room's grid, each grid partition is assigned a color intensity (e.g., dark color indicates high activity and light color indicates low activity) to generate a partition activity heatmap, where each grid cell displays its corresponding activity value.
[0110] This solution obtains a set of location coordinates for all speakers based on their location coordinates, ensuring the completeness and operability of speaker location information and providing a foundation for mapping to grid partitions. According to a pre-defined conference room grid, the set of location coordinates is mapped to grid partitions, providing a spatial framework for statistical analysis of speaking data in each partition and ensuring that speaking distribution analysis is based on a unified partitioning standard. Based on the speaking content, the total speaking time and the number of speaking sessions per unit time within each grid partition are calculated, quantifying the speaking activity index of each grid partition and providing specific numerical input for generating heatmaps. Weighted averages of speaking sessions and total speaking time are used to generate a partition activity heatmap, visually displaying the speaking activity levels in different areas of the meeting and achieving a visual representation of the discussion distribution.
[0111] In some embodiments, the content of the speech is analyzed to determine the topic keywords; based on the speech distribution heatmap, the discussion duration and number of speakers for the topic keywords are determined; if the number of speakers per unit time is lower than a preset value and the discussion duration exceeds a threshold, it is marked as an open topic; when the duration of the open topic exceeds the limit, the content of the speech is analyzed to determine fragmented viewpoints; the fragmented viewpoints are clustered to generate several structured advancement options.
[0112] Topic keywords can be keywords extracted from the content of a speech.
[0113] Discussion duration can be the cumulative speaking time for different topic keywords.
[0114] The number of speakers can be the number of speakers who participated in discussions on the relevant keywords.
[0115] The preset value can be a pre-set threshold parameter used for comparing and analyzing the number of speakers. It is stored in the server and called when needed.
[0116] Open-ended topics can represent issues with low participation and high time consumption.
[0117] Fragmented viewpoints can be scattered and inconsistent expressions of opinion in open-ended topics.
[0118] The structured advancement options can be several guiding options generated after clustering fragmented viewpoints.
[0119] Specifically, traverse the speech content; segment the speech content and remove stop words (such as "de", "shi" in Chinese), then calculate the frequency or importance score of each word; select the word with the highest score as the topic keyword. For example, when the speech content is "The cost is too high, it is recommended to optimize the design; there is not enough time, a delay is needed", after segmenting and removing stop words, calculate the word frequency. The frequencies of "cost", "optimize", "time", and "delay" are relatively high, so they are extracted as topic keywords.
[0120] Traverse the speech distribution heat map to identify high-activity partitions; then, retrieve all speech events in the speech content that contain the topic keyword (a speech record in the speech content, including the speech text, speaker identity, start time, and end time of the speech); according to the speech events, extract the speaker's identity and speech timestamp (start time and end time) from the speech status; then accumulate the speech durations of all relevant speech events to obtain the discussion duration of the topic keyword. Furthermore, for each topic keyword, count the number of speakers who mention the keyword in the speech content. For example, for the topic keyword "cost" in the high-activity partition, the speech content shows that speakers A and B mention "cost", the cumulative discussion duration is 50 seconds, and the number of speakers is 2.
[0121] Check whether the number of speakers per unit time is lower than a preset value (flexibly configured according to the meeting scale, such as raising the preset value for large meetings); at the same time, check whether the discussion duration exceeds a threshold (such as setting a lower threshold for small meetings); if both conditions are met, mark the topic keyword as an open topic.
[0122] When the continuous duration of an open topic exceeds a preset overrun threshold (used to determine whether to trigger fragmented view analysis), extract all speech events related to the topic keyword in the speech content; then scan the speech events to identify sentences expressing views or opinions (such as sentences containing keywords like "think", "suggest", etc.), and regard them as fragmented views. For example, the open topic is "budget" with a continuous duration of 6 minutes (the overrun threshold is 5 minutes), and the relevant speech content is "It is recommended to increase the budget" and "Think that the cost should be reduced"; after analysis, the fragmented views are "It is recommended to increase the budget" and "Think that the cost should be reduced".
[0123] Convert the fragmented views into numerical vectors; use a clustering algorithm to cluster the numerical vectors and extract the central views or representative phrases in the clusters to form several structured promotion options. For example, the fragmented views are "insufficient resources", "time adjustment", "increase manpower", and after clustering, the structured promotion options are generated: Option 1: Optimize resource allocation (the cluster center contains "insufficient resources", "increase manpower"); Option 2: Adjust time management (the cluster center contains "time adjustment").
[0124] This solution analyzes speech content to identify topic keywords, addressing the limitation of not being able to analyze topic keywords based on speech content, and providing an analytical anchor for proactive strategies. Based on a speech distribution heatmap, it determines the discussion duration and number of speakers for each topic keyword, outputting a numerical indicator for each keyword as input for open-ended topic determination. If the number of speakers per unit time is lower than a preset value and the discussion duration exceeds a threshold, it is marked as an open-ended topic, triggering fragmented viewpoint analysis to avoid the problem of open-ended topics continuously exceeding the limit without being processed. When the duration of an open-ended topic continuously exceeds the limit, the speech content is analyzed to identify fragmented viewpoints, addressing the deficiency of fragmented viewpoints not converging. Fragmented viewpoints are clustered to generate several structured advancement options, fulfilling the requirements for generating follow-up questions or clustering viewpoints.
[0125] In some embodiments, the content of the speech is analyzed to determine the viewpoint; the status of the speech is analyzed to determine the degree of matching between the status and the viewpoint; if the degree of matching is lower than a preset matching value, a contradiction marker is generated, and the activity level of the discussion in the partition is determined according to the heat map of the speech distribution; based on the contradiction marker, different questioning strategies for different partitions are generated according to the viewpoint and the activity level of the discussion in the partition.
[0126] Opinions expressed can be subjective statements representing the speaker's opinions, stance, or suggestions.
[0127] Matching degree can be used to represent the degree of match between the type of micro-expression in a speech and the viewpoint expressed.
[0128] The preset matching value can be a fixed threshold used for comparison with the matching degree, which is stored in the server in advance and called when needed.
[0129] A contradiction marker can be used to record events where the viewpoint expressed is inconsistent with the state of the statement.
[0130] The activity level of a discussion can be quantified as the activity level of each section in the heatmap of speech distribution.
[0131] Different zones can be different areas within a pre-defined conference room grid.
[0132] The probing strategy can be based on contradiction markers and the activity level of partitioned discussions, which can be used as instructions for the robot's proactive interaction.
[0133] Specifically, the process iterates through each message record in the speech content (including the speech text, speaker identity, and speech timestamp); uses a text analysis algorithm to scan the speech text and identify statements containing keywords expressing opinions (such as "I think," "suggest," "agree," and "disagree"); and extracts these statements as the speech opinions. For example, if the speech content is "I think the budget should be increased, but there is not enough time," then the statement "I think the budget should be increased" is identified as the speech opinion.
[0134] For each viewpoint expressed, the corresponding micro-expression type (such as smiling, frowning, or confused) is extracted from the speaking state. Then, based on the micro-expression type, the matching degree between the speaking state and the speaking viewpoint is determined through preset rules. For example, when the speaking viewpoint is positive (such as agreeing or supporting) and the micro-expression is positive (such as smiling), the matching degree is high; if the speaking viewpoint is positive but the micro-expression is negative (such as frowning), the matching degree is low.
[0135] The matching score is compared with a preset matching value (used for comparison to generate contradiction markers) based on historical meeting data. If the matching score is lower than the preset matching value, a contradiction marker is generated (containing speaker identity, viewpoint, micro-expression type, and matching score value, used to identify inconsistencies between the speech content and the speaking state). The speech distribution heatmap is analyzed to extract the activity level (e.g., high, medium, low) of each partition (a region in the preset meeting room grid) as the partition discussion activity level. The viewpoints in the contradiction markers are extracted and transformed into questions (e.g., why do you support this viewpoint?). Based on the partition discussion activity level, the strategy content is adjusted (adding detailed follow-up questions in high-activity areas and using open-ended questions in low-activity areas), generating follow-up questioning strategies for different partitions. For example, for high-activity partitions, follow-up questioning strategies are generated to deepen the discussion (e.g., asking specific questions based on the viewpoints); for low-activity partitions, follow-up questioning strategies are generated to activate participation (e.g., simplifying questions or guiding the speech).
[0136] This solution analyzes speech content, identifies viewpoints, and extracts core intentions to provide input for matching analysis, addressing the problem of not being able to deduce key viewpoints from speech content. It analyzes speech status to determine the matching degree between the status and viewpoints, achieving multi-dimensional data fusion and resolving the inability to assess the consistency between speech content and facial expressions, accurately identifying potential contradictions. If the matching degree is lower than a preset value, a contradiction marker is generated, and the activity level of discussion zones is determined based on a speech distribution heatmap, providing visual data support for discussion distribution and addressing the problem of missing viewpoints due to the inability to identify low-activity areas, laying the foundation for strategic zoned push notifications. Based on the contradiction markers, and according to the viewpoints and discussion activity levels of different zones, follow-up questioning strategies are generated for different zones, addressing the problem of guidance failure due to the lack of proactive interaction mechanisms, and achieving closed-loop feedback.
[0137] In some embodiments, a preset micro-expression-semantic mapping table is obtained, which stores the correspondence between words and positive micro-expressions; the micro-expression types in the speaking state are analyzed to determine the actual micro-expression polarity; based on the micro-expression-semantic mapping table, keywords in the speaking opinions are parsed to determine the expected micro-expression polarity; if the actual micro-expression polarity is inconsistent with the expected micro-expression polarity, the matching degree is determined to be lower than the preset matching value.
[0138] A micro-expression-semantic mapping table can be a mapping table that stores the correspondence between different words and positive micro-expression types.
[0139] Positive micro-expressions can be a type of micro-expression that is considered positive.
[0140] The correspondence can be a rule of association between words and positive micro-expression types.
[0141] Actual micro-expression polarity can be derived from the micro-expression type field in the speech state, representing the polarity label of the speaker's actual emotional state.
[0142] Keywords can be different words identified in a statement's viewpoint.
[0143] Expected microexpression polarity can be derived from keywords and represents the theoretical emotional response that a speaker should exhibit when expressing their views.
[0144] Specifically, a pre-defined micro-expression-semantic mapping table is obtained from a micro-expression-semantic data repository. This table stores the correspondence between different words (such as "think," "suggest," "agree," and "disagree") and their corresponding positive micro-expressions. For example, the word "agree" is mapped to a smiling positive micro-expression, and the word "disagree" is mapped to a serious positive micro-expression.
[0145] Micro-expression types (such as smiling, frowning, and confused) are extracted from the speech state data. Then, based on the pre-micro-expression polarity rule (used to classify the extracted micro-expression types into actual micro-expression polarities), the polarity (positive or negative) of the extracted micro-expression type is determined, thereby determining the actual micro-expression polarity. For example, when the micro-expression type is smiling, the actual micro-expression polarity is determined to be positive; when the micro-expression type is frowning or confused, the actual micro-expression polarity is determined to be negative.
[0146] The text of the expressed opinions is scanned to identify keywords containing opinions (such as support, opposition, and agreement). Based on the identified keywords, a micro-expression-semantic mapping table is consulted to obtain the expected micro-expression polarity corresponding to the keywords. For example, when the keyword is support, the micro-expression-semantic mapping table indicates that it is associated with a positive micro-expression, so the expected micro-expression polarity is determined to be positive; when the keyword is opposition, the micro-expression-semantic mapping table indicates that it is associated with a positive micro-expression, so the expected micro-expression polarity is also determined to be positive.
[0147] The actual micro-expression polarity is compared with the expected micro-expression polarity to determine whether the matching degree is lower than the preset matching value directly set based on the actual micro-expression polarity and the expected micro-expression polarity (used to compare the matching degree to determine the consistency between the two). For example, when the two are consistent (such as the actual micro-expression polarity being positive and the expected micro-expression polarity being positive), the matching degree is set to high matching degree and the matching degree is determined to be higher than the preset matching value; when the two are inconsistent (such as the actual micro-expression polarity being negative but the expected micro-expression polarity being positive), the matching degree is set to low matching degree and the matching degree is determined to be lower than the preset matching value.
[0148] This solution obtains a pre-defined micro-expression-semantic mapping table, storing the correspondence between words and positive micro-expressions. This avoids real-time calculations or reliance on external data, standardizing and repeating the process of determining the expected polarity. It analyzes the types of micro-expressions in the speech state to determine the actual micro-expression polarity, ensuring processing speed and reliability, and capturing the speaker's nonverbal cues as a key basis for matching degree calculation. Based on the micro-expression-semantic mapping table, it parses keywords in the speech's viewpoint to determine the expected micro-expression polarity, standardizing the emotional interpretation of the speech's viewpoint, avoiding semantic ambiguity, and providing a consistent benchmark for comparison with the actual polarity, ensuring efficient execution and providing expected reference values for matching degree determination. If the actual micro-expression polarity is inconsistent with the expected micro-expression polarity, the matching degree is determined to be lower than the pre-defined matching value, achieving a quantitative assessment of emotional consistency, solving the problem of being unable to assess the consistency between speech content and facial expressions, and ensuring real-time response.
[0149] In some embodiments, entity words and action words are extracted from fragmented viewpoints; entity words are parsed to determine the correlation between at least two entity words; action words and entity words are parsed to determine the action relationship between any action word and any entity word; a semantic graph is constructed based on the correlation, where nodes represent entity words and edges represent action relationships; entities with the same parent node are merged based on the semantic graph to generate candidate viewpoint clusters; the frequency of occurrence of each candidate viewpoint cluster is counted, and several structured advancement options are obtained based on the statistical results.
[0150] Entity words can be nouns that represent specific things or concepts, extracted from fragmented opinion texts.
[0151] Action words can be verbal words that express actions or operations, extracted from fragmented opinion texts.
[0152] Relevance can be a relevance score between entity words.
[0153] Action relationships can be a combination of action words and entity words in a binary tuple.
[0154] Semantic graphs can be network structures with entity words as nodes and action relationships as edges.
[0155] Entities can be nodes in a semantic graph.
[0156] Candidate viewpoint clusters can be topic groups formed by merging entity words that share the same parent node in a semantic graph.
[0157] Frequency of occurrence can be the total number of times a candidate opinion cluster is mentioned in fragmented opinion texts.
[0158] The statistical results can be a list sorted by the frequency of occurrence of candidate opinion clusters.
[0159] Specifically, part-of-speech tagging is used to identify noun words (such as product and cost) in fragmented viewpoints as entity words; at the same time, part-of-speech tagging is used to identify verb words (such as improve and reduce) in fragmented viewpoints as action words; for example, if the input fragmented viewpoint is to optimize user experience, the extracted entity word is user experience and the action word is optimization.
[0160] The semantic relationships (logical connections) between entity words are analyzed. By calculating the co-occurrence frequency of entity words in fragmented viewpoints, the degree of association between at least two entity words is determined. For example, if two entity words (such as product and market) appear simultaneously in multiple fragmented viewpoints, the degree of association is high; otherwise, it is low.
[0161] Analyze the semantic collocation of action words and entity words (the grammatical and semantic combination relationship between action words and entity words). By matching the combination relationship between action words and entity words in fragmented viewpoints, determine the action relationship between any action word and any entity word. For example, when the action word "support" is combined with the entity word "solution", the action relationship is defined as support-solution.
[0162] A semantic graph is constructed based on relevance, where entity words are created as nodes in the semantic graph; action relationships are used as edges to connect related entity word nodes. All nodes in the semantic graph are traversed, and the parent node of each node is identified (the parent node that directly connects multiple child nodes, such as a product node connecting a function node and an interface node). Then, for each parent node, all its direct child nodes (i.e., entity words) are merged into a group (e.g., the child nodes (functions and interfaces) of the parent node (product) are merged into a cluster (product-function-interface)). Finally, a candidate viewpoint cluster is generated for each merged group (based on the combination of parent and child nodes).
[0163] Traverse the fragmented viewpoints and count the frequency of occurrence of each candidate viewpoint cluster (e.g., the cluster (product-function-interface) appears in 10 fragmented viewpoints); then, sort the candidate viewpoint clusters from high to low frequency; subsequently, select the top N high-frequency clusters (N is a natural number) and directly map them to several structured advancement options.
[0164] This solution extracts entity words and action words from fragmented viewpoints, focusing on core vocabulary and ignoring irrelevant text to improve processing efficiency and accuracy, avoiding clustering bias caused by text noise. Entity words are analyzed to determine the correlation between at least two entity words, ensuring the semantic graph accurately reflects the tightness of entity word density, thus reducing interference from irrelevant entities and improving clustering accuracy. Action words and entity words are analyzed to determine the action relationship between any action word and any entity word, ensuring the semantic graph represents the dynamic semantic structure of viewpoints and supports entity merging. A semantic graph is constructed based on correlation, where nodes represent entity words and edges represent action relationships, enabling the semantic structure and interaction relationships in fragmented viewpoints to be represented digitally, providing a unified framework for entity merging and simplifying clustering operations. Based on the semantic graph, entities with the same parent node are merged to generate candidate viewpoint clusters, achieving preliminary clustering of viewpoints and integrating scattered entity words into coherent groups, thereby reducing redundancy. The frequency of occurrence of each candidate viewpoint cluster is statistically analyzed, and based on the statistical results, several structured advancement options are obtained, converging fragmented viewpoints into actionable options, thus supporting meeting-guided decision-making.
[0165] In some embodiments, the semantic conflict degree of any two structured advancement options is analyzed; if there is a semantically conflicting option pair, a conflict warning mark is generated; based on the speech distribution heatmap, the partition where the supporter of the conflict warning mark is located is located; based on fragmented viewpoints, a debate guidance strategy is generated for the partition where the supporter is located.
[0166] Semantic conflict can be a quantification of the degree of semantic opposition between two structured advancement options.
[0167] Semantic conflict can be a qualitative state in which there is a semantic opposition between two structured advancement options.
[0168] An option pair can be any combination of two structured advancement options.
[0169] Conflict warning tags can be tags generated when semantic conflicts exist.
[0170] The supporter's section can be the sub-section of speakers who support the conflict option.
[0171] Debate guidance strategies can be based on generating text instructions from fragmented viewpoints for proactive robot interaction.
[0172] Specifically, it iterates through all combinations of structured advancement options; for each combination, it extracts relevant viewpoints from fragmented viewpoints (e.g., if option A is to optimize cost, then it extracts fragmented viewpoints containing both cost and optimization); then it checks whether there are expressions in the fragmented viewpoints that support one option but oppose the other (e.g., if the viewpoint supports optimizing cost but opposes increasing the budget, and option B is to increase the budget, then it detects opposing action words); then it counts the total number of fragmented viewpoints related to the current option pair, and then counts the number of viewpoints containing semantically contradictory action words (e.g., opposing, contradictory, conflicting action words); based on the total number of fragmented viewpoints and the number of semantically contradictory viewpoints, it determines the semantic conflict degree between any two structured advancement options.
[0173] If the semantic conflict between any two structured advancement options exceeds a conflict threshold set based on historical meeting data (used to control the generation conditions of conflict warning tags), the option pair is determined to have a semantic conflict, and a conflict warning tag is generated. The index of relevant fragmented viewpoints is extracted from the conflict warning tags to locate supporters (speakers expressing semantic opposition in relevant fragmented viewpoints), and the location coordinates of the supporters are obtained. Then, a speech distribution heatmap is used to map the supporters' locations to partitions, and the partitions where all supporters are located are summarized.
[0174] For each supporter's section, the fragmented viewpoints expressed by speakers in that section are filtered out, and semantic content related to conflict warning markers is focused. Based on the relevant semantic content, entity words and action words are combined to form a debate guidance strategy for the supporter's section. For example, if the viewpoint includes opposition to increasing the budget, a strategy is generated to initiate a debate: optimize costs or increase the budget, please explain the reasons for opposition in the section.
[0175] This solution analyzes the semantic conflict between any two structured advancement options, addressing the issue of missed contradictions due to the lack of proactive interaction mechanisms, and providing data for generating conflict warning markers. If a pair of options exhibits semantic conflict, a conflict warning marker is generated, fulfilling the prerequisite for generating follow-up questioning strategies and transforming abstract conflict into actionable machine instruction trigger signals. Based on a heatmap of speech distribution, the solution locates the partitions where supporters of conflict warning markers reside, fulfilling the requirement of heatmap-based support area location. Based on fragmented viewpoints, a debate guidance strategy is generated for the partitions where supporters reside, addressing the shortcomings of passive recording without guidance.
[0176] In some embodiments, when a contradiction marker is detected, the opinions expressed are analyzed to determine the point of contradiction; the low-activity zone to which the contradiction marker belongs is located based on the activity level of the discussion in the zone; the point of contradiction is pushed to the high-activity zone and a follow-up question is generated in the high-activity zone; feedback audio data from the high-activity zone is received, the feedback audio data is parsed, and a feedback summary is generated; the feedback summary is pushed back to the original low-activity zone.
[0177] The points of contention can be identified by analyzing the content of the statements made.
[0178] The low-activity zone can be a meeting room grid zone where the discussion activity value is lower than a preset activity threshold.
[0179] The high-activity zone can be a meeting room grid zone where the discussion activity value is higher than a preset activity threshold.
[0180] Follow-up commands in high-activity zones can be text commands pushed to high-activity zones.
[0181] Feedback audio data can be the raw audio of new speech captured by distributed microphones after a high-activity zone responds to a follow-up question.
[0182] Feedback summaries can be structured text summaries generated after parsing feedback audio data.
[0183] The original low-activity area can be a low-activity area used to receive feedback digests.
[0184] Specifically, when a contradiction marker is detected, the content of the statement is analyzed to identify key controversial entities and action words (such as opposing themes like cost optimization and budget increase), thereby determining the point of contradiction (such as the contradictory theme of cost optimization or budget increase).
[0185] The activity level of the discussion zones is compared with a preset activity threshold set based on historical meeting data (used to identify low-activity and high-activity zones, pre-stored on the server and retrieved when needed). When the activity level of a discussion zone is lower than the preset activity threshold, the zone to which the conflict marker belongs is identified as a low-activity zone. The meeting room grid zones with discussion activity values higher than the preset activity threshold are identified from the speech distribution heatmap, and the zones to which the conflict marker belongs are identified as high-activity zones. Then, the conflicting issues are sent to the high-activity zones via the integrated display screen device (e.g., displaying the conflicting issue as text: the opposition between cost optimization and budget increase). Based on the conflicting issues, entity words and action words are combined to form follow-up instructions for the high-activity zones, such as: "Please discuss in the high-activity zones: the opposition between cost optimization and budget increase."
[0186] Feedback audio data from high-activity areas is captured using distributed microphones covering a pre-defined grid of the meeting room; speech recognition technology is used to convert the feedback audio data into text; then the text content is analyzed to extract key viewpoints (highly frequent controversial keywords, such as cost optimization and budget increase), and a feedback summary is generated (e.g., the summary indicates that the high-activity area supports cost optimization and opposes budget increase).
[0187] Feedback summaries are pushed to the original low-activity areas via partitioned display screens (e.g., supporting cost optimization and opposing budget increases for high-activity areas).
[0188] This solution analyzes the viewpoints when a contradiction marker is detected, identifies the points of contention, and ensures that the follow-up questioning strategy is based on a clear and actionable topic of contention. This prevents the strategy from becoming ineffective due to ambiguity or generalization, directly supporting the accuracy of the follow-up questioning strategy. Based on the activity level of the discussion in each area, the solution locates the low-activity zone to which the contradiction marker belongs, identifying areas with insufficient participation. This ensures that the strategy is designed specifically for this low-activity zone, avoiding wasted resources on non-target areas, thus establishing a zone-level positioning foundation for the follow-up questioning strategy. The points of contention are then pushed to the high-activity zone, generating follow-up questioning instructions for that zone. This activates the discussion focus in the high-activity zone, guiding participants to provide targeted feedback and laying the data foundation for feedback collection, thereby enabling the follow-up questioning strategy to be executed in the high-activity zone. The solution receives audio feedback data from the high-activity zone, parses the audio data, generates feedback summaries, extracts the essence of the feedback, and transforms it into concise and easily disseminated summaries, avoiding information overload and simplifying the information in the follow-up questioning strategy. The feedback summaries are then pushed back to the original low-activity zone, stimulating participants in the low-activity zone to reflect on or respond to the viewpoints in the high-activity zone, promoting the activation of the discussion in that zone, thus completing the closed-loop guidance of the follow-up questioning strategy.
[0189] Figure 3 A schematic diagram of the structure of an AI-based robot control system for conference research, as provided in one embodiment of this application, is shown below. Figure 3 As shown, the AI-based robot control system 300 for conference research in this embodiment includes: a data acquisition module 301, a data analysis module 302, a heat map generation module 303, and a strategy generation module 304.
[0190] The data acquisition module 301 is used to acquire audio data and monitoring data of the conference environment;
[0191] Data analysis module 302 is used to analyze the monitoring data and the audio data to determine the speaking status and speaking content;
[0192] The heatmap generation module 303 is used to determine a heatmap of speech distribution based on the speech status and speech content;
[0193] The strategy generation module 304 is used to determine the robot's proactive interaction strategy based on the speech content and the speech distribution heatmap, and send it to the robot so that the robot executes the proactive interaction strategy to guide the meeting discussion.
[0194] Optionally, when the data analysis module 302 analyzes the monitoring data and the audio data to determine the speaking status, it is used to: analyze the monitoring data to determine the facial muscle movement characteristics and the speaker's position coordinates; determine the micro-expression type based on the movement characteristics; parse the audio data to determine the voiceprint characteristics and speaking content; match the voiceprint characteristics with a meeting record database to determine the speaker's identity; and determine the micro-expression type, the position coordinates, and the speaker's identity as the speaker's speaking status.
[0195] Optionally, when the heatmap generation module 303 determines the heatmap distribution based on the speaking status and speaking content, it is used to: obtain a set of position coordinates of all speakers based on the position coordinates; map the set of position coordinates to grid partitions according to a preset conference room planar grid; calculate the total speaking time and the number of speaking sessions per unit time in each grid partition according to the speaking content; and generate a partition activity heatmap by weighting the number of speaking sessions and the total speaking time.
[0196] Optionally, when the strategy generation module 304 determines the robot's active interaction strategy based on the speech content and the speech distribution heatmap, it is used to: parse the speech content and determine topic keywords; based on the speech distribution heatmap, determine the discussion duration and number of speakers for the topic keywords; if the number of speakers per unit time is lower than a preset value and the discussion duration exceeds a threshold, mark it as an open topic; when the duration of the open topic exceeds the limit, analyze the speech content and determine fragmented viewpoints; cluster the fragmented viewpoints and generate several structured advancement options.
[0197] Optionally, the AI-based meeting survey robot control system further includes a follow-up questioning strategy generation module 305, used for: analyzing the speech content and determining the speech viewpoint; analyzing the speech state and determining the matching degree between the speech state and the speech viewpoint; if the matching degree is lower than a preset matching value, generating a contradiction marker and determining the discussion activity level of the partition based on the speech distribution heatmap; and generating follow-up questioning strategies for different partitions based on the contradiction marker, the speech viewpoint, and the discussion activity level of the partition.
[0198] Optionally, when the questioning strategy generation module 305 analyzes the speaking state and determines the matching degree between the speaking state and the speaking viewpoint, it is used to: obtain a preset micro-expression-semantic mapping table, which stores the correspondence between words and positive micro-expressions; analyze the micro-expression types in the speaking state and determine the actual micro-expression polarity; based on the micro-expression-semantic mapping table, parse the keywords in the speaking viewpoint and determine the expected micro-expression polarity; if the actual micro-expression polarity is inconsistent with the expected micro-expression polarity, then the matching degree is determined to be lower than the preset matching value.
[0199] Optionally, when the strategy generation module 304 clusters the fragmented viewpoints and generates several structured advancement options, it is used to: extract entity words and action words from the fragmented viewpoints; parse the entity words to determine the correlation between at least two entity words; parse the action words and the entity words to determine the action relationship between any action word and any entity word; construct a semantic graph based on the correlation, where nodes represent entity words and edges represent action relationships; merge entities with the same parent node based on the semantic graph to generate candidate viewpoint clusters; count the frequency of occurrence of each candidate viewpoint cluster, and obtain several structured advancement options based on the statistical results.
[0200] Optionally, the AI-based conference survey robot control system further includes a guidance strategy generation module 306, used for: analyzing the semantic conflict degree of any two structured advancement options; generating conflict warning markers if there are semantically conflicting option pairs; locating the partition where the supporters of the conflict warning markers are located based on the speech distribution heatmap; and generating a debate guidance strategy for the partition where the supporters are located based on the fragmented viewpoints.
[0201] Optionally, when the follow-up questioning strategy generation module 305 generates follow-up questioning strategies for different partitions based on the contradiction marker, the speaking opinions, and the partition discussion activity, it is used to: when a contradiction marker is detected, analyze the speaking opinions to determine the contradiction point; locate the low-activity area to which the contradiction marker belongs based on the partition discussion activity; push the contradiction point to the high-activity area and generate a follow-up questioning instruction for the high-activity area; receive feedback audio data from the high-activity area, parse the feedback audio data, and generate a feedback summary; and push the feedback summary to the original low-activity area.
[0202] The system in this embodiment can be used to execute the methods of any of the above embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
Claims
1. A robot control method for conference research based on artificial intelligence, characterized in that, include: Acquire audio and monitoring data of the meeting environment; Analyze the monitoring data and the audio data to determine the speaking status and speaking content; Based on the speaking status and speaking content, a speaking distribution heatmap is determined; Based on the content of the speeches and the heatmap of the speech distribution, a proactive interaction strategy for the robot is determined and sent to the robot so that the robot can execute the proactive interaction strategy to guide the meeting discussion; The step of determining the robot's proactive interaction strategy based on the speech content and the speech distribution heatmap further includes: Analyze the content of the speech to determine the viewpoint; Analyze the speaking status to determine the degree of matching between the speaking status and the speaking viewpoint; If the matching degree is lower than the preset matching value, a contradiction marker is generated, and the activity level of the discussion in the partition is determined according to the heat map of the speech distribution. Based on the aforementioned contradiction markers, and according to the stated viewpoints and the activity level of the discussion in each partition, different follow-up questioning strategies are generated for each partition. The analysis of the speaking state and the determination of the matching degree between the speaking state and the speaking viewpoint include: Obtain a preset micro-expression-semantic mapping table, which stores the correspondence between words and positive micro-expressions; Analyze the types of micro-expressions in the speaking state to determine the actual polarity of the micro-expressions; Based on the micro-expression-semantic mapping table, keywords in the expressed opinions are analyzed to determine the expected micro-expression polarity; If the actual micro-expression polarity is inconsistent with the expected micro-expression polarity, then the matching degree is determined to be lower than the preset matching value; Based on the contradiction markers, and according to the expressed viewpoints and the activity level of the discussion in each partition, the following follow-up questioning strategies are generated for different partitions: When a contradiction marker is detected, the stated viewpoints are analyzed to determine the points of contention. Based on the activity level of the discussion in each partition, locate the low-activity zone to which the contradiction marker belongs; The problematic points are directed to a high-activity area, and follow-up questions are generated in the high-activity area. Receive feedback audio data from high-activity areas, parse the feedback audio data, and generate a feedback summary; The feedback summary is then pushed to the original low-activity area.
2. The method according to claim 1, characterized in that, The analysis of the monitoring data and the audio data to determine the speaking status includes: Analyze the monitoring data to determine the facial muscle movement characteristics and the speaker's position coordinates; Based on the aforementioned motion characteristics, the type of micro-expression is determined; Analyze the audio data to determine the voiceprint characteristics and content of the speech; Based on the aforementioned voiceprint characteristics, the speaker's identity is determined by matching them against the meeting record database. The micro-expression type, the location coordinates, and the speaker's identity are determined as the speaker's speaking status.
3. The method according to claim 2, characterized in that, The step of determining the speech distribution heatmap based on the speech status and speech content includes: Based on the aforementioned location coordinates, a set of location coordinates for all speakers is obtained; Based on the preset conference room plan grid, the set of location coordinates is mapped to the grid partitions; Based on the content of the speech, the total speaking time and the number of speeches per unit time in each grid partition are calculated; By weighting the number of messages and the total duration of the messages, a heatmap of the activity level of the partition is generated.
4. The method according to claim 1, characterized in that, The step of determining the robot's proactive interaction strategy based on the speech content and the speech distribution heatmap includes: Analyze the content of the speech to identify topic keywords; Based on the heatmap of the speech distribution, the discussion duration and number of speakers for the topic keywords are determined; If the number of speakers per unit time is lower than the preset value and the discussion time exceeds the threshold, it is marked as an open topic; When an open-ended topic exceeds the time limit, the content of the speech is analyzed to identify fragmented viewpoints. Clustering the fragmented viewpoints generates several structured advancement options.
5. The method according to claim 4, characterized in that, The clustering of fragmented viewpoints generates several structured advancement options, including: Extract entity words and action words from the fragmented viewpoints; Analyze the entity words to determine the degree of association between at least two entity words; Analyze the action words and entity words to determine the action relationship between any action word and any entity word; A semantic graph is constructed based on the correlation degree, where nodes represent entity words and edges represent action relationships; Based on the semantic graph, entities with the same parent node are merged to generate candidate viewpoint clusters; The frequency of occurrence of each candidate viewpoint cluster is counted, and several structured advancement options are obtained based on the statistical results.
6. The method according to claim 5, characterized in that, After obtaining several structured advancement options based on statistical results, the following is also included: Analyze the semantic conflict degree between any two structured advancement options; If there are semantically conflicting option pairs, generate a conflict warning flag; Based on the aforementioned heatmap of speech distribution, locate the partition where the supporters of the conflict warning markers are located; Based on the fragmented viewpoints, generate debate guidance strategies tailored to the supporters' respective sections.
7. An AI-based robot control system for conference research, characterized in that: The method applied to any one of claims 1-6 includes: The data acquisition module is used to acquire audio data and monitoring data from the meeting environment; The data analysis module is used to analyze the monitoring data and the audio data to determine the speaking status and speaking content; A heatmap generation module is used to determine a heatmap of speech distribution based on the speech status and speech content; The strategy generation module is used to determine the robot's proactive interaction strategy based on the content of the speech and the heatmap of the speech distribution, and send it to the robot so that the robot can execute the proactive interaction strategy to guide the meeting discussion.