Video conference data processing method and system

By recognizing the verbal intentions and operational behaviors of participants in video conferences, predicting the focus of the meeting, and allocating network resources, this technology solves the problem of insufficient identification of the importance of video streams in existing technologies, and achieves high-definition transmission of key information and improved meeting efficiency.

CN121486524APending Publication Date: 2026-02-06ZHEJIANG TUXUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511546912.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing video conferencing systems cannot identify the actual importance of different video streams at the current stage of the meeting, resulting in the erroneous reduction of the clarity of key information and failing to ensure the high-definition transmission of key information in scenarios where the importance of content is uneven.

Method used

By recognizing participants' verbal intentions and user actions, the system can predict the focus of the meeting, identify the core video stream, allocate high-priority network transmission resources to it, and degrade the quality of non-core video streams.

Benefits of technology

Ensuring high-definition and smooth transmission of critical information under limited network resources, while appropriately downgrading relatively unimportant auxiliary information, improves meeting efficiency and user experience, and achieves automated identification and processing of meeting focus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121486524A_ABST
    Figure CN121486524A_ABST
Patent Text Reader

Abstract

The invention discloses a video conference data processing method and system, relates to the field of video conference data processing, is used for guaranteeing the quality presentation of a current most critical video stream, and comprises the following steps: identifying a verbal intention instruction of a conference participant; the participants are participants of the video conference; monitoring user operation behaviors of the participants; pre-judging a conference focus based on the oral intention instruction and the user operation behavior to obtain a core video stream; and allocating high-priority network transmission resources to the core video stream, and performing quality degradation processing on non-core video streams except the core video stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video conferencing data processing, and more particularly to a video conferencing data processing method and system. Background Technology

[0002] In modern remote collaboration environments, video conferencing systems have become indispensable tools for team communication and project advancement. These systems typically ensure remote communication by capturing, encoding, transmitting, and decoding audio and video data. To cope with complex network environments, existing systems generally employ an adaptive data transmission strategy based on network conditions. This involves real-time monitoring of the user's network connection quality (such as available bandwidth, data transmission latency, and packet loss) and automatically reducing the video stream's resolution or frame rate when network conditions are poor, prioritizing the smoothness of the meeting. This general strategy can maintain a basic call experience in most cases.

[0003] However, this general approach reveals significant limitations in certain work scenarios. For example, in video conferences containing multiple types of video streams (such as high-resolution static drawings, changing screen presentation videos, and participant camera feeds), existing data processing methods primarily rely on network connection quality for uniform, indiscriminate video quality adjustment. This approach fails to recognize the actual importance of different video streams at the current stage of the meeting. Consequently, in scenarios where content importance is uneven (such as chip layouts presented in a design review meeting), the clarity of critical information (such as design drawings) is incorrectly reduced, while relatively unimportant auxiliary information (such as participants' facial expressions) may remain clear. Summary of the Invention

[0004] This invention provides a video conferencing data processing method to ensure the quality presentation of the most critical video streams.

[0005] In a first aspect, to solve the above-mentioned technical problems, the present invention provides a video conferencing data processing method, comprising: identifying the verbal intent instructions of participants; identifying the participants as participants in the video conference; monitoring the user operation behavior of the participants; predicting the focus of the meeting based on the verbal intent instructions and user operation behavior to obtain the core video stream; allocating high-priority network transmission resources to the core video stream, and performing quality degradation processing on non-core video streams other than the core video stream.

[0006] Optionally, based on verbal intent instructions and user operation behavior, the focus of the meeting is predicted to obtain the core video stream, including: when a verbal intent instruction is recognized that the participant intends to display important content, the client application presets an intent evidence score for each monitored preparatory behavior and accumulates the intent evidence score; when the participant's accumulated intent evidence score is greater than or equal to the preset prediction locking threshold, the client application sends a quasi-focus locking request signaling to the meeting server; after receiving the quasi-focus locking request signaling, the meeting server marks the participant as the unique quasi-focus and determines the video stream to be emitted by the unique quasi-focus as the core video stream.

[0007] Optionally, high-priority network transmission resources are allocated to the core video stream, and quality degradation processing is performed on non-core video streams other than the core video stream. This includes: entering a locked state based on the unique quasi-focus; the locked state is a state that ignores conflict signals from other participants outside the unique quasi-focus; the conference server pre-allocates network transmission resources for the core video stream to be emitted by the unique quasi-focus and sends a degradation instruction to the client applications of participants who are not the unique quasi-focus; the degradation instruction is used to require the video encoders of the client applications of participants who are not the unique quasi-focus to adjust their encoding parameters and perform quality degradation processing; when the unique quasi-focus clicks the share screen button on the conference application interface, the conference server activates the pre-allocated network channel and encoding parameters.

[0008] Optionally, the client application presets an intent evidence score for each monitored pre-preparation behavior and accumulates the intent evidence score, including: when a participant's pre-preparation behavior is monitored, assigning an initial intent evidence score to each pre-preparation behavior and starting a timer associated with the pre-preparation behavior; the initial intent evidence score decays over time during the timer's operation; when a participant performs a new pre-preparation behavior, accumulating the initial intent evidence score of the new pre-preparation behavior and starting a new timer for the new pre-preparation behavior.

[0009] Optionally, an initial intent evidence score may be assigned to each preparatory behavior, including: identifying the meeting type or meeting stage; and dynamically adjusting the initial intent evidence score of the preparatory behavior based on the identified meeting type or meeting stage.

[0010] Optionally, the client application presets an intent evidence score for each monitored pre-preparation behavior and accumulates the intent evidence score, including: when a participant's pre-preparation behavior is monitored, assigning an initial intent evidence score to each pre-preparation behavior and starting a timer associated with that pre-preparation behavior; the initial intent evidence score decays over time during the timer's operation; identifying sequence patterns or combinations of participant pre-preparation behaviors; adjusting the initial intent evidence score with weights based on the identified sequence patterns or combinations; and when a participant performs a new pre-preparation behavior, accumulating the weighted initial intent evidence score of the new pre-preparation behavior and starting a new timer for the new pre-preparation behavior.

[0011] Optionally, identifying the sequence or combination patterns of participants' pre-preparation behaviors includes: receiving pre-preparation behavior event streams from multiple participants; extracting behavior patterns from each participant's pre-preparation behavior event stream; and resolving conflicts between similar behavior patterns when multiple participants exhibit similar behavior patterns simultaneously to obtain a unique shared intent pattern.

[0012] Optionally, conflict resolution is performed on similar behavioral patterns to obtain a unique shared intent pattern, including: identifying the activity level of each similar behavioral pattern, where activity level is determined based on the frequency of the behavior within the time window in which the behavior occurs and the recentity of the behavior; identifying the context relevance of each similar behavioral pattern, where context relevance is determined based on the degree of matching between the behavioral pattern and the current meeting type or meeting stage; performing a weighted evaluation on each similar behavioral pattern based on activity level and context relevance to obtain a pattern priority score; and selecting the behavioral pattern with the highest pattern priority score as the unique shared intent pattern.

[0013] Optionally, the activity level of each similar behavioral pattern is identified, including: when monitoring the pre-event preparation behavior of participants, generating an event record for each behavioral event containing the behavior type, timestamp, and client application identifier; the client application synchronizing and calibrating its local system clock before sending the event record; after receiving event records from multiple client applications, the conference server uniformly calibrating the timestamp of each event record; based on the calibrated timestamps, the conference server statistically analyzes the frequency of occurrence of behavioral events within a preset time window and calculates the recentity of the behavioral events; the recentity is used to characterize how close the behavioral event is to the current moment; based on the frequency of occurrence and the recentity, combined with a preset decay function, the conference server calculates the activity contribution of each behavioral event and accumulates them to obtain the activity level.

[0014] In a second aspect, the present invention provides a video conferencing data processing system, the system comprising: The instruction recognition module is used to recognize the verbal intentions of participants; the participants are those involved in the video conference. The behavior monitoring module is used to monitor the user operation behavior of participants; The focus prediction module is used to predict the focus of the meeting based on verbal intentions and user actions, and obtain the core video stream. The resource allocation module is used to allocate high-priority network transmission resources to the core video stream and to perform quality degradation processing on non-core video streams other than the core video stream.

[0015] Compared with the prior art, the present invention has the following beneficial effects: The video conferencing data processing method disclosed in this application can intelligently predict the focus of the meeting by recognizing the verbal intentions of participants and monitoring user operation behavior, and obtain the core video stream accordingly. Based on this, high-priority network transmission resources are allocated to the core video stream, while non-core video streams undergo quality degradation processing. This method effectively solves the problem in existing video conferencing systems that cannot identify the actual importance of different video streams at the current stage of the meeting, leading to an erroneous reduction in the clarity of key information. Through intelligent prediction and differentiated transmission, this application can ensure that, under limited network resources, key information of the meeting (such as presentations, design drawings, etc.) can be transmitted with high definition and high smoothness, while relatively unimportant auxiliary information is appropriately downgraded, thereby significantly improving the efficiency of the meeting and the user experience. Furthermore, this method achieves automated identification and processing of the meeting focus, avoiding the tediousness and inefficiency of traditional manual intervention, allowing the meeting host to focus more on the meeting content itself and improving the natural flow of the meeting. In summary, this application provides an intelligent and efficient video conferencing data processing solution that overcomes the shortcomings of existing technologies in understanding meeting content and interaction status, and realizes the transformation from a passive network pipeline optimizer to an intelligent meeting collaboration assistant. Attached Figure Description

[0016] Figure 1 This is a schematic flowchart of a video conferencing data processing method provided in an embodiment of the present invention; Figure 2 This is a schematic flowchart of another video conferencing data processing method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a video conferencing data processing system provided in an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0018] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0019] The following specific embodiments will provide a detailed description and explanation of a video conferencing data processing method provided in this application.

[0020] The implementation environment of this application typically includes one or more client applications (running on the participants' terminal devices, such as computers, mobile phones, etc.) and a conference server. The client applications are responsible for collecting the participants' verbal intentions and user actions, and performing preliminary processing; the conference server is responsible for receiving signals from the clients, predicting the focus of the meeting, and coordinating the allocation of network resources and the quality management of the video stream.

[0021] Reference Figure 1 This invention provides a video conferencing data processing method, comprising the following steps: S1 identifies the verbal intentions of the participants.

[0022] The participants are those who are involved in the video conference.

[0023] Verbal intention instructions refer to instructions expressed verbally by participants during a meeting, indicating their intentions. Examples include "I want to share my screen" or "Please look at this document."

[0024] As one possible implementation, the audio processing unit in the video conferencing system continuously performs low-power scanning of the voice input from all participants, receiving audio data streams from all microphones in the conference in real time. Employing acoustic model-based keyword detection technology, the recognition process identifies whether preset keyword groups representing intended commands exist in the speech.

[0025] For example, by pre-training a small acoustic model to identify the speech features of these specific phrases, the keyword is considered to be identified when the acoustic features of the input audio match the features of a keyword in the model more than a preset threshold (e.g., 0.85).

[0026] This allows for the rapid capture of verbal announcements from participants about to share their screens or switch presentation content with minimal computational overhead, which is particularly common in engineering design review meetings where frequent handovers of presentation rights are required.

[0027] S2. Monitor the user behavior of attendees.

[0028] User actions refer to the various actions performed by participants on the video conferencing client application. Examples include clicking the screen sharing button, opening a document, switching presentation pages, mouse movement, and keyboard input. These actions can reflect the participants' underlying intentions.

[0029] As one possible implementation, upon recognizing a participant's verbal intent, the system could immediately mark the participant's internal state as "pending switch state" and monitor specific user interface operation events on the participant's terminal to monitor the participant's user behavior.

[0030] In this context, specific user interface operation events are generated by the participant's client application (e.g., desktop video conferencing software) and sent to the conference server when the user performs a specific operation.

[0031] For example, when participant A clicks the "Share Screen" button on the interface, the client application immediately sends a signaling packet containing the user ID and action type (e.g., {"user_id": "A", "action": "start_screen_share_button_clicked"}). The system sets a short time window to wait for these action events to occur. This time window can be a configurable parameter, such as 2 to 5 seconds after successful recognition of the intent instruction. This time range is set based on the average human reaction time to perform an action after a verbal announcement, capturing the true intent while avoiding long waiting times.

[0032] It should be noted that when the system successfully identifies a participant (e.g., participant A) as having issued the aforementioned "intent command" keyword group, the system will immediately mark participant A's internal state as "pending switch state." This mark can be represented in the conference server's memory as a simple Boolean variable (e.g., is_pending_share[participant_A_ID] = true) or a state machine variable. At this point, the system will not immediately adjust network resources but will proceed to the next verification phase. This delayed processing is to avoid misjudgments caused by slips of the tongue or informal expressions, ensuring the accuracy of subsequent resource allocation.

[0033] S3. Based on verbal instructions and user actions, predict the focus of the meeting and obtain the core video stream.

[0034] The focus of the meeting refers to the core content or speaker that is of general interest or about to be of interest to the participants at a particular stage of the meeting.

[0035] The core video stream can refer to the video stream that is anticipated to be the focus of the meeting, such as the speaker's video stream, the shared presentation video stream, or the video stream of a specific document.

[0036] Conversely, non-core video streams can refer to video streams other than the core video stream, such as camera feeds of other participants, supplementary documents that are not part of the current discussion, etc.

[0037] As one possible implementation, the system can pre-determine an "intent evidence score" for each verbal instruction and user action. These scores are accumulated when the corresponding instruction or action is recognized. For example, the verbal instruction "I want to share" might receive a high score, while opening presentation software might receive a medium score, and moving the mouse over the presentation might receive a low score. When the accumulated intent evidence score reaches a pre-determined threshold, the system predicts that the participant will become the focus of the meeting and identifies their upcoming video stream as the core video stream.

[0038] As another possible implementation, the system can employ machine learning models to learn the correlation patterns between verbal intentions and user actions and meeting focus by training on a large amount of meeting data, thereby predicting meeting focus and obtaining the core video stream. For example, the model can identify that in a specific meeting type, a certain sequence of actions (such as saying "Please look," then opening the document, and then clicking to share the screen) is more likely to indicate a shift in meeting focus.

[0039] In one example, if participant A says "I'd like to share" and clicks the "Share Screen" button within 5 seconds, the two-factor authentication is successful. At this point, the meeting server will immediately determine that a meeting focus shift is about to occur, and the new focus will be the screen content that participant A is about to share.

[0040] S4. Allocate high-priority network transmission resources to the core video stream and perform quality degradation processing on non-core video streams other than the core video stream.

[0041] High-priority network transmission resources refer to the priority bandwidth, lower transmission latency, or higher transmission guarantee level allocated to the core video stream during network transmission to ensure its transmission quality.

[0042] Quality degradation refers to adjusting the encoding parameters of non-core video streams, such as reducing resolution, frame rate, or compression quality, to reduce their network bandwidth usage and free up resources for core video streams.

[0043] As one possible implementation, once the core video stream is determined, the conferencing server sends instructions to network devices (such as routers and switches) to set a higher Quality of Service (QoS) priority for the core video stream's data packets, ensuring its priority transmission during network congestion. Simultaneously, the conferencing server sends degradation instructions to the client applications sending non-core video streams, requiring their video encoders to adjust their encoding parameters.

[0044] For example, the video resolution can be reduced from 1080p to 720p, or the frame rate can be reduced from 30fps to 15fps to reduce its bandwidth usage.

[0045] For example, when a presenter's presentation is identified as the core video stream, the system will ensure that the video stream of that presentation is transmitted at the highest quality, while the video streams from other participants' cameras may be appropriately downgraded.

[0046] Understandably, by intelligently recognizing and analyzing the verbal intentions and user actions of participants, the focus of the meeting can be automatically predicted, and the transmission priority and quality of the video stream can be dynamically adjusted accordingly, thereby effectively solving the problem of blind adjustment of video stream quality in existing technologies.

[0047] In one possible design, such as Figure 2 As shown, in order to predict the focus of the meeting and obtain the core video stream based on verbal instructions and user actions, this application may further include the following steps: S101. When a verbal intent instruction is detected indicating that the participant intends to present important content, the client application presets an intent evidence score for each monitored preparatory behavior and accumulates the intent evidence scores.

[0048] In the context of recognizing verbal intent instructions indicating that a participant intends to present important content, the system uses speech recognition technology to analyze the participant's speech and determine whether it contains explicit expressions such as "I want to share my screen," "I'm going to show you something," or "Please look at my presentation," indicating an intention to present important content. Important content can be understood as any form of information that a participant plans to present to other participants during the meeting, such as presentations, documents, images, videos, application interfaces, or screen-sharing content.

[0049] Preparatory actions refer to a series of actions that participants may take during a meeting, indicating that they are about to present important content. These actions may include, but are not limited to: opening presentation software (such as PowerPoint or Keynote), opening document editing software (such as Word or Google Docs), connecting external display devices, adjusting camera or microphone settings, typing prompts such as "Wait a moment, I'll share right away" in the chat box, or preparatory actions before clicking the "Share" or "Present" button in the meeting application.

[0050] Each preparatory action is assigned a preset intent evidence score, reflecting its relevance or importance to the intended presentation of key content. For example, opening presentation software might be assigned a higher score, while adjusting microphone settings might be assigned a lower score. The client application continuously monitors these actions and accumulates their corresponding intent evidence scores.

[0051] As one possible implementation, when monitoring participants' pre-preparation behaviors, the system can assign an initial intent evidence score to each pre-preparation behavior and start a timer associated with the pre-preparation behavior; when a participant performs a new pre-preparation behavior, the initial intent evidence score of the new pre-preparation behavior is accumulated, and a new timer is started for the new pre-preparation behavior.

[0052] The initial intent evidence score decays over time during the timer's operation.

[0053] In one specific implementation, suppose a participant is preparing to share content during a video conference.

[0054] First, at time T1, the participant opened a presentation (preparatory behavior A), the system assigned an initial intent evidence score of 50 to the participant, and started a timer.

[0055] During the timer's operation, the 50-point value begins to decay at a rate of 1 point per second.

[0056] Subsequently, at time T2 (T2 > T1), the participant opened a webpage link related to the presentation (preparatory behavior B). The system assigned an initial intent evidence score of 30 to the participant and started a new timer for behavior B, and the 30 score began to decay.

[0057] At this point, the cumulative intent evidence score will be (50 - (T2-T1) * 1) + 30.

[0058] If the participant clicks the "Share Screen" button on the meeting application interface at time T3 (T3 > T2) (preparatory behavior C), the system will immediately assign an initial intent evidence score of 100 and start a new timer.

[0059] At this point, the cumulative intent evidence score will be (50 - (T3-T1) * 1) + (30 - (T3-T2) * 1) + 100.

[0060] In this way, the more recent the action and the more relevant it is to the current intent, the greater its contribution to the cumulative intent evidence score. This allows the system to more accurately determine that an attendee is about to become the focus of the meeting and to make timely predictions and resource allocations for the core video stream. If an attendee opens the presentation at time T1 and does not perform any other related actions for a long period of time, the 50-point score will continue to decay until its contribution to the total score becomes negligible, thus avoiding the situation where a single early action leads to a continuous misjudgment of the focus.

[0061] By introducing a "time decay" mechanism for intent evidence scores, the problem of delayed or inaccurate intent judgment that may occur with traditional simple accumulation methods is effectively solved. Because the intent evidence score of each preparatory action gradually decays over time, the contribution of earlier actions with lower relevance to the current intent to the total score gradually decreases, thus ensuring that the accumulated intent evidence score more accurately reflects the participant's current, real-time intent. Simultaneously, when a participant performs a new preparatory action, a new initial intent evidence score is immediately assigned and a new timer is started. This ensures that the latest and most relevant actions are reflected in the accumulated score promptly and fully, enabling the meeting focus prediction system to respond quickly to changes in participant intent.

[0062] S102. When the cumulative intent evidence score of the participants is greater than or equal to the preset pre-judgment locking threshold, the client application sends a quasi-focus locking request signaling to the conference server.

[0063] Among them, the pre-set threshold is a pre-defined value used to determine whether the attendee's intention to present important content is strong and clear enough.

[0064] When the accumulated intent evidence score reaches or exceeds this threshold, it indicates that the system has preliminarily confirmed that the participant is about to become the focus of the meeting. At this point, the client application generates and sends a quasi-focus lock request signaling to the meeting server. This signaling is used to notify the server that the participant has been pre-judged by the client as a potential focus of the meeting.

[0065] S103. After receiving the quasi-focus lock request signaling, the conference server marks the participants as the unique quasi-focus and determines the video stream that the unique quasi-focus is about to send as the core video stream.

[0066] In this context, "unique focus" refers to a situation where, at a given moment, the conference server, based on client requests and its own judgment, identifies a particular participant as the most likely to become the focus of the meeting. Once marked as the unique focus, the video stream that participant is about to send, such as their webcam or screen-shared video stream, will be designated as the core video stream by the system.

[0067] In some preferred embodiments, a specific example is given below. Suppose that in an online product launch conference, attendee A plans to give a product demonstration. When attendee A says during the conference, "Next, I will demonstrate the core features of the new product," the system recognizes their verbal intent. Subsequently, the client application begins monitoring attendee A's preparatory behavior.

[0068] Specifically, if participant A performed the following actions in sequence: 1. The local "Product Presentation PPT" file was opened (preset intent evidence score: 30 points).

[0069] 2. In the meeting chat box, type "Please wait a moment, I am preparing to share my screen" (preset intent evidence score: 20 points).

[0070] 3. The “Share Screen” button on the meeting application interface was clicked, but no specific content to be shared has been selected (preset intent evidence score: 40 points).

[0071] Assume the preset prediction and locking threshold is 80 points.

[0072] When participant A completes the first step, their accumulated intent evidence score is 30 points. After completing the second step, the accumulated score increases to 50 points. When participant A completes the third step, their accumulated score reaches 90 points. At this point, 90 points exceeds the preset pre-determined focus lock threshold of 80 points. The client application immediately sends a quasi-focus lock request signaling to the conference server. Upon receiving this signaling, the conference server marks participant A as the unique quasi-focus and determines their upcoming screen-sharing video stream as the core video stream. In this way, before participant A actually begins their presentation, the system has already predicted and locked the conference focus, laying the foundation for subsequent optimized allocation of network resources.

[0073] Through the aforementioned technical solution, this application significantly improves the accuracy and timeliness of predicting meeting focus in video conferencing. By comprehensively analyzing participants' verbal intentions and preparatory behaviors, and introducing an accumulation of intent evidence scores and a threshold judgment mechanism, the system can identify participants who are about to become the focus of the meeting earlier and more accurately. This proactive prediction allows the core video stream to be determined before the actual content is presented, thus providing sufficient time and accurate basis for subsequent allocation of high-priority network transmission resources and quality degradation processing of non-core video streams. Therefore, it effectively avoids problems such as video stuttering and blurring caused by delayed focus identification, greatly optimizing the meeting experience for participants, especially in key information sharing and presentation sessions, ensuring the smooth and high-quality presentation of core content.

[0074] In one possible design, high-priority network transmission resources are allocated to the core video stream, and non-core video streams are degraded in quality. This application also includes the following steps: S201, Based on the unique quasi-focus, enter the locked state.

[0075] The locked state is a state in which conflict signals from other participants, except for the single quasi-focus, are ignored.

[0076] S202, The conference server pre-allocates network transmission resources for the core video stream that will be sent from the single quasi-focus, and sends a degradation instruction to the client applications of participants who are not the single quasi-focus.

[0077] The degradation instruction is used to require the video encoder of a participant's client application that is not a unique focal point to adjust the encoding parameters and perform quality degradation processing.

[0078] As one possible implementation, the conferencing server reserves the necessary network bandwidth, QoS (Quality of Service) priority, or other transmission channel resources for the single focal point before it actually begins transmitting high-bandwidth content (such as screen sharing or high-definition video). Simultaneously, the conferencing server sends degradation instructions to the non-single focal point participant client applications. These degradation instructions can be understood as control signals that require the video encoders of the non-single focal point participant client applications to adjust their encoding parameters.

[0079] The encoding parameters may include video resolution, frame rate, bit rate, etc.

[0080] This pre-allocation mechanism aims to reduce latency at the start of actual transmission, ensuring that core video streams receive the necessary transmission guarantees immediately. By adjusting encoding parameters, the quality of non-core video streams can be degraded, thereby freeing up network resources and prioritizing the transmission of core video streams.

[0081] S203. When the only quasi-focused user clicks the share screen button on the conference application interface, the conference server activates the pre-assigned network channel and encoding parameters.

[0082] In some preferred embodiments, a specific example is given below. Suppose that in an online product launch video conference, the presenter (participant A) is preparing to present detailed slides of a new product. Before presenting, presenter A may perform a series of preparatory actions, such as opening the slide file, adjusting the microphone volume, and clearing their throat. The system identifies presenter A's verbal intent instructions and user actions based on these actions and accumulates their intent evidence score. When presenter A's accumulated intent evidence score reaches a preset pre-judgment locking threshold, the client application sends a quasi-focus locking request signal to the conference server. Upon receiving this signal, the conference server marks presenter A as the unique quasi-focus.

[0083] At this point, the system enters a locked state, meaning the conference server ignores any conflict signals that might come from other participants (such as participant B or C), such as them also attempting to open files or adjust settings. The conference server immediately pre-allocates high-priority network transmission resources for the core video stream that will be sent by speaker A (such as their camera stream or the screen stream to be shared), reserving specific bandwidth or QoS channels. Simultaneously, the conference server sends degradation instructions to the client applications of all participants who are not the sole focus (such as participants B and C), requiring their video encoders to adjust encoding parameters, such as reducing their video resolution from 1080p to 720p or lower, to free up network resources.

[0084] When speaker A finally clicks the screen-sharing button on the meeting application interface, the meeting server immediately activates the network channel and encoding parameters previously pre-assigned to him. This means that speaker A's screen-sharing content will be transmitted to all participants with the highest quality and lowest latency, while the video streams of other participants remain in a degraded state, thus ensuring that speaker A's presentation is smooth and clear, while optimizing the overall network load.

[0085] Through the above technical solutions, this application can significantly improve the transmission quality and stability of the core video stream in video conferencing. By introducing a locking state, it effectively avoids frequent switching or conflicts of focus among multiple participants, ensuring the authority of a single quasi-focus. Pre-allocating network transmission resources and combining them with a precise activation mechanism ensures that the core video stream receives immediate, high-quality transmission at critical moments (such as screen sharing), greatly reducing latency and stuttering. Simultaneously, quality degradation processing of non-core video streams optimizes the overall network resource utilization efficiency, avoiding unnecessary bandwidth waste, thus providing a smoother meeting experience even under limited network conditions. Therefore, the solution of this application provides a more intelligent, efficient, and stable focus management and resource allocation mechanism for video conferencing.

[0086] In one possible design, in order to assign an initial intent evidence score to each preparatory action, this application further includes the following steps: S301. Identify the meeting type or meeting stage.

[0087] Among them, identifying the meeting type or meeting stage means that the system can automatically or manually obtain the attribute information of the current video conference, such as the topic of the meeting, the identities of the participants, the agenda of the meeting, and the duration of the meeting, so as to determine whether the meeting belongs to different types such as formal reports, brainstorming, daily stand-up meetings, and technical discussions, or is in different stages such as the opening, topic discussion, summary, and Q&A.

[0088] As one possible implementation, the system can use this information, such as meeting invitations, meeting titles, preset tags, or by analyzing the meeting agenda or verbal instructions using natural language processing (NLP) technology, to determine the meeting type or stage.

[0089] S302. Based on the identified meeting type or meeting stage, dynamically adjust the initial intent evidence score of the preparatory behavior.

[0090] The system dynamically adjusts the initial intent evidence score of pre-conference preparation behaviors based on the identified meeting type or stage. This can be understood as the system assigning different initial intent evidence scores to different pre-conference preparation behaviors (such as opening a presentation, sharing the screen, turning on the camera, adjusting the microphone, clearing the throat, etc.) according to the specific context of the current meeting.

[0091] For example, in a formal presentation meeting, opening a presentation or sharing a screen might be assigned a higher initial intent evidence score because this usually indicates that the speaker is about to present important content; while in an informal brainstorming session, these actions might receive a relatively lower initial score because their importance may not be as great as a formal presentation. The aim is to make the allocation of intent evidence scores more intelligent and contextual, thereby more accurately reflecting the true intent of the participants.

[0092] In some preferred embodiments, a specific example is given below. Assume a video conferencing system that, at the start of a meeting, first identifies the current meeting type as a "quarterly performance review meeting." In this type of meeting, the system pre-sets high initial intent evidence scores for preparatory actions such as "opening the presentation" and "sharing the screen," for example, 80 and 90 points respectively. Actions like "adjusting microphone volume" or "clearing the throat" may have lower initial scores, for example, 10 and 5 points respectively. When the meeting enters the "Q&A session," the system recognizes a change in meeting stage and may dynamically adjust the scores, for example, increasing the initial intent evidence score for "raising a hand" or "turning on the camera and speaking," while decreasing the score for "opening the presentation." When a participant opens the presentation and then clicks the screen-sharing button during the review session, the system accumulates the high-score intent evidence based on the current "quarterly performance review meeting" meeting type, quickly raising the accumulated intent evidence score to the pre-defined locking threshold, thus rapidly identifying the participant as the sole potential focus.

[0093] Through the aforementioned technical solution, the system can intelligently adjust the intent evidence score of pre-conference preparation behaviors based on the actual context of the meeting, such as the meeting type or stage. This significantly improves the accuracy and robustness of meeting focus prediction, avoiding misjudgments or omissions caused by fixed scores. Specifically, in different meeting scenarios, the system can more accurately identify signals that a participant is about to become the focus of the meeting, thereby making the determination of the core video stream more timely and accurate. This, in turn, optimizes the allocation of network resources and the quality management of video streams, enhancing the overall user experience of video conferencing.

[0094] In one possible design, in order to accumulate intent evidence scores, this application also includes: S401. When a participant’s pre-preparation behavior is monitored, an initial intent evidence score is assigned to each pre-preparation behavior, and a timer associated with that pre-preparation behavior is started.

[0095] For example, when monitoring attendees' preparatory behaviors, such as opening a presentation, connecting to a projector, adjusting microphone volume, or opening a screen-sharing application, the client application assigns an initial intent evidence score to each behavior. Simultaneously, a separate timer is started for each behavior. This initial intent evidence score gradually decays over time as the timer runs, meaning that earlier preparatory behaviors have a weaker indicative effect on the current intent.

[0096] S402. Identify the sequence or combination patterns of participants' pre-event preparation behaviors.

[0097] Among them, identifying the sequence pattern or combination pattern of participants' pre-preparation behavior refers to the analysis of a series of consecutive or simultaneous pre-preparation behaviors by the client application or conference server in order to identify behavioral patterns with specific meanings.

[0098] Sequence patterns refer to the order in which actions occur, such as "opening a file" followed immediately by "clicking to share the screen"; combination patterns refer to multiple actions occurring simultaneously or overlapping within a short period of time, such as "opening the camera" and "adjusting the microphone" happening at the same time.

[0099] As one possible implementation, the system can identify the sequence or combination patterns of participants' pre-preparation behaviors through a pre-set rule base, machine learning model, or behavioral graph analysis.

[0100] As one possible implementation, the system can identify the sequence or combination patterns of participants' pre-event preparation behaviors based on the following methods: S4021, Receive pre-preparation behavior event streams from multiple participants.

[0101] Receiving pre-conference preparation event streams from multiple participants refers to the continuous collection and aggregation of various pre-conference preparation behavior data generated by all participants in the video conference by the conference server or client application. This includes actions such as mouse clicks, keyboard input, file opening, and screen sharing previews. This behavior data is typically transmitted in real-time as event streams. Each event stream contains the behavior type of a specific participant, a timestamp of the occurrence, and other relevant contextual information.

[0102] S4022. Extract behavioral patterns from the pre-event preparation behavior flow of each participant.

[0103] Among them, the behavior pattern extraction of the pre-preparation behavior event stream of each participant can be understood as the system analyzing the behavior event stream of each participant received to identify whether there is a specific behavior sequence or combination.

[0104] For example, a participant may perform a series of actions in a short period of time, such as "opening the presentation" -> "adjusting the microphone volume" -> "clicking the screen sharing button". These actions constitute a potential behavioral pattern of "preparing to speak" or "preparing to present".

[0105] It should be noted that behavioral patterns can be extracted through preset rules, machine learning algorithms (such as hidden Markov models and recurrent neural networks), or statistical analysis based on time windows.

[0106] S4023. When multiple participants exhibit similar behavioral patterns, conflict resolution is performed on these similar behavioral patterns to obtain a unique shared intent pattern.

[0107] As one possible implementation, the system can perform a weighted evaluation of each similar behavior pattern based on activity level and context relevance to obtain a pattern priority score; and select the behavior pattern with the highest pattern priority score as the unique shared intent pattern.

[0108] S403. Adjust the initial intent evidence score by weighting based on the identified sequence pattern or combination pattern.

[0109] In one example, if a preparatory action is part of a pre-defined, strongly indicative sequence or combination pattern, then the initial intent evidence score of that action will be given a higher weight, thus making a greater contribution to the total cumulative intent evidence score.

[0110] For example, if the actions of "opening a presentation" and "clicking to share the screen" occur sequentially within a short period of time, the intent evidence score for "clicking to share the screen" may be significantly increased. This weighted adjustment aims to reflect the deeper intent information implied in the behavioral pattern.

[0111] S404. When a participant performs a new preparatory action, the weighted adjusted initial intent evidence score of the new preparatory action is accumulated, and a new timer is started for the new preparatory action.

[0112] In some preferred embodiments, a specific example is given below. Suppose that in a video conference, participant A is preparing to give a screen-sharing presentation.

[0113] First, participant A performed the preparatory action of "opening the presentation". The client application assigned an initial intent evidence score to A, for example, 10 points, and started a timer.

[0114] Next, participant A performed the preparatory action of "clicking the share screen button." At this point, the client application recognized the sequence pattern of "opening the presentation" followed immediately by "clicking the share screen button." Since this sequence pattern is preset as a strongly indicative "preparing for presentation" mode, the initial intent evidence score of "clicking the share screen button" (e.g., 15 points) will be weighted according to this pattern, for example, multiplied by a weighting coefficient of 1.5, resulting in a weighted score of 22.5 points.

[0115] Subsequently, participant A performed the preparatory action of "adjusting microphone volume". The client application recognized the combination pattern of "clicking the screen sharing button" and "adjusting microphone volume", which may have been preset as an auxiliary pattern for "preparing to speak". Therefore, the initial intent evidence score of "adjusting microphone volume" (e.g., 5 points) will be weighted according to this pattern, for example, multiplied by a weighting coefficient of 1.2, so that its weighted adjusted score is 6 points.

[0116] Ultimately, the client application accumulates these weighted and adjusted intent evidence scores (e.g., an initial 10 points decayed to 22.5 points, plus another 6 points) to obtain a higher cumulative intent evidence score. When this cumulative score reaches or exceeds a preset prediction lock-in threshold, participant A will be more quickly and accurately predicted as the meeting focus, and their upcoming video stream (i.e., screen-sharing stream) will be identified as the core video stream and receive high-priority network transmission resources. This approach significantly improves the accuracy of meeting focus prediction and ensures the smooth transmission of critical content.

[0117] Through the aforementioned technical solution, this application can more accurately predict the focus of a meeting. By identifying and weighting the sequence or combination patterns of preparatory behaviors, the system can distinguish between accidental behaviors and behaviors with clear intent, thereby avoiding misjudging the focus due to a single behavior. This intelligent identification and score adjustment mechanism based on behavior patterns makes the accumulation of intent evidence scores more scientific and reasonable, significantly improving the accuracy and robustness of meeting focus prediction. This ensures more accurate identification of core video streams, thereby allocating high-priority network transmission resources to core video streams and degrading the quality of non-core video streams, thus optimizing the overall video conferencing experience. Especially under conditions of limited network resources, it can effectively guarantee the transmission quality of critical content.

[0118] In one possible design, in order to resolve conflicts among similar behavioral patterns and obtain a unique shared intent pattern, this application also includes: S501. Identify the activity level of each similar behavioral pattern.

[0119] Activity level is determined based on the frequency of the behavior within the time window in which the behavior occurred and the recentity of the behavior. The higher the frequency of the behavior and the closer the time of the behavior is to the current moment, the higher the activity level of the behavior pattern.

[0120] S502. Identify the scene relevance of each similar behavioral pattern.

[0121] Among them, the relevance of the scenario is determined based on the degree of matching between the behavioral pattern and the current meeting type or meeting stage.

[0122] For example, it can be determined based on how well the behavioral patterns match the type of the current meeting (e.g., product launch, internal discussion, or technical review) or the stage of the meeting (e.g., opening presentation, thematic discussion, or closing Q&A).

[0123] The aim is to ensure that the identified patterns of sharing intent are aligned with the actual progress and objectives of the meeting. For example, during a product launch, behavioral patterns related to "sharing a presentation" are more relevant to the context than those related to "opening personal notes."

[0124] S503. Based on activity level and context relevance, each similar behavior pattern is weighted and evaluated to obtain a pattern priority score.

[0125] S504. Select the behavior pattern with the highest priority score as the unique shared intent pattern.

[0126] In one possible design, to identify the activity level of each similar behavioral pattern, this application also includes: S601. When monitoring the pre-event preparation behavior of participants, generate an event record for each behavior event, including the behavior type, timestamp, and client application identifier.

[0127] For example, when a participant clicks the screen-sharing button, opens a presentation, or adjusts the microphone volume, an event log is immediately generated for that action. This event log includes not only the type of action but also the precise timestamp when the action occurred and a unique identifier of the client application that performed the action.

[0128] S602: Before sending event logs, the client application synchronizes and calibrates the local system clock.

[0129] For example, the Network Time Protocol (NTP) or other time synchronization mechanisms can be used to synchronize the local clock with an authoritative time server to eliminate time discrepancies caused by client clock drift or initial setting differences.

[0130] S603: After receiving event records from multiple client applications, the conference server performs unified calibration on the timestamp of each event record.

[0131] For example, Network Time Protocol (NTP) or other time synchronization mechanisms can be used to synchronize the local clock with an authoritative time server to eliminate time discrepancies caused by client clock drift or differences in initial settings. Furthermore, after receiving event logs from multiple client applications, the conference server performs a unified timestamp calibration on each event log. Even if the client applications have performed local calibration, the server may still perform a secondary calibration to address network latency or minor time synchronization errors, ensuring that all event logs have a consistent and accurate time reference on the server side.

[0132] S604. The conference server uses the calibrated timestamp to count the frequency of behavioral events within a preset time window and calculates the recentity of the behavioral events.

[0133] Among them, the degree of relevance is used to characterize how close or distant a behavioral event is from the current moment.

[0134] For example, the system can use a decrementing function to calculate the recentity of a behavioral event.

[0135] S605: The conference server calculates the activity contribution of each behavioral event based on its frequency and recentity, combined with a preset decay function, and accumulates them to obtain the activity score.

[0136] The decay function can be an exponential or linear decay function, used to quantify the impact of time on behavioral activity. For example, recently occurring, high-frequency behaviors will receive a higher activity contribution. The activity contributions of all behavioral events are summed to obtain the final activity score for similar behavioral patterns.

[0137] In some preferred embodiments, a specific example is given below. Suppose that in a video conference, participant A and participant B simultaneously exhibit similar behavioral patterns of "preparing to share their screen." Specifically, participant A clicks "Open Presentation" at 10:00:05, clicks "Adjust Microphone" at 10:00:15, and clicks "Share Screen Preview" at 10:00:25. Participant B clicks "Open Presentation" at 10:00:10, clicks "Adjust Microphone" at 10:00:20, and clicks "Share Screen Preview" at 10:00:30.

[0138] Prior to these events, the client applications of both participant A and participant B had synchronized with the time server via the NTP protocol to ensure the accuracy of their local clocks. When these events occurred, the client applications immediately generated an event log containing the event type, timestamp, and client application identifier, and sent it to the conference server.

[0139] For example, participant A's "Open Presentation" event log might be {Type: Open Presentation, Timestamp: 10:00:05.123, Client ID: A}. After receiving these event logs, the meeting server will recalibrate all timestamps to eliminate minor errors that might be introduced by network transmission delays, ensuring that all events have a consistent time base on the server side. Assuming the current time is 10:00:35, the meeting server sets a 30-second time window to count activity. For participant A's behavior patterns: "Open Presentation" (10:00:05): Frequency 1, relatively recent (30 seconds ago). "Adjust Microphone" (10:00:15): Frequency 1, relatively recent (20 seconds ago). "Share Screen Preview" (10:00:25): Frequency 1, very recent (10 seconds ago). For participant B's behavior patterns: "Open presentation" (10:00:10): Frequency 1, relatively recent (25 seconds ago). "Adjust microphone" (10:00:20): Frequency 1, relatively recent (15 seconds ago). "Share screen preview" (10:00:30): Frequency 1, very recent (5 seconds ago). The meeting server calculates the activity contribution of each behavior event based on a preset decay function (e.g., higher recentity, greater contribution; higher frequency, greater contribution), and sums them to obtain the activity of each participant A and participant B's behavior patterns. For example, if the decay function gives greater weight to more recent behaviors, then participant B's "Share screen preview" behavior (10:00:30) will receive a higher recentity score than participant A's (10:00:25). Ultimately, through summation, participant B's overall activity may be slightly higher than participant A's because their key behavior (share screen preview) occurred later, closer to the current moment. In this way, even if two participants exhibit similar behavioral sequences, the system can quantify the activity level more accurately based on the precise time, frequency, and recentity of their behavior. This allows for a more reasonable judgment in subsequent conflict resolution of which party's intention is stronger or more timely, ultimately determining the unique shared intention pattern.

[0140] like Figure 3As shown in the figure, this embodiment of the invention also provides a video conferencing data processing system. The system includes: The instruction recognition module is used to recognize the verbal intentions of participants; the participants are those involved in the video conference. The behavior monitoring module is used to monitor the user operation behavior of participants; The focus prediction module is used to predict the focus of the meeting based on verbal intentions and user actions, and obtain the core video stream. The resource allocation module is used to allocate high-priority network transmission resources to the core video stream and to perform quality degradation processing on non-core video streams other than the core video stream.

[0141] For example, a computer program can be divided into one or more modules / units, one or more of which are stored in memory and executed by a processor to complete the present invention. One or more modules / units can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in a terminal device.

[0142] Terminal devices can be computing devices such as desktop computers, laptops, PDAs, and smart tablets. Terminal devices may include, but are not limited to, processors and memory. Those skilled in the art will understand that the above-described components are merely examples of terminal devices and do not constitute a limitation on the terminal device. The device may include more or fewer components than described above, or a combination of certain components, or different components. For example, a terminal device may also include input / output devices, network access devices, buses, etc.

[0143] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device through various interfaces and lines.

[0144] Memory can be used to store computer programs and / or modules. The processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area can store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). In addition, memory can include high-speed random access memory, and can also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, Flash Card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0145] If the modules / units integrated into the terminal device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or system capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0146] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0147] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention in detail. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A video conferencing data processing method, characterized in that, include: Identify verbal instructions from participants; the participants are those involved in the video conference. Monitor the user actions of the participants; Based on the verbal intent and the user's actions, the focus of the meeting is predicted, and the core video stream is obtained. Allocate high-priority network transmission resources to the core video stream and perform quality degradation processing on non-core video streams other than the core video stream.

2. The video conferencing data processing method according to claim 1, characterized in that, The process of predicting the focus of the meeting and obtaining the core video stream based on the verbal intent and the user's operational behavior includes: When the verbal intent instruction indicates that the participant intends to display important content, the client application presets an intent evidence score for each monitored preparatory behavior and accumulates the intent evidence score. When the cumulative intent evidence score of the participants is greater than or equal to the preset pre-judgment locking threshold, the client application sends a quasi-focus locking request signaling to the conference server; After receiving the quasi-focus lock request signaling, the conference server marks the participant as the unique quasi-focus and determines the video stream that the unique quasi-focus is about to send as the core video stream.

3. The video conferencing data processing method according to claim 2, characterized in that, The process of allocating high-priority network transmission resources to the core video stream and performing quality degradation processing on non-core video streams other than the core video stream includes: Based on the unique quasi-focus, a locked state is entered; the locked state is a state that ignores conflict signals from other participants outside the unique quasi-focus. The conference server pre-allocates network transmission resources for the core video stream to be sent by the unique quasi-focus, and sends a degradation instruction to the client applications of non-unique quasi-focus participants; the degradation instruction is used to require the video encoder of the client applications of non-unique quasi-focus participants to adjust the encoding parameters and perform quality degradation processing. When the single focus clicks the share screen button on the conference application interface, the conference server activates the pre-allocated network channel and encoding parameters.

4. The video conferencing data processing method according to claim 2, characterized in that, The client application presets an intent evidence score for each monitored preparatory behavior and accumulates the intent evidence scores, including: When a participant’s pre-preparation behavior is monitored, an initial intent evidence score is assigned to each pre-preparation behavior, and a timer associated with the pre-preparation behavior is started; the initial intent evidence score decays over time during the timer’s operation. When a participant performs a new preparatory action, the initial intent evidence score for the new preparatory action is accumulated, and a new timer is started for the new preparatory action.

5. A video conferencing data processing method according to claim 4, characterized in that, Assigning an initial intent evidence score to each preparatory action includes: Identify the meeting type or meeting stage; The initial intent evidence score for preparatory behaviors is dynamically adjusted based on the identified meeting type or stage.

6. A video conferencing data processing method according to claim 2, characterized in that, The client application presets an intent evidence score for each monitored preparatory behavior and accumulates the intent evidence scores, including: When a participant’s pre-preparation behavior is monitored, an initial intent evidence score is assigned to each pre-preparation behavior, and a timer associated with that pre-preparation behavior is started; the initial intent evidence score decays over time during the timer’s operation. Identify the sequence or combination patterns of the participants' pre-event preparation behaviors; The initial intent evidence score is weighted and adjusted based on the identified sequence pattern or combination pattern. When a participant performs a new preparatory action, the weighted adjusted initial intent evidence score of the new preparatory action is accumulated, and a new timer is started for the new preparatory action.

7. A video conferencing data processing method according to claim 6, characterized in that, The sequence pattern or combination pattern for identifying the pre-event preparation behavior of the participants includes: Receive a stream of pre-construction behavior events from multiple participants; Behavioral patterns are extracted from the pre-event preparation behavior flow of each participant; When multiple participants exhibit similar behavioral patterns, conflict resolution is performed on these similar behavioral patterns to obtain a unique shared intent pattern.

8. A video conferencing data processing method according to claim 7, characterized in that, The process of resolving conflicts among the similar behavioral patterns to obtain a unique shared intent pattern includes: The activity level of each similar behavioral pattern is identified, and the activity level is determined based on the frequency of the behavior occurring within a time window and the recentity of the behavior. Identify the scenario relevance of each similar behavioral pattern, the scenario relevance being determined based on the degree of matching between the behavioral pattern and the current meeting type or meeting stage; Based on the activity level and the relevance of the scenario, each similar behavior pattern is weighted and evaluated to obtain a pattern priority score; The behavior pattern with the highest pattern priority score is selected as the sole shared intent pattern.

9. A video conferencing data processing method according to claim 8, characterized in that, The identification of the activity level of each similar behavioral pattern includes: When monitoring participants' pre-event preparation behavior, generate an event record for each behavior event, including the behavior type, timestamp, and client application identifier; Before sending the event record, the client application synchronizes and calibrates the local system clock. After receiving the event records from multiple client applications, the conference server uniformly calibrates the timestamp of each event record. The conference server uses the calibrated timestamps to count the frequency of behavioral events within a preset time window and calculates the recentity of the behavioral events; the recentity is used to characterize how close or far the behavioral event is from the current moment. The conference server calculates the activity contribution of each behavioral event based on the occurrence frequency and the recentity, combined with a preset decay function, and accumulates them to obtain the activity score.

10. A video conferencing data processing system, characterized in that, The system includes: The instruction recognition module is used to recognize the verbal intentions of participants; the participants are those involved in the video conference. The behavior monitoring module is used to monitor the user operation behavior of the participants; The focus prediction module is used to predict the focus of the meeting based on the verbal intention instructions and the user's operation behavior, and obtain the core video stream. The resource allocation module is used to allocate high-priority network transmission resources to the core video stream and to perform quality degradation processing on non-core video streams other than the core video stream.