An OPS framework conference communication system for virtual enhanced digital conferences

Through the multimodal sensing array and dynamic window allocation unit, the screen ratio between the screen and the figure is automatically adjusted, which solves the problem of low efficiency in gesture occlusion and manual layout in traditional digital meetings, and realizes clear display of screen content and efficient information transmission.

CN120263931BActive Publication Date: 2025-08-05CHENGDU TIANTANGYUN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510703466.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-05
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In traditional digital conferences, speaker gestures or tools obscures the content of the screen, resulting in unclear information, manual adjustment of window layout is inefficient, and the inability to intelligently identify speaker intentions, affecting the efficiency of the conference and information transmission.

Method used

The multimodal sensing array unit is used to detect the occlusion in real time, generate a translucent virtual structure covering the occlusion area, combine it with the dynamic window allocation unit to automatically adjust the ratio of the screen and the character screen based on behavioral intention, use the Bezier curve to achieve smooth scaling, and generate virtual guide lines through the virtual and real fusion rendering unit to enhance interactivity and information communication.

Benefits of technology

Ensure that the screen content is clear and visible, reduce manual intervention, improve meeting fluency and information communication efficiency, avoid unclear content display problems caused by unreasonable screen proportions, and enhance immersive prompt effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263931B_ABST
    Figure CN120263931B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image communication technology, and in particular to an OPS framework conference communication system for virtual enhanced digital conferences. The system comprises a multimodal sensor array unit, a dynamic window allocation unit and a virtual-reality fusion rendering unit. The present invention can automatically switch between content explanation mode and free speech mode based on behavioral intention, use Bezier curves to achieve smooth window scaling, automatically adjust the screen and character image ratio, reduce manual intervention, and at the same time, the screen window splitting module detects the edges of multiple screens, divides independent windows and renders them separately, ensuring that the content of each screen is clearly displayed and avoiding image compression caused by distance. In addition, during conference communications, the multimodal sensor array is used to detect obstructions in real time, generate a translucent virtual structure to cover the obstructed area, and synchronize gesture trajectories to ensure virtual demonstration meetings and improve the quality of screen display content and dynamic demonstrations of digital conferences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image communication, and in particular to an OPS framework conference communication system for virtual enhanced digital conference. Background Art

[0002] With the rapid development of science and technology and society, people are exposed to an increasing amount of information in their daily lives and work. Therefore, information exchange and communication between people are becoming more frequent and more important. Business negotiations, product demonstrations, guest meetings, and the issuance of government orders are all forms of communication between people. Among them, digital conferencing utilizes computer, digital, and network technologies to network various systems. It is an automated conference management system that integrates computer, communication, automatic control, multimedia, image, and audio technologies. The system organically connects independent subsystems such as conference registration, speaking, voting, video, audio, display, and network access. A central control computer coordinates the work of each subsystem according to the meeting agenda, providing the most accurate and timely information and services for various lectures, academic conferences, and remote conferences.

[0003] However, in traditional digital conferences, the main speaker usually delivers a speech by displaying content on a display screen. However, during the speech, the speaker often uses auxiliary tools such as arms or pointing sticks to point or control operations. This operation method can easily cause the arm to block the content on the display screen, affecting the viewing experience of other participants. In addition, digital conferences have poor mobility and cannot automatically switch the screen layout according to the speaker's gestures or behaviors, resulting in an unreasonable screen ratio distribution between the speaker and the display content, further reducing meeting efficiency. In view of this, we propose an OPS framework conference communication system for virtual enhanced digital conferences. Summary of the Invention

[0004] The purpose of the present invention is to provide an OPS framework conference communication system for virtual enhanced digital conferencing. Through a closed loop of multimodal perception, dynamic decision-making, and virtual-reality rendering, the system significantly improves the naturalness of interaction and the efficiency of information transmission in digital conferencing, thereby solving the following problems raised in the above background technology:

[0005] 1. The speaker's movements obscure the screen content, making the information unclear;

[0006] 2. Manually adjusting window layouts is inefficient, affecting meeting smoothness. Furthermore, content displayed in multi-screen meetings is too small to be effectively presented.

[0007] 3. Traditional systems are unable to intelligently identify the speaker's intent and switch display modes. Furthermore, the virtual overlay effect is not precise enough, which may cause interference or delays. This results in inaccurate guidance when pointing with tools such as pointers, affecting the delivery of key points.

[0008] To achieve the above-mentioned object, the present invention provides an OPS framework conference communication system for virtual enhanced digital conference, comprising a multimodal sensor array unit, a dynamic window allocation unit and a virtual-reality fusion rendering unit;

[0009] The multimodal sensor array unit is used to collect speaker behavior data and screen occlusion status;

[0010] The dynamic window allocation unit is used to extract behavioral data features, fuse the attention model to predict the focus area and output behavioral intentions, and use Bezier curves to achieve smooth scaling of the conference view window. The behavioral intentions include content explanation mode and free speech mode.

[0011] When the content explanation mode is detected, the window is smoothly scaled to prioritize the screen video perspective, and when the virtual-reality fusion rendering unit senses that the screen occlusion state is a signal of being blocked, the cover feature is generated into a semi-transparent virtual feature structure to cover the screen video perspective, so that the virtual feature structure is synchronized with the behavior data;

[0012] When the free speech mode is detected, the window is smoothly scaled to prioritize the person's video perspective, and the virtual-reality fusion rendering unit is enabled to form a virtual guide line according to the speaker's behavior data.

[0013] As a further improvement of the present technical solution, the multimodal sensing array unit includes a video stream acquisition module, a motion capture module and a screen occlusion detection module;

[0014] The video stream acquisition module uses an RGB camera to acquire the speaker's depth information and video stream, and renders the resolution through the OPS framework to align the RGB image with the depth map, combining color and depth information to generate a three-dimensional point cloud;

[0015] The motion capture module extracts a feature graph formed by the speaker's posture data in the three-dimensional point cloud based on a feature extractor, detects key points of the human body in the feature graph, connects the key points using a graph model, and constructs human posture graph data;

[0016] The screen occlusion detection module is used to use an edge detection algorithm to identify screen edge features in the video stream, extract straight line edges, fit the screen quadrilateral boundary, calculate the homography matrix based on the detected quadrilateral vertices, map the screen viewing area into a rectangular ROI, and lock the screen features within the screen video viewing area. Within the rectangular ROI, the current frame is compared with the background frame to generate a differential binary image, perceive the features of the occlusion object other than the screen features, and issue an occlusion warning signal.

[0017] As a further improvement of the present technical solution, the dynamic window allocation unit includes a behavior feature extraction module, an intention analysis module and a window switching module;

[0018] The behavior feature extraction module is used to receive the posture graph data of the motion capture module and convert it into a time series, each time step contains key point data, adopts the LSTM model to train the model using the labeled data set, receives the key point data of each time step at the input layer, extracts the time-dependent features at the LSTM layer, and generates the classified action at the output layer;

[0019] The intention analysis module is used to establish a mapping relationship between classified actions and behavioral intentions based on historical data, input the classified actions of the behavioral feature extraction module into the mapping relationship, and output the behavioral intention;

[0020] The window switching module is used to receive the behavioral intention output by the intention analysis module to implement Bezier curve smooth scaling, define the current layout and target layout of the window scaling path, set control points to adjust the curve smoothness, and share window switching, including the following gestures:

[0021] Posture 1, content explanation mode: the screen window ratio is increased to 70%, and the character window is reduced to 30%;

[0022] Posture 2, Free Speech Mode: The character window is expanded to 80% and the screen window is reduced to 20%.

[0023] As a further improvement of the present technical solution, the output end of the intention analysis module is connected to a sound source localization module, which forms a microphone array to collect digital conference sound communications, calculates the time alignment cost of the active voice interval and the action sequence, finds the minimum cumulative cost path, aligns the voice peak with the start and end time of the action, and enables the sound source localization module and the behavior feature extraction module to fuse the time series data of the aligned active voice interval and the action capture. If the human posture graph data is detected to have an interactive action with the screen when the voice is active, it is weightedly determined to be in "content explanation mode". If the voice is active but there is no screen interactive action, and the transition time period is exceeded, it is weightedly determined to be in the state of "free speech mode";

[0024] As a further improvement of the present technical solution, the sound source localization module also includes calculating the sound source direction through the phase difference of the microphone array, locating the sound source coordinates based on the signal arrival time difference, inputting the sound source coordinates and visual positioning into the Kalman filter, outputting the optimal estimated position, and driving the RGB camera of the video stream acquisition module to track the weighted determination position.

[0025] As a further improvement of the present technical solution, the multimodal sensing array unit also includes a screen window splitting module, which is used to detect the number of straight edges output by the screen occlusion detection module, calculate the number of screens, divide the number of screen windows of the window switching module to match the number of screens, and make the multiple rectangular ROIs formed correspond to the screen window displays respectively.

[0026] As a further improvement of this technical solution, the virtual-reality fusion rendering unit includes an occlusion feature analysis module, a virtual structure generation module and a trajectory modeling module;

[0027] The occlusion feature analysis module is used to receive the occlusion warning signal from the screen occlusion detection module, segment the occluders in the screen area in real time, and extract the contour features of the occluders;

[0028] The virtual structure generation module is used to virtually generate a semi-transparent AR layer that matches the outline features of the occluder as a virtual feature structure, receive the human posture graph data provided by the motion capture module, extract the motion path in the human posture graph data, and synchronously render it to form a motion trajectory of the semi-transparent AR layer;

[0029] The trajectory modeling module is used to locate the hand joint point motion path that matches the human body posture diagram data of the motion capture module in the character window of the window switching module, and dynamically generate the light spot trajectory movement in the hand joint point motion path. The color of the light spot trajectory changes with time, and the color intensity is associated with the gesture amplitude.

[0030] As a further improvement to the present technical solution, the occlusion feature analysis module further includes a contour analysis module, which is used to detect the degree of match between the occluder's motion trajectory and the human posture graph data output by the motion capture module. If the occluder's motion trajectory overlaps with the speaker's hand trajectory, it is determined to be active occlusion and transmitted to the virtual structure generation module for matching with the semi-transparent AR layer. If the occluder's motion trajectory does not overlap with the speaker's hand trajectory, it is determined to be passive occlusion, and the occluder's contour area and aspect ratio are calculated to filter out occluder interference.

[0031] As a further improvement of the present technical solution, the virtual structure generation module also includes an end light module, which is used to receive the contour area and aspect ratio of the obstruction output by the contour analysis module. Since the arm occlusion usually appears in irregular blocks, while the pointer is a slender strip, the aspect ratio and area ratio of the contour are calculated to distinguish the arm and the pointer. When outputting the pointer, the pointer is morphologically skeletonized, the center line is extracted, and the end coordinates are marked at the skeleton endpoints, and a particle emitter is generated at the end coordinates.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] 1. In traditional conferencing systems, a speaker's gestures or tools (such as arms or pointers) can obstruct screen content, affecting attendees' access to key information. This system uses a multimodal sensor array to detect obstructions in real time, generate a semi-transparent virtual structure to cover the obscured area, and synchronize the gesture trajectory to ensure clear on-screen content. For example, an arm or pointer can be transformed into a semi-transparent arrow to dynamically indicate the content being pointed to.

[0034] 2. To address the problem of existing systems relying on manual switching of screen layouts, which disrupts meeting flow, the dynamic window allocation unit uses Bezier curves to achieve smooth window scaling based on behavioral intent (such as content explanation mode and free speech mode), automatically adjusting the screen-to-person screen ratio (e.g., 70% screen + 30% person), reducing manual intervention. The screen window splitting module also detects the edges of multiple screens and renders them in separate windows, ensuring clear display of content on each screen and avoiding image compression caused by distance.

[0035] 3. To address the issue of virtual guidance delaying or interfering with real content, the trajectory modeling module dynamically adjusts virtual features based on depth information, generating gradient-colored particles at the end of the pointer based on movement speed and depth. If the pointer is behind the screen, the brightness is reduced to avoid visual interference. At the same time, the end lighting module skeletonizes the pointer, generates a dynamic halo at the endpoint, and sprays particles based on the waving direction to enhance the immersive prompting of the pointed content. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a block diagram of the overall structural principle of the present invention;

[0037] Figure 2 A detailed schematic diagram of the overall structure of the present invention;

[0038] Figure 3 A conference communication window demonstration diagram for analysis of the intent of the present invention.

[0039] The meaning of each number in the figure is:

[0040] 100, multimodal sensor array unit; 110, video stream acquisition module; 120, motion capture module; 130, screen occlusion detection module;

[0041] 200, dynamic window allocation unit; 210, behavior feature extraction module; 220, intention analysis module; 230, window switching module;

[0042] 300, virtual-reality fusion rendering unit; 310, occlusion feature analysis module; 320, virtual structure generation module; 330, trajectory modeling module. DETAILED DESCRIPTION

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0044] See also Figure 1-Figure 3 As shown, this embodiment provides an OPS framework conference communication system for virtual enhanced digital conference, including a multimodal sensor array unit 100, a dynamic window allocation unit 200 and a virtual-reality fusion rendering unit 300;

[0045] The multimodal sensor array unit 100 is used to collect speaker behavior data and screen occlusion status;

[0046] Specifically, the multimodal sensor array unit 100 includes a video stream acquisition module 110, a motion capture module 120, and a screen occlusion detection module 130;

[0047] The video stream acquisition module 110 uses an RGB camera to obtain the depth information and video stream of the speaker. The RGB camera can capture color images, provide color information of the scene, and simultaneously capture depth images. Each pixel represents the distance from the camera. The depth is calculated by projecting light of a known pattern and analyzing the deformation. The resolution is rendered through the OPS framework to align the RGB image with the depth map. The color and depth information are combined to generate a three-dimensional point cloud. The transmission is optimized through the OPS framework, the rendering resolution is adjusted, the quality and bandwidth are balanced, and the synchronization of RGB and depth data is ensured. The synchronization mechanism provided by the camera SDK is used to ensure the synchronization of RGB and depth frames. When rendering the resolution, the RGB and depth data are superimposed to generate an enhanced video stream. The resolution is adjusted using the image pyramid technology. The video is then compressed using H.264 encoding. A transmission channel is established through WebRTC for real-time transmission to the cloud for convenient processing in the cloud.

[0048] The motion capture module 120 extracts a feature graph formed by the speaker's posture data from the three-dimensional point cloud based on the VGG-19 feature extractor, detects key points of the human body (such as hands, head, shoulders, etc.) in the feature graph, connects the key points using a graph model, and constructs human posture graph data;

[0049] The screen occlusion detection module 130 is used to use an edge detection algorithm to identify screen edge features in the video stream, extract straight edges, fit the screen quadrilateral boundary, calculate the homography matrix based on the detected quadrilateral vertices, map the screen viewing area into a rectangular ROI, and lock the screen features within the screen video viewing area. Within the rectangular ROI, it compares the current frame with the background frame (the screen image when there is no occlusion) to generate a differential binary image. It senses the features of obstructions other than the screen features and issues an occlusion warning signal. This helps to remind the speaker if they unconsciously block the screen content, so as to avoid affecting other conference participants' viewing of the screen content.

[0050] It is worth noting that the dynamic window allocation unit 200 is used to extract behavioral data features, fuse the attention model, predict the focus area, and output behavioral intentions, and use the Bezier curve to achieve smooth scaling of the conference view window. This is conducive to dynamically adjusting the screen layout based on the speaker's behavioral characteristics, making the screen ratio distribution of the speaker and the display content more reasonable, ensuring the clear display of the speaker and the display content, avoiding blurred content display or information loss due to unreasonable screen ratio, and improving meeting efficiency. At the same time, it realizes the intelligent switching of the speaker's control instructions and conference scenes, reduces manual operations, and improves user experience. Behavioral intentions include content explanation mode (the speaker focuses on the content of the display screen, frequently points to and looks at the screen) and free speech mode (the speaker does not interact with the display screen for a certain period of time, the speaker moves away from the screen, and the amplitude of gestures increases);

[0051] When a content explanation mode is detected, the window is smoothly scaled to prioritize the screen video perspective. When the virtual-reality fusion rendering unit 300 senses a signal indicating that the screen is blocked, the obstructing object feature (which may be the speaker's arm or pointing tool) is converted into a semi-transparent virtual feature structure to cover the screen video perspective. Synchronizing the virtual feature structure with behavioral data facilitates converting the speaker's pointing movements into transparent or semi-transparent virtual feature structures, preventing the arm from blocking the display screen content and improving the clarity of the content displayed.

[0052] When the free speech mode is detected, the window is smoothly scaled to prioritize the character video perspective, and the virtual-reality fusion rendering unit 300 forms a virtual guide line based on the speaker's behavior data, which is conducive to the virtual demonstration of the speaker's behavior and enables the meeting participants to observe the speaker's speech more vividly.

[0053] Based on the above content, the specific working principle is disclosed in detail:

[0054] First, the dynamic window allocation unit 200 includes a behavior feature extraction module 210, an intention analysis module 220 and a window switching module 230;

[0055] The behavioral feature extraction module 210 is used to receive the posture graph data from the motion capture module 120 and convert it into a time series. Each time step contains key point data. The LSTM model is used to train the model using a labeled dataset (video data containing different speech scenes is collected and the action categories and timestamps are manually annotated). The input layer receives the key point data of each time step, extracts time-dependent features in the LSTM layer, and generates classified actions (such as pointing or looking at the screen) in the output layer. Specifically, the key point data of consecutive frames are arranged in chronological order to form a time series of key point data, which is input into the LSTM model and the classified actions are output. Note that the key point coordinates are normalized (mean is zeroed and variance is normalized) to prevent scale differences from affecting model training.

[0056] The intent analysis module 220 is used to establish a mapping relationship between classified actions and behavioral intentions based on historical data, input the classified actions from the behavioral feature extraction module 210 into the mapping relationship, and output the behavioral intention. If the speaker is not pointing or looking at the screen within a time period, it means that the speaker continues to speak but does not interact with the display. Conversely, if the speaker points or looks at the screen within the time period, it means that the speaker is focused on the content on the display and frequently points and looks at the screen for conference communication. Combining the deep learning model of the LSTM model with time series analysis improves the accuracy of action recognition, further accurately classifies the speaker's status in real time, and enhances the robustness of the system in complex environments.

[0057] Specifically, when establishing the mapping relationship between classified actions and behavioral intentions, the classified action sequence is input into the Bi-LSTM model to capture long-term dependencies, output pattern probabilities, integrate the rule engine with the model prediction results (weighted voting), collect user feedback data (such as manual correction of pattern classification results), and fine-tune the model parameters to achieve the prediction of behavioral intentions based on classified actions. Note that the prediction results have a certain degree of probability. People in this technical field can easily associate real-time updates to ensure the accuracy of the mapping relationship.

[0058] The window switching module 230 is used to receive the behavioral intent output by the intent analysis module 220 to implement Bezier curve smooth scaling, define the current layout and target layout of the window scaling path, set control points to adjust the curve smoothness, and share window switching, including the following gestures:

[0059] Posture 1, content explanation mode: the screen window ratio is increased to 70%, and the character window is reduced to 30%;

[0060] Posture 2, Free Speech Mode: The character window is expanded to 80% and the screen window is reduced to 20%;

[0061] Therefore, dynamically adjusting the window layout and achieving smooth transition based on behavioral intentions can achieve efficient and natural conference screen layout adjustments.

[0062] To further significantly improve the accuracy and robustness of the intent analysis module 220 in identifying behavioral intentions, the output end of the intent analysis module 220 is connected to a sound source localization module. The sound source localization module forms a microphone array to collect digital conference audio communications. On the one hand, the time alignment cost of the active speech interval and the action sequence is calculated, and the minimum cumulative cost path is found. The speech peak and the start and end time of the action are aligned to solve the time offset problem of the speech and action data. The sound source localization module and the behavior feature extraction module 210 are integrated with the aligned speech active interval and the time series data of the action capture. If the human posture graph data is detected to have interaction with the screen when the speech is active, it is weightedly determined to be in "content explanation mode". If the speech is active but there is no screen interaction, and the transition time period is exceeded (to avoid switching modes during the speaker's explanation transition stage, to ensure that the speaker actually interacts with the screen but does not interact for more than the transition time period), it is weightedly determined to be in "free speech mode", thereby implementing an anti-misjudgment mechanism.

[0063] On the other hand, the sound source localization module also includes calculating the sound source direction through the phase difference of the microphone array, and locating the sound source coordinates based on the signal arrival time difference, which is conducive to accurately locating the speaker's sound source coordinates through the microphone array, supplementing the shortcomings of visual positioning, and inputting the sound source coordinates and visual positioning into the Kalman filter, outputting the optimal estimated position, and driving the RGB camera of the video stream acquisition module 110 to track the weighted determination position. Specifically, according to the optimal estimated position coordinates output by the Kalman filter, the pan-tilt rotation angle and zoom are controlled, and the PID controller is used to adjust the pan-tilt movement speed to avoid jitter, so as to keep the screen and the speaker centered in the content explanation mode, and keep the speaker centered in the free speech mode. The camera control and layout animation improve the smoothness of the meeting, which is suitable for complex meeting scenarios, especially in remote collaboration, educational speeches and other fields. It has broad application prospects.

[0064] Furthermore, considering that during conference communications, speakers may use multiple screens to conduct digital conferences simultaneously, if the screen occlusion detection module 130 uses two screen edges that are farther apart as the boundary lines when detecting the screen boundary lines, the content displayed on multiple screens in the window corresponding to the screen video perspective will be smaller, affecting the conference effect. Therefore, in a multi-screen conference scenario, the number of screens is dynamically detected and the windows are split to ensure that the content of each screen is displayed independently, thereby improving the picture clarity and readability. The multimodal sensor array unit 100 also includes a screen window splitting module. The screen window splitting module is used to detect the number of straight edges output by the screen occlusion detection module 130, calculate the number of screens, and match the number of screen windows of the window switching module 230 with the number of screens, and make the multiple rectangular ROIs formed correspond to the screen windows for display. This is beneficial for when giving a speech on multiple screens. A window can be divided according to each screen, and the content of each screen can be rendered separately, avoiding the text / image being too small due to the merged display, and realizing the separate execution of digital conferences on each screen, avoiding the distance between the two screens causing the video capture screen to be small, affecting the content display effect.

[0065] Then, the virtual-reality fusion rendering unit 300 includes an occlusion feature analysis module 310 , a virtual structure generation module 320 , and a trajectory modeling module 330 ;

[0066] The occlusion feature analysis module 310 is used to receive the occlusion warning signal from the screen occlusion detection module 130, segment the occluders (arms / pointers) in the screen area in real time, and extract the contour features of the occluders. Specifically, it receives the differential binary image and background subtraction result output by the screen occlusion detection module 130, and obtains the color image and depth image captured by the RGB camera. When segmenting the occluders:

[0067] Perform morphological processing (such as corrosion and dilation) on the difference binary image to remove noise and smooth the boundaries. Use the depth image to extract the depth range of the occluder, separate the foreground occluder from the background content, and generate the occluder mask.

[0068] When extracting contour features, use edge detection algorithms (such as the Canny operator) to extract the contour of the occluder, calculate the key geometric features of the contour (such as area, perimeter, main axis direction, etc.), and output the mask and contour features of the occluder;

[0069] The virtual structure generation module 320 is used to virtually generate a semi-transparent AR layer that matches the outline features of the occluder as a virtual feature structure, receive the human posture graph data provided by the motion capture module 120, extract the motion path in the human posture graph data, and synchronously render it to form a motion trajectory of the semi-transparent AR layer. When generating the AR layer, a semi-transparent AR layer that matches the outline features of the occluder is generated, the color, transparency, and shape parameters of the AR layer are set, the motion path of the hand joint points is converted into the motion trajectory of the AR layer, and the trajectory lines or arrows are dynamically drawn on the AR layer. The color gradually changes with the gesture amplitude, which is conducive to improving the clarity and interactivity of the virtual enhanced digital conference system. It is particularly suitable for complex conference scenes and remote collaboration environments, and does not affect the real-time viewing of screen content and the guidance of the speaker by the conference participants.

[0070] The trajectory modeling module 330 is used to locate the hand joint motion path that matches the human posture diagram data of the motion capture module 120 in the character window of the window switching module 230, and dynamically generate the movement of light spot trajectories in the hand joint motion path. The color of the light spot trajectory changes with time, and the color intensity is related to the amplitude of the gesture. When the gesture is swung to the right, the translucent arrow is generated 0.2 seconds in advance to avoid visual delay, which is beneficial for the speaker to display gestures more dynamically when using gestures to demonstrate the content of the meeting.

[0071] Considering that the interference may be from a small area other than the speaker (such as a mosquito in front of the camera), it is meaningless to perform virtual enhancement at this time. Therefore, the occlusion feature analysis module 310 also includes a contour analysis module. The contour analysis module is used to detect the matching degree between the movement trajectory of the occluding object and the human posture map data output by the motion capture module 120. If the movement trajectory of the occluding object overlaps with the movement trajectory of the speaker's hand, it is determined to be active occlusion (such as occlusion caused by teaching aids or arms) and transmitted to the virtual structure generation module 320 for matching with the semi-transparent AR layer. If the movement trajectory of the occluding object does not overlap with the movement trajectory of the speaker's hand, it is determined to be passive occlusion (occlusion caused by mosquitoes, environmental impurities, etc.). The area and aspect ratio of the occluding object's contour are calculated, and the interference of the occluding object is filtered out. The small area interference such as mosquitoes is effectively identified and filtered out, avoiding unnecessary AR enhancement.

[0072] Furthermore, when the types of occluders are different, the virtual feature structures generated by the virtual structure generation module 320 are all of the same type, which in turn affects the speaker's direction. For example, when a speaker is giving a speech with a pointer, the direction of the pointer's end indicates the content that needs to be focused on, but the meeting participants cannot accurately distinguish it. Therefore, the virtual structure generation module 320 also includes an end lighting module, which is used to receive the occluder contour area and aspect ratio output by the contour analysis module. Since arm occlusion usually presents an irregular block shape, while the pointer is a slender strip, the aspect ratio and area ratio of the contour are calculated to distinguish between the arm and the pointer. If the aspect ratio is greater than 5:1 and the area is less than 200 pixels², it is determined to be a pointer. If the aspect ratio is less than 3:1 and the area is greater than 500 pixels², it is determined to be an arm. When outputting the pointer, the pointer is morphologically skeletonized, the center line is extracted, and the skeleton endpoints are marked as end coordinates. A particle emitter is generated at the end coordinates. The parameter settings are as follows:

[0073] Number of particles: 50-100 / second (adaptive according to movement speed);

[0074] Color gradient: transitions from bright yellow in the center (RGB: 255, 255, 100) to translucent red at the periphery (RGB: 255, 0, 0, α = 0.3);

[0075] Movement mode: spray in the direction of the pointer swing, speed = pointer tip speed × 0.8;

[0076] The depth value of the pointer tip is compared with the depth of the screen content. If the pointer is in front of the screen, the halo is rendered normally. If the pointer is behind the screen (such as pointing outside the edge of the screen), the halo brightness is reduced (brightness coefficient = 0.3) to avoid visual interference. The AlphaBlending formula is used to blend the virtual halo and the real picture to render a golden halo with a diameter of 30px at the endpoint, and the tail length is 50px as the pointer moves. When the pointer points to the resistor symbol on the circuit diagram, the halo is translucently superimposed, and the symbol content is clearly visible, significantly improving the immersion and information communication efficiency of remote collaboration.

[0077] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. An OPS framework conference communication system for virtual enhanced digital conference, characterized by: It includes a multimodal sensor array unit (100), a dynamic window allocation unit (200), and a virtual-reality fusion rendering unit (300); The multimodal sensor array unit (100) is used to collect speaker behavior data and screen occlusion status; The dynamic window allocation unit (200) is used to extract behavioral data features, fuse the attention model, predict the focus area, and output behavioral intentions, and use Bezier curves to achieve smooth scaling of the conference view window. The behavioral intentions include content explanation mode and free speech mode. When the content explanation mode is detected, the window is smoothly scaled to prioritize the screen video perspective, and when the virtual-reality fusion rendering unit (300) senses that the screen occlusion state is a signal of being occluded, the covering object feature is generated into a semi-transparent virtual feature structure to cover the screen video perspective, so that the virtual feature structure is synchronized with the behavior data; When the free speech mode is detected, the window is smoothly scaled to prioritize the person's video perspective, and the virtual-reality fusion rendering unit (300) forms a virtual guide line according to the speaker's behavior data.

2. The OPS framework conference communication system for virtual enhanced digital conference according to claim 1 is characterized by: The multimodal sensing array unit (100) comprises a video stream acquisition module (110), a motion capture module (120), and a screen occlusion detection module (130); The video stream acquisition module (110) uses an RGB camera to acquire the depth information and video stream of the speaker, and renders the resolution through the OPS framework to align the RGB image with the depth map, combining the color and depth information to generate a three-dimensional point cloud; The motion capture module (120) extracts a feature graph formed by the posture data of the speaker in the three-dimensional point cloud based on a feature extractor, detects key points of the human body in the feature graph, connects the key points using a graph model, and constructs human posture graph data; The screen occlusion detection module (130) is used to use an edge detection algorithm to identify screen edge features in a video stream, extract straight line edges, fit the screen quadrilateral boundary, calculate a homography matrix based on the detected quadrilateral vertices, map the screen viewing area to a rectangular ROI, and lock the screen features within the screen video viewing area. Within the rectangular ROI, the current frame is compared with the background frame to generate a differential binary image, and the features of the occlusion object other than the screen features are sensed, and an occlusion warning signal is issued.

3. The OPS framework conference communication system for virtual enhanced digital conference according to claim 2 is characterized by: The dynamic window allocation unit (200) includes a behavior feature extraction module (210), an intention analysis module (220) and a window switching module (230); The behavior feature extraction module (210) is used to receive the posture graph data of the motion capture module (120) and convert it into a time series, each time step contains key point data, adopts an LSTM model to train the model using a labeled data set, receives the key point data of each time step at the input layer, extracts time-dependent features at the LSTM layer, and generates classified actions at the output layer; The intention analysis module (220) is used to establish a mapping relationship between classified actions and behavioral intentions based on historical data, input the classified actions of the behavioral feature extraction module (210) into the mapping relationship, and output the behavioral intention; The window switching module (230) is used to receive the behavioral intention output by the intention analysis module (220) to implement Bezier curve smooth scaling, define the current layout and target layout of the window scaling path, set control points to adjust the curve smoothness, and share window switching, including the following gestures: Posture 1, content explanation mode: the screen window ratio is increased to 70%, and the character window is reduced to 30%; Posture 2, Free Speech Mode: The character window is expanded to 80% and the screen window is reduced to 20%.

4. The OPS framework conference communication system for virtual enhanced digital conference according to claim 3 is characterized by: The output end of the intention analysis module (220) is connected to a sound source localization module, which forms a microphone array to collect digital conference sound communications, calculates the time alignment cost of the voice active interval and the action sequence, finds the minimum cumulative cost path, aligns the voice peak and the action start and end time, and enables the sound source localization module and the behavior feature extraction module (210) to fuse the aligned voice active interval and the action capture time series data. If the human body posture graph data is detected to have an interactive action with the screen when the voice is active, it is weighted to be judged as the "content explanation mode". If the voice is active but there is no screen interactive action, and the transition time period is exceeded, it is weighted to be judged as the state in the "free speech mode".

5. The OPS framework conference communication system for virtual enhanced digital conference according to claim 4 is characterized by: The sound source localization module also includes calculating the sound source direction through the phase difference of the microphone array, locating the sound source coordinates based on the signal arrival time difference, inputting the sound source coordinates and the visual positioning into the Kalman filter, outputting the optimal estimated position, and driving the RGB camera of the video stream acquisition module (110) to track the weighted determination position.

6. The OPS framework conference communication system for virtual enhanced digital conference according to claim 4 is characterized by: The multimodal sensing array unit (100) further includes a screen window splitting module, the screen window splitting module being used to detect the number of straight line edges output by the screen occlusion detection module (130), calculate the number of screens, divide the number of screen windows of the window switching module (230) to match the number of screens, and make the formed multiple rectangular ROIs correspond to the screen windows for display respectively.

7. The OPS framework conference communication system for virtual enhanced digital conference according to claim 6 is characterized by: The virtual-reality fusion rendering unit (300) includes an occlusion feature analysis module (310), a virtual structure generation module (320), and a trajectory modeling module (330); The occlusion feature analysis module (310) is used to receive the occlusion warning signal from the screen occlusion detection module (130), segment the occluders in the screen area in real time, and extract the contour features of the occluders; The virtual structure generation module (320) is used to virtually generate a semi-transparent AR layer that matches the outline features of the occluder as a virtual feature structure, receive the human body posture graph data provided by the motion capture module (120), extract the motion path in the human body posture graph data, and synchronously render it to form a motion trajectory of the semi-transparent AR layer; The trajectory modeling module (330) is used to locate the hand joint point motion path of the human body posture diagram data of the motion capture module (120) in the character window of the window switching module (230), and dynamically generate the light spot trajectory movement in the hand joint point motion path. The color of the light spot trajectory changes with time, and the color intensity is associated with the gesture amplitude.

8. The OPS framework conference communication system for virtual enhanced digital conference according to claim 7 is characterized by: The occlusion feature analysis module (310) further includes a contour analysis module, which is used to detect the matching degree between the occlusion movement trajectory and the human body posture diagram data output by the motion capture module (120). If the occlusion movement trajectory overlaps with the speaker's hand trajectory, it is determined to be active occlusion and transmitted to the virtual structure generation module (320) to match the semi-transparent AR layer. If the occlusion movement trajectory does not overlap with the speaker's hand trajectory, it is determined to be passive occlusion, and the occlusion contour area and aspect ratio are calculated to filter the occlusion interference.

9. The OPS framework conference communication system for virtual enhanced digital conference according to claim 8 is characterized by: The virtual structure generation module (320) also includes an end lighting module, which is used to receive the area and aspect ratio of the occluder contour output by the contour analysis module. Since the arm occlusion usually presents an irregular block shape, while the pointer is a slender strip, the aspect ratio and area ratio of the calculated contour are used to distinguish the arm and the pointer. When outputting the pointer, the pointer is morphologically skeletonized, the center line is extracted, and the end coordinates are marked at the skeleton endpoints, and a particle emitter is generated at the end coordinates.

Citation Information

Patent Citations

  • Meeting interaction method and system based on meta universe

    CN119536602A

  • Individual video conferencing spaces with shared virtual channels and immersive users

    US11317060B1