OPS framework conference communication system of virtual enhanced digital conference
Through multimodal sensing array and dynamic window allocation unit, the problem of unreasonable gesture occlusion and picture proportions in traditional digital conferences is solved, and the screen content is clearly displayed and intelligent layout adjustment is realized, which improves the interactivity and information communication efficiency of virtual enhanced digital conferences.
Patent Information
- Application Number
- CN202510703466.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-29
AI Technical Summary
In traditional digital conferences, speaker gestures or tools obscures the screen content, affecting the viewing experience, manually adjusting the window layout is inefficient, and cannot intelligently identify the speaker's intentions, resulting in unreasonable screen proportions and inaccurate virtual superposition effects, which affect information transmission.
The multimodal sensing array unit is used to detect occlusions in real time, generate a translucent virtual structure covering the occlusion area, combine it with the dynamic window allocation unit to smoothly scale the window based on behavioral intention, use the Bezier curve to adjust the proportion of the screen and the character screen, and the virtual and real fusion rendering unit generates virtual guide lines and translucent AR layers to enhance interactivity and information communication.
Ensure that the screen content is clear and visible, reduce manual intervention, improve meeting fluency, improve information communication efficiency, enhance immersive prompts, avoid visual interference, and is suitable for complex meeting scenarios and remote collaboration.
Smart Images

Figure CN120263931A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image communication, and more specifically, to an OPS framework conference communication system for virtual enhanced digital conferences. Background Art
[0002] Today, with the rapid development of technology and society, the amount of information that people possess and come into contact with in their daily lives and work is increasing. Therefore, information exchange and communication among people have become more and more frequent and important. Business negotiations, product demonstrations, guest meetings, government order announcements, etc. are all forms of communication between people. Among them, digital conferences use computer, digital, and network technologies to form a network of various systems, integrating technologies such as computers, communications, automatic control, multimedia, images, and audio into a conference automation management system. The system organically connects each independent subsystem such as conference registration, speech, voting, camera, audio, display, and network access into a whole. The central control computer coordinates the work of each subsystem according to the conference agenda, providing the most accurate and timely information and services for various speech conferences, academic reports, and remote conferences. However, in traditional digital conferences, the main speaker usually gives a speech by displaying content on a display screen. However, during the speech, the speaker often uses auxiliary tools such as the arm or a pointing stick for pointing or control operations. This operation method easily causes the arm to block the content on the display screen, affecting the viewing experience of other participants. In addition, the mobility of digital conferences is poor, and the screen layout cannot be automatically switched according to the speaker's gestures or behaviors, resulting in an unreasonable allocation of the screen ratio between the speaker and the display screen content, further reducing the conference efficiency. In view of this, we propose an OPS framework conference communication system for virtual enhanced digital conferences. Summary of the Invention
[0003] The purpose of the present invention is to provide an OPS framework conference communication system for virtual enhanced digital conferences, which significantly improves the interaction naturalness and information transmission efficiency of digital conferences through a closed loop of multi-modal perception, dynamic decision-making, and virtual-real rendering, so as to solve the following problems raised in the above background art: 1. The speaker's actions block the screen content, resulting in unclear information; 2. Manually adjusting the window layout is inefficient, affecting the smoothness of the conference, and when there are multiple screens, the content is displayed too small to be effectively presented; 3. Traditional systems cannot intelligently recognize the speaker's intentions and switch the display mode, and the virtual overlay effect is not accurate enough, which may cause interference or delay, resulting in insufficient accurate prompts when using tools such as pointers, affecting the conveyance of key content.
[0004] To achieve the above object, the present invention provides an OPS framework conference communication system for virtual augmented digital conferences, including a multimodal sensing array unit, a dynamic window allocation unit, and a virtual-real fusion rendering unit; The multimodal sensing array unit is used to collect the behavior data of the speaker and the screen occlusion state; The dynamic window allocation unit is used to extract the feature fusion attention model of the behavior data to predict the focus area and output the behavior intention, and use the Bezier curve to realize the smooth zooming of the conference perspective window. The behavior intention includes the content explanation mode and the free speech mode; When the content explanation mode is detected, the window is smoothly zoomed to prioritize the screen video perspective, and when the virtual-real fusion rendering unit senses that the screen occlusion state is an occluded signal, a semi-transparent virtual feature structure is generated from the occluder feature to cover the screen video perspective, and the virtual feature structure synchronizes the behavior data; When the content explanation mode is detected, the window is smoothly zoomed to prioritize the person video perspective, and the virtual-real fusion rendering unit forms a virtual guiding line according to the speaker's behavior data.
[0005] As a further improvement of this technical solution, the multimodal sensing array unit includes a video stream acquisition module, an action capture module, and a screen occlusion detection module; The video stream acquisition module uses an RGB camera to obtain the depth information and video stream of the speaker, and through the OPS framework rendering resolution, aligns the RGB image with the depth map, combines the color and depth information, and generates a three-dimensional point cloud; The action capture module forms a feature map based on the pose data of the speaker in the three-dimensional point cloud extracted by the feature extractor, detects the human key points in the feature map, and uses a graph model to connect the key points to construct human pose graph data; The screen occlusion detection module is used to identify the screen edge features in the video stream by using an edge detection algorithm, extract the straight edge, fit the screen quadrilateral boundary, calculate the homography matrix according to the detected quadrilateral vertices, map the screen perspective area to a rectangular ROI, and lock the screen features within the screen video perspective area. Within the rectangular ROI, compare the current frame with the background frame to generate a differential binary image, sense the occluder features other than the screen features, and issue an occlusion warning signal.
[0006] As a further improvement of this technical solution, the dynamic window allocation unit includes a behavior feature extraction module, an intention analysis module, and a window switching module; The behavior feature extraction module is used to receive the pose map data of the motion capture module and convert it into a time series. Each time step contains key point data. An LSTM model is used to train the model with an annotated data set. At the input layer, the key point data of each time step is received. At the LSTM layer, time-dependent features are extracted. At the output layer, a classified action is generated. The intention analysis module is used to establish a mapping relationship between the classified action and the behavior intention based on historical data. The classified action of the behavior feature extraction module is input into the mapping relationship, and the behavior intention is output. The window switching module is used to receive the behavior intention output by the intention analysis module, implement Bezier curve smoothing scaling to define the current layout and target layout of the window scaling path, set control points to adjust the curve smoothness, and share window switching, including the following postures: Posture 1, content explanation mode: The screen window proportion is increased to 70%, and the character window is reduced to 30%. Posture 2, free speech mode: The character window is expanded to 80%, and the screen window is reduced to 20%.
[0007] As a further improvement of this technical solution, the output end of the intention analysis module is connected to a sound source localization module. The sound source localization module forms a microphone array to collect digital conference voice communication, calculates the time alignment cost between the voice active interval and the action sequence, finds the minimum cumulative cost path, aligns the voice peak with the start and end times of the action, so that the sound source localization module and the behavior feature extraction module are fused to align the voice active interval and the timing data of the motion capture. If a human body posture map data and the screen interaction action are detected during voice activity, it is weighted and determined as the "content explanation mode". If there is voice activity but no screen interaction action and it exceeds the transition time period, it is weighted and determined as the state in the "free speech mode". As a further improvement of this technical solution, the sound source localization module also includes calculating the sound source direction through the phase difference of the microphone array, locating the sound source coordinates based on the time difference of arrival of the signal, inputting the sound source coordinates and the visual positioning into a Kalman filter, outputting the optimal estimated position, and driving the RGB camera of the video stream acquisition module to track the weighted determined position.
[0008] As a further improvement of this technical solution, the multi-modal sensing array unit also includes a screen window splitting module. The screen window splitting module is used to detect the number of straight edges output by the screen occlusion detection module, calculate the number of screens, divide the number of screen windows of the window switching module to match the number of screens, and make the formed multiple rectangular ROIs respectively correspond to the screen window display.
[0009] As a further improvement of this technical solution, the virtual-real fusion rendering unit includes an occlusion feature analysis module, a virtual structure generation module, and a trajectory modeling module. The occlusion feature analysis module is used to receive the occlusion warning signal from the screen occlusion detection module, segment the occluder in the screen area in real time, and extract the contour features of the occluder; The virtual structure generation module is used to virtually generate a semi-transparent AR layer that matches the contour features of the occluder as a virtual feature structure, receive the human body pose map data provided by the motion capture module, and extract the motion trajectory of the semi-transparent AR layer that is synchronously rendered according to the motion path in the human body pose map data; The trajectory modeling module is used to locate the motion path of the hand joint points in the person window of the window switching module that match the human body pose map data of the motion capture module, generate a light point trajectory to move dynamically on the motion path of the hand joint points, and the color of the light point trajectory changes with time, and the color intensity is associated with the gesture amplitude.
[0010] As a further improvement of this technical solution, the occlusion feature analysis module further includes a contour analysis module. The contour analysis module is used to detect the matching degree between the motion trajectory of the occluder and the human body pose map data output by the motion capture module. If the moving trajectory of the occluder overlaps with the hand trajectory of the speaker, it is determined as an active occlusion and transmitted to the virtual structure generation module to match the semi-transparent AR layer. If the moving trajectory of the occluder does not overlap with the hand trajectory of the speaker, it is determined as a passive occlusion, and the contour area and aspect ratio of the occluder are calculated to filter out the interference of the occluder.
[0011] As a further improvement of this technical solution, the virtual structure generation module further includes an end light module. The end light module is used to receive the contour area and aspect ratio of the occluder output by the contour analysis module. Since the arm occlusion usually presents an irregular block shape, while the pointer is a slender strip shape, it is possible to calculate the aspect ratio and area ratio of the contour to distinguish between the arm and the pointer. When the pointer is output, morphological skeletonization processing is performed on the pointer to extract the center line, mark the end coordinates at the skeleton endpoints, and generate a particle emitter at the end coordinates.
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Aiming at the problem that in a traditional conference system, the gestures or tools of the speaker (such as the arm, pointer) may occlude the screen content, affecting the participants' acquisition of key information, the multi-modal sensing array is used to detect the occluder in real time, generate a semi-transparent virtual structure to cover the occlusion area, and synchronize the gesture trajectory to ensure that the screen content is clearly visible. For example, the arm or pointer is transformed into a semi-transparent arrow to dynamically prompt the pointed content; 2. To address the problem that the existing system relies on manual switching of the screen layout, interrupting the smoothness of the meeting, the dynamic window allocation unit uses Bezier curves to achieve smooth window scaling based on behavioral intentions (such as content explanation mode, free speech mode), automatically adjusts the screen-to-person ratio (such as 70% screen + 30% person), reduces manual intervention. At the same time, the screen window splitting module detects the edges of multiple screens, divides independent windows for separate rendering, ensures clear display of the content on each screen, and avoids image compression caused by distance. 3. To address the problem of virtual guidance delay or interference with real content, the trajectory modeling module dynamically adjusts virtual features in combination with depth information, generates gradient color particles at the end of the pointer according to the moving speed and depth. If it is behind the screen, the brightness is reduced to avoid visual interference. At the same time, the end light module skeletonizes the pointer, generates a dynamic halo at the end point, and sprays particles in combination with the waving direction to enhance the immersive hint of the pointed content. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is the overall structural principle block diagram of the present invention; Figure 2 is the detailed schematic diagram of the overall structure of the present invention; Figure 3 is the meeting communication window demonstration diagram of the intention analysis of the present invention.
[0014] The meanings of the reference numerals in the figures are as follows: 100, multi-modal sensing array unit; 110, video stream acquisition module; 120, motion capture module; 130, screen occlusion detection module; 200, dynamic window allocation unit; 210, behavior feature extraction module; 220, intention analysis module; 230, window switching module; 300, virtual-real fusion rendering unit; 310, occlusion feature analysis module; 320, virtual structure generation module; 330, trajectory modeling module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0016] Please refer to Figures 1 - 3 As shown, this embodiment provides an OPS framework meeting communication system for virtual augmented digital meetings, including a multi-modal sensing array unit 100, a dynamic window allocation unit 200, and a virtual-real fusion rendering unit 300; The multimodal sensing array unit 100 is used to collect the behavior data of the speaker and the screen occlusion state; Specifically, the multimodal sensing array unit 100 includes a video stream acquisition module 110, an action capture module 120, and a screen occlusion detection module 130; The video stream acquisition module 110 uses an RGB camera to obtain the depth information and video stream of the speaker. The RGB camera can capture color images, provide color information of the scene, and at the same time capture depth images. Each pixel represents the distance from the camera. By projecting light with a known pattern and analyzing the deformation to calculate the depth, and rendering the resolution through the OPS framework to align the RGB image with the depth map, combining the color and depth information to generate a three-dimensional point cloud, optimizing the transmission through the OPS framework, adjusting the rendering resolution, balancing quality and bandwidth, ensuring the synchronization of RGB and depth data, using the synchronization mechanism provided by the camera SDK to ensure the synchronization of RGB and depth frames. When rendering the resolution, overlay the RGB and depth data to generate an enhanced video stream, and use the image pyramid technology to adjust the resolution, and then use H.264 encoding to compress the video, and establish a transmission channel through WebRTC for real-time transmission to the cloud for convenient processing in the cloud; The action capture module 120 extracts the feature map formed by the pose data of the speaker in the three-dimensional point cloud based on the feature extractor of VGG-19, and detects the human key points (such as hands, head, shoulders, etc.) in the feature map, and uses a graph model to connect the key points to construct human pose graph data; The screen occlusion detection module 130 is used to identify the screen edge features in the video stream by using an edge detection algorithm, extract the straight edges, fit the screen quadrilateral boundary, calculate the homography matrix according to the detected quadrilateral vertices, map the screen viewing area to a rectangular ROI, and lock the screen features within the screen video viewing area. Within the rectangular ROI, compare the current frame with the background frame (the screen image without occlusion) to generate a differential binary image, perceive the occlusion feature other than the screen feature, and send out an occlusion warning signal, which is beneficial to reminding the speaker to avoid occluding the screen content unconsciously and preventing it from affecting other meeting participants from viewing the content on the screen.
[0017] It should be noted that the dynamic window allocation unit 200 is used to extract the behavior data feature fusion attention model to predict the focus area and output the behavior intention, and uses the Bezier curve to realize the smooth zooming of the meeting perspective window. This is beneficial to dynamically adjust the screen layout based on the speaker's behavior characteristics, making the screen ratio distribution of the speaker and the display content more reasonable, ensuring the clear display of the speaker and the display content, avoiding the blurred display of content or information loss caused by unreasonable screen ratio, improving the meeting efficiency. At the same time, it realizes the intelligent switching of the speaker's control instructions and the meeting scene, reduces manual operations, and improves the user experience. The behavior intentions include the content explanation mode (the speaker focuses on the display content, frequently points to and gazes at the screen) and the free speech mode (the speaker has no interaction with the display within a certain time period, the speaker is far away from the screen and the gesture amplitude increases); When the content explanation mode is detected, the window smoothly zooms to prioritize the screen video perspective. When the virtual-real fusion rendering unit 300 senses that the screen occlusion state is an occluded signal, it generates a semi-transparent virtual feature structure of the occlusion feature (which may be the speaker's arm or pointing tool) to cover the screen video perspective, making the virtual feature structure synchronize with the behavior data. This is beneficial to convert the speaker's pointing action into a transparent or semi-transparent virtual feature structure, avoiding the arm from occluding the display content and improving the clarity of content display; When the content explanation mode is detected, the window smoothly zooms to prioritize the person video perspective, and the virtual-real fusion rendering unit 300 forms a virtual guiding line according to the speaker's behavior data, which is beneficial to virtually demonstrate the speaker's behavior actions and enables the participants in the meeting to more vividly observe the speaker's speech.
[0018] Based on the above content, the specific working principle is detailed as follows: First of all, the dynamic window allocation unit 200 includes a behavior feature extraction module 210, an intention analysis module 220, and a window switching module 230; The behavior feature extraction module 210 is used to receive the pose graph data of the motion capture module 120 and convert it into a time series. Each time step contains key point data. The LSTM model is trained using the labeled data set (collecting video data containing different speech scenarios and manually labeling the action categories and timestamps). At the input layer, it receives the key point data of each time step, extracts the time-dependent features at the LSTM layer, and generates classification actions (such as pointing and gazing at the screen) at the output layer. Specifically, the key point data of consecutive frames are arranged in chronological order to form a time series of key point data and input it into the LSTM model to output classification actions. Note that the key point coordinates are normalized (the mean is set to zero and the variance is normalized) to avoid the influence of scale differences on model training; The intention analysis module 220 is used to establish a mapping relationship between classification actions and behavioral intentions based on historical data. The classification actions of the input behavior feature extraction module 210 are input into the mapping relationship, and the behavioral intentions are output. If it is detected within a time period that the speaker does not point to or gaze at the screen, it means that the speaker continues to speak without interacting with the display screen. On the contrary, if the speaker points to or gazes at the screen within the time period, it means that the speaker focuses on the content of the display screen and frequently points to and gazes at the screen for conference communication. Combining the deep learning model of the LSTM model and time series analysis can improve the accuracy of action recognition, further classify the state of the speaker in real time more accurately, and enhance the robustness of the system in complex environments; Specifically, when establishing the mapping relationship between classification actions and behavioral intentions, the classification action sequence is input into the Bi-LSTM model to capture long-term dependence relationships and output pattern probabilities. The rule engine and the model prediction results (weighted voting) are fused, and user feedback data (such as manually correcting the mode classification results) is collected to fine-tune the model parameters to enable predicting behavioral intentions based on classification actions. Note that the prediction results have a certain probability, and it is easy for those skilled in the art to think of real-time updates to ensure the accuracy of the mapping relationship; The window switching module 230 is used to receive the behavioral intentions output by the intention analysis module 220 to implement Bezier curve smoothing scaling, define the current layout and target layout of the window scaling path, set control points to adjust the curve smoothness, and share window switching, including the following postures: Posture 1, content explanation mode: The screen window accounts for 70% and the person window is reduced to 30%; Posture 2, free speech mode: The person window is expanded to 80% and the screen window is reduced to 20%; Therefore, according to the behavioral intentions, dynamically adjusting the window layout and achieving smooth transitions can realize efficient and natural adjustment of the conference screen layout.
[0019] To further significantly improve the accuracy and robustness of the intent analysis module 220 in behavioral intent recognition, connect the output end of the intent analysis module 220 to the sound source localization module. The sound source localization module forms a microphone array to collect digital conference voice communications. On the one hand, calculate the time alignment cost between the voice active interval and the action sequence, find the minimum cumulative cost path, align the voice peak with the start and end times of the action, solve the time offset problem of voice and action data, and enable the sound source localization module and the behavioral feature extraction module 210 to fuse and align the timing data of the voice active interval and action capture. If human body pose map data and screen interaction actions are detected during voice activity, it is weighted and determined as the "content explanation mode". If there is voice activity but no screen interaction action, and it exceeds the transition time period (to avoid switching modes during the speaker's transition stage of explanation and ensure that the situation where the speaker actually interacts with the screen but has no interaction action will not last longer than the transition time period), it is weighted and determined as the state in the "free speech mode", realizing an anti-misjudgment mechanism; On the other hand, the sound source localization module also includes calculating the sound source direction through the phase difference of the microphone array and locating the sound source coordinates based on the time difference of arrival of signals, which is beneficial to accurately locate the speaker's sound source coordinates through the microphone array, supplement the deficiencies of visual localization, input the sound source coordinates and visual localization into the Kalman filter, and output the optimal estimated position, driving the RGB camera of the video stream acquisition module 110 to track the weighted determined position. Specifically, according to the optimal estimated position coordinates output by the Kalman filter, control the rotation angle and zoom of the pan-tilt head, and use a PID controller to adjust the movement speed of the pan-tilt head to avoid jitter, so as to keep the screen and the speaker centered in the content explanation mode and keep the speaker centered in the free speech mode. The camera control and layout animation improve the meeting fluency and are applicable to complex meeting scenarios, especially in the fields of remote collaboration, educational lectures, etc., with broad application prospects.
[0020] Furthermore, considering that during conference communication, the speaker will use multiple screens to cooperate and conduct digital conferences simultaneously. If the screen occlusion detection module 130 uses the edges of two relatively distant screens as the boundary lines when detecting the screen boundary lines, the content on multiple screens will be displayed small in the window corresponding to the screen video perspective, affecting the conference effect. Therefore, in a multi-screen conference scenario, the number of screens is dynamically detected and the window is split to ensure that the content of each screen is independently displayed, improving the clarity and readability of the picture. The multi-modal sensing array unit 100 further includes a screen window splitting module. The screen window splitting module is used to detect the number of straight edges output by the screen occlusion detection module 130, calculate the number of screens, divide the number of screen windows of the window switching module 230 to match the number of screens, and make the formed multiple rectangular ROIs correspond to the screen windows for display, which is beneficial for giving a speech on multiple screens. Each screen can be divided into a window, and the content of each screen is rendered separately, avoiding the text / images being too small due to combined display, realizing separate execution of digital conferences for each screen, and avoiding the small video capture picture caused by the distance between two screens, which affects the content display effect.
[0021] Then, the virtual-real fusion rendering unit 300 includes an occlusion feature analysis module 310, a virtual structure generation module 320, and a trajectory modeling module 330; The occlusion feature analysis module 310 is used to receive the occlusion warning signal of the screen occlusion detection module 130, real-time segment the occluders (arms / pointers) in the screen area, and extract the contour features of the occluders. Specifically, it receives the differential binary image and the background subtraction result output by the screen occlusion detection module 130, and obtains the color image and depth image captured by the RGB camera. When segmenting the occluder: Perform morphological processing (such as erosion and dilation) on the differential binary image to remove noise and smooth the boundary, use the depth image to extract the depth range of the occluder, separate the foreground occluder from the background content, and generate an occluder mask; When extracting the contour features, use an edge detection algorithm (such as the Canny operator) to extract the contour of the occluder, calculate the key geometric features of the contour (such as area, perimeter, major axis direction, etc.), and output the mask and contour features of the occluder; The virtual structure generation module 320 is used to virtually generate a semi-transparent AR layer that matches the contour features of the occluder as a virtual feature structure. It receives the human body posture map data provided by the motion capture module 120, extracts the motion trajectory of the semi-transparent AR layer that is synchronously rendered with the motion path in the human body posture map data. When generating the AR layer, a semi-transparent AR layer that matches the contour features of the occluder is generated, and the color, transparency, and shape parameters of the AR layer are set. The motion path of the hand joint points is converted into the motion trajectory of the AR layer, and trajectory lines or arrows are dynamically drawn on the AR layer, and the color changes gradually with the gesture amplitude, which is beneficial to improving the clarity and interactivity of the virtual augmented digital conference system, especially suitable for complex conference scenarios and remote collaboration environments, without affecting the real-time viewing of the screen content by conference participants and guiding the speaker throughout the speech; The trajectory modeling module 330 is used to locate the motion path of the hand joint points that match the human body posture map data of the motion capture module 120 in the person window of the window switching module 230. When the motion path of the hand joint points is dynamically generated, the light spot trajectory moves. The color of the light spot trajectory changes with time, and the color intensity is associated with the gesture amplitude. When the gesture swings to the right, a semi-transparent arrow is generated 0.2 seconds in advance to avoid visual delay, which is beneficial for the speaker to more dynamically display the gesture when demonstrating the conference content with gestures.
[0022] Considering that there may be small area interferences other than the speaker (such as mosquitoes in front of the camera, etc.), if virtual augmentation is still performed at this time, it is meaningless. Therefore, the occlusion feature analysis module 310 also includes a contour analysis module. The contour analysis module is used to detect the matching degree between the motion trajectory of the occluder and the human body posture map data output by the motion capture module 120. If the moving trajectory of the occluder overlaps with the hand trajectory of the speaker, it is determined as active occlusion (such as occlusion caused by teaching aids, arms, etc.) and transmitted to the virtual structure generation module 320 to match the semi-transparent AR layer. If the moving trajectory of the occluder does not overlap with the hand trajectory of the speaker, it is determined as passive occlusion (occlusion by mosquitoes, environmental impurities, etc.), and the contour area and aspect ratio of the occluder are calculated to filter out the occluder interference, effectively identifying and filtering out small area interferences such as mosquitoes to avoid unnecessary AR augmentation.
[0023] Furthermore, when the types of occluders are different, the virtual feature structures generated by the virtual structure generation module 320 are all of the same type, which instead affects the speaker's pointing. For example, when the speaker gives a speech with a pointer in hand, the pointing of the tip of the pointer indicates the content that needs to be focused on, but the meeting participants cannot accurately distinguish it. Therefore, the virtual structure generation module 320 further includes a tip lighting module. The tip lighting module is used to receive the occluder contour area and aspect ratio output by the contour analysis module. Based on the fact that the arm occlusion usually presents an irregular block shape, while the pointer is a slender strip shape, it calculates the aspect ratio and area ratio of the contour to distinguish between the arm and the pointer. If the aspect ratio > 5:1 and the area < 200 pixel² → it is determined as a pointer. If the aspect ratio < 3:1 and the area > 500 pixel² → it is determined as an arm. When outputting a pointer, perform morphological skeletonization on the pointer, extract the center line, mark the end coordinates at the skeleton endpoints, and generate a particle emitter at the end coordinates. The parameter settings are as follows: Number of particles: 50 - 100 per second (adaptive according to the moving speed); Color gradient: transitioning from bright yellow at the center (RGB: 255, 255, 100) to semi - transparent red at the periphery (RGB: 255, 0, 0, α = 0.3); Motion mode: eject along the waving direction of the pointer, speed = the speed of the pointer tip × 0.8; And compare the depth value of the pointer tip with the depth of the screen content. If the pointer is in front of the screen, render the halo normally. If the pointer is behind the screen (such as pointing outside the screen edge), reduce the halo brightness (brightness coefficient = 0.3) to avoid visual interference. Use the AlphaBlending formula to blend the virtual halo with the real - world image, and render a golden - colored halo with a diameter of 30px at the endpoint. The trailing length is 50px as the pointer moves. When the pointer points to the resistor symbol on the circuit diagram, the halo is semi - transparently superimposed, and the symbol content is clearly visible, significantly enhancing the immersion and information transmission efficiency of remote collaboration.
[0024] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above - mentioned embodiments. The above - mentioned embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. An OPS framework conference communication system for virtual enhanced digital conferences, characterized in that: It includes a multi-modal sensing array unit (100), a dynamic window allocation unit (200), and a virtual-real fusion rendering unit (300); The multi-modal sensing array unit (100) is used to collect the presenter's behavior data and the screen occlusion state; The dynamic window allocation unit (200) is used to extract the behavior data features, fuse the attention model to predict the focus area, output the behavior intention, and use the Bezier curve to realize the smooth zooming of the meeting perspective window. The behavior intention includes the content explanation mode and the free speech mode; When the content explanation mode is detected, the window is smoothly zoomed to prioritize the screen video perspective, and when the virtual-real fusion rendering unit (300) senses that the screen occlusion state is an occluded signal, a semi-transparent virtual feature structure is generated from the occluder feature to cover the screen video perspective, so that the virtual feature structure synchronizes with the behavior data; When the content explanation mode is detected, the window is smoothly zoomed to prioritize the person video perspective, and the virtual-real fusion rendering unit (300) forms a virtual guiding line according to the presenter's behavior data.
2. The OPS framework conference communication system for virtual augmented digital conferences according to claim 1, wherein: The multi-modal sensing array unit (100) includes a video stream acquisition module (110), an action capture module (120), and a screen occlusion detection module (130); The video stream acquisition module (110) uses an RGB camera to obtain the depth information and video stream of the presenter, renders the resolution through the OPS framework, aligns the RGB image with the depth map, combines the color and depth information, and generates a three-dimensional point cloud; The action capture module (120) forms a feature map based on the pose data of the presenter in the three-dimensional point cloud extracted by the feature extractor, detects the human key points in the feature map, and uses a graph model to connect the key points to construct the human pose graph data; The screen occlusion detection module (130) is used to identify the screen edge features in the video stream by using an edge detection algorithm, extract the straight edges, fit the screen quadrilateral boundary, calculate the homography matrix according to the detected quadrilateral vertices, map the screen perspective area to a rectangular ROI, lock the screen features within the screen video perspective area, compare the current frame with the background frame within the rectangular ROI, generate a differential binary map, sense the occluder features other than the screen features, and send out an occlusion warning signal.
3. The OPS framework conference communication system for virtual enhanced digital conferences according to claim 2, characterized in that: The dynamic window allocation unit (200) includes a behavior feature extraction module (210), an intention analysis module (220), and a window switching module (230); The behavior feature extraction module (210) is used to receive the pose graph data of the action capture module (120) and convert it into a time series. Each time step contains key point data. The LSTM model is used to train the model with the labeled data set. The key point data of each time step is received at the input layer, the time-dependent features are extracted at the LSTM layer, and the classification action is generated at the output layer; The intention analysis module (220) is used to establish the mapping relationship between the classification action and the behavior intention based on historical data, input the classification action of the behavior feature extraction module (210) into the mapping relationship, and output the behavior intention; The window switching module (230) is used to receive the current layout and target layout of the Bessel curve smooth zoom definition window zoom path implemented by the behavior intention output by the intention analysis module (220), set control points to adjust the curve smoothness, and share window switching, including the following postures: Posture 1, content explanation mode: The screen window accounts for 70% and the person window is reduced to 30%; Posture 2, free speech mode: The person window is expanded to 80% and the screen window is reduced to 20%.
4. The OPS framework conference communication system for virtual enhanced digital conferences according to claim 3, characterized in that: The output end of the intention analysis module (220) is connected to the sound source localization module. The sound source localization module forms a microphone array to collect digital conference voice communication, calculates the time alignment cost between the voice active interval and the action sequence, finds the minimum cumulative cost path, aligns the voice peak with the start and end times of the action, and enables the sound source localization module and the behavior feature extraction module (210) to fuse and align the timing data of the voice active interval and action capture. If human pose map data and screen interaction actions are detected during voice activity, it is weighted and determined as the "content explanation mode". If there is voice activity but no screen interaction action and it exceeds the transition time period, it is weighted and determined as the state in the "free speech mode".
5. The OPS framework conference communication system for virtual enhanced digital conferences according to claim 4, characterized in that: The sound source localization module also includes calculating the sound source direction through the phase difference of the microphone array, positioning the sound source coordinates based on the time difference of arrival of signals, inputting the sound source coordinates and visual positioning into the Kalman filter, outputting the optimal estimated position, and driving the RGB camera of the video stream acquisition module (110) to track the weighted determined position.
6. The OPS framework conference communication system for virtual enhanced digital conferences according to claim 4, characterized in that: The multi-modal sensing array unit (100) also includes a screen window splitting module. The screen window splitting module is used to detect the number of straight edges output by the screen occlusion detection module (130), calculate the number of screens, divide the number of screen windows of the window switching module (230) to match the number of screens, and enable the formed multiple rectangular ROIs to correspond to the screen window display respectively.
7. The OPS framework conference communication system for virtual enhanced digital conferences according to claim 6, characterized in that: The virtual-real fusion rendering unit (300) includes an occlusion feature analysis module (310), a virtual structure generation module (320), and a trajectory modeling module (330); The occlusion feature analysis module (310) is used to receive the occlusion warning signal of the screen occlusion detection module (130), real-time segment the occluder in the screen area, and extract the contour features of the occluder; The virtual structure generation module (320) is used to virtually generate a semi-transparent AR layer matching the contour features of the occluder as a virtual feature structure, receive the human pose map data provided by the action capture module (120), and extract the motion path in the human pose map data to synchronously render the motion trajectory of the semi-transparent AR layer; The trajectory modeling module (330) is used to locate the motion path of the hand joint points in the person window of the window switching module (230) that match the human pose map data of the action capture module (120), dynamically generate a light point trajectory to move in the motion path of the hand joint points, and the color of the light point trajectory changes with time, and the color intensity is associated with the gesture amplitude.
8. The OPS framework conference communication system for virtual enhanced digital conferences according to claim 7, characterized in that: The occlusion feature analysis module (310) further includes a contour analysis module, which is used to detect the matching degree between the movement trajectory of the occluder and the human body pose map data output by the motion capture module (120). If the movement trajectory of the occluder overlaps with the hand trajectory of the speaker, it is determined as active occlusion and transmitted to the virtual structure generation module (320) to match the semi-transparent AR layer. If the movement trajectory of the occluder does not overlap with the hand trajectory of the speaker, it is determined as passive occlusion, and the contour area and aspect ratio of the occluder are calculated to filter out the interference of the occluder.
9. The OPS framework conference communication system for virtual enhanced digital conferences according to claim 8, characterized in that: The virtual structure generation module (320) further includes an end lighting module, which is used to receive the contour area and aspect ratio of the occluder output by the contour analysis module. Since the arm occlusion usually presents an irregular block shape, while the pointer is a slender strip shape, it is possible to distinguish the arm and the pointer by calculating the aspect ratio and area ratio of the contour. When outputting the pointer, morphological skeletonization processing is performed on the pointer to extract the center line, and the end coordinates are marked at the skeleton endpoints, and a particle emitter is generated at the end coordinates.
Citation Information
Patent Citations
Video acquisition method and device
CN114143496A
Bullet screen presentation method and device, equipment and storage medium
CN117998154A
Meeting interaction method and system based on meta universe
CN119536602A
Smart interactive white board apparatus using Wi-Fi location tracking technology and its control method
KR102590257B1
Individual video conferencing spaces with shared virtual channels and immersive users
US11317060B1
Cited By
Intelligent picture control method and related device
CN121547642A