Video conference shared picture processing system, method, device and equipment
By using smart glasses to collect gaze data in video conferencing and marking gaze hotspots in the shared screen, the problem of participants having difficulty finding the speaker's content is solved, resulting in a more efficient video conferencing experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI QIANWEN ZHILIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
- Filing Date
- 2026-01-06
- Publication Date
- 2026-04-28
AI Technical Summary
Participants may find it difficult to locate the speaker's current content in the shared screen during a video conference, leading to missed information.
The system receives user gaze data collected by the first smart glasses through the first client, obtains gaze hotspot areas, marks them in the shared screen, and sends the marked screen to the second client through the server.
The system displays the speaker's gaze hotspots on the shared screen in real time, enabling participants to quickly locate the content being presented and improving video conferencing efficiency and user experience.
Smart Images

Figure CN121940504A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video conferencing technology, specifically to video conferencing shared image processing systems, methods and apparatus, and electronic devices. Background Technology
[0002] The person initiating the video conference, acting as the speaker, frequently uses the screen sharing function. During the sharing process, the speaker will review the content while explaining it to the participants, thereby improving the efficiency of the online meeting.
[0003] However, the applicant of this application has found that the existing technology has at least the following problems: when participants are watching, their attention may be distracted at some point, or the speaker may jump around in their explanations, making it difficult for them to find the content the speaker is currently explaining in the shared screen, resulting in missing a lot of key information. Therefore, how to help video conference participants find the content the speaker is explaining in the shared screen, thereby improving the efficiency of video conferences, is an urgent problem that needs to be researched and solved. Summary of the Invention
[0004] This application provides a video conferencing screen sharing processing system to solve the problem in existing technologies where participants have difficulty finding the speaker's content in the shared screen. This application also provides a video conferencing screen sharing processing method and apparatus, as well as electronic equipment.
[0005] This application provides a video conferencing screen sharing processing system, including: A first client is used to display the shared screen of a video conference; receive gaze data of a first user viewing the shared screen through first smart glasses; obtain the first user's gaze hotspot area based on the first user's gaze data; mark the first user's gaze hotspot area in the shared screen; and send the shared screen marked with the first user's gaze hotspot area to a second client through a server. The first smart glasses are used to collect and transmit the first user's gaze data.
[0006] Optionally, the shared screen includes image content, and the first user's gaze hotspot area includes partial image content.
[0007] Optionally, the first client can also be used to annotate the explanatory trajectory of multiple image local contents.
[0008] Optionally, the first client is also used to annotate the currently explained local content of the image with a first display attribute, and to annotate the explained local content of the image with a second display attribute.
[0009] Optionally, the first client is also used to store the correspondence between user gaze data and gaze hotspot areas; and to switch and annotate the local content of the multiple images according to the correspondence and the first user gaze data.
[0010] Optionally, the image content includes video footage, and the partial image content includes partial content of the video footage; The first client is also configured to: obtain a request to enable the annotation mode; respond to the request to enable the annotation mode and annotate the first user's visual hotspot area in the shared screen; and obtain a request to disable the annotation mode; respond to the request to disable the annotation mode and stop annotating the first user's visual hotspot area.
[0011] Optionally, the first smart glasses are also used to obtain a labeling mode enable command and send a labeling mode enable request; and to obtain a labeling mode disable command and send a labeling mode disable request.
[0012] Optionally, the first smart glasses are also used to display annotation mode operation options; and receive annotation mode enable or disable commands through the operation options.
[0013] Optionally, the first smart glasses include a labeling mode operation button; through the operation option, it can receive a labeling mode enable command or a labeling mode disable command.
[0014] Optionally, the first client is further configured to obtain a annotation mode start instruction and generate the annotation mode start request; and to obtain a annotation mode end instruction and generate the annotation mode end request.
[0015] Optionally, the first client is also configured to display multiple calibration points in the shared screen; receive calibration point gaze data of viewing the calibration points through the first smart glasses; and specifically, to obtain the first user gaze hotspot area based on the first user gaze data and the multiple calibration point gaze data of the multiple calibration points; The first smart glasses are also used to acquire and send the calibration point gaze data.
[0016] Optionally, the first smart glasses are also used to collect user gaze data on multiple calibration points displayed in the shared screen, and send the user gaze data; The first client is also configured to display the plurality of calibration points in the shared screen; receive the user's line-of-sight data; receive a calibration point line-of-sight determination instruction input by the user, and use the current user's line-of-sight data as the calibration point line-of-sight data of the corresponding calibration point; and specifically, to obtain the first user's line-of-sight hotspot area based on the first user's line-of-sight data and the plurality of calibration point line-of-sight data.
[0017] Optionally, the first client is also used to display prompts indicating that the calibration point should be viewed directly.
[0018] Optionally, the first smart glasses are also used to obtain the annotation mode entry instruction and send the annotation mode entry request; The first client is also configured to receive the annotation mode entry request; and specifically, to respond to the annotation mode entry request and obtain the first user's gaze hotspot area based on the first user's gaze data.
[0019] Optionally, the first client is further configured to obtain an instruction to invite the first smart glasses to enter the annotation mode, send an invitation to enter the annotation mode request; and receive the annotation mode entry request; and specifically, in response to the annotation mode entry request, obtain the first user's gaze hotspot area based on the first user's gaze data. The first smart glasses are also used to receive the invitation to enter the annotation mode request and send the annotation mode entry request.
[0020] Optionally, the first client is also used to obtain prompt information for prompting entry into the annotation mode, and send the prompt information to the first smart glasses; The first smart glasses are also used to receive and display the prompt information.
[0021] Optionally, the first smart glasses and the first client transmit data through the server; The first smart glasses are also used to connect to the Internet via wireless access points or smart terminals; obtain the access password for video conferences, and send access requests to the server; The server is used to determine whether the first smart glasses are devices allowed to enter the meeting based on the first smart glasses identifier carried in the request; if the determination result is yes, the first smart glasses are used as the gaze tracking device of the first client.
[0022] Optionally, the first smart glasses and the first client can transmit data via short-range communication.
[0023] Optional, also includes: The second client is used to receive the shared screen and display the shared screen; receive second user gaze data of the second user viewing the shared screen through second smart glasses; obtain location data of the second gaze hotspot area based on the second user gaze data; and send the location data to the first client through the server. The second smart glasses are used to collect and send the second user's gaze data. The first client is specifically used to mark the first line-of-sight hotspot area in the shared screen; the first client is also used to receive the location data; and mark the second line-of-sight hotspot area in the shared screen according to the location data.
[0024] This application provides a method for processing shared video feeds in video conferencing, including: Displays the shared screen of the video conference; Receives gaze data from a first user viewing the shared screen through first smart glasses; Based on the first user's gaze data, obtain the first user's gaze hotspot area; The first user's line-of-sight hotspot area is marked in the shared screen; The server sends the shared screen, which marks the hotspot area in the first user's line of sight, to the second client.
[0025] Optional, also includes: The calibration points are displayed in the shared screen; Receive calibration point line-of-sight data when viewing the calibration point through the first smart glasses; The step of obtaining the first user gaze hotspot area based on the first user gaze data includes: Based on the first user's line-of-sight data and the calibration point's line-of-sight data, the first user's line-of-sight hotspot area is obtained.
[0026] Optional, also includes: Receive the annotation mode entry request sent by the first smart glasses; The step of obtaining the first user gaze hotspot area based on the first user gaze data includes: In response to the annotation mode entry request, the first user's gaze hotspot area is obtained based on the first user's gaze data.
[0027] This application provides a video conferencing image sharing processing method for a first smart glasses, comprising: Collect gaze data of the first user who is viewing the shared video conference feed displayed on the first client. Send the first user's gaze data.
[0028] Optional, also includes: Collect calibration point line-of-sight data for viewing calibration points in the shared screen; Send the line-of-sight data for the calibration point.
[0029] This application provides a video conferencing screen sharing processing device, including: The shared screen display unit is used to display the shared screen of the video conference; The first user gaze data receiving unit is used to receive first user gaze data when viewing the shared screen through the first smart glasses; The first user gaze hotspot area determination unit is used to obtain the first user gaze hotspot area based on the first user gaze data. The first user gaze hotspot area marking unit is used to mark the first user gaze hotspot area in the shared screen; The shared screen sending unit is used to send the shared screen marked with the hot spot area in the first user's line of sight to the second client through the server.
[0030] This application provides a video conferencing screen sharing processing device, including: The first user gaze data acquisition unit is used to collect the gaze data of the first user who is watching the shared video conference screen displayed by the first client. The first user gaze data sending unit is used to send the first user gaze data.
[0031] This application provides an electronic device, including: Processor; and A memory for storing a program for implementing the method described in any of the preceding methods, wherein the device is powered on and the program of the method is executed by the processor.
[0032] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the various methods described above.
[0033] This application also provides a computer program product including instructions that, when run on a computer, cause the computer to perform the various methods described above.
[0034] Compared with the prior art, this application has the following advantages: The video conferencing shared screen processing system provided in this application embodiment displays the shared screen of a video conference through a first client; receives first user gaze data provided by first smart glasses viewing the shared screen through the first smart glasses; obtains first user gaze hotspot areas based on the first user gaze data; marks the first user gaze hotspot areas in the shared screen; and sends the shared screen marked with the first user gaze hotspot areas to a second client through a server. The first smart glasses collect the first user gaze data and provide it to the first client; the second client receives the shared screen marked with the first user gaze hotspot areas and displays the shared screen marked with the first user gaze hotspot areas. This processing method allows the speaker's gaze hotspot areas to be drawn on the shared screen in real time, so participants can easily find the content being explained by the speaker in the shared screen at any time; therefore, it can effectively improve video conferencing efficiency and thus enhance user experience. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of device interaction of an embodiment of the video conferencing sharing screen processing system provided in this application; Figure 2 This is another device interaction diagram of an embodiment of the video conferencing screen sharing processing method provided in this application; Figure 3 This is a flowchart illustrating an embodiment of the video conferencing screen sharing processing method provided in this application. Detailed Implementation
[0036] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.
[0037] This application provides a video conferencing sharing screen processing system, method, and apparatus, as well as an electronic device. The various solutions are described in detail below in each embodiment.
[0038] First Embodiment Please refer to Figure 1 This is a schematic diagram of device interaction for the video conferencing sharing and processing system of this application. In this embodiment, the system may include: a first client and a first smart glasses.
[0039] The first client is used to display the shared screen of the video conference; receive the gaze data of a first user viewing the shared screen through a first smart glasses; obtain the first user gaze hotspot area based on the first user gaze data; mark the first user gaze hotspot area in the shared screen; and send the shared screen marked with the first user gaze hotspot area to the second client through the server; the first smart glasses are used to collect the first user gaze data and send the first user gaze data.
[0040] The meeting initiator (first user) initiates a video conference through a first client, enables screen sharing on the first client, and opens the file to be shared. The first client then sends the shared screen to a second client via the meeting server. Participants (second users) view the shared screen through their second clients. The file content shared via screen sharing can include at least one of text, images, and video. While sharing the file content, the first user simultaneously reviews the file and explains it to the second user, thus improving the efficiency of the online meeting.
[0041] In this embodiment, the first client corresponds to the first smart glasses. The first smart glasses can collect the first user's gaze on the shared screen in real time through an image acquisition device, thereby achieving accurate tracking of the first user's gaze and sending the first user's gaze data to the first client. Accordingly, based on the first user's gaze data, the first client can determine which local area in the shared screen the gaze falls on, designate that local area as the first user's gaze hotspot area, and mark the hotspot area in the shared screen, such as by drawing a bounding box around the hotspot area. In this way, even if the second user's attention is distracted at some point or the speaker jumps around in their explanation, the second user can still quickly see the content the speaker is currently explaining, avoiding missing key information, thereby further improving the efficiency of video conferencing.
[0042] In one example, the shared screen includes image content, and the first user explains this content to the second user while referring to it. In this case, the first user's focal point can be a specific area of the image. The image content can be a picture or a video. For example, if a speech recognition model diagram based on a complex neural network is displayed on the shared screen, the presenter can explain in detail how each module in the diagram works. In this case, the specific area of the image could be a particular module or data transmitted between modules. Another example is a video played on the shared screen, which is watched together by the presenter and participants. The presenter can explain specific areas of the video at any time while watching.
[0043] In one example, the image content includes a video frame, and the partial image content includes a partial content of the video frame. The first client is further configured to: acquire a request to enable annotation mode; respond to the request to enable annotation mode by annotating the first user's gaze hotspot area in the shared screen; acquire a request to disable annotation mode; and respond to the request to disable annotation mode by stopping the annotation of the first user's gaze hotspot area. This approach allows the presenter to enable and disable annotation of the user's gaze hotspot area in the video frame at any time, avoiding meaningless annotations caused by continuously annotating the user's gaze hotspot area throughout the video conference; therefore, it effectively improves the user experience.
[0044] In one example, the first smart glasses are also used to obtain a labeling mode enable command and send a labeling mode enable request; and to obtain a labeling mode disable command and send a labeling mode disable request. This processing method allows the labeling mode to be enabled and disabled at any time via the smart glasses; therefore, it effectively improves the convenience of user operation.
[0045] In specific implementation, the first smart glasses are also used to display annotation mode operation options; through these operation options, they receive annotation mode enable or disable commands. For example, when annotation mode is not enabled, the annotation mode enable operation option is displayed; when the user clicks the annotation mode enable operation option, the smart glasses receive the annotation mode enable command; when annotation mode is enabled, the annotation mode disable operation option is displayed; when the user clicks the annotation mode disable operation option, the smart glasses receive the annotation mode disable command.
[0046] In specific implementation, the first smart glasses are also used to display the annotation mode enable operation option and the annotation mode disable operation option; receive the annotation mode enable command through the enable operation option; and receive the annotation mode disable command through the disable operation option.
[0047] In specific implementation, the first smart glasses includes a labeling mode operation button; through the operation option, it receives a labeling mode enable command or a labeling mode disable command. For example, when the labeling mode is not enabled, pressing the labeling mode operation button receives a labeling mode enable command; when the labeling mode is enabled, pressing the labeling mode operation button receives a labeling mode disable command.
[0048] In specific implementation, the first smart glasses include a labeling mode on / off operation button and a labeling mode off operation button; the on operation button receives a labeling mode on / off command; the off operation button receives a labeling mode off command.
[0049] In another example, the first client is also used to obtain a annotation mode enable command and generate an annotation mode enable request; and to obtain a annotation mode disable command and generate an annotation mode disable request. This approach allows the annotation mode to be enabled and disabled at any time via the first client, thus effectively improving user convenience.
[0050] In one example, the first client is also used to annotate the narration trajectory of multiple image segments. For instance, the narration trajectory is the input data of modules D, F, A, and C, annotated sequentially in the narration order in the speech recognition model diagram above. This processing method allows the narration trajectory of the image to be displayed, helping the second user better understand the image content, such as the working principle of the speech recognition model; therefore, it can effectively improve the efficiency of video conferencing.
[0051] In one example, the first client is also used to annotate the currently explained image portion using a first display attribute, and to annotate other image portions using a second display attribute. Display attributes can include the shape and color of the bounding box, line style, and thickness, etc. The first and second display attributes can have different values for certain attributes; for example, both the first and second display attributes may include bounding box color, with the first attribute using red and the second using green. This approach effectively distinguishes the currently explained image portion from other image portions that have already been explained; therefore, it effectively improves the differentiation between the two types of image portions, thereby enhancing the user experience.
[0052] In one example, the first client is also used to store the correspondence between user gaze data and gaze hotspot areas; based on the correspondence and the first user gaze data, the multiple image local contents are switched and labeled. Specifically, the correspondence between the first user gaze data and gaze hotspot areas can be stored. When a change in the first user gaze data is detected, the hotspot area corresponding to the current gaze is obtained based on the correspondence, and the labeling of that hotspot area is directly switched. This processing method allows for direct switching to previously determined image local contents based on the first user gaze data when repeatedly switching between multiple contents in a shared screen, avoiding the need to calculate gaze hotspot areas each time; therefore, it effectively improves the switching efficiency of gaze hotspot areas.
[0053] In one example, the first client is also used to display the plurality of calibration points in the shared screen; receive calibration point gaze data viewed through the first smart glasses; and specifically, to obtain the first user gaze hotspot area based on the first user gaze data and the plurality of calibration point gaze data; the first smart glasses are also used to obtain the calibration point gaze data and send the calibration point gaze data. In specific implementation, the initial angle of the pupil relative to the five calibration points (center point, upper left corner, upper right corner, lower left corner, and lower right corner) can be confirmed through gaze angle calibration and the real-time pupil tracking capability of the glasses; after angle calibration, the offset is subsequently calculated based on the real-time pupil angle and the calibration angles of the five points, and finally the area that the user is looking at on the screen is drawn on the screen. This processing method allows for the acquisition of calibration point gaze data for each calibration point displayed in the shared screen through the first smart glasses. The calibration point gaze data is then determined using the first smart glasses. Based on multiple calibration point gaze data and the first user's video data, the first user's gaze hotspot area can be more accurately identified. Even if the first user's gaze is only focused on a very small area in the shared screen, that small area can still be accurately identified. Therefore, the accuracy of hotspot area labeling can be effectively improved, further enhancing the efficiency of video conferencing.
[0054] In another example, the first smart glasses are also used to collect user gaze data on multiple calibration points displayed in the shared screen and send the user gaze data; the first client is also used to display the multiple calibration points in the shared screen; receive the user gaze data; receive a calibration point gaze determination instruction input by the user, and use the current user gaze data as the calibration point gaze data for the corresponding calibration point; and specifically, to obtain the first user gaze hotspot area based on the first user gaze data and the multiple calibration point gaze data. This processing method allows for the acquisition of calibration point gaze data for each calibration point displayed in the shared screen viewed through the first smart glasses, and the determination of the calibration point gaze data through the first client. Based on the multiple calibration point gaze data and the first user video data, the first user gaze hotspot area can be determined more accurately, even if the first user's gaze is only aimed at a very small area in the shared screen, this small area can still be accurately identified; therefore, the accuracy of hotspot area labeling can be effectively improved, further enhancing video conferencing efficiency.
[0055] In one example, the first client is also used to display a prompt message on the shared screen to encourage the user to look directly at the calibration point. This approach allows the first user to see the prompt and guides them to focus their gaze on the calibration point in a timely manner, facilitating the collection of the user's calibration point gaze data; thus, it effectively improves the user experience.
[0056] In one example, the first smart glasses are also used to acquire a labeling mode entry command and send a labeling mode entry request; the first client is also used to receive the labeling mode entry request; and specifically, to respond to the labeling mode entry request by acquiring the first user's gaze hotspot area based on the first user's gaze data. This processing method allows the first user to input a labeling mode entry command into the first smart glasses, which then forwards the command to the first client. This enables the first client to know that the first smart glasses have entered labeling mode, allowing it to track the user's gaze in real time and label the gaze hotspot area; thus, it effectively enables the labeling process for gaze hotspot areas.
[0057] In another example, the first client is also configured to receive an instruction to invite the first smart glasses to enter annotation mode, send an invitation request to enter annotation mode, and receive the annotation mode entry request; specifically, in response to the annotation mode entry request, it obtains the first user's gaze hotspot area based on the first user's gaze data; the first smart glasses are also configured to receive the invitation request to enter annotation mode and send the annotation mode entry request. This processing method allows the first user to input the annotation model entry instruction to the first client; therefore, it effectively improves the convenience of inputting the annotation model entry instruction, thereby enhancing the user experience.
[0058] In one example, the first client is further configured to acquire a prompt message indicating entry into annotation mode and send the prompt message to the first smart glasses; the first smart glasses are further configured to receive the prompt message and display it. This processing method allows the first user to see the prompt message indicating entry into annotation mode, guiding the first user to promptly enter annotation mode with the first smart glasses, thus enabling the collection of the first user's gaze data while viewing the shared screen; therefore, it effectively improves the user experience.
[0059] It should be noted that entering, enabling, and disabling annotation modes are different annotation mode processes. Entering annotation mode means using the first smart glasses as the gaze tracking device for the first client. Enabling annotation mode means starting to annotate the user's gaze hotspot areas in the shared screen. Disabling annotation mode, corresponding to enabling annotation mode, means stopping the annotation of the user's gaze hotspot areas in the shared screen.
[0060] In one example, the first client is further configured to obtain a request to enable annotation mode; in response to the request to enable annotation mode, annotate the first user's visual hotspot area in the shared screen; and obtain a request to disable annotation mode; in response to the request to disable annotation mode, stop annotating the first user's visual hotspot area.
[0061] In one example, the first smart glasses are also used to obtain a labeling mode enable command and send a labeling mode enable request; and to obtain a labeling mode disable command and send a labeling mode disable request.
[0062] In practice, the first smart glasses are also used to display annotation mode operation options; through the operation options, they receive annotation mode start or annotation mode stop commands.
[0063] In specific implementation, the first smart glasses include a labeling mode operation key; through the operation option, it receives a labeling mode enable command or a labeling mode disable command.
[0064] In another example, the first client is also used to obtain a annotation mode enable instruction and generate the annotation mode enable request; and to obtain a annotation mode end instruction and generate the annotation mode end request.
[0065] In one example, the first smart glasses and the first client transmit data through a conference server. In this case, the first smart glasses are also used to connect to the network via a wireless access point (AP) or a smart terminal (such as a smartphone or tablet); obtain the video conference entry password, and send an entry request to the server; the server is used to determine whether the first smart glasses are allowed to enter the conference based on the first smart glasses identifier carried in the request; if the determination result is yes, the first smart glasses are used as the eye-tracking device for the first client. In specific implementation, the smart glasses connect to the network and connect to the conference system via the conference password. If the first client system does not have Bluetooth, the smart glasses can also connect to the network via "phone Bluetooth" or Wi-Fi and then connect to the conference system via the conference password. This approach allows the glasses to connect to the network via an AP or smart terminal when the first client cannot establish a direct short-range connection with the first smart glasses, thereby directly establishing a connection channel with the server and providing the entry password to the server, thus enabling them to act as the eye-tracking device corresponding to the first client. The server stores the correspondence between smart glasses identifiers and clients, and determines the first client corresponding to the glasses based on the smart glasses identifier carried in the request. Furthermore, for data sent from the first smart glasses to the first client (such as user gaze data, annotation mode entry requests, etc.), the first smart glasses can hand over the data to be transmitted to the server for forwarding to the first client; for data sent from the first client to the first smart glasses (such as invitations to enter annotation mode, prompt messages, etc.), the first client can hand over the data to be transmitted to the server for forwarding to the first smart glasses.
[0066] In another example, the first smart glasses and the first client transmit data via short-range communication, such as Bluetooth. In practice, the first client's video conferencing app monitors the system's Bluetooth connection status in real time. If a smart glasses device is detected connecting via Bluetooth, the app interacts with the smart glasses via Bluetooth to confirm that the smart glasses support "labeling mode," meaning the first smart glasses enters labeling mode.
[0067] Please refer to Figure 2 This is a schematic diagram illustrating the specific device interaction of the video conferencing shared screen processing system of this application. In one example, the system provided in this embodiment may further include: a second client, configured to receive and display the shared screen; receive second user gaze data viewed through second smart glasses; obtain location data of a second gaze hotspot area based on the second user gaze data; and send the location data to the first client via a server; the second smart glasses, configured to collect and send the second user gaze data; and the first client, further configured to receive the location data; and mark the second gaze hotspot area in the shared screen based on the location data. This processing method allows the areas of gaze of participants other than the presenter to be displayed in the shared screen, enabling the presenter to know the area where the participants' questions are located; therefore, it can further improve the efficiency of video conferencing.
[0068] In practice, the first client can label the first line-of-sight hotspot area with a first display attribute in the shared screen; and label the second line-of-sight hotspot area with a third display attribute in the shared screen. This approach allows for the labeling of the first user's and the second user's line-of-sight hotspot areas with different display attributes in the shared screen; therefore, it can further improve video conferencing efficiency and enhance the user experience.
[0069] As can be seen from the above embodiments, the video conferencing shared screen processing system provided in this application displays the shared screen of a video conference through a first client; receives gaze data of a first user viewing the shared screen through first smart glasses; obtains a first user gaze hotspot area based on the first user gaze data; marks the first user gaze hotspot area in the shared screen; and sends the shared screen marked with the first user gaze hotspot area to a second client through a server; the first smart glasses collect the first user gaze data and provide it to the first client; the second client receives the shared screen marked with the first user gaze hotspot area and displays it. This processing method allows the speaker's gaze hotspot area to be drawn on the shared screen in real time, so participants can easily find the content being explained by the speaker in the shared screen; therefore, it can effectively improve video conferencing efficiency and thus enhance the user experience.
[0070] Second Embodiment In the above embodiments, a video conferencing screen sharing processing system is provided. Correspondingly, this application also provides a video conferencing screen sharing processing method for a first client. This method corresponds to the embodiments of the above system. Since the method embodiments are basically similar to the system embodiments, they are described simply, and relevant details can be found in the descriptions of the system embodiments. The method embodiments described below are merely illustrative.
[0071] Please refer to Figure 3 This is a schematic diagram of the device interaction of the video conferencing sharing screen processing system of this application. This application also provides a video conferencing sharing screen processing method, which may include the following steps: Step S301: Display the shared screen of the video conference.
[0072] Step S303: Receive first user gaze data when viewing the shared screen through the first smart glasses.
[0073] Step S305: Obtain the first user's gaze hotspot area based on the first user's gaze data.
[0074] Step S307: Mark the first user's line-of-sight hotspot area in the shared screen.
[0075] Step S309: The server sends the shared screen, which marks the hotspot area in the first user's line of sight, to the second client.
[0076] In one example, the method provided in this application embodiment may further include the following steps: displaying calibration points in the shared screen; receiving calibration point gaze data viewed through first smart glasses; step S305 may be implemented in the following manner: obtaining the first user gaze hotspot area based on the first user gaze data and the calibration point gaze data.
[0077] In one example, the method provided in this application embodiment may further include the following steps: receiving a labeling mode entry request sent by the first smart glasses; step S305 may be implemented in the following manner: responding to the labeling mode entry request, obtaining the first user's gaze hotspot area based on the first user's gaze data.
[0078] In one example, the shared screen includes image content, and the first user's gaze hotspot area includes partial image content.
[0079] In one example, the method provided in this application embodiment may further include the following step: marking the explanation trajectory for multiple local contents of images. The explanation trajectory can reflect the explanation order of multiple local contents of images, such as the explanation trajectory including the direction of arrows.
[0080] In one example, step S307 can be implemented as follows: annotate the local content of the image currently being explained with a first display attribute, and annotate the local content of the image that has already been explained with a second display attribute.
[0081] In one example, the method provided in this application embodiment may further include the following steps: storing the correspondence between user gaze data and gaze hotspot areas; and switching and labeling the multiple image local contents according to the correspondence and the first user gaze data.
[0082] Third Embodiment In the above embodiments, a method for processing shared video conferencing footage is provided. Correspondingly, this application also provides a device for processing shared video conferencing footage. This device corresponds to the embodiments of the method described above. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant details can be found in the descriptions of the method embodiments. The device embodiments described below are merely illustrative.
[0083] This application also provides a video conferencing shared screen processing device, including: a shared screen display unit, a first user gaze data receiving unit, a first user gaze hotspot area determination unit, a first user gaze hotspot area marking unit, and a shared screen sending unit.
[0084] The system includes a shared screen display unit for displaying the shared screen of a video conference; a first user gaze data receiving unit for receiving first user gaze data from a first smart glasses user viewing the shared screen; a first user gaze hotspot area determination unit for obtaining a first user gaze hotspot area based on the first user gaze data; a first user gaze hotspot area marking unit for marking the first user gaze hotspot area in the shared screen; and a shared screen sending unit for sending the shared screen marked with the first user gaze hotspot area to a second client via a server.
[0085] In one example, the apparatus provided in this application embodiment may further include: a calibration point processing unit, configured to display calibration points in the shared screen; receive calibration point viewing data through first smart glasses; and a first user viewing hotspot area determination unit, specifically configured to obtain the first user viewing hotspot area based on the first user viewing data and the calibration point viewing data.
[0086] In one example, the apparatus provided in this application embodiment may further include: a labeling mode entry unit, configured to receive a labeling mode entry request sent by a first smart glasses; and a first user gaze hotspot area determination unit, specifically configured to respond to the labeling mode entry request and obtain a first user gaze hotspot area based on the first user gaze data.
[0087] In one example, the shared screen includes image content, and the first user's gaze hotspot area includes partial image content.
[0088] In one example, the apparatus provided in this application embodiment may further include: a narration trajectory annotation unit, used to annotate the narration trajectory of multiple image local contents. The narration trajectory can reflect the narration order of multiple image local contents, such as including the direction of arrows in the narration trajectory.
[0089] In one example, the first user-view hotspot area annotation unit is specifically used to annotate the local content of the image currently being explained with a first display attribute, and to annotate the local content of the image that has already been explained with a second display attribute.
[0090] In one example, the apparatus provided in this application embodiment may further include: a switching annotation unit, used to store the correspondence between user gaze data and gaze hotspot areas; and to switch annotations on the plurality of image local contents according to the correspondence and the first user gaze data.
[0091] Fourth embodiment In the above embodiments, a video conferencing shared image processing system is provided. Correspondingly, this application also provides a video conferencing shared image processing method for smart glasses. This method corresponds to the embodiments of the above system. Since the method embodiments are basically similar to the system embodiments, they are described simply, and relevant details can be found in the descriptions of the system embodiments. The method embodiments described below are merely illustrative.
[0092] This application also provides a method for processing shared video conferencing footage, which may include the following steps: Step S401: Collect the gaze data of the first user who is watching the shared video conference screen displayed on the first client.
[0093] Step S401: Send the first user's gaze data.
[0094] In one example, the method provided in this application embodiment may further include the following steps: acquiring calibration point line-of-sight data of viewing calibration points in the shared screen; and sending the calibration point line-of-sight data.
[0095] Fifth embodiment In the above embodiments, a method for processing shared video conferencing footage is provided. Correspondingly, this application also provides a device for processing shared video conferencing footage. This device corresponds to the embodiments of the method described above. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant details can be found in the descriptions of the method embodiments. The device embodiments described below are merely illustrative.
[0096] This application also provides a video conferencing shared screen processing device, including: a first user gaze data acquisition unit, used to acquire first user gaze data of watching the video conferencing shared screen displayed on a first client; and a first user gaze data transmission unit, used to transmit the first user gaze data.
[0097] In one example, the apparatus provided in this application embodiment may further include: a calibration point line-of-sight data processing unit, used to collect calibration point line-of-sight data of viewing calibration points in the shared screen; and to send the calibration point line-of-sight data.
[0098] Sixth Embodiment In the above embodiments, a video conferencing screen sharing processing method is provided. Correspondingly, this application also provides an electronic device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant details can be found in the description of the method embodiments. The device embodiments described below are merely illustrative.
[0099] The electronic device of this embodiment includes: a memory and a processor; the memory is used to store a program for implementing the above-described video conferencing screen sharing processing method, and the device is powered on and runs the program of the above-described video conferencing screen sharing processing method through the processor.
[0100] Memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0101] In specific implementations, the electronic device may further include one or more of the following components: a power supply component, an input / output (I / O) interface, and a communication component. The power supply component provides power to various components of the electronic device. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device. The I / O interface provides an interface between the processor 503 and peripheral interface modules, which may be a keyboard, click wheel, buttons, etc. The communication component is configured to facilitate wired or wireless communication between the electronic device and a first user device (such as a smartphone, tablet, etc.).
[0102] Seventh Embodiment This application also provides a computer-readable storage medium. Since the embodiments of the computer-readable storage medium are substantially similar to the method embodiments, the description is relatively simple; relevant details can be found in the description of the method embodiments. The computer-readable storage medium embodiments described below are merely illustrative.
[0103] In this embodiment, a non-transitory computer-readable storage medium including instructions is provided, such as a memory including instructions. These instructions can be executed by a processor of an electronic device to complete the video conferencing sharing image processing method provided in this disclosure. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.
[0104] It should be noted that the embodiments of this application may involve the use of the first user's data. In practical applications, the first user's specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., the first user has given explicit consent, the first user has been properly notified, etc.).
[0105] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.
[0106] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0107] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0108] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.
[0109] 2. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
Claims
1. A video conferencing shared image processing system, characterized in that, include: A first client is used to display the shared screen of a video conference; receive gaze data of a first user viewing the shared screen through first smart glasses; obtain the first user's gaze hotspot area based on the first user's gaze data; and mark the first user's gaze hotspot area in the shared screen. The server sends the shared screen, which marks the hotspot area in the first user's line of sight, to the second client. The first smart glasses are used to collect and transmit the first user's gaze data.
2. The system according to claim 1, characterized in that, The shared screen includes image content, and the first user's gaze hotspot area includes partial image content.
3. The system according to claim 2, characterized in that, The first client is also used to annotate the explanation trajectory of multiple local contents of images.
4. The system according to claim 2, characterized in that, The first client is also used to annotate the local content of the image currently being explained with a first display attribute, and to annotate the local content of the image that has already been explained with a second display attribute.
5. The system according to claim 2, characterized in that, The first client is also used to store the correspondence between user gaze data and gaze hotspot areas; and to switch and annotate the local content of the multiple images according to the correspondence and the first user gaze data.
6. The system according to claim 2, characterized in that, The image content includes video footage, and the partial image content includes partial content of the video footage; The first client is also configured to: obtain a request to enable the annotation mode; respond to the request to enable the annotation mode and annotate the first user's visual hotspot area in the shared screen; and obtain a request to disable the annotation mode; respond to the request to disable the annotation mode and stop annotating the first user's visual hotspot area.
7. The system according to claim 6, characterized in that, The first smart glasses are also used to obtain instructions to enable annotation mode and send annotation mode enable requests; and to obtain instructions to disable annotation mode and send annotation mode disable requests.
8. The system according to claim 1, characterized in that, The first client is also configured to display multiple calibration points in the shared screen; receive calibration point gaze data viewed through the first smart glasses; and specifically, to obtain the first user gaze hotspot area based on the first user gaze data and the multiple calibration point gaze data of the multiple calibration points. The first smart glasses are also used to acquire and send the calibration point gaze data.
9. The system according to claim 1, characterized in that, Also includes: The second client is used to receive the shared screen and display the shared screen. Receive second user gaze data from viewing the shared screen through second smart glasses; Based on the second user gaze data, obtain the location data of the second gaze hotspot area; The location data is sent from the server to the first client. The second smart glasses are used to collect and send the second user's gaze data. The first client is specifically used to mark the first line-of-sight hotspot area in the shared screen; the first client is also used to receive the location data; and mark the second line-of-sight hotspot area in the shared screen according to the location data.
10. A method for processing shared video feeds in a video conferencing conference, characterized in that, include: Displays the shared screen of the video conference; Receives gaze data from a first user viewing the shared screen through first smart glasses; Based on the first user's gaze data, obtain the first user's gaze hotspot area; The first user's line-of-sight hotspot area is marked in the shared screen; The server sends the shared screen, which marks the hotspot area in the first user's line of sight, to the second client.
11. A video conferencing shared image processing method for a first smart glasses, characterized in that, include: Collect gaze data of the first user who is viewing the shared video conference feed displayed on the first client. Send the first user's gaze data.
12. A video conferencing shared image processing device, characterized in that, include: The shared screen display unit is used to display the shared screen of the video conference; The first user gaze data receiving unit is used to receive first user gaze data when viewing the shared screen through the first smart glasses; The first user gaze hotspot area determination unit is used to obtain the first user gaze hotspot area based on the first user gaze data. The first user gaze hotspot area marking unit is used to mark the first user gaze hotspot area in the shared screen; The shared screen sending unit is used to send the shared screen marked with the hot spot area in the first user's line of sight to the second client through the server.
13. A video conferencing shared image processing device, characterized in that, include: The first user gaze data acquisition unit is used to collect the gaze data of the first user who is watching the shared video conference screen displayed by the first client. The first user gaze data sending unit is used to send the first user gaze data.
14. An electronic device, characterized in that, include: processor; as well as A memory for storing a program for implementing the method according to any one of claims 10 to 11, wherein the device is powered on and the program for running the method is executed by the processor.