Video conference method and apparatus, device, and storage medium
Near-eye display devices solve the problem of unclear audio data in multi-speaker situations by tracking the user's gaze focus, filtering and broadcasting target audio data in video conferences, thus improving the auditory experience and broadcasting effect of video conferences.
Patent Information
- Application Number
- PCT/CN2025/102846
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-13
- Filing Date
- 2025-06-23
- Publication Date
- 2026-02-19
AI Technical Summary
In video conferencing, when multiple speakers speak at the same time, users may have difficulty hearing the audio data of the speaker they are interested in, resulting in poor audio data playback quality.
By tracking the user's gaze focus through a near-eye display device, the system identifies the speaker the user is interested in and extracts and outputs the speaker's target audio data from preset audio data. This includes filtering audio data based on the speaker's vocal range and voiceprint information according to the gaze focus, and adjusting the playback priority and volume appropriately.
It improves the audio data playback effect in multi-speaker scenarios, ensuring that users can more easily hear the speakers they are interested in, thus enhancing the audio experience and playback accuracy of video conferencing.
Smart Images

Figure CN2025102846_19022026_PF_FP_ABST
Abstract
Description
Video conference method, device, equipment and storage medium
[0001] The present application claims priority from the Chinese patent application No. 202411110284X filed on August 13, 2024, and entitled "Video conference method, device, equipment and storage medium", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the technical field of video conference, and in particular to a video conference method, device, equipment and storage medium. BACKGROUND
[0003] At present, when a user participates in a video conference, if there are too many speakers speaking at the same time in the video conference, the user may not be able to hear the voice data of the speaker he / she is interested in, resulting in poor voice data broadcasting effect of the video conference. SUMMARY
[0004] The main purpose of the present application is to provide a video conference method, device, equipment and storage medium, which aims to solve the technical problem of poor voice data broadcasting effect of the video conference due to the user being unable to hear the voice data of the speaker he / she is interested in when there are too many speakers speaking at the same time in the video conference.
[0005] In a first aspect, the present application provides a video conference method applied to a near-eye display device, the video conference method comprising:
[0006] displaying a video picture of a video conference, the video picture comprising at least one speaker;
[0007] obtaining a visual line focal point of a user wearing the near-eye display device on the video picture;
[0008] determining target voice data from preset voice data of at least one speaker corresponding to the video conference according to the visual line focal point, and outputting the target voice data.
[0009] In a second aspect, the present application provides a video conference device, the video conference device comprising:
[0010] a picture display module configured to display a video picture of a video conference, the video picture comprising at least one speaker;
[0011] a focal point obtaining module configured to obtain a visual line focal point of a user wearing the near-eye display device on the video picture;
[0012] The target determining module is configured to determine target sound data from preset sound data of at least one speaker corresponding to the video conference according to the line-of-sight focus, and output the target sound data.
[0013] In a third aspect, the present application provides a near-eye display device, comprising a memory and a processor;
[0014] The memory is configured to store a computer program.
[0015] The processor is configured to execute the computer program and implement the steps of the video conference method as described above when executing the computer program.
[0016] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the video conference method as described above are implemented.
[0017] The present application provides a video conference method, device, equipment and storage medium. The video conference method comprises the following steps: displaying a video picture of a video conference, wherein the video picture comprises at least one speaker; obtaining a line-of-sight focus of a user wearing a near-eye display device on the video picture; determining target sound data from preset sound data of at least one speaker corresponding to the video conference according to the line-of-sight focus, and outputting the target sound data.
[0018] During the process of displaying the video picture of the video conference, the near-eye display device can determine a speaker that the user focuses on from at least one speaker included in the video picture according to the line-of-sight focus of the user on the video picture. For example, the speaker that the user focuses on is the speaker whose line-of-sight focus is focused on the video picture. Based on this, the near-eye display device can obtain target sound data of the speaker that the user focuses on from preset sound data of at least one speaker corresponding to the video conference. Accordingly, the near-eye display device can output the target sound data, such as playing the target sound data, so that the user can more conveniently hear the target sound data of the speaker that the user focuses on, thereby facilitating improvement of the playing effect of the target sound data of the video conference by the near-eye display device. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0020] FIG. 1 is a flowchart of a video conference method according to an embodiment of the present application;
[0021] FIG. 2 is a schematic diagram of a video screen of a video conference according to an embodiment of the present application;
[0022] FIG. 3 is a schematic diagram of a video screen of a video conference according to another embodiment of the present application;
[0023] FIG. 4 is a schematic block diagram of a video conference device according to an embodiment of the present application;
[0024] FIG. 5 is a schematic block diagram of a near-eye display device according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0026] The flowcharts shown in the drawings are only exemplary and do not necessarily include all contents and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps can be further divided, combined or partially merged, so the actual execution order may be changed according to the actual situation.
[0027] The embodiments of the present application provide an interaction method, device and equipment of a near-eye display device and a storage medium. The interaction method of the near-eye display device can be applied to the near-eye display device. The near-eye display device can include an extended reality (XR) device. The XR device can include an augmented reality (AR) glasses, a virtual reality (VR) glasses, a mixed reality (MR) glasses, an AR helmet, a VR helmet, an MR helmet, etc., without limitation.
[0028] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. In the case of no conflict, the embodiments described below and the features in the embodiments can be combined with each other.
[0029] Please refer to FIG. 1, which is a flowchart of a video conference method according to an embodiment of the present application. It should be noted that the video conference method provided by the embodiments of the present application can be used in a near-eye display device.
[0030] As shown in FIG. 1, the video conference method includes steps S101 to S104.
[0031] S101, display a video picture of a video conference, the video picture comprising at least one speaker.
[0032] For example, the near-eye display device accesses the video conference and displays a video picture of the video conference upon detecting an access instruction for the video conference. For example, a user wearing the near-eye display device can send an access instruction for the video conference to the near-eye display device through gesture information, voice information, device pose information, touch information, etc., without limitation.
[0033] For example, the speaker can be determined according to at least one of a speaker in an offline conference scene or a speaker identifier of an online participant.
[0034] In some embodiments, the video conference in which the user participates can have a corresponding offline conference, and the video picture displayed by the near-eye display device can comprise a picture related to the offline conference scene. As shown in FIG. 2, the video picture can comprise a picture corresponding to a participant in the offline conference scene. Since the participants in the offline conference scene can choose whether to speak according to their own needs, the video picture can comprise at least one speaker.
[0035] In some embodiments, the video conference in which the user participates adopts a conference form of an online conference. In the case where the conference form of the video conference is an online conference, the video picture displayed by the near-eye display device can comprise a speaker identifier of an online participant. As shown in FIG. 3, the speaker identifier can comprise a preset icon corresponding to a speaker, a virtual image corresponding to the speaker, etc., without limitation. The preset icon can comprise a head portrait of the speaker, a name, etc., without limitation. Since the participants in the online participant list can choose whether to speak according to their own needs, the video picture can comprise at least one speaker.
[0036] Of course, it is not limited thereto, for example, the video conference can also adopt a conference form in which an online conference and an offline conference are performed in parallel. Accordingly, the video picture of the video conference can comprise at least one of a picture corresponding to a participant in an offline conference scene or a speaker identifier of an online participant, without limitation.
[0037] In the case where the near-eye display device displays a video picture of a video conference, the user can determine a speaker of interest from at least one speaker included in the video picture, for subsequent determination of target speech data by the near-eye display device.
[0038] In this way, the near-eye display device displays a video picture of the video conference, and the video picture includes at least one speaker, so that the user can intuitively view the speaker of the video conference, thereby improving the visualization effect of the speaker in the video conference. Accordingly, the video picture of the video conference displayed by the near-eye display device can be used by the near-eye display device to subsequently determine the target voice data, thereby improving the convenience of the near-eye display device in determining the target voice data.
[0039] S102, acquire a visual focus point of the user wearing the near-eye display device on the video picture.
[0040] For example, in the process of participating in the video conference, if the user is interested in a speaker, the user can focus the visual line on the speaker that the user is interested in. Based on this, the near-eye display device can acquire the visual focus point of the user on the video picture, so as to subsequently determine the voice data of the speaker that the user is interested in according to the visual focus point of the user.
[0041] In some embodiments, an eye tracking sensor can be arranged on the near-eye display device. The eye tracking sensor can be used to identify the visual fixation position (also referred to as the visual focus position) of the user. In the case of identifying the visual fixation position of the user, the near-eye display device can determine the visual focus point of the user on the video picture in combination with the video picture displayed by the near-eye display device.
[0042] The visual focus point can be used by the near-eye display device to subsequently determine the target voice data.
[0043] S103, determine the target voice data from the preset voice data of at least one speaker corresponding to the video conference according to the visual focus point, and output the target voice data.
[0044] For example, in the process of participating in the video conference, the near-eye display device can acquire the preset voice data of at least one speaker corresponding to the video conference. For example, in the case where a speaker is speaking, the near-eye display device can acquire the corresponding voice data, and the near-eye display device can determine the corresponding voice data as the preset voice data of the speaker. In addition, the near-eye display device can integrate the preset voice data of all speakers as the preset voice data of at least one speaker corresponding to the video conference.
[0045] In the case where the near-eye display device acquires the visual focus point of the user on the video picture, the near-eye display device can determine that the visual line of the user is focused on at least one speaker according to the visual focus point. Accordingly, the near-eye display device can determine that the speaker focused by the visual line of the user is a target speaker that the user is interested in. Accordingly, the near-eye display device can determine the voice data of the target speaker included in the preset voice data as the target voice data.
[0046] In this way, the near-eye display device can determine the target speaker according to the visual focus of the user, thereby improving the accuracy of determining the target speaker.
[0047] In some embodiments, the target speaker is determined from the speakers included in the video image according to the visual focus; the sound emitting orientation is determined to be within a target sound emitting range corresponding to the target speaker, and / or the voiceprint information is determined to be consistent with target voiceprint information of the target speaker, and the preset sound data corresponding to the consistent voiceprint information is the target sound data.
[0048] For example, when the visual focus is within the image range of the speaker in the video image, the near-eye display device can determine that the speaker is the target speaker. The target speaker is used to indicate the speaker that the user is paying attention to. Of course, this is not limited, and when the visual focus is within the image range of each of the multiple speakers in the video image, the near-eye display device can also determine that the multiple speakers are target speakers. For example, when the near-eye display device displays a video image under a certain preset shooting angle, the image ranges of the multiple speakers in the video image can have overlapping image ranges in the case that the seats of the multiple speakers are relatively close to each other. Accordingly, when the visual focus is within the overlapping image range, the near-eye display device can determine that the multiple speakers corresponding to the overlapping image range are target speakers, without limitation.
[0049] In the case of determining the target speaker, at least one of the target sound emitting range corresponding to the target speaker and the target voiceprint information of the target speaker can be obtained. The target sound emitting range is used to indicate a position range in which at least one sound source position corresponding to the sound data collected by the near-eye display device when the target speaker speaks is located. The target voiceprint information is used to indicate the voiceprint information of the target speaker. The target sound emitting range and the target voiceprint information can be determined in advance or updated in real time, without limitation.
[0050] For example, when the speaker speaks, the speaker can change the head posture, such as at least one of changing the head position and changing the head direction. For example, in the case of a speaker in an offline meeting, the speaker can lower the head to check the file or raise the head to look at the participant, and the head posture of the speaker will change. In response to the change of the head posture of the speaker, the mouth posture of the speaker will also change, which is equivalent to the change of the sound source position corresponding to the sound data. Based on this, the near-eye display device can determine the target sound emitting range corresponding to the target speaker according to the change range of the head posture.
[0051] At least one of the target sound emitting range and the target voiceprint information can be used by the near-eye display device to filter the target sound data from the preset sound data corresponding to the video conference. The target sound data is used to indicate the sound data of the target speaker.
[0052] For example, the near-eye display device can determine whether there is sound data in the preset sound data in which the sound orientation is in the target sound range corresponding to the target speaker. If there is sound data in the preset sound data in which the sound orientation is in the target sound range, and the possibility that the sound data is the sound data of the target speaker is greater than or equal to the preset probability threshold, the near-eye display device can determine that the sound data is the target sound data, and filter the target sound data from the preset sound data according to the target sound range.
[0053] For example, the near-eye display device can determine whether there is sound data in the preset sound data in which the voiceprint information is the target voiceprint information of the target speaker. If there is sound data in the preset sound data in which the voiceprint information is the target voiceprint information of the target speaker, and the possibility that the sound data is the sound data of the target speaker is greater than or equal to the preset probability threshold, the near-eye display device can determine that the sound data is the target sound data, and filter the target sound data from the preset sound data according to the target voiceprint information.
[0054] Of course, the near-eye display device can also filter the target sound data from the preset sound data by comprehensively considering the target sound range corresponding to the target speaker and the target voiceprint information.
[0055] In this way, when determining the target sound data, the near-eye display device can determine the target speaker according to the user's visual focus, which is beneficial to improve the convenience of determining the target speaker. Accordingly, when the target speaker is determined, the near-eye display device can flexibly filter the preset sound data according to at least one of the target sound range corresponding to the target speaker and the target voiceprint information to obtain the target sound data, which is beneficial to improve the flexibility of determining the target sound data.
[0056] In some embodiments, when the target sound data is determined, the near-eye display device can output the target sound data, such as playing the target sound data.
[0057] For example, the near-eye display device can set at least one of the playing order and the playing volume of the target sound data, so that the user can better hear the target sound data of the target speaker who is the focus of the user, thereby improving the playing effect of the target sound data by the near-eye display device.
[0058] For example, in a case where there are multiple speakers speaking at the same time, the multiple speakers include the target speaker and speakers other than the target speaker, the near-eye display device can determine that the preset playing priority of the target sound data of the target speaker is the highest, and then the near-eye display device can set the playing order of the target sound data to be earlier than the playing order of the sound data of the speakers other than the target speaker, so as to play the target sound data preferentially.
[0059] Of course, it is not limited to this, for example, in the case where multiple speakers exist and multiple speakers include target speakers and speakers other than target speakers, the near-eye display device can determine that the target sound data of the target speakers needs to be highlighted. Based on this, the near-eye display device can broadcast the target sound data at a preset broadcast volume. Accordingly, the near-eye display device can reduce the broadcast volume of the sound data of the speakers other than the target speakers, and can not broadcast the sound data of the speakers other than the target speakers before the target sound data is broadcast, and the like, which is not limited herein.
[0060] In some embodiments, when multiple target speakers exist, the target sound data of at least one target speaker is broadcast according to a preset broadcast priority between the multiple target speakers.
[0061] For example, during a video conference, a user can gaze at multiple speakers, and the number of target speakers determined by the user's line of sight focus of the near-eye display device can also be multiple. In the case where multiple target speakers speak in the same time period, if the near-eye display device simultaneously broadcasts the target sound data of each of the multiple target speakers, it is easy to cause the user to not be able to clearly hear the multiple target sound data. Based on this, the near-eye display device can broadcast the target sound data of at least one target speaker according to a preset broadcast priority between the multiple target speakers. The preset broadcast priority can be preset or set by the user, which is not limited herein. For example, the near-eye display device can broadcast the target sound data of a target speaker with a higher priority first, and then broadcast the target sound data of a target speaker with a lower priority according to the preset broadcast priority between the multiple target speakers. Of course, it is not limited to this, in the case where the near-eye display device broadcasts the multiple target sound data according to the preset broadcast priority, the near-eye display device can stop broadcasting the remaining target sound data in response to receiving an end broadcast instruction. The near-eye display device can also determine a current broadcast duration of the target sound data to stop broadcasting the remaining target sound data when the current broadcast duration reaches a preset duration threshold, which is not limited herein.
[0062] In this way, when multiple target speakers exist, the target sound data of at least one target speaker is broadcast according to a preset broadcast priority between the multiple target speakers, the near-eye display device can determine the target sound data that the user is most interested in according to the preset broadcast priority, and preferentially broadcast the target sound data that the user is most interested in, which is beneficial to improve the broadcast accuracy and flexibility of the near-eye display device for the target sound data, and improve the auditory experience of the user for the target sound data.
[0063] In some embodiments, when the sound volume corresponding to the target sound data is less than or equal to a preset volume threshold, the target sound data is played according to a preset playing volume, and the preset playing volume is greater than the preset volume threshold.
[0064] For example, when the speaking voice of the target speaker is small, the sound volume corresponding to the target sound data is also small accordingly. When the sound volume corresponding to the target sound data is small, the user is likely to be unable to hear the target sound data clearly. Based on this, the near-eye display device can determine whether the sound volume corresponding to the target sound data is less than or equal to a preset volume threshold. If the sound volume corresponding to the target sound data is less than or equal to the preset volume threshold, it can be determined that the user is likely to be unable to hear the target sound data clearly. Based on this, the near-eye display device can play the target sound data according to a preset playing volume, and the preset playing volume is greater than the preset volume threshold, so that the user can hear the target sound data clearly. The preset playing volume can be preset or set by the user, which is not limited herein.
[0065] In this way, when the sound volume corresponding to the target sound data is less than or equal to the preset volume threshold, the target sound data is played according to the preset playing volume, and the preset playing volume is greater than the preset volume threshold, so that the user can hear the target sound data played by the near-eye display device clearly, which is beneficial to improve the playing flexibility of the near-eye display device for the target sound data and improve the auditory experience of the user for the target sound data.
[0066] In some embodiments, when the target sound data is output, target text information corresponding to the target sound data is displayed.
[0067] For example, when the target sound data is output by the near-eye display device, if the user's speaking speed corresponding to the target sound data is fast, the user is likely to be unable to hear the target sound data played currently. Based on this, in order to improve the auditory experience of the user for the target sound data, the near-eye display device can display the target text information corresponding to the target sound data to assist the user to know the target sound data. The text type of the target text information, such as font type, font size, language type, etc., can be preset or set by the user, which is not limited herein.
[0068] In this way, the near-eye display device displays the target text information corresponding to the target sound data when playing the target sound data, so that the user can view the target text information corresponding to the target sound data when listening to the target sound data played by the near-eye display device, which is beneficial to improve the accuracy of the user to know the target sound data and improve the flexibility of the near-eye display device to convey the target sound data.
[0069] For example, the near-eye display device can acquire preset sound data corresponding to the video conference. For example, the preset sound data corresponding to the video conference can be stored in a cloud corresponding to the near-eye display device. Accordingly, the cloud can send the preset sound data corresponding to the video conference to the near-eye display device by data broadcasting, without limitation.
[0070] The near-eye display device can determine the preset sound data in advance before acquiring the preset sound data.
[0071] In some embodiments, initial sound data is acquired; when a sound emitting orientation of the initial sound data is within a preset sound emitting range corresponding to a speaker included in a video picture, and / or voiceprint information of the initial sound data matches preset voiceprint information of the speaker, the initial sound data is processed according to a preset identity identifier of the speaker to obtain initial sound data carrying the preset identity identifier; and preset sound data of at least one speaker corresponding to the video conference is determined according to the initial sound data carrying the preset identity identifier.
[0072] For example, during a video conference, the near-eye display device can collect initial sound data. The near-eye display device can also determine preset sound emitting ranges corresponding to each speaker in the video picture, and preset voiceprint information of each speaker. For example, the near-eye display device can be provided with a positioning sensor. The near-eye display device can locate the position of the speaker through the positioning sensor. The near-eye display device can determine the preset sound emitting range corresponding to the speaker according to the position of the speaker. For example, when the speaker speaks, the speaker's head posture may change. In response to the change in the head posture of the speaker, the position of the speaker's mouth changes, and the sound source position corresponding to the sound data changes. Based on this, the near-eye display device can determine the preset sound emitting range corresponding to the preset speaker according to the change range of the head posture. Accordingly, the near-eye display device can determine whether the initial sound data is the sound data of the preset speaker according to at least one of the preset sound emitting range corresponding to the preset speaker and the preset voiceprint information. For example, when the sound emitting direction of the initial sound data is within the preset sound emitting range corresponding to one of the speakers included in the video picture, and / or the voiceprint information of the initial sound data matches the preset voiceprint information of the speaker, the near-eye display device can determine that the possibility that the initial sound data is the sound data of the speaker is greater than or equal to a preset probability threshold. Then, the near-eye display device can process the initial sound data according to the preset identity identifier of the speaker to obtain initial sound data carrying the preset identity identifier. The preset identity identifier can include a device identification number of the speaker, account information of the speaker, a name of the speaker, and the like, which are not limited herein. The device identification number can include a media access control address (Media Access Control Address, MAC address), an identity document (Identity Document, ID), and the like, which are not limited herein.
[0073] Accordingly, the near-eye display device can determine the preset sound data of at least one speaker corresponding to the video conference according to all initial sound data carrying the preset identity identifier.
[0074] For example, in a case where the preset sound data corresponding to the video conference is determined according to all initial sound data carrying the preset identity identifier, the near-eye display device can perform subsequent processing on the initial sound data according to the preset identity identifier. For example, when storing the preset sound data, the preset sound data can be stored in different storage locations according to different preset identity identifiers. For example, when outputting the preset sound data, the preset sound data can be translated into a corresponding translation language according to language information corresponding to the preset identity identifier, and the translated preset sound data can be played, and the like, which are not limited herein.
[0075] In some embodiments, the preset voice data is converted into preset text according to the preset identity, and conference record information of the video conference is obtained.
[0076] For example, the near-eye display device can distinguish the respective speaker corresponding to each preset voice data based on the preset identity corresponding to the preset voice data. Based on this, the conference record information of the video conference obtained by converting the preset voice data into preset text according to the preset identity can determine the preset text corresponding to the preset voice data of each speaker.
[0077] When the user needs to record or review the related conference information of the video conference, the user can obtain the conference record information of the video conference through the near-eye display device to determine the specific speech of the speaker according to the conference record information, so that the user does not need to write the conference record information by himself, which is beneficial to improve the convenience of determining the conference record information.
[0078] The video conference method provided in the above embodiments includes: displaying a video picture of a video conference, the video picture including at least one speaker; obtaining a visual focus point of a user wearing a near-eye display device on the video picture; determining target voice data from preset voice data of at least one speaker corresponding to the video conference according to the visual focus point, and outputting the target voice data.
[0079] During the process of displaying the video picture of the video conference, the near-eye display device can determine the speaker that the user focuses on from the at least one speaker included in the video picture according to the visual focus point of the user on the video picture. For example, the speaker that the visual focus point of the user focuses on in the video picture is the speaker that the user focuses on. Based on this, the near-eye display device can obtain the target voice data of the speaker that the user focuses on from the preset voice data corresponding to the video conference. Accordingly, the near-eye display device can output the target voice data, such as playing the target voice data, so that the user can more conveniently hear the target voice data of the speaker that the user focuses on, thereby being beneficial to improve the playing effect of the target voice data of the video conference by the near-eye display device.
[0080] Please refer to FIG. 4, which is a schematic block diagram of a video conference device provided in an embodiment of the present application. The video conference device can be configured in a near-eye display device for executing the aforementioned video conference method. The near-eye display device can include an XR device. The XR device can include AR glasses, VR glasses, MR glasses, AR helmet, VR helmet, MR helmet, etc., without limitation here.
[0081] As shown in FIG. 4, the video conference device includes a picture display module 110, a focus point obtaining module 120, and a target determining module 130.
[0082] The picture display module 110 is configured to display a video picture of the video conference, the video picture comprising at least one speaker.
[0083] The focus obtaining module 120 is configured to obtain a line-of-sight focus of a user wearing a near-eye display device on the video picture.
[0084] The target determining module 130 is configured to determine target sound data from preset sound data of at least one speaker corresponding to the video conference according to the line-of-sight focus, and output the target sound data.
[0085] In an example, the target determining module 130 comprises a speaker determining sub-module and a sound data determining sub-module.
[0086] The speaker determining sub-module is configured to determine a target speaker from the speakers included in the video picture according to the line-of-sight focus.
[0087] The sound data determining sub-module is configured to determine that the sound emitting orientation is within a target sound emitting range corresponding to the target speaker, and / or the voiceprint information conforms to target voiceprint information of the target speaker, and the preset sound data is the target sound data.
[0088] In an example, the target determining module 130 comprises a first broadcasting sub-module.
[0089] The first broadcasting sub-module is configured to, when there are multiple target speakers, broadcast target sound data of at least one target speaker according to a preset broadcasting priority between the multiple target speakers.
[0090] In an example, the target determining module 130 comprises a second broadcasting sub-module.
[0091] The second broadcasting sub-module is configured to, when a sound volume corresponding to the target sound data is less than or equal to a preset volume threshold, broadcast the target sound data at a preset broadcasting volume, the preset broadcasting volume being greater than the preset volume threshold.
[0092] In an example, the video conference device further comprises a data obtaining sub-module, a first data processing sub-module, and a second data processing sub-module.
[0093] The data obtaining sub-module is configured to obtain initial sound data.
[0094] The first data processing sub-module is configured to, when the sound emitting orientation of the initial sound data is within a preset sound emitting range corresponding to a speaker included in the video frame, and / or the voiceprint information of the initial sound data matches preset voiceprint information of the speaker, process the initial sound data according to a preset identity of the speaker to obtain initial sound data carrying the preset identity.
[0095] The second data processing sub-module is configured to determine preset sound data of at least one speaker of the video conference according to the initial sound data carrying the preset identity.
[0096] The video conference device further includes a conference recording sub-module.
[0097] The conference recording sub-module is configured to convert the preset sound data into preset text according to the preset identity to obtain conference recording information of the video conference.
[0098] The video conference device further includes a text display sub-module.
[0099] The text display sub-module is configured to display target text information corresponding to the target sound data when the target sound data is played.
[0100] It should be noted that, for the convenience and brevity of description, the specific working processes of the above-described device and each module and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein.
[0101] The method of the present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0102] Exemplarily, the method and the device described above can be implemented in the form of a computer program, which can run on the near-eye display device to participate in the video conference through the near-eye display device. Exemplarily, the near-eye display device can include an XR device, such as a VR glasses, an AR glasses, an MR glasses, a VR helmet, an AR helmet, an MR helmet, and the like, without limitation.
[0103] Please refer to FIG. 5, which is a structural schematic block diagram of a near-eye display device provided by an embodiment of the present application.
[0104] As shown in FIG. 5, the near-eye display device includes a memory and a processor. The memory and the processor can be connected through a system bus. The memory can include a storage medium and an internal memory.
[0105] The storage medium can store an operating system and a computer program. The computer program, when executed, can make the processor execute any kind of video conference method.
[0106] The processor is configured to provide computing and control capabilities to support the operation of the entire near-eye display device.
[0107] The internal memory provides an environment for the execution of the computer program in the storage medium. The computer program, when executed by the processor, can make the processor execute any kind of video conference method.
[0108] Those skilled in the art can understand that the structure shown in FIG. 5 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the near-eye display device to which the scheme of the present application is applied. The specific near-eye display device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0109] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0110] In one embodiment, the processor is configured to execute the computer program and implement the following steps when executing the computer program:
[0111] displaying a video picture of a video conference, the video picture comprising at least one speaker;
[0112] obtaining a line-of-sight focus point of a user wearing a near-eye display device on the video picture;
[0113] determining target sound data from preset sound data of the at least one speaker corresponding to the video conference according to the line-of-sight focus point;
[0114] playing the target sound data.
[0115] It should be noted that, for the convenience and brevity of description, the above description of the specific working process of the video conference can refer to the corresponding process in the foregoing video conference method embodiments, and will not be described here.
[0116] The embodiments of the present application further provide a computer readable storage medium, and the computer readable storage medium stores a computer program. The method realized by executing the computer program by a processor can refer to each embodiment of the video conference method of the present application.
[0117] The computer readable storage medium can be an internal storage unit of the near-eye display device, for example, a hard disk or a memory of the near-eye display device. The computer readable storage medium can also be an external storage device of the near-eye display device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0118] It should be understood that the terms used herein in the specification and the appended claims are merely used for the purpose of describing particular embodiments and are not intended to limit the present application. As used in the specification and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0119] It should also be understood that, in the specification and the appended claims, the terms "and / or" is used to mean one or more of the associated listed items, as well as any combination of any of the associated listed items. It will be apparent to those skilled in the art that various modifications and variations can be made to the present application without departing from the spirit or scope of the application. Thus, it is intended that the present application cover modifications and variations of this application provided they come within the scope of the appended claims and their equivalents.
[0120] The above-mentioned embodiment serial numbers of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments. The above is only a specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A video conference method applied to a near-eye display device, the video conference method comprising: displaying a video picture of a video conference, the video picture comprising at least one speaker; obtaining a visual line focus point of a user wearing the near-eye display device on the video picture; determining target sound data from preset sound data of at least one speaker corresponding to the video conference according to the visual line focus point, and outputting the target sound data.
2. The video conferencing method of claim 1, wherein, The determining target sound data from preset sound data of at least one speaker corresponding to the video conference according to the visual line focus point comprises: determining a target speaker from the speakers included in the video picture according to the visual line focus point; determining that preset sound data in which a sound emitting orientation is within a target sound emitting range corresponding to the target speaker and / or voiceprint information is consistent with target voiceprint information of the target speaker is the target sound data.
3. The video conferencing method of claim 2, wherein, The outputting the target sound data comprises: when there are multiple target speakers, outputting target sound data of at least one target speaker according to a preset announcement priority between the multiple target speakers.
4. The video conferencing method of claim 1, wherein, The outputting the target sound data comprises: when a sound emitting volume corresponding to the target sound data is less than or equal to a preset volume threshold, outputting the target sound data at a preset announcement volume, the preset announcement volume being greater than the preset volume threshold.
5. The video conferencing method of claim 1, wherein, Before the determining target sound data from preset sound data of at least one speaker corresponding to the video conference according to the visual line focus point, the video conference method further comprises: obtaining initial sound data; when a sound emitting orientation of the initial sound data is within a preset sound emitting range corresponding to a speaker included in the video picture and / or voiceprint information of the initial sound data is consistent with preset voiceprint information of the speaker, processing the initial sound data according to a preset identity identifier of the speaker to obtain initial sound data carrying the preset identity identifier; determining preset sound data of at least one speaker corresponding to the video conference according to the initial sound data carrying the preset identity identifier.
6. The video conferencing method of claim 5, wherein, After the determining preset sound data of at least one speaker corresponding to the video conference according to the initial sound data carrying the preset identity identifier, the video conference method further comprises: converting the preset sound data into preset text according to the preset identity identifier to obtain conference record information of the video conference.
7. The video conferencing method of claim 2, wherein, Before the determining target sound data from preset sound data of at least one speaker corresponding to the video conference according to the visual line focus point, the video conference method further comprises: obtaining initial sound data; when a sound emitting orientation of the initial sound data is within a preset sound emitting range corresponding to a speaker included in the video picture and / or voiceprint information of the initial sound data is consistent with preset voiceprint information of the speaker, processing the initial sound data according to a preset identity identifier of the speaker to obtain initial sound data carrying the preset identity identifier; According to the initial sound data carrying the preset identity, the preset sound data of at least one speaker corresponding to the video conference is determined.
8. The video conferencing method of claim 7, wherein, After the initial sound data carrying the preset identity is obtained, the video conference method further comprises: According to the preset identity, the preset sound data is converted into preset text to obtain conference record information of the video conference.
9. The video conferencing method of claim 3, wherein, Before the target sound data is determined from the preset sound data of at least one speaker corresponding to the video conference according to the line-of-sight focus, the video conference method further comprises: Obtaining initial sound data; When the sound emitting orientation of the initial sound data is within the preset sound emitting range corresponding to one of the speakers included in the video picture, and / or the voiceprint information of the initial sound data matches the preset voiceprint information of the speaker, the initial sound data is processed according to the preset identity of the speaker to obtain initial sound data carrying the preset identity. According to the initial sound data carrying the preset identity, the preset sound data of at least one speaker corresponding to the video conference is determined.
10. The video conferencing method of claim 9, wherein, After the initial sound data carrying the preset identity is obtained, the video conference method further comprises: According to the preset identity, the preset sound data is converted into preset text to obtain conference record information of the video conference.
11. The video conferencing method of claim 4, wherein, Before the target sound data is determined from the preset sound data of at least one speaker corresponding to the video conference according to the line-of-sight focus, the video conference method further comprises: Obtaining initial sound data; When the sound emitting orientation of the initial sound data is within the preset sound emitting range corresponding to one of the speakers included in the video picture, and / or the voiceprint information of the initial sound data matches the preset voiceprint information of the speaker, the initial sound data is processed according to the preset identity of the speaker to obtain initial sound data carrying the preset identity. According to the initial sound data carrying the preset identity, the preset sound data of at least one speaker corresponding to the video conference is determined.
12. The video conferencing method of claim 11, wherein, After the initial sound data carrying the preset identity is obtained, the video conference method further comprises: According to the preset identity, the preset sound data is converted into preset text to obtain conference record information of the video conference.
13. The video conferencing method of claim 1, wherein, The video conference method further comprises: When the target sound data is output, the target text information corresponding to the target sound data is displayed.
14. The video conferencing method of claim 2, wherein, The video conference method further comprises: When the target sound data is output, the target text information corresponding to the target sound data is displayed.
15. The video conferencing method of claim 3, wherein, The video conference method further comprises: When the target sound data is output, the target text information corresponding to the target sound data is displayed.
16. The video conferencing method of claim 4, wherein, The video conference method further comprises: When the target sound data is output, target text information corresponding to the target sound data is displayed. 17.A video conference device, comprising: a picture display module configured to display a video picture of a video conference, the video picture comprising at least one speaker; a focus obtaining module configured to obtain a line-of-sight focus of a user wearing a near-eye display device on the video picture; a target determining module configured to determine target sound data from preset sound data of at least one speaker corresponding to the video conference according to the line-of-sight focus, and output the target sound data.
18. The video conferencing apparatus of claim 17, wherein, The target determining module comprises a speaker determining sub-module and a sound data determining sub-module. The speaker determining sub-module is configured to determine a target speaker from the speakers included in the video picture according to the line-of-sight focus. The sound data determining sub-module is configured to determine that the preset sound data in which a sound emitting direction is within a target sound emitting range corresponding to the target speaker and / or voiceprint information is consistent with target voiceprint information of the target speaker is the target sound data. 19.A near-eye display device, comprising a memory and a processor; the memory is configured to store a computer program; the processor is configured to execute the computer program and implement the following steps when the computer program is executed: display a video picture of a video conference, the video picture comprising at least one speaker; obtain a line-of-sight focus of a user wearing a near-eye display device on the video picture; determine target sound data from preset sound data of at least one speaker corresponding to the video conference according to the line-of-sight focus, and output the target sound data. 20.A computer readable storage medium, having a computer program stored thereon, the computer program, when executed by a processor, implements the following steps: display a video picture of a video conference, the video picture comprising at least one speaker; obtain a line-of-sight focus of a user wearing a near-eye display device on the video picture; determine target sound data from preset sound data of at least one speaker corresponding to the video conference according to the line-of-sight focus, and output the target sound data.
Citation Information
Patent Citations
Video conference picture adjusting method
CN113473066A
Video conference picture adjusting method and device, electronic equipment and medium
CN116614598A
Video conference method and device, equipment and storage medium
CN119182879A
Picture display device provided with line-of-sight detecting function
JP1996076289A
Organic conversations in a virtual group setting
US20230308501A1