Conference content display method and device based on wireless earphone and wireless earphone

By using wireless earphones to convert online meeting audio into text and display summary information, the problem of users having difficulty understanding meeting content in noisy environments or with hearing impairments is solved, resulting in a better meeting experience.

CN116095552BActive Publication Date: 2026-02-03NEWLINE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211634618.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2026-02-03
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

In online meetings, users often struggle to understand audio information due to noisy environments or hearing impairments, resulting in poor meeting outcomes.

Method used

The system converts conference audio into text using wireless earphones and displays a summary of the meeting on the earphone storage device's screen, optimizing the information display using a speech-to-text model and facial recognition technology.

Benefits of technology

In noisy environments or when hearing impairments occur, users can better understand the content of online meetings, thus improving the meeting experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116095552B_ABST
    Figure CN116095552B_ABST
Patent Text Reader

Abstract

The present disclosure provides a conference content display method and device based on a wireless earphone and a wireless earphone, the wireless earphone comprising an earphone body and an earphone storage device, the earphone body and the earphone storage device being connected through a wireless communication network, and a display screen being arranged on the earphone storage device, the method being implemented through the earphone storage device, and the method comprising: receiving conference voice information sent by the earphone body; converting the conference voice information into conference text information; generating conference summary information of the conference text information, and controlling the display screen on the earphone storage device to play the conference summary information. Through the present disclosure, the conference voice information can be converted into conference text information, and conference summary information can be extracted to be displayed on the display screen of the earphone storage device, so that when a user is inconvenient to listen to voice, is in an environment with relatively noisy sound, or has hearing impairment, the user can understand the information in the voice in an online conference, and the participation effect is better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, and wireless earphone for displaying meeting content based on a wireless earphone. Background Technology

[0002] This section is intended to provide background or context for the embodiments of this disclosure as set forth in the claims. The description herein is not intended to be a prior art simply because it is included in this section.

[0003] With the development of science and technology, online meetings have gradually replaced some offline meetings and are now common in work and study.

[0004] In online meetings, projects to be discussed are typically presented on the terminal screen in the form of presentation slides such as PowerPoint, while communication is conducted via voice.

[0005] However, if users are unable to listen to audio, are in noisy environments, or have hearing impairments, they may not be able to fully understand the information in the audio during online meetings, resulting in a poor meeting experience. Summary of the Invention

[0006] In view of this, the purpose of this disclosure is to provide a method, device and wireless earphone for displaying meeting content based on wireless earphones.

[0007] To achieve the above objectives, this exemplary embodiment provides a method for displaying meeting content based on a wireless headset. The wireless headset includes a headset body and a headset storage device. The headset body and the headset storage device can be connected via a wireless communication network. The headset storage device is equipped with a display screen. The method is implemented through the headset storage device and includes:

[0008] Receive conference voice information sent by the headset itself;

[0009] Convert the conference audio information into conference text information;

[0010] Generate a meeting summary of the meeting text information, and control the display screen on the headphone storage device to play the meeting summary information.

[0011] Based on the same inventive concept, an exemplary embodiment of this disclosure also provides a meeting content display device based on a wireless headset. The wireless headset includes a headset body and a headset storage device. The headset body and the headset storage device can be connected via a wireless communication network. The headset storage device is provided with a display screen. The device is implemented through the headset storage device. The device includes:

[0012] The conference voice information acquisition device is configured to receive conference voice information sent by the headset body;

[0013] The conference text information conversion device is configured to convert the conference audio information into conference text information;

[0014] The meeting summary information display device is configured to generate meeting summary information of the meeting text information and control the display screen on the headphone storage device to play the meeting summary information.

[0015] Based on the same inventive concept, an exemplary embodiment of this disclosure also provides a wireless earphone, including an earphone body and an earphone storage device, wherein the earphone body and the earphone storage device can be connected via a wireless communication network, and the earphone storage device is provided with a display screen;

[0016] The headphone storage device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method described in any of the above descriptions.

[0017] As can be seen from the above description, the present disclosure provides a method, apparatus, and wireless earphone for displaying meeting content based on a wireless earphone. The wireless earphone includes an earphone body and an earphone storage device. The earphone body and the earphone storage device can be connected wirelessly via a communication network. The earphone storage device is equipped with a display screen. The method is implemented through the earphone storage device and includes: receiving meeting voice information sent by the earphone body; converting the meeting voice information into meeting text information; generating meeting summary information from the meeting text information; and controlling the display screen on the earphone storage device to play the meeting summary information. Through this disclosure, meeting voice information can be converted into meeting text information, and meeting summary information can be extracted and displayed on the display screen of the earphone storage device. When users have difficulty listening to the voice, are in a noisy environment, or have hearing impairments, they can still understand the information in the voice during online meetings, resulting in a better meeting experience. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A schematic diagram illustrating an application scenario of the meeting content display method based on wireless headphones provided in this embodiment of the disclosure;

[0020] Figure 2 A flowchart illustrating a meeting content display method based on wireless headphones provided in this embodiment of the present disclosure;

[0021] Figure 3 A schematic diagram of a meeting content display device based on a wireless headset provided in an embodiment of this disclosure;

[0022] Figure 4 This is a more specific hardware structure diagram of the headphone storage device provided in an embodiment of the present disclosure. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this disclosure clearer, the principles and spirit of this disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0024] In this article, it is important to understand that any number of elements in the accompanying figures is for illustrative purposes and not for limitation, and any naming is for distinction only and has no limiting meaning.

[0025] The principles and spirit of this disclosure will be explained in detail below with reference to several representative embodiments.

[0026] As described in the background section, in related technologies, online meetings typically involve displaying the items to be discussed on a terminal screen in the form of presentation slides such as PowerPoint, while communication is conducted via voice.

[0027] However, the inventors of this disclosure have found that if users are unable to listen to audio, are in noisy environments, or have hearing impairments, they cannot properly understand the information in the audio during online meetings, resulting in a poor meeting experience.

[0028] To address the aforementioned issues, this disclosure provides a meeting content display solution based on wireless earphones. Specifically, the wireless earphones include an earphone body and an earphone storage device. The earphone body and the earphone storage device can be connected wirelessly via a communication network. The earphone storage device is equipped with a display screen. The method is implemented through the earphone storage device and includes: receiving meeting audio information sent by the earphone body; converting the meeting audio information into meeting text information; generating meeting summary information from the meeting text information; and controlling the display screen on the earphone storage device to play the meeting summary information. This disclosure allows meeting audio information to be converted into meeting text information and meeting summary information to be extracted and displayed on the display screen of the earphone storage device. When users have difficulty listening to audio, are in noisy environments, or have hearing impairments, they can still understand the information in the audio during online meetings, resulting in a better meeting experience.

[0029] After introducing the basic principles of this disclosure, various non-limiting embodiments of this disclosure will be described in detail below.

[0030] refer to Figure 1 This is a schematic diagram of an application scenario of the meeting content display method based on wireless headphones provided in the embodiments of this disclosure.

[0031] This application scenario includes an earphone storage device 110, earphones 120, and a meeting terminal 130. The earphone storage device 110, earphones 120, and meeting terminal 130 can all be connected wirelessly via a communication network. The meeting terminal 130 includes, but is not limited to, desktop computers, mobile phones, portable computers, tablets, media players, smart wearable devices, personal digital assistants (PDAs), or other electronic devices capable of performing the aforementioned functions.

[0032] The participating terminal 130 sends the conference audio information from the online conference to the headset 120, the headset 120 sends the conference audio information to the headset storage device 110, the headset storage device 110 converts the conference audio information into conference text information, generates conference summary information of the conference text information, and controls the display screen 112 to play the conference summary information.

[0033] The following is combined with Figure 1 The above application scenarios are used to describe the meeting content display method based on wireless earphones according to exemplary embodiments of this disclosure. It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of this disclosure, and the embodiments of this disclosure are not limited in any way. Rather, the embodiments of this disclosure can be applied to any applicable scenario.

[0034] refer to Figure 2 This is a flowchart illustrating a method for displaying meeting content based on wireless headphones, as provided in an embodiment of this disclosure.

[0035] The wireless earphone includes an earphone body and an earphone storage device. The earphone body and the earphone storage device can be connected wirelessly via a communication network. The earphone storage device is equipped with a display screen, and the meeting content display method based on the wireless earphone is implemented through the earphone storage device.

[0036] The method for presenting meeting content based on wireless headphones includes the following steps:

[0037] Step S210: Receive conference voice information sent by the headset body.

[0038] The wireless headset receives conference audio information from participating terminals through the headset body and sends the conference audio information to the headset storage device for processing.

[0039] In some exemplary embodiments, participating terminals include, but are not limited to, desktop computers, mobile phones, mobile computers, tablet computers, media players, smart wearable devices, personal digital assistants (PDAs), or other electronic devices capable of performing the above functions.

[0040] Users participate in online meetings through a conference terminal, using wireless headphones as an external device to receive conference audio information and act as a microphone.

[0041] The conference audio information may include the speech of the user currently speaking via voice, as well as the audio being played in the online conference attended by the participating terminals.

[0042] Step S220: Convert the conference audio information into conference text information.

[0043] This involves converting voice information into text information for display on the headphone storage device's screen.

[0044] In some exemplary embodiments, S220 specifically includes:

[0045] The conference audio information is input into a pre-trained speech-to-text model to obtain the conference text information corresponding to the conference audio information output by the speech-to-text model.

[0046] The speech-to-text model includes several speech-to-text networks, with different networks corresponding to different domains; the speech information of the meeting is converted into the text information of the meeting through at least one of the speech-to-text networks.

[0047] The different speech-to-text networks are trained using training conference speech information from different domains and training conference text information corresponding to the training conference speech information.

[0048] In some exemplary embodiments, the training method for the speech-to-text model includes:

[0049] Construct several sample sets, each containing a number of samples; wherein different sample sets target different domains; each sample includes: sample data and label data; the sample data includes training conference audio information; the label data includes training conference text information corresponding to the training conference audio information;

[0050] Based on each of the sample sets, the speech-to-text network corresponding to the sample set is constructed and trained using a predetermined machine learning algorithm to obtain the speech-to-text model.

[0051] The predetermined machine learning algorithm can be selected from one or more of the following: Naive Bayes algorithm, decision tree algorithm, support vector machine algorithm, kNN algorithm, neural network algorithm, deep learning algorithm, and logistic regression algorithm.

[0052] In this scenario, since different sample sets target different domains, the resulting speech-to-text networks are also designed for different domains. Furthermore, each speech-to-text network iterates through specialized corpora for its target domain, continuously improving its accuracy in converting speech to text for that domain.

[0053] In some exemplary embodiments, S220 specifically includes:

[0054] Determine the domain corresponding to the conference voice information;

[0055] The conference audio information is input into the speech-to-text network corresponding to the domain in the speech-to-text model, and the output of the speech-to-text network is determined as the conference text information output by the speech-to-text model.

[0056] Among them, domain-specific speech-to-text networks have higher accuracy.

[0057] In some exemplary embodiments, a domain determination page is provided on the display screen of the headphone storage device, through which the user can input or select from several options to obtain the domain of the current online meeting.

[0058] In some exemplary embodiments, before inputting the conference audio information into the speech-to-text network, the basic conference information corresponding to the conference audio information is first input into the domain determination network to obtain the domain corresponding to the conference audio information output by the domain determination network. Then, a speech-to-text network is selected according to the domain, and the conference audio information is input into the speech-to-text network to obtain the conference text information output by the speech-to-text network.

[0059] Optionally, basic meeting information may include at least one of the following: meeting name, meeting topic, meeting description, and participants.

[0060] In some exemplary embodiments, before inputting the conference audio information into the speech-to-text network, the conference audio information is first input into a domain determination network to obtain the domain corresponding to the conference audio information output by the domain determination network. Then, a speech-to-text network is selected according to the domain, and the conference audio information is input into the speech-to-text network to obtain the conference text information output by the speech-to-text network.

[0061] In some exemplary embodiments, the domain determines a method for training the network, including:

[0062] Construct a sample set comprising several samples; wherein, the samples include: sample data and label data; the sample data includes training conference audio information; the label data includes the training domain corresponding to the training conference audio information;

[0063] Based on the sample set, the domain determination network is constructed and trained using a predetermined machine learning algorithm.

[0064] The predetermined machine learning algorithm can be selected from one or more of the following: Naive Bayes algorithm, decision tree algorithm, support vector machine algorithm, kNN algorithm, neural network algorithm, deep learning algorithm, and logistic regression algorithm.

[0065] However, this disclosure is not limited to this. In some cases, such as when the domain features in the conference audio information are not obvious or even contain almost no domain features, the accuracy of the above-mentioned scheme of first determining the domain of the conference audio information and then processing the conference audio information through the speech-to-text network corresponding to that domain will decrease. To solve this problem, this disclosure provides the following solution:

[0066] In some exemplary embodiments, after obtaining the domain corresponding to the conference voice information output by the domain determination network, the method further includes:

[0067] Determine the confidence level of the domain and determine whether the confidence level is greater than a confidence threshold;

[0068] In response to determining that the confidence level is greater than a confidence threshold, the domain is identified as the domain corresponding to the conference voice information;

[0069] In response to determining that the confidence level is less than or equal to the confidence level threshold, the domain corresponding to the previous meeting's audio information in the same meeting is determined as the domain corresponding to the current meeting's audio information.

[0070] In response to the fact that the basic information of two meetings is the same, the two meetings are determined to be the same meeting.

[0071] Optionally, basic meeting information may include at least one of the following: meeting name, meeting topic, meeting description, and participants.

[0072] In this disclosure, considering that the fields corresponding to the audio information in the same meeting are basically the same or similar, and the accuracy of the currently determined field is low, the field corresponding to the audio information of the previous meeting in the same meeting can be used to ensure basic accuracy.

[0073] In some exemplary embodiments, the input of the speech-to-text model and the output of the previous speech-to-text network are used together as the input of the next speech-to-text network.

[0074] In view of the fact that a single speech-to-text network is difficult to meet the requirements of speech-to-text in complex domain dimensions, this disclosure provides a speech-to-text model with multiple speech-to-text networks stacked together. The input of the speech-to-text model and the output of the previous speech-to-text network are used as the input of the next speech-to-text network. This ensures that deep features of conference speech information can be extracted in multi-domain dimensions, and increases the nonlinear fitting capability of the speech-to-text model.

[0075] Step S230: Generate meeting summary information of the meeting text information, and control the display screen on the headphone storage device to play the meeting summary information.

[0076] In particular, considering the limited size of the display screen on the headphone storage device, if the display screen is controlled to play meeting text information, there will be problems with untimely playback when there is a lot of meeting text information. Therefore, this disclosure generates meeting summary information of meeting text information and controls the display screen to play meeting summary information, which helps users grasp the key information.

[0077] In some exemplary embodiments, the headphone storage device is also provided with a camera;

[0078] After generating the meeting summary information of the meeting text information, the method further includes:

[0079] The camera acquires video information and detects whether facial image information exists in the video information;

[0080] The control of the display screen on the headphone storage device to play the meeting summary information includes:

[0081] In response to determining that facial image information exists in the video information, the display screen on the headphone storage device is controlled to play the meeting summary information;

[0082] In response to determining that no facial image information is present in the video information, the display screen is controlled to pause playing the meeting summary information until, in response to determining that facial image information is present in the video information, the display screen is controlled to continue playing the meeting summary information.

[0083] Considering the limited size of the display screen on the headphone storage device, it is usually difficult to display all the meeting summary information on the screen at the same time. This disclosure provides a solution for scrolling through the meeting summary information. In this case, the timing of playback is obviously crucial; if the user is not watching during playback, they may only see a portion of the meeting summary information. To solve this problem, this disclosure uses facial recognition technology to play the meeting summary information when a face is detected and pause playback when no face is detected, thus ensuring, to a certain extent, the completeness of the meeting summary information seen by the user.

[0084] In some exemplary embodiments, after determining that facial image information exists in the video information, the method further includes:

[0085] Obtain eye image information from the facial image information;

[0086] A reference point and an observation point are determined in the eye image information, and based on the positional relationship between the reference point and the observation point, it is determined whether the gaze target in the face image information is the display screen;

[0087] The control of the display screen on the headphone storage device to play the meeting summary information includes:

[0088] In response to determining that the gaze target in the facial image information is the display screen, the display screen on the headphone storage device is controlled to play the meeting summary information;

[0089] In response to determining that the gaze target in the facial image information is not the display screen, the display screen is controlled to pause playing the meeting summary information until it is determined that the gaze target in the facial image information is the display screen, at which point the display screen is controlled to continue playing the meeting summary information.

[0090] In addition, considering that although the user is facing the display screen on the headphone storage device, their line of sight may not be on the display screen, this disclosure further provides a scheme to detect whether the user's line of sight is on the display screen, thereby further ensuring the integrity of the meeting summary information seen by the user.

[0091] In some exemplary embodiments, the reference point is the outer corner of the eye, and the observation point is the iris.

[0092] Methods for determining gaze targets in facial image information include:

[0093] The user's eye image is acquired using a visible light imaging module;

[0094] Determine the location of the inner corner of the eye and the location of the iris in the eye image;

[0095] Based on the positional relationship between the outer corner of the eye and the iris, the gaze target in the facial image information is determined.

[0096] In some exemplary embodiments, the reference point is the corneal reflection position, and the observation point is the pupil position;

[0097] Methods for determining gaze targets in facial image information include:

[0098] The user's eye image is acquired using an infrared imaging module;

[0099] Determine the corneal reflection position and pupil position in the eye image;

[0100] The gaze target in the face image information is determined based on the positional relationship between the corneal reflection position and the pupil position.

[0101] In some exemplary embodiments, after generating the meeting summary information of the meeting text information, the method further includes:

[0102] Detect whether the headphone storage device has moved;

[0103] The control of the display screen on the headphone storage device to play the meeting summary information includes:

[0104] In response to determining that the headphone storage device is not moving, the display screen on the headphone storage device is controlled to play the meeting summary information;

[0105] In response to determining that the headphone storage device has moved, the display screen is controlled to pause playing the meeting summary information until the headphone storage device stops moving, at which point the display screen is controlled to resume playing the meeting summary information.

[0106] In addition, considering that users may move the headphone storage device because they cannot see the meeting summary information displayed on the screen of the headphone storage device, the playback of the meeting summary information is paused when the user moves the headphone storage device, which further ensures the integrity of the meeting summary information seen by the user.

[0107] In some exemplary embodiments, methods for detecting whether the headphone storage device has moved, detecting faces, and detecting gaze targets can be used in combination to ensure the integrity of the meeting summary information seen by the user as much as possible.

[0108] In some exemplary embodiments, in response to a screen projection command for the meeting summary information, the meeting summary information is sent from the headphone storage device to the target terminal, and the display screen on the target terminal is controlled to play the meeting summary information.

[0109] Optionally, the target terminal can be a desktop computer, mobile phone, mobile computer, tablet computer, media player, smart wearable device, personal digital assistant (PDA) or other electronic device capable of performing the above functions.

[0110] This disclosure provides a screen projection solution as described in the above embodiments to facilitate users in sharing meeting summary information.

[0111] In some exemplary embodiments, in response to determining that the current online meeting session has ended, the meeting text information is summarized to obtain meeting records, and the meeting summary information is summarized to obtain meeting minutes.

[0112] This disclosure allows for the automatic generation of meeting minutes and summaries, saving the cost of generating them separately.

[0113] As can be seen from the above, the meeting content display method based on wireless earphones provided in this disclosure includes an earphone body and an earphone storage device. The earphone body and the earphone storage device can be connected wirelessly via a communication network. The earphone storage device is equipped with a display screen. The method is implemented through the earphone storage device and includes: receiving meeting voice information sent by the earphone body; converting the meeting voice information into meeting text information; generating meeting summary information of the meeting text information; and controlling the display screen on the earphone storage device to play the meeting summary information. This disclosure allows meeting voice information to be converted into meeting text information and meeting summary information to be extracted and displayed on the display screen of the earphone storage device. When users have difficulty listening to the voice, are in a noisy environment, or have hearing impairments, they can still understand the information in the voice during online meetings, resulting in a better meeting experience.

[0114] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.

[0115] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0116] Based on the same inventive concept, corresponding to any of the above embodiments, this disclosure also provides a conference content display device based on wireless headphones.

[0117] The wireless earphone includes an earphone body and an earphone storage device. The earphone body and the earphone storage device can be connected via a wireless communication network. The earphone storage device is equipped with a display screen. The device is implemented through the earphone storage device.

[0118] refer to Figure 3 The meeting content display device based on wireless headphones includes:

[0119] The conference voice information acquisition device 310 is configured to receive conference voice information sent by the headset body;

[0120] The conference text information conversion device 320 is configured to convert the conference audio information into conference text information;

[0121] The meeting summary information display device 330 is configured to generate meeting summary information of the meeting text information and control the display screen on the headphone storage device to play the meeting summary information.

[0122] In some exemplary embodiments, the meeting text information conversion device 320 is configured to:

[0123] The conference audio information is input into a pre-trained speech-to-text model to obtain the conference text information corresponding to the conference audio information output by the speech-to-text model.

[0124] The speech-to-text model includes several speech-to-text networks, with different networks corresponding to different domains; the speech information of the meeting is converted into the text information of the meeting through at least one of the speech-to-text networks.

[0125] In some exemplary embodiments, the speech-to-text model further includes a domain determination network;

[0126] The meeting text information conversion device 320 is configured as follows:

[0127] The conference audio information is input into the domain determination network to obtain the domain corresponding to the conference audio information output by the domain determination network;

[0128] The conference audio information is input into the speech-to-text network corresponding to the domain, and the conference text information corresponding to the conference audio information output by the speech-to-text network is obtained.

[0129] In some exemplary embodiments, the meeting text information conversion device 320 is configured to:

[0130] Determine the confidence level of the domain and determine whether the confidence level is greater than a confidence threshold;

[0131] In response to determining that the confidence level is greater than a confidence threshold, the domain is identified as the domain corresponding to the conference voice information;

[0132] In response to determining that the confidence level is less than or equal to the confidence level threshold, the domain corresponding to the previous meeting's audio information in the same meeting is determined as the domain corresponding to the meeting's audio information.

[0133] In some exemplary embodiments, the meeting text information conversion device 320 is configured to:

[0134] The input of the speech-to-text model and the output of the previous speech-to-text network are used together as the input of the next speech-to-text network.

[0135] In some exemplary embodiments, the headphone storage device is also provided with a camera;

[0136] The meeting summary information display device 330 is configured as follows:

[0137] The camera acquires video information and detects whether facial image information exists in the video information;

[0138] In response to determining that facial image information exists in the video information, the display screen on the headphone storage device is controlled to play the meeting summary information;

[0139] In response to determining that no facial image information is present in the video information, the display screen is controlled to pause playing the meeting summary information until, in response to determining that facial image information is present in the video information, the display screen is controlled to continue playing the meeting summary information.

[0140] In some exemplary embodiments, the meeting summary information display device 330 is configured to:

[0141] Obtain eye image information from the facial image information;

[0142] A reference point and an observation point are determined in the eye image information, and based on the positional relationship between the reference point and the observation point, it is determined whether the gaze target in the face image information is the display screen;

[0143] In response to determining that the gaze target in the facial image information is the display screen, the display screen on the headphone storage device is controlled to play the meeting summary information;

[0144] In response to determining that the gaze target in the facial image information is not the display screen, the display screen is controlled to pause playing the meeting summary information until it is determined that the gaze target in the facial image information is the display screen, at which point the display screen is controlled to continue playing the meeting summary information.

[0145] In some exemplary embodiments, the meeting summary information display device 330 is configured to:

[0146] Detect whether the headphone storage device has moved;

[0147] In response to determining that the headphone storage device is not moving, the display screen on the headphone storage device is controlled to play the meeting summary information;

[0148] In response to determining that the headphone storage device has moved, the display screen is controlled to pause playing the meeting summary information until the headphone storage device stops moving, at which point the display screen is controlled to resume playing the meeting summary information.

[0149] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0150] The apparatus described above is used to implement the corresponding meeting content display method based on wireless headphones in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0151] Based on the same inventive concept, corresponding to any of the above embodiments, this disclosure also provides a wireless earphone, including an earphone body and an earphone storage device, wherein the earphone body and the earphone storage device can be connected through a wireless communication network, and the earphone storage device is provided with a display screen;

[0152] The earphone storage device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the meeting content display method based on wireless earphones as described in any of the above embodiments.

[0153] Figure 4 This embodiment illustrates a more specific hardware structure diagram of an earphone storage device. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0154] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0155] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0156] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0157] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0158] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0159] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0160] The electronic devices described above are used to implement the corresponding meeting content display method based on wireless headphones in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0161] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, which are used to cause the earphone storage device to execute the meeting content display method based on wireless earphones as described in any of the above embodiments when the earphone storage device is used as a computer.

[0162] The aforementioned non-transitory computer-readable storage media can be any available medium or data storage device that a computer can access, including but not limited to magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MOs), etc.), optical storage (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs)).

[0163] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the meeting content display method based on wireless headphones as described in any of the embodiments in the exemplary method section above, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0164] Those skilled in the art will recognize that embodiments of this disclosure can be implemented as a system, method, or computer program product. Therefore, this disclosure can be implemented as entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this disclosure can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.

[0165] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example,, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (not exhaustive) of a computer-readable storage medium may include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0166] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0167] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0168] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0169] It should be understood that each block of a flowchart and / or block diagram, as well as combinations of blocks in a flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine that, when executed by a computer or other programmable data processing device, creates means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0170] These computer program instructions may also be stored in a computer-readable medium that enables a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce a product comprising an instruction apparatus that implements the functions / operations specified in the boxes of a flowchart and / or block diagram.

[0171] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions that execute on the computer or other programmable apparatus can provide a process for implementing the functions / operations specified in the boxes of a flowchart and / or block diagram.

[0172] Furthermore, although the operations of the methods of this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Rather, the steps depicted in the flowcharts may be executed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0173] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0174] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be interpreted in the broadest sense, thereby encompassing all such modifications and equivalent structures and functions.

Claims

1. A method for displaying meeting content based on wireless headphones, characterized in that, The wireless earphone includes an earphone body and an earphone storage device. The earphone body and the earphone storage device are connected via a wireless communication network. The earphone storage device is equipped with a display screen and a camera. The method is implemented through the earphone storage device, and the method includes: Receive conference voice information sent by the headset itself; Convert the conference audio information into conference text information; The system generates a meeting summary of the meeting text information, acquires video information through the camera, and detects whether there is a facial image in the video information; in response to determining that there is a facial image in the video information, it controls the display screen on the headphone storage device to play the meeting summary information; in response to determining that there is no facial image in the video information, it controls the display screen to pause playing the meeting summary information until it determines that there is a facial image in the video information, at which point it controls the display screen to continue playing the meeting summary information.

2. The method according to claim 1, characterized in that, The step of converting the conference audio information into conference text information includes: The conference audio information is input into a pre-trained speech-to-text model to obtain the conference text information corresponding to the conference audio information output by the speech-to-text model. The speech-to-text model includes several speech-to-text networks, with different networks corresponding to different domains; the speech information of the meeting is converted into the text information of the meeting through at least one of the speech-to-text networks.

3. The method according to claim 2, characterized in that, The speech-to-text model also includes a domain determination network; The step of inputting the conference audio information into a pre-trained speech-to-text model to obtain the conference text information corresponding to the conference audio information output by the speech-to-text model includes: The conference audio information is input into the domain determination network to obtain the domain corresponding to the conference audio information output by the domain determination network; The conference audio information is input into the speech-to-text network corresponding to the domain, and the conference text information corresponding to the conference audio information output by the speech-to-text network is obtained.

4. The method according to claim 3, characterized in that, After inputting the conference audio information into the domain identification network to obtain the domain corresponding to the conference audio information output by the domain identification network, the method further includes: Determine the confidence level of the domain and determine whether the confidence level is greater than a confidence threshold; In response to determining that the confidence level is greater than a confidence threshold, the domain is identified as the domain corresponding to the conference voice information; In response to determining that the confidence level is less than or equal to the confidence level threshold, the domain corresponding to the previous meeting's audio information in the same meeting is determined as the domain corresponding to the meeting's audio information.

5. The method according to claim 2, characterized in that, The input of the speech-to-text model and the output of the previous speech-to-text network are used together as the input of the next speech-to-text network.

6. The method according to claim 1, characterized in that, After determining that facial image information exists in the video information, the method further includes: Obtain eye image information from the facial image information; A reference point and an observation point are determined in the eye image information, and based on the positional relationship between the reference point and the observation point, it is determined whether the gaze target in the face image information is the display screen; The control of the display screen on the headphone storage device to play the meeting summary information includes: In response to determining that the gaze target in the facial image information is the display screen, the display screen on the headphone storage device is controlled to play the meeting summary information; In response to determining that the gaze target in the facial image information is not the display screen, the display screen is controlled to pause playing the meeting summary information until it is determined that the gaze target in the facial image information is the display screen, at which point the display screen is controlled to continue playing the meeting summary information.

7. The method according to claim 1, characterized in that, After generating the meeting summary information of the meeting text information, the method further includes: Detect whether the headphone storage device has moved; The control of the display screen on the headphone storage device to play the meeting summary information includes: In response to determining that the headphone storage device is not moving, the display screen on the headphone storage device is controlled to play the meeting summary information; In response to determining that the headphone storage device has moved, the display screen is controlled to pause playing the meeting summary information until the headphone storage device stops moving, at which point the display screen is controlled to resume playing the meeting summary information.

8. A conference content display device based on wireless headphones, characterized in that, The wireless earphone includes an earphone body and an earphone storage device. The earphone body and the earphone storage device are connected via a wireless communication network. The earphone storage device is equipped with a display screen and a camera. The device is implemented through the earphone storage device and includes: The conference voice information acquisition device is configured to receive conference voice information sent by the headset body; The conference text information conversion device is configured to convert the conference audio information into conference text information; A meeting summary information display device is configured to generate meeting summary information of the meeting text information, acquire video information through the camera, and detect whether facial image information exists in the video information; in response to determining that facial image information exists in the video information, control the display screen on the headphone storage device to play the meeting summary information; in response to determining that facial image information does not exist in the video information, control the display screen to pause playing the meeting summary information until in response to determining that facial image information exists in the video information, control the display screen to continue playing the meeting summary information.

9. A wireless earphone, characterized in that, The device includes an earphone body and an earphone storage device. The earphone body and the earphone storage device are connected via a wireless communication network. The earphone storage device is equipped with a display screen. The headphone storage device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Translation method and device, earphone and earphone storage device

    CN111696554A