Information display method and device
By displaying the user ID in the voice recognition interface and combining the user ID in the voice recognition interface, the problem of low discrimination efficiency of speakers caused by similar tone in the prior art is solved, and the rapid and accurate distinction of the speaker's position is achieved.
Patent Information
- Application Number
- CN202510486125.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-24
AI Technical Summary
When distinguishing voice data from different users, it is difficult to accurately distinguish when tone similarity occurs when electronic devices are less efficient in distinguishing speakers in multi-person meetings or social gatherings.
By displaying N user identifiers in the voice recognition interface, the display position of the user identifier corresponds to the actual position, and when the voice of the speaker is collected, the text converted from the voice text and the user identifier are displayed, and the user identifier of the voice recognition interface is quickly distinguished.
Improve the efficiency of electronic devices to distinguish speakers in multi-person meetings or social gatherings, ensuring that hearing-impaired users can quickly understand the speech content of different speakers.
Smart Images

Figure CN120199253A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of electronic devices, and particularly relates to an information display method and apparatus Background Art
[0002] As more and more hearing-impaired people participate in modern social work, when hearing-impaired people are in a meeting, it is difficult for them to obtain external voice information in a timely and effective manner, which brings great inconvenience to their lives and work.
[0003] Generally, in the scenario of a multi-person offline meeting or social gathering, a hearing-impaired user can collect voice data emitted by different users in the current scenario through an electronic device. Then, the different voice data are distinguished by timbre to associate the different voice data with the corresponding users. Then, the electronic device performs text conversion on these voice data, and arranges and displays the converted text on the display screen of the electronic device according to different timbres, so that the hearing-impaired user can know the speech content of different speakers in the current scenario by viewing the text.
[0004] However, the current algorithm for timbre discrimination is not yet mature. If there are similar timbres among different users, when the electronic device distinguishes the voice data of different users, it may process the voice data of different users as the voice data of the same user, thus unable to accurately distinguish the speakers corresponding to different voice data, and still requires the user to determine the speaker corresponding to the voice data by himself, resulting in a low efficiency of the electronic device in distinguishing speakers. Summary of the Invention
[0005] The purpose of the embodiments of this application is to provide an information display method and apparatus, which can improve the efficiency of an electronic device in distinguishing speakers.
[0006] In a first aspect, the embodiments of this application provide an information display method, which includes: displaying a voice recognition interface, where the voice recognition interface includes N user identifiers, and the display positions of the N user identifiers correspond to the actual positions among the N users indicated by the N user identifiers; when the first voice of the first speaker is collected, displaying the first text and the first user identifier on the voice recognition interface, where the first text is the text obtained by performing text conversion on the first voice, the first speaker is one of the N users, and the first user identifier is the user identifier corresponding to the first speaker among the N user identifiers; where N is an integer greater than 1; the user identifier includes at least one of the following: the seat identifier of the user's seat, the user nickname, and the user avatar.
[0007] Second aspect, an embodiment of the present application provides an information display device, which includes a display module; the display module is configured to display a voice recognition interface, the voice recognition interface includes N user identifiers, and the display positions of the N user identifiers correspond to the actual positions among the N users indicated by the N user identifiers; the above display module is further configured to, when the first voice of the first speaker is collected, display the first text and the first user identifier on the voice recognition interface, the first text is the text obtained by text conversion of the first voice, the first speaker is one of the N users, and the first user identifier is the user identifier corresponding to the first speaker among the N user identifiers; where N is an integer greater than 1; the user identifier includes at least one of the following: the seat identifier of the user's seat, the user nickname, and the user avatar.
[0008] Third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, it implements the steps of the method described in the first aspect.
[0009] Fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, it implements the steps of the method described in the first aspect.
[0010] Fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.
[0011] Sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium, and the program product is executed by at least one processor to implement the method described in the first aspect.
[0012] In an embodiment of the present application, a voice recognition interface is displayed. The voice recognition interface includes N user identifiers, and the display positions of the N user identifiers correspond to the actual positions among the N users indicated by the N user identifiers. When the first voice of the first speaker is collected, the first text and the first user identifier are displayed on the voice recognition interface. The first text is the text obtained by converting the first voice text, the first speaker is one of the N users, and the first user identifier is the user identifier corresponding to the first speaker among the N user identifiers. Here, N is an integer greater than 1. The user identifier includes at least one of the following: the seat identifier of the user's seat, the user nickname, and the user avatar. In this solution, the electronic device intuitively shows the actual positions among the N users to the user by displaying the N user identifiers in the voice recognition interface. At the same time, when the electronic device collects the voice of the speaker, the text after converting the first voice text can be associated and displayed with the user identifier on the voice recognition interface. Combining with the N user identifiers in the voice recognition interface, the orientation of the speaker can be quickly distinguished, thereby improving the efficiency of distinguishing the speaker. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a schematic flowchart of an information display method provided by an embodiment of the present application;
[0014] Figure 2 is a schematic display diagram of a seat relationship diagram provided by an embodiment of the present application;
[0015] Figure 3 is a schematic display diagram of a voice recognition interface provided by an embodiment of the present application;
[0016] Figure 4 is a schematic display diagram in which a pointing arrow points to a user identifier provided by an embodiment of the present application;
[0017] Figure 5 is a schematic display order diagram of text and user identifiers provided by an embodiment of the present application;
[0018] Figure 6 is a schematic interface diagram of user identifier input of a second electronic device provided by an embodiment of the present application;
[0019] Figure 7 is a schematic structural diagram of an information display device provided by an embodiment of the present application;
[0020] Figure 8 is a schematic structural diagram of an information display device provided by an embodiment of the present application;
[0021] Figure 9 is a schematic structural diagram of an information display device provided by an embodiment of the present application;
[0022] Figure 10 One of the schematic diagrams of the hardware structure of an electronic device provided by an embodiment of the present application;
[0023] Figure 11 Another one of the schematic diagrams of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.
[0025] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.
[0026] The terms "at least one (item)", "at least one of", etc. in the specification and claims of the present application refer to any one, any two or more combinations of the objects they contain. For example, at least one (item) of a, b, and c can represent: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two (items)" means two or more, and its expressed meaning is similar to that of "at least one (item)".
[0027] The identifiers in the present application are used to indicate information such as characters, symbols, images, etc., and can use identifiers or other containers as carriers for displaying information, including but not limited to text identifiers, image identifiers, symbol identifiers, etc.
[0028] It should be noted that for the information display method provided by the embodiments of the present application, the execution subject can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted electronic device, etc. In some embodiments of the present application, taking the electronic device as the execution subject to execute the information display method as an example, the information display method provided by the embodiments of the present application is described.
[0029] The information display method, device, electronic device, and readable storage medium provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.
[0030] The information display method provided by the embodiments of the present application can be applied to scenarios where hearing-impaired users participate in offline multi-person meetings or social gatherings.
[0031] Generally, in offline multi-person meetings and social gatherings, hearing-impaired users obtain the speech content of the speakers in the current environment by watching the text converted from the speech data collected by the electronic device. However, at this time, most of the attention of the hearing-impaired users is on the electronic device. Therefore, it is difficult for hearing-impaired users to quickly find the speaker corresponding to the text. Moreover, it is easy to miss some important information during the search process because they do not see the text on the electronic device, thus reducing the work efficiency of hearing-impaired users.
[0032] In the related art, the electronic device converts the collected speech data of the speaker into text and displays it for the hearing-impaired user to view. However, there may be a problem that the hearing-impaired user does not know which user said the text and cannot respond in time and cannot make a corresponding connection.
[0033] Although most hearing-impaired users can distinguish the speakers by themselves, they still cannot quickly perceive who said it. At this time, the hearing-impaired users still have to look around to align which user is speaking, which is likely to miss important information. The current method is to distinguish the characters in the speech data by tone color and display the text corresponding to the speech data and the user after tone color distinction on the display screen of the electronic device, so that the user knows which user said the current displayed text content. However, this method has the following problems:
[0034] Problem 1: Currently, the tone color algorithm is not yet mature. Therefore, the accuracy of tone color distinction is relatively low, and it is difficult to achieve the effect of accurately distinguishing characters.
[0035] Problem 2: If two users speak simultaneously, it is very likely to be recognized as the tone color of one user, equivalent to recognizing a new user. For example, there are originally 6 people in the meeting, but due to the speaking speed and tone problems, when the electronic device recognizes a certain sentence, it may recognize a 7th tone color, resulting in a 7th user, making it impossible for the hearing-impaired user to find the corresponding speaker. Another example is that when the 1st user starts speaking and the 2nd user suddenly interrupts in the middle, the electronic device may recognize the speech data of the 2nd user as the silver emitted by the 1st user, thus unable to distinguish different users, and further resulting in a poor user experience.
[0036] Problem 3: The recording ability of the electronic device is limited. If other speakers are located slightly farther away from the electronic device used by the hearing-impaired user, such as when the teacher is on the podium and the hearing-impaired user is sitting in the last row, it may be difficult to record clearly when using the electronic device for recording, and it is difficult to obtain satisfactory voice separation and text recognition effects.
[0037] The object of the invention of the embodiments of the present application is to create a scenario for the hearing-impaired population to quickly know the content of the speaker in a meeting communication or small gathering scenario, and know who is speaking, and accurately distinguish who the speaker is. When multiple speakers speak at the same time, they can also be distinguished. The specific implementation method is as follows:
[0038] Display a voice recognition interface, where the voice recognition interface includes N user identifiers, and the display positions of the N user identifiers correspond to the actual positions among the N users indicated by the N user identifiers; when the first voice of the first speaker is collected, display the first text and the first user identifier on the voice recognition interface, the first text is the text obtained by text conversion of the first voice, the first speaker is one of the N users, and the first user identifier is the user identifier corresponding to the first speaker among the N user identifiers; where N is an integer greater than 1; the user identifier includes at least one of the following: the seat identifier of the user's seat, the user nickname, and the user avatar. In this solution, the electronic device intuitively shows the actual positions among the N users by displaying the N user identifiers in the voice recognition interface. At the same time, when the electronic device collects the voice of the speaker, the text after text conversion of the first voice can be associated and displayed with the user identifier on the voice recognition interface, and combined with the N user identifiers in the voice recognition interface, the orientation of the speaker can be quickly distinguished, thereby improving the efficiency of distinguishing the speaker.
[0039] The execution subject of the information display method provided by the embodiments of the present application may be an information display device. Exemplarily, the information display device may be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. Hereinafter, the information display method provided by the embodiments of the present application will be exemplarily described by taking an electronic device as an example.
[0040] The embodiments of the present application provide an information display method, Figure 1 shows a flowchart of an information display method provided by the embodiments of the present application, and this method can be applied to a first electronic device. As Figure 1 shown, the information display method provided by the embodiments of the present application may include the following steps 201 and 202.
[0041] Step 201, the first electronic device displays a voice recognition interface.
[0042] In some embodiments of the present application, the above voice recognition interface includes N user identifiers, where N is an integer greater than 1.
[0043] In some embodiments of the present application, the display positions of the above N user identifiers correspond to the actual positions among the N users indicated by the N user identifiers.
[0044] In some embodiments of the present application, each of the above N user identifiers corresponds to a user.
[0045] In some embodiments of the present application, the user identifier includes at least one of the following: the seat identifier of the user's seat, the user nickname, and the user avatar.
[0046] Exemplarily, the seat identifier of the above user's seat can be displayed in the form of a digital number or an English number. For example, 1, 2 or one, two.
[0047] Exemplarily, the above user nickname can be a name customized by the user, the name of the electronic device used by the user, or the name of the user.
[0048] Exemplarily, the above user avatar can be customized by the user or an image of the user taken currently.
[0049] Exemplarily, the above N user identifiers can be displayed in the form of a seat relationship diagram.
[0050] Example 1, as Figure 2 shown, taking the scenario where a hearing-impaired user is in an offline multi-person meeting as an example, in the upper half area of the voice recognition interface 21, the avatar and name of the hearing-impaired user are displayed, and the avatars, names, and seat numbers of five users are displayed in the order of their actual positions. For example, User Zhang San with seat number 1, User Li Si with seat number 2, User Xiao Gao with seat number 3, User Chen with seat number 4, and User General Manager Wang with seat number 5.
[0051] Step 202: When the first electronic device collects the first voice of the first speaker, display the first text and the first user identifier on the voice recognition interface.
[0052] In some embodiments of the present application, the above first text is the text obtained by converting the first voice.
[0053] Exemplarily, the above text conversion can use a voice translation model to convert the voice into text and display it on the display screen of the electronic device.
[0054] In some embodiments of the present application, the above first speaker is one of the N users.
[0055] In some embodiments of the present application, the above-mentioned first user identifier is the user identifier corresponding to the first speaker among the N user identifiers.
[0056] In one example, the first electronic device can uniformly collect the voice emitted by the current speaker through the recording device of this electronic device, that is, the first voice of the first speaker.
[0057] In another example, each user has an electronic device with a recording function. When the speaker is speaking, the corresponding electronic device transmits the speech content as voice data to the electronic device used by the hearing-impaired user, that is, the above-mentioned first electronic device.
[0058] In a possible embodiment, the above-mentioned speech recognition interface includes a first display area and a second display area.
[0059] In some embodiments of the present application, the N user identifiers are displayed in the above-mentioned first display area, and specific reference can be made to the above Figure 2 shown content.
[0060] Further, in some embodiments of the present application, the above-mentioned step 202 can be implemented through the following step 202a.
[0061] Step 202a: The first electronic device displays the first text and the first user identifier in the second display area.
[0062] In some embodiments of the present application, the above-mentioned second display area can be any area other than the above-mentioned first display area.
[0063] In some embodiments of the present application, the above-mentioned second display area is used to associatively display the text after text conversion and the user identifier.
[0064] For example, "User No. 1 Zhang San: The meeting starts." is displayed in the above-mentioned second display area.
[0065] Example 2, in combination with Figure 2 , as Figure 3 shown, taking the scenario where the hearing-impaired user is in an offline multi-person meeting as an example, in the upper half area 31 of the speech recognition interface 21, that is, the above-mentioned first display area, N user identifiers are displayed. And in the lower half area 32 of the speech recognition interface 21, that is, the above-mentioned second display area, "User No. 1 Zhang San: The meeting starts." is displayed, that is, the above-mentioned first text and the first user identifier.
[0066] In this way, the electronic device can associatively display the seat relationship diagram of the scene and the user identifier corresponding to the text through two display areas, which can solve the problem in the current solution that only the name of the speaker is displayed, and the hearing-impaired user still has to figure out who said this sentence when trying to correspond to the person, and then go to the corresponding person. Thus, the hearing-impaired user can quickly find the contact corresponding to the text at the scene according to the association relationship between the user identifiers established by the two display areas.
[0067] In the information display method provided in the embodiment of the present application, a voice recognition interface is displayed. The voice recognition interface includes N user identifiers, and the display positions of the N user identifiers correspond to the actual positions among the N users indicated by the N user identifiers. When the first voice of the first speaker is collected, a first text and a first user identifier are displayed on the voice recognition interface. The first text is the text obtained by converting the first voice into text, the first speaker is one of the N users, and the first user identifier is the user identifier corresponding to the first speaker among the N user identifiers. Wherein, N is an integer greater than 1; the user identifier includes at least one of the following: the seat identifier of the seat where the user is located, the user nickname, the user avatar. In this solution, the electronic device intuitively shows the actual positions among the N users to the user by displaying the N user identifiers in the voice recognition interface. At the same time, when the electronic device collects the voice of the speaker, the text after converting the first voice text and the user identifier can be associatively displayed on the voice recognition interface, and combined with the N user identifiers in the voice recognition interface, the orientation of the speaker can be quickly distinguished, thereby improving the efficiency of distinguishing the speaker.
[0068] Optionally, in some embodiments of the present application, after the step 202 “the first electronic device displays the first text and the first user identifier on the voice recognition interface”, the information display method provided in the embodiment of the present application further includes step 301.
[0069] Step 301: The first electronic device displays the first user identifier in a preset display manner in the first display area.
[0070] In some embodiments of the present application, the above preset display manner may include but is not limited to at least one of the following: highlighting, flashing, magnifying.
[0071] Exemplarily, when the first speaker is the user corresponding to the first user identifier, the user identifier corresponding to Zhang San among the N user identifiers is displayed in a preset display manner.
[0072] For example, in combination with Example 1, when the first speaker is Zhang San No. 1, the user identifier corresponding to Zhang San No. 1 is highlighted with a brightness greater than that of other user identifiers.
[0073] For example, in combination with Example 1, when the first speaker is Zhang San with ID 1, the user identifier corresponding to Zhang San with ID 1 is flashed and displayed in an alternating light and dark form.
[0074] For example, in combination with Example 1, when the first speaker is Zhang San with ID 1, the user identifier corresponding to Zhang San with ID 1 is enlarged. For example, if the diameter of other user identifiers is 1 cm, the diameter of the user identifier corresponding to Zhang San with ID 1 can be enlarged to 2 cm.
[0075] In this way, the electronic device highlights the current speaker through the preset display method, enabling the hearing-impaired user to quickly find the position of the speaker corresponding to the current text.
[0076] Optionally, in some embodiments of the present application, after the above step 202 "the first electronic device displays the first text and the first user identifier on the speech recognition interface", the information display method provided by the embodiments of the present application further includes step 302.
[0077] Step 302: The first electronic device displays orientation prompt information in the first display area.
[0078] In the embodiments of the present application, the above orientation prompt information is used to prompt the orientation of the first speaker relative to the target user.
[0079] In some embodiments of the present application, the above target user is one of the N users.
[0080] In some embodiments of the present application, the above orientation prompt information includes at least one of the following: displaying distance information in text form, displaying direction information in text form, and indicating direction information in icon form.
[0081] Exemplarily, the above icon form indicating direction information can indicate the direction in the form of an arrow pointing.
[0082] It can be understood that the first display area can display all the above orientation prompt information, or any combination of the prompt information at the same time.
[0083] Example 3, in combination with Figure 3 , such as Figure 4 shown, taking the scenario where the hearing-impaired user is in an offline multi-person meeting as an example, in the upper half area 31 of the speech recognition interface 21, that is, the above first display area, N user identifiers are displayed. And in the lower half area 32 of the speech recognition interface 21, that is, the above second display area, "Zhang San with ID 1: The meeting starts." is displayed, that is, the above first text and the first user identifier. At this time, a pointing arrow 41 pointing to Zhang San is synchronously displayed in the first display area, that is, the above orientation prompt information.
[0084] In this way, the electronic device quickly prompts the user with the location information of the current speaker through simple and clear orientation prompt information, enabling the hearing-impaired user to quickly locate the corresponding speaker if they want to find the speaker at the scene.
[0085] Optionally, in some embodiments of the present application, before the above step 201 "the first electronic device displays the speech recognition interface", the information display method provided by the embodiments of the present application further includes step 401 and step 402.
[0086] Step 401: The first electronic device obtains the user location information.
[0087] In some embodiments of the present application, the above user location information is used to indicate the location relationship among N users.
[0088] In some embodiments of the present application, the above user location information includes at least one of the following: the location information of N electronic devices, the actual location picture among the N users.
[0089] Exemplarily, one of the above N electronic devices corresponds to one of the above N users. In other words, one electronic device belongs to one user.
[0090] In some embodiments of the present application, the location information of the above N electronic devices is the location information of N - 1 second electronic devices and the location information of the first electronic device.
[0091] Exemplarily, the location information of the above N - 1 second electronic devices and the location information of the first electronic device, that is, the location information of the above N electronic devices, may be that the first electronic device obtains the positioning information of the second electronic device, such as GPS positioning information.
[0092] Exemplarily, the actual location picture among the above N users may be hand-drawn by the hearing-impaired user after observing the scene, or may be a picture of the scene taken by the electronic device.
[0093] Exemplarily, the above method of "obtaining the user location information" may include obtaining from the picture hand-drawn by the hearing-impaired user after observing the scene, or obtaining from the picture taken of the scene, or obtaining from the positioning information of the location where the second electronic device is located.
[0094] Optionally, before the first electronic device obtains the user location information, other users except the hearing-impaired user can input their own information in the second electronic device, such as seat number, nickname, or avatar, etc.
[0095] Exemplarily, other users except the hearing-impaired users input their own information on the personal electronic device, which is made to correspond one by one with the numbers established by the hearing-impaired users in the first electronic device, that is, the seats are numbered synchronously starting from the hearing-impaired user in a clockwise or counterclockwise direction. Then, after the meeting is started, when the corresponding speaker is speaking, the second electronic device records the voice of the speaker and transmits the recorded voice to the first electronic device used by the hearing-impaired user.
[0096] It can be understood that when the second electronic device transmits the voice, it will transmit user identifiers, such as seat numbers and nicknames, together with the voice to the first electronic device, so that the first electronic device can clearly distinguish the user identifier corresponding to the speaker.
[0097] In this way, because the sound pickup process, that is, the above-mentioned voice recording process, is carried out on the electronic device of the hearing-normal user, the problem of poor long-distance sound pickup for hearing-impaired users can be solved. Moreover, the hearing-normal user manually inputs user identifiers, for example, numbers and names, which can most accurately indicate who the speaker bound to this voice is, solving the problem that the speaker cannot be accurately distinguished only by the algorithm of tone color, so that the speaker can be accurately distinguished.
[0098] Step 402: The first electronic device displays N user identifiers on the voice recognition interface based on the user location information.
[0099] Exemplarily, the first electronic device first generates corresponding user identifiers based on different user location information and displays the user identifiers on the voice recognition interface according to the actual positions among the users.
[0100] Taking the scenario where the hearing-impaired user is in an offline multi-person meeting as an example, the upper half area of the voice recognition interface is set as the area for displaying N user identifiers. Before the meeting starts, the hearing-impaired user can use the electronic device to take a photo of the scene including all the participants. Or, the hearing-impaired user can use the note graffiti function to hand-draw a simple picture to distinguish all the participants, write down the names by hand, and generate a seat relationship diagram according to the actual orientation of the seats, that is, the above-mentioned N user identifiers. In addition, the hearing-impaired user numbers each participant, for example, starting from the hearing-impaired user in a clockwise or counterclockwise direction. Such a final seat relationship diagram will always be displayed in the top area of the voice recognition interface and can be pointed at the speaker in real time with a pointer. In this way, the problem that only showing the names may cause the hearing-impaired user to wrongly associate the name of the speaker with the on-site participants can be solved, so that the hearing-impaired user can quickly react and find the corresponding speaker.
[0101] In this way, by obtaining the user's location information, a seating relationship diagram reflecting the actual positions among users is generated, so as to clearly show the seating relationship among users to the hearing-impaired users.
[0102] Optionally, in some embodiments of the present application, the information display method provided by the embodiments of the present application further includes steps 501 to 503.
[0103] Step 501: If the first electronic device receives the second speech of the second speaker within the acquisition duration of the first speech, the first speech is split into M speech segments.
[0104] In some embodiments of the present application, the above-mentioned second speaker is one of N users.
[0105] In some embodiments of the present application, the above-mentioned acquisition duration is the voice data within the duration from the start to the end of the speech of the first speaker.
[0106] Exemplarily, the above-mentioned "receiving the second speech of the second speaker within the acquisition duration of the first speech" can be understood as a scenario where the second speaker interrupts during the speech of the first speaker.
[0107] In one example, the electronic device can split the first speech into M speech segments according to the time points at the pauses in the first speech.
[0108] For example, the electronic device uses voice silence detection technology to break the voice of a speaker into individual speech segments. For example, when speaker 1 is speaking and the electronic device starts recording, the voice is continuous. However, when there is a pause during the speech, the first speech can be split into M speech segments according to the pause time points.
[0109] In another example, the electronic device can split the first speech into M speech segments according to a preset duration.
[0110] In another example, the electronic device can split the first speech according to the time coincidence point between the second speech and the first speech to obtain M speech segments.
[0111] Example 4: The first speech issued by Zhang San No. 1 within the acquisition duration is "Speech 1". The electronic device can split the first speech into: Speech Segment 1, Speech Segment 2, and Speech Segment 3. At the acquisition time point of Speech Segment 2, the electronic device also receives the second speech "Speech 2" issued by Li Si No. 2.
[0112] It should be noted that after the first electronic device splits the first speech, it binds the split speech segments with the user identifier corresponding to the first speech, so that each speech segment is associated with a user identifier.
[0113] Step 502: The first electronic device sorts the M voice segments and the second voice based on the acquisition timestamps corresponding to each voice segment and the acquisition timestamp of the second voice, and generates a voice queue.
[0114] In some embodiments of the present application, the above acquisition timestamp is the time point when the user emits the voice.
[0115] In some embodiments of the present application, the first voice and the second voice received by the first electronic device respectively carry corresponding acquisition timestamps.
[0116] In some embodiments of the present application, the first electronic device sorts according to the chronological order of the acquisition timestamps to generate a voice queue.
[0117] Example 5, combined with Example 4, the acquisition timestamp of Voice 1 is 9:00:00. Among them, the acquisition timestamp of Voice Segment 1 is 9:00:00, the acquisition timestamp of Voice Segment 2 is 9:00:05, the acquisition timestamp of Voice Segment 3 is 9:00:10, and the acquisition timestamp of Voice 2 is 9:00:03. Then the arrangement order of the generated voice queue is: Voice Segment 1, Voice 2, Voice Segment 2, Voice Segment 3.
[0118] It can be understood that after the hearing-impaired user edits the seat relationship diagram, voice monitoring is started, and voice data from each speaker is received. When parsing the voice data, according to a certain voice threshold, the start and end of speaking are checked, and the voices from different speakers are switched into individual voice segments and placed in the queue to be converted. The basis for the order of placement in the queue is the acquisition timestamp. The first electronic device establishes an association relationship between the acquisition timestamp and the user identifier and the above voice segments. Then, after the first electronic device converts the voice segments into text, the user identifier and the text are displayed together on the voice mouse interface.
[0119] Thus, on the one hand, since there is an association relationship between the user identifier and the voice segment, there is no need to distinguish different speakers by differentiating voices, that is, there is no need to access the voice differentiation algorithm, thus saving costs. On the other hand, based on the voice noise reduction technology, each user uses their own electronic device to record voice, which can eliminate the voices of others speaking from other directions as much as possible, and can improve the effect of converting voice into text. Since, after one person starts speaking and the second person interrupts, because the words of the first person have not yet entered the queue and been completed, the words of the second person can only wait to enter the queue again. And after the voice of the first person is noise-reduced, most of the voice part is their own voice. Similarly, for the second person, only their own voice is recorded, separating the voices of the two people at the source and adding them to the recognition queue in the order of priority, not only solving the problem of inaccurate speaker differentiation when multiple users speak simultaneously, but also solving the problem of low accuracy of the text content obtained by text conversion due to long distance and simultaneous speaking.
[0120] Step 503: The first electronic device displays M second texts and the first user identifier, as well as the third text and the second user identifier on the voice recognition interface according to the voice queue.
[0121] In some embodiments of the present application, the above-mentioned M second texts are texts obtained by converting M voice segments into text.
[0122] In some embodiments of the present application, the above-mentioned third text is a text obtained by converting the second voice into text.
[0123] Exemplarily, each second text corresponds to a first user identifier.
[0124] In some embodiments of the present application, the above-mentioned second user identifier is the identifier of the second speaker among the N user identifiers.
[0125] In some embodiments of the present application, the first electronic device performs text conversion according to the sorting of each voice segment in the voice queue to obtain the corresponding text, and binds each text to the corresponding user identifier to associate and display the text and the user identifier on the voice recognition interface.
[0126] Example 6, in combination with Example 5, refer to Figure 3 , such as Figure 5 shown, taking the scenario where a hearing-impaired user is in an offline multi-person meeting as an example, in the upper half area 31 of the voice recognition interface 21, that is, the above-mentioned first display area, N user identifiers are displayed. And in the lower half area 32 of the voice recognition interface 21, that is, the above-mentioned second display area, the following are displayed in sequence: Zhang San No. 1: Voice segment 1, Li Si No. 2: Voice 2, Zhang San No. 1: Voice segment 2, Zhang San No. 1: Voice segment 3.
[0127] In this way, in an interruption scenario or a scenario where multiple users speak simultaneously, the accuracy of the electronic device in distinguishing users can be improved.
[0128] The following uses specific examples to exemplarily illustrate the information display method provided by the embodiments of the present application.
[0129] Embodiment 1: Take the electronic device of a hearing-impaired user as the first electronic device, and the electronic devices of other normal users as the second electronic devices, and take the scenario of a multi-person offline meeting as an example.
[0130] The above information display method may include the following steps A1 to A4.
[0131] Step A1: The first electronic device displays a seat relationship diagram on the speech recognition interface.
[0132] Exemplarily, the upper half area of the speech recognition interface is set as the area for displaying N user identifiers. Before the meeting starts, the hearing-impaired user can use the electronic device to take a live photo of all the participants. Or, the hearing-impaired user can use the note doodle function to hand-draw a simple picture to distinguish all the participants, and write down the names by hand, generating a seat relationship diagram according to the actual seat orientation, that is, the above N user identifiers. In addition, the hearing-impaired user numbers each participant. For example, starting from the hearing-impaired user, numbering is carried out clockwise or counterclockwise. This final seat relationship diagram will always be displayed in the top area of the speech recognition interface, and the pointer can point to the speaker in real time. In this way, the problem that only showing the name may cause the hearing-impaired user to wrongly associate the name of the speaker with the on-site participants can be solved, so that the hearing-impaired user can quickly respond and find the corresponding speaker.
[0133] Step A2: When the hearing-normal user participates in the meeting, enter their seat number and name information in the application of the second electronic device.
[0134] It can be understood that the seat numbers entered by the above hearing-normal users need to correspond one by one with the numbers established by the hearing-impaired user, that is, synchronously numbering clockwise or counterclockwise starting from the hearing-impaired user.
[0135] Exemplarily, the second electronic device receives a touch input for the user to click to start the meeting control. In response to this touch input, after starting the meeting, the second electronic device starts the recording function of the second electronic device, synchronously collects the voice data of the speech content when the user starts speaking, and sends it to the first electronic device used by the hearing-impaired user in real time. It should be emphasized that when the voice data is transmitted, the seat number and name information previously input by the user will be transmitted to the first electronic device used by the hearing-impaired user together.
[0136] For example, as Figure 6 shown, user Zhang San inputs his seat number "1" and name information "Zhang San" in interface 61. After clicking the meeting start control 62, Zhang San's speech content will be transmitted to the hearing-impaired user as voice data. At the same time, the seat number and name will be transmitted together with the voice.
[0137] In this way, since each second electronic device corresponds to a hearing-normal user and the second electronic device is located in a relatively close position near the user, it can solve the problem of poor voice pickup caused by only using the electronic device used by the hearing-impaired user for long-distance voice pickup. At the same time, based on the noise reduction function, it is also possible to exclude the interference of the environment and the voices of other speakers as much as possible when the user is speaking. Moreover, the hearing-normal user manually inputs the seat number and name, which most accurately indicates who the speaker associated with this voice is, solving the problem that it is impossible to accurately distinguish only by the algorithmic voice timbre, thereby improving the accuracy of the electronic device in distinguishing the speaker corresponding to the voice data.
[0138] Step 103: After the hearing-impaired user edits the seat relationship diagram, start voice monitoring and receive voice data from each speaker.
[0139] Exemplarily, when the first electronic device analyzes the voice data, it checks the start and end of the voice according to a certain sound threshold, and cuts the different voices from different speakers into individual voice segments and puts them into the queue to be converted. The basis for the order of putting them into the queue is the acquisition timestamp.
[0140] Then, the first electronic device will establish a corresponding relationship between the acquisition timestamp and the user identifier corresponding to the voice source and the above-mentioned segmented voice segments. After the first electronic device converts the voice segment into text, it will display the user identifier and the text at the same time.
[0141] Specifically, the first electronic device uses the voice mute detection technology to break the voice of a speaker into individual voice segments. For example, when speaker 1 is speaking and the electronic device turns on the recording function, the voice is continuous at this time. However, when there is a pause during his speech, the first electronic device can break the voice into segments according to the pause time point.
[0142] For example, speaker 1, Zhang San, said two sentences. When transmitting the voice, the serial number 1 and Zhang San are transmitted at the same time. During the process of segmenting the voice segments, the segmented voice segments are associated with the serial number 1 and Zhang San. Similarly, speaker 3, Xiao Gao, also said two sentences, and the serial number 3 and Xiao Gao are transmitted to the first electronic device together with the voice. At this time, the first electronic device can also segment the voice spoken by Xiao Gao into voice segments.
[0143] Exemplarily, each segmented voice clip carries a corresponding acquisition timestamp. Therefore, on the side of the hearing-impaired user, the first electronic device can sequentially place the voice clips into the voice queue according to the order of the acquisition timestamps, and perform text conversion and display. Moreover, since each voice clip corresponds to a serial number and a name, in this way, the converted text can be accurately associated with the serial number and the person's name and displayed together.
[0144] Step 104: The first electronic device displays a pointing arrow in the first display area, pointing to the user identifier corresponding to the text and the user identifier displayed in the second display area.
[0145] Exemplarily, after the first electronic device adds the voice clips to the voice queue according to the acquisition timestamps of the voice clips, when displaying in the second display area, the serial number and name of the voice clip are bound and displayed together. At the same time, the serial number in the second display area is matched with the user identifier in the above first display area, and the position of the pointer is dynamically changed. In this way, the hearing-impaired user can quickly know who is speaking according to the position indicated by the pointer. In this way, not only can the problems of unclear recognition in the interjection scenario, incorrect role differentiation, and poor speech-to-text conversion effect caused by poor long-distance sound pickup be solved, but also the speaker can be pointed to in real time, allowing the hearing-impaired user to look at the speaker in time according to the pointer indication, which has strong interaction convenience.
[0146] In this way, this embodiment provides a method to improve the experience of hearing-impaired users during conference communication. By inputting numbers and names by remote speakers to start recording and hearing-impaired users to start recognition, hearing-impaired users can accurately know who is speaking, and even at a relatively long distance, they can have better speech parsing and role differentiation effects. When others are speaking, the hearing-impaired user can look at the speaker in time according to the pointer, greatly reducing the obstacles for hearing-impaired users in scenarios such as meetings and small-scale gatherings.
[0147] Each of the above method embodiments, or various possible implementation manners in each method embodiment, can be executed independently, or any two or more of them can be combined with each other. It can be specifically determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0148] For the information display method provided by the embodiments of the present application, the execution subject can be an electronic device or an information display device. In the embodiments of the present application, taking the information display device executing the information display method as an example, the information display device provided by the embodiments of the present application is described.
[0149] Figure 7 Shows a possible structural schematic diagram of the information display device involved in the embodiments of the present application. As Figure 7 shown, the information display device 700 may include: a display module 701.
[0150] Optionally, in some embodiments of the present application, the above display module 701 is configured to display a voice recognition interface, where the voice recognition interface includes N user identifiers, and the display positions of the N user identifiers correspond to the actual positions among the N users indicated by the N user identifiers; the above display module 701 is further configured to, when the first voice of the first speaker is collected, display the first text and the first user identifier on the voice recognition interface, where the first text is the text obtained by text-converting the first voice, the first speaker is one of the N users, and the first user identifier is the user identifier corresponding to the first speaker among the N user identifiers; where N is an integer greater than 1; the above user identifier includes at least one of the following: the seat identifier of the user's seat, the user nickname, and the user avatar.
[0151] Optionally, in some embodiments of the present application, the above voice recognition interface includes a first display area and a second display area, and the N user identifiers are displayed in the first display area; specifically, the display module 701 is configured to display the first text and the first user identifier in the second display area.
[0152] Optionally, in some embodiments of the present application, after the above display module 701 displays the first text and the first user identifier on the voice recognition interface, the first user identifier is displayed in the first display area in a preset display manner.
[0153] Optionally, in some embodiments of the present application, after the above display module 701 displays the first text and the first user identifier on the voice recognition interface, in the first display area, orientation prompt information is displayed, and the orientation prompt information is used to prompt the orientation of the first speaker relative to the target user, where the target user is one of the N users; where the orientation prompt information includes at least one of the following: displaying distance information in text form, displaying direction information in text form, and indicating direction information in icon form.
[0154] Optionally, in some embodiments of the present application, in combination Figure 7 , as Figure 8 shown, the above device 700 further includes: an acquisition module 702; the acquisition module 702 is configured to acquire user position information before the above display module 701 displays the first interface in the first electronic device, and the user position information is used to indicate the positional relationship among the N users; the above display module 701 is further configured to, based on the user position information acquired by the above acquisition module 702, display the N user identifiers on the voice recognition interface; the above user position information includes at least one of the following: the position information of the N electronic devices, and the actual position picture among the N users; where one of the N electronic devices corresponds to one of the N users.
[0155] Optionally, in some embodiments of the present application, in combination Figure 7, such as Figure 9 As shown, the above device 700 further includes: a processing module 703; the processing module 703 is configured to, if a second voice of a second speaker is received within the acquisition duration of the first voice, split the first voice into M voice segments, where the second speaker is one of N users; the processing module 703 is further configured to sort the M voice segments and the second voice based on the acquisition timestamps corresponding to each voice segment and the acquisition timestamp of the second voice to generate a voice queue; the display module 701 is further configured to display M second texts and a first user identifier, as well as a third text and a second user identifier on the voice recognition interface according to the voice queue generated by the processing module 703; where the M second texts are texts obtained by text-converting the M voice segments, the third text is a text obtained by text-converting the second voice; the second user identifier is the identifier of the second speaker among the N user identifiers.
[0156] In the information display device provided in the embodiment of the present application, a voice recognition interface is displayed, and the voice recognition interface includes N user identifiers, and the display positions of the N user identifiers correspond to the actual positions among the N users indicated by the N user identifiers; when the first voice of the first speaker is acquired, a first text and a first user identifier are displayed on the voice recognition interface, the first text is a text obtained by text-converting the first voice, the first speaker is one of N users, and the first user identifier is the user identifier corresponding to the first speaker among the N user identifiers; where N is an integer greater than 1; the user identifier includes at least one of the following: the seat identifier of the user's seat, the user nickname, the user avatar. In this solution, the information display device intuitively shows the actual positions among the N users to the user by displaying the N user identifiers in the voice recognition interface. At the same time, when the information display device acquires the voice of the speaker, the text obtained by text-converting the first voice can be associated and displayed with the user identifier on the voice recognition interface, and combined with the N user identifiers in the voice recognition interface, the orientation of the speaker can be quickly distinguished, thereby improving the efficiency of distinguishing the speaker.
[0157] The information display device in the embodiments of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than terminals. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0158] The information display device in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0159] The information display device provided in the embodiments of the present application can implement each process implemented by the embodiments of the information display method and achieve the same technical effects. To avoid repetition, details are not described herein again.
[0160] Optionally, as Figure 10 shown, the embodiments of the present application further provide an electronic device 800, including a processor 801 and a memory 802. A program or instruction that can run on the processor 801 is stored on the memory 802. When the program or instruction is executed by the processor 801, it implements each step of the above-mentioned information display method embodiments and can achieve the same technical effects. To avoid repetition, details are not described herein again.
[0161] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0162] Figure 11 FIG. is a schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.
[0163] The electronic device 100 includes, but is not limited to, components such as a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110.
[0164] Those skilled in the art can understand that the electronic device 100 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 110 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 11 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here. The above electronic device may be a first electronic device.
[0165] Among them, the above display unit 106 is used to display a voice recognition interface. The voice recognition interface includes N user identifiers. The display positions of the N user identifiers correspond to the actual positions among the N users indicated by the N user identifiers. The above display unit 106 is further used to, when the first voice of the first speaker is collected, display the first text and the first user identifier on the voice recognition interface. The first text is the text obtained by text conversion of the first voice. The first speaker is one of the N users, and the first user identifier is the user identifier corresponding to the first speaker among the N user identifiers. Wherein, N is an integer greater than 1. The user identifier includes at least one of the following: the seat identifier of the user's seat, the user nickname, and the user avatar.
[0166] Optionally, in some embodiments of the present application, the above voice recognition interface includes a first display area and a second display area. The N user identifiers are displayed in the first display area. The display unit 106 is specifically used to display the first text and the first user identifier in the second display area.
[0167] Optionally, in some embodiments of the present application, the above display unit 106 is further used to, after displaying the first text and the first user identifier on the voice recognition interface, display the first user identifier in the first display area in a preset display manner.
[0168] Optionally, in some embodiments of the present application, the above display unit 106 is further used to, after displaying the first text and the first user identifier on the voice recognition interface, display orientation prompt information in the first display area. The orientation prompt information is used to prompt the orientation of the first speaker relative to the target user. The target user is one of the N users. Wherein, the orientation prompt information includes at least one of the following: displaying distance information in text form, displaying direction information in text form, and indicating direction information in icon form.
[0169] Optionally, in some embodiments of the present application, the above-mentioned processor 110 is configured to obtain user location information before the above-mentioned display unit 106 displays a first interface in a first electronic device, where the user location information is used to indicate the positional relationship among N users; the above-mentioned display unit 106 is further configured to display N user identifiers on the voice recognition interface based on the user location information; the user location information includes at least one of the following: the location information of N electronic devices, the actual position pictures among N users; wherein, one of the N electronic devices corresponds to one of the N users.
[0170] Optionally, in some embodiments of the present application, the above-mentioned processor 110 is further configured to, if a second voice of a second speaker is received within the acquisition duration of the first voice, split the first voice into M voice segments, where the second speaker is one of the N users; the above-mentioned processor 110 is further configured to sort the M voice segments and the second voice based on the acquisition timestamps corresponding to each voice segment and the acquisition timestamp of the second voice to generate a voice queue; the above-mentioned display unit 106 is further configured to display M second texts and a first user identifier, as well as a third text and a second user identifier on the voice recognition interface according to the above-mentioned voice queue; wherein, the M second texts are texts obtained by text-converting the M voice segments, the third text is a text obtained by text-converting the second voice; the second user identifier is the identifier of the second speaker among the N user identifiers.
[0171] In the electronic device provided in the embodiments of the present application, a voice recognition interface is displayed, and the voice recognition interface includes N user identifiers, and the display positions of the N user identifiers correspond to the actual positions among the N users indicated by the N user identifiers; in the case where a first voice of a first speaker is collected, a first text and a first user identifier are displayed on the voice recognition interface, the first text is a text obtained by text-converting the first voice, the first speaker is one of the N users, and the first user identifier is the user identifier corresponding to the first speaker among the N user identifiers; wherein, N is an integer greater than 1; the user identifier includes at least one of the following: the seat identifier of the seat where the user is located, the user nickname, the user avatar. In this solution, the electronic device intuitively shows the actual positions among N users to the user by displaying N user identifiers in the voice recognition interface. At the same time, in the case where the electronic device collects the voice of a speaker, the text obtained by text-converting the first voice can be associated and displayed with the user identifier on the voice recognition interface, and combined with the N user identifiers in the voice recognition interface, the orientation of the speaker can be quickly distinguished, thereby improving the efficiency of distinguishing the speaker.
[0172] It should be understood that in the embodiments of the present application, the input unit 104 may include a Graphics Processing Unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes the image data of static pictures or videos obtained by an image capturing device (such as a camera) in a video capturing mode or an image capturing mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. The other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.
[0173] The memory 109 can be used to store software programs and various data. The memory 109 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 can include a volatile memory or a non-volatile memory, or the memory 109 can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 109 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.
[0174] The processor 110 may include one or more processing units; optionally, the processor 110 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor 110 either.
[0175] The embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned information display method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0176] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk, or an optical disc, etc.
[0177] The embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above-mentioned information display method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0178] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.
[0179] The embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement each process of the above-mentioned information display method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0180] It should be noted that, in this text, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising such element. In addition, it should be pointed out that the scope of the methods and apparatuses in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0181] From the description of the above embodiments, those skilled in the art can clearly understand that the above-described method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0182] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit of the present application and the scope protected by the claims, can also make many forms, all of which fall within the protection scope of the present application.
Claims
1. An information display method, executed by a first electronic device, characterized in that: The method comprises: Displaying a voice recognition interface, the voice recognition interface including N user identifiers, wherein display positions of the N user identifiers correspond to actual positions between the N users indicated by the N user identifiers; In the case where a first voice of a first speaker is collected, a first text and a first user identifier are displayed on the voice recognition interface, wherein the first text is a text obtained by converting the first voice into text, the first speaker is one of the N users, and the first user identifier is a user identifier corresponding to the first speaker among the N user identifiers; Wherein, N is an integer greater than 1; the user identifier includes at least one of the following: a seat identifier of the user's seat, a user nickname, and a user avatar.
2. The method according to claim 1, characterized in that The speech recognition interface includes a first display area and a second display area; Displaying the N user identifiers in the first display area; The displaying the first text and the first user identifier on the voice recognition interface includes: The first text and the first user identification are displayed in the second display area.
3. The method according to claim 2, characterized in that After the voice recognition interface displays the first text and the first user identifier, the method further includes: The first user identifier is displayed in the first display area in a preset display manner.
4. The method according to claim 2 or 3, characterized in that: After the voice recognition interface displays the first text and the first user identifier, the method further includes: In the first display area, position prompt information is displayed, where the position prompt information is used to indicate the position of the first speaker relative to a target user, where the target user is one of the N users; The direction prompt information includes at least one of the following: displaying distance information in the form of text, displaying direction information in the form of text, and indicating direction information in the form of an icon.
5. The method according to claim 1, characterized in that Before displaying the first interface in the first electronic device, the method further includes: Acquire user location information, where the user location information is used to indicate a location relationship between the N users; Based on the user location information, displaying the N user identifiers on the voice recognition interface; The user location information includes at least one of the following: location information of N electronic devices, and actual location pictures of the N users; Among them, one electronic device among the N electronic devices corresponds to one user among the N users.
6. The method according to claim 1, characterized in that The method further comprises: If a second voice of a second speaker is received within the collection time of the first voice, the first voice is split into M voice segments, and the second speaker is one of the N users; Based on the collection timestamp corresponding to each voice segment and the collection timestamp of the second voice, sort the M voice segments and the second voice to generate a voice queue; According to the voice queue, displaying M second texts and the first user identifier, as well as a third text and a second user identifier on the voice recognition interface; Among them, the M second texts are texts obtained by converting the M voice segments into text, the third text is a text obtained by converting the second voice into text; and the second user identifier is the identifier of the second speaker among the N user identifiers.
7. An information display device, characterized in that: The information display device comprises: a display module; The display module is used to display a voice recognition interface, wherein the voice recognition interface includes N user identifiers, and the display positions of the N user identifiers correspond to the actual positions of the N users indicated by the N user identifiers; The display module is further configured to display a first text and a first user identifier on the voice recognition interface when a first voice of a first speaker is collected, wherein the first text is a text obtained by converting the first voice into text, the first speaker is one of the N users, and the first user identifier is a user identifier corresponding to the first speaker among the N user identifiers; Wherein, N is an integer greater than 1; the user identifier includes at least one of the following: a seat identifier of the user's seat, a user nickname, and a user avatar.
8. The device according to claim 7, characterized in that The voice recognition interface includes a first display area and a second display area, and the N user identifiers are displayed in the first display area; The display module is specifically configured to display the first text and the first user identifier in the second display area.
9. The device according to claim 8, characterized in that The display module is further configured to display the first user identifier in a preset display mode in the first display area after the first text and the first user identifier are displayed on the voice recognition interface.
10. The device according to claim 8 or 9, characterized in that The display module is further used to display direction prompt information in the first display area after the first text and the first user identifier are displayed on the voice recognition interface, and the direction prompt information is used to prompt the direction of the first speaker relative to the target user, and the target user is one of the N users; wherein the direction prompt information includes at least one of the following: displaying distance information in the form of text, displaying direction information in the form of text, and indicating direction information in the form of an icon.
11. The device according to claim 7, characterized in that The device also includes: an acquisition module; The acquisition module is used to acquire user location information before the display module displays the first interface in the first electronic device, where the user location information is used to indicate the location relationship between the N users; The display module is further configured to display the N user identifiers on the voice recognition interface based on the user location information acquired by the acquisition module; The user location information includes at least one of the following: location information of N electronic devices, and actual location pictures of the N users; Among them, one electronic device among the N electronic devices corresponds to one user among the N users.
12. The device according to claim 7, characterized in that The device further comprises: a processing module; The processing module is configured to split the first voice into M voice segments if a second voice of a second speaker is received within the collection time of the first voice, and the second speaker is one of the N users; The processing module is further used to sort the M voice segments and the second voice based on the collection timestamp corresponding to each voice segment and the collection timestamp of the second voice to generate a voice queue; The display module is further configured to display M second texts and the first user identifier, as well as a third text and a second user identifier on the voice recognition interface according to the voice queue generated by the processing module; Among them, the M second texts are texts obtained by converting the M voice segments into text, the third text is a text obtained by converting the second voice into text; and the second user identifier is the identifier of the second speaker among the N user identifiers.