Information display method and device, electronic equipment and storage medium

By displaying participant identification and speaking information in the meeting interface, and combining voiceprint feature comparison and user selection, the problem of indistinguishable subtitle text in multi-user meetings is solved, achieving more flexible and accurate information display.

CN122293785APending Publication Date: 2026-06-26VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2026-03-31
Publication Date
2026-06-26

Smart Images

  • Figure CN122293785A_ABST
    Figure CN122293785A_ABST
Patent Text Reader

Abstract

This application discloses an information display method, apparatus, electronic device, storage medium, and program product, belonging to the field of artificial intelligence technology. The method includes: displaying a conference interface of a first conference, the first conference including N participants, where N is an integer greater than 1; when the received first audio includes the speaking voices of M participants, displaying M participant identifiers and M speaking information on the conference interface; wherein, the N participants include M participants; one participant identifier is used to indicate one participant; M is a positive integer less than or equal to N.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to an information display method, device, electronic device, and storage medium. Background Technology

[0002] Currently, with the development of communication technology, when users conduct conference calls through conference applications on electronic devices, the electronic devices can convert the voices of different conference participants into subtitles and display the subtitles on the conference interface.

[0003] In related technologies, electronic devices can display the aforementioned subtitle text in the display area corresponding to the meeting account of the meeting participant, so that users can quickly identify the meeting participant corresponding to the subtitle text displayed by the electronic device.

[0004] However, in the above method, if multiple users use a conference account for a conference call, the electronic device will display the caption text corresponding to all users in the same display area, making it impossible for users to distinguish the conference party corresponding to the caption text. Thus, the electronic device has poor flexibility in displaying caption text. Summary of the Invention

[0005] The purpose of this application is to provide an information display method, apparatus, electronic device, and storage medium that can improve the flexibility of electronic devices in displaying subtitle text.

[0006] In a first aspect, embodiments of this application provide an information display method, which includes: displaying a conference interface of a first conference, the first conference including N participants, where N is an integer greater than 1; when the received first audio includes the speaking voices of M participants, displaying M participant identifiers and M speaking information on the conference interface; wherein, the N participants include M participants; a participant identifier is used to indicate a participant; and M is a positive integer less than or equal to N.

[0007] In the information display method provided in this application embodiment, since the electronic device can display the participant identifier and the speaking information corresponding to each participant in the meeting interface of the first meeting, the user can determine the participant to which the speaking information belongs by the participant identifier, avoiding the user's inability to determine the participant corresponding to the subtitle during the meeting, and improving the flexibility of the electronic device in displaying subtitle text.

[0008] In some embodiments of this application, the above-mentioned display of participant identifiers and M speech information corresponding to M participants in the meeting interface includes: displaying participant information corresponding to each parameter person among the M participants in the identifier display area of ​​the meeting interface; the participant information includes at least one of the following: participant avatar, participant name, meeting name of the first meeting, participant's speech time, and participant's location information; and displaying the speech information corresponding to each participant in the text display area corresponding to the participant identifier of each of the M participants.

[0009] In this embodiment, the electronic device displays detailed participant information for each of the M participants in the conference interface through the corresponding identifier display area, and displays the speaking information of each participant in the corresponding text display area. This allows users to quickly view the participants and their speaking information based on the identifier display area and text display area corresponding to each participant, thus improving the flexibility of the information displayed by the electronic device.

[0010] In some embodiments of this application, when the received first audio includes the speaking voices of M participants, before displaying the M participant identifiers and M speaking information on the conference interface, the method further includes: extracting voiceprint features from the first audio to obtain M voiceprint feature information; comparing each voiceprint feature information with the voiceprint feature information stored in the database; and determining the participant name corresponding to the voiceprint feature information with a matching degree greater than a first threshold as the participant name of each participant.

[0011] In this embodiment, the electronic device compares the voiceprint feature information extracted from the first audio with the voiceprint feature information stored in the database to determine the name of the participant corresponding to the voiceprint feature information in the first audio, thereby improving the accuracy of the electronic device in determining the name of the participant.

[0012] In some embodiments of this application, the meeting interface includes audio identifiers corresponding to N participants; the method further includes: obtaining the participant names of the N participants; if at least one participant among the N participants is detected to not match any of the M voiceprint features, displaying the participant names of K participants in the user name display area corresponding to the audio identifier of at least one participant in the meeting interface, where K is a positive integer less than M; receiving a first input for a first participant name and a first audio identifier; the first participant name is any one of the K participant names corresponding to the K participants; the first audio identifier is any one of the K audio identifiers corresponding to the K participants; in response to the first input, determining the participant name corresponding to the voiceprint feature of the audio of the first audio identifier, and obtaining the K participant names corresponding to the voiceprint feature of the audio indicated by the K audio identifiers.

[0013] In this embodiment of the application, when the electronic device cannot determine the participant's name through voiceprint features, the electronic device can determine the participant's name corresponding to the audio through user selection, thereby improving the flexibility of the electronic device in determining the participant's name.

[0014] In some embodiments of this application, before displaying the participant identifiers and M speaking information corresponding to the M participants on the meeting interface, the method further includes: performing text conversion processing on the first audio to obtain the first text; determining the industry field corresponding to the first text based on the semantic information of the first text; and updating the first text based on the industry field corresponding to the first text to obtain the speaking information corresponding to the first text.

[0015] In this embodiment, the electronic device adjusts the speech information corresponding to the first text based on the industry field corresponding to the first text, thereby improving the accuracy of the speech information generated by the electronic device.

[0016] In some embodiments of this application, the above-mentioned text update based on the industry field corresponding to the first text to obtain the speech information corresponding to the first text includes: displaying at least one field identifier based on the industry field corresponding to the first text; receiving a second input for a target field identifier among the at least one field identifier; and, in response to the second input, updating the first text based on the target industry field corresponding to the target field identifier to obtain the speech information corresponding to the first text.

[0017] In this embodiment of the application, when the electronic device cannot determine the accurate industry sector, the electronic device can determine the accurate industry sector by the user's selection, thereby improving the accuracy of the electronic device in determining the industry sector.

[0018] In some embodiments of this application, the above-mentioned text update of the first text based on the target industry field corresponding to the target field identifier to obtain the speech information corresponding to the first text includes: replacing the first text segment in the first text with the professional name based on the professional term corresponding to the target industry field to obtain the speech information; wherein, the semantic matching between the first text segment and the professional term.

[0019] In this embodiment, the electronic device can replace the first text segment in the first text with professional terms corresponding to the target industry field, thereby improving the accuracy and professionalism of the speech information generated by the electronic device.

[0020] In some embodiments of this application, the meeting interface includes a caption summary control; after displaying M participant identifiers and M speech information on the meeting interface, the method further includes: receiving a third input to the caption summary control; responding to the third input, performing a text summary based on the speech information corresponding to the M participants to obtain a summary text; and displaying the summary text.

[0021] In this embodiment, the electronic device can summarize the speech information of M participants by the user inputting the subtitle summary control, obtain and display the summary text, thereby improving the flexibility of the electronic device in obtaining the summary text.

[0022] In some embodiments of this application, before displaying the participant identifiers and M speaking information corresponding to the M participants on the conference interface when the received first audio is detected to include the speaking voices of M participants out of N participants, the method further includes: performing echo cancellation on the received second audio to obtain a third audio; performing sound source localization on the third audio to obtain the sound incidence angle corresponding to the third audio; performing noise reduction processing on the third audio based on the audio incidence angle to obtain a fourth audio; and performing de-reverberation on the fourth audio to obtain a first audio.

[0023] In this embodiment of the application, the electronic device can obtain a noise-free first audio by performing audio processing on the second audio, thereby improving the audio quality of the first audio acquired by the electronic device.

[0024] Secondly, embodiments of this application provide an information display device, which includes a display module. The display module is used to display a meeting interface of a first meeting, the first meeting including N participants, where N is an integer greater than 1; and when the received first audio includes the speaking voices of M participants, the meeting interface displays M participant identifiers and M speaking information; wherein, the N participants include M participants; one participant identifier is used to indicate one participant; and M is a positive integer less than or equal to N.

[0025] In one possible implementation, the aforementioned display module is specifically used to display participant information corresponding to each parameter person among the M participants in the identifier display area of ​​the meeting interface; the participant information includes at least one of the following: participant avatar, participant name, meeting name of the first meeting, participant speaking time, and participant location information; and to display the speaking information corresponding to each participant in the text display area corresponding to the participant identifier of each of the M participants.

[0026] In one possible implementation, the information display device further includes: a processing module; the processing module is configured to, when the received first audio includes the speaking voices of M participants, before displaying the M participant identifiers and M speaking information on the conference interface, extract voiceprint features from the first audio to obtain M voiceprint feature information; compare each voiceprint feature information with the voiceprint feature information stored in the database; and determine the participant name corresponding to the voiceprint feature information with a matching degree greater than a first threshold as the participant name of each participant.

[0027] In one possible implementation, the conference interface includes audio identifiers corresponding to N participants; the information display device further includes: a receiving module; a processing module, further configured to acquire the participant names of the N participants; a display module, further configured to, when detecting that at least one participant among the N participants does not match any of the M voiceprint features, display the participant names of K participants in the user name display area corresponding to the audio identifier of at least one participant in the conference interface, where K is a positive integer less than M; the receiving module, configured to receive a first input for a first participant name and a first audio identifier; the first participant name is any one of the K participant names corresponding to the K participants; the first audio identifier is any one of the K audio identifiers corresponding to the K participants; the processing module, further configured to, in response to the first input received by the receiving module, determine the participant name corresponding to the voiceprint feature of the audio of the first audio identifier, and obtain the K participant names corresponding to the voiceprint feature of the audio indicated by the K audio identifiers.

[0028] In one possible implementation, the information display device further includes: a processing module; the processing module is further configured to perform text conversion processing on the first audio to obtain the first text before the display module displays the participant identifiers corresponding to M participants and the M speaking information on the conference interface; and determine the industry field corresponding to the first text based on the semantic information of the first text; and update the first text based on the industry field corresponding to the first text to obtain the speaking information corresponding to the first text.

[0029] In one possible implementation, the information display device further includes: a receiving module; a display module, specifically configured to display at least one domain identifier based on the industry domain corresponding to the first text; a receiving module, configured to receive a second input to a target domain identifier among the at least one domain identifier; and a processing module, specifically configured to update the first text based on the target industry domain corresponding to the target domain identifier in response to the second input received by the receiving module, thereby obtaining the speech information corresponding to the first text.

[0030] In one possible implementation, the aforementioned processing module is specifically used to replace the first text segment in the first text with the professional name based on the professional terminology corresponding to the target industry field, thereby obtaining the speech information; wherein, the semantic matching between the first text segment and the professional terminology is performed.

[0031] In one possible implementation, the meeting interface includes a caption summary control; the information display device further includes a receiving module and a processing module; the receiving module is used to receive a third input to the caption summary control after the display module displays M participant identifiers and M speech information on the meeting interface; the processing module is used to perform a text summary based on the speech information corresponding to the M participants in response to the third input received by the receiving module, and obtain a summary text; the display module is also used to display the summary text.

[0032] In one possible implementation, the information display device further includes: a processing module; the processing module is further configured to, before the display module detects that the received first audio includes the speaking voices of M participants out of N participants, perform echo cancellation on the received second audio to obtain a third audio; perform sound source localization on the third audio to obtain the sound incident angle corresponding to the third audio; perform noise reduction processing on the third audio based on the audio incident angle to obtain a fourth audio; and perform de-reverberation on the fourth audio to obtain the first audio.

[0033] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0034] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0035] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0036] In a sixth aspect, embodiments of this application provide a computer program / program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0037] In this embodiment, since the electronic device can display the participant identifier and the speaking information of each participant in the meeting interface of the first meeting, the user can determine the participant to which the speaking information belongs by the participant identifier, thus avoiding the user being unable to determine the participant corresponding to the subtitle during the meeting and improving the flexibility of the electronic device in displaying subtitle text. Attached Figure Description

[0038] Figure 1 This is a flowchart illustrating an information display method provided in some embodiments of this application;

[0039] Figure 2 This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0040] Figure 3A This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0041] Figure 3B This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0042] Figure 4 This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0043] Figure 5A This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0044] Figure 5B This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0045] Figure 6 This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0046] Figure 7 This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0047] Figure 8 This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0048] Figure 9A This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0049] Figure 9B This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0050] Figure 9C This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0051] Figure 9D This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0052] Figure 9E This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0053] Figure 9F This is a schematic diagram illustrating an example of a meeting interface provided by some embodiments of this application;

[0054] Figure 10 This is a flowchart illustrating an information display method provided in some embodiments of this application;

[0055] Figure 11 This is a flowchart illustrating an information display method provided in some embodiments of this application;

[0056] Figure 12 This is a schematic diagram of the structure of an information display device provided in some embodiments of this application;

[0057] Figure 13 This is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of this application;

[0058] Figure 14 This is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of this application. Detailed Implementation

[0059] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0060] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects. For example, a first object can be one or more, where "more" means at least two. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0061] The terms "at least one" and "at least one of" in this application's specification refer to any one, any two, or a combination of two or more of the included objects. For example, "at least one of a, b, and c" can mean "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple, and multiple means at least two. Similarly, "at least two" means two or more, and its meaning is similar to "at least one". The identifiers in this application are text, symbols, images, etc., used to indicate information, and can use controls or other containers as carriers for displaying information, including but not limited to text identifiers, symbol identifiers, and image identifiers.

[0062] The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. The terminology involved in the embodiments of this application is explained below.

[0063] Controls: Elements in a graphical user interface that can receive user input to perform corresponding processing or display relevant data. Controls can include, but are not limited to, virtual buttons, sliders, progress bars, and checkboxes.

[0064] Interface: Refers to the medium through which users interact with electronic devices. The interface allows users to send commands to the system via input devices and receive feedback information via output devices. Input devices can be keyboards, mice, touchscreens, etc.; monitors, speakers, etc.

[0065] Speech recognition: also known as automatic speech recognition, computer speech recognition, or speech-to-text recognition, aims to enable computers to automatically convert human speech into corresponding text.

[0066] The information display method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0067] The information display method provided in this application can be applied to meeting scenarios.

[0068] Scenario 1: Users Zhang San, Li Si, and Wang Wu are having a meeting via a conferencing application. If the electronic device receives Zhang San's voice, it can convert Zhang San's audio into a first text file, associate Zhang San with the first text file, and display it in the meeting interface. Then, if the electronic device receives Wang Wu's voice, it can convert Wang Wu's audio into a second text file, associate Wang Wu with the second text file, and display it in the meeting interface. Next, if the electronic device receives Li Si's voice, it can convert Li Si's audio into a third text file, associate Li Si with the third text file, and display it in the meeting interface.

[0069] Based on the above-mentioned scenarios applied in the embodiments of this application, the information display method provided in the embodiments of this application allows the electronic device to display the participant identifier and the speaking information corresponding to each participant in the meeting interface of the first meeting. Therefore, the user can determine the participant to which the speaking information belongs by the participant identifier, avoiding the user's inability to determine the participant corresponding to the subtitle during the meeting, and improving the flexibility of the electronic device in displaying subtitle text.

[0070] The information display method provided in this application is executed by an information display device, which can be an electronic device, or a functional module or entity within an electronic device. This application does not limit the specific implementation of this method. The following will use an electronic device as an example to illustrate the information display method provided in this application.

[0071] This application provides an information display method. Figure 1 A flowchart illustrating an information display method provided in an embodiment of this application is shown. Figure 1 As shown, the information display method provided in this application embodiment may include the following steps 201 and 202.

[0072] Step 201: The electronic device displays the conference interface of the first meeting.

[0073] In some embodiments of this application, the first meeting mentioned above includes N participants, where N is an integer greater than 1.

[0074] In some embodiments of this application, the aforementioned first meeting may be a video conference or an audio conference, etc. The specific method can be determined according to actual usage needs, and this application does not impose any limitations.

[0075] In some embodiments of this application, the meeting interface may include at least one of the following: microphone control, screenshot control, role-based control, role editing control, simultaneous interpretation control, and reminder setting control.

[0076] For example, the microphone control described above is used to record the audio of the participants.

[0077] For example, the above screenshot control is used to take a screenshot of the content of the meeting interface.

[0078] For example, the aforementioned role-based control is used to trigger the electronic device to display M participant identifiers and M speaking messages in the meeting interface.

[0079] For example, the above-described edit role control is used to edit the identifiers of M participants.

[0080] For example, the simultaneous interpretation control described above is used to translate the speech of M participants.

[0081] For example, the above-mentioned reminder control is used to remind participants of the to-do items they have set.

[0082] For example, in scenario 1, such as Figure 2 As shown, taking a mobile phone as an example and Zhang San as a mobile phone user, for example, if a solution review meeting is to be organized, Zhang San and Li Si are at their workstations, and Wang Wu and Zhao Liu are in the meeting room. Before the meeting, all participants open the meeting software on their mobile phones. At this time, Zhang San's mobile phone can display the meeting interface 10 of the first meeting, which includes Zhang San, Li Si, Wang Wu and Zhao Liu. The first display area 11 of the meeting interface 10 displays a microphone control 12, a screenshot control 13, a role-based control 14, a role-editing control 15, a simultaneous interpretation control 16 and a reminder setting control 17.

[0083] Step 202: If the received first audio includes the voices of M participants, the electronic device displays the identifiers of M participants and the speech information of M participants on the conference interface.

[0084] In some embodiments of this application, the aforementioned N participants include M participants; a participant identifier is used to indicate a participant; M is a positive integer less than or equal to N.

[0085] It is understandable that the above M speech information can be subtitles corresponding to the voices of the M participants.

[0086] In some embodiments of this application, the electronic device can receive a fifth input from the user to the aforementioned role-based control, thereby displaying M participant identifiers and M speaking information on the conference interface when the received first audio includes the speaking voices of M participants.

[0087] In some embodiments of this application, the fifth input can be user click input, long press input, or voice input on the role-specific control. The specific input can be determined according to actual usage needs, and this application does not impose any limitations.

[0088] For example, the fifth input mentioned above can be the user's click input on the role control.

[0089] In some embodiments of this application, the electronic device can use a speaker separation algorithm to separate the audio corresponding to the speech of each of the M participants from the first audio.

[0090] For example, the speaker separation algorithm described above can be any of the following: clustering algorithm, deep learning-based speech separation algorithm, or large speech model separation algorithm, etc. The specific algorithm can be determined according to actual usage requirements, and this application embodiment does not impose any limitations.

[0091] In some embodiments of this application, the electronic device may display M participant identifiers and M speech messages in the second display area of ​​the conference interface; or, the electronic device may display M participant identifiers and M speech messages in the conference interface via a pop-up window.

[0092] For example, the second display area can be a preset display area or a blank display area in the conference interface.

[0093] For example, the blank display area mentioned above can be an area in the meeting interface that does not include controls or information.

[0094] It should be noted that the specific implementation process of step 202 above can be found in the following embodiments, and will not be repeated here to avoid repetition.

[0095] In some embodiments of this application, the step 202 above, "the electronic device displays the participant identifiers and M speaking information corresponding to the M participants on the conference interface", can be specifically implemented through the following steps 202a and 202b.

[0096] Step 202a: The electronic device displays the participant information corresponding to each parameter person among the M participants in the identification display area of ​​the conference interface.

[0097] In some embodiments of this application, the aforementioned participant information includes at least one of the following: participant avatar, participant name, meeting name of the first meeting, participant speaking time, and participant location information.

[0098] In some embodiments of this application, each of the M participants can correspond to an identifier display area.

[0099] In some embodiments of this application, the aforementioned identification display area can be a preset area in the conference interface.

[0100] For example, the aforementioned M identification display areas can be associated with the audio identifications of the M participants.

[0101] In some embodiments of this application, the aforementioned audio identifier may include at least one of the following: image identifier, text identifier, and special symbol identifier, etc. The specific identifier can be determined according to actual usage requirements, and this application does not impose any limitations.

[0102] In some embodiments of this application, the aforementioned participant avatars can be user-defined or preset by the electronic device. The specific method can be determined based on actual usage needs, and this application does not impose any limitations.

[0103] For example, the aforementioned participant avatars can be real photos of the participants or cartoon avatars.

[0104] In some embodiments of this application, the names of the participants can be user-defined or preset by the electronic device. The specific details can be determined based on actual usage needs, and this application does not impose any limitations.

[0105] For example, the names of the participants mentioned above can be their real names or virtual names.

[0106] It should be noted that, regardless of whether the participant's name is a real name or a virtual name, it is acceptable as long as it can uniquely and accurately identify the corresponding participant.

[0107] In some embodiments of this application, the speaking time of the aforementioned participants can be the current system time of the electronic device.

[0108] In some embodiments of this application, the location information of the participants may include: the city where the participants are located, and the district or county of that city.

[0109] In some embodiments of this application, the name of the first meeting can be user-defined or preset by the electronic device. The specific name can be determined according to actual usage needs, and this application does not impose any limitations.

[0110] For example, the name of the first meeting mentioned above could be: Solution Review Meeting.

[0111] In some embodiments of this application, the participant's avatar, participant's name, the name of the first meeting, the participant's speaking time, and the participant's location information can all be modified by the user via electronic devices.

[0112] Step 202b: The electronic device displays the speaking information corresponding to each participant in the text display area corresponding to the participant identifier of each of the M participants.

[0113] In some embodiments of this application, the text display area corresponding to the participant identifier of each participant may also include the audio identifier of the audio corresponding to each participant.

[0114] For example, combined Figure 2 ,like Figure 3A As shown, assuming user Zhang San needs the subtitle function, he can click on the role-specific control 14 in the meeting interface 10 to input the subtitles, such as... Figure 3BAs shown, when the mobile phone detects that the first audio received includes the voices of Li Si, Zhang San, and Wang Wu, and Li Si is the first speaker, Zhang San is the second speaker, and Wang Wu is the third speaker, the mobile phone can display Li Si's corresponding participant avatar 22 in the speaker area 20 of the meeting interface 10, Li Si's location information: City A, the name of the first meeting: Meeting Room 901, participant name: Li Si, Li Si's speaking time: 8:33; and display Li Si's speaking information in the text display area 23 below the Li Si corresponding identifier display area 21: Today we will discuss the usability of solution A, and everyone can speak freely.

[0115] Then, the mobile phone can display Zhang San's corresponding participant avatar 31, Zhang San's location information 32 (City A, Meeting Name of the First Meeting: Meeting Room 901, Participant Name: Zhang San, Zhang San's Speaking Time 35: 8:50) in the Zhang San corresponding identifier display area 30 of the speaker area 20 of the meeting interface 10; and display Zhang San's speaking information in the text display area 32 below the Li Si corresponding identifier display area 30: I think the logic of the A plan is insufficient and the possibility of its actual implementation is not high.

[0116] Finally, the mobile phone can display Wang Wu's corresponding participant avatar 41 in the speaker area 20 of the meeting interface 10, Wang Wu's location information: City A, meeting name of the first meeting: Meeting Room 901, participant name: Wang Wu, Wang Wu's speaking time: 9:00; and display Wang Wu's speaking information in the text display area 42 below Wang Wu's corresponding avatar display area 40: I agree with Zhang San's point of view, and we can further investigate Plan A.

[0117] In the information display method provided in this application embodiment, the electronic device displays the conference interface of a first conference, which includes N participants, where N is an integer greater than 1. When the received first audio includes the speaking voices of M participants, the conference interface displays M participant identifiers and M speech messages. Here, the N participants include M participants; one participant identifier indicates one participant; and M is a positive integer less than or equal to N. In this solution, since the electronic device can display the participant identifier and speech messages corresponding to each participant in the conference interface of the first conference, the user can determine the participant to whom the speech messages belong through the participant identifier. This avoids the user being unable to determine the participant corresponding to the subtitles during the conference, improving the flexibility of the electronic device in displaying subtitle text.

[0118] In some embodiments of this application, before step 202 above, the information display method provided in this application embodiment further includes steps 301 to 303 as described below.

[0119] Step 301: The electronic device extracts voiceprint features from the first audio signal to obtain M voiceprint feature information.

[0120] In some embodiments of this application, the electronic device can perform audio preprocessing, multi-domain feature fusion, and cascaded processing of artificial intelligence (AI) feature optimization on the first audio to extract voiceprint feature information that is unique to each individual and resistant to interference.

[0121] For example, the aforementioned voiceprint feature information can be a voiceprint feature vector.

[0122] For example, the above audio preprocessing can specifically be: Voice Activity Detection (VAD): using the energy + zero-crossing rate dual threshold method, valid speech segments in the first audio are detected, invalid components such as silence and breathing sounds are removed, and valid speech segments are output.

[0123] It should be noted that the duration of the above-mentioned effective speech segments is greater than or equal to 2 seconds, which meets the minimum speech length requirement for voiceprint extraction.

[0124] Then, amplitude normalization is performed on the effective speech segments, that is, the amplitude of the effective speech segments is mapped to the [-1,1] interval, and pre-emphasis filtering is used to compensate for high-frequency attenuation of speech and improve the stability of feature extraction.

[0125] It should be noted that the above pre-emphasis filtering can use a transfer function, which can be: .

[0126] Finally, the normalized effective speech segments are converted into first frequency domain signals using short-time Fourier transform, retaining the core speech frequency band of 100Hz~8kHz.

[0127] The short-time Fourier transform has a frame length of 20ms, a frame shift of 10ms, and uses a Hanning window plus windowing method for sampling.

[0128] For example, the above-mentioned multi-domain feature fusion can specifically be: fusing acoustic features and prosodic features to construct a multi-dimensional basic feature set to ensure voiceprint distinguishability.

[0129] For example, the electronic device extracts 40-dimensional Mel-frequency cepstral coefficients and their first and second-order difference coefficients corresponding to the first frequency domain signal, totaling 120 dimensions; and extracts 64-dimensional log-Mel-spectral features corresponding to the first frequency domain signal, covering the detailed texture of the speech spectrum. Then, the electronic device extracts the fundamental frequency, duration, and rate of energy change of the first frequency domain signal to construct a 10-dimensional prosodic feature vector. Next, the electronic device extracts the linear prediction coefficients and linear prediction cepstral coefficients corresponding to the first frequency domain signal; these linear prediction coefficients and linear prediction cepstral coefficients can reflect the differences in the physiological structure of the speaker's vocal tract. Finally, the electronic device concatenates the 40-dimensional Mel-frequency cepstral coefficients and their first and second-order difference coefficients, the 64-dimensional log-Mel-spectral features, the 10-dimensional prosodic feature vector, the linear prediction coefficients, and the linear prediction cepstral coefficients to form a (120+64+10+12+12=218)-dimensional basic feature matrix.

[0130] For example, the above AI feature optimization can specifically be as follows: reduce the dimensionality and enhance the 218 feature base feature matrix to obtain the enhanced feature matrix; then, generate speaker embedding vectors from the enhanced feature matrix to obtain the embedded speaker vectors; and finally, perform feature standardization on the embedded speaker vectors to obtain the voiceprint feature vectors.

[0131] For example, the above feature dimensionality reduction and enhancement can be achieved by: inputting the 218 basic features into a lightweight feature encoder, selecting a lightweight variant of the deep residual network ResNet-18 or an efficient channel attention temporal deep neural network, and optimizing the features through the following process:

[0132] Temporal attention mechanism: Introducing a self-attention module, which assigns dynamic weights to frame-level features, enhances the feature contribution of key speech frames, such as vowel segments, and suppresses interference from irrelevant frames such as consonants and transitional sounds.

[0133] Channel attention fusion: The weights of different feature channels are adaptively adjusted through the channel attention mechanism to highlight feature channels that are strongly correlated with the voiceprint, such as Mel frequency cepstral coefficients and linear prediction cepstral coefficients, while weakening noise-sensitive channels.

[0134] For example, the above speaker embedding vector generation can be achieved by: outputting frame-level feature vectors through the fully connected layer of the encoder, using statistical pooling, such as mean and variance pooling, to aggregate the frame-level features and generate a fixed-dimensional, preferably 256-dimensional or 512-dimensional, voiceprint embedding vector. This vector has removed irrelevant factors such as speech content and emotion, and only retains the individual physiological characteristics of the speaker.

[0135] For example, the feature standardization described above can be achieved by performing L2 normalization on the speaker embedding vector, i.e., ,in, Here, E is the speaker embedding vector after feature standardization, ensuring the scale consistency of the feature vectors and providing a unified benchmark for subsequent comparisons.

[0136] In some embodiments of this application, after obtaining the feature-standardized speaker embedding vector, the electronic device can perform feature robustness verification and optimization on the feature-standardized speaker embedding vector to obtain the voiceprint feature vector.

[0137] For example, the above feature robustness verification and optimization can specifically be as follows: the electronic device calculates the discriminative index of the speaker embedding vector after feature standardization, such as the intra-class distance / inter-class distance ratio, which is required to be ≥3.5, and the noise robustness index, such as feature stability under signal-to-noise ratio changes, where the feature fluctuation is ≤5% within the signal-to-noise ratio range of 10~30dB; then, if the feature discriminative index does not meet the requirements, it backtracks to the basic feature extraction stage, increases the Mel frequency cepstral coefficient dimension or optimizes the attention module weights; if the robustness is insufficient, it adds slight noise enhancement training to the basic features, such as data augmentation, to improve the encoder's anti-interference ability.

[0138] Thus, electronic devices achieve feature discrimination with an intra-class distance / inter-class distance ratio ≥4.0, feature stability ≥95% in scenarios with a signal-to-noise ratio of 5~30dB and different speech content, and an embedding vector dimension of only 256, balancing storage efficiency and discrimination.

[0139] Step 302: The electronic device compares each voiceprint feature with the voiceprint feature stored in the database.

[0140] In some embodiments of this application, the electronic device can compare each of the M voiceprint feature information with the voiceprint feature information stored in the database by traversing through them.

[0141] Step 303: The electronic device determines the participant name corresponding to the voiceprint feature information with a matching degree greater than the first threshold as the participant name for each participant.

[0142] In some embodiments of this application, the aforementioned first threshold may be determined by the electronic device or defined by the user. The specific threshold can be determined according to actual usage requirements, and this application does not impose any limitations.

[0143] For example, the value range of the first threshold can be [80%-90%].

[0144] For example, the first threshold mentioned above can be 80%, 85%, or 90%.

[0145] For example, electronic devices can identify the participant's name based on voiceprint feature information with a matching degree greater than 85%.

[0146] In this embodiment, the electronic device compares the voiceprint feature information extracted from the first audio with the voiceprint feature information stored in the database to determine the name of the participant corresponding to the voiceprint feature information in the first audio, thereby improving the accuracy of the electronic device in determining the name of the participant.

[0147] In some embodiments of this application, the meeting interface includes audio identifiers corresponding to N participants.

[0148] For example, the information display method provided in this application embodiment further includes the following steps 401 to 404.

[0149] Step 401: The electronic device obtains the names of the N participants.

[0150] In some embodiments of this application, the electronic device can obtain the names of N participants from the to-do list stored in the electronic device.

[0151] In some embodiments of this application, the to-do list content can be the to-do list content stored in any application in the electronic device.

[0152] Step 402: If at least one of the N participants does not match any of the M voiceprint features, the electronic device displays the names of K participants in the user name display area corresponding to the audio identifier of at least one participant in the conference interface, where K is a positive integer less than M.

[0153] In some embodiments of this application, the electronic device may display the aforementioned user name display area via a pop-up window; or, the electronic device may display the aforementioned user name display area in the blank area corresponding to the audio identifier of each of at least one participant.

[0154] In some embodiments of this application, the electronic device can compare the names of N participants with the names of participants stored in the database, and identify the participant name not stored in the database as at least one of the participants.

[0155] For example, combined Figure 3B ,like Figure 4 As shown, if the mobile phone cannot determine the corresponding participant's name through voiceprint feature information, the mobile phone can display an audio identifier 50 in the speaker area 20 of the conference interface 10, and display a user name display area 51 on the audio identifier 50 through a pop-up window. The user name display area 51 includes: Zhao Liu, Li Qi, Zhao Ba, and Zhou Jiu.

[0156] Step 403: The electronic device receives the first input for the name of the first participant and the first audio identifier.

[0157] In some embodiments of this application, the first participant name is any one of the K participant names corresponding to the K participants; the first audio identifier is any one of the K audio identifiers corresponding to the K participants.

[0158] In some embodiments of this application, the aforementioned first input can be user input via click, long press, or preset trajectory input, etc., of the first participant's name and the first audio identifier. The specific input can be determined according to actual usage requirements, and this application does not impose any limitations.

[0159] For example, the first input mentioned above can be a user's click input of the first participant's name and the first audio identifier.

[0160] In some embodiments of this application, since the user name display area is displayed in the area corresponding to the audio identifier of each participant in at least one participant, the first input described above can also be an input only for the name of the first participant.

[0161] Step 404: The electronic device responds to the first input, determines the participant name corresponding to the voiceprint feature of the audio of the first audio identifier, and obtains the K participant names corresponding to the voiceprint features of the audio indicated by the K audio identifiers.

[0162] In some embodiments of this application, the electronic device may associate the voiceprint features of the first audio identifier with the participant name of the first participant and store them in a database.

[0163] It is understandable that for each participant name that the electronic device cannot recognize, the electronic device can obtain the K participant names corresponding to the voiceprint features of the audio indicated by the K audio identifiers through the above method.

[0164] In some embodiments of this application, after the electronic device determines the user corresponding to the first audio, it can display the participant identifier corresponding to the first audio in the conference interface.

[0165] For example, combined Figure 4 ,like Figure 5A As shown, users can slide to input Zhao Liu and the first audio 50, such as... Figure 5BAs shown, the mobile phone can display Zhao Liu's avatar 55, the meeting name of the first meeting: Meeting Room 901, the participant name: Zhao Liu, and Zhao Liu's speaking time: 9:05 in the identification display area 54 corresponding to the first audio 50; and display Zhao Liu's speaking information: That concludes today's meeting in the text display area 56 below the identification display area 54 corresponding to Zhao Liu, and cancel the display of the user name display area 51.

[0166] In this embodiment of the application, when the electronic device cannot determine the participant's name through voiceprint features, the electronic device can determine the participant's name corresponding to the audio through user selection, thereby improving the flexibility of the electronic device in determining the participant's name.

[0167] In some embodiments of this application, before "displaying the participant identifiers and M speaking information corresponding to M participants on the meeting interface" in step 202 above, the information display method provided in this application embodiment further includes the following steps 501 to 503.

[0168] Step 501: The electronic device performs text conversion processing on the first audio to obtain the first text.

[0169] In some embodiments of this application, the electronic device can perform audio-to-text processing on the first audio to obtain the first text.

[0170] It should be noted that the specific implementation process of the above audio-to-text conversion can be found in the description of related technologies, and will not be repeated here to avoid repetition.

[0171] Step 502: The electronic device determines the industry sector corresponding to the first text based on the semantic information of the first text.

[0172] In some embodiments of this application, an electronic device can determine the semantic information of a first text using a first model.

[0173] For example, the first model described above can be an artificial intelligence model, a neural network model, or a large language model, etc. The specific model can be determined according to actual usage requirements, and this application embodiment does not impose any limitations.

[0174] In some embodiments of this application, the electronic device can perform semantic matching with all industry sectors pre-stored in the electronic device based on semantic information, thereby determining the industry sector corresponding to the first text.

[0175] Step 503: The electronic device updates the first text based on the industry field corresponding to the first text to obtain the speech information corresponding to the first text.

[0176] It should be noted that the specific implementation process of step 503 above can be found in the following embodiments, and will not be repeated here to avoid repetition.

[0177] For example, such as Figure 6 As shown, the mobile phone can update the first text based on the technology field corresponding to the first text, obtaining the corresponding speaking information 61. In addition, the conference has specially set up a thematic forum to focus on discussing ethical issues crucial to AI development, such as data privacy, federated learning, and AI guardrails. Industry leaders and innovators will share cutting-edge insights into zero-shot and few-shot learning, jointly drawing a blueprint for responsible embodied intelligence development and promoting interdisciplinary collaboration and exchange.

[0178] In this embodiment, the electronic device adjusts the speech information corresponding to the first text based on the industry field corresponding to the first text, thereby improving the accuracy of the speech information generated by the electronic device.

[0179] In some embodiments of this application, step 503 can be specifically implemented by steps 503a to 503c as described below.

[0180] Step 503a: The electronic device displays at least one domain identifier based on the industry domain corresponding to the first text.

[0181] It is understood that the above-mentioned industry sectors can be at least one, and the at least one industry sector can correspond to at least one sector identifier, that is, one industry sector corresponds to one sector identifier.

[0182] In some embodiments of this application, the electronic device may display at least one field identifier in the conference interface via a pop-up window; or, the electronic device may display at least one field identifier in a blank area of ​​the conference interface.

[0183] In some embodiments of this application, the aforementioned field identifier may include at least one of the following: text identifier, numerical identifier, or special symbol identifier, etc. The specific identifier can be determined according to actual usage requirements, and this application does not impose any limitations.

[0184] For example, such as Figure 7 As shown, the mobile phone can display five domain identifiers in the meeting interface 10 via pop-up window 70. The five domain identifiers are: general domain identifier 71, technology domain identifier 72, economic domain identifier 73, legal domain identifier 74, and health domain identifier 76.

[0185] In some embodiments of this application, the electronic device can enable the domain recognition function in the domain settings interface.

[0186] For example, such as Figure 8As shown, the mobile phone can display the settings interface 80 of the meeting interface 10. The settings interface 80 includes a voice settings control 81, a subtitle size control 82, and an industry domain control 83. At this time, the user can click on the industry domain control 83 to input, so that the mobile phone can enable the domain recognition function.

[0187] Step 503b: The electronic device receives a second input for a target field identifier in at least one field identifier.

[0188] In some embodiments of this application, the second input can be a user's click input, long press input, or voice input on a target domain identifier in at least one domain identifier. The specific input can be determined according to actual usage needs, and this application does not impose any limitations.

[0189] For example, the second input mentioned above can be a user's click input on the target area identifier.

[0190] Step 503c: The electronic device responds to the second input and updates the first text based on the target industry field corresponding to the target field identifier to obtain the speech information corresponding to the first text.

[0191] It should be noted that the specific implementation process of step 503c above can be found in the following embodiments, and will not be repeated here to avoid repetition.

[0192] In this embodiment of the application, when the electronic device cannot determine the accurate industry sector, the electronic device can determine the accurate industry sector by the user's selection, thereby improving the accuracy of the electronic device in determining the industry sector.

[0193] In some embodiments of this application, step 503c can be specifically implemented by step 503c1 as described below.

[0194] Step 503c1: The electronic device replaces the first text segment in the first text with the professional name corresponding to the target industry field to obtain the speaking information.

[0195] In some embodiments of this application, the first text segment described above is semantically matched with technical terms.

[0196] In some embodiments of this application, the aforementioned technical terms may be one or more. When there are multiple technical terms, there are also multiple first text segments, that is, one technical term may correspond to one first text segment.

[0197] In some embodiments of this application, the electronic device can perform a traversal operation on the first text to determine a first text segment that semantically matches the professional terminology corresponding to the target industry field, and then replace the first text segment in the first text with the professional terminology to obtain the speech information.

[0198] In this embodiment, the electronic device can replace the first text segment in the first text with professional terms corresponding to the target industry field, thereby improving the accuracy and professionalism of the speech information generated by the electronic device.

[0199] In some embodiments of this application, the meeting interface described above includes a caption summary control.

[0200] For example, after step 202 above, the information display method provided in this application embodiment further includes steps 601 to 603 as described below.

[0201] Step 601: The electronic device receives a third input to the subtitle summary control.

[0202] In some embodiments of this application, the aforementioned third input can be user input via clicking, long-pressing, or inputting a preset trajectory on the subtitle summary control. The specific input can be determined according to actual usage requirements, and this application does not impose any limitations.

[0203] For example, the third input mentioned above can be the user's click input on the subtitle summary control.

[0204] Step 602: The electronic device responds to the third input and summarizes the text based on the speaking information of the M participants to obtain the summary text.

[0205] In some embodiments of this application, the electronic device can input the speaking information of M participants into the first model described above, so as to summarize the speaking information of the M participants into a text using the first model, and obtain a summary text.

[0206] It should be noted that the specific process of the electronic device summarizing text through the first model can be found in the description in the relevant technology, and will not be repeated here to avoid repetition.

[0207] Step 603: The electronic device displays the summary text.

[0208] In some embodiments of this application, the electronic device can display summary text in the meeting interface via a pop-up window.

[0209] In some embodiments of this application, the electronic device can display the document identifier of the document corresponding to the summary text in the meeting interface, and then the user can click on the document identifier to enable the electronic device to jump to the document interface of the document application to display the summary text.

[0210] In some embodiments of this application, the electronic device can summarize the speech information of all the above users using summary templates corresponding to different industry fields.

[0211] For example, such as Figure 9A As shown, when the aforementioned industry sectors are general sectors, the summary template may include: Sector Template Name 90: General Template, Recording Time 91: Year A, Month B, Day C, Attendees 92: Li Si, Zhang San, Wang Wu, Zhao Liu, Summary 93, Key Points 94, Conclusion 95, and To-Do Items 96. For example... Figure 9B As shown, in the case of meeting minutes in the aforementioned industry sectors, the summary template may include: Sector Template Name 101: Meeting Minutes, Recording Time 102: Year A, Month B, Day C, Participants 103: Li Si, Zhang San, Wang Wu, Zhao Liu, Summary 104, Final Conclusion 105, Discussion Points 106, and To-Do Items 107. Figure 9C As shown, in the case of call records in the aforementioned industry sectors, the summary template may include: Sector Template Name 201: Call Record, Recording Time 202: Year A, Month B, Day C, Participants 203: Li Si, Zhang San, Wang Wu, Zhao Liu, Summary 204, Key Points of the Call 205, Conclusion 206, and To-Do Items 207. Figure 9D As shown, in the case of lectures and popular science presentations in the aforementioned industry sectors, the summary template can include: Sector Template Name 301: Lecture Popular Science, Recording Time 302: Year A, Month B, Day C, Attendees 303: Li Si, Zhang San, Wang Wu, Zhao Liu, Summary 304, Core Viewpoints 305, and Knowledge Extension 306. For example... Figure 9E As shown, in the case of the aforementioned industry field being education and learning, the summary template may include: Field Template Name 401: Education and Learning, Recording Time 402: Year A, Month B, Day C, Participants 403: Li Si, Zhang San, Wang Wu, Zhao Liu, Summary 404, Classroom Knowledge 405, and Key Points Marking 406. For example... Figure 9F As shown, in the case of interview records in the above-mentioned industry fields, the above summary template may include: Field Template Name 501: Interview Record, Recording Time 502: Year A Month B Day C, Participants 503: Li Si, Zhang San, Wang Wu, Zhao Liu, Summary 504, Interviewee 505, Question and Answer Record 506, and Analysis Summary 507.

[0212] In this embodiment, the electronic device can summarize the speech information of M participants by the user inputting the subtitle summary control, obtain and display the summary text, thereby improving the flexibility of the electronic device in obtaining the summary text.

[0213] In some embodiments of this application, before step 202 above, the information display method provided in the embodiments of this application further includes steps 701 to 704 as described below.

[0214] Step 701: The electronic device performs echo cancellation on the received second audio to obtain the third audio.

[0215] In some embodiments of this application, the second audio may be acquired through a microphone in an electronic device, or the second audio may be received through a conferencing application.

[0216] In some embodiments of this application, the electronic device can perform signal synchronization, linear echo cancellation, and nonlinear echo residual suppression on the received second audio to obtain the aforementioned third audio.

[0217] For example, the above signal synchronization can specifically be achieved by using timestamp alignment or an adaptive synchronization algorithm to ensure the timing consistency between the near-end audio signal x(n) and the far-end reference signal r(n) in the second audio, thereby eliminating misalignment interference caused by signal transmission delay. Here, x(n) represents the sound acquired by the microphone in the electronic device from the on-site participants and the sound played aloud by the conference software, and r(n) represents the echo source signal, such as the voice from the other end of the call.

[0218] For example, the above-mentioned linear echo cancellation can specifically be as follows: Initialize adaptive filtering parameters: Set the order L of the Normalized Least Mean Square (NLMS) adaptive filter, configured according to the actual echo path length, preferably 64~512 order; set the step size factor μ, with an initial value of 0.01~0.1, supporting dynamic adjustment; and set the filter coefficient vector. It is initialized as a zero vector, and the regularization parameter δ takes a value of 1e-6 to 1e-4 to avoid the denominator being zero.

[0219] Then, the filter input is constructed: a filter input vector is generated based on the far-end reference signal r(n). Linear echo estimation: Estimating the echo signal using an NLMS filter. ,in Transpose of the filter coefficient vector; adaptive coefficient update: calculate the linear error signal. Update filter coefficients based on the NLMS criterion: ,in The second norm square of the input vector r(n) is used to balance the filtering convergence speed and steady-state error through this update strategy.

[0220] For example, the above-mentioned nonlinear echo residual suppression can specifically be: Feature extraction: for the error signal after linear echo cancellation Feature engineering is performed to extract time-domain features, such as short-time energy and zero-crossing rate; frequency-domain features, such as Mel-frequency cepstral coefficients and power spectral density; and statistical features, such as signal-to-noise ratio estimates, to construct a multidimensional feature matrix F(n). Nonlinear echo modeling and suppression: The feature matrix F(n) is input into a pre-trained AI model, preferably a lightweight deep learning model, such as a convolutional neural network, a gated recurrent unit, or a lightweight variant of the Transformer. This AI model learns nonlinear echoes during the training phase, such as speaker nonlinear distortion, nonlinear reflections from the acoustic environment, and feature differences from near-end speech, and outputs a nonlinear echo estimate. Pure signal output: The target signal after echo cancellation is obtained through residual error calculation. This achieves full-dimensional suppression of both linear and nonlinear echoes.

[0221] In some embodiments of this application, the electronic device can also adaptively adjust and robustly optimize the target signal after echo cancellation to obtain a third audio.

[0222] For example, the above adaptive adjustment and robust optimization can specifically be as follows: the electronic device can monitor the signal-to-noise ratio and echo return loss enhancement index of the target signal s(n) in real time, and dynamically adjust the step size factor μ of the NLMS filter and the output weight of the AI ​​model: when the echo return loss enhancement index is <15dB, increase μ to accelerate filter convergence; when the signal-to-noise ratio is >25dB, reduce the computational complexity of the AI ​​model to save computing power and ensure the stability of echo cancellation under different acoustic scenarios.

[0223] Thus, this step eliminates more than 80% of linear echoes through NLMS filtering, and the AI ​​model further suppresses nonlinear echoes. The final echo return loss enhancement index reaches 20~35dB, meeting the echo suppression requirements of real-time voice communication, such as video conferencing and in-vehicle calls.

[0224] Step 702: The electronic device locates the sound source of the third audio and obtains the sound incident angle corresponding to the third audio.

[0225] Step 703: The electronic device performs noise reduction processing on the third audio based on the audio incident angle to obtain the fourth audio.

[0226] Step 704: The electronic device performs de-reverberation on the fourth audio signal to obtain the first audio signal.

[0227] It should be noted that the specific implementation process of steps 702 to 704 above can be found in the following embodiments, and will not be repeated here to avoid repetition.

[0228] In this embodiment of the application, the electronic device can obtain a noise-free first audio by performing audio processing on the second audio, thereby improving the audio quality of the first audio acquired by the electronic device.

[0229] This application provides a method. Figure 10 A flowchart illustrating an information display method provided in an embodiment of this application is shown. The following will use a meeting scenario as an example to exemplify the information display method provided in this application. Figure 10 As shown, the information display method provided in this application embodiment may include the following steps 101 to 144.

[0230] Step 101: Participants open the meeting software and log in to their accounts. The account name is displayed by default.

[0231] For example, suppose we are organizing a solution review meeting. Zhang San and Li Si are at their workstations, while Wang Wu and Zhao Liu are in the meeting room. Before the meeting, Zhang San opens the meeting software on his mobile phone and displays the meeting waiting screen 60. The meeting waiting screen includes Zhang San's avatar, Zhang San's name, a join meeting control, a quick meeting control, a schedule meeting control, and a screen sharing control.

[0232] Step 102: Participants may selectively open the subtitle software on their mobile phones.

[0233] For example, participants who need to use the mobile phone's subtitle function can open the subtitle function of the mobile phone's subtitle software; if they do not need this function, they can directly use the subtitle function built into the meeting software.

[0234] Step 103, Audio Settings: In the same conference room, each offline participant can only use one mobile device to record audio, and the audio recording function of the conference room screen should be turned off.

[0235] For example, if both parties are logged into the conference system, the microphone function of the conference screen needs to be turned off in order to utilize the microphone array of the mobile phone. This ensures that there is no live sound in the reference signal for echo cancellation, preventing the live sound from being eliminated. At the same time, to avoid sound resonance, only one mobile phone device is allowed to record sound.

[0236] Step 104, Audio Settings: In the same conference room, each offline participant can only use one mobile phone to play audio, or only play audio through the conference room's large screen.

[0237] For example, if the meeting room does not have a conference system, one of the mobile phones can be projected onto the TV screen, and the TV can display the meeting PPT and act as the speaker. If the TV does not support external audio output, the mobile phone can be used to output the audio. In short, there should only be one external audio source.

[0238] The sound r(n) played by the conference system is transmitted through the speaker in one direction and through the system channel in another direction to step 106 to provide a reference signal for echo cancellation.

[0239] Step 105: The electronic device acquires the audio of all participants in the offline meeting room and the audio played out by the meeting software.

[0240] For example, the microphone of the mobile phone responsible for recording can acquire the voices of the participants on site and the voices played by the conference software, which is called the near-end audio signal x(n), including near-end speech, linear echo, and ambient noise.

[0241] Step 106: The electronic equipment performs echo cancellation, eliminating the sound from the conference system's external speakers and retaining only the voice of the offline speaker.

[0242] For example, the near-end audio signal collected in step 105 and the far-end reference signal r(n) provided in step 104, i.e. the echo source signal, such as the voice of the other end in the call, are compared, and the sound signal in step 104 is echo-cancelled, retaining only the voice of the online reference person.

[0243] It should be noted that the specific implementation process can be found in the above embodiments, and will not be repeated here to avoid repetition.

[0244] Step 107: Electronic devices perform audio signal processing: multi-microphone sound source localization, multi-microphone noise reduction, and reverb removal.

[0245] For example, the mobile phone can select different recording modes, each corresponding to a different signal processing scheme. Recording modes include standard mode, conference mode, speaker mode, and interview mode.

[0246] For example, as shown in Table 1, Table 1 illustrates the differences between the four modes mentioned above.

[0247] Table 1

[0248]

[0249] For example, in order to improve the signal processing effect, the mobile phone integrates 4 microphones and adopts a 3+1 microphone, that is, 3 main array microphones + 1 reference microphone.

[0250] For example, the sound source localization algorithm is deployed after echo cancellation and before multi-microphone noise reduction. Through the collaborative localization of 3 main arrays + 1 reference microphone, it improves the angle estimation accuracy in complex reverberation or noisy environments, providing a precise steering vector for subsequent beamforming. The specific steps are as follows:

[0251] Step 1: Preprocessing of positioning signals:

[0252] Receive 4 signals after echo cancellation (3 main array signals) The reference microphone signal X_r(n) is first filtered by a bandpass filter (100Hz~8kHz, covering the core speech frequency band) to remove irrelevant frequency band interference, and then converted into a frequency domain signal by a short-time Fourier transform (STFT, frame length 10~20ms, frame shift 5~10ms). and (k is the frequency point, l is the frame index), to reduce the interference of temporal reverberation on positioning.

[0253] Step 2: Building the foundation for localization through multi-domain feature fusion:

[0254] Step 2.1: Calculate the frequency domain covariance matrix of the 3-channel main array. By diagonal loading Optimize the numerical stability of the covariance matrix to avoid matrix singularity caused by reverberation.

[0255] Step 2.2: Extract auxiliary features of the reference microphone: Calculate the signal power ratio between the reference microphone and the 3-channel main array. ,when If the frequency point is determined to be dominated by strong noise, its weight is reduced during localization.

[0256] Step 3, MUSIC-ESPRIT Fusion Positioning (Main Array Core Calculation):

[0257] Step 3.1, Eigenvalue decomposition and subspace separation: For Perform eigenvalue decomposition ,in, This is the signal subspace (corresponding to the target speech). This is the noise subspace (corresponding to reverberation + noise).

[0258] Step 3.2, Coarse Localization (MUSIC Algorithm): Constructing the Spatial Spectral Function ,in Domain-oriented vector c is the speed of sound, 340 m / s. (where k corresponds to the frequency), and the angle corresponding to the peak value of the search spatial spectrum. Complete coarse positioning (error ≤ ±8°).

[0259] Step 3.3, Fine localization (ESPRIT algorithm): using Centered on a target area, a fine search range of ±10° is defined. Utilizing the rotation invariance of the ESPRIT algorithm, the search is conducted through the signal subspace. Solve for the exact value of the incident angle. Positioning accuracy ≤ ±2°.

[0260] Step 4: Refer to microphone-assisted verification and robustness optimization:

[0261] Step 4.1, Verification of positioning results: Calculate the target angle. The ratio of the main array signal energy to the reference microphone signal energy in the direction ,like If the positioning result is determined to be contaminated by noise, the current angle is updated using the moving average of the positioning results of the previous 3 frames.

[0262] Step 4.2, Dynamic Angle Tracking: An extended Kalman filter is used to track the angle. Smoothing is performed, and the state equation is set as follows: ( (for angular variation noise), the observation equation is set as follows: ( (to reduce observation noise), outputting the final stable target speech incidence angle. .

[0263] Step 5, Output the positioning results: Output the results for each frame. It is transmitted in real time to the subsequent MVDR beamforming module to update the steering vector and blocking matrix, so as to realize the dynamic coordination of positioning and beamforming.

[0264] For example, the following is a patented technology description of a 3+1 microphone (3 main array + 1 reference microphone) MVDR+GSC noise reduction module that adapts to the "sound source localization → multi-microphone noise reduction" process logic and enhances the "localization-beamforming dynamic coordination". It embeds a real-time localization result feedback mechanism to ensure that the beamforming and sound source position are dynamically matched.

[0265] This module is deployed after the sound source localization module and receives the real-time target speech incidence angle output by the localization module. (l is the frame index). Through the collaborative design of "3-channel main array dynamic beamforming + 1-channel reference microphone noise assistance + GSC residual suppression", a closed loop of "real-time feedback of positioning results - dynamic update of beam weights - precise noise suppression" is achieved. The specific steps are as follows:

[0266] Step 1, 3+1 Microphone Signal Synchronization and Preprocessing (Input from Positioning Module):

[0267] Step 1.1, Signal Reception and Synchronization: Receive the three main array signals after echo cancellation. (Including target speech and ambient noise), 1 reference microphone signal (Far from the sound source, primarily ambient noise), while simultaneously receiving the real-time incident angle output by the sound source localization module for each frame. Hardware synchronization pulses and timestamps are used to ensure the timing consistency of the four signals, and the angle data is strictly synchronized with the signal frame (frame length 10~20ms, consistent with the frame length of the positioning module).

[0268] Step 1.2, Preprocessing Optimization: Perform bandpass filtering (100Hz~8kHz, matching the positioning module frequency band), high-pass filtering (cutoff frequency 80~120Hz, filtering out low-frequency vibration noise) and impulse noise removal (short-time energy threshold method) on the four signals, and output the preprocessed main array signal. With reference microphone signal To avoid noise interference with beamforming weight calculation.

[0269] Step 2, Dynamic Steering Vector Update (Location-Beamforming Coordination Core):

[0270] Step 2.1, Frequency Domain Conversion: For Perform a short-time Fourier transform (STFT, with parameters consistent with the positioning module: frame length 10~20ms, frame shift 5~10ms, Hanning window) to obtain the frequency domain signal. (k is the frequency point).

[0271] Step 2.2, Angle Adaptive Guidance Vector Generation: Based on the real-time positioning angle of each frame. Calculate the dynamic frequency domain steering vector corresponding to each frequency point k:

[0272]

[0273] in The spacing between the main array elements (c=340m / s is the speed of sound) (where k is the frequency corresponding to the frequency point), ensuring that the steering vector matches the current sound source position in real time.

[0274] Step 2.3, Dynamic Update of Covariance Matrix: Based on the current frame's main array frequency domain signal, calculate and update the covariance matrix. Matrix stability is optimized using sliding window averaging (window length 3-5 frames), while diagonal loading (load factor) is added. This avoids matrix singularities caused by sound source movement.

[0275] Step 3: MVDR Dynamic Beamforming (Spatial Filtering Based on Real-Time Positioning):

[0276] Step 3.1, Frame-level optimal weight solution: For each frame of signal, using the real-time steering vector... and dynamic covariance matrix Design an MVDR beamformer with the following input and constraints: (To ensure the target speech is distortion-free), the objective function is to minimize the output power. The optimal weights in the frequency domain for each frame are obtained by solving:

[0277]

[0278] in It is the inverse of the covariance matrix, and the weights vary with the positioning angle. Real-time updates enable "sound source motion - beam tracking".

[0279] Step 3.2, Dynamic Beam Output: The three main array frequency domain signals are weighted and summed with the frame-level optimal weights to obtain the spatially filtered main beam signal.

[0280]

[0281] right Perform inverse STFT conversion to time domain signal At this time, the signal gain in the target speech direction is ≥15dB, and the suppression ratio of coherent noise in other directions (such as side background voices) is ≥22dB.

[0282] Step 4: GSC Generalized Sidelobe Cancellation (Location Assist + Reference Microphone Enhancement):

[0283] Step 4.1 Construction of Dynamic Blocking Matrix: Based on Real-Time Guiding Vector Design an adaptive blocking matrix (Dimension 2×3), satisfying To ensure complete blocking of the target speech and generate sidelobe reference signals. .

[0284] Step 4.2, Dual-noise reference fusion: The sidelobe reference signal is fused... With reference microphone frequency domain signal The fusion yields an enhanced noise reference signal. It utilizes the pure noise characteristics of a reference microphone to improve the accuracy of noise estimation, especially suitable for incoherent noise scenarios.

[0285] Step 4.3, Dynamic Training of Adaptive Noise Canceller: For input, Given the desired signal, the weights of an adaptive filter are trained based on the Normalized Least Mean Square (NLMS) criterion. The weight update formula is:

[0286]

[0287] in For adaptive step size (dynamically adjusted according to the rate of change of positioning angle: increase when the rate of change of angle > 5° / frame). (Accelerate convergence) For error signals, This is the regularization factor.

[0288] Step 4.4, Residual Noise Suppression Output: The frequency domain signal after GSC processing is obtained through noise cancellation.

[0289]

[0290] Perform inverse STFT conversion to time domain signal At this point, the residual noise suppression ratio is ≥18dB and the target speech distortion (PESQ) is ≥3.9.

[0291] Step 5, Dynamic Robustness Optimization and Final Output:

[0292] Step 5.1, Positioning Assistance Adjustment: Real-time monitoring of the rate of change of positioning angle. ,when When the sound source moves rapidly, increase the MVDR covariance matrix update frequency (shorten the window length to 2 frames) and increase the GSC step size. To ensure the beam quickly tracks the sound source; when When the sound source is stationary, reduce the update frequency to save computing power.

[0293] Step 5.2, AI Fine Noise Reduction (Optional): For Extract the fused features of "array phase difference + time-frequency features + positioning angle correlation", input them into a lightweight CNN model, further suppress non-steady residual noise (such as keyboard noise, sudden interference), and output an optimized signal. .

[0294] Step 5.3, Signal Normalization: For Amplitude normalization (mapped to the [-1,1] interval) and DC component removal are performed to output the final noise-reduced signal. This is then passed to the subsequent de-reverb module.

[0295] Thus, when the sound source moving speed is ≤1m / s, the beam tracking delay is ≤10ms, the angle tracking error is ≤±2°, and the noise suppression ratio (NSR) in moving scenarios is still ≥28dB; noise reduction and speech preservation: steady-state noise suppression ratio ≥35dB, non-steady-state noise suppression ratio ≥25dB, target speech intelligibility (STOI) ≥0.90, speech distortion (PESQ) ≥4.0; real-time performance: the overall algorithm latency is ≤45ms, it is compatible with portable devices such as mobile phones, in-vehicle terminals, and smart speakers, and the computing power consumption is ≤50MIPS.

[0296] For example, the dreverberation algorithm is deployed after 3+1 microphone MVDR+GSC noise reduction to precisely suppress indoor reverberation (speech trailing and blurring caused by reflected sound). It adopts a cascaded design of "linear dreverberation + nonlinear residual reverberation suppression", and the specific steps are as follows:

[0297] Step 1: Reverberation signal preprocessing and feature analysis:

[0298] Receive signal after multi-microphone noise reduction Perform STFT conversion to frequency domain signal (Consistent with the STFT parameters for sound source localization to ensure frame synchronization), calculate the reverberation characteristics of each frame.

[0299] Reverberation time estimation: By fitting the energy decay curve, the boundary between early reflections and late reverberation of the speech signal is estimated. (Reverberation time, which is the time it takes for the energy to decay by 60 dB).

[0300] Frequency domain features: Extract subband energy ratio (early reflection energy / late reverberation energy), spectral flatness, and phase distortion to construct a reverberation feature vector. .

[0301] Step 2, Linear De-reverberation (Weighted Prediction Error, WPE Algorithm):

[0302] Step 2.1: Constructing the prediction model: Based on the linear prediction characteristics of speech signals, a frequency-domain weighted prediction error model is established, assuming the current frame signal... It can be linearly predicted from the signal of the previous p frames (p=4~8 frames, adapting to common reverberation delays):

[0303]

[0304] in For frequency domain prediction coefficients, This is the estimated reverberation signal.

[0305] Step 2.2, Optimization of weighting and prediction coefficients: Introducing a weighting matrix ( (For noise variance estimation of the first m frames), by minimizing the weighted prediction error. Solve for the optimal prediction coefficients:

[0306]

[0307] in The autocorrelation matrix is... This is a cross-correlation vector.

[0308] Step 2.3, Linear Reverberation Cancellation: Calculate the signal after linear dereverberation. At this point, the early reverberation suppression ratio is ≥25dB.

[0309] Step 3: Nonlinear residual reverberation suppression (AI-assisted modeling):

[0310] Step 3.1, Residual Reverberation Feature Extraction: For Extract cross-domain features, including: STFT amplitude / phase spectrum, Mel-frequency cepstral coefficients (MFCC, 16-dimensional), and reverberation feature vector. The phase consistency characteristics of the three main arrays are fused to construct a multi-dimensional feature matrix. .

[0311] Step 3.2, AI residual reverberation modeling: ... Input a pre-trained lightweight deep learning model (preferably a mini variant of U-Net or a TCN temporal convolutional network). The model learns the difference between the "clean speech spectrum + residual reverberation spectrum" during the training phase (especially for late-stage nonlinear reverberation), and outputs a frequency domain estimate of the residual reverberation. .

[0312] Step 3.3, Spectral Subtraction Optimization: Adaptive spectral subtraction is used to cancel residual reverberation, and the de-reverberated frequency domain signal is output.

[0313]

[0314] in The pure speech spectrum estimated by the model. For adaptive inhibition factor (according to) Dynamic adjustment The larger, The closer to 1), Set a floor threshold (to avoid speech distortion). To prevent zero factor.

[0315] Step 4: Signal Reconstruction and Robustness Adjustment

[0316] Step 4.1, Inverse STFT Reconstruction: For Perform inverse STFT conversion to obtain the time-domain dereverberation signal. Inter-frame distortion is eliminated by overlapping and adding.

[0317] Step 4.2, Dynamic Adaptive Optimization: Monitor the Speech Proficiency Index (STOI) in real time. When STOI < 0.85, increase the prediction order p of WPE and improve the suppression weight of the AI ​​model. When STOI > 0.95, reduce the computational complexity of the algorithm to balance performance and real-time performance.

[0318] Step 5, Final Signal Output:

[0319] right Perform amplitude normalization and DC component removal to output a clean, reverberant target speech signal. Complete the entire process.

[0320] Step 108: The electronic device requests the cloud service framework to obtain speech-to-text and role separation information.

[0321] For example, the cloud service framework consists of three parts: preprocessing, algorithm factory, and postprocessing.

[0322] For example, the algorithm factory includes three types of algorithms: role separation, speech recognition, and speech translation. Each type of algorithm integrates some of the best APIs available on the market. The performance of each API varies depending on the specific scenario.

[0323] For example, the algorithm factory section is described in detail below:

[0324] Start the algorithm service and establish a WebSocket network connection, including authentication and handshake; deploy cascading APIs on the server side. Cascading APIs refer to the sequential execution of Automatic Speech Recognition (ASR) and Text Translation (MT).

[0325] For example, inputting audio, speech recognition produces the text: "I'd like to ask if the current system supports centralized mode?", and then the text is translated from Chinese to English to get: "I'd like to ask if the current system supports centralized mode?".

[0326] LID (Linguistic Identification) for Source Language: Since subtitle software supports "Chinese-English self-recognition," when people from different countries appear in a meeting, it is not necessary to frequently switch the source language. Therefore, in this case, the result language identification (LID) is needed to determine the source language, thereby configuring the source and target language parameters of the translation MT engine and the end-to-end engine.

[0327] For example, the aforementioned preprocessing refers to the ability to distinguish different usage scenarios through language and domain identification, allowing different algorithms to be matched to different scenarios. The domains supported by this application include general, technological, economic, legal, and health fields.

[0328] This application also supports automatic domain matching. However, because domain identification may be incorrect, automatic switching is not possible. Instead, we use a suggested matching strategy, providing users with suggestions for selection. The specific suggested matching trigger threshold rules are as follows:

[0329] When the number of transcribed words exceeds 500, the transcribed content is summarized by AI to determine whether it conforms to the above-mentioned professional fields (the algorithm needs to optimize the summary prompt here, and further strategies will be added depending on the judgment); if the field can be determined, suggestions will be made.

[0330] If the AI ​​summary cannot determine the professional field, then when the number of transcribed words reaches 1000, 1500, and 2000 (these will be dynamically adjusted based on the actual assessment results), the AI ​​will summarize and judge the transcribed content again; if it can determine the field, then suggestions will be made; no suggestions are needed for general fields.

[0331] If a judgment still cannot be made, no further judgment or advice will be given at this time.

[0332] If the user switches scenarios midway, the transcription count will be recalculated and the system will return to the general context.

[0333] If the user accepts the suggestion to switch professional fields, the previously generated transcript will not be changed.

[0334] For example, the post-processing described above refers to addressing some domain-specific recognition errors, further improving the performance of speech recognition, speech translation, and role separation. Specifically, it involves first determining the domain of the recognized text. This can be done using a text classification model to ascertain the current topic, such as finance, technology, or business. Then, keyword extraction algorithms like NER or LLM are used to extract domain-specific keywords. These keywords are then matched against keywords in our dictionary, such as through pinyin matching, text matching, or fuzzy pinyin matching. If a keyword in the dictionary is matched but does not match the extracted recognized text, then keyword replacement is performed, thereby improving the accuracy of the text.

[0335] For example, OpenAI launched the large-scale dialogue model ChatGPD. However, the word "ChatGPD" was incorrectly identified. We replaced it with "ChatGPT" from the technology lexicon, thus obtaining the correct result "OpenAI launched the large-scale dialogue model ChatGPT".

[0336] Hot word matching relies heavily on the construction of a domain-specific thesaurus, including its coverage and quality. Inevitably, some user hot words may not be covered, leading to keyword recognition errors. Therefore, we support users inputting their own keywords to improve accuracy. It's important to note that users do not need to modify the keywords themselves; they simply need to enter them into the hot word list.

[0337] Step 109: The electronic device determines whether there are multiple speakers.

[0338] For example, if yes, then step 111 is executed; if no, then step 110 is executed.

[0339] For example, the algorithm engine will return role information, such as speaker 1, speaker 2. If the number of speakers is greater than 1, then it is a scenario with multiple speakers.

[0340] Step 110: Synchronize the subtitle software and conference software account name in the electronic device.

[0341] For example, if a single person is participating in the meeting, they can choose to use the account name of the meeting software, such as Zhang San logging into the meeting, and then the captioning software will display that the speaker is Zhang San.

[0342] Step 111: The electronic device prompts the user to change the meeting room name.

[0343] For example, if multiple people are attending the meeting, users should be prompted to change the name of the subtitle software to the meeting room name, such as Meeting Room 901 in City A. The meeting room name should be distinctive, and the general format is: Location + Meeting Room Number.

[0344] Step 112: The electronic device automatically or manually fills in the speaker's name based on the role memory function.

[0345] For example, if the names of Speaker 1 and Speaker 2 are not yet determined, one approach is to allow users to manually modify them, and another is a semi-automatic filling method.

[0346] For example, such as Figure 11 As shown, step 1121: The electronic device obtains the individual account of the historical meeting.

[0347] For example, if a meeting was attended in the past and there was an individual's account, then there would be that individual's name, and this name information can be remembered; for example, Zhang San and Wang Wu. Because there is an error in separating roles under multiple accounts, only the name and audio of the individual account are used for memorization.

[0348] Step 1122: Extract account names and data from electronic devices and store them in the database.

[0349] Extract the username of this account, for example, Zhang San, and then create a storage unit in the database. This can be achieved using Table 2.

[0350] Table 2

[0351]

[0352] Step 1123: The electronic device saves the speaker's audio clips and voiceprint features.

[0353] In the database, an audio clip of this account is stored, and voiceprint features are extracted from this audio clip using a cloud-based voiceprint algorithm. The algorithm for extracting voiceprints is described below.

[0354] The voiceprint extraction algorithm is deployed after the dereverberation module to obtain the clean speech signal after dereverberation. Using this as input, a cascaded design of "preprocessing - multi-domain feature fusion - AI feature optimization" is employed to extract voiceprint feature vectors that are unique to each individual and resistant to interference. The specific steps are as follows:

[0355] Step 1: Preprocessing before voiceprint extraction:

[0356] Step 1.1, Voice Activity Detection (VAD): A dual threshold method of energy and zero-crossing rate is used to detect the signal. The valid speech segments are processed, and invalid components such as silence and breathing sounds are removed to output the valid speech segments. (Duration ≥ 2s, meeting the minimum speech length requirement for voiceprint extraction).

[0357] Step 1.2, Normalization: For Perform amplitude normalization (mapping the signal amplitude to the [-1,1] interval) and pre-emphasis filtering (transfer function) This compensates for high-frequency attenuation in speech and improves the stability of feature extraction.

[0358] Step 1.3, Time-Frequency Conversion: Short-Time Fourier Transform (STFT, frame length 20ms, frame shift 10ms, Hanning window plus windowing) is used to convert the time-domain speech into frequency-frequency content. Convert to frequency domain signal (k is the frequency point, l is the frame index), reserve the 100Hz~8kHz core voice frequency band.

[0359] Step 2: Extraction of multi-domain basic voiceprint features:

[0360] By integrating acoustic and prosodic features, a multi-dimensional basic feature set is constructed to ensure voiceprint distinguishability.

[0361] Step 2.1, Spectral Features: Extract 40-dimensional Mel frequency cepstral coefficients (MFCC) and their first and second order difference coefficients, for a total of 120 dimensions; extract 64-dimensional log-Mel spectrogram features to cover the detailed texture of the speech spectrum.

[0362] Step 2.2, Prosodic Features: Extract the fundamental frequency (F0, estimated by autocorrelation), duration (duration of each speech frame), and energy change rate to construct a 10-dimensional prosodic feature vector.

[0363] Step 2.3, Vocal Tract Features: Extract linear prediction coefficients (LPC, 12th order) and linear prediction cepstral coefficients (LPCC, 12th order) to reflect the differences in the physiological structure of the speaker's vocal tract.

[0364] Step 2.4, Basic Feature Fusion: Concatenate the above features to form a (120+64+10+12+12=218) dimensional basic feature matrix. (Dimension is "frames × 218").

[0365] Step 3: AI-enhanced voiceprint feature optimization

[0366] Step 3.1, Feature Dimensionality Reduction and Enhancement: The basic feature matrix... Input a lightweight feature encoder (preferably a lightweight variant of the deep residual network ResNet-18 or ECAPA-TDNN (Efficient Channel Attention Temporal Deep Neural Network)) and optimize the features through the following process:

[0367] Temporal attention mechanism: Introducing a self-attention module, which assigns dynamic weights to frame-level features, enhances the feature contribution of key speech frames (such as vowel segments), and suppresses interference from irrelevant frames such as consonants and transitional sounds.

[0368] Channel attention fusion: The weights of different feature channels are adaptively adjusted through the channel attention mechanism to highlight feature channels that are strongly correlated with the voiceprint, such as MFCC and LPCC, while weakening noise-sensitive channels.

[0369] Step 3.2, Speaker Embedding Vector Generation: The frame-level feature vectors output from the fully connected layer of the encoder are aggregated using statistical pooling (mean + variance pooling) to generate speaker embedding vectors of fixed dimensions (preferably 256 or 512 dimensions). (D=256 / 512) This vector has removed irrelevant factors such as speech content and emotions, and only retains the individual physiological characteristics of the speaker.

[0370] Step 3.3, Feature Standardization: Perform L2 normalization on the embedding vector E, i.e. This ensures the consistency of feature vector scale, providing a unified benchmark for subsequent comparisons.

[0371] Step 4: Feature robustness verification and optimization:

[0372] Step 4.1, Feature Quality Assessment: Calculate the embedding vector The discrimination index (such as intra-class distance / inter-class distance ratio, which requires ≥3.5) and the noise robustness index (such as feature stability under signal-to-noise ratio changes, with feature fluctuation ≤5% in the SNR range of 10~30dB).

[0373] Step 4.2, Dynamic Adjustment: If the feature discrimination does not meet the requirements, backtrack to the basic feature extraction stage, increase the MFCC dimension or optimize the attention module weights; if the robustness is insufficient, add slight noise to the basic features to enhance training (data augmentation) and improve the encoder's anti-interference ability.

[0374] Step 5, Voiceprint Feature Output: Output the standardized voiceprint embedding vector. It can be stored in the voiceprint database (registration stage) or directly input into the voiceprint comparison module (verification stage).

[0375] Thus, the feature discrimination ratio (intra-class / inter-class distance ratio) of this voiceprint extraction algorithm is ≥4.0. In scenarios with SNR=5~30dB and different speech content, the feature stability is ≥95%, and the embedding vector dimension is only 256, balancing storage efficiency and discrimination.

[0376] Step 1124: Electronic devices construct database triples: name, audio, and voiceprint.

[0377] For example, the person's name, audio clip, and voiceprint features are stored in the same storage unit to form a triple. Since the voiceprint has already been stored, why is it necessary to store the audio? Because the voiceprint algorithm is not fixed and may be upgraded. If the voiceprint algorithm changes, then the voiceprint needs to be re-prepared.

[0378] Step 1125: Electronic devices acquire the accounts of multiple people in the current meeting.

[0379] In the next meeting, if multiple people are attending, you can use subtitle software to separate the roles and generate subtitles. By default, it will return Speaker 1 and Speaker 2.

[0380] Step 1126: The electronic device obtains the names of the participants based on the meeting notification.

[0381] A list of attendees can be obtained from SMS messages or meeting notices in the schedule.

[0382] For example, the following schedule:

[0383] [Fixed Agenda Item 2]: Synchronizing Project Progress and Risks

[0384] [Duration]: 20 minutes

[0385] [Key stakeholders]: Must-attendees and representatives of demand

[0386] Zhang San, Li Si, Wang Wu, Zhao Liu, Li Qi, Zhao Ba, Zhou Jiu

[0387] The extracted personnel list is shown in Table 3:

[0388] Table 3

[0389]

[0390] Step 1127: The electronic device determines whether there is stored memory.

[0391] For example, the electronic device compares the names in the list of attendees with the names in the memory database to determine whether they are in the memory database; if they are in the memory database, proceed to step 309; if they are not in the memory database, proceed to step 308.

[0392] Step 1128: The electronic device generates a list of candidate names.

[0393] If the speaker's name is not in the role memory, a list of alternative names is generated, as shown in Table 4. These alternative names are convenient for us to match unknown speakers.

[0394] Table 4

[0395]

[0396] Step 1129: The electronic device determines whether the voiceprint algorithm has been upgraded.

[0397] For example, the electronic device determines whether the voiceprint algorithm has been upgraded based on the parameters. If it has been upgraded, proceed to step 1130; otherwise, proceed to step 1131.

[0398] Step 1130: If the voiceprint algorithm is upgraded, the electronic device extracts the audio segments of the participants from the stored memory, extracts the voiceprint features using the new algorithm, and saves the new features into the database.

[0399] For example, as shown in Table 5.

[0400] Table 5

[0401]

[0402] Step 1131: If the voiceprint algorithm remains unchanged, the electronic device directly extracts the voiceprint features from the memory bank.

[0403] Step 1132: Open the subtitle software on the electronic device.

[0404] For example, a user opens a mobile phone subtitle software.

[0405] Step 1133: Separate roles for electronic devices.

[0406] For example, the electronic device turns on the role separation function, performs voice recognition, and separates roles to obtain speaker 1, speaker 2, speaker N, etc.

[0407] Step 1134: The electronic device extracts the voiceprint features of multiple speakers respectively.

[0408] For example, the electronic device extracts the voiceprints of speaker 1, speaker 2, and speaker N respectively, and creates a new temporary data table, as shown in Table 6.

[0409] Table 6

[0410]

[0411] Step 1135: The electronic device determines whether the voiceprint matches.

[0412] For example, an electronic device can match the voiceprint obtained from the character's separation with the voiceprint in the character's memory bank.

[0413] The following describes the algorithm for extracting and comparing voiceprints.

[0414] The voiceprint comparison algorithm receives "template voiceprint features stored during the registration phase" and "voiceprint features to be compared extracted during the verification phase." Through a process of "similarity calculation - confidence calibration - dynamic decision-making," it achieves accurate verification of the speaker's identity. The specific steps are as follows:

[0415] Step 1: Feature preprocessing before alignment.

[0416] Step 1.1, Feature Reading and Alignment: Read the registered template features of the target speaker from the voiceprint database. (K≥3, representing the feature vector set from multiple registrations to improve template reliability), read the features to be compared extracted during the verification phase. .

[0417] Step 1.2, Feature Consistency Verification: Verification and The dimensions are consistent (both are D-dimensional). If the dimensions do not match, feature interpolation or re-extraction is performed to ensure the feasibility of the comparison.

[0418] Step 2: Calculate the similarity of multiple metrics.

[0419] By integrating multiple similarity measurement methods, this approach balances comparison accuracy and robustness against interference, avoiding the limitations of a single measurement method.

[0420] Step 2.1, Cosine Similarity: Calculate the cosine similarity between the feature to be compared and each template feature, reflecting the directional consistency of the feature vectors.

[0421]

[0422] in , The closer the value is to 1, the higher the similarity.

[0423] Step 2.2, Euclidean Distance Similarity: Calculate the normalized Euclidean distance between feature vectors, reflecting the distance differences in the feature space.

[0424]

[0425] in The maximum possible distance (normalized to the [0,1] interval). The closer to 1, the higher the similarity.

[0426] Step 2.2, Mahalanobis Distance Similarity: Introducing a feature covariance matrix to eliminate interference from correlations between feature dimensions, suitable for high-dimensional feature comparison:

[0427]

[0428] in To register the covariance matrix of the template features, It is the trace (normalization factor) of the inverse of the covariance matrix.

[0429] Step 2.4, Weighted Fusion Similarity: Based on the reliability of the measurement in different scenarios, dynamic weights are assigned to each similarity score. (satisfy The final similarity score is obtained by fusion:

[0430]

[0431] Preferred weighting configuration: It can be dynamically adjusted according to the actual scenario.

[0432] Step 3: Confidence calibration and dynamic threshold decision.

[0433] Step 3.1, Confidence Mapping: Using a pre-trained calibration model (such as logistic regression or a lightweight neural network), the fused similarity scores are mapped... Mapped to confidence scores Where Conf≥80 indicates a high-confidence match, and Conf≤40 indicates a low-confidence mismatch.

[0434] Step 3.2, Dynamic Threshold Generation: Discarding fixed thresholds, the decision threshold Th is dynamically adjusted based on the following factors:

[0435] Registration template quality: The larger K is (the more registrations), the lower the threshold Th is by 2 to 5 percentage points.

[0436] Verify speech quality: When the SNR of the verified speech is ≥20dB, the threshold is reduced by 3 percentage points; when the SNR is ≤10dB, the threshold is increased by 5 percentage points.

[0437] Scene adaptation: Quiet indoor scene Th=65, noisy outdoor scene Th=75, in-vehicle scene Th=70.

[0438] Step 4, Identity Decision and Result Output:

[0439] Step 4.1, Decision Logic: If the confidence score The system determines that "the speaker to be compared matches the registered speaker" (match successful); if... The system determines that "identity does not match" (matching failed).

[0440] Step 4.2, Result Optimization: If (Critical interval), initiate a second comparison: re-extract the voiceprint features of the speech to be compared (change the feature extraction parameters, such as adjusting the MFCC dimension), repeat steps 2-3, and determine the final decision based on the results of the second comparison to avoid critical misjudgment.

[0441] Thus, the correct recognition rate (CRR) of this voiceprint matching algorithm is ≥99.2%, the false acceptance rate (FAR) is ≤0.01% (when FAR=0.1%, the false rejection rate FRR≤0.5%), and the matching latency is ≤100ms, which meets the requirements of real-time identity verification.

[0442] Step 1136: If a match is found, the electronic device will directly replace the account with a person's name.

[0443] For example, if a match is found, the person's name is directly replaced.

[0444] Step 1137: If no match is found, the electronic device provides a list of alternative names and allows users to edit the names.

[0445] If the algorithm is inaccurate and fails to find a match, a list of alternative names is provided, allowing users to manually select a name. Manual editing of the speaker's name is also supported. Manually edited names are guaranteed to be correct and will be automatically saved to the memory dictionary. This can further improve the accuracy of character separation.

[0446] Step 113: The electronic device shares the subtitles from the mobile phone subtitle software with the conferencing software and obtains subtitles from other parties from the conferencing software.

[0447] For example, the shared captioning software provides multi-person captions to the conferencing software, while simultaneously retrieving captions from other parties within the conferencing software. It is suggested that in other meeting rooms where multiple people are also participating, mobile captioning software can be used simultaneously, allowing for speaker separation, including the text segments and start / end times of each speaker's remarks.

[0448] Step 114: The electronic device collects the captions of all speakers and summarizes and analyzes the meeting.

[0449] For example, electronic devices can generate summaries, mind maps, to-do lists, etc., according to templates.

[0450] For example, the common fields in the above templates include: title, summary, minutes, conclusion, and tasks.

[0451] In this embodiment, the present invention combines the single-person captioning capability of conferencing software with the multi-person captioning capability of mobile terminals. Conferencing software has poor ability to separate multiple roles, which is precisely the advantage of mobile terminals; at the same time, conferencing software can separate multiple single-person signals, which is precisely the disadvantage of mobile terminals. The former is suitable for online single-person voice transcription, while the latter is suitable for offline multi-person voice transcription, and the combination of the two can achieve complementary advantages.

[0452] Furthermore, online audio undergoes signal processing upon transmission to a mobile phone over the network, including noise reduction, packet loss, and poor signal strength, resulting in degraded audio quality. Therefore, mobile devices are not natively suitable for transcribing online audio into text. In contrast, offline meeting audio undergoes signal processing via a mobile phone's multi-microphone array, further enhancing audio quality. This direct transmission to the speech recognition engine without network overhead further improves the results. Thus, the mobile phone's advantage lies in offline interactions, while the meeting software's advantage lies in online interactions; the two complement each other, resulting in optimal text transcription accuracy.

[0453] Furthermore, conferencing software can directly display the names of participating speakers, while mobile captioning software can only separate multiple people and cannot determine the speaker's name. Conferencing software thus compensates for the shortcomings of mobile captioning software. Determining the speaker's name after role separation requires utilizing role memory capabilities.

[0454] Finally, this application developed an end-to-end subtitle software that integrates the microphone array of a mobile phone. Through a combination of hardware and software, it can achieve sound source localization and speaker separation for multiple people offline, thereby improving speech recognition and role separation capabilities. It can also simultaneously obtain the interface of the conferencing software to acquire online single-person subtitles and upload multi-person subtitles to the conferencing software.

[0455] To illustrate the various scenarios in which the embodiments of this application can be applied, and in conjunction with the various implementation schemes of the embodiments of this application described above, specific examples are given below to explain the implementation process of the embodiments of this application in various scenarios. A mobile phone is used as an example for illustration.

[0456] It should be noted that the above-described method embodiments, or the various possible implementations of the method embodiments, can be executed individually, or, provided there are no contradictions, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0457] It should be noted that the information display method provided in this application embodiment can be executed by an information display device. This application embodiment uses an information display device executing the information display method as an example to illustrate the information display device provided in this application embodiment.

[0458] Figure 12 A schematic diagram of a possible structure of the information display device involved in an embodiment of this application is shown. For example... Figure 12 As shown, the information display device 70 may include a display module 71.

[0459] The display module 71 is used to display the conference interface of the first conference, which includes N participants, where N is an integer greater than 1; and when the received first audio includes the speaking voices of M participants, the conference interface displays M participant identifiers and M speaking information; wherein, the N participants include M participants; one participant identifier is used to indicate one participant; and M is a positive integer less than or equal to N.

[0460] In one possible implementation, the display module 71 is specifically used to display the participant information corresponding to each parameter person among the M participants in the identifier display area of ​​the meeting interface; the participant information includes at least one of the following: participant avatar, participant name, meeting name of the first meeting, participant speaking time, and participant location information; and to display the speaking information corresponding to each participant in the text display area corresponding to the participant identifier of each of the M participants.

[0461] In one possible implementation, the information display device 70 further includes a processing module. The processing module is configured to, when the received first audio includes the speaking voices of M participants, before displaying M participant identifiers and M speaking information on the conference interface, extract voiceprint features from the first audio to obtain M voiceprint feature information; compare each voiceprint feature information with voiceprint feature information stored in the database; and determine the participant name corresponding to the voiceprint feature information with a matching degree greater than a first threshold as the participant name for each participant.

[0462] In one possible implementation, the conference interface includes audio identifiers corresponding to N participants; the information display device 70 further includes a receiving module. The processing module is further configured to acquire the participant names of the N participants. The display module 71 is further configured to, when detecting that at least one participant among the N participants does not match any of the M voiceprint features, display the participant names of K participants in the user name display area corresponding to the audio identifier of at least one participant in the conference interface, where K is a positive integer less than M. The receiving module is configured to receive a first input for a first participant name and a first audio identifier; the first participant name is any one of the K participant names corresponding to the K participants; the first audio identifier is any one of the K audio identifiers corresponding to the K participants. The processing module is further configured to, in response to the first input received by the receiving module, determine the participant name corresponding to the voiceprint feature of the audio of the first audio identifier, and obtain the K participant names corresponding to the voiceprint feature of the audio indicated by the K audio identifiers.

[0463] In one possible implementation, the information display device 70 further includes a processing module. The processing module is further configured to: perform text conversion processing on the first audio to obtain first text before the display module displays the participant identifiers corresponding to the M participants and the M speaking information on the conference interface; determine the industry sector corresponding to the first text based on the semantic information of the first text; and update the first text based on the industry sector corresponding to the first text to obtain the speaking information corresponding to the first text.

[0464] In one possible implementation, the information display device 70 further includes a receiving module. The display module 71 is specifically configured to display at least one domain identifier based on the industry domain corresponding to the first text. The receiving module is configured to receive a second input for a target domain identifier among the at least one domain identifier. The processing module is specifically configured to, in response to the second input received by the receiving module, update the first text based on the target industry domain corresponding to the target domain identifier, to obtain the speech information corresponding to the first text.

[0465] In one possible implementation, the aforementioned processing module is specifically used to replace the first text segment in the first text with the professional name based on the professional terminology corresponding to the target industry field, thereby obtaining the speech information; wherein, the semantic matching between the first text segment and the professional terminology is performed.

[0466] In one possible implementation, the conference interface includes a caption summary control; the information display device 70 further includes a receiving module and a processing module. The receiving module is used to receive a third input to the caption summary control after the display module has displayed M participant identifiers and M speech information on the conference interface. The processing module is used to, in response to the third input received by the receiving module, perform a text summary based on the speech information corresponding to the M participants, and obtain a summary text. The display module 71 is also used to display the summary text.

[0467] In one possible implementation, the information display device 70 further includes a processing module. The processing module is further configured to, before the display module detects that the received first audio includes the speaking voices of M participants out of N participants, perform echo cancellation on the received second audio to obtain a third audio; perform sound source localization on the third audio to obtain the sound incidence angle corresponding to the third audio; perform noise reduction processing on the third audio based on the audio incidence angle to obtain a fourth audio; and perform de-reverberation on the fourth audio to obtain the first audio.

[0468] In the information display device provided in this application embodiment, since the electronic device can display the participant identifier and the speaking information corresponding to each participant in the conference interface of the first conference, the user can determine the participant to which the speaking information belongs by the participant identifier, avoiding the user's inability to determine the participant corresponding to the subtitle during the conference, and improving the flexibility of the electronic device in displaying subtitle text.

[0469] The information display device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0470] The information display device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0471] The information display device provided in this application embodiment can implement the various processes implemented in the above method embodiments, and will not be described again here to avoid repetition.

[0472] Optionally, such as Figure 13 As shown, this application embodiment also provides an electronic device 90, including a processor 91 and a memory 92. The memory 92 stores a program or instructions that can run on the processor 91. When the program or instructions are executed by the processor 91, they implement the various steps of the above-described information display method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0473] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0474] Figure 14 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0475] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0476] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 14 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0477] The display unit 106 is used to display the conference interface of the first conference, which includes N participants, where N is an integer greater than 1; and when the received first audio includes the speaking voices of M participants, the conference interface displays M participant identifiers and M speaking information; wherein, the N participants include M participants; one participant identifier is used to indicate one participant; and M is a positive integer less than or equal to N.

[0478] In some embodiments of this application, the display unit 106 is specifically used to display participant information corresponding to each parameter person among the M participants in the identifier display area of ​​the meeting interface; the participant information includes at least one of the following: participant avatar, participant name, meeting name of the first meeting, participant speaking time, and participant location information; and to display the speaking information corresponding to each participant in the text display area corresponding to the participant identifier of each of the M participants.

[0479] In some embodiments of this application, the processor 110 is further configured to, when the received first audio includes the speaking voices of M participants, before displaying the M participant identifiers and M speaking information on the conference interface, extract voiceprint features from the first audio to obtain M voiceprint feature information; compare each voiceprint feature information with the voiceprint feature information stored in the database; and determine the participant name corresponding to the voiceprint feature information with a matching degree greater than a first threshold as the participant name of each participant.

[0480] In some embodiments of this application, the conference interface includes audio identifiers corresponding to the N participants; the processor 110 is further configured to obtain the participant names of the N participants. The display unit 106 is further configured to, when detecting that at least one of the N participants does not match any of the M voiceprint features, display the participant names of the K participants in the user name display area corresponding to the audio identifier of the at least one participant in the conference interface, where K is a positive integer less than M. The user input unit 107 is configured to receive a first input for a first participant name and a first audio identifier; the first participant name is any one of the K participant names corresponding to the K participants; the first audio identifier is any one of the K audio identifiers corresponding to the K participants. The processor 110 is further configured to, in response to the first input, determine the participant name corresponding to the voiceprint feature of the audio of the first audio identifier, and obtain the K participant names corresponding to the voiceprint feature of the audio indicated by the K audio identifiers.

[0481] In some embodiments of this application, the processor 110 is further configured to perform text conversion processing on the first audio to obtain the first text before displaying the participant identifiers and M speech information corresponding to the M participants on the conference interface; and determine the industry field corresponding to the first text based on the semantic information of the first text; and update the first text based on the industry field corresponding to the first text to obtain the speech information corresponding to the first text.

[0482] In some embodiments of this application, the processor 110 is specifically configured to display at least one domain identifier based on the industry domain corresponding to the first text. The user input unit 107 is configured to receive a second input for a target domain identifier among the at least one domain identifier. The processor 110 is further configured to, in response to the second input, update the first text based on the target industry domain corresponding to the target domain identifier, to obtain the speech information corresponding to the first text.

[0483] In some embodiments of this application, the processor 110 is specifically used to replace the first text segment in the first text with the professional name based on the professional name corresponding to the target industry field to obtain the speech information; wherein, the first text segment and the professional name are semantically matched.

[0484] In some embodiments of this application, the meeting interface includes a caption summary control; the user input unit 107 is further configured to receive a third input to the caption summary control after displaying M participant identifiers and M speech information on the meeting interface. The processor 110 is further configured to, in response to the third input, perform a text summary based on the speech information corresponding to the M participants to obtain a summary text. The display unit 106 is further configured to display the summary text.

[0485] In some embodiments of this application, the processor 110 is further configured to, upon detecting that the received first audio includes the speaking voices of M participants out of N participants, perform echo cancellation on the received second audio to obtain a third audio before displaying the participant identifiers and M speaking information corresponding to the M participants on the conference interface; perform sound source localization on the third audio to obtain the sound incidence angle corresponding to the third audio; perform noise reduction processing on the third audio based on the audio incidence angle to obtain a fourth audio; and perform de-reverberation on the fourth audio to obtain a first audio.

[0486] In the electronic device provided in this application embodiment, since the electronic device can display the participant identifier and the speaking information of each participant in the meeting interface of the first meeting, the user can determine the participant to which the speaking information belongs by the participant identifier, avoiding the user's inability to determine the participant corresponding to the subtitle during the meeting, and improving the flexibility of the electronic device in displaying subtitle text.

[0487] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0488] For details on the beneficial effects of the various implementation methods in this embodiment, please refer to the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, these will not be repeated here.

[0489] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0490] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0491] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0492] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0493] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0494] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0495] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0496] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.

[0497] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0498] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0499] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An information display method, characterized in that, The method includes: The meeting interface of the first meeting is displayed. The first meeting includes N participants, where N is an integer greater than 1. If the first audio received includes the voices of M participants, the conference interface displays M participant identifiers and M speech messages. Wherein, the N participants include the M participants; a participant identifier is used to indicate a participant; M is a positive integer less than or equal to N.

2. The method according to claim 1, characterized in that, The meeting interface displays the participant identifiers and M speech information corresponding to the M participants, including: In the identification display area of ​​the meeting interface, the participant information corresponding to each parameter person among the M participants is displayed; the participant information includes at least one of the following: participant avatar, participant name, meeting name of the first meeting, participant speaking time, and participant location information; In the text display area corresponding to the participant identifier of each of the M participants, the speaking information of each participant is displayed.

3. The method according to claim 2, characterized in that, In the case that the received first audio includes the speaking voices of M participants, before displaying the M participant identifiers and M speaking information on the conference interface, the method further includes: The first audio is subjected to voiceprint feature extraction to obtain M voiceprint feature information; Each voiceprint feature is compared with the voiceprint feature information stored in the database; The names of the participants whose voiceprint features match the first threshold are identified as the participant names for each participant.

4. The method according to claim 3, characterized in that, The meeting interface includes audio identifiers corresponding to the N participants; the method further includes: Obtain the names of the N participants; If at least one of the N participants is detected to be inconsistent with all M voiceprint features, the names of the K participants, where K is a positive integer less than M, are displayed in the user name display area corresponding to the audio identifier of the at least one participant in the meeting interface. Receive a first input for the name of the first participant and a first audio identifier; the first participant name is any one of the K participant names corresponding to the K participants; the first audio identifier is any one of the K audio identifiers corresponding to the K participants; In response to the first input, the participant name corresponding to the voiceprint feature of the audio of the first audio identifier is determined, and K participant names corresponding to the voiceprint features of the audio indicated by K audio identifiers are obtained.

5. The method according to claim 1, characterized in that, Before displaying the participant identifiers and M speaking messages corresponding to the M participants on the meeting interface, the method further includes: The first audio is converted into text to obtain the first text; Based on the semantic information of the first text, the industry sector corresponding to the first text is determined; Based on the industry sector corresponding to the first text, the first text is updated to obtain the corresponding speaking information.

6. The method according to claim 5, characterized in that, The step of updating the first text based on the industry sector corresponding to the first text to obtain the speech information corresponding to the first text includes: Based on the industry sector corresponding to the first text, display at least one sector identifier; Receive a second input for a target domain identifier among the at least one domain identifier; In response to the second input, the first text is updated based on the target industry field corresponding to the target field identifier to obtain the speech information corresponding to the first text.

7. The method according to claim 6, characterized in that, The step of updating the first text based on the target industry domain corresponding to the target domain identifier to obtain the speech information corresponding to the first text includes: Based on the professional terms corresponding to the target industry field, the first text segment in the first text is replaced with the professional name to obtain the speaking information; The first text segment is semantically matched with the technical term.

8. The method according to claim 1, characterized in that, The meeting interface includes a caption summary control; After displaying M participant identifiers and M speaking messages on the meeting interface, the method further includes: Receive a third input to the subtitle summary control; In response to the third input, a text summary is generated based on the speaking information of the M participants to obtain the summary text. Display the summary text.

9. The method according to claim 1, characterized in that, Before displaying the participant identifiers and M speech messages corresponding to the M participants on the conference interface, when the received first audio is detected to include the speaking voices of M participants out of the N participants, the method further includes: Echo cancellation is performed on the received second audio to obtain the third audio; The sound source of the third audio is located to obtain the sound incident angle corresponding to the third audio. Based on the audio incident angle, the third audio is subjected to noise reduction processing to obtain the fourth audio; The fourth audio signal is de-reverberated to obtain the first audio signal.

10. An information display device, characterized in that, The information display device: The display module is used to display the meeting interface of the first meeting, which includes N participants, where N is an integer greater than 1. And if the first audio received includes the speaking voices of M participants, the conference interface displays M participant identifiers and M speaking messages; Wherein, the N participants include the M participants; a participant identifier is used to indicate a participant; M is a positive integer less than or equal to N.

11. The apparatus according to claim 10, characterized in that, The display module is specifically used to display the participant information corresponding to each parameter person among the M participants in the identification display area of ​​the meeting interface; the participant information includes at least one of the following: participant avatar, participant name, meeting name of the first meeting, participant speaking time, and participant location information; And in the text display area corresponding to the participant identifier of each of the M participants, the speaking information of each participant is displayed.

12. The apparatus according to claim 11, characterized in that, The information display device further includes: a processing module; The processing module is used to extract voiceprint features from the first audio received by the display module, before displaying M participant identifiers and M speech information on the conference interface, to obtain M voiceprint feature information. Each voiceprint feature is then compared with the voiceprint feature information stored in the database. And the participant names corresponding to the voiceprint feature information with a matching degree greater than the first threshold are determined as the participant names of each participant.

13. The apparatus according to claim 12, characterized in that, The conference interface includes audio identifiers corresponding to the N participants; the information display device further includes a receiving module. The processing module is also used to obtain the names of the N participants; The display module is further configured to, when detecting that at least one of the N participants does not match any of the M voiceprint features, display the names of the K participants in the user name display area corresponding to the audio identifier of the at least one participant in the conference interface, where K is a positive integer less than M; The receiving module is configured to receive a first input of the first participant's name and a first audio identifier; the first participant's name is any one of the K participant names corresponding to the K participants; the first audio identifier is any one of the K audio identifiers corresponding to the K participants; The processing module is further configured to, in response to the first input received by the receiving module, determine the participant name corresponding to the voiceprint feature of the audio of the first audio identifier, and obtain the K participant names corresponding to the voiceprint features of the audio indicated by the K audio identifiers.

14. The apparatus according to claim 10, characterized in that, The information display device further includes: a processing module; The processing module is further configured to perform text conversion processing on the first audio to obtain the first text before the display module displays the participant identifiers and M speaking information corresponding to the M participants on the conference interface; Based on the semantic information of the first text, the industry sector corresponding to the first text is determined; And based on the industry sector corresponding to the first text, the first text is updated to obtain the speech information corresponding to the first text.

15. The apparatus according to claim 14, characterized in that, The information display device further includes: a receiving module; The display module is specifically used to display at least one industry identifier based on the industry field corresponding to the first text; The receiving module is configured to receive a second input for the target domain identifier in the at least one domain identifier; The processing module is specifically used to respond to the second input received by the receiving module, and update the first text based on the target industry field corresponding to the target field identifier, so as to obtain the speech information corresponding to the first text.

16. The apparatus according to claim 15, characterized in that, The processing module is specifically used to replace the first text segment in the first text with the professional name corresponding to the target industry field to obtain the speech information; wherein, the first text segment and the professional name are semantically matched.

17. The apparatus according to claim 10, characterized in that, The conference interface includes a caption summary control; the information display device further includes a receiving module and a processing module. The receiving module is used to receive a third input to the subtitle summary control after the display module displays M participant identifiers and M speaking information on the conference interface; The processing module is used to respond to the third input received by the receiving module, and to perform a text summary based on the speaking information of the M participants to obtain a summary text; The display module is also used to display the summary text.

18. The apparatus according to claim 10, characterized in that, The information display device further includes: a processing module; The processing module is further configured to perform echo cancellation on the received second audio to obtain a third audio before the display module displays the participant identifiers and M speech information corresponding to the M participants on the conference interface when the display module detects that the received first audio includes the speaking voices of M participants out of the N participants; The sound source of the third audio is located to obtain the sound incident angle corresponding to the third audio. Based on the audio incident angle, the third audio is subjected to noise reduction processing to obtain the fourth audio; The fourth audio signal is de-reverberated to obtain the first audio signal.

19. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the information display method as described in any one of claims 1 to 9.

20. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the information display method as described in any one of claims 1 to 9.