Audio processing method, apparatus, device, storage medium and computer program product

By acquiring users' audio content and temporal correlation information during online meetings to create virtual groups, the problem of inconsistent discussion topics in online meetings was solved, meeting efficiency was improved, and resource waste was reduced.

CN119276644BActive Publication Date: 2025-11-18MIGU CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411266595.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-11-18
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

In existing online conferencing systems, participants in the same meeting room often discuss different topics, leading to inefficiency and requiring frequent group discussions and meeting room merging, resulting in a waste of resources and time.

Method used

By acquiring information on the correlation and temporal correlation of audio content among users in online meetings, virtual groups are formed, and the audio of group members is played to ensure that participants only hear the voices of their own group members.

Benefits of technology

It enables automatic grouping based on discussion topics within the same meeting room, improving meeting efficiency, avoiding the need for frequent changes in meeting rooms, and enhancing resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119276644B_ABST
    Figure CN119276644B_ABST
Patent Text Reader

Abstract

The application provides an audio processing method, device, equipment, storage medium and computer program product, wherein the audio processing method comprises: obtaining audio content correlation information and audio timing correlation information between a first user and a second user in an online conference; performing virtual group division on the first user and the second user according to the audio content correlation information and the audio timing correlation information; and playing audio of each member in a virtual group in which the first user is located, the first user being a terminal user, and the second user being other terminal user in the online conference except the first user. The application can support division of different groups based on analysis of audio of participants, and all participants in the online conference only need to enter the same conference room to only hear the speech of the members in the group in which the participant is located, without repeatedly entering and exiting different conference rooms, thereby improving the conference efficiency and solving the problem of low efficiency of the online conference solution in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method, apparatus, device, storage medium, and computer program product. Background Technology

[0002] Current meeting systems typically involve a host setting a meeting topic and creating an online meeting room, then notifying relevant participants. However, during the meeting, participants from different project teams may speak based on their own perspectives, causing the meeting topic to diverge from the main theme of the meeting room. This results in multiple topics being discussed simultaneously within a single meeting room, and the simultaneous presentations of people on different topics can interfere with each other. For example, in a troubleshooting review meeting, after the host describes the problem, front-end developers might discuss issues with the app (application), while back-end developers might discuss issues with service nodes. This creates two separate groups discussing these two topics, causing the audio in the online meeting room to interfere with each other, leading to inefficiency or even preventing further discussion. In such cases, it is necessary to create additional meeting rooms and distribute some participants to these new rooms.

[0003] In other words, in the existing technology, participants in the same online meeting room may need to discuss different topics on the same issue. Specifically, different online meeting rooms are set up for people who want to solve the same issue but have different focuses on the discussion topics. After the discussion in the separate meeting rooms is completed, they enter the same meeting room to discuss. If disagreements arise again at this time, it is necessary to repeat the cycle of discussion in separate meeting rooms and then entering the same meeting room. This will result in waste of resources, waste of time, and low efficiency.

[0004] As shown above, existing online conferencing solutions suffer from inefficiencies and other problems. Summary of the Invention

[0005] The purpose of this application is to provide an audio processing method, apparatus, device, storage medium, and computer program product to solve the problem of low efficiency in existing online conferencing solutions.

[0006] To address the aforementioned technical problems, embodiments of this application provide an audio processing method, including:

[0007] Obtain audio content correlation information and audio temporal correlation information between the first user and the second user in an online meeting;

[0008] Based on the audio content correlation information and audio time sequence correlation information, the first user and the second user are divided into virtual groups;

[0009] Play the audio of each member in the virtual group to which the first user is located. The first user is the user of this terminal, and the second user is the other terminal users in the online meeting besides the first user.

[0010] Optionally, obtain audio content relevance information between the first user and the second user in the online meeting, including:

[0011] Obtain a first set of words from a first text, wherein the first text is the text of the first audio data of the first user in the online meeting, and the first set of words includes: the words in the first text whose frequency of occurrence is ranked in the top Z positions, where Z is an integer greater than or equal to 1;

[0012] Obtain a second vocabulary set from the second text, wherein the second text is the text of the second audio data of the second user in the online meeting, and the second vocabulary set includes: the words whose frequency of occurrence is ranked in the top Z positions in the second text; the second vocabulary set has the same total number of elements as the first vocabulary set;

[0013] Obtain the number of words that match between the first vocabulary set and the second vocabulary set;

[0014] Based on the number of words and the total number of elements, the similarity between the first word set and the second word set is determined, and the similarity is used as the audio content relevance information between the first user and the second user.

[0015] Optionally, obtain audio temporal correlation information between the first user and the second user in the online meeting, including:

[0016] Calculate the audio continuity between the first user and the second user;

[0017] Based on the audio continuity, audio temporal correlation information between the first user and the second user is obtained.

[0018] Optionally, calculating the audio continuity between the first user and the second user includes:

[0019] Obtain a first time sequence consisting of the end times of each speech segment of the first user in the first audio data; the first audio data is the audio data of the first user in the online meeting;

[0020] Obtain a second time series consisting of the start times of each speech segment of the second user in the second audio data; the second audio data is the audio data of the second user in the online meeting;

[0021] Based on the first time series and the second time series, obtain the audio interval value between the first user and the second user;

[0022] The audio continuity between the first user and the second user is calculated based on the audio interval value.

[0023] Optionally, obtaining the second time series formed by the start times of each speech interval of the second user in the second audio data includes:

[0024] Determine the silence interval of the first user corresponding to the speaking voice interval of the second user; wherein, the speaking voice interval of the second user refers to the speaking voice interval of the second user in the second audio data of the online meeting, and the silence interval of the first user refers to the silence interval of the first user in the first audio data of the online meeting.

[0025] Based on the time relationship between the silence interval and the speech interval, determine the start time of the second user's speech in the second audio data.

[0026] Based on the determined start time of the speech, the second time sequence is obtained.

[0027] Optionally, determining the start time of the second user's speech in the second audio data based on the time relationship between the silence interval and the speech interval includes at least one of the following:

[0028] If the speaking voice interval overlaps with the silent interval, and the actual speaking start time of the second user corresponding to the speaking voice interval is later than the start time of the silent interval, and the interval between the actual speaking start time and the start time of the silent interval is less than a first threshold, then the actual speaking start time is taken as the speaking start time of the speaking voice interval.

[0029] If the speaking voice interval overlaps with the silent interval, and the actual start time of the second user corresponding to the speaking voice interval is earlier than the start time of the silent interval, then the actual start time of the speaking voice interval shall be taken as the start time of the speaking voice interval.

[0030] If the speaking voice interval overlaps with the silent interval, and the actual speaking start time of the second user corresponding to the speaking voice interval is later than the start time of the silent interval, and the interval between the actual speaking start time and the start time of the silent interval is greater than or equal to a second threshold, then the start time of the silent interval is taken as the speaking start time of the speaking voice interval.

[0031] If the speaking voice interval does not overlap with the silence interval, then the start time of the silence interval is taken as the start time of the speaking voice interval.

[0032] This application also provides an audio processing apparatus, including:

[0033] The first acquisition module is used to acquire audio content correlation information and audio timing correlation information between the first user and the second user in the online meeting.

[0034] The first processing module is used to divide the first user and the second user into virtual groups based on the audio content correlation information and the audio time sequence correlation information.

[0035] The first playback module is used to play the audio of each member in the virtual group to which the first user is located. The first user is the user of this terminal, and the second user is the other terminal users in the online meeting besides the first user.

[0036] Optionally, obtain audio content relevance information between the first user and the second user in the online meeting, including:

[0037] Obtain a first set of words from a first text, wherein the first text is the text of the first audio data of the first user in the online meeting, and the first set of words includes: the words in the first text whose frequency of occurrence is ranked in the top Z positions, where Z is an integer greater than or equal to 1;

[0038] Obtain a second vocabulary set from the second text, wherein the second text is the text of the second audio data of the second user in the online meeting, and the second vocabulary set includes: the words whose frequency of occurrence is ranked in the top Z positions in the second text; the second vocabulary set has the same total number of elements as the first vocabulary set;

[0039] Obtain the number of words that match between the first vocabulary set and the second vocabulary set;

[0040] Based on the number of words and the total number of elements, the similarity between the first word set and the second word set is determined, and the similarity is used as the audio content relevance information between the first user and the second user.

[0041] Optionally, obtain audio temporal correlation information between the first user and the second user in the online meeting, including:

[0042] Calculate the audio continuity between the first user and the second user;

[0043] Based on the audio continuity, audio temporal correlation information between the first user and the second user is obtained.

[0044] Optionally, calculating the audio continuity between the first user and the second user includes:

[0045] Obtain a first time sequence consisting of the end times of each speech segment of the first user in the first audio data; the first audio data is the audio data of the first user in the online meeting;

[0046] Obtain a second time series consisting of the start times of each speech segment of the second user in the second audio data; the second audio data is the audio data of the second user in the online meeting;

[0047] Based on the first time series and the second time series, obtain the audio interval value between the first user and the second user;

[0048] The audio continuity between the first user and the second user is calculated based on the audio interval value.

[0049] Optionally, obtaining the second time series formed by the start times of each speech interval of the second user in the second audio data includes:

[0050] Determine the silence interval of the first user corresponding to the speaking voice interval of the second user; wherein, the speaking voice interval of the second user refers to the speaking voice interval of the second user in the second audio data of the online meeting, and the silence interval of the first user refers to the silence interval of the first user in the first audio data of the online meeting.

[0051] Based on the time relationship between the silence interval and the speech interval, determine the start time of the second user's speech in the second audio data.

[0052] Based on the determined start time of the speech, the second time sequence is obtained.

[0053] Optionally, determining the start time of the second user's speech in the second audio data based on the time relationship between the silence interval and the speech interval includes at least one of the following:

[0054] If the speaking voice interval overlaps with the silent interval, and the actual speaking start time of the second user corresponding to the speaking voice interval is later than the start time of the silent interval, and the interval between the actual speaking start time and the start time of the silent interval is less than a first threshold, then the actual speaking start time is taken as the speaking start time of the speaking voice interval.

[0055] If the speaking voice interval overlaps with the silent interval, and the actual start time of the second user corresponding to the speaking voice interval is earlier than the start time of the silent interval, then the actual start time of the speaking voice interval shall be taken as the start time of the speaking voice interval.

[0056] If the speaking voice interval overlaps with the silent interval, and the actual speaking start time of the second user corresponding to the speaking voice interval is later than the start time of the silent interval, and the interval between the actual speaking start time and the start time of the silent interval is greater than or equal to a second threshold, then the start time of the silent interval is taken as the speaking start time of the speaking voice interval.

[0057] If the speaking voice interval does not overlap with the silence interval, then the start time of the silence interval is taken as the start time of the speaking voice interval.

[0058] This application also provides an audio processing device, including a memory, a processor, and a program stored in the memory and executable on the processor; when the processor executes the program, it implements the above-described audio processing method.

[0059] This application also provides a readable storage medium storing a program that, when executed by a processor, implements the steps in the audio processing method described above.

[0060] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the above-described audio processing method.

[0061] The beneficial effects of the above technical solution in this application are as follows:

[0062] In the above solution, the audio processing method obtains audio content correlation information and audio timing correlation information between the first user and the second user in the online meeting; based on the audio content correlation information and audio timing correlation information, it divides the first user and the second user into virtual groups; it plays the audio of each member in the virtual group to which the first user belongs, where the first user is the terminal user and the second user is the other terminal user in the online meeting besides the first user; it can support the division of participants into groups with different topics based on the analysis of the participants' audio, and all participants in the online meeting only need to enter the same meeting room to hear only the voices of the members in their own group, without having to repeatedly enter and exit different meeting rooms, thereby improving meeting efficiency and solving the problem of low efficiency in existing online meeting solutions. Attached Figure Description

[0063] Figure 1 This is a schematic diagram of the audio processing method according to an embodiment of this application;

[0064] Figure 2 This is a schematic diagram illustrating the specific implementation flow of the audio processing method according to an embodiment of this application;

[0065] Figure 3 This is a schematic diagram of single-person audio discretization in an embodiment of this application;

[0066] Figure 4 This is a schematic diagram of multi-person audio discretization in an embodiment of this application;

[0067] Figure 5 This is a schematic diagram illustrating the timing analysis of user speech in an embodiment of this application.

[0068] Figure 6 This is a schematic diagram of the audio processing device according to an embodiment of this application. Detailed Implementation

[0069] To make the technical problems, technical solutions and advantages of this application clearer, a detailed description will be provided below in conjunction with the accompanying drawings and specific embodiments.

[0070] This application addresses the problem of low efficiency in existing online conferencing solutions by providing an audio processing method, such as... Figure 1 As shown, it includes:

[0071] Step 11: Obtain the audio content correlation information and audio timing correlation information between the first user and the second user in the online meeting;

[0072] Step 12: Based on the audio content relevance information and audio temporal relevance information, divide the first user and the second user into virtual groups;

[0073] Step 13: Play the audio of each member in the virtual group to which the first user is located. The first user is the user of this terminal, and the second user is the other terminal users in the online meeting besides the first user.

[0074] The audio content relevance information can represent the similarity of topics discussed in the meeting, such as the proportion of similar words between content 1 published by user A and content 2 published by user B; the audio temporal relevance information can represent the continuity between speeches in the meeting, such as the degree of correlation between the speech intervals of user A and user B, etc., which are not limited here.

[0075] The audio processing method provided in this embodiment acquires audio content correlation information and audio timing correlation information between a first user and a second user in an online meeting; based on the audio content correlation information and audio timing correlation information, it divides the first user and the second user into virtual groups; it plays the audio of each member in the virtual group to which the first user belongs, where the first user is the terminal user and the second user is the other terminal user in the online meeting besides the first user; it can support the division of participants into groups with different topics based on the analysis of the participants' audio (such as the first audio data and the second audio data mentioned above), and all participants in the online meeting only need to enter the same meeting room to hear only the voices of the members in their own group, without having to repeatedly enter and exit different meeting rooms, thereby improving meeting efficiency and solving the problem of low efficiency in existing online meeting solutions.

[0076] The method for obtaining audio content relevance information between a first user and a second user in an online meeting includes: obtaining a first vocabulary set from a first text, where the first text is the text of the first audio data of the first user in the online meeting, and the first vocabulary set includes the words with the highest frequency (ranked by Z) in the first text, where Z is an integer greater than or equal to 1; obtaining a second vocabulary set from a second text, where the second text is the text of the second audio data of the second user in the online meeting, and the second vocabulary set includes the words with the highest frequency (ranked by Z) in the second text; the second vocabulary set has the same total number of elements as the first vocabulary set; obtaining the number of matching words between the first vocabulary set and the second vocabulary set; determining the similarity between the first vocabulary set and the second vocabulary set based on the number of words and the total number of elements, and using the similarity as the audio content relevance information between the first user and the second user. This method can accurately obtain audio content relevance information. The "number of matching words" can include the number of completely identical and / or similar words, such as the same word or synonyms, and is not limited here.

[0077] In this embodiment, obtaining audio temporal correlation information between a first user and a second user in an online meeting includes: calculating the audio coherence between the first user and the second user; and obtaining audio temporal correlation information between the first user and the second user based on the audio coherence. This allows for accurate acquisition of audio temporal correlation information. Specifically, "calculating the audio coherence between the first user and the second user" may include: discretizing the audio data of the first user and the second user based on temporal sequence to obtain discretized speech data; and in the discretized speech data, marking speech segments within a preset frequency range and exceeding a preset time as a first value, and marking speech segments outside the preset frequency range and / or not exceeding a preset time as a second value; and calculating the audio coherence between the first user and the second user based on the first and second values ​​marked in the temporal sequence, which can also be referred to as speech coherence. Regarding "obtaining audio temporal correlation information between the first user and the second user based on the audio coherence degree", it may include: obtaining audio temporal correlation information between the first user and the second user based on the audio coherence degree and a first formula; wherein, the first formula is: P2 = 1 - DX, where P2 represents the audio temporal correlation information and DX represents the audio coherence degree, but is not limited thereto.

[0078] The calculation of the audio continuity between the first user and the second user includes: obtaining a first time series consisting of the end times of each speech interval of the first user in the first audio data; the first audio data being the audio data of the first user in the online meeting; obtaining a second time series consisting of the start times of each speech interval of the second user in the second audio data; the second audio data being the audio data of the second user in the online meeting; obtaining an audio interval value between the first user and the second user based on the first time series and the second time series; and calculating the audio continuity between the first user and the second user based on the audio interval value. This allows for a more specific and accurate acquisition of the audio continuity. The operation of obtaining the first time series and / or the second time series can be performed based on the aforementioned first and second values, but is not limited to them; furthermore, a person's speech throughout the meeting can be divided into at least one speech interval based on silence, for example, the first audio data may include at least one silence interval and at least one speech interval.

[0079] In this embodiment, obtaining the second time series formed by the start times of each speech interval of the second user in the second audio data includes: determining the silence interval of the first user corresponding to the speech interval of the second user; wherein, the speech interval of the second user refers to the speech interval of the second user in the second audio data of the online meeting, and the silence interval of the first user refers to the silence interval of the first user in the first audio data of the online meeting; determining the speech start time of the speech interval of the second user in the second audio data according to the time relationship between the silence interval and the speech interval; and obtaining the second time series according to the determined speech start time. This allows for an accurate second time series. The interval can be a time interval, such as a speech interval being a time interval of speech; and / or, the time relationship includes time overlap, or, the time relationship includes time overlap, and the difference in time between the actual speech start time of the second user corresponding to the speech interval and the start time of the silence interval; but it is not limited thereto. The phrase "determining the silence interval of the first user corresponding to the speech interval of the second user" can include: determining the silence interval of the first user corresponding to the speech interval of the second user by sorting by time; for example, after sorting the speech intervals of the first user and the speech intervals of the second user respectively (which can be sorted by time sequence), the first silence interval of the first user is obtained, and the first silence interval is used as the silence interval corresponding to the first speech interval of the second user, but it is not limited to this.

[0080] The step of determining the start time of the second user's speech in the second audio data based on the time relationship between the silence interval and the speech interval includes at least one of the following: (1) if the speech interval overlaps with the silence interval, and the actual start time of the second user corresponding to the speech interval is later than the start time of the silence interval, and the interval between the actual start time of the speech and the start time of the silence interval is less than a first threshold, then the actual start time of the speech is taken as the start time of the speech interval; (2) if the speech interval overlaps with the silence interval, and the second user corresponding to the speech interval... If the actual start time of the speech is earlier than the start time of the silence interval, then the actual start time of the speech is taken as the start time of the speech voice interval; (3) If the speech voice interval overlaps with the silence interval, and the actual start time of the second user corresponding to the speech voice interval is later than the start time of the silence interval, and the interval between the actual start time of the speech and the start time of the silence interval is greater than or equal to the second threshold, then the start time of the silence interval is taken as the start time of the speech voice interval; (4) If the speech voice interval does not overlap with the silence interval, then the start time of the silence interval is taken as the start time of the speech voice interval. In this way, the start time of the second user's speech voice interval in the second audio data can be accurately obtained. Examples of the above are as follows:

[0081] Item (1), for example: the second user's speech interval 1 in the second audio data corresponds to the first user's silence interval 1 in the first audio data; the silence interval 1 is t1~t3, the actual speech start time of the second user corresponding to the speech interval 1 is t2, the interval between t2 and t1 is less than the first threshold, then the actual speech start time is t2 as the speech start time of the speech interval.

[0082] Item (2), for example: the second user's speech interval 1 in the second audio data corresponds to the first user's silence interval 1 in the first audio data; the silence interval 1 is t5~t6, the actual speech start time of the second user corresponding to the speech interval 1 is t4, and the actual speech end time of the second user corresponding to the speech interval 1 is between t5 and t6, then the actual speech start time is t4 as the speech start time of the speech interval.

[0083] Item (3), for example: the second user's speech interval 1 in the second audio data corresponds to the first user's silence interval 1 in the first audio data; the silence interval 1 is t7~t9, the actual speech start time of the second user corresponding to the speech interval 1 is t8, the interval between t8 and t7 is greater than or equal to the second threshold, then the start time t7 of the silence interval is taken as the speech start time of the speech interval.

[0084] For example, in item (4), the second user's speech interval 1 in the second audio data corresponds to the first user's silence interval 1 in the first audio data; the silence interval 1 is t10~t11, and the actual speech start time of the second user corresponding to the speech interval 1 is t12. Then, the start time t10 of the silence interval is taken as the speech start time of the speech interval.

[0085] In this embodiment, the first time series and the second time series contain the same number of time information items. The step of obtaining the audio interval value between the first user and the second user based on the first time series and the second time series includes: obtaining the average of the first differences between the time information at corresponding sorting positions in the first and second time series, as the audio interval value between the first user and the second user. This allows for accurate acquisition of the audio interval value. Specifically, the "average of the first differences" may include, but is not limited to, the average of the absolute values ​​of the first differences.

[0086] The step of calculating the audio continuity between the first user and the second user based on the audio interval value includes: obtaining a second difference between an empirical interval value and the audio interval value; and calculating the audio continuity between the first user and the second user based on the first difference, the second difference, and the empirical interval value. This allows for an accurate determination of the audio continuity. The first difference may specifically include, but is not limited to, the absolute value of the first difference.

[0087] In this embodiment, the step of virtually grouping the first user and the second user based on the audio content relevance information and the audio temporal relevance information includes: obtaining the audio similarity between the first user and the second user based on the audio content relevance information and the audio temporal relevance information; and virtually grouping the first user and the second user based on the audio similarity and a grouping threshold. This allows for precise grouping.

[0088] The step of obtaining audio content relevance information and audio timing relevance information between the first user and the second user in the online meeting includes: obtaining the audio content relevance information and audio timing relevance information according to a second method; the second method includes at least one of a periodic method and a condition-triggered method; and / or, the step of dividing the first user and the second user into virtual groups based on the audio content relevance information and audio timing relevance information includes: dividing the first user and the second user into virtual groups according to a third method; the third method includes at least one of a periodic method and a condition-triggered method. This can support dynamic grouping and further improve the accuracy of grouping. For example, a condition-triggered method might detect that the content of a member's speech in a virtual group has a low relevance to the content of other members' speech, i.e., the audio content relevance information indicates a relevance below a threshold value, which can trigger a re-division of the groups, such as returning the operation of "obtaining audio content relevance information and audio timing relevance information between the first user and the second user in the online meeting".

[0089] Furthermore, the audio processing method further includes: canceling the virtual group division according to the instruction information and playing the audio of all users in the online meeting. This makes the solution more user-friendly, supporting active or passive cancellation of virtual groups to meet diverse user needs. The instruction information may include: instructions input by the first user, and / or instructions input by other participating users, etc., which are not limited here.

[0090] The audio processing method provided in the embodiments of this application will be illustrated below with examples.

[0091] To address the aforementioned technical problems, this application provides an audio processing method, specifically an audio extraction method based on a multi-person virtual conference. This method supports dynamically dividing participants into groups based on audio analysis, allowing all participants to hear only the audio related to their chosen topic (i.e., their own group) simply by entering the same meeting room, eliminating the need for repeated entry and exit from different meeting rooms. This application primarily involves: extracting individual audio files from each participant in the conference; calculating the topic relevance factor (including content relevance factor P1 and timing relevance factor P2) for the content discussed by two participants; then, combining the weights of P1 and P2 to divide all participants in the same meeting room into different virtual groups based on the topic relevance; ensuring that each member of a virtual group can only hear the audio of their own group members; this operation corresponds to obtaining the audio content relevance information and audio timing relevance information between the first and second users in an online conference; dividing the first and second users into virtual groups based on the audio content relevance information and audio timing relevance information; and playing the audio of each member in the virtual group of the first user. Among them, the speech content relevance factor P1 corresponds to the above audio content relevance information, and the speech time sequence relevance factor P2 corresponds to the above audio time sequence relevance information.

[0092] Specifically, such as Figure 2 As shown, the specific contents of the embodiments of this application may include the following:

[0093] 1. Audio preprocessing;

[0094] (1) Process the voice data of each participant separately. For example, extract the audio data of the first participant and then extract high-frequency words; perform the same processing on the voice data of the second participant and also extract high-frequency words; the frequency of word occurrence can be obtained using the following formula:

[0095] Where TF(t,d) represents word frequency; the occurrence count(t,d) represents the number of times the word appears within the text length d, and the text length d corresponds to the extracted audio data of the person.

[0096] (2) P1 analysis of the relevance factor of the speech content;

[0097] The vocabulary of each participant is sorted from highest to lowest by word frequency TF(t,d), and the top 5 words are extracted as the high-frequency word groups for each participant. Then, a participant's high-frequency word groups are compared with those of other participants. For each matching word, the hit count N is incremented by 1. The relevance factor P1 between the participant's speech and that of other participants can be calculated using the following formula: N is the hit count of the high-frequency word groups (corresponding to the number of words mentioned above), and 5 is the total number of comparison data (corresponding to the total number of elements mentioned above). This results in a numerical value between 0 and 1, with a higher value indicating a higher degree of thematic similarity. This part of the operation corresponds to obtaining the first vocabulary set in the first text, wherein the first text is the text of the first audio data of the first user in the online meeting, and the first vocabulary set includes: the words with the highest frequency of occurrence in the first text, Z being an integer greater than or equal to 1; obtaining the second vocabulary set in the second text, wherein the second text is the text of the second audio data of the second user in the online meeting, and the second vocabulary set includes: the words with the highest frequency of occurrence in the second text; the second vocabulary set has the same total number of elements as the first vocabulary set; obtaining the number of matching words between the first vocabulary set and the second vocabulary set; determining the similarity between the first vocabulary set and the second vocabulary set based on the number of words and the total number of elements, and using the similarity as the audio content relevance information between the first user and the second user.

[0098] Formula 1 above is: P1 = N / 5.

[0099] For example, suppose participant A's high-frequency word grouping is {server, startup, weak network, front-end, crash}, and participant B's high-frequency word grouping is {weak network, front-end, crash, log, startup}. Then A and B have the four word groups "weak network", "front-end", "crash" and "startup" matched. Their content similarity is 80% (4 / 5), that is, the relevance factor P1 of their speech content is 80%, which is highly similar.

[0100] 2. Relevance factor P2 to the speaking time sequence;

[0101] This section mainly involves calculating the connectivity between each speech segment among participants to obtain temporal correlation parameters; specifically including:

[0102] (1) Audio data is discretized according to time sequence;

[0103] The audio data of each participant is discretized and normalized to obtain a discretized result, which is used to mark whether each person is speaking at each moment. Specifically, considering that the Hertz range of human speech is 85Hz to 1100Hz, the raw audio data of each person can be judged in chronological order. If the Hertz is within 85Hz to 1100Hz, it means that the person is speaking, which can be marked as 1; if it is outside this range, it means that the person is not speaking (i.e., silent), which can be marked as 0. Based on this, we obtain... Figure 3 The discretization result of the single-person audio is shown. Speech periods shorter than 1 second can be marked as 0 (no speaking) to exclude interference from factors such as coughing. Figure 3 The effect of discretizing single-person audio data is shown in a line graph. Non-zero values ​​indicate that the person is speaking, and zero values ​​indicate silence. The user in the graph is speaking from 0 to t1, silent from t1 to t2, speaking from t2 to t3, silent from t3 to t4, and speaking from t4 to t5.

[0104] In this embodiment of the application, to facilitate subsequent analysis, the audio results of all participants in the meeting (i.e., the aforementioned discretization results) can be merged and displayed intuitively on the same image. This allows for a longitudinal observation of the speaking situation of all participants from a timeline perspective. For example... Figure 4 As shown, at time t1, only users B and C were speaking in the entire conference room; at time t2, users A and B were silent, while users C and D were speaking; from Figure 4 As can be seen from this, when users B and D were speaking, the other was silent, and they formed a question-and-answer situation, which likely indicated that they were discussing the same issue. This time interval can be quantified in the future. Figure 4 In the (multi-person discretized data display), all horizontal lines represent non-zero data, corresponding to users speaking; different line types correspond to different users.

[0105] (2) P2 analysis of the correlation factor between speaking time sequence;

[0106] In a typical meeting room, participants on the same topic speak in turn, meaning only one person can speak on the same topic at a time. There may be brief overlaps and periods of silence (i.e., quiet periods). Therefore, this embodiment of the application determines the continuity between other users' speech and the speaker's speech based on the duration of silence for each user. Specifically, this can be quantified based on the time interval between the start and end times of the speaker's speech. Taking participant A (corresponding to the first user mentioned above) as an example, their discussion topic is tentatively set as Topic 1. Figure 5 The diagram shows his four silent time intervals during the meeting: Q1, Q2, Q3, and Q4. Other participants in Topic 1 (corresponding to user B in the diagram) mainly spoke during these four silent windows (i.e., silent time intervals). Figure 5 In the (user speaking time sequence analysis diagram): all horizontal lines represent non-zero data, corresponding to the user speaking at this time; different line types correspond to different users; Q1...Q4 are the 4 silent periods of participant A (i.e., user A), that is, the silent time intervals, which can be corresponding to the above-mentioned silent intervals.

[0107] In this embodiment, from the perspective of participant A, other potential participants with the common topic Topic 1 are identified, and the temporal similarity between participant A and participant B is calculated based on the silent periods of participant A; the temporal similarity is the speaking temporal correlation factor P2. Assuming participant A has M silent periods (Q1-Q4 in the diagram, meaning M is 4), the status of other participants is checked at the end of each speech by participant A; at this time, any other participant (such as participant B, which can correspond to the second user mentioned above) has only two states:

[0108] Status 1, Speaking: Figure 5 The Q2 time zone corresponds to the situation where participant B is speaking when participant A begins to remain silent. This usually occurs when participant B interrupts participant A's speech or when participant B is participating in a discussion on another topic. In this case, the start time of participant B's current speech segment (which can correspond to the start time of the second user's speech segment mentioned above) can be denoted as TY. i =t4, where the start time of this speech segment refers to the start time of this current speech, not the start time of participant B's first speech in the entire meeting. A person's speech throughout the meeting can be divided into at least one speech segment based on periods of silence. Based on this, in this state, the end time TX of participant A's current speech segment can be marked. i =t5, the start time of participant B's speech segment TY i =t4; The processing in this state can correspond to the above: if the speaking voice interval overlaps with the silent interval, and the actual speaking start time of the second user corresponding to the speaking voice interval is earlier than the start time of the silent interval, then the actual speaking start time (corresponding to t4 in the figure) is taken as the speaking start time of the speaking voice interval.

[0109] State 2, silent, can include the following three situations:

[0110] Scenario 1: After a short period of time, participant B begins to speak, corresponding to interval Q1 in the diagram; in this case, the end time of participant A's current speech segment can be marked as TX. i =t1, marking the start time TY of participant B's current speech segment. i=t2; The processing in this case can correspond to the above: if the speaking voice interval overlaps with the silence interval, and the actual speaking start time of the second user corresponding to the speaking voice interval is later than the start time of the silence interval, and the interval between the actual speaking start time and the start time of the silence interval is less than the first threshold, then the actual speaking start time (corresponding to t2 in the figure) is taken as the speaking start time of the speaking voice interval.

[0111] Scenario ②: After a long pause, participant B begins to speak, corresponding to interval Q3 in the diagram. In this case, after participant A muted and before participant B spoke, a third person might have spoken. In this embodiment, TS = 10s can be set based on data statistics to represent the time threshold for the presence of a third person speaking. Based on this, if the pause time (i.e., t8-t7 in the diagram) is determined to be greater than TS in interval Q3, then the end time of participant A's current speech segment is marked as TX. i =t7, marking the start time of participant B's speech segment as the end time of participant A's speech segment, i.e., TY i =t7; The processing in this case can correspond to the overlap between the speaking voice interval and the silence interval, and the actual speaking start time of the second user corresponding to the speaking voice interval (corresponding to t8 in the figure) is later than the start time of the silence interval (corresponding to t7 in the figure), and the interval between the actual speaking start time and the start time of the silence interval is greater than or equal to the second threshold. In this case, the start time of the silence interval is taken as the speaking start time of the speaking voice interval. This can help avoid the situation where the discussion topics of participant A and participant B are highly related, but the timing leads to a very low error P2.

[0112] Scenario 3: Before participant A speaks next, participant B remains silent, corresponding to interval Q4 in the diagram. In this case, participant B does not speak during the entire silent interval Q4, which may indicate that a third person is speaking; the end time of participant A's speech segment can be marked as TX. i =t10, marking the start time of participant B's speech segment as the end time of participant A's speech segment, i.e., TY i =t10; The processing in this case can correspond to the above: if the speech interval and the silence interval do not overlap, then the start time of the silence interval (corresponding to t10 in the figure) is taken as the start time of the speech interval; this can also help avoid the situation where the discussion topics of participant A and participant B are highly related, but the timing leads to a very low error P2.

[0113] Based on the above processing, the time sequence TX1...TX1 at the end of each speech by participant A can be obtained. m (This can correspond to the first time sequence mentioned above), and the time sequence TY1...TY at the beginning of each speech by participant B.m (This can correspond to the second time series mentioned above); TY1...TY m The values ​​in the table are different values ​​taken according to the above situation, and are not necessarily the actual start time of each speech by participant B.

[0114] Then, using the two sets of sequences mentioned above, we substitute them into Formula 2 below to obtain the average value T of the subtraction results of the two sets of data. This operation can correspond to the average value of the first difference between the time information of the corresponding sorting positions in the first time series and the second time series, as the audio interval value between the first user and the second user.

[0115] Formula 2 above is:

[0116] In this embodiment, TU can be set to 2s based on statistical analysis, where TU represents the empirical value of the voice interval between two people during normal communication. To facilitate the calculation of the dispersion of the statistical data time sequence relative to the empirical value TU, the difference ΔT between the mean T and the empirical value TU can be calculated using the following formula three; this operation corresponds to obtaining the second difference between the empirical value of the interval and the audio interval value as described above.

[0117] Formula 3 above is: ΔT = TU - T;

[0118] Based on the above processing, the variance of the difference (absolute value) between the end time of participant A's speech and the start time of participant B's speech, relative to the empirical value TU of the speaking interval, can be calculated (i.e., the variance of the speaking interval relative to the empirical value). This corresponds to the calculation of the audio continuity between the first user and the second user based on the first difference, the second difference, and the empirical value of the interval. Variance helps assess the stability and variability of the data. In a set of data, a smaller variance indicates a high degree of continuity between participants A and B's speech, suggesting they are likely discussing the same topic; a larger variance indicates a lower degree of continuity between participants A and B's speech, suggesting they are likely not discussing the same topic. Specifically, the variance DX relative to TU can be calculated using the following formula:

[0119]

[0120] Wherein, DX corresponds to the aforementioned audio coherence; the smaller the DX, the higher the speech coherence, and the larger the DX, the lower the speech coherence. Finally, the speech sequence correlation factor P2 can be obtained through the following inverse formula:

[0121] P2 = 1 - DX; this corresponds to obtaining the audio timing correlation information between the first user and the second user based on the aforementioned audio continuity.

[0122] 3. Virtual group division;

[0123] After the above processing, the content relevance factor P1 and the time sequence relevance factor P2 can be calculated. The similarity division of P1 involves coarse-grained division, which can be combined with P2 for calibration to achieve accurate grouping.

[0124] In this embodiment of the application, after statistical analysis of a large number of samples, it was found that the speech similarity P between participants A and B is composed of P1 and P2 in a certain proportion (that is, P1 and P2 are used to calculate the similarity P between participants A and B in a certain proportion), where the proportion of P1 can be 0.4 and the proportion of P2 can be 0.6. Based on this, the calculation formula for speech similarity P (which can correspond to the above audio similarity) is as follows:

[0125] P = P1 × 0.4 + P2 × 0.6;

[0126] Furthermore, an empirical value PU = 0.7 can be set. This value can be derived from statistical data and corresponds to the aforementioned grouping threshold. If the similarity P between two speakers is greater than or equal to PU, it can be confirmed that the two participants are discussing the same topic and can be grouped into the same virtual group. This operation corresponds to obtaining the audio similarity between the first user and the second user based on the audio content relevance information and audio temporal relevance information; and dividing the first user and the second user into virtual groups based on the audio similarity and the grouping threshold.

[0127] In this embodiment, the similarity P value of the speaker's speech to that of all other participants can be calculated. Those with P ≥ PU are grouped into the same virtual group; that is, participants are grouped according to the similarity P. For example, if there are participants A, B, C, and D, from the perspective of participant A, the similarity P relative to participant B needs to be calculated. B Similarity P relative to participant C C The similarity P relative to participant D D Then P B P C P D Compare each with PU. Those that satisfy the condition (P≥PU) are grouped together. For example, {A, B} are grouped together.

[0128] Based on the above division of virtual groups, any participant's terminal (which can correspond to the aforementioned terminal, i.e., the terminal corresponding to the first user or the terminal used by the first user) can already know the members belonging to the same virtual group as the user on its terminal side. After the conferencing software receives the voice data, it can filter and process it according to the virtual group to which the member belongs, sending only the voice data of the corresponding group members to the speaker. This operation corresponds to playing the audio of each member in the virtual group to which the first user belongs. For example, in the above analysis, participant A can only hear participant B's voice, while the voices of participants C and D are not played, thus avoiding interference with the discussion between participants A and B. At this time, participants C and D may also be in other groups, and they can only hear each other's voices. In this way, participants can conduct their own discussions simultaneously in different groups within the same online meeting room without affecting each other.

[0129] Therefore, the solution provided in this application can determine different voice discussion groups based on the speaking time sequence and content relevance of participants. This involves: calculating the content relevance factor and time sequence relevance factor between the audio data of all participants based on their audio data in an online meeting; and dividing the participants into virtual groups based on these factors. The audio of non-group members is then muted within each virtual group. The content relevance factor represents the similarity of the topics discussed in the meeting, and the time sequence relevance factor represents the continuity between speeches. If the content relevance factor indicates similarity in topics, and the time sequence relevance factor indicates sequential continuity in speeches, then it can be determined that the relevant participants are discussing the same topic and are subsequently grouped into the same virtual group. This method allows for group division based on the actual discussion situation in the meeting. Specifically:

[0130] In this embodiment, audio data of each participant in an online meeting can be extracted; high-frequency words can be extracted from the audio data of each participant, and a content relevance factor among all participants can be calculated based on the high-frequency words; the audio data of each participant can be discretized based on time sequence to obtain discretized speech data, and in the discretized speech data, speech segments within a preset frequency range and with a duration exceeding a preset time are marked as 1, and speech segments outside the preset frequency range or with a duration not exceeding the preset time are marked as 0; based on the 0 and 1 markings of each participant in time sequence, the speech connectivity between all participants can be obtained, and a time sequence relevance factor among all participants can be obtained based on the speech connectivity; the participants can be divided into virtual groups according to the content relevance factor and the time sequence relevance factor, and the speech of non-group members can be masked in each virtual group.

[0131] In summary, the embodiments of this application introduce a comprehensive analysis of the timing of participants' speeches and the relevance of their speeches, which can dynamically divide participants into groups and play corresponding audio.

[0132] This application also provides an audio processing device, such as... Figure 6 As shown, it includes:

[0133] The first acquisition module 61 is used to acquire audio content correlation information and audio timing correlation information between the first user and the second user in an online meeting.

[0134] The first processing module 62 is used to divide the first user and the second user into virtual groups based on the audio content correlation information and the audio time sequence correlation information.

[0135] The first playback module 63 is used to play the audio of each member in the virtual group to which the first user is located. The first user is the user of this terminal, and the second user is the other terminal users in the online meeting besides the first user.

[0136] The audio processing device provided in this embodiment acquires audio content correlation information and audio timing correlation information between a first user and a second user in an online meeting; based on the audio content correlation information and audio timing correlation information, it divides the first user and the second user into virtual groups; it plays the audio of each member in the virtual group to which the first user belongs, where the first user is the terminal user and the second user is the other terminal user in the online meeting besides the first user; it can support the division of participants into groups with different topics based on the analysis of the participants' audio, and all participants in the online meeting only need to enter the same meeting room to hear only the voices of the members in their own group, without having to repeatedly enter and exit different meeting rooms, thereby improving meeting efficiency and solving the problem of low efficiency in existing online meeting solutions.

[0137] In this embodiment, obtaining audio content relevance information between a first user and a second user in an online meeting includes: obtaining a first vocabulary set in a first text, wherein the first text is the text of the first audio data of the first user in the online meeting, and the first vocabulary set includes: the words whose frequency of occurrence in the first text is ranked in the top Z positions, where Z is an integer greater than or equal to 1; obtaining a second vocabulary set in a second text, wherein the second text is the text of the second audio data of the second user in the online meeting, and the second vocabulary set includes: the words whose frequency of occurrence in the second text is ranked in the top Z positions; the second vocabulary set and the first vocabulary set have the same total number of elements; obtaining the number of matching words between the first vocabulary set and the second vocabulary set; determining the similarity between the first vocabulary set and the second vocabulary set based on the number of words and the total number of elements, and using the similarity as the audio content relevance information between the first user and the second user.

[0138] The process of obtaining audio temporal correlation information between a first user and a second user in an online meeting includes: calculating the audio continuity between the first user and the second user; and obtaining audio temporal correlation information between the first user and the second user based on the audio continuity.

[0139] In this embodiment of the application, calculating the audio continuity between the first user and the second user includes: obtaining a first time series consisting of the end times of each speech segment of the first user in the first audio data; the first audio data being the audio data of the first user in the online meeting; obtaining a second time series consisting of the start times of each speech segment of the second user in the second audio data; the second audio data being the audio data of the second user in the online meeting; obtaining an audio interval value between the first user and the second user based on the first time series and the second time series; and calculating the audio continuity between the first user and the second user based on the audio interval value.

[0140] The step of obtaining the second time series formed by the start times of each speech segment of the second user in the second audio data includes: determining the silence segment of the first user corresponding to the speech segment of the second user; wherein, the speech segment of the second user refers to the speech segment of the second user in the second audio data of the online meeting, and the silence segment of the first user refers to the silence segment of the first user in the first audio data of the online meeting; determining the speech start time of the speech segment of the second user in the second audio data according to the time relationship between the silence segment and the speech segment; and obtaining the second time series according to the determined speech start time.

[0141] In this embodiment of the application, determining the start time of the second user's speech in the second audio data based on the time relationship between the silence interval and the speech interval includes at least one of the following: (1) If the speech interval overlaps with the silence interval, and the actual start time of the second user corresponding to the speech interval is later than the start time of the silence interval, and the interval between the actual start time of the speech interval and the start time of the silence interval is less than a first threshold, then the actual start time of the speech is taken as the start time of the speech in the speech interval; (2) If the speech interval overlaps with the silence interval, and the second user corresponding to the speech interval... (3) If the actual start time of the user's speech is earlier than the start time of the silence interval, then the actual start time of the speech is taken as the start time of the speech in the speech interval; (4) If the speech interval overlaps with the silence interval, and the actual start time of the second user corresponding to the speech interval is later than the start time of the silence interval, and the interval between the actual start time of the speech and the start time of the silence interval is greater than or equal to the second threshold, then the start time of the silence interval is taken as the start time of the speech in the speech interval; (5) If the speech interval does not overlap with the silence interval, then the start time of the silence interval is taken as the start time of the speech in the speech interval.

[0142] The implementation embodiments of the above-described audio processing method are all applicable to the embodiments of the audio processing device and can achieve the same technical effect.

[0143] This application also provides an audio processing device, including a memory, a processor, and a program stored in the memory and executable on the processor; when the processor executes the program, it implements the above-described audio processing method.

[0144] The implementation embodiments of the above-described audio processing method are all applicable to the embodiments of the audio processing device and can achieve the same technical effect.

[0145] This application also provides a readable storage medium storing a program that, when executed by a processor, implements the steps in the audio processing method described above.

[0146] The implementation embodiments of the above audio processing method are all applicable to the embodiments of the readable storage medium and can achieve the same technical effect.

[0147] This application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the various processes of the above-described audio processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0148] It should be noted that many of the functional components described in this specification are referred to as modules in order to more specifically emphasize the independence of their implementation.

[0149] In this embodiment, the module can be implemented in software so that it can be executed by various types of processors. For example, an identified executable code module may include one or more physical or logical blocks of computer instructions, which may be constructed as objects, procedures, or functions. Nevertheless, the executable code of the identified module does not need to be physically located together, but may include different instructions stored in different bits, which, when logically combined, constitute the module and achieve the module's intended purpose.

[0150] In practice, an executable code module can be a single instruction or many instructions, and can even be distributed across multiple different code segments, different programs, and across multiple memory devices. Similarly, operational data can be identified within the module and can be implemented in any suitable form and organized within any suitable type of data structure. This operational data can be collected as a single dataset or distributed across different locations (including different storage devices), and can exist, at least in part, solely as electronic signals within the system or network.

[0151] When a module can be implemented using software, considering the current level of hardware technology, modules that can be implemented in software can be implemented using hardware circuits by those skilled in the art to achieve the corresponding functions, without considering cost. These hardware circuits include conventional very-large-scale integrated circuits (VLSI) or gate arrays, as well as existing semiconductors such as logic chips and transistors, or other discrete components. Modules can also be implemented using programmable hardware devices, such as field-programmable gate arrays, programmable array logic, and programmable logic devices.

[0152] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0153] The above describes the preferred embodiments of this application. It should be noted that those skilled in the art can make several improvements and modifications without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. An audio processing method, characterized in that, include: Obtain audio content correlation information and audio temporal correlation information between the first user and the second user in an online meeting; Based on the audio content correlation information and audio time sequence correlation information, the first user and the second user are divided into virtual groups; Play the audio of each member in the virtual group to which the first user is located. The first user is the user of this terminal, and the second user is the other terminal users in the online meeting besides the first user.

2. The audio processing method according to claim 1, characterized in that, Obtain audio content correlation information between the first user and the second user in an online meeting, including: Obtain a first set of words from a first text, wherein the first text is the text of the first audio data of the first user in the online meeting, and the first set of words includes: the words in the first text whose frequency of occurrence is ranked in the top Z positions, where Z is an integer greater than or equal to 1; Obtain a second vocabulary set from the second text, wherein the second text is the text of the second audio data of the second user in the online meeting, and the second vocabulary set includes: the words whose frequency of occurrence is ranked in the top Z positions in the second text; the second vocabulary set has the same total number of elements as the first vocabulary set; Obtain the number of words that match between the first vocabulary set and the second vocabulary set; Based on the number of words and the total number of elements, the similarity between the first word set and the second word set is determined, and the similarity is used as the audio content relevance information between the first user and the second user.

3. The audio processing method according to claim 1 or 2, characterized in that, Obtain audio temporal correlation information between the first and second users in an online meeting, including: Calculate the audio continuity between the first user and the second user; Based on the audio continuity, audio temporal correlation information between the first user and the second user is obtained.

4. The audio processing method according to claim 3, characterized in that, The calculation of the audio continuity between the first user and the second user includes: Obtain a first time sequence consisting of the end times of each speech segment of the first user in the first audio data; the first audio data is the audio data of the first user in the online meeting; Obtain a second time series consisting of the start times of each speech segment of the second user in the second audio data; the second audio data is the audio data of the second user in the online meeting; Based on the first time series and the second time series, obtain the audio interval value between the first user and the second user; The audio continuity between the first user and the second user is calculated based on the audio interval value.

5. The audio processing method according to claim 4, characterized in that, The step of obtaining the second time series composed of the start times of each speech segment of the second user in the second audio data includes: Determine the silence interval of the first user corresponding to the speaking voice interval of the second user; wherein, the speaking voice interval of the second user refers to the speaking voice interval of the second user in the second audio data of the online meeting, and the silence interval of the first user refers to the silence interval of the first user in the first audio data of the online meeting. Based on the time relationship between the silence interval and the speech interval, determine the start time of the second user's speech in the second audio data. Based on the determined start time of the speech, the second time sequence is obtained.

6. The audio processing method according to claim 5, characterized in that, Determining the start time of the second user's speech in the second audio data based on the time relationship between the silence interval and the speech interval includes at least one of the following: If the speaking voice interval overlaps with the silent interval, and the actual speaking start time of the second user corresponding to the speaking voice interval is later than the start time of the silent interval, and the interval between the actual speaking start time and the start time of the silent interval is less than a first threshold, then the actual speaking start time is taken as the speaking start time of the speaking voice interval. If the speaking voice interval overlaps with the silent interval, and the actual start time of the second user corresponding to the speaking voice interval is earlier than the start time of the silent interval, then the actual start time of the speaking voice interval shall be taken as the start time of the speaking voice interval. If the speaking voice interval overlaps with the silent interval, and the actual speaking start time of the second user corresponding to the speaking voice interval is later than the start time of the silent interval, and the interval between the actual speaking start time and the start time of the silent interval is greater than or equal to a second threshold, then the start time of the silent interval is taken as the speaking start time of the speaking voice interval. If the speaking voice interval does not overlap with the silence interval, then the start time of the silence interval is taken as the start time of the speaking voice interval.

7. An audio processing device, characterized in that, include: The first acquisition module is used to acquire audio content correlation information and audio timing correlation information between the first user and the second user in the online meeting. The first processing module is used to divide the first user and the second user into virtual groups based on the audio content correlation information and the audio time sequence correlation information. The first playback module is used to play the audio of each member in the virtual group to which the first user is located. The first user is the user of this terminal, and the second user is the other terminal users in the online meeting besides the first user.

8. An audio processing device, comprising a memory, a processor, and a program stored in the memory and executable on the processor; characterized in that, When the processor executes the program, it implements the audio processing method as described in any one of claims 1 to 6.

9. A readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the audio processing method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the audio processing method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Smart classroom audio control system and method and storage medium

    CN116866783A

  • Conference page display method and device, electronic equipment and storage medium

    CN118264842A