Audio group intelligent scheduling system and method based on artificial intelligence
By adopting an intelligent audio group scheduling system based on artificial intelligence in online meetings, the problem of mutual interference between multiple audio streams is solved, and the quality and efficiency of meetings are improved.
Patent Information
- Application Number
- CN202411994271.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-09
AI Technical Summary
In an online meeting, multiple audio streams interfere with each other, affecting the quality and efficiency of the meeting.
The intelligent scheduling system of audio groups based on artificial intelligence is adopted, and the conference attributes and audio priorities are determined through the intelligent analysis module, and the audio stream is reduced and abnormal characteristics are processed using the artificial intelligence model, scheduling strategies are generated, and the intelligent scheduling of the audio stream is realized through the scheduling management module.
Improves the fluency and speech quality of the audio stream, reduces the impact of invalid audio on the meeting, and improves the efficiency of the meeting.
Smart Images

Figure CN119964591A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of audio processing and relates to intelligent scheduling technology of audio groups, specifically an audio group intelligent scheduling system and method based on artificial intelligence. Background Art
[0002] An audio group is an audio mixer that can combine multiple audio objects together. It can control the properties of multiple audio signals at the same time and allow users to centrally manage and process multiple audio signals, such as adjusting the volume, applying sound effects, or splitting the audio data stream.
[0003] With the development of computer technology and the enhancement of people's sense of time, online meetings are gradually increasing. While online meetings provide convenience, they are also affected by various factors, such as network quality, equipment quality, etc. In addition, when there are many people involved in online meetings, multiple audio streams may be generated at the same time. If multiple audio streams are transmitted and played in sequence according to the receiving time, it is easy to cause poor conference effect quality due to the mutual influence between audio streams. How to intelligently schedule audio groups is a technical problem that needs to be solved urgently.
[0004] The present application provides an audio group intelligent scheduling system and method based on artificial intelligence to solve the above technical problems. Summary of the invention
[0005] The present application aims to solve at least one of the technical problems existing in the prior art; to this end, the present application proposes an audio group intelligent scheduling system and method based on artificial intelligence, which is used to solve the technical problem that multiple audio streams generated simultaneously in existing online meetings interfere with each other, affecting the quality and efficiency of the meeting.
[0006] To achieve the above-mentioned purpose, the first aspect of the present application provides an audio group intelligent scheduling system based on artificial intelligence, including an intelligent analysis module and a scheduling management module connected thereto;
[0007] Intelligent analysis module: used to determine the meeting attributes of the online meeting and set audio priorities for participants based on the meeting attributes; the meeting attributes include the meeting type and meeting process; and,
[0008] Receive an audio group, determine the playback priority of each audio stream in the audio group according to the conference attributes and the audio priority, and generate a scheduling strategy after preprocessing each audio stream through an artificial intelligence model; wherein the preprocessing includes noise reduction and synthesis;
[0009] Scheduling management module: used to determine the audio scheduling method according to network resources, and implement intelligent scheduling of each audio stream in the audio group according to the audio scheduling method and the scheduling strategy; wherein the audio scheduling method includes dynamic scheduling or scheduling based on load status.
[0010] Preferably, determining the conference attributes of the online conference includes:
[0011] Determine the attributes of the meeting by identifying the meeting materials prepared by the online meeting initiator; wherein the meeting materials include text materials, audio materials or video materials; or,
[0012] The opening remarks of the online meeting are analyzed by natural language processing technology to determine the attributes of the meeting; wherein the opening remarks include text materials or audio materials.
[0013] Preferably, the preprocessing of each audio stream by using an artificial intelligence model includes:
[0014] Performing noise reduction processing on each of the audio streams in the audio group using a noise reduction algorithm to obtain a noise-reduced audio stream; wherein the noise reduction algorithm includes WebRtc Noise Suppression, wavelet transform, adaptive filter or spectrum gating;
[0015] An artificial intelligence model is used to identify abnormal features of each audio stream after noise reduction, and a new audio stream is generated after optimizing the abnormal features; wherein the abnormal features refer to audio segments that affect the accuracy of audio content or audio rhythm, and the artificial intelligence model includes a voice activity detection algorithm or a keyword detection algorithm.
[0016] Preferably, the identifying abnormal features of each audio stream after noise reduction by an artificial intelligence model includes:
[0017] The audio signal of the audio stream to be processed is segmented, and each segment of the audio signal is converted into an audio vector, and the frequency spectrum feature of the audio vector is extracted; wherein the frequency spectrum feature includes a Mel frequency cepstrum coefficient;
[0018] The spectral features of each of the audio vectors are analyzed by an artificial intelligence model to determine whether there are abnormal features; if yes, the location of the abnormal features is identified; wherein the artificial intelligence model is obtained through training with a data training set.
[0019] Preferably, generating a new audio stream after optimizing the abnormal features includes:
[0020] Extracting the position of each abnormal feature, qualitatively analyzing the abnormal feature in each audio signal segment, and optimizing the audio signal according to the qualitative result; wherein the optimization processing includes abnormal feature removal, abnormal feature adjustment or audio signal continuity optimization;
[0021] The optimized audio signals are synthesized by a preset synthesis method to generate a new audio stream; wherein the preset synthesis method includes an overall synthesis method or a segmented synthesis method.
[0022] Preferably, the overall synthesis method refers to synthesizing all audio signals into a new audio stream, and the segmented synthesis method refers to integrating all audio signals into multiple audio signals and merging the multiple audio signals into a new audio stream.
[0023] Preferably, before generating the scheduling strategy, noise reduction processing is performed on each audio stream in the audio group according to the conference process, including:
[0024] When the conference process allows multiple audio streams in the audio group to coexist and play, the audio streams are firstly subjected to individual noise reduction, and then the multiple audio streams are subjected to joint noise reduction in combination with the play priority; otherwise,
[0025] The audio streams in the audio group are individually denoised according to the playback priority.
[0026] Preferably, the joint noise reduction is performed on a plurality of audio streams in combination with the playback priority, including:
[0027] Determine a playback order of each audio stream according to the playback priority, and select at least one audio stream as a reference audio according to the playback order;
[0028] Detect the voice activity area of each audio stream, and set the time difference between adjacent audio streams according to the voice activity area of each audio stream; and fine-tune each audio stream in the audio group according to the time difference based on the reference video.
[0029] Preferably, after fine-tuning each audio stream in the audio group, the volume of each audio stream is adjusted differently.
[0030] The second aspect of the present application provides an audio group intelligent scheduling method based on artificial intelligence, comprising:
[0031] Determine the meeting attributes of the online meeting and set audio priorities for participants based on the meeting attributes;
[0032] receiving an audio group, and determining a playback priority of each audio stream in the audio group according to a conference attribute and the audio priority;
[0033] After preprocessing each of the audio streams through an artificial intelligence model, a scheduling strategy is generated, and each of the audio streams in the audio group is intelligently scheduled according to the audio scheduling method and the scheduling strategy.
[0034] Compared with the prior art, the beneficial effects of this application are:
[0035] 1. In this application, the intelligent analysis module first determines the conference attributes based on the conference materials, and determines the audio priority of the participants based on the conference attributes; after receiving the audio group, based on the conference attributes and the audio priority of the participants in each audio stream, the artificial intelligence model is used to reduce noise and remove abnormal features for each audio stream in the audio group, so that the processed audio stream is natural and smooth, and finally a scheduling strategy is generated for execution by the scheduling management module; this application sets the playback priority to adjust the speaking order of the participants, and only plays audio streams with high playback priority to avoid interference from others and improve meeting efficiency; and by eliminating abnormal features in the audio stream through the artificial intelligence model, the fluency of the audio stream can be improved, thereby improving the speech quality of the participants and reducing the impact of invalid audio on the entire meeting.
[0036] 2. Before generating the scheduling strategy, the present application performs noise reduction processing on each audio stream in the audio group according to the conference process. When the conference process allows several audio streams in the audio group to be played coexistently, the audio stream is first subjected to individual noise reduction, and then the several audio streams are jointly subjected to noise reduction in combination with the playback priority; otherwise, the audio streams in the audio group are subjected to individual noise reduction according to the playback priority. The present application takes into account the influence between the audio streams in the audio group. Therefore, when the duration of the mixed audio is sufficient, the mutual influence between the audio streams is avoided by staggering the voice activity areas of each audio stream. When the duration is insufficient, the audio streams with overlapping voice activity areas are distinguished by volume or timbre, so as to ensure that the participants can hear their opinions clearly at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0038] Figure 1 This is a schematic diagram of the principle of the audio group intelligent scheduling system in Example 1 of the present application;
[0039] Figure 2 Schematic diagram of the flow of the audio group intelligent scheduling method in the embodiment of the present application;
[0040] Figure 3 This is a schematic diagram of a process for identifying abnormal features in an audio signal in Embodiment 1 of the present application;
[0041] Figure 4This is a flow chart of optimizing the abnormal features in the second embodiment of the present application. DETAILED DESCRIPTION
[0042] The technical solution of the present application will be described clearly and completely in conjunction with the embodiments below. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application.
[0043] With the development of computer technology and the enhancement of people's sense of time, many things that originally required offline communication can be completed through online meetings, such as task arrangement, communication of contract details between supply and demand parties, etc. Online meetings are more convenient than offline communication and reduce communication costs.
[0044] In an online meeting (or other scenarios where multiple audio streams can be generated simultaneously), multiple participants may speak at the same time and generate multiple audio streams. When these audio streams are played in the order in which they are generated, they will definitely interfere with each other, which will not only affect the quality of speech of each participant, but also affect the efficiency of the entire online meeting.
[0045] Embodiment 1: This embodiment intelligently schedules the audio streams corresponding to the participants through conference attributes to avoid the speeches of the participants from affecting each other and improve the efficiency of the online conference.
[0046] See also Figure 1-Figure 2 , the first aspect of the present application provides an audio group intelligent scheduling system based on artificial intelligence, including an intelligent analysis module and a scheduling management module connected thereto;
[0047] Intelligent analysis module: used to determine the meeting attributes of the online meeting and set audio priorities for participants based on the meeting attributes; the meeting attributes include the meeting type and meeting process; and,
[0048] Receive an audio group, determine the playback priority of each audio stream in the audio group according to the conference attributes and the audio priority, and generate a scheduling strategy after preprocessing each audio stream through an artificial intelligence model; wherein the preprocessing includes noise reduction and synthesis;
[0049] Scheduling management module: used to determine the audio scheduling method according to network resources, and implement intelligent scheduling of each audio stream in the audio group according to the audio scheduling method and the scheduling strategy; wherein the audio scheduling method includes dynamic scheduling or scheduling based on load status.
[0050] The audio group intelligent scheduling system provided in this embodiment can be used as an audio scheduling system alone, or it can be used in conjunction with online conference software. The system mainly includes an intelligent analysis module and a scheduling management module, which are connected by wired or wireless means; the intelligent analysis module is mainly responsible for data processing, and needs to be connected to relevant equipment to collect relevant data for audio scheduling, and schedule the audio streams in the audio group based on the relevant data; the scheduling management module is responsible for specific scheduling work, such as playing the processed audio streams according to the sequence in the scheduling strategy, so that the playback of each audio stream is orderly and the playback quality of each audio stream is guaranteed.
[0051] This embodiment uses online meetings as an application scenario to explain this technical solution in detail. Online meetings generally have a large number of participants, and there are many types of meetings. The speaking requirements and speaking order of different meetings are different. For example, for meetings to arrange urgent tasks, the meeting generally only requires the task scheduler to distribute the task content to each participant and explain the task content in detail. Other participants do not need to speak too much or even speak at all. Other meetings, such as discussion meetings on a certain issue, will have multiple meeting processes, such as at least including an introduction to the problem. At this time, relevant personnel are required to give a comprehensive introduction to the content of the problem, and other participants do not need to speak; it is also necessary to include a discussion session, then all participants have the possibility of speaking.
[0052] Although the existing online conference software gives the initiative of speaking to participants, and participants speak by turning on the microphone, in reality there are still cases where the microphone is not turned on when it is time to speak, or is turned on when it is not time to speak, which affects the overall quality of the meeting.
[0053] In this embodiment, the meeting attributes of the online meeting need to be determined first. The meeting attributes are mainly the meeting type and the meeting process. The above-mentioned arrangement of urgent tasks or discussion of a certain issue is the meeting type, and the above-mentioned discussion of a certain issue and the problem introduction and problem discussion correspond to two meeting processes. According to the meeting attributes, it can be determined in what time period which participants can speak, and the audio priority of each participant can be determined.
[0054] Next, after receiving the collected audio group, the playback priority of each audio stream in the audio group is set according to the conference attributes and the determined audio priority. After pre-processing each audio stream, a scheduling strategy is generated according to the playback priority, thereby realizing intelligent scheduling of multiple audio streams.
[0055] It should be noted that playback priority refers to adjusting the playback order of each audio stream in the audio group, which can be played one by one, or of course, each audio stream can be adjusted according to the playback priority to ensure that the audio streams do not interfere with each other during playback. The conference attributes and audio priority jointly determine the playback priority of the audio stream, such as the speaking priority of a participant in different conference processes is different. Of course, the audio priority can also be associated with the conference process, that is, different conference processes set different audio priorities, so that the playback priority of each audio stream in the audio group is directly determined according to the audio priority. In some other preferred embodiments, the audio priority and the playback priority can also be equivalent.
[0056] Finally, the audio streams in the audio group are preprocessed through the artificial intelligence model, and the scheduling strategy of each preprocessed audio stream is generated based on the playback priority. The emergency management module realizes intelligent scheduling of audio streams according to the status of network resources and scheduling resources, which can ensure that the speeches of participants in each meeting process can be played in an orderly manner to avoid mutual influence between speeches of multiple people.
[0057] The conference attributes in this embodiment specifically include the conference type and conference process. The conference attributes can be determined in the following two ways:
[0058] 1. Determine the attributes of the meeting by identifying the meeting materials prepared by the online meeting initiator. Meeting materials include text materials, audio materials or video materials, such as the meeting outline, audio materials or video materials prepared in advance by the meeting initiator. Based on OCR recognition technology and text processing technology, these materials are identified to determine the type of meeting and the meeting process. After the meeting attributes are completed, the integrity of the meeting attributes should also be verified, such as whether there are any problems with the content and continuity of the meeting type and meeting process. If there are any problems, the meeting initiator should be reminded to supplement, or the content should be supplemented through method 2.
[0059] 2. Analyze the opening remarks of the online meeting through natural language processing technology to determine the meeting attributes. Compared with method 1, which determines the meeting attributes by identifying information before the meeting starts, method 2 determines the meeting attributes through the opening remarks at the beginning of the meeting. For example, the meeting host will outline the meeting theme and meeting process in the opening remarks. The meeting attributes can be determined by identifying the opening remarks.
[0060] In addition to the above two methods, in some other preferred embodiments, the conference attributes can also be determined by other methods. Of course, the contents of the conference attributes can also be similar, and the determination and contents of the conference attributes will not be repeated here.
[0061] The audio groups received in this embodiment are the audios of the speeches of the participants in the online conference. Some audios meet the requirements of the online conference, and some audios will affect other audios. The audio groups can be divided according to time periods, such as determining a time period, and the audios generated in each time period are regarded as an audio group; of course, other methods can also be used to generate audio groups.
[0062] See also Figure 3 After receiving the audio group, the intelligent analysis module needs to pre-process each audio stream in the audio group. The specific pre-processing includes:
[0063] The audio streams in the audio group are subjected to noise reduction processing by using a noise reduction algorithm to obtain a noise-reduced audio stream. The noise reduction algorithms used here include WebRtc Noise Suppression, wavelet transform, adaptive filter or spectrum gating, which are mainly used to remove noise carried in the audio stream, such as the ambient noise of the participants and the noise generated by the equipment used.
[0064] The artificial intelligence model identifies abnormal features of each audio stream after noise reduction, optimizes the abnormal features and generates a new audio stream. The abnormal features here refer to audio segments that affect the accuracy of the audio content or the audio rhythm. The artificial intelligence model includes a voice activity detection algorithm or a keyword detection algorithm.
[0065] When participants speak through microphones, they generate an audio stream that carries various types of noise. The audio stream is first denoised using a noise reduction algorithm to remove the noise. The artificial intelligence model is then used to detect anomalies in the denoised audio stream and identify abnormal features. The presence of abnormal features will affect the subscriber's understanding of the audio stream content, and participants may use many invalid words such as "um" and "ah" due to nervousness, habit, etc., which will affect the coherence and accuracy of the content expressed in the audio stream. Even when the participants' sentences are incoherent, it will affect the fluency of the entire meeting.
[0066] The method provided in this embodiment for identifying abnormal features of each audio stream after noise reduction by using an artificial intelligence model includes the following steps:
[0067] 1) Split the audio file into smaller segments;
[0068] 2) Convert each segment into an audio vector;
[0069] 3) Extracting Mel frequency cepstrum (MFCC) coefficients from these vectors as spectral features;
[0070] 4) Use an artificial intelligence model trained with the data training set to analyze these spectral features to identify whether there are anomalies;
[0071] 5) If anomalies are detected, determine the specific locations of these abnormal features in the audio.
[0072] Online meetings should be transmitted in real time, but this is difficult to achieve in practice. Therefore, a delay is preset to make the meeting smoother overall. For details, please refer to the following solutions:
[0073] After the online meeting starts, the speeches of all participants are collected to form an audio stream. The audio streams within a certain period of time are regarded as an audio group. Some time is reserved for the processing of each audio group. If the computing power is sufficient, this part of time will not be very long. This period of time can be used as a delay, and the delay is reminded to the participants.
[0074] After identifying the abnormal features in each audio stream, these abnormal features are optimized to obtain a new audio stream. Optimization refers to eliminating abnormal features. For example, if the abnormal features are invalid words, the invalid words are eliminated; if the abnormal features are semantically incoherent, the semantics are appropriately adjusted.
[0075] In the process of identifying abnormal features of the audio stream, various types of AI big models can be used to determine whether there are abnormal features. Of course, the audio stream can also be optimized through the AI big model, such as removing invalid words, reorganizing the audio segments, etc. It should be noted that the optimized audio stream should conform to the speaking characteristics of the participants, so the audio stream should sound smooth. This verification can be determined by analyzing the audio features of the audio stream, and of course it can also be achieved through the AI big model.
[0076] This embodiment first determines the conference attributes based on the conference materials, and determines the audio priority of the participants based on the conference attributes; after receiving the audio group, based on the conference attributes and the audio priority of the participants in each audio stream, each audio stream in the audio group is subjected to noise reduction and abnormal feature removal through an artificial intelligence model, so that the processed audio stream is natural and smooth. This embodiment sets the playback priority to adjust the speaking order of the participants, and only plays the audio stream with a high playback priority to avoid interference from others and improve the efficiency of the meeting; and by removing abnormal features in the audio stream through the artificial intelligence model, the fluency of the audio stream can be improved, thereby improving the speech quality of the participants and reducing the impact of invalid audio on the entire meeting.
[0077] See also Figure 4 , Embodiment 2: Based on Embodiment 1, this embodiment provides a method for optimizing the above abnormal features, which can ensure that the audio stream is smooth and natural.
[0078] This embodiment generates a new audio stream after optimizing the abnormal features, including:
[0079] The position of each abnormal feature is extracted, the abnormal feature is qualitatively characterized in each audio signal segment, and the audio signal is optimized according to the qualitative result; each segment of the optimized audio signal is synthesized by a preset synthesis method to generate a new audio stream.
[0080] After determining the location of the abnormal features in the audio stream through Example 1 or other methods, the abnormal features are first characterized, and the qualitative results determine the subsequent processing methods. If the abnormal features have no substantial impact on the content of the audio stream, such as repeated modal particles, etc., optimization is performed by removing abnormal features. If the abnormal features have an impact on the content of the audio stream, such as the inversion of characters in a professional vocabulary, optimization is performed through adjustment. Whether it is abnormal feature removal or abnormal feature adjustment, it will affect the fluency of the audio stream, so it is also necessary to perform signal continuity processing on the audio stream, mainly to adjust the coordination of the audio segment of the audio stream with the previous and next audio segments, such as adjusting the audio, which can also be achieved through the AI large model.
[0081] The identification and optimization of the above-mentioned abnormal features are for each segment of audio signals. Each segment of audio signals is segmented from an audio stream, so it is also necessary to integrate each audio signal into an audio stream. This embodiment provides two integration methods, namely, an overall synthesis method and a segmented synthesis method. The overall synthesis method refers to synthesizing all audio signals into a new audio stream; the segmented synthesis method refers to integrating all audio signals into multiple audio signals, and merging multiple audio signals into a new audio stream. The above two synthesis methods are selected according to the specific situation. If the number of audio signals is small, the audio stream can be obtained by the overall synthesis method; if the audio signals are large, they can be synthesized in segments to finally form an audio stream.
[0082] Embodiment 3: Based on Embodiment 1 or Embodiment 2, this embodiment provides a processing method after optimizing the audio streams in each pair of audio groups. This processing method is to eliminate the mutual influence between the audio streams in the audio group and is suitable for scenarios where the audio streams of multiple participants are played simultaneously.
[0083] Before generating the scheduling strategy, each audio stream in the audio group is subjected to noise reduction processing according to the conference process, including: when the conference process allows several audio streams in the audio group to be played coexistently, the audio stream is first subjected to noise reduction individually, and then the several audio streams are subjected to joint noise reduction in combination with the playback priority; otherwise, the audio streams in the audio group are subjected to noise reduction individually according to the playback priority.
[0084] The meeting process may be a single speech session or a group discussion session. No matter which session it is, it is possible to receive audio streams from multiple participants in a period of time. In the single speech session, the audio streams of other participants should be blocked, while in the group discussion session, the audio streams of each participant need to be played, and the mutual influence between the audio streams should be minimized.
[0085] If the conference process allows multiple audio streams in the audio group to be played at the same time, such as during a conference discussion, each audio stream will be noise-reduced separately first, and then the mutual influence between the audio streams needs to be considered. If there is any mutual influence, the influence should be eliminated or reduced as much as possible. If the conference process does not allow audio streams to be played at the same time, the audio streams to be played will be selected according to the playback priority and noise will be reduced separately.
[0086] For joint noise reduction, joint noise reduction is performed on several audio streams in combination with the playback priority, including: determining the playback order of each audio stream according to the playback priority, and selecting at least one audio stream as a reference audio according to the playback order; detecting the voice activity area of each audio stream, and setting the time difference between adjacent audio streams according to the voice activity area of each audio stream; and fine-tuning each audio stream in the audio group according to the time difference based on the reference video.
[0087] The calculation of the time difference can refer to the following steps:
[0088] 1. Audio feature extraction: Extract the time series features of the reference audio and its preceding and following audio streams (determined by playback priority), such as MFCC. These features can represent the key information of the audio signal;
[0089] 2. Detect voice activity: Use the voice activity detection (AVD) algorithm to detect the voice activity area in each audio and determine the start and end time of the voice;
[0090] 3. Compare the speech activity areas of each audio stream, calculate the time difference with the least mutual influence, and fine-tune the playback time difference of the reference audio and its previous and subsequent audios based on the time difference.
[0091] 4. Audio synthesis and volume adjustment: synthesize the fine-tuned audio streams to generate mixed audio; at the same time, adjust the sound of each audio stream in the mixed audio to ensure that the audio stream corresponding to each participant can be heard clearly without affecting other audio streams. In other preferred embodiments, the audio streams may not be superimposed and synthesized, and only the playback time of each audio stream needs to be marked.
[0092] It should be noted that the duration of the generated mixed audio cannot exceed the duration of the audio group, so as not to squeeze the processing time of the subsequent audio group. Therefore, if the duration is sufficient, the voice activity areas of each audio stream will be staggered. If the duration is insufficient, the voice activity areas can overlap, but it should be ensured that the overlapping parts can be heard clearly, which can be achieved by distinguishing the volume or timbre, etc.
[0093] After processing each audio stream in the audio group separately, this embodiment takes into account the influence between the audio streams. Therefore, when the duration of the mixed audio is sufficient, the mutual influence between the audio streams is avoided by staggering the voice activity areas of each audio stream. When the duration is insufficient, the audio streams with overlapping voice activity areas are distinguished by volume or timbre, so as to ensure that the participants can hear their opinions at the same time.
[0094] See also Figure 2 The second aspect of the present application provides an audio group intelligent scheduling method based on artificial intelligence, including:
[0095] Determine the meeting attributes of the online meeting and set audio priorities for participants based on the meeting attributes;
[0096] receiving an audio group, and determining a playback priority of each audio stream in the audio group according to a conference attribute and the audio priority;
[0097] After preprocessing each of the audio streams through an artificial intelligence model, a scheduling strategy is generated, and each of the audio streams in the audio group is intelligently scheduled according to the audio scheduling method and the scheduling strategy.
[0098] The scheduling strategy in this application can be understood as the order in which the audio streams in the audio group are played after processing. If they are combined into a mixed audio, the mixed audio can be played directly. If there are several separate audio streams, the fine-tuned audio streams are played according to the playback priority.
[0099] The above embodiments are only used to illustrate the technical method of the present application and are not intended to limit it. Although the present application has been described in detail with reference to the preferred embodiments, a person of ordinary skill in the art should understand that the technical method of the present application may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present application.
Claims
1. An audio group intelligent scheduling system based on artificial intelligence, comprising an intelligent analysis module and a scheduling management module connected thereto; characterized in that: Intelligent analysis module: used to determine the meeting attributes of the online meeting and set audio priorities for participants based on the meeting attributes; the meeting attributes include the meeting type and meeting process; and, Receive an audio group, determine the playback priority of each audio stream in the audio group according to the conference attributes and the audio priority, and generate a scheduling strategy after preprocessing each audio stream through an artificial intelligence model; wherein the preprocessing includes noise reduction and synthesis; Scheduling management module: used to determine the audio scheduling method according to network resources, and implement intelligent scheduling of each audio stream in the audio group according to the audio scheduling method and the scheduling strategy; wherein the audio scheduling method includes dynamic scheduling or scheduling based on load status.
2. The audio group intelligent scheduling system based on artificial intelligence according to claim 1 is characterized in that: The determining of the conference attributes of the online conference includes: Determine the attributes of the meeting by identifying the meeting materials prepared by the online meeting initiator; wherein the meeting materials include text materials, audio materials or video materials; or, The opening remarks of the online meeting are analyzed by natural language processing technology to determine the attributes of the meeting; wherein the opening remarks include text materials or audio materials.
3. The audio group intelligent scheduling system based on artificial intelligence according to claim 1 is characterized in that: The preprocessing of each audio stream by the artificial intelligence model includes: Performing noise reduction processing on each of the audio streams in the audio group using a noise reduction algorithm to obtain a noise-reduced audio stream; wherein the noise reduction algorithm includes WebRtc Noise Suppression, wavelet transform, adaptive filter or spectrum gating; An artificial intelligence model is used to identify abnormal features of each audio stream after noise reduction, and a new audio stream is generated after optimizing the abnormal features; wherein the abnormal features refer to audio segments that affect the accuracy of audio content or audio rhythm, and the artificial intelligence model includes a voice activity detection algorithm or a keyword detection algorithm.
4. The audio group intelligent scheduling system based on artificial intelligence according to claim 3 is characterized in that: The method of identifying abnormal features of each audio stream after noise reduction by using an artificial intelligence model includes: The audio signal of the audio stream to be processed is segmented, and each segment of the audio signal is converted into an audio vector, and the frequency spectrum feature of the audio vector is extracted; wherein the frequency spectrum feature includes a Mel frequency cepstrum coefficient; The spectral features of each of the audio vectors are analyzed by an artificial intelligence model to determine whether there are abnormal features; if yes, the location of the abnormal features is identified; wherein the artificial intelligence model is obtained through training with a data training set.
5. The audio group intelligent scheduling system based on artificial intelligence according to claim 3 is characterized in that: The step of generating a new audio stream after optimizing the abnormal features includes: Extracting the position of each abnormal feature, qualitatively analyzing the abnormal feature in each audio signal segment, and optimizing the audio signal according to the qualitative result; wherein the optimization processing includes abnormal feature removal, abnormal feature adjustment or audio signal continuity optimization; The optimized audio signals are synthesized by a preset synthesis method to generate a new audio stream; wherein the preset synthesis method includes an overall synthesis method or a segmented synthesis method.
6. The audio group intelligent scheduling system based on artificial intelligence according to claim 5 is characterized in that: The overall synthesis method refers to synthesizing all audio signals into a new audio stream, and the segmented synthesis method refers to integrating all audio signals into multiple audio signals and merging the multiple audio signals into a new audio stream.
7. The audio group intelligent scheduling system based on artificial intelligence according to claim 1 is characterized in that: Before generating the scheduling strategy, performing noise reduction processing on each audio stream in the audio group according to the conference process includes: When the conference process allows multiple audio streams in the audio group to coexist and play, the audio streams are firstly subjected to individual noise reduction, and then the multiple audio streams are subjected to joint noise reduction in combination with the play priority; otherwise, The audio streams in the audio group are individually denoised according to the playback priority.
8. The audio group intelligent scheduling system based on artificial intelligence according to claim 7 is characterized in that: Combined noise reduction is performed on multiple audio streams in combination with the playback priority, including: Determine a playback order of each audio stream according to the playback priority, and select at least one audio stream as a reference audio according to the playback order; Detect the voice activity area of each audio stream, and set the time difference between adjacent audio streams according to the voice activity area of each audio stream; and fine-tune each audio stream in the audio group according to the time difference based on the reference video.
9. The audio group intelligent scheduling system based on artificial intelligence according to claim 8, characterized in that: After fine-tuning each audio stream in the audio group, the volume of each audio stream is adjusted differently.
10. An audio group intelligent scheduling method based on artificial intelligence, based on the operation of an audio group intelligent scheduling system based on artificial intelligence according to any one of claims 1 to 9, characterized in that: include: Determine the meeting attributes of the online meeting and set audio priorities for participants based on the meeting attributes; receiving an audio group, and determining a playback priority of each audio stream in the audio group according to a conference attribute and the audio priority; After preprocessing each of the audio streams through an artificial intelligence model, a scheduling strategy is generated, and each of the audio streams in the audio group is intelligently scheduled according to the audio scheduling method and the scheduling strategy.
Citation Information
Cited By
Display screen integrated structure and microphone control system for paperless conference
CN121218065A