Online conference voice real-time enhancement method and device based on generative artificial intelligence, and electronic equipment

Through the real-time enhancement method of online conference voice based on generative artificial intelligence, the voice model is trained and intelligently adjusted, and the problem of difficulty in identifying and optimizing online conference voice quality in the existing technology is solved, and high-quality voice communication and personalized optimization effects are achieved.

CN120048279APending Publication Date: 2025-05-27TIANFU JIANGXI LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510200097.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing technology is difficult to identify the specific reasons for the decline in voice quality in online conferences in real time and accurately. It cannot be targetedly optimized based on user personalized voice characteristics, and there is a lack of solutions for in-depth analysis and intelligent optimization using artificial intelligence technology.

Method used

Through the online conference voice real-time enhancement method based on generative artificial intelligence, the speech model is trained to adapt to different voice scenarios of users, spectrum analysis is performed to obtain baseline voice waveforms, evaluate network and microphone performance before the meeting, obtain voice data and network indicators in real time during the meeting, judge whether voice enhancement is needed, and use the trained speech model to intelligently adjust or replace.

Benefits of technology

It realizes accurate identification and resolution of voice quality problems, provides personalized voice optimization, improves the voice communication quality of online meetings, and ensures high-quality voice communication under different network conditions and user status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048279A_ABST
    Figure CN120048279A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an online conference voice real-time enhancement method and device based on generative artificial intelligence and electronic equipment, and relates to the technical field of voice enhancement, and the method comprises the steps: carrying out the training of a generative artificial intelligence voice model based on the audio data of a user in different voice scenes; performing spectral analysis on the audio data of the user in different voice scenes to obtain a baseline voice waveform; before a conference, network performance and microphone performance are evaluated, and an initial voice quality evaluation score is obtained according to an evaluation result; in the conference process, real-time voice data, a real-time voice quality evaluation score and a real-time network index are obtained, and whether voice enhancement needs to be carried out or not is judged; and if yes, performing voice enhancement by using the trained artificial intelligence voice model. According to the technical scheme of the invention, real-time monitoring, intelligent analysis and accurate optimization of voice are realized, and clear and natural voice communication experience is provided for an online conference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice enhancement technology. Specifically, it relates to an online meeting voice real-time enhancement method, device, and electronic device based on generative artificial intelligence. Background Art

[0002] In the wave of today's digital communication, online meetings have become an important way for people to collaborate and communicate remotely. When users participate in online meetings, the quality of their voices is often interfered with by various factors. The user's own health condition may affect voice output. For example, being sick can cause changes in voice. Network environment factors have a particularly significant impact on voice quality. When participating in a meeting in an area with limited network bandwidth, or when there are high packet loss rates, high latencies, etc. in the network connection, problems such as data loss, interruption, or distortion are likely to occur during the transmission of voice signals. These factors combined make the voice quality received by other meeting participants greatly reduced, and phenomena such as unclear voices, unstable volumes, noise, or stuttering may occur, seriously affecting the communication effect of online meetings.

[0003] Currently, there is a lack of a comprehensive and intelligent solution for the problem of online meeting voice quality. Existing technologies often cannot accurately identify the specific reasons for the decline in voice quality in real time, nor can they perform targeted optimization according to the user's personalized voice characteristics. When the network condition is poor, generally, one can only passively accept the decline in voice quality and cannot actively take measures to enhance the voice signal. For the problem of voice quality caused by changes in the user's own voice, there are also no effective means to improve it in real time, and it is impossible to make the voice output reach the best state without changing the user's natural speaking habits. Most existing solutions do not fully utilize the advantages of artificial intelligence technology, cannot perform in-depth analysis and intelligent optimization of voices, and cannot meet the needs of users for high-quality voice communication under different network conditions and their own states. Summary of the Invention

[0004] Embodiments of this application provide an online meeting voice real-time enhancement method, device, and electronic device based on generative artificial intelligence to solve the technical problems existing in the prior art.

[0005] Other features and advantages of this application will become apparent through the following detailed description, or will be partially learned through the practice of this application.

[0006] According to the first aspect of the embodiments of this application, an online meeting voice real-time enhancement method based on generative artificial intelligence is provided, including:

[0007] Training a generative artificial intelligence voice model based on audio data of the user in different voice scenarios;

[0008] Perform spectral analysis on the audio data of the user in different voice scenarios to obtain the baseline speech waveform;

[0009] Before the meeting, evaluate the network performance and microphone performance, and obtain the initial speech quality evaluation score according to the evaluation results;

[0010] During the meeting, obtain the real-time speech data, real-time speech quality evaluation score, and real-time network metrics, and determine whether speech enhancement is required;

[0011] If so, use the trained artificial intelligence speech model for speech enhancement.

[0012] In some embodiments of the present application, based on the foregoing solution, training the generative artificial intelligence speech model based on the audio data of the user in different voice scenarios includes:

[0013] Collect the audio data of the user in different voice scenarios and the text corresponding to the audio data;

[0014] Perform data preprocessing on the collected audio data, and convert the audio data into training segments suitable for model learning through noise removal technology, audio level normalization operation, and audio data segmentation;

[0015] Train the generative artificial intelligence speech model based on the training segments and the text corresponding to the audio data.

[0016] In some embodiments of the present application, based on the foregoing solution, performing spectral analysis on the audio data of the user in different voice scenarios to obtain the baseline speech waveform includes:

[0017] Decompose the audio data into numerous frequency components in detail, and calculate the statistical indicators of each frequency component;

[0018] Determine the baseline speech waveform according to the statistical indicators of each frequency component.

[0019] In some embodiments of the present application, based on the foregoing solution, obtaining the real-time speech data, real-time speech quality evaluation score, and real-time network metrics, and determining whether speech enhancement is required includes:

[0020] Compare the real-time speech data with the baseline speech waveform to determine whether speech enhancement is required;

[0021] Compare the real-time speech quality evaluation score with the initial speech quality evaluation score to determine whether speech enhancement is required;

[0022] Compare the real-time network metrics with the network metric requirements to determine whether speech enhancement is required.

[0023] In some embodiments of the present application, based on the foregoing solution, the comparison of the real-time voice data with the baseline voice waveform to determine whether voice enhancement is required includes:

[0024] Obtain real-time voice data, perform spectral analysis on the real-time voice data, extract the frequency components of the current voice waveform, and compare the frequency components of the current voice waveform with the baseline voice waveform. If the difference between the two exceeds a preset threshold, it is determined that voice enhancement is required.

[0025] In some embodiments of the present application, based on the foregoing solution, the comparison of the real-time voice quality assessment score with the initial voice quality assessment score to determine whether voice enhancement is required includes:

[0026] Calculate the real-time voice quality assessment score, and compare the real-time voice quality assessment score with the initial voice quality assessment score. If the real-time voice quality assessment score is less than the initial voice quality assessment score, it is determined that voice enhancement is required.

[0027] In some embodiments of the present application, based on the foregoing solution, the comparison of the real-time network metrics with the network metric requirements to determine whether voice enhancement is required includes:

[0028] Comprehensively evaluate the current network connection status to obtain the real-time network metrics;

[0029] Compare the real-time network metrics with the network metric requirements. If the real-time network metrics are lower than the network metric requirements, it is determined that voice enhancement is required.

[0030] In some embodiments of the present application, based on the foregoing solution, the use of the trained artificial intelligence voice model for voice enhancement includes:

[0031] If it is determined that voice enhancement is required through the real-time voice data, use the trained artificial intelligence voice model to intelligently adjust the frequency components of the current voice waveform so that the frequency components of the current voice waveform are accurately matched with the baseline voice waveform;

[0032] If it is determined that voice enhancement is required through the real-time voice quality assessment score or the real-time network metrics, convert the user voice data into text data and input it into the trained artificial intelligence voice model to be converted into voice and then output.

[0033] According to the second aspect of the embodiments of the present application, there is provided an online conference voice real-time enhancement device based on generative artificial intelligence, including:

[0034] A training unit for training a generative artificial intelligence voice model based on audio data of a user in different voice scenarios;

[0035] An analysis unit for performing spectral analysis on the audio data of the user in different voice scenarios to obtain a baseline voice waveform;

[0036] An evaluation unit for evaluating network performance and microphone performance before a meeting and obtaining an initial voice quality evaluation score according to the evaluation results;

[0037] A judgment unit for obtaining real-time voice data, real-time voice quality evaluation scores, and real-time network metrics during a meeting and judging whether voice enhancement is required;

[0038] An enhancement unit for performing voice enhancement using the trained artificial intelligence voice model.

[0039] According to a third aspect of an embodiment of the present application, there is provided a computer-readable storage medium storing computer instructions, which, when running on a computer, cause the computer to execute the method described in the first aspect above.

[0040] According to a fourth aspect of an embodiment of the present application, there is provided an electronic device including: a memory and a processor;

[0041] The memory for storing computer instructions;

[0042] The processor for calling the computer instructions stored in the memory, causing the electronic device to execute the method described in the first aspect above.

[0043] The technical solution of the present application accurately identifies the root cause of voice quality problems by comparing the baseline voice waveform and combining network condition evaluation, and at the same time uses the trained generative artificial intelligence voice model to perform personalized enhancement or replacement of the voice, realizing real-time monitoring, intelligent analysis, and precise optimization of the voice, and providing a clear and natural voice communication experience for online meetings.

[0044] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts. In the drawings:

[0046] Figure 1 Shows a schematic flowchart of a method for real-time enhancement of online meeting voice based on generative artificial intelligence according to an embodiment of the present application;

[0047] Figure 2 Shows a block diagram of a device for real-time enhancement of online meeting voice based on generative artificial intelligence according to an embodiment of the present application;

[0048] Figure 3 Shows a block diagram of an electronic device according to an embodiment of the present application;

[0049] Figure 4 Shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners

[0050] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0051] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will recognize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be employed. In other instances, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0052] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0053] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0054] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0055] The following will describe in detail some embodiments of the present application in conjunction with the accompanying drawings. Without conflict, the embodiments and features in the following embodiments may be combined with each other.

[0056] See Figure 1 , which shows a schematic flowchart of a method for real-time enhancement of online meeting voice based on generative artificial intelligence according to an embodiment of the present application.

[0057] As Figure 1 shown, a method for real-time enhancement of online meeting voice based on generative artificial intelligence is presented, which specifically includes steps S100 to S500.

[0058] Referring to Figure 1 , in step S100, the generative artificial intelligence voice model is trained based on the audio data of the user in different voice scenarios.

[0059] It can be understood that this step is implemented in the offline stage, that is, before the meeting.

[0060] In some feasible embodiments, based on the foregoing solution, the training of the generative artificial intelligence voice model based on the audio data of the user in different voice scenarios includes:

[0061] Collect the audio data of the user in different voice scenarios and the text corresponding to the audio data;

[0062] Perform data preprocessing on the collected audio data, and convert the audio data into training segments suitable for model learning through noise removal technology, audio level normalization operation, and audio data segmentation;

[0063] Train the generative artificial intelligence voice model based on the training segments and the text corresponding to the audio data.

[0064] It can be understood that different voice scenarios refer to voice scenarios such as when the user is speaking normally, at different speaking speeds, and at different volumes.

[0065] It is understandable that during the learning process, the generative AI voice model deeply explores the complex and subtle relationship between text and voice, accurately masters the voice detail features including intonation, rhythm, pronunciation, etc., and thus has the ability to simulate the user's personalized voice style.

[0066] Continue to refer to Figure 1 , step S200, perform spectrum analysis on the audio data of the user in different voice scenarios to obtain the baseline voice waveform.

[0067] It should be noted that the baseline voice waveform is used as an accurate representative of the user's voice features in the normal state and becomes the key benchmark for subsequent voice comparison.

[0068] In some feasible embodiments, based on the foregoing solution, the performing spectrum analysis on the audio data of the user in different voice scenarios to obtain the baseline voice waveform includes:

[0069] Break down the audio data into numerous frequency components in detail and calculate the statistical indicators of each frequency component;

[0070] Determine the baseline voice waveform according to the statistical indicators of each frequency component.

[0071] In this embodiment, based on the user's rich historical voice data and general voice quality standards, accurately determine the reasonable range of each frequency component, and then identify the baseline voice waveform, providing a more accurate reference basis for subsequent voice enhancement work.

[0072] Continue to refer to Figure 1 , step S300, before the meeting, evaluate the network performance and microphone performance, and obtain the initial voice quality evaluation score according to the evaluation results.

[0073] It should be noted that the evaluation of network performance specifically refers to testing multiple key indicators such as network speed, bandwidth availability, latency, jitter, packet loss rate, etc., and comparing these indicators with preset thresholds.

[0074] It should be noted that the evaluation of microphone performance specifically refers to checking core parameters such as the microphone's audio capture ability and sensitivity.

[0075] It should be noted that the initial voice quality evaluation score can be represented by the PESQ score (Perceptual Evaluation of Speech Quality score).

[0076] Continue to refer to Figure 1 , step S400, during the meeting, obtain real-time voice data, real-time voice quality evaluation score, and real-time network metrics, and determine whether voice enhancement is required.

[0077] In some feasible embodiments, based on the foregoing solution, the obtaining of the real-time voice data, the real-time voice quality evaluation score, and the real-time network metrics, and determining whether voice enhancement is required includes:

[0078] Comparing the real-time voice data with the baseline voice waveform to determine whether voice enhancement is required;

[0079] Comparing the real-time voice quality evaluation score with the initial voice quality evaluation score to determine whether voice enhancement is required;

[0080] Comparing the real-time network metrics with the network metric requirements to determine whether voice enhancement is required.

[0081] In some feasible embodiments, based on the foregoing solution, the comparing the real-time voice data with the baseline voice waveform to determine whether voice enhancement is required includes:

[0082] Obtaining the real-time voice data, performing spectral analysis on the real-time voice data, extracting the frequency components of the current voice waveform, and comparing the frequency components of the current voice waveform with the baseline voice waveform. If the difference between the two exceeds a preset threshold, it is determined that voice enhancement is required.

[0083] It can be understood that in this embodiment, the main focus is on analyzing the differences in frequency, amplitude, etc. between the frequency components of the current voice waveform and the baseline voice waveform.

[0084] In some feasible embodiments, based on the foregoing solution, the comparing the real-time voice quality evaluation score with the initial voice quality evaluation score to determine whether voice enhancement is required includes:

[0085] Calculating the real-time voice quality evaluation score, and comparing the real-time voice quality evaluation score with the initial voice quality evaluation score. If the real-time voice quality evaluation score is less than the initial voice quality evaluation score, it is determined that voice enhancement is required.

[0086] In some feasible embodiments, based on the foregoing solution, the comparing the real-time network metrics with the network metric requirements to determine whether voice enhancement is required includes:

[0087] Comprehensively evaluating the current network connection status to obtain the real-time network metrics;

[0088] Comparing the real-time network metrics with the network metric requirements. If the real-time network metrics are lower than the network metric requirements, it is determined that voice enhancement is required.

[0089] Continue to refer to Figure 1, step S500, if so, use the trained artificial intelligence voice model for voice enhancement.

[0090] In some feasible embodiments, based on the foregoing solution, the use of the trained artificial intelligence voice model for voice enhancement includes:

[0091] If it is determined from the real-time voice data that voice enhancement is required, use the trained artificial intelligence voice model to intelligently adjust the frequency components of the current voice waveform so that the frequency components of the current voice waveform are precisely matched with the baseline voice waveform;

[0092] If it is determined from the real-time voice quality assessment score or real-time network metrics that voice enhancement is required, convert the user voice data into text data and input it into the trained artificial intelligence voice model, convert it into voice and then output.

[0093] Exemplarily, if during a meeting, it is found that there are problems with the frequency components of the voice waveform and voice enhancement is required, use the well-trained generative artificial intelligence voice model to intelligently adjust the frequency components of the current voice waveform to precisely match the frequency components of the baseline voice waveform, thereby effectively enhancing the clarity and naturalness of the voice.

[0094] For example, when certain frequency components of the user's voice are interfered due to a noisy environment, the generative artificial intelligence voice model can targetedly adjust these frequencies to make the voice clear and intelligible again.

[0095] If the microphone fails or there are serious network connection problems that cause the voice to not be transmitted normally, promptly receive the text information input by the user (for example, through the convenient text input box of the online meeting software), and quickly input it into the generative artificial intelligence voice model. Relying on the accurate grasp of the user's personalized voice style, the generative artificial intelligence voice model quickly converts the input text into a natural and smooth voice output that conforms to the user's style, realizes efficient voice replacement, and ensures that the continuity of the meeting communication is not affected. Throughout the online meeting process, the system always records the user's voice data and detailed processing results in real time to generate a comprehensive operation record for subsequent in-depth analysis and further optimization.

[0096] In summary, the technical solution of this application can accurately locate the specific reasons for the decline in voice quality by comparing the user's current voice waveform with the baseline voice waveform in real time and combining a comprehensive evaluation of the network conditions. Whether it is the change in the user's own health status resulting in a change in voice or the transmission problems caused by fluctuations in a complex network environment, corresponding measures can be taken quickly and specifically. When the user has a cold and a hoarse voice, the differences in frequency, amplitude, etc. between the current voice waveform and the baseline waveform can be keenly identified, and the trained artificial intelligence voice model is used to optimize the voice, making the output voice closer to the clear voice of the user in a healthy state, enabling other meeting participants to easily understand the user's speech content. The generative artificial intelligence voice model carefully trained based on the user's specific training materials deeply learns the user's unique personalized voice features, including unique intonation, rhythm, and pronunciation habits, etc. During the voice enhancement process, accurate optimization can be carried out according to these personalized features, ensuring that the enhanced voice perfectly maintains the user's natural speaking style and effectively avoiding the unnatural or mechanical feeling that may be brought by general voice optimization methods. For users with a specific accent or a fast speaking speed, the clarity and intelligibility of the voice can be skillfully improved on the premise of fully respecting their accent and speaking speed characteristics, allowing users to express themselves naturally without restraint in an online meeting, while ensuring that the voice quality fully meets the requirements of efficient communication. Under different network conditions, the technical solution of this application demonstrates strong adaptability and can dynamically adjust the voice processing strategy. When the network bandwidth is sufficient and the connection is stable, more computing resources can be invested in further optimizing the detailed quality of the voice, such as improving the timbre of the voice and enriching the emotional expressiveness of the voice; while in a poor network environment, such as in a severe scenario with a high packet loss rate or low bandwidth, the basic clarity and coherence of the voice are prioritized, and the frequency components of the voice waveform are intelligently adjusted to effectively reduce the voice distortion caused by network transmission. By real-time evaluating the network metrics and comparing them with the baseline, the impact of network changes on the voice can be detected in a timely manner and a precise response can be made quickly, ensuring that high-quality voice communication services can be provided to users in various complex network environments, thereby effectively improving the overall communication efficiency and user experience of online meetings. When the microphone fails or the network connection is severely damaged, resulting in the inability to transmit voice normally, the voice replacement function of the technical solution of this application can intervene in a timely and effective manner. The user only needs to input text information and, with the help of a well-trained artificial intelligence voice model, can output voice in their familiar and natural voice style, perfectly avoiding the embarrassing situation of the interruption of meeting communication caused by equipment or network problems. This efficient fault response mechanism greatly ensures the continuity of online meetings, enabling users to still participate in meeting communication smoothly when encountering sudden equipment or network problems, greatly improving the reliability and stability of the online meeting system, and providing users with a more reliable and efficient online meeting voice communication solution.

[0097] The following introduces the device embodiments of the present application, which can be used to execute a method for real-time enhancement of online meeting voice based on generative artificial intelligence in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application above.

[0098] Referring to Figure 2 As shown, a device 200 for real-time enhancement of online meeting voice based on generative artificial intelligence according to an embodiment of the present application includes:

[0099] A training unit 201 for training a generative artificial intelligence voice model based on audio data of a user in different voice scenarios;

[0100] An analysis unit 202 for performing spectral analysis on audio data of a user in different voice scenarios to obtain a baseline voice waveform;

[0101] An evaluation unit 203 for evaluating network performance and microphone performance before a meeting, and obtaining an initial voice quality evaluation score according to the evaluation results;

[0102] A judgment unit 204 for obtaining real-time voice data, real-time voice quality evaluation scores, and real-time network metrics during a meeting, and judging whether voice enhancement is required;

[0103] An enhancement unit 205 for performing voice enhancement using the trained artificial intelligence voice model.

[0104] As Figure 3 shown, an embodiment of the present application also provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, the steps of the above method for real-time enhancement of online meeting voice based on generative artificial intelligence are implemented.

[0105] Since the electronic device introduced in this embodiment is the device adopted for implementing a device for real-time enhancement of online meeting voice based on generative artificial intelligence in the embodiments of the present application, based on the method introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present application will not be described in detail here. As long as the device adopted by those skilled in the art to implement the method in the embodiments of the present application belongs to the scope protected by the present application.

[0106] In the specific implementation process, when the computer program 311 is executed by the processor, it can implement any implementation manner in the corresponding embodiments of the first aspect.

[0107] Figure 4 The figure shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application.

[0108] It should be noted that Figure 4 The computer system 400 of the shown electronic device is only an example, and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0109] As Figure 4 shown, the computer system 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage section 408 into the random access memory (RAM) 403, such as executing the methods described in the above embodiments. In the RAM 403, various programs and data required for system operation are also stored. The CPU 401, ROM 402, and RAM 403 are connected to each other via a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.

[0110] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as required. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as required, so that the computer program read from it can be installed into the storage section 408 as required.

[0111] In particular, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, various functions defined in the system of the present application are executed.

[0112] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0114] The units involved in the embodiments described in the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.

[0115] As another aspect, the present application also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes a method for real-time enhancement of online conference voice based on generative artificial intelligence described in the above embodiments.

[0116] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device is enabled to implement a method for real-time enhancement of online conference voice based on generative artificial intelligence described in the above embodiments.

[0117] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0118] Those skilled in the art can easily understand from the description of the above embodiments that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0119] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A real-time speech enhancement method for online conferences based on generative artificial intelligence, characterized in that: include: Train the generative AI speech model based on the user’s audio data in different speech scenarios; Perform spectrum analysis on the user's audio data in different speech scenarios to obtain a baseline speech waveform; Before the meeting, evaluate the network performance and microphone performance, and obtain an initial voice quality evaluation score based on the evaluation results; During the conference, obtain real-time voice data, real-time voice quality assessment scores, and real-time network indicators, and determine whether voice enhancement is needed; If so, the trained artificial intelligence speech model is used for speech enhancement.

2. The method according to claim 1, characterized in that The training of the generative artificial intelligence speech model based on the audio data of the user in different speech scenarios includes: Collecting audio data of the user in different voice scenarios and text corresponding to the audio data; Preprocess the collected audio data and convert it into training segments suitable for model learning through noise removal technology, audio level normalization operation and audio data segmentation; A generative artificial intelligence speech model is trained based on the training segments and text corresponding to the audio data.

3. The method according to claim 1, characterized in that The method of performing spectrum analysis on the audio data of the user in different speech scenarios to obtain a baseline speech waveform includes: Decomposing the audio data into multiple frequency components, and calculating the statistical index of each frequency component; The baseline speech waveform is determined based on the statistical indicators of each frequency component.

4. The method according to claim 1, characterized in that The obtaining of real-time voice data, real-time voice quality evaluation scores, and real-time network indicators, and determining whether voice enhancement is required, includes: Comparing the real-time speech data with the baseline speech waveform to determine whether speech enhancement is required; Comparing the real-time speech quality assessment score with the initial speech quality assessment score to determine whether speech enhancement is required; Compare the real-time network indicators with the network indicator requirements to determine whether voice enhancement is needed.

5. The method according to claim 4, characterized in that The comparing the real-time speech data with the baseline speech waveform to determine whether speech enhancement is required includes: Real-time speech data is acquired, spectrum analysis is performed on the real-time speech data, frequency components of the current speech waveform are extracted, and the frequency components of the current speech waveform are compared with the baseline speech waveform. If the difference between the two exceeds a preset threshold, it is determined that speech enhancement is required.

6. The method according to claim 5, characterized in that The comparing the real-time speech quality evaluation score with the initial speech quality evaluation score to determine whether speech enhancement is required includes: A real-time speech quality evaluation score is calculated, and the real-time speech quality evaluation score is compared with the initial speech quality evaluation score. If the real-time speech quality evaluation score is less than the initial speech quality evaluation score, it is determined that speech enhancement is required.

7. The method according to claim 6, characterized in that The comparing the real-time network indicator with the network indicator requirement to determine whether voice enhancement is required includes: Comprehensively evaluate the current network connection status to obtain the real-time network indicators; The implemented network index is compared with the network index requirement, and if the real-time network index is lower than the network index requirement, it is determined that voice enhancement is required.

8. The method according to claim 7, characterized in that The method of using the trained artificial intelligence speech model to perform speech enhancement includes: If it is determined through real-time speech data that speech enhancement is required, the trained artificial intelligence speech model is used to intelligently adjust the frequency components of the current speech waveform so that the frequency components of the current speech waveform accurately match the baseline speech waveform; If the real-time voice quality assessment score or real-time network indicators determine that voice enhancement is required, the user voice data is converted into text data and input into the trained artificial intelligence voice model for conversion into voice and output.

9. A real-time speech enhancement device for online conferences based on generative artificial intelligence, characterized in that: include: A training unit, used to train a generative artificial intelligence speech model based on the user's audio data in different speech scenarios; An analysis unit, used to perform spectrum analysis on the audio data of the user in different speech scenarios to obtain a baseline speech waveform; An evaluation unit is used to evaluate the network performance and microphone performance before the meeting, and obtain an initial voice quality evaluation score based on the evaluation results; A judgment unit, used to obtain real-time voice data, real-time voice quality evaluation scores and real-time network indicators during the conference, and to judge whether voice enhancement is required; The enhancement unit is used to perform speech enhancement using the trained artificial intelligence speech model.

10. An electronic device, characterized in that: include: Memory and processor; The memory is used to store computer instructions; The processor is used to call the computer instructions stored in the memory so that the electronic device executes the method as described in any one of claims 1-8.