Voice injection method and device based on virtual sound card, computer equipment and medium

Through the voice injection method based on virtual sound card, the problems of high cost, high complexity and low broadcast accuracy in the prior art are solved, efficient and flexible voice injection are achieved, and the work efficiency and accuracy of the operators are improved.

CN119993116APending Publication Date: 2025-05-13PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226634.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has high cost, high technical complexity and operators are prone to missed or missed broadcasts when realizing voice injection, which affects work efficiency and accuracy.

Method used

The voice injection method based on the virtual sound card is adopted to convert the voice by obtaining the text content to be broadcast, and a voice file is generated, and the virtual sound card is used for signal processing, output audio signals, and finally transmitted to the client through a soft phone for broadcasting.

Benefits of technology

It realizes the simplicity and efficiency of voice injection, improves the flexibility of voice injection, enhances the work efficiency and work accuracy of workers, and reduces the overall development difficulty and maintenance cost of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993116A_ABST
    Figure CN119993116A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a voice injection method and device based on a virtual sound card, computer equipment and a medium, and the method comprises the steps: obtaining a to-be-broadcasted text content, and carrying out the voice conversion of the text content, and obtaining a to-be-broadcasted voice file; inputting the voice file into a virtual sound card, and performing signal processing on the voice file by using the virtual sound card to obtain an audio signal; and transmitting the audio signal to a client based on a soft phone, so that the client broadcasts the audio signal. According to the invention, the to-be-broadcasted voice file is input into the virtual sound card, the virtual sound card performs internal conversion and outputs the corresponding audio signal, and then the audio signal is transmitted to the client for broadcasting through the soft phone, so that voice injection can be simply and efficiently realized, the flexibility of voice injection is improved, and the user experience is improved. Therefore, the working efficiency and the working accuracy of operators are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a voice injection method, device, computer equipment and medium based on a virtual sound card. Background Art

[0002] Voice injection (also called remote voice injection or telephone injection) generally refers to the transmission of sound signals to a telephone system through a network in order to play pre-recorded messages, music, or other audio content during a call. This technology is widely used in scenarios such as medical calls, call centers, and telemarketing IVR (interactive voice response) systems. At the same time, with the rise of centralized operations, the workload of call center operators has gradually become saturated. Operators need to make and receive calls every day and broadcast standard scripts in accordance with business specifications. This type of standard script is a standardized fixed text statement with the characteristics of single oral content, long broadcast time, and high repetitiveness. With the continuous growth of business, this type of script is also increasing. This makes it easy for operators to miss or misbroadcast when broadcasting, which in turn leads to operational errors in the business.

[0003] In order to solve the above problems, the prior art combines the telephone system for docking development, and stores the voice to be injected in the telephone system in advance. When injection is required, the telephone system triggers the playback of the injected voice, converts the digital voice signal into an analog model, and then transmits it to an ordinary telephone through a traditional telephone line. Or use the SIP protocol (SessionInitiation Protocol) to send the voice data packet to a phone that supports SIP. However, both solutions have certain disadvantages. For example, by combining the telephone system of the call center for docking development, the entire telephone system needs to be transformed. Since the telephone systems used by various companies are different, the customization cost is high, and the initial investment and long-term operation costs will continue to increase. For another example, the use of the SIP protocol requires that both the phone and the voice injection system support SIP communication, and it is also necessary to develop a voice injection system that supports SIP communication, which leads to an increase in the overall development difficulty and technical complexity, and also requires a higher technical level to implement and maintain. Therefore, how to simply and efficiently realize voice injection, improve the flexibility of voice injection, and improve the work efficiency and operation accuracy of operators is a problem that technicians in this field need to solve. Summary of the invention

[0004] The embodiments of the present invention provide a voice injection method, device, computer equipment and storage medium based on a virtual sound card, aiming to simply and efficiently implement voice injection, improve the flexibility of voice injection, and thus improve the work efficiency and accuracy of operators.

[0005] In a first aspect, an embodiment of the present invention provides a voice injection method based on a virtual sound card, comprising:

[0006] Acquire the text content to be broadcast, and perform voice conversion on the text content to obtain the voice file to be broadcast;

[0007] Inputting the voice file into a virtual sound card, and performing signal processing on the voice file using the virtual sound card to obtain an audio signal;

[0008] Based on the softphone, the audio signal is transmitted to the client, so that the client broadcasts the audio signal.

[0009] In a second aspect, an embodiment of the present invention provides a voice injection device based on a virtual sound card, comprising:

[0010] A text acquisition unit, used to acquire the text content to be broadcast, and perform voice conversion on the text content to obtain a voice file to be broadcast;

[0011] A signal processing unit, used for inputting the voice file into a virtual sound card, and performing signal processing on the voice file using the virtual sound card to obtain an audio signal;

[0012] The audio broadcast unit is used to transmit the audio signal to the client based on the soft phone, so that the client broadcasts the audio signal.

[0013] In a third aspect, an embodiment of the present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the voice injection method based on the virtual sound card as described in the first aspect is implemented.

[0014] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the voice injection method based on the virtual sound card as described in the first aspect is implemented.

[0015] The embodiment of the present invention provides a voice injection method, device, computer equipment and storage medium based on a virtual sound card, the method comprising: obtaining text content to be broadcast, and performing voice conversion on the text content to obtain a voice file to be broadcast; inputting the voice file into a virtual sound card, and using the virtual sound card to perform signal processing on the voice file to obtain an audio signal; based on a softphone, transmitting the audio signal to a client, so that the client broadcasts the audio signal. The embodiment of the present invention inputs the voice file to be broadcast into a virtual sound card, so that the virtual sound card performs internal conversion and outputs the corresponding audio signal, and then transmits the audio signal to the client through the softphone for broadcasting, so that voice injection can be realized simply and efficiently, and the flexibility of voice injection can be improved, so as to improve the work efficiency and accuracy of the operator. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.

[0017] Figure 1 A schematic diagram of an application environment of a speech generation method based on a conditional stream matching model provided by an embodiment of the present invention;

[0018] Figure 2 A flow chart of a voice injection method based on a virtual sound card provided by an embodiment of the present invention;

[0019] Figure 3 A schematic diagram of a sub-flow of step S101 in a voice injection method based on a virtual sound card provided by an embodiment of the present invention;

[0020] Figure 4 A schematic diagram of a sub-flow of step S102 in a voice injection method based on a virtual sound card provided by an embodiment of the present invention;

[0021] Figure 5 A schematic diagram of a sub-flow of step S103 in a voice injection method based on a virtual sound card provided by an embodiment of the present invention;

[0022] Figure 6 A schematic block diagram of a voice injection device based on a virtual sound card provided by an embodiment of the present invention;

[0023] Figure 7 A first sub-schematic block diagram of a voice injection device based on a virtual sound card provided by an embodiment of the present invention;

[0024] Figure 8 A second sub-schematic block diagram of a voice injection device based on a virtual sound card provided by an embodiment of the present invention;

[0025] Fig. 9 A third sub-schematic block diagram of a voice injection device based on a virtual sound card provided by an embodiment of the present invention;

[0026] Fig.10 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0029] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0030] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0031] The voice injection method based on the virtual sound card provided by the embodiment of the present invention can be applied in Figure 1In an application environment, the client communicates with the server through a network, and the middle end communicates with the client and the server respectively. The middle end can obtain the text content to be broadcast, and perform voice conversion on the text content to obtain a voice file to be broadcast; input the voice file into a virtual sound card, and use the virtual sound card to perform signal processing on the voice file to obtain an audio signal; based on a soft phone, transmit the audio signal to the client, so that the client broadcasts the audio signal. In an embodiment of the present invention, by inputting the voice file to be broadcast into a virtual sound card, so that the virtual sound card performs internal conversion and outputs the corresponding audio signal, and then transmits the audio signal to the client through a soft phone for broadcast, voice injection can be realized simply and efficiently, and the flexibility of voice injection can be improved, so as to improve the work efficiency and accuracy of the operators.

[0032] See below Figure 2 , an embodiment of the present invention provides a voice injection method based on a virtual sound card, specifically comprising: steps S101 to S103.

[0033] Step S101, obtaining text content to be broadcast, and performing voice conversion on the text content to obtain a voice file to be broadcast;

[0034] Step S102: input the voice file into a virtual sound card, and use the virtual sound card to perform signal processing on the voice file to obtain an audio signal;

[0035] Step S103: Based on the softphone, the audio signal is transmitted to the client, so that the client broadcasts the audio signal.

[0036] In this embodiment, the text file to be broadcast is first obtained and converted into a corresponding voice file, and then the voice file is input into the virtual sound card for signal processing to obtain a corresponding audio signal, and then the audio signal is sent to the client through the soft phone, so that the client broadcasts the audio signal to the corresponding customer. In this embodiment, by inputting the voice file to be broadcast into the virtual sound card, the virtual sound card performs internal conversion and outputs the corresponding audio signal, and then the audio signal is transmitted to the client through the soft phone for broadcast, voice injection can be realized simply and efficiently, the flexibility of voice injection is improved, and the work efficiency and accuracy of the operator are improved.

[0037] In the actual operation center scenario, the operators at the seats can use the Chrome plug-in to extract the text information to be broadcast from the business system of the Chrome browser. Of course, other methods can also be used to obtain text data. For the acquired text, the voice assistant client plug-in of the seat can process it and generate a voice file. In addition, operators can also independently record standardized fixed text statements according to personal needs, and combine them with the operation process in the business system. When oral broadcast is required, the corresponding voice broadcast function is triggered by controlling the plug-in.

[0038] In one embodiment, the voice injection method based on the virtual sound card further includes:

[0039] The voice file is input into a physical sound card, and the voice file is synchronously transmitted to a server through the physical sound card, so that the server monitors the voice file.

[0040] In this embodiment, the voice file is transmitted to the server through the physical sound card, so that the server staff can monitor the audio signal broadcast by the client, so that when the customer makes a request through the client, the server staff can respond quickly. For example, when the customer interrupts temporarily, the staff can communicate with the customer in real voice in time based on monitoring, avoiding continuing to broadcast the voice file, thereby improving the customer experience. At the same time, it can also release the seat staff, so that the staff can get enough rest, so as to better provide customers with high-quality services.

[0041] In actual application scenarios, when monitoring through a physical sound card, first ensure the reliability of the physical sound card, such as ensuring that the physical sound card has been correctly installed on the corresponding device, and the installed driver can ensure that the various functions of the sound card can be realized, as well as good compatibility with the operating system and other software. When transmitting voice files through a physical sound card, if the server and the local device are in the same local area network or connected through the Internet, the network protocol can be used to realize the transmission of voice files. For example, the transmission method based on the TCP / IP protocol has reliable connection characteristics and can ensure that the voice file data is accurately transmitted to the server. After determining the transmission method and related parameters, start the transmission operation of the voice file audio signal, and send the processed audio signal to the server according to the set transmission path through the physical sound card. During the transmission process, monitor the transmission status in real time to determine whether there are abnormal situations such as data loss and transmission interruption. For example, when using the TCP / IP protocol for transmission, you can check whether the network connection status is normal and whether there are retransmission data packets. Once a problem is detected, it is necessary to take corresponding solutions in time, such as checking the network connection, restarting the transmission, etc., to ensure that the voice file audio signal can be smoothly and synchronously transmitted to the server. Subsequently, the server starts the monitoring function. For example, in an audio monitoring system, the server will start a dedicated audio receiving thread and configure the monitoring port number and other parameters, waiting to receive the audio signal of the voice file transmitted from the physical sound card. After the server receives the audio signal of the voice file transmitted from the physical sound card, it first performs operations such as data unpacking and decoding to restore it to the original audio signal form. Then, according to the specific monitoring needs, corresponding processing is performed, such as recording the relevant parameters of the audio signal (such as volume, frequency, etc.), storing the audio data to the local disk (if archiving is required), and playing the audio signal in real time for monitoring. During the entire monitoring process, the server must also continuously monitor the quality of the audio signal to check whether there are problems such as noise and signal interruption, so as to troubleshoot and solve the problems to ensure that the audio signal of the voice file can be monitored smoothly.

[0042] In one embodiment, if Figure 3 As shown, the step S101 includes: steps S201 to S205.

[0043] Step S201, obtaining the current business scenario;

[0044] Step S202: query the corresponding operation text according to the business scenario;

[0045] Step S203: judging whether there is a corresponding recording file according to the operation text;

[0046] Step S204: if it is determined that there is a recording file, setting the recording file as the voice file;

[0047] Step S205: If it is determined that there is no recording file, voice conversion is performed on the operation text, and the result of the voice conversion is set as the voice file.

[0048] In the process of obtaining text content and obtaining a voice file based on the text content, this embodiment first obtains the corresponding job file according to the business scenario, and then obtains the voice file according to the job file. For example, the agent voice assistant client plug-in can capture the current business scenario, such as checking bank card balance, credit card installment business, modifying user password and other businesses, and then access the URL and extract the key fields of the page to perform rule judgment. If it is identified as a bank card balance inquiry, the local pre-recorded bank card balance recording file list is retrieved, so that the agent can easily select the recording file to be played from the list.

[0049] In actual application scenarios, when obtaining text content, the corresponding operation text can be obtained based on the business scenario, or the operator can manually enter the corresponding text information according to actual needs. For example, in a medical call scenario, the obtained operation text may include asking for basic information, such as asking the patient's age and gender to understand the patient's basic situation in order to judge the condition and prepare corresponding rescue measures, asking what the specific symptoms are to further confirm the patient's detailed symptoms and provide a basis for diagnosis, and asking where the incident occurred to clarify the rescue location and ensure that the rescue personnel can arrive quickly. For another example, in a financial call scenario, the obtained operation text may include opening remarks, business inquiries and introductions, identity verification, information confirmation, etc. When the text content is obtained, if there is a corresponding recording file, it can be directly used as a voice file, and if there is no recording file, the text content needs to be further processed to obtain a voice file. Specifically, the text content can be preprocessed first, such as formatting the obtained text content, removing redundant spaces, line breaks, tabs, etc., so that the text presents a standardized format for subsequent processing. For example, when there is a text encoding inconsistency, the encoding conversion is performed to unify it into an encoding format supported by the system or application to ensure that the text can be correctly recognized and processed to avoid garbled characters and other problems. Next, the language used in the text content is determined, and the language recognition algorithm in the natural language processing technology is used to analyze the vocabulary, grammatical structure and other features in the text to determine the language type, such as Chinese, English, French, etc. For example, when a large number of Chinese characters and sentence structures that conform to Chinese grammatical habits appear in the text, it can be determined as Chinese, and then prepare for calling the Chinese speech synthesis module. Then the rhythmic features in the text are extracted, including stress, intonation, pauses and other information. For example, by analyzing the structure of the sentence, punctuation marks and keywords, it is determined which words need to be stressed, where there should be pauses, whether the tone is rising or falling, etc. For example, in a medical call scenario, when asking the user questions such as "whether there is coma, difficulty breathing, chest pain, bleeding", it is necessary to pause for different situations according to the comma in the text. For example, in a financial call scenario, when introducing different financial products to users, there is also a need for a certain pause.

[0050] After that, a speech synthesis model can be constructed to convert the text content into a speech file. For example, a suitable speech synthesis model can be selected based on factors such as the previously recognized language and the specific application scenario of the text, such as a rule-based model, a statistical parameter model, a deep learning model (such as a speech synthesis model of a Transformer architecture, etc.). Here, for simple voice prompts, such as the device startup voice prompt, a rule-based speech synthesis model may be sufficient to meet the needs, because it is simple and can quickly generate speech; while for scenarios such as audiobook synthesis that requires highly natural and emotional expression, it is more appropriate to choose a deep learning model, which can generate more realistic and more human voice characteristics. When constructing a speech synthesis model, the corresponding synthesis parameters are set, such as setting the timbre of speech synthesis, that is, the characteristics of the sound. Usually, different male voices, female voices, children's voices, etc. can be selected. You can also select a timbre with a specific style (such as gentle, calm, lively, etc.) according to specific needs. Different timbres correspond to different acoustic parameter configurations, which will affect the frequency, amplitude and other characteristics of the synthesized speech. For example, in the scenario of medical calls, a relatively calm and steady voice is usually required so that the user can have a calm conversation after answering the call to ensure timely and accurate information exchange. In the scenario of financial calls, when it is necessary to recommend financial products to users, emotional voice can be used to interact with users so that users can feel the advantages and characteristics of the products and increase their willingness to buy.

[0051] Another example is to set parameters such as speech speed and intonation, adjust the speech speed of speech synthesis, and set it reasonably according to factors such as the urgency and length of the text, so that the voice broadcast can clearly convey the content without appearing too fast or too slow; at the same time, set the basic mode of intonation, such as a steady intonation (such as a medical call for help) or an intonation with certain fluctuations (such as a financial product introduction) to meet the broadcast requirements of different texts. Using the selected speech synthesis model, the previously extracted text features, set parameters, etc. are used as input to perform model inference operations. For example, the deep learning model will process the input text information through its internal neural network layer according to the pre-trained weights and algorithms, and gradually generate corresponding speech acoustic features, such as Mel spectrum. Then, based on these acoustic features, it is converted into actual speech waveform signals through components such as vocoders, and finally generates speech audio data. This process involves complex mathematical operations and acoustic conversion mechanisms, and different models have different specific implementation methods.

[0052] Furthermore, according to application requirements and system compatibility, select a suitable audio format to save the voice file. Common audio formats include WAV, MP3, AAC, etc. Different audio formats have different encoding methods and characteristics. For example, the WAV format is a lossless audio format that can well preserve the original quality of the voice, but the file is relatively large; while the MP3 format is a lossy compression format with a smaller file size, which is easier to store and transmit. Comprehensive considerations should be made when selecting. According to the encoding requirements of the selected audio format, encode the generated voice audio data and convert it into an audio file of the corresponding format. This process involves operations such as audio data compression and format encapsulation to ensure that the voice file meets the format specifications.

[0053] In one embodiment, if Figure 4 As shown, the step S102 includes: steps S301 to S303.

[0054] Step S301, receiving the voice file through a virtual sound card, and performing format processing on the voice file to convert the format of the voice file into a pulse code modulation format;

[0055] Step S302: assigning channels to the formatted voice files according to a preset assignment rule, and associating the channel-assigned voice files with designated virtual devices;

[0056] Step S303: Acquire the audio format of the client, and convert the format of the voice file according to the audio format to obtain an audio signal that can be output by the virtual device to the client.

[0057] When outputting an audio signal through a virtual sound card, this embodiment first converts the format of the input voice file, that is, converts its format into a pulse code modulation format, then allocates a channel to the voice file, and associates it with a virtual device (such as a virtual microphone, etc.) to output the audio signal through the associated virtual device.

[0058] In a specific embodiment, the virtual sound card will allocate channels for audio data according to the routing rules pre-set by the user. For example, audio from different applications can be allocated to different virtual channels, such as allocating game audio to the left channel and music audio to the right channel, etc., to facilitate subsequent mixing operations; or according to the settings of the virtual device, the audio is directed to a specific virtual output port to simulate the functions of different interfaces of the real sound card. The virtual sound card can also associate audio data with corresponding virtual audio devices, such as virtual speakers, virtual headphones, etc., to determine which virtual device the audio should be output from, and this process is actually preparing for subsequent transmission to the real audio output device, so that the operating system and application programs can treat these virtual devices according to the logic of conventional audio devices. Furthermore, the virtual sound card can also change the timbre characteristics of the audio by adjusting the gain or attenuation of different frequency bands. For example, enhancing the high frequency band makes the sound clearer and brighter, and enhancing the low frequency band makes the sound heavier, so as to meet the user's needs for different sound styles.

[0059] In some optional embodiments, in order to avoid the delay problem of the virtual sound card, the buffer mechanism of the virtual sound card is optimized, such as setting the buffer size reasonably. Virtual sound card software usually allows users to adjust the size of the audio buffer by themselves. A smaller buffer can reduce audio delay and allow the sound to be output faster, but if the buffer is too small, audio jams may occur due to untimely data supply; while a larger buffer can avoid jams, but it will increase delays. Therefore, a suitable buffer value can be set based on comprehensive considerations such as device performance and the complexity of audio data to find a balance between delay and audio playback fluency. For example, in scenarios with high real-time requirements such as simple voice calls, the buffer can be appropriately reduced; when playing more complex audio such as high-bitrate music, the buffer can be appropriately increased to ensure smooth playback while controlling delay as much as possible. At the same time, the buffer is dynamically adjusted to monitor the transmission of audio data and the occupancy of computer system resources in real time. When it is found that the audio data transmission is smooth and the system resources are sufficient, the buffer size is automatically reduced to reduce delay; and when there is pressure on transmission and system resources are tight, such as running multiple resource-consuming programs at the same time, the buffer is correspondingly increased to avoid audio jams, and dynamically adapt to different usage scenarios to optimize delay performance. In addition, the data transmission efficiency can be optimized. Specifically, the virtual sound card driver can be optimized so that it can interact with the operating system and other related software more efficiently, speeding up the transmission of audio data in the system. You can also use efficient transmission protocols to select appropriate and efficient transmission protocols between the virtual sound card and the audio input source and output device. For example, some virtual sound cards support low-latency network audio transmission protocols, which can reduce the waiting time during data packaging, unpacking and transmission when transmitting audio data, ensuring that audio data flows quickly and accurately, thereby effectively reducing latency.

[0060] In one embodiment, if Figure 5 As shown, the step S103 includes: steps S401 to S403.

[0061] Step S401: encoding and compressing the audio signal through the softphone;

[0062] Step S402: encapsulating and packaging the encoded and compressed audio signal according to a preset protocol to obtain an audio data packet;

[0063] Step S403: Send the audio data packet to the client through the softphone, so that the client receives and decodes the audio data packet, and broadcasts an audio signal based on the decoded audio data packet.

[0064] In this embodiment, when transmitting an audio signal through a softphone, the softphone first needs to obtain the audio signal to be transmitted from the virtual sound card. In order to reduce the amount of audio data, improve transmission efficiency and reduce network bandwidth occupancy, the softphone usually encodes and compresses the acquired audio signal. Common audio encoding formats include G.711, G.723.1, G.729, AAC, MP3, Opus, etc., and the softphone will select a suitable encoding method based on factors such as network conditions and encoding formats supported by the client. The encoded audio data will be encapsulated and packaged according to a specific protocol for transmission on the network. For example, in VoIP communication, the real-time transport protocol (RTP) is usually used to encapsulate audio data packets, and the real-time transport control protocol (RTCP) may also be used to provide transmission quality feedback and control information. During the encapsulation process, some header information, such as the sequence number and timestamp of the data packet, will also be added for data reorganization and synchronization at the receiving end. Then, the softphone sends the packaged audio data packet to the client through the network. Specifically, it can occur based on a network connection based on the TCP / I protocol, the UDP protocol, etc. During the transmission process, the softphone can dynamically adjust the transmission strategy according to the network conditions to ensure the real-time and integrity of the audio data, such as automatically selecting the appropriate network route and adjusting the transmission rate.

[0065] In one embodiment, the voice injection method based on the virtual sound card further includes:

[0066] In response to the broadcast instruction sent by the server, the audio signal transmitted to the client is broadcast controlled.

[0067] Specifically, in response to the broadcast instruction sent by the server, the broadcast control of the audio signal transmitted to the client includes:

[0068] The broadcast instruction is parsed, and the broadcast parameters of the audio signal are extracted from the parsed broadcast instruction; wherein the broadcast parameters include broadcast time, broadcast volume and sound effect parameters.

[0069] The audio signal is transmitted and broadcast according to the broadcast parameters.

[0070] In this embodiment, when the operator monitors the voice file of the broadcast through the server, instructions about the broadcast content can be sent at any time, such as adjusting the broadcast volume, pausing the broadcast, and setting the broadcast sound effect, etc. Therefore, when the broadcast instruction sent by the server is received, it is parsed to determine the instruction information therein, and then the broadcast is controlled according to the instruction information. For example, in a medical call scene, when asking the user for patient information, the operator can decide whether to increase the volume by monitoring the user's voice to ensure the accurate transmission of the information, or when the user is emotionally excited, send instructions through the server to reduce the volume of the broadcast to avoid further stimulating the user. For example, in a financial call scene, when introducing financial products to users, the operator can send instructions through the server to adjust the speed and tone of the broadcast according to user feedback, so as to better attract the user's attention and highlight the advantages and characteristics of the product. In addition, specific sound effect parameters can also be set, such as adding background music, environmental sound effects, etc., to enhance the atmosphere and effect of voice broadcast. For example, in a medical call scenario, some calming and soothing background music can be added to help users stay calm; in a financial call scenario, some energetic sound effects can be added to stimulate users' desire to buy. After the broadcast parameters are extracted, the audio signal is transmitted and broadcasted according to these parameters to ensure that the broadcast effect meets the expectations of the operator.

[0071] In a specific embodiment, when a broadcast instruction is detected to be sent, the corresponding communication protocol parsing mechanism is used to extract the specific content of the broadcast instruction from the received message data. For example, the instruction may contain key information such as the start time and end time of the broadcast, the related identification of the broadcast audio signal, and the volume control requirements. For example, the instruction format is identified. Different servers may use different instruction formats, so it is necessary to identify the received instruction format. Common formats include JSON format, XML format, etc. For example, if it is a broadcast instruction in JSON format, it can be parsed according to the grammatical rules of JSON to extract the meaning and parameter values ​​corresponding to each field. Extract key parameters for audio broadcast control from the parsed instructions, which usually include audio selection parameters, that is, clearly knowing which audio signal needs to be broadcast, which can be determined by audio file number, audio stream identifier, etc., so as to accurately find the corresponding audio data for processing; time control parameters, that is, determine the start time and end time of the broadcast. If the instruction requires the broadcast to start from a specific time point, then the audio data must be located and intercepted according to this time point. For the end time, it is also necessary to be able to accurately control the audio to stop playing at the corresponding time to achieve precise duration control; volume and sound effect parameters, that is, obtain the setting requirements of the volume, sound effects (such as whether to add reverberation, echo and other special effects) in the instruction, so as to prepare for the subsequent corresponding processing of the audio signal.

[0072] In addition, according to the instruction information such as audio selection and time control in the broadcast instruction, the received audio signal is positioned. If it is a complete audio file, the corresponding audio start position can be found by means of timestamps, etc.; if it is an audio stream, it can be intercepted from the corresponding stream data according to the time range to ensure that only the part of the audio data that meets the requirements is broadcasted later. The audio signal can also be processed accordingly according to the volume and sound effect parameters in the broadcast instruction. For example, by calling the audio processing module, the amplitude of the audio signal is adjusted according to the volume required by the instruction, and the audio gain algorithm is used to adjust the volume to the specified level, such as increasing or decreasing the sampling value of the audio signal to achieve an increase or decrease in the volume. In addition, if the broadcast instruction requires the addition of specific sound effects, such as adding reverberation effects, the corresponding audio special effects algorithm can be used to simulate the acoustic environment to create a reverberation feeling for the audio signal; conversely, if certain sound effects are required to be removed, the original timbre characteristics of the audio can also be restored by corresponding inverse operations and other means.

[0073] In addition, during or after the audio broadcast, the playback status information can be fed back to the server based on the actual situation, such as whether the playback is successfully started, whether there are any abnormalities during the playback, whether the playback is finally completed, etc., so that the server can understand the execution status of the audio broadcast in a timely manner.

[0074] In a specific embodiment, in order to ensure the broadcast quality of the audio signal during the broadcast process, a reliable transmission protocol can be selected. For example, when the data integrity requirement is high, TCP (Transmission Control Protocol) can be used for audio data transmission. TCP has a reliable connection mechanism. It establishes a connection through a three-way handshake. During the transmission process, the data packet will be confirmed, retransmitted, etc., which can effectively avoid data loss and ensure that the client receives a complete audio data packet, laying the foundation for high-quality audio broadcasting. Even if a reliable transmission protocol is used, a small amount of erroneous data may appear due to the complex network environment. Therefore, an error detection and correction function can be set, such as using a cyclic redundancy check (CRC) and other technologies to verify the audio data packet. Once erroneous data is found, it is promptly retransmitted or repaired by an error correction algorithm to prevent the audio quality from being affected by data errors.

[0075] Furthermore, the audio signal processing can also be optimized, for example, according to the encoding format of the audio signal, the corresponding decoder is selected, and some parameters of the decoder are reasonably set. For example, when decoding some variable bit rate audio, the bit rate adaptive parameters of the decoder are adjusted so that it can better track the changes in the original audio bit rate, accurately restore the audio content, and prevent the audio quality from being degraded due to inaccurate decoding. For example, when adjusting the volume, in order to avoid the situation where the volume is too high and the audio signal is clipped and distorted, or the volume is too low and the sound is not clear, an audio processing algorithm is used to reasonably set the gain value according to the dynamic range of the audio signal to achieve a smooth increase or decrease in the volume, ensure that the audio is played at a suitable loudness, and maintain good sound quality. If sound effects need to be added, such as reverberation, echo, equalization, etc., it is necessary to ensure that the addition of sound effects is appropriate and in line with the characteristics of the audio content. Excessive use of sound effects may cover up the original sound details of the audio, resulting in poor sound quality. For example, when adding reverberation, according to the type of audio (such as voice, music, etc.) and scene requirements, the reverberation time, attenuation and other parameters are accurately controlled to create a natural acoustic effect that helps to enhance the auditory experience.

[0076] Furthermore, during the broadcast of the audio signal, key indicators of the audio signal are monitored in real time, such as volume, whether the audio is stuck, whether there is noise, etc. The audio analysis algorithm is used to analyze the audio signal being played in real time. Once an abnormal indicator is found, the corresponding solution is taken in time. At the same time, pay attention to the network connection status, because network fluctuations may affect the transmission and playback quality of subsequent audio data. By monitoring indicators such as network bandwidth, delay, and packet loss rate, when there is a problem with the network, such as narrowing of the bandwidth, the audio playback strategy can be adaptively adjusted, such as reducing the audio bit rate, pausing audio playback and waiting for network recovery, etc., to ensure the smoothness and quality of audio playback.

[0077] When an abnormal situation is detected, the audio quality-related issues will be promptly fed back, such as reporting audio playback freezes, noise, and other detailed information such as the corresponding time nodes. The cause will be analyzed based on this and corresponding adjustments will be made, such as resending audio data, optimizing transmission strategies, etc.

[0078] Specifically, a cyclic redundancy check (CRC) algorithm is used to control the quality of the audio signal. CRC is an error detection algorithm based on polynomial division. First, a generating polynomial is selected (for example, the commonly used polynomial for CRC-16 is x 16 +x 15 +x 2 +1), the audio data to be sent is regarded as a polynomial coefficient sequence, and then the data polynomial is divided by the generator polynomial, and the remainder is the CRC check code. The check code is attached to the end of the audio data and sent to the server. After receiving the data, the server uses the same generator polynomial to perform another division operation on the data containing the check code. If the remainder is 0, it is considered that there is a high probability that there is no error in the data; if the remainder is not 0, it indicates that an error has occurred during the data transmission process.

[0079] A parity check algorithm can also be used. Parity check is divided into odd check and even check. For odd check, the sender counts the number of "1" in the data before sending the audio data. If the number is even, a "1" is added to the end of the data to make the total number of "1" an odd number; if the number is odd, a "0" is added. Even check is the opposite. If the number of "1" in the data is odd, a "1" is added to make it an even number, and a "0" is added if it is an even number. The receiving end checks the received data according to the same parity check rule. If it does not meet the check rule, the data is judged to be wrong. Alternatively, the Hamming code algorithm can be used. Hamming code is a coding method that can correct single bit errors. It achieves error control by inserting specific redundant check bits into the original audio data. The position and value of these check bits are determined according to certain mathematical rules, which can detect and locate the bit position where the error occurs in the data, and then correct the error bit. For example, for a 7-bit Hamming code (containing 4 bits of original data and 3 check bits), these bits are generated and checked through a specific check equation. When the received data does not conform to the check equation, the position of the error bit can be calculated and corrected.

[0080] Figure 6 A schematic block diagram of a voice injection device 500 based on a virtual sound card provided in an embodiment of the present invention, the device 500 includes:

[0081] The text acquisition unit 501 is used to acquire the text content to be broadcast, and perform voice conversion on the text content to obtain the voice file to be broadcast;

[0082] The signal processing unit 502 is used to input the voice file into the virtual sound card, and use the virtual sound card to perform signal processing on the voice file to obtain an audio signal;

[0083] The audio broadcast unit 503 is used to transmit the audio signal to the client based on the softphone, so that the client broadcasts the audio signal.

[0084] In this embodiment, the text file to be broadcast is first obtained and converted into a corresponding voice file, and then the voice file is input into the virtual sound card for signal processing to obtain a corresponding audio signal, and then the audio signal is sent to the client through the soft phone, so that the client broadcasts the audio signal to the corresponding customer. In this embodiment, by inputting the voice file to be broadcast into the virtual sound card, the virtual sound card performs internal conversion and outputs the corresponding audio signal, and then the audio signal is transmitted to the client through the soft phone for broadcast, voice injection can be realized simply and efficiently, the flexibility of voice injection is improved, and the work efficiency and accuracy of the operator are improved.

[0085] In the actual operation center scenario, the operators at the seats can use the Chrome plug-in to extract the text information to be broadcast from the business system of the Chrome browser. Of course, other methods can also be used to obtain text data. For the acquired text, the voice assistant client plug-in of the seat can process it and generate a voice file. In addition, operators can also independently record standardized fixed text statements according to personal needs, and combine them with the operation process in the business system. When oral broadcast is required, the corresponding voice broadcast function is triggered by controlling the plug-in.

[0086] In one embodiment, the virtual sound card-based voice injection device 500 further includes:

[0087] The voice monitoring unit inputs the voice file into the physical sound card, and synchronously transmits the voice file to the server through the physical sound card, so that the server monitors the voice file.

[0088] In this embodiment, the voice file is transmitted to the server through the physical sound card, so that the server staff can monitor the audio signal broadcast by the client, so that when the customer makes a request through the client, the server staff can respond quickly. For example, when the customer interrupts temporarily, the staff can communicate with the customer in real voice in time based on monitoring, avoiding continuing to broadcast the voice file, thereby improving the customer experience. At the same time, it can also release the seat staff, so that the staff can get enough rest, so as to better provide customers with high-quality services.

[0089] In actual application scenarios, when monitoring through a physical sound card, first ensure the reliability of the physical sound card, such as ensuring that the physical sound card has been correctly installed on the corresponding device, and the installed driver can ensure that the various functions of the sound card can be realized, as well as good compatibility with the operating system and other software. When transmitting voice files through a physical sound card, if the server and the local device are in the same local area network or connected through the Internet, the network protocol can be used to realize the transmission of voice files. For example, the transmission method based on the TCP / IP protocol has reliable connection characteristics and can ensure that the voice file data is accurately transmitted to the server. After determining the transmission method and related parameters, start the transmission operation of the voice file audio signal, and send the processed audio signal to the server according to the set transmission path through the physical sound card. During the transmission process, monitor the transmission status in real time to determine whether there are abnormal situations such as data loss and transmission interruption. For example, when using the TCP / IP protocol for transmission, you can check whether the network connection status is normal and whether there are retransmission data packets. Once a problem is detected, it is necessary to take corresponding solutions in time, such as checking the network connection, restarting the transmission, etc., to ensure that the voice file audio signal can be smoothly and synchronously transmitted to the server. Subsequently, the server starts the monitoring function. For example, in an audio monitoring system, the server will start a dedicated audio receiving thread and configure the monitoring port number and other parameters, waiting to receive the audio signal of the voice file transmitted from the physical sound card. After the server receives the audio signal of the voice file transmitted from the physical sound card, it first performs operations such as data unpacking and decoding to restore it to the original audio signal form. Then, according to the specific monitoring needs, corresponding processing is performed, such as recording the relevant parameters of the audio signal (such as volume, frequency, etc.), storing the audio data to the local disk (if archiving is required), and playing the audio signal in real time for monitoring. During the entire monitoring process, the server must also continuously monitor the quality of the audio signal to check whether there are problems such as noise and signal interruption, so as to troubleshoot and solve the problems to ensure that the audio signal of the voice file can be monitored smoothly.

[0090] In one embodiment, if Figure 7 As shown, the text acquisition unit 501 includes:

[0091] A scene acquisition unit 601 is used to acquire the current business scene;

[0092] A job query unit 602 is used to query the corresponding job text according to the business scenario;

[0093] A file determination unit 603, used to determine whether there is a corresponding recording file according to the operation text;

[0094] The first determination unit 604 is configured to set the recording file as the voice file if it is determined that there is a recording file;

[0095] The second determination unit 605 is configured to perform voice conversion on the operation text if it is determined that there is no recording file, and set the result of the voice conversion as the voice file.

[0096] In the process of obtaining text content and obtaining a voice file based on the text content, this embodiment first obtains the corresponding job file according to the business scenario, and then obtains the voice file according to the job file. For example, the agent voice assistant client plug-in can capture the current business scenario, such as checking bank card balance, credit card installment business, modifying user password and other businesses, and then access the URL and extract the key fields of the page to perform rule judgment. If it is identified as a bank card balance inquiry, the local pre-recorded bank card balance recording file list is retrieved, so that the agent can easily select the recording file to be played from the list.

[0097] In actual application scenarios, when obtaining text content, the corresponding operation text can be obtained based on the business scenario, or the operator can manually enter the corresponding text information according to actual needs. For example, in a medical call scenario, the obtained operation text may include asking for basic information, such as asking the patient's age and gender to understand the patient's basic situation in order to judge the condition and prepare corresponding rescue measures, asking what the specific symptoms are to further confirm the patient's detailed symptoms and provide a basis for diagnosis, and asking where the incident occurred to clarify the rescue location and ensure that the rescue personnel can arrive quickly. For another example, in a financial call scenario, the obtained operation text may include opening remarks, business inquiries and introductions, identity verification, information confirmation, etc. When the text content is obtained, if there is a corresponding recording file, it can be directly used as a voice file, and if there is no recording file, the text content needs to be further processed to obtain a voice file. Specifically, the text content can be preprocessed first, such as formatting the obtained text content, removing redundant spaces, line breaks, tabs, etc., so that the text presents a standardized format for subsequent processing. For example, when there is a text encoding inconsistency, the encoding conversion is performed to unify it into an encoding format supported by the system or application to ensure that the text can be correctly recognized and processed to avoid garbled characters and other problems. Next, the language used in the text content is determined, and the language recognition algorithm in the natural language processing technology is used to analyze the vocabulary, grammatical structure and other features in the text to determine the language type, such as Chinese, English, French, etc. For example, when a large number of Chinese characters and sentence structures that conform to Chinese grammatical habits appear in the text, it can be determined as Chinese, and then prepare for calling the Chinese speech synthesis module. Then the rhythmic features in the text are extracted, including stress, intonation, pauses and other information. For example, by analyzing the structure of the sentence, punctuation marks and keywords, it is determined which words need to be stressed, where there should be pauses, whether the tone is rising or falling, etc. For example, in a medical call scenario, when asking the user questions such as "whether there is coma, difficulty breathing, chest pain, bleeding", it is necessary to pause for different situations according to the comma in the text. For example, in a financial call scenario, when introducing different financial products to users, there is also a need for a certain pause.

[0098] After that, a speech synthesis model can be constructed to convert the text content into a speech file. For example, a suitable speech synthesis model can be selected based on factors such as the previously recognized language and the specific application scenario of the text, such as a rule-based model, a statistical parameter model, a deep learning model (such as a speech synthesis model of a Transformer architecture, etc.). Here, for simple voice prompts, such as the device startup voice prompt, a rule-based speech synthesis model may be sufficient to meet the needs, because it is simple and can quickly generate speech; while for scenarios such as audiobook synthesis that requires highly natural and emotional expression, it is more appropriate to choose a deep learning model, which can generate more realistic and more human voice characteristics. When constructing a speech synthesis model, the corresponding synthesis parameters are set, such as setting the timbre of speech synthesis, that is, the characteristics of the sound. Usually, different male voices, female voices, children's voices, etc. can be selected. You can also select a timbre with a specific style (such as gentle, calm, lively, etc.) according to specific needs. Different timbres correspond to different acoustic parameter configurations, which will affect the frequency, amplitude and other characteristics of the synthesized speech. For example, in the scenario of medical calls, a relatively calm and steady voice is usually required so that the user can have a calm conversation after answering the call to ensure timely and accurate information exchange. In the scenario of financial calls, when it is necessary to recommend financial products to users, emotional voice can be used to interact with users so that users can feel the advantages and characteristics of the products and increase their willingness to buy.

[0099] Another example is to set parameters such as speech speed and intonation, adjust the speech speed of speech synthesis, and set it reasonably according to factors such as the urgency and length of the text, so that the voice broadcast can clearly convey the content without appearing too fast or too slow; at the same time, set the basic mode of intonation, such as a steady intonation (such as a medical call for help) or an intonation with certain fluctuations (such as a financial product introduction) to meet the broadcast requirements of different texts. Using the selected speech synthesis model, the previously extracted text features, set parameters, etc. are used as input to perform model inference operations. For example, the deep learning model will process the input text information through its internal neural network layer according to the pre-trained weights and algorithms, and gradually generate corresponding speech acoustic features, such as Mel spectrum. Then, based on these acoustic features, it is converted into actual speech waveform signals through components such as vocoders, and finally generates speech audio data. This process involves complex mathematical operations and acoustic conversion mechanisms, and different models have different specific implementation methods.

[0100] Furthermore, according to application requirements and system compatibility, select a suitable audio format to save the voice file. Common audio formats include WAV, MP3, AAC, etc. Different audio formats have different encoding methods and characteristics. For example, the WAV format is a lossless audio format that can well preserve the original quality of the voice, but the file is relatively large; while the MP3 format is a lossy compression format with a smaller file size, which is easier to store and transmit. Comprehensive considerations should be made when selecting. According to the encoding requirements of the selected audio format, encode the generated voice audio data and convert it into an audio file of the corresponding format. This process involves operations such as audio data compression and format encapsulation to ensure that the voice file meets the format specifications.

[0101] In one embodiment, if Figure 8 As shown, the signal processing unit 502 includes:

[0102] The format processing unit 701 is used to receive the voice file through the virtual sound card and perform format processing on the voice file to convert the format of the voice file into a pulse code modulation format;

[0103] The channel allocation unit 702 is used to allocate channels to the format-processed voice files according to a preset allocation rule, and associate the channel-allocated voice files with the specified virtual devices;

[0104] The format conversion unit 703 is used to obtain the audio format of the client, and convert the format of the voice file according to the audio format to obtain an audio signal that can be output by the virtual device to the client.

[0105] When outputting an audio signal through a virtual sound card, this embodiment first converts the format of the input voice file, that is, converts its format into a pulse code modulation format, then allocates a channel to the voice file, and associates it with a virtual device (such as a virtual microphone, etc.) to output the audio signal through the associated virtual device.

[0106] In a specific embodiment, the virtual sound card will allocate channels for audio data according to the routing rules pre-set by the user. For example, audio from different applications can be allocated to different virtual channels, such as allocating game audio to the left channel and music audio to the right channel, etc., to facilitate subsequent mixing operations; or according to the settings of the virtual device, the audio is directed to a specific virtual output port to simulate the functions of different interfaces of the real sound card. The virtual sound card can also associate audio data with corresponding virtual audio devices, such as virtual speakers, virtual headphones, etc., to determine which virtual device the audio should be output from, and this process is actually preparing for subsequent transmission to the real audio output device, so that the operating system and application programs can treat these virtual devices according to the logic of conventional audio devices. Furthermore, the virtual sound card can also change the timbre characteristics of the audio by adjusting the gain or attenuation of different frequency bands. For example, enhancing the high frequency band makes the sound clearer and brighter, and enhancing the low frequency band makes the sound heavier, so as to meet the user's needs for different sound styles.

[0107] In some optional embodiments, in order to avoid the delay problem of the virtual sound card, the buffer mechanism of the virtual sound card is optimized, such as setting the buffer size reasonably. Virtual sound card software usually allows users to adjust the size of the audio buffer by themselves. A smaller buffer can reduce audio delay and allow the sound to be output faster, but if the buffer is too small, audio jams may occur due to untimely data supply; while a larger buffer can avoid jams, but it will increase delays. Therefore, a suitable buffer value can be set based on comprehensive considerations such as device performance and the complexity of audio data to find a balance between delay and audio playback fluency. For example, in scenarios with high real-time requirements such as simple voice calls, the buffer can be appropriately reduced; when playing more complex audio such as high-bitrate music, the buffer can be appropriately increased to ensure smooth playback while controlling delay as much as possible. At the same time, the buffer is dynamically adjusted to monitor the transmission of audio data and the occupancy of computer system resources in real time. When it is found that the audio data transmission is smooth and the system resources are sufficient, the buffer size is automatically reduced to reduce delay; and when there is pressure on transmission and system resources are tight, such as running multiple resource-consuming programs at the same time, the buffer is correspondingly increased to avoid audio jams, and dynamically adapt to different usage scenarios to optimize delay performance. In addition, the data transmission efficiency can be optimized. Specifically, the virtual sound card driver can be optimized so that it can interact with the operating system and other related software more efficiently, speeding up the transmission of audio data in the system. You can also use efficient transmission protocols to select appropriate and efficient transmission protocols between the virtual sound card and the audio input source and output device. For example, some virtual sound cards support low-latency network audio transmission protocols, which can reduce the waiting time during data packaging, unpacking and transmission when transmitting audio data, ensuring that audio data flows quickly and accurately, thereby effectively reducing latency.

[0108] In one embodiment, if Fig. 9 As shown, the audio broadcast unit 503 includes:

[0109] The coding and compression unit 801 is used to perform coding and compression processing on the audio signal through the softphone;

[0110] The encapsulation and packaging unit 802 is used to encapsulate and package the audio signal after the encoding and compression processing according to a preset protocol to obtain an audio data packet;

[0111] The sending and decoding unit 803 is used to send the audio data packet to the client through the softphone, so that the client receives and decodes the audio data packet, and broadcasts an audio signal based on the decoded audio data packet.

[0112] In this embodiment, when transmitting an audio signal through a softphone, the softphone first needs to obtain the audio signal to be transmitted from the virtual sound card. In order to reduce the amount of audio data, improve transmission efficiency and reduce network bandwidth occupancy, the softphone usually encodes and compresses the acquired audio signal. Common audio encoding formats include G.711, G.723.1, G.729, AAC, MP3, Opus, etc., and the softphone will select a suitable encoding method based on factors such as network conditions and encoding formats supported by the client. The encoded audio data will be encapsulated and packaged according to a specific protocol for transmission on the network. For example, in VoIP communication, the real-time transport protocol (RTP) is usually used to encapsulate audio data packets, and the real-time transport control protocol (RTCP) may also be used to provide transmission quality feedback and control information. During the encapsulation process, some header information, such as the sequence number and timestamp of the data packet, will also be added for data reorganization and synchronization at the receiving end. Then, the softphone sends the packaged audio data packet to the client through the network. Specifically, it can occur based on a network connection based on the TCP / I protocol, the UDP protocol, etc. During the transmission process, the softphone can dynamically adjust the transmission strategy according to the network conditions to ensure the real-time and integrity of the audio data, such as automatically selecting the appropriate network route and adjusting the transmission rate.

[0113] In one embodiment, the virtual sound card-based voice injection device 500 further includes:

[0114] The broadcast control unit is used to control the broadcast of the audio signal transmitted to the client in response to the broadcast instruction sent by the server.

[0115] In one embodiment, the broadcast control list includes:

[0116] The instruction parsing unit is used to parse the broadcast instruction and extract the broadcast parameters of the audio signal from the parsed broadcast instruction; wherein the broadcast parameters include broadcast time, broadcast volume and sound effect parameters.

[0117] The parameter broadcasting unit is used to transmit and broadcast the audio signal according to the broadcasting parameters.

[0118] In this embodiment, when the operator monitors the voice file of the broadcast through the server, instructions about the broadcast content can be sent at any time, such as adjusting the broadcast volume, pausing the broadcast, and setting the broadcast sound effect, etc. Therefore, when the broadcast instruction sent by the server is received, it is parsed to determine the instruction information therein, and then the broadcast is controlled according to the instruction information. For example, in a medical call scene, when asking the user for patient information, the operator can decide whether to increase the volume by monitoring the user's voice to ensure the accurate transmission of the information, or when the user is emotionally excited, send instructions through the server to reduce the volume of the broadcast to avoid further stimulating the user. For example, in a financial call scene, when introducing financial products to users, the operator can send instructions through the server to adjust the speed and tone of the broadcast according to user feedback, so as to better attract the user's attention and highlight the advantages and characteristics of the product. In addition, specific sound effect parameters can also be set, such as adding background music, environmental sound effects, etc., to enhance the atmosphere and effect of voice broadcast. For example, in a medical call scenario, some calming and soothing background music can be added to help users stay calm; in a financial call scenario, some energetic sound effects can be added to stimulate users' desire to buy. After the broadcast parameters are extracted, the audio signal is transmitted and broadcasted according to these parameters to ensure that the broadcast effect meets the expectations of the operator.

[0119] In a specific embodiment, when a broadcast instruction is detected to be sent, the corresponding communication protocol parsing mechanism is used to extract the specific content of the broadcast instruction from the received message data. For example, the instruction may contain key information such as the start time and end time of the broadcast, the related identification of the broadcast audio signal, and the volume control requirements. For example, the instruction format is identified. Different servers may use different instruction formats, so it is necessary to identify the received instruction format. Common formats include JSON format, XML format, etc. For example, if it is a broadcast instruction in JSON format, it can be parsed according to the grammatical rules of JSON to extract the meaning and parameter values ​​corresponding to each field. Extract key parameters for audio broadcast control from the parsed instructions, which usually include audio selection parameters, that is, clearly knowing which audio signal needs to be broadcast, which can be determined by audio file number, audio stream identifier, etc., so as to accurately find the corresponding audio data for processing; time control parameters, that is, determine the start time and end time of the broadcast. If the instruction requires the broadcast to start from a specific time point, then the audio data must be located and intercepted according to this time point. For the end time, it is also necessary to be able to accurately control the audio to stop playing at the corresponding time to achieve precise duration control; volume and sound effect parameters, that is, obtain the setting requirements of the volume, sound effects (such as whether to add reverberation, echo and other special effects) in the instruction, so as to prepare for the subsequent corresponding processing of the audio signal.

[0120] In addition, according to the instruction information such as audio selection and time control in the broadcast instruction, the received audio signal is positioned. If it is a complete audio file, the corresponding audio start position can be found by means of timestamps, etc.; if it is an audio stream, it can be intercepted from the corresponding stream data according to the time range to ensure that only the part of the audio data that meets the requirements is broadcasted later. The audio signal can also be processed accordingly according to the volume and sound effect parameters in the broadcast instruction. For example, by calling the audio processing module, the amplitude of the audio signal is adjusted according to the volume required by the instruction, and the audio gain algorithm is used to adjust the volume to the specified level, such as increasing or decreasing the sampling value of the audio signal to achieve an increase or decrease in the volume. In addition, if the broadcast instruction requires the addition of specific sound effects, such as adding reverberation effects, the corresponding audio special effects algorithm can be used to simulate the acoustic environment to create a reverberation feeling for the audio signal; conversely, if certain sound effects are required to be removed, the original timbre characteristics of the audio can also be restored by corresponding inverse operations and other means.

[0121] In addition, during or after the audio broadcast, the playback status information can be fed back to the server based on the actual situation, such as whether the playback is successfully started, whether there are any abnormalities during the playback, whether the playback is finally completed, etc., so that the server can understand the execution status of the audio broadcast in a timely manner.

[0122] In a specific embodiment, in order to ensure the broadcast quality of the audio signal during the broadcast process, a reliable transmission protocol can be selected. For example, when the data integrity requirement is high, TCP (Transmission Control Protocol) can be used for audio data transmission. TCP has a reliable connection mechanism. It establishes a connection through a three-way handshake. During the transmission process, the data packet will be confirmed, retransmitted, etc., which can effectively avoid data loss and ensure that the client receives a complete audio data packet, laying the foundation for high-quality audio broadcasting. Even if a reliable transmission protocol is used, a small amount of erroneous data may appear due to the complex network environment. Therefore, an error detection and correction function can be set, such as using a cyclic redundancy check (CRC) and other technologies to verify the audio data packet. Once erroneous data is found, it is promptly retransmitted or repaired by an error correction algorithm to prevent the audio quality from being affected by data errors.

[0123] Furthermore, the audio signal processing can also be optimized, for example, according to the encoding format of the audio signal, the corresponding decoder is selected, and some parameters of the decoder are reasonably set. For example, when decoding some variable bit rate audio, the bit rate adaptive parameters of the decoder are adjusted so that it can better track the changes in the original audio bit rate, accurately restore the audio content, and prevent the audio quality from being degraded due to inaccurate decoding. For example, when adjusting the volume, in order to avoid the situation where the volume is too high and the audio signal is clipped and distorted, or the volume is too low and the sound is not clear, an audio processing algorithm is used to reasonably set the gain value according to the dynamic range of the audio signal to achieve a smooth increase or decrease in the volume, ensure that the audio is played at a suitable loudness, and maintain good sound quality. If sound effects need to be added, such as reverberation, echo, equalization, etc., it is necessary to ensure that the addition of sound effects is appropriate and in line with the characteristics of the audio content. Excessive use of sound effects may cover up the original sound details of the audio, resulting in poor sound quality. For example, when adding reverberation, according to the type of audio (such as voice, music, etc.) and scene requirements, the reverberation time, attenuation and other parameters are accurately controlled to create a natural acoustic effect that helps to enhance the auditory experience.

[0124] Furthermore, during the broadcast of the audio signal, key indicators of the audio signal are monitored in real time, such as volume, whether the audio is stuck, whether there is noise, etc. The audio analysis algorithm is used to analyze the audio signal being played in real time. Once an abnormal indicator is found, the corresponding solution is taken in time. At the same time, pay attention to the network connection status, because network fluctuations may affect the transmission and playback quality of subsequent audio data. By monitoring indicators such as network bandwidth, delay, and packet loss rate, when there is a problem with the network, such as narrowing of the bandwidth, the audio playback strategy can be adaptively adjusted, such as reducing the audio bit rate, pausing audio playback and waiting for network recovery, etc., to ensure the smoothness and quality of audio playback.

[0125] When an abnormal situation is detected, the audio quality-related issues will be promptly fed back, such as reporting audio playback freezes, noise, and other detailed information such as the corresponding time nodes. The cause will be analyzed based on this and corresponding adjustments will be made, such as resending audio data, optimizing transmission strategies, etc.

[0126] Specifically, a cyclic redundancy check (CRC) algorithm is used to control the quality of the audio signal. CRC is an error detection algorithm based on polynomial division. First, a generating polynomial is selected (such as the commonly used polynomial for CRC-16 is x16+x15+x2+1), and the audio data to be sent is regarded as a coefficient sequence of a polynomial. Then, this data polynomial is divided by the generating polynomial, and the remainder obtained is the CRC check code. The check code is attached to the audio data and sent to the server together. After the server receives the data, it uses the same generating polynomial to perform another division operation on the data containing the check code. If the remainder is 0, it is considered that the data is most likely not erroneous; if the remainder is not 0, it indicates that an error occurred during data transmission.

[0127] A parity check algorithm can also be used. Parity check is divided into odd check and even check. For odd check, the sender counts the number of "1" in the data before sending the audio data. If the number is even, a "1" is added to the end of the data to make the total number of "1" an odd number; if the number is odd, a "0" is added. Even check is the opposite. If the number of "1" in the data is odd, a "1" is added to make it an even number, and a "0" is added if it is an even number. The receiving end checks the received data according to the same parity check rule. If it does not meet the check rule, the data is judged to be wrong. Alternatively, the Hamming code algorithm can be used. Hamming code is a coding method that can correct single bit errors. It achieves error control by inserting specific redundant check bits into the original audio data. The position and value of these check bits are determined according to certain mathematical rules, which can detect and locate the bit position where the error occurs in the data, and then correct the error bit. For example, for a 7-bit Hamming code (containing 4 bits of original data and 3 check bits), these bits are generated and checked through a specific check equation. When the received data does not conform to the check equation, the position of the error bit can be calculated and corrected.

[0128] See also Fig.10 , Fig.10 The present invention provides a schematic block diagram of a computer device provided in an embodiment of the present invention. The computer device is a device with wireless communication and wired communication.

[0129] The computer device includes a processor 902 , a memory, and a network interface 905 connected via a system bus 901 , wherein the memory may include a non-volatile storage medium 903 and an internal memory 904 .

[0130] The non-volatile storage medium 903 can store an operating system 9031 and a computer program 9032. When the computer program 9032 is executed, the processor 902 can execute a voice injection method based on a virtual sound card.

[0131] The processor 902 is used to provide computing and control capabilities to support the operation of the entire computer device.

[0132] The internal memory 904 provides an environment for the operation of the computer program 9032 in the non-volatile storage medium 903. When the computer program 9032 is executed by the processor 902, the processor 902 can execute a voice injection method based on a virtual sound card.

[0133] The network interface 905 is used to communicate with other devices over the network. Fig.10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0134] The processor 902 is used to run a computer program 9032 stored in the memory to implement any embodiment of the above-mentioned voice injection method based on a virtual sound card.

[0135] It should be understood that in the embodiment of the present invention, the processor 902 may be a central processing unit (CPU), and the processor 902 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0136] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:

[0137] Acquire the text content to be broadcast, and perform voice conversion on the text content to obtain the voice file to be broadcast;

[0138] Inputting the voice file into a virtual sound card, and performing signal processing on the voice file using the virtual sound card to obtain an audio signal;

[0139] Based on the softphone, the audio signal is transmitted to the client, so that the client broadcasts the audio signal.

[0140] The embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed, the steps provided in the above embodiment can be implemented. The storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0141] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0142] Acquire the text content to be broadcast, and perform voice conversion on the text content to obtain the voice file to be broadcast;

[0143] Inputting the voice file into a virtual sound card, and performing signal processing on the voice file using the virtual sound card to obtain an audio signal;

[0144] Based on the softphone, the audio signal is transmitted to the client, so that the client broadcasts the audio signal.

[0145] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0146] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0147] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0148] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A voice injection method based on a virtual sound card, characterized in that: include: Acquire the text content to be broadcast, and perform voice conversion on the text content to obtain the voice file to be broadcast; Inputting the voice file into a virtual sound card, and performing signal processing on the voice file using the virtual sound card to obtain an audio signal; Based on the softphone, the audio signal is transmitted to the client, so that the client broadcasts the audio signal.

2. The voice injection method based on a virtual sound card according to claim 1, characterized in that: Also includes: The voice file is input into a physical sound card, and the voice file is synchronously transmitted to a server through the physical sound card, so that the server monitors the voice file.

3. The voice injection method based on a virtual sound card according to claim 1, characterized in that: The step of obtaining the text content to be broadcast and performing voice conversion on the text content to obtain the voice file to be broadcast includes: Get the current business scenario; Query the corresponding job text according to the business scenario; Determining whether there is a corresponding recording file according to the operation text; If it is determined that there is a recording file, setting the recording file as the voice file; If it is determined that there is no recording file, the operation text is voice-converted, and the result of the voice conversion is set as the voice file.

4. The voice injection method based on a virtual sound card according to claim 1, characterized in that: The step of inputting the voice file into a virtual sound card and performing signal processing on the voice file using the virtual sound card to obtain an audio signal comprises: The voice file is received through a virtual sound card, and format processing is performed on the voice file to convert the format of the voice file into a pulse code modulation format; Allocate channels for the format-processed voice files according to preset allocation rules, and associate the channel-allocated voice files with designated virtual devices; The audio format of the client is obtained, and the format of the voice file is converted according to the audio format to obtain an audio signal that can be output by the virtual device to the client.

5. The voice injection method based on a virtual sound card according to claim 1, characterized in that: The method of transmitting the audio signal to the client based on the softphone so that the client broadcasts the audio signal includes: Performing encoding and compression processing on the audio signal through the softphone; Encapsulating and packaging the encoded and compressed audio signal according to a preset protocol to obtain an audio data packet; The audio data packet is sent to the client through the softphone, so that the client receives and decodes the audio data packet and broadcasts an audio signal based on the decoded audio data packet.

6. The voice injection method based on a virtual sound card according to claim 1, characterized in that: Also includes: In response to the broadcast instruction sent by the server, the audio signal transmitted to the client is broadcast controlled.

7. The voice injection method based on a virtual sound card according to claim 6, characterized in that: The step of controlling the broadcast of the audio signal transmitted to the client in response to the broadcast instruction sent by the server includes: Parsing the broadcast instruction and extracting the broadcast parameters of the audio signal from the parsed broadcast instruction; wherein the broadcast parameters include broadcast time, broadcast volume and sound effect parameters; The audio signal is transmitted and broadcast according to the broadcast parameters.

8. A voice injection device based on a virtual sound card, characterized in that: include: A text acquisition unit, used to acquire the text content to be broadcast, and perform voice conversion on the text content to obtain a voice file to be broadcast; A signal processing unit, used for inputting the voice file into a virtual sound card, and performing signal processing on the voice file using the virtual sound card to obtain an audio signal; The audio broadcast unit is used to transmit the audio signal to the client based on the soft phone, so that the client broadcasts the audio signal.

9. A computer device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the voice injection method based on the virtual sound card as claimed in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the voice injection method based on a virtual sound card according to any one of claims 1 to 7 is implemented.