A wireless communication method, system, device and medium of an intelligent voice gateway
By using an intelligent voice gateway to perform time matching and similarity judgment on audio information from the voice receiving and playback ends, the problem of content transmission caused by poor network in wireless calls is solved, and the complete display of content is achieved when the network is unstable, thus improving the user experience.
Patent Information
- Application Number
- CN202211488608.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-11-25
AI Technical Summary
In existing wireless communication processes, poor network conditions may cause the call content to fail to be transmitted smoothly, or even result in lost or silent call content.
The intelligent voice gateway performs time matching and similarity comparison on the audio information of the voice receiver and the playback end, identifies voice pauses and marks the time, predicts the delay time, and converts low similarity audio into text display to ensure the content is conveyed.
When the network is unstable, converting voice to text ensures the complete transmission of the call content and improves the user experience.
Smart Images

Figure CN115695384B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication technology, and in particular to a wireless calling method, system, device, and computer-readable storage medium for an intelligent voice gateway. Background Technology
[0002] A VoIP gateway (Voice Over IP) is a device that converts analog telephone signals into digital signals to enable voice calls. However, in existing wireless calling systems, poor network conditions can lead to communication breakdowns, lost calls, or even complete silence. Summary of the Invention
[0003] In order to overcome the shortcomings of the prior art, one of the objectives of this invention is to provide a wireless calling method for an intelligent voice gateway, which can convert the user's voice into text for transmission when the call content cannot be transmitted smoothly, thus maintaining normal call usage.
[0004] The second objective of this invention is to provide a wireless communication system for an intelligent voice gateway.
[0005] The third objective of this invention is to provide an electronic device.
[0006] The fourth objective of this invention is to provide a computer-readable storage medium.
[0007] One of the objectives of this invention is achieved through the following technical solution:
[0008] A wireless calling method for an intelligent voice gateway, comprising:
[0009] Acquire the first audio information collected by the voice receiver and the second audio information played by the voice player, and compare the time-matched first audio information with the second audio information;
[0010] If it is determined that the similarity between the first audio information and the second audio information is less than a preset ratio, then the first audio information is used to generate text information and pushed to the display terminal for display.
[0011] Furthermore, when acquiring the first audio information, the method also includes:
[0012] In real time, determine whether there is a natural speech pause in the first audio information, and record the corresponding speech acquisition time for the audio record before the speech pause.
[0013] Furthermore, the method for obtaining the second audio information is as follows:
[0014] The second audio information being played is recorded using the voice playback terminal, and the corresponding recording time is marked on the audio before the pause in the voice.
[0015] Furthermore, before comparing the time-matched first audio information with the second audio information, the process further includes:
[0016] The audio delay time between the audio receiver and the audio player is tested in advance.
[0017] Furthermore, the method for comparing the time-matched first audio information with the second audio information is as follows:
[0018] When a natural speech pause is detected in the first audio information collected by the voice receiver, the audio is marked as the first target audio, and the corresponding speech collection time is superimposed with the speech delay time to obtain the target time.
[0019] The audio recording time that is closest to the target time in the second audio information is retrieved as the second target audio, and it is compared with the first target audio to determine whether the voice similarity between the two reaches a preset ratio.
[0020] Furthermore, the method for determining the speech similarity is as follows:
[0021] Extract audio from multiple time periods in the first target audio and the second target audio, determine whether the audio in each time period is the same, and calculate the speech similarity by combining the comparison results of the audio in each time period and the weight.
[0022] Furthermore, the method for determining the speech similarity is as follows:
[0023] Both the first target audio and the second target audio are forwarded as speech-text, and the speech similarity is calculated by speech-text comparison.
[0024] The second objective of this invention is achieved by the following technical solution:
[0025] A wireless calling system for an intelligent voice gateway includes a VoIP voice gateway, which is connected to a voice receiver for acquiring real-time user voice, a voice playback terminal for recording relayed voice, and a display terminal for displaying text information converted from voice. The VoIP voice gateway is used to execute the wireless calling method of the intelligent voice gateway described above.
[0026] The third objective of this invention is achieved by the following technical solution:
[0027] An electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the wireless calling method of the intelligent voice gateway as described above.
[0028] The fourth objective of this invention is achieved by the following technical solution:
[0029] A computer-readable storage medium having a computer program stored thereon, which, when executed, implements the above-described wireless communication method for an intelligent voice gateway.
[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0031] This invention collects the audio played by the voice playback terminal and compares it with the user's voice call. When the similarity is lower than a preset ratio, it means that the audio played by the voice playback terminal has problems such as no sound, stuttering, or missing voice. In this case, the collected user voice call is converted into text and displayed, thereby ensuring that the content of the call can be conveyed smoothly and improving the user experience. Attached Figure Description
[0032] Figure 1 This is a flowchart illustrating the wireless calling method of the intelligent voice gateway of the present invention.
[0033] Figure 2 This is a schematic diagram of the wireless communication system of the intelligent voice gateway of the present invention. Detailed Implementation
[0034] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.
[0035] Example 1
[0036] This embodiment provides a wireless calling method using a smart voice gateway, wherein the smart voice gateway used is a VoIP voice gateway, which can support multiple digital channels for signal transmission. (Refer to...) Figure 2 As shown, a voice receiver and a voice playback device are connected to a voice gateway. The voice receiver can be a telephone, mobile phone, or other device that can collect the user's speech during a call, convert the analog voice signal into a digital signal, and transmit it to the voice gateway. Then, the voice playback device converts the digital signal back into an analog signal for playback. The voice playback device also needs to have a recording function and can be a smart device such as a mobile phone. The voice playback device records the audio while playing it, thereby determining whether any parts of the transmitted voice have been lost or are silent based on the recorded audio content.
[0037] like Figure 1 As shown, the wireless calling method of the intelligent voice gateway specifically includes the following steps:
[0038] Acquire the first audio information collected by the voice receiver and the second audio information played by the voice player, and compare the time-matched first audio information with the second audio information;
[0039] If it is determined that the similarity between the first audio information and the second audio information is less than a preset ratio, then the first audio information is used to generate text information and pushed to the display terminal for display.
[0040] When a user initiates a voice call with another user, the user speaks into the voice receiver. The voice receiver captures the user's voice during the call and converts it into first audio information, which is then transmitted to the voice gateway. After acquiring the first audio information, the voice gateway determines in real time whether there are any natural pauses in the first audio information, i.e., whether there are any pauses in the call content that exceed a preset time. If a pause occurs, a pause tag is marked at the pause location, and the corresponding voice acquisition time / segment is recorded for the audio information before the pause.
[0041] The voice gateway transmits the first audio information to another user's voice playback terminal via the Internet for playback. While playing the call audio, the voice playback terminal records the second audio information being played, using the recorded audio as the second audio information. After acquiring the second audio information, the voice gateway also performs voice pause analysis on the second audio information and marks the corresponding voice recording time for the audio before the pause.
[0042] Since there may be a certain time delay between the user's voice and the playback during a voice call, it is necessary to test the voice delay time between the voice receiving end and the voice playing end before comparing the first audio information with the second audio information that matches the time.
[0043] When the voice gateway detects a natural pause in the first audio information collected by the voice receiver, it marks the audio as the first target audio. At this time, the voice acquisition time and voice delay time corresponding to the first target audio are superimposed to obtain the target time. Then, the audio with the voice recording time closest to the target time is retrieved from the second audio information as the second target audio. At this time, the second target audio can be identified as the call audio played after the first target audio is transmitted through the voice gateway. The audio content of the second target audio and the first target audio are compared to determine whether the voice similarity between the two reaches a preset ratio.
[0044] In this embodiment, the method for determining the speech similarity is as follows: the first target audio and the second target audio are divided into multiple sub-audio segments according to the same time interval, it is determined whether the audio segments in each same time segment are the same, and the speech similarity is calculated by combining the comparison results of the audio segments in each same time segment and the weight.
[0045] For example, extract sub-audio segments from the first to the third second and from the fifth to the sixth second of the first and second target audio. Compare the pronunciation of the sub-audio segments from the first to the third second of the first and second target audio segments to determine if the pronunciations within the sub-audio segments are identical. If they are identical, assign a value of "1" to the comparison result for that time segment; if they are completely different, assign a value of "0" to the comparison result for that time segment. If the pronunciations of the two sub-audio segments partially overlap, assign a value between "0" and "1" based on the percentage of overlap. Repeat the sub-audio segment comparison for multiple time segments, and multiply the comparison result of each time segment by the weight corresponding to each time segment to calculate the speech similarity value. The weight can be set according to the duration of the sub-audio segments.
[0046] In some embodiments, both the first target audio and the second target audio can be directly forwarded as speech-text, and the speech similarity can be calculated by comparing the speech-text one by one.
[0047] When it is determined that the voice similarity is lower than the preset value, it means that there is a large difference between the second target audio and the first target audio. At this time, it can be assumed that there may be omissions or interruptions in the transmission of the voice. Then, the first target audio is converted into text and transmitted to the display terminal for display. The display terminal can be a third-party terminal connected to the voice gateway, or a voice playback terminal with recording and display functions.
[0048] If there are any interruptions, omissions, or silences when the smart gateway is transmitting user call audio, it will convert the user's call content into text for display, ensuring that the content spoken during the call can be fully presented to another user, thus improving the user experience.
[0049] Example 2
[0050] This embodiment provides a wireless communication system for an intelligent voice gateway, such as... Figure 2 As shown, the system includes a VoIP voice gateway, a voice receiver connected to the VoIP voice gateway, and a voice playback terminal; it can also be matched with a corresponding display terminal for the system.
[0051] The voice receiver is used to collect the user's real-time voice; the voice playback terminal is used to play and record the broadcast voice; the display terminal is used to display text information converted from voice; and the VoIP voice gateway is used to execute the wireless calling method of the intelligent voice gateway as described in Embodiment 1.
[0052] In some embodiments, an electronic device is also provided, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the wireless communication method of the smart voice gateway in Embodiment 1. In addition, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed, it implements the aforementioned wireless communication method of the smart voice gateway.
[0053] The above-described system, device, and storage medium are based on multiple aspects of the same inventive concept as the methods in the foregoing embodiments. The implementation process of the method has been described in detail above, so those skilled in the art can clearly understand the structure and implementation process of the system, device, and storage medium in this embodiment based on the foregoing description. For the sake of brevity, it will not be described again here.
[0054] The above embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of protection of the present invention. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention shall fall within the scope of protection claimed by the present invention.
Claims
1. A wireless calling method for an intelligent voice gateway, characterized in that, include: Acquire the first audio information collected by the voice receiver and the second audio information played by the voice player, and compare the time-matched first audio information with the second audio information; If it is determined that the similarity between the first audio information and the second audio information is less than a preset ratio, then the first audio information is generated into text information and pushed and displayed. When obtaining the first audio information, the method further includes: Real-time determination of whether there are natural speech pauses in the first audio information, and recording the corresponding speech acquisition time for the audio record before the speech pause; The method for obtaining the second audio information is as follows: The second audio information being played is recorded using the voice playback terminal, and the corresponding voice recording time is marked on the audio before the voice pause. Before comparing the time-matched first audio information with the second audio information, the method further includes: Pre-test the audio delay time between the audio receiver and the audio playback device; The method for comparing the time-matched first audio information with the second audio information is as follows: When a natural speech pause is detected in the first audio information collected by the voice receiver, the audio is marked as the first target audio, and the corresponding speech collection time is superimposed with the speech delay time to obtain the target time. The audio recording time that is closest to the target time in the second audio information is retrieved as the second target audio, and it is compared with the first target audio to determine whether the voice similarity between the two reaches a preset ratio.
2. The wireless calling method of the intelligent voice gateway according to claim 1, characterized in that, The method for determining the speech similarity is as follows: Extract audio from multiple time periods in the first target audio and the second target audio, determine whether the audio in each time period is the same, and calculate the speech similarity by combining the comparison results of the audio in each time period and the weight.
3. The wireless calling method of the intelligent voice gateway according to claim 1, characterized in that, The method for determining the speech similarity is as follows: Both the first target audio and the second target audio are forwarded as speech-text, and the speech similarity is calculated by speech-text comparison.
4. A wireless communication system for an intelligent voice gateway, characterized in that, include: A VoIP voice gateway, which is connected to a voice receiver for collecting real-time voice from users and a voice playback terminal for recording broadcast voice. The VoIP voice gateway is used to perform the wireless calling method of the smart voice gateway as described in any one of claims 1 to 3.
5. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the wireless calling method of the smart voice gateway according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed, implements the wireless calling method of the intelligent voice gateway according to any one of claims 1 to 3.
Citation Information
Patent Citations
Method and device for transferring voice messages
CN102710539A