Intercom voice quality optimization method, device and equipment and storage medium
By acquiring the voice status of the intercom device in real time and combining it with the status of the other end device, the processing method is dynamically adjusted, which solves the problems of echo cancellation and noise suppression during two-way intercom, and improves the clarity of voice intercom and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN FANGWEI COMM TECH CO TD
- Filing Date
- 2021-08-07
- Publication Date
- 2026-05-29
Smart Images

Figure CN115706875B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intercom equipment technology, and in particular to a method, apparatus, device and storage medium for optimizing intercom voice quality. Background Technology
[0002] With the rapid development of the modern information technology industry, VoIP communication technology has also developed rapidly. Various voice communication devices, such as video conferencing systems, hands-free phones, mobile communications, and hearing aids, have emerged, making people's communication more convenient and comfortable.
[0003] In VoIP communication, the requirements for voice call quality are becoming increasingly stringent. Currently, most products use dedicated DSP processing chips to achieve echo cancellation and noise reduction. However, in embedded systems and IoT products, due to considerations of product size and cost, abandoning traditional dedicated DSP chips and implementing DSP algorithms in software on general-purpose chips is gradually becoming an industry trend.
[0004] However, while there are many DSP algorithms that achieve echo cancellation and noise reduction, most of them can only be applied to a specific situation. For example, some echo cancellation algorithms have a good echo cancellation effect when one end is speaking, but when both ends are speaking at the same time, the echo cancellation effect will also cancel the effective speech of the one end, resulting in the other end hearing intermittent sound and causing poor voice communication. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for optimizing intercom voice quality, in order to solve the technical problem of poor intercom voice quality in existing intercom devices.
[0006] To address the aforementioned problems, this application provides a method for optimizing intercom voice quality. This method utilizes a first intercom device, which is communicatively connected to a second intercom device. The method includes: acquiring sound information collected by the first intercom device in real time, and analyzing the sound information to obtain a first voice state of the first intercom device, including a call state and a mute state; recording the first voice state and sending it to the second intercom device, and receiving a second voice state sent by the second intercom device in real time; performing echo cancellation or noise suppression processing on the sound information according to the first and second voice states using a preset sound processing method, and then sending it to the second intercom device.
[0007] As a further improvement of this application, the method of acquiring sound information collected by the first intercom device in real time and analyzing the sound information to obtain the first voice state of the first intercom device includes: acquiring continuous sound information segments of the first intercom device itself according to a preset period, and confirming the first voice state according to the sound information segments.
[0008] As a further improvement of this application, the method of acquiring continuous audio information segments of the first intercom device itself according to a preset period and confirming the first voice state according to the audio information segments includes: acquiring continuous audio information segments of the first intercom device itself according to a preset period and confirming whether the currently recorded first intercom device is in a call state or a silent state; when the first intercom device is in a call state, sequentially identifying whether each audio information segment is voice or non-voice, and when all continuous audio information segments are non-voice, re-marking the voice state of the first intercom device as silent state, otherwise not changing the voice state of the first intercom device; when the first intercom device is in a silent state, sequentially identifying whether each audio information segment is voice or non-voice, and when all continuous audio information segments are voice, re-marking the voice state of the first intercom device as call state, otherwise not changing the voice state of the first intercom device.
[0009] As a further improvement to this application, the recognition of audio information segments is implemented based on VAD technology.
[0010] As a further improvement of this application, the sound information is processed by echo cancellation or noise suppression according to a preset sound processing method based on the first voice state and the second voice state before being sent to the second intercom device, including: acquiring the first voice state and the second voice state; when the first voice state is a call state and the second voice state is a silent state, or when both the first voice state and the second voice state are silent states, performing strong noise reduction processing on the sound information collected by the first intercom device before sending it to the second intercom device; when the first voice state is a silent state and the second voice state is a call state, or when both the first voice state and the second voice state are a call state, performing echo cancellation processing on the sound information collected by the first intercom device before sending it to the second intercom device.
[0011] As a further improvement of this application, when the first voice state is a mute state and the second voice state is a call state, or when both the first and second voice states are call states, echo cancellation processing is performed on the sound information collected by the first intercom device, including: when the first voice state is a mute state and the second voice state is a call state, strong echo cancellation processing is performed on the sound information collected by the first intercom device; when both the first and second voice states are call states, weak echo cancellation processing is performed on the sound information collected by the first intercom device, wherein the echo cancellation capability of the weak echo cancellation processing is weaker than that of the strong echo cancellation processing.
[0012] As a further improvement of this application, when both the first voice state and the second voice state are in a call state, while performing weak echo cancellation processing on the sound information collected by the first intercom device, the method also includes: receiving the voice information sent by the second intercom device, and suppressing the voice information of the second intercom device before playing it.
[0013] To address the aforementioned issues, this application provides, in another aspect, a voice quality optimization device for intercom systems, comprising: a status confirmation module, configured to acquire sound information collected by a first intercom device in real time and analyze the sound information to obtain a first voice status of the first intercom device, the voice status including a call status and a mute status; a status transceiver module, configured to record the first voice status and transmit the first voice status to a second intercom device, and receive a second voice status transmitted by the second intercom device in real time; and a sound processing module, configured to perform echo cancellation or noise suppression processing on the sound information according to the first and second voice statuses using a preset sound processing method, and then transmit the processed information to the second intercom device.
[0014] To address the aforementioned problems, this application provides a computer device in another aspect, the computer device including a processor and a memory coupled to the processor, the memory storing program instructions, which, when executed by the processor, cause the processor to perform the steps of the intercom voice quality optimization method as described above.
[0015] To address the aforementioned problems, this application provides, in another aspect, a storage medium storing program instructions capable of implementing the intercom voice quality optimization method as described above.
[0016] The beneficial effects of this application are as follows: The intercom voice quality optimization method of this application analyzes the sound information collected by the intercom device to determine whether the local intercom device is in a call or silent state. Then, based on the voice states of the local and remote devices, echo cancellation or noise suppression is performed on the sound information collected by the local intercom device according to a preset sound processing method. That is, whether the local intercom device performs echo cancellation or noise suppression is related to the voice states of the local and remote intercom devices, rather than being fixed. This allows the local intercom device to adjust the sound processing algorithm in real time according to the voice state, and to use the appropriate sound processing algorithm at the appropriate time to avoid affecting the voice quality. At the same time, it effectively achieves echo cancellation and noise reduction processing of the sound. Attached Figure Description
[0017] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the present disclosure are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:
[0018] Figure 1 This is a flowchart illustrating the method for optimizing intercom voice quality according to an embodiment of the present invention;
[0019] Figure 2 This is a schematic diagram of the functional modules of the intercom voice quality optimization device according to an embodiment of the present invention;
[0020] Figure 3 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0021] Figure 4 This is a schematic diagram of the structure of the storage medium according to an embodiment of the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0023] The specific embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0024] Figure 1 This is a flowchart illustrating the intercom voice quality optimization method according to an embodiment of the present invention. In this embodiment, the intercom voice quality optimization method is applied to a first intercom device, which is communicatively connected to a second intercom device. It should be noted that in this embodiment, the first intercom device and the second intercom device constitute a two-way intercom network. The first intercom device refers to the local device in the two-way intercom network, and the second intercom device is the peer device in the two-way intercom network. It should be understood that either end in the two-way intercom network can serve as the first intercom device, and the other end serves as the second intercom device. That is, for the two-way intercom network, any local device is itself a first intercom device, and any peer device is a second intercom device.
[0025] like Figure 1 As shown, the method for optimizing the voice quality of intercom includes:
[0026] Step S101: Acquire the sound information collected by the first intercom device in real time, and analyze the sound information to obtain the first voice state of the first intercom device, which includes the call state and the mute state.
[0027] Typically, during two-way communication, situations may arise such as one-way speaking, two-way speaking, or two-way muting. One-way speaking can be further divided into situations where the user speaks while the other is muted, or vice versa. Therefore, for the first intercom device, its corresponding voice state includes both a call state and a muted state. In step S101, to adjust the sound processing algorithm based on the voice state of the intercom device, this embodiment needs to acquire the currently collected sound information of the first intercom device in real time. This sound information may be the user's voice or meaningless noise. After acquiring the sound information, it is analyzed to determine whether it is voice or noise. If the sound information includes voice, it indicates that the user is currently sending voice to the second intercom device through the first intercom device, meaning the first intercom device is in a call state. If the sound information does not include voice, it indicates that the user is not inputting voice into the first intercom device, meaning the first intercom device is in a muted state.
[0028] In some embodiments, in order to ensure the accuracy of identification, step S101 may further be: acquiring continuous audio information segments of the first intercom device itself according to a preset period, and confirming the first voice state according to the audio information segments.
[0029] Specifically, by identifying multiple consecutive audio information segments separately, it can be confirmed whether each audio information segment corresponds to speech information or noise information. Then, the identification results of multiple audio information segments can be used to confirm the first voice state of the first intercom device, reducing the probability of misjudgment caused by a single detection.
[0030] It should be understood that, in order to prevent interference with the intercom process, the preset period is usually set to a very short time period, such as 10 milliseconds or 20 milliseconds. The number of consecutive audio information segments can also be preset, such as 8 segments or 10 segments. When a certain number of the consecutive audio information segments are voice information segments, the first intercom device can be considered to be in a call state; otherwise, it is in a mute state.
[0031] Furthermore, the voice status of the first intercom device can be comprehensively judged in conjunction with the current voice status of the first intercom device. Therefore, the step of obtaining continuous audio information segments of the first intercom device according to a preset period and confirming the first voice status based on the audio information segments specifically includes:
[0032] 1. Obtain continuous audio information segments of the first intercom device according to a preset period, and confirm whether the currently recorded first intercom device is in a call state or a silent state.
[0033] 2. When the first intercom device is in call mode, identify whether each audio information segment is voice or non-voice in turn. When all consecutive audio information segments are non-voice, re-mark the voice status of the first intercom device as silent. Otherwise, do not change the voice status of the first intercom device.
[0034] Specifically, it needs to be understood that the first voice state of the first intercom device must be recorded after confirmation. In this embodiment, when it is confirmed that the currently recorded first voice state of the first intercom device is a call state, each audio information segment is identified sequentially according to the time sequence to determine whether each audio information segment is audio or non-audio. When consecutive audio information segments are all non-audio, it is considered that the voice state of the first intercom device has changed, and the first voice state of the first intercom device is re-marked as silent and recorded; otherwise, the first voice state of the first intercom device is not changed. Non-audio includes noise information, etc.
[0035] 3. When the first intercom device is in a silent state, identify whether each audio information segment is voice or non-voice in turn. When all consecutive audio information segments are voice, re-mark the voice status of the first intercom device as call status; otherwise, do not change the voice status of the first intercom device.
[0036] Specifically, when it is confirmed that the first intercom device is in a silent state, multiple audio information segments are identified. If all of them are voice, the first voice state of the first intercom device is re-marked as a call state and recorded. Otherwise, no change is made.
[0037] Preferably, in this embodiment, the recognition of audio information segments can be achieved through VAD technology.
[0038] VAD technology, also known as speech activity detection, speech endpoint detection, or speech boundary detection, is mainly used to accurately locate the start and end points of speech in noisy speech.
[0039] Step S102: Record the first voice status and send the first voice status to the second intercom device, and receive the second voice status sent by the second intercom device in real time.
[0040] In step S102, after obtaining the first voice state of the first intercom device based on the sound information, the first voice state is recorded. At the same time, the first voice state needs to be sent to the second intercom device, and the second voice state sent by the second intercom device is received, so that the first intercom device can know the voice state of the second intercom device.
[0041] Step S103: Based on the first voice state and the second voice state, perform echo cancellation or noise suppression processing on the sound information according to the preset sound processing method, and then send it to the second intercom device.
[0042] In step S103, after obtaining the first voice state and the second voice state, the sound processing method of the first intercom device is adjusted in real time according to the first voice state and the second voice state. For example, when the first intercom device is in a call state and the second intercom device is in a silent state, the sound information collected by the first intercom device can be noise-reduced and then sent to the second intercom device. When the first intercom device is in a silent state and the second intercom device is in a call state, the first intercom device can be echo-cancelled to prevent echoes from being transmitted to the second intercom device and avoid affecting the user experience of the second intercom device.
[0043] Furthermore, step 103 specifically includes:
[0044] 1. Obtain the first voice state and the second voice state.
[0045] Specifically, the first voice state is the voice state of the first intercom device, and the second intercom state is the second voice state of the second voice device.
[0046] 2. When the first voice state is in call state and the second voice state is in mute state, or when both the first and second voice states are in mute state, the sound information collected by the first intercom device is subjected to strong noise reduction processing before being sent to the second intercom device.
[0047] Specifically, when the first voice state is in a call state and the second voice state is in a mute state, or when both the first and second voice states are in a mute state, the second intercom device remains in a mute state and does not send voice information to the first intercom device for playback. At this time, the sound acquisition sensor of the first intercom device collects very little echo and has a lot of noise. At this time, regardless of whether the first intercom device is in a call state or a mute state, strong noise reduction processing can be performed on the sound information collected by the first intercom device. If the first intercom device is in a call state, strong noise reduction processing can effectively remove the noise in the sound information, making the voice information transmitted to the second intercom device clearer and the voice quality higher. If the first intercom device is in a mute state, strong noise reduction processing can prevent the first intercom device from transmitting the noise it collects to the second intercom device, thus improving the user experience of the second intercom device.
[0048] 3. When the first voice state is in a mute state and the second voice state is in a call state, or when both the first and second voice states are in a call state, the sound information collected by the first intercom device is processed for echo cancellation before being sent to the second intercom device.
[0049] Specifically, when the first voice state is muted and the second voice state is in call state, or when both the first and second voice states are in call state, the second intercom device is always in call state, sending voice information to the first intercom device and playing it on the first intercom device. At this time, the sound information collected by the sound acquisition sensor of the first intercom device contains a lot of echo information, which has a significant impact on the voice quality. At this time, regardless of whether the first intercom device is in call state or muted state, echo cancellation processing can be performed on the sound information collected by the first intercom device. If the first intercom device is also in call state at this time, echo cancellation processing can remove the echo in the sound information, making the voice information transmitted to the second intercom device clearer and the voice quality higher. If the first intercom device is in muted state at this time, echo cancellation processing can prevent the first intercom device from transmitting its own collected echo to the second intercom device, improving the user experience of the second intercom device.
[0050] Furthermore, when the first voice state is a mute state and the second voice state is a call state, or both the first and second voice states are call states, the step of performing echo cancellation processing on the sound information collected by the first intercom device specifically includes:
[0051] 3.1 When the first voice state is mute and the second voice state is call state, the sound information collected by the first intercom device is processed for strong echo cancellation.
[0052] 3.2 When both the first voice state and the second voice state are in call state, weak echo cancellation processing is performed on the sound information collected by the first intercom device.
[0053] In this embodiment, echo cancellation processing includes two methods: strong echo cancellation processing and weak echo cancellation processing. The echo cancellation capability of weak echo cancellation processing is weaker than that of strong echo cancellation processing.
[0054] Specifically, when the first voice state is mute and the second voice state is call state, the first intercom device does not need to collect voice information to send to the second intercom device. Therefore, a strong echo cancellation method is used to cancel the strong echo in the voice information collected by the first intercom device, thereby improving the echo cancellation effect. When both the first and second voice states are call states, the first intercom device also needs to send voice information to the second intercom device. In this case, to avoid deleting the normal voice information collected by the first intercom device, a weak echo cancellation method is used to cancel the weak echo in the voice information collected by the first intercom device. This can eliminate the echo in the voice information to a certain extent while ensuring the quality of the voice information and avoiding intermittent playback.
[0055] Furthermore, in this embodiment, while performing weak echo cancellation processing on the sound information collected by the first intercom device when both the first voice state and the second voice state are in a call state, it also includes:
[0056] 3.3 Receive voice information sent by the second intercom device, and suppress the voice information of the second intercom device before playing it.
[0057] Specifically, in order to avoid excessive echo in the sound information collected by the first intercom device, which would prevent the echo from being effectively eliminated, in this embodiment, after receiving the voice information sent by the second intercom device, the first intercom device suppresses the voice information sent by the second intercom device and then plays it on the first intercom device, so that the first intercom device does not collect excessive echo when collecting sound information.
[0058] The intercom voice quality optimization method of this invention analyzes the sound information collected by the intercom device to determine whether the local intercom device is in a call or silent state. Then, based on the voice states of the local and remote devices, it performs echo cancellation or noise suppression processing on the sound information collected by the local intercom device according to a preset sound processing method. That is, whether the local intercom device performs echo cancellation or noise suppression processing is related to the voice states of the local and remote intercom devices, rather than being fixed. This allows the local intercom device to adjust the sound processing algorithm in real time according to the voice state, and use the appropriate sound processing algorithm at the appropriate time to avoid affecting the voice quality, while effectively realizing echo cancellation and noise reduction processing of the sound.
[0059] Figure 2 A schematic diagram of the functional modules of the intercom voice quality optimization device according to an embodiment of the present invention is shown. For example... Figure 2 As shown, the intercom voice quality optimization device 20 includes: a status confirmation module 21, a status transceiver module 22, and a sound processing module 23.
[0060] The status confirmation module 21 is used to acquire the sound information collected by the first intercom device in real time and analyze the sound information to obtain the first voice status of the first intercom device, which includes a call status and a silent status; the status transceiver module 22 is used to record the first voice status and send the first voice status to the second intercom device, and receive the second voice status sent by the second intercom device in real time; the sound processing module 23 is used to perform echo cancellation or noise suppression processing on the sound information according to the first voice status and the second voice status according to a preset sound processing method, and then send it to the second intercom device.
[0061] Preferably, the operation of the status confirmation module 21 to acquire the sound information collected by the first intercom device in real time and analyze the sound information to obtain the first voice status of the first intercom device can also be: acquiring continuous sound information segments of the first intercom device itself according to a preset period, and confirming the first voice status according to the sound information segments.
[0062] Preferably, the operation of the status confirmation module 21 to obtain continuous audio information segments of the first intercom device itself according to a preset period and to confirm the first voice status according to the audio information segments can also be as follows: obtain continuous audio information segments of the first intercom device itself according to a preset period and confirm whether the currently recorded first intercom device is in a call state or a silent state; when the first intercom device is in a call state, sequentially identify whether each audio information segment is voice or non-voice, and when all continuous audio information segments are non-voice, re-mark the voice status of the first intercom device as silent state, otherwise do not change the voice status of the first intercom device; when the first intercom device is in a silent state, sequentially identify whether each audio information segment is voice or non-voice, and when all continuous audio information segments are voice, re-mark the voice status of the first intercom device as call state, otherwise do not change the voice status of the first intercom device.
[0063] Preferably, the recognition of audio information segments is achieved based on VAD technology.
[0064] Preferably, the operation of the sound processing module 23 to perform echo cancellation or noise suppression processing on the sound information according to the first voice state and the second voice state in accordance with a preset sound processing method before sending it to the second intercom device can also be as follows: acquiring the first voice state and the second voice state; when the first voice state is a call state and the second voice state is a silent state, or when both the first voice state and the second voice state are silent states, performing strong noise reduction processing on the sound information collected by the first intercom device before sending it to the second intercom device; when the first voice state is a silent state and the second voice state is a call state, or when both the first voice state and the second voice state are a call state, performing echo cancellation processing on the sound information collected by the first intercom device before sending it to the second intercom device.
[0065] Preferably, the operation of the sound processing module 23 to perform echo cancellation processing on the sound information collected by the first intercom device when the first voice state is a mute state and the second voice state is a call state, or when both the first voice state and the second voice state are call states, can also be as follows: when the first voice state is a mute state and the second voice state is a call state, perform strong echo cancellation processing on the sound information collected by the first intercom device; when both the first voice state and the second voice state are a call state, perform weak echo cancellation processing on the sound information collected by the first intercom device, wherein the echo cancellation capability of the weak echo cancellation processing is weaker than that of the strong echo cancellation processing.
[0066] Preferably, when the first voice state and the second voice state are both in a call state, the sound processing module 23 performs weak echo cancellation processing on the sound information collected by the first intercom device, and at the same time, it is also used to: receive the voice information sent by the second intercom device, and suppress the voice information of the second intercom device before playing it.
[0067] For other details regarding the implementation techniques of each module in the intercom voice quality optimization device of the above embodiments, please refer to the description in the intercom voice quality optimization method of the above embodiments, which will not be repeated here.
[0068] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0069] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Figure 3 As shown, the computer device 30 includes a processor 31 and a memory 32 coupled to the processor 31. The memory 32 stores program instructions. When the program instructions are executed by the processor 31, the processor 31 performs the steps of the intercom voice quality optimization method described in any of the above embodiments.
[0070] The processor 31 can also be referred to as a CPU (Central Processing Unit). The processor 31 may be an integrated circuit chip with signal processing capabilities. The processor 31 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor.
[0071] See Figure 4 , Figure 4This is a schematic diagram of the structure of a storage medium according to an embodiment of the present invention. The storage medium of this embodiment stores program instructions 41 capable of implementing all the above methods. These program instructions 41 can be stored in the storage medium in the form of a software product, including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or computer devices such as computers, servers, mobile phones, and tablets.
[0072] In the several embodiments provided in this application, it should be understood that the disclosed computer devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0073] In the foregoing description of this specification, unless otherwise expressly specified and limited, the terms "fixed," "installed," "connected," or "linked" should be interpreted broadly. For example, the term "linked" can refer to a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; or it can refer to the internal communication of two components or the interaction between two components. Therefore, unless otherwise expressly limited in this specification, those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0074] Based on the above description in this specification, those skilled in the art will also understand that terms used, such as "upper," "lower," "front," "rear," "left," "right," "length," "width," "thickness," "vertical," "horizontal," "top," "bottom," "inner," "outer," "axial," "radial," "circumferential," "center," "longitudinal," "transverse," "clockwise," or "counterclockwise," are terms indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings of this specification. They are only for the purpose of facilitating the explanation of the present invention and simplifying the description, and do not imply that the devices or elements involved must have the specific orientation, or be constructed and operated in a specific orientation. Therefore, the above-mentioned orientation or positional relationship terms should not be understood or interpreted as limitations on the present invention.
[0075] Furthermore, the terms "first" or "second," etc., used in this specification to refer to numbers or ordinal numbers are for descriptive purposes only and should not be construed as indicating, explicitly or implicitly, relative importance or specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this specification, "a plurality of" means at least two, such as two, three, or more, unless otherwise explicitly specified.
[0076] While various embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of the invention. The appended claims are intended to define the scope of protection of the invention and therefore cover modular compositions, equivalents, or alternatives within the scope of these claims.
Claims
1. A method for optimizing the quality of intercom voice messages, characterized in that, The method utilizes a first intercom device, which is communicatively connected to a second intercom device; the method includes: The system acquires the sound information collected by the first intercom device itself in real time and analyzes the sound information to obtain the first voice state of the first intercom device itself, which includes a call state and a mute state. Record the first voice status and send the first voice status to the second intercom device, and receive the second voice status sent by the second intercom device in real time; Based on the first voice state and the second voice state, the sound information is processed by echo cancellation or noise suppression according to a preset sound processing method, including: When the first voice state is the call state and the second voice state is the mute state, or both the first and second voice states are the mute state, the sound information collected by the first intercom device is subjected to strong noise reduction processing before being sent to the second intercom device; when the first voice state is the mute state and the second voice state is the call state, or both the first and second voice states are the call state, the sound information collected by the first intercom device is subjected to echo cancellation processing before being sent to the second intercom device.
2. The method for optimizing intercom voice quality according to claim 1, characterized in that, The step of acquiring and analyzing the sound information collected by the first intercom device in real time to obtain the first voice state of the first intercom device includes: The first intercom device acquires continuous audio information segments of itself according to a preset period, and confirms the first voice state based on the audio information segments.
3. The method for optimizing intercom voice quality according to claim 2, characterized in that, The step of acquiring continuous audio information segments of the first intercom device itself according to a preset period, and confirming the first voice state based on the audio information segments, includes: According to a preset period, the first intercom device itself obtains continuous audio information segments and confirms whether the currently recorded first intercom device is in a call state or a silent state. When the first intercom device is in the call state, each sound information segment is identified as either voice or non-voice. When all consecutive sound information segments are non-voice, the voice state of the first intercom device is re-marked as the mute state; otherwise, the voice state of the first intercom device is not changed. When the first intercom device is in the silent state, each audio information segment is sequentially identified as either speech or non-speech. When all consecutive audio information segments are speech, the speech state of the first intercom device is relabeled as the call state; otherwise, the speech state of the first intercom device is not changed.
4. The method for optimizing intercom voice quality according to claim 3, characterized in that, The recognition of the audio information segment is achieved based on VAD technology.
5. The method for optimizing intercom voice quality according to claim 1, characterized in that, When the first voice state is the mute state and the second voice state is the call state, or both the first voice state and the second voice state are the call state, echo cancellation processing is performed on the sound information collected by the first intercom device, including: When the first voice state is the silent state and the second voice state is the call state, the sound information collected by the first intercom device is subjected to strong echo cancellation processing. When both the first voice state and the second voice state are in the call state, the sound information collected by the first intercom device is subjected to weak echo cancellation processing. The echo cancellation capability of the weak echo cancellation processing is weaker than that of the strong echo cancellation processing.
6. The method for optimizing intercom voice quality according to claim 5, characterized in that, When both the first voice state and the second voice state are in the call state, the process of performing weak echo cancellation processing on the sound information collected by the first intercom device also includes: The system receives voice information sent by the second intercom device, suppresses the voice information from the second intercom device, and then plays it back.
7. A device for optimizing the voice quality of intercom, characterized in that, It includes: The status confirmation module is used to acquire the sound information collected by the first intercom device itself in real time, and analyze the sound information to obtain the first voice status of the first intercom device itself, including the call status and the mute status. The status transceiver module is used to record the first voice status and send the first voice status to the second intercom device, and to receive the second voice status sent by the second intercom device in real time. The sound processing module is used to perform echo cancellation or noise suppression processing on the sound information according to the first voice state and the second voice state in accordance with a preset sound processing method, including: when the first voice state is the call state and the second voice state is the silent state, or both the first voice state and the second voice state are the silent state, performing strong noise reduction processing on the sound information collected by the first intercom device, and then sending it to the second intercom device; when the first voice state is the silent state and the second voice state is the call state, or both the first voice state and the second voice state are the call state, performing echo cancellation processing on the sound information collected by the first intercom device, and then sending it to the second intercom device.
8. A computer device, characterized in that, The computer device includes a processor and a memory coupled to the processor, the memory storing program instructions that, when executed by the processor, cause the processor to perform the steps of the intercom voice quality optimization method as described in any one of claims 1-6.
9. A storage medium, characterized in that, The storage medium stores program instructions capable of implementing the intercom voice quality optimization method as described in any one of claims 1-6.