Earphone, earphone chip and intercom method for intercom

Through the microphone, filter processor, buffer and switch in the walkie-talkie headset chip, the voice data is recognized to be earlier than the PTT signal and compressed and cached or marked, which solves the problem of voice loss in the walkie-talkie and realizes the complete transmission of voice.

CN119815241BActive Publication Date: 2025-09-16SHENZHEN VOCODER TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510303309.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-09-16
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

In existing intercoms, the speaking and PTT button operations are not coordinated, resulting in the problem that the spoken voice is lost at the initial and end stages.

Method used

The system uses the microphone, filter processor, buffer, switch and recognition controller in the earphone chip to perform voice compression processing and cache when the voice data is recognized before the PTT control signal appears, and outputs the compressed voice data when the PTT signal appears, or adds an initial identifier and an end identifier at the beginning of the voice to ensure complete transmission.

Benefits of technology

It effectively solves the problem of voice loss caused by the lack of coordination between speaking and PTT button operation, ensuring that the spoken voice is sent completely at the initial and end stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119815241B_ABST
    Figure CN119815241B_ABST
Patent Text Reader

Abstract

The present application discloses a headset, headset chip, and intercom method for an intercom. The headset includes: a microphone, a filter processor, a buffer, a switch, and a first recognition controller, as well as a PTT input terminal and a voice output terminal. The filter processor is used to perform digital-to-analog conversion on the voice signal output by the microphone to obtain raw voice data and perform filtering processing. The first recognition controller recognizes that the raw voice data appears before the PTT control signal at the PTT input terminal, controls the raw voice data to be stored in a buffer after voice compression processing, and controls the switch to be connected to the buffer. When the PTT control signal appears, the compressed voice data in the buffer is output to the voice output terminal. Based on this application, the problem of lost speech at the initial and final stages of intercom speech due to the lack of coordination between speaking and PTT button operation in the prior art can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of intercoms, and in particular to an earphone, an earphone chip, and an intercom method for an intercom. Background Art

[0002] The two-way radio's voice communication mode is simplex, requiring the PTT (Push To Talk) button to switch between sending and receiving. Typically, the user presses the PTT button before speaking, and stops speaking before releasing it. This ensures that the voice is transmitted while the PTT button is held down. Therefore, the user must be able to coordinate speaking with the PTT button.

[0003] However, in actual use, speaking and pressing the PTT button often fail to work together. In particular, if you speak first and then press the PTT button, some of the first words you spoke won't be sent. Alternatively, if you lift the PTT button before you've finished speaking, the rest of your speech won't be sent. This prevents the other party from receiving the complete voice message, forcing them to repeat themselves. Summary of the Invention

[0004] The main technical problem solved by this application is to provide headphones, headphone chips and intercom methods for intercoms, so as to solve the problem in the prior art that the spoken voice is lost in the initial and ending stages due to the lack of coordination between speaking and PTT button operation.

[0005] To solve the above technical problems, the present application adopts a technical solution to provide a headset for a walkie-talkie, comprising: a microphone, a filter processor, a buffer, a switch, and a first recognition controller, as well as a PTT input terminal and a voice output terminal. The switch is respectively connected to the filter processor, the buffer, the first recognition controller, and the voice output terminal. The first recognition controller is also electrically connected to the filter processor and the PTT input terminal, and the filter processor is also electrically connected to the microphone and the buffer, and is configured to perform digital-to-analog conversion on the voice signal output by the microphone to obtain original voice data and filter the original voice data. The first recognition controller recognizes the original voice data generated by the filter processor and also recognizes a PTT control signal from the PTT input terminal. If the first recognition controller recognizes that the original voice data appears earlier than the PTT control signal, it controls the original voice data to be stored in the buffer after voice compression processing, and controls the switch to be connected to the buffer. When it is recognized that the PTT control signal appears, the compressed voice data in the buffer is output to the voice output terminal.

[0006] In some embodiments, the filtering processor includes an audio amplification unit, an AD sampling unit, a compression filtering unit, and a power detection unit. The audio amplification unit is used to perform distortion-free amplification on the audio signal input from the microphone to obtain an analog audio signal that has undergone power amplification, and then enters the AD sampling unit for analog-to-digital conversion to obtain the original voice data; the original voice data enters the compression filtering unit, and compression filtering is performed on the original voice data to remove low-power components in the original voice data, thereby obtaining the compressed voice data.

[0007] In some embodiments, the power detection unit is used to perform power detection on the original voice data output by the AD sampling unit to detect whether there is a voice signal input.

[0008] In some embodiments, the filtering processor further includes a de-interference filtering unit configured to filter out background interference from the original voice data output by the AD sampling unit and output de-interferenced voice data.

[0009] In some embodiments, the first recognition controller includes a first MCU, a first voice recognition unit, a first PPT recognition unit, and a time difference calculation unit electrically connected to the first MCU; the first voice recognition unit recognizes the original voice data from the filtering processor, and then forwards the recognition result to the first MCU. After the first MCU receives the valid recognition signal from the first voice recognition unit, it records the first voice moment when the valid recognition signal arrives; the first PPT recognition unit is also electrically connected to the PTT input end, recognizes the PTT control signal connected to the PTT input end, and sends the recognition result to the first MCU. After the first MCU receives the valid PTT control signal from the first PPT recognition unit, it indicates that the PTT button has been pressed, and records the first PTT moment when the valid PTT control signal arrives; the first MCU is electrically connected to the time difference calculation unit, and is used to send both the first voice moment and the first PTT moment to the time difference calculation unit, which is used to determine the order of recognizing the first voice moment and the first PTT moment, and calculate the time difference between the two moments, that is, the first time difference.

[0010] In some embodiments, the first recognition controller also includes an identification generating unit, and the first MCU is electrically connected to the identification generating unit. After the first MCU determines that the first voice moment is earlier than the first PTT moment, the identification generating unit generates an initial voice identification, which is added to the front of the original voice data as the starting part of the voice, and then output through the voice output end.

[0011] In some embodiments, the switching switch includes a first input end, a second input end, a control end and an output end, the first input end is electrically connected to the buffer, the second input end is electrically connected to the filtering processor, the control end is electrically connected to the first recognition controller, and the output end is electrically connected to the voice output end.

[0012] The present application also provides an earphone chip, the interior of the earphone chip includes a buffer, a filter processor, a first recognition controller and a switch, as well as a PTT input connection line and a voice output connection line; the pins of the earphone chip include a power pin, a ground pin, a voice output pin, a PTT input pin and a pickup input pin, the power pin and the ground pin are used to connect to the positive and negative poles of direct current, the voice output pin is electrically connected to the voice output connection line, the PTT input pin is electrically connected to the PTT input connection line, and the pickup input pin is electrically connected to the filter processor for connecting to a pickup; the switch is respectively connected to the filter processor, the buffer, the first recognition controller and the voice output connection line, the first recognition controller is also electrically connected to the filter processor, the PTT The input connection line and the filter processor are also electrically connected to the sound pickup input pin and the buffer, and are used to perform digital-to-analog conversion on the voice signal of the sound pickup input pin to obtain original voice data, and filter the original voice data; the first recognition controller recognizes the original voice data generated by the filter processor, and also recognizes the PTT control signal from the PTT input pin; if the first recognition controller recognizes that the original voice data appears earlier than the PTT control signal, it controls the original voice data to be stored in the buffer after voice compression processing, and controls the switch to be connected to the buffer. When it is recognized that the PTT control signal appears, the compressed voice data in the buffer is output to the voice output pin.

[0013] The present application also provides an intercom method, which utilizes the aforementioned earphone for an intercom to conduct intercom speech, comprising the steps of:

[0014] The first recognition controller recognizes the presence of a valid recognition signal indicating the presence of a voice signal and records the first voice moment of its presence , and the PTT control signal has not appeared at this time, the first recognition controller controls the filter processor to perform compression processing and sends the compressed voice data to the buffer for caching; the first recognition controller recognizes the PTT valid signal of the PTT pressing and records the first PTT moment of its appearance , calculate the first speech moment With the first PTT moment The time difference between ; The buffer caches the compressed voice data from the filtering processor by frame, and the first recognition controller controls the switching switch to connect with the buffer, and inputs the cached compressed voice data in the buffer to the voice output end frame by frame; after all the compressed voice data cached in the buffer are output, the first recognition controller controls the switching switch to be electrically connected with the filtering processor, and inputs the de-interferenced voice data or original voice data output by the filtering processor to the voice output end.

[0015] In some embodiments, the buffer buffering the compressed speech data from the filter processor by frame comprises: dividing the original speech data into frames of length , that is, the frame length of the original voice data frame is , the corresponding frame length of the compressed voice data frame is , the frame compression duration of the two is ; To compensate for the first time difference , then after pressing the PTT button, a total of Original speech data frames, where the symbol Indicates rounding up; after these M original voice data frames, the compressed voice data frames are no longer used to output to the voice output end via the buffer, but the de-interferenced voice data or original voice data output by the filtering processor is directly input to the voice output end.

[0016] The beneficial effects of this application are as follows: This application discloses an earphone, earphone chip, and intercom method for an intercom, the earphone comprising: a microphone, a filter processor, a buffer, a switch, and a first recognition controller, as well as a PTT input terminal and a voice output terminal. The filter processor is used to perform digital-to-analog conversion on the voice signal output by the microphone to obtain raw voice data and perform filtering processing; the first recognition controller recognizes that the raw voice data appears before the PTT control signal at the PTT input terminal, controls the raw voice data to be stored in a buffer after voice compression processing, and controls the switch to be connected to the buffer. When the PTT control signal appears, the compressed voice data in the buffer is output to the voice output terminal. Based on this application, the problem of lost speech at the initial and final stages of speech due to the lack of coordination between speaking and PTT button operation in the prior art can be effectively solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of the components of the headset and intercom in one embodiment of the present application;

[0018] Figure 2 This is a schematic diagram of the internal components of a filter processor of an earphone in one embodiment of the present application;

[0019] Figure 3 is a waveform diagram of original voice data in an embodiment of the present application;

[0020] Figure 4 is a waveform diagram of compressed voice data in an embodiment of the present application;

[0021] Figure 5 This is a schematic diagram of the internal components of the first identification controller of the headset in one embodiment of the present application;

[0022] Figure 6 1 is a schematic diagram of the composition of the switch of the headset in one embodiment of the present application;

[0023] Figure 7 Schematic diagram of the composition of the earphone chip in one embodiment of the present application;

[0024] Figure 8 is a schematic diagram of a headset connector in one embodiment of the present application;

[0025] Figure 9 This is a flow chart of an intercom method on the earphone side in one embodiment of the present application;

[0026] Figure 10 This is a schematic diagram of the principle of compressing and storing original voice data and compressed voice data by data frame in one embodiment of the present application;

[0027] Figure 11 This is a flow chart of an intercom method on the earphone side in one embodiment of the present application;

[0028] Figure 12 Schematic diagram of a speech waveform with an initial speech mark in an embodiment of the present application;

[0029] Figure 13 This is a flow chart of an intercom method on the earphone side in one embodiment of the present application;

[0030] Figure 14 1 is a schematic diagram of a speech waveform with an end speech mark in an embodiment of the present application;

[0031] Figure 15 This is a circuit diagram of a second identification controller of an intercom in one embodiment of the present application;

[0032] Figure 16 This is a flow chart of an intercom method on the intercom side in one embodiment of the present application;

[0033] Figure 17 This is a flow chart of an intercom method on the intercom side in one embodiment of the present application. DETAILED DESCRIPTION

[0034] To facilitate understanding of the present application, the present application is described in more detail below with reference to the accompanying drawings and specific embodiments. The accompanying drawings provide preferred embodiments of the present application. However, the present application can be implemented in many different forms and is not limited to the embodiments described in this specification. Rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of the present application.

[0035] It should be noted that, unless otherwise defined, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the relevant listed items.

[0036] The following describes each embodiment in detail with reference to the accompanying drawings.

[0037] Figure 1 This is a schematic diagram of the composition of an embodiment of the headset and intercom based on this application. Figure 1 It can be seen that on the earphone side, it includes: a microphone 11, a buffer 12, a filter processor 13, a first recognition controller 14 and a switch 15, as well as a PTT input terminal 16 and a voice output terminal 17. The switch 15 is respectively connected to the filter processor 13, the buffer 12, the first recognition controller 14 and the voice output terminal 17. The first recognition controller 14 is also electrically connected to the filter processor 13 and the PTT input terminal 16. The filter processor 13 is also electrically connected to the microphone 11, and is used to perform digital-to-analog conversion on the voice signal output by the microphone 11 to obtain original voice data, and perform multiple filtering processes on the original voice data.

[0038] Furthermore, the first recognition controller 14 recognizes the original voice data generated by the filter processor 13, and also recognizes the PTT control signal from the PTT input terminal 16; when the first recognition controller 14 recognizes that the original voice data appears earlier than the PTT control signal, it controls the original voice data to be stored in the buffer 12 for caching after voice compression processing, and controls the switching switch 15 to be connected to the buffer 12. When it is recognized that the PTT control signal appears, the compressed voice data in the buffer 12 is output to the voice output terminal 17.

[0039] Based on the above-described earphone configuration and operating method, it can be seen that when the earphone user begins speaking before the PTT control signal is issued, the microphone 11, buffer 12, filter processor 13, and first recognition controller 14 can effectively recognize and store the input raw voice data. When the PTT control signal is detected, compressed voice data is output from the buffer. This ensures that voice signals before the PTT control signal arrives can also be collected and transmitted, avoiding the problem of voice signal "clipped" due to delayed PTT button operation.

[0040] Further, Figure 2 The internal components of the filter processor 13 are shown, including an audio amplifier unit 131, an AD sampling unit 132, a compression filter unit 133, a power detection unit 134, and preferably a de-interference filter unit 135. The audio amplifier unit 131 is used to perform distortion-free amplification on the audio signal input from the microphone to obtain a power-amplified analog audio signal, which then enters the AD sampling unit 132 for analog-to-digital conversion to obtain the original voice data.

[0041] The original speech data can be output in three ways, entering the compression filter unit 133, the power detection unit 134, and the descrambling filter unit 135 respectively. Each of these will be described below. The compression filter unit 133 primarily performs compression filtering on the original speech data to remove low-power components, primarily removing silence or whispers in speech.

[0042] Combine Figure 3 and Figure 4 The embodiment shown, wherein Figure 3 is the original speech data A1 input to the compression filter unit 133, Figure 4 The compressed voice data A2 is output from the compression filter unit 133. It can be clearly seen that the digital voice with low audio amplitude in the original voice data A1 is filtered out. Figure 4 The diagram in the middle shows that the low-amplitude audio surrounded by multiple red boxes in the original speech data A1 is removed, so that the effective speech components in time distribution are compressed and the time length of the speech is shortened, but the effective speech content is not reduced. Figure 3 and Figure 4The time scale is marked in the figure. You can see that the original speech data A1 lasts approximately 11 seconds (time scale 00:02 to 00:12), while the compressed speech data A2 lasts approximately 9 seconds, resulting in a compressed time of 2 seconds. Therefore, by properly selecting the time scales of the original speech data A1 and the compressed speech data A2, you can achieve the desired compression time. Of course, this also depends on the speaker's voice characteristics, and the appropriate compression time scale can be selected based on the speaker's voice characteristics.

[0043] Figure 2 The compression filter unit 133 in Figure 1 The buffer 12 in the memory is electrically connected to store the compressed speech data output by the compression filter unit 133 in the buffer 12.

[0044] exist Figure 2 In the example, power detection unit 134 is used to perform power detection on the raw voice data output by AD sampling unit 132 to detect whether there is a voice signal input by the intercom user, primarily to distinguish between noise, background sound, etc. Therefore, it is necessary to detect both the amplitude and duration of the raw voice data to effectively determine whether it is a voice signal. Therefore, after power detection on the raw voice data, power detection unit 134 outputs a valid recognition signal, such as a high voltage signal, if it determines that it is voice data. If it determines that it is not voice data, it outputs an invalid recognition signal, such as a low voltage signal. The low voltage signal can be used as a default. Only when the power detection is performed on the raw voice data input and valid raw voice data is identified, a valid recognition signal is output.

[0045] Figure 2 The power detection unit 134 and Figure 1 The first recognition controller 14 is electrically connected to the power detection unit 134, and the power detection unit 134 outputs a valid recognition signal or an invalid recognition signal to the first recognition controller 14, and then the first recognition controller 14 further identifies and judges. Figure 2 In the example, the de-interference filter unit 135 is used to filter background interference from the original voice data output by the AD sampling unit 132, eliminating background noise and background interference in the voice, making the voice content clearer, and outputting de-interferenced voice data. Of course, the de-interference filter unit 135 can also be used selectively, that is, only when there is obvious background noise. For example, when the headset is used at a construction site with very high background noise, the unit can be selected to perform de-interference filtering on the original voice data. In general use, the original voice data can be directly output.

[0046] like Figure 5As shown, the first recognition controller 14 includes a first MCU 141, a first voice recognition unit 142 electrically connected to the first MCU 141, a first PPT recognition unit 143, an identification generating unit 144 and a time difference calculating unit 145. Figure 2 The power detection unit 134 is electrically connected to receive the valid recognition signal or invalid recognition signal output by the power detection unit 134, and then forwarded to the first MCU141. After the first MCU141 receives the valid recognition signal from the first voice recognition unit 142, it records the first voice moment when the valid recognition signal arrives.

[0047] In addition, the power detection unit 134 may also forward the raw voice data output from the AD sampling unit 132 to the first voice recognition unit 142 without processing it, and perform voice feature recognition to identify the speaker's identity and thus decide whether to transmit the speaker's voice. In this way, the intelligence level of the headset is enhanced, and a corresponding mapping relationship can be established based on the identified speaker's identity and his / her operation and use characteristics of the intercom. For example, it can be identified whether the speaker operates the PTT button in a standard manner, whether he / she often speaks too early but presses the PTT button later, or continues to speak after the PTT button is released, and the time difference between such delayed pressing of the PTT button and premature release of the PTT button.

[0048] The first PPT recognition unit 143 is also electrically connected to the PTT input terminal 16, recognizes the PTT control signal received by the PTT input terminal 16, and sends the recognition result to the first MCU 141. After the first MCU 141 receives the valid PTT control signal from the first PPT recognition unit 143, it indicates that the PTT button has been pressed, and records the first PTT time when the valid PTT control signal arrives.

[0049] The first MCU 141 is electrically connected to the time difference calculation unit 145 and is configured to transmit both the first speech moment and the first PTT moment to the time difference calculation unit 145. This is used to determine which of the first speech moment and the first PTT moment is the earlier time. If the first PTT moment is earlier than the first speech moment, it indicates a normal intercom mode. If the first PTT moment is later than the first speech moment, it indicates an abnormal intercom mode. For abnormal intercom modes, the time difference between the first speech moment and the first PTT moment, i.e., the first time difference, must be calculated. This first time difference is used to further control the length of the aforementioned compressed time.

[0050] The first MCU 141 is electrically connected to a marker generating unit 144. After determining that the first speech moment is earlier than the first PTT moment, the marker generating unit 144 generates an initial speech marker, such as an audio signal with a single frequency f1 within the speech frequency range. This initial speech marker is added to the beginning of the original speech data as the beginning of the speech, and then output to the intercom via the speech output terminal 17. Furthermore, the marker generating unit 144 can also generate an ending speech marker, such as an audio signal with a single frequency f2 within the speech frequency range, where f2 is not equal to f1. This ending speech marker is added to the end of the original speech data as the end of the speech. The marker generating unit 144 is electrically connected to the switch 15. Based on the recognition of the first speech moment and the first PTT moment, the marker generating unit 144 outputs a switching control signal to the switch 15. Furthermore, the marker generating unit 144 can output the initial speech marker to the switch 15 as needed. Therefore, optionally, the first recognition controller 14 and the switching switch 15 are connected not only by a control line, but also by a signal line for transmitting the initial voice identifier and the end voice identifier, which is connected to the output end of the identifier generating unit 144 of the first recognition controller 14.

[0051] like Figure 6 As shown, combined with the aforementioned Figure 1 、 Figure 2 and Figure 5 The switch 15 includes a first input terminal 151, a second input terminal 152, a control terminal 153, and an output terminal 154. The first input terminal 151 is electrically connected to the buffer 12, the second input terminal 152 is electrically connected to the filtering processor 13, specifically to the de-interference filtering unit 135, the control terminal 153 is electrically connected to the first recognition controller 14, specifically to the identifier generating unit 144, and the output terminal 154 is electrically connected to the voice output terminal 17. If necessary, a third input terminal 155 can also be provided for connecting to the identifier generating unit 144 in the first recognition controller 14. When it is necessary to output an initial voice identifier or an end voice identifier, it can be switched to the third input terminal 155 under the control of the control terminal 153.

[0052] Based on the above Figure 1 、 Figure 2 、 Figure 5 、 Figure 6 The circuit in the earphones of this application can also be integrated in the form of a chip to achieve the same function. Figure 7 As shown, the earphone chip 3 includes a buffer 12, a filter processor 13, a first recognition controller 14 and a switch 15, as well as a PTT input connection line 36 and a voice output connection line 37, which correspond to Figure 1The buffer 12, the filter processor 13, the first recognition controller 14 and the switch 15, as well as the PTT input terminal 16 and the voice output terminal 17, have the same functions and internal components, which will not be described in detail here.

[0053] The pins of the earphone chip 3 include a power pin 301, a ground pin 302, a voice output pin 303, a PTT input pin 304, and a pickup input pin 305. The power pin 301 and the ground pin 302 are used to receive the positive and negative poles of direct current (DC). This DC power can come from the earphone's own power supply or from a compatible intercom via a cable. The voice output pin 303 is electrically connected to the voice output connection line 37 inside the earphone chip 3 and is used to output valid voice data and the initial voice identifier / end voice identifier to the intercom. The PTT input pin 304 is electrically connected to the PTT input connection line 36 inside the earphone chip 3 and is used to receive the PTT control signal from the intercom. The pickup input pin 305 is electrically connected to the filter processor 13 inside the earphone chip 3 and is used to connect to an external pickup and receive the analog voice signal from the pickup.

[0054] The headphone chip based on single-chip integration can realize the various functions disclosed above in this application, which is conducive to realizing the functions of the above-mentioned various control processing circuits in the miniaturization of headphones, ensuring the optimal design of multiple factors such as volume, energy consumption, and cost.

[0055] Figure 8 The headset further shows a connector for connecting to a walkie-talkie, that is, the headset includes a connector, and correspondingly, the walkie-talkie adapted to the headset includes a jack for connecting to the connector. The connector includes a voice output unit 403 located at the end, and the voice output unit 403 corresponds to Figure 1 The physical implementation of the voice output terminal 17 in the embodiment realizes the electrical connection between the intercom and the headset. And the PTT signal input part 404 (corresponding to Figure 1The physical implementation of the PTT input 16 in the walkie-talkie is shown here, as well as the power connector 402 (used to provide power from within the intercom to the headset), and the ground connector 401 (used to connect the common ground between the intercom's internal circuits and the headset's internal circuits). The order in which these functional components (i.e., 401-404) are arranged is not limited to this embodiment and can be adjusted based on design requirements. These functional components are all made of conductive metal materials, including a first insulating portion 411 provided between the voice output portion 403 and the PTT signal input portion 404, a second insulating portion 412 provided between the PTT signal input portion 404 and the power connector 402, a third insulating portion 413 provided between the power connector 402 and the ground connector 401, and a fourth insulating portion 414 provided behind the ground connector 401. These insulating portions isolate and insulate these functional components from each other.

[0056] Furthermore, cables corresponding to these functional components are located within the connector. These cables include a third cable 423 electrically connected to the voice output unit 403, a fourth cable 424 electrically connected to the PTT signal input unit 404, a second cable 422 electrically connected to the power connection unit 402, and a first cable 421 electrically connected to the ground connection unit 401. These cables are also insulated from each other, each with an insulating material exterior and conductive wires encased within. These cables are then encased in a headphone cable made of insulating material and reinforced with nylon.

[0057] The following combination Figure 9 and Figure 10 Further explanation is given on how the headset of the present application can achieve complete transmission of the voice signal when there is a voice signal before the PTT is pressed. This can be achieved in two ways, and the first implementation method is introduced first.

[0058] like Figure 9 As shown, the following steps are included:

[0059] Step S101: The first recognition controller 14 recognizes the presence of a valid recognition signal indicating the presence of a voice signal and records the first voice moment of the signal. , and the PTT control signal has not appeared at this time, the first recognition controller 14 controls the filter processor 13 to perform compression processing and send the compressed voice data to the buffer 12 for buffering.

[0060] Step S102: The first recognition controller 14 recognizes the PTT valid signal of the PTT button being pressed and records the first PTT moment of the signal. , calculate the first speech moment With the first PTT moment The time difference between .

[0061] In step S103, the buffer 12 buffers the compressed speech data from the filter processor 13 frame by frame, and the first recognition controller 14 controls the switch 15 to connect to the buffer 12, and inputs the buffered compressed speech data to the speech output terminal frame by frame.

[0062] Step S104: After all the compressed voice data buffered in the buffer 12 are output, the first recognition controller 14 controls the switch 15 to be electrically connected to the filter processor 13, and inputs the de-interferenced voice data or the original voice data output by the filter processor 13 to the voice output terminal.

[0063] In step S103, the buffer 12 also includes a frame buffer for the compressed speech data from the filter processor 13, such as Figure 10 As shown, the frame length of the original speech data is , that is, the frame length of the original voice data frame is , after such Figure 4 After compression using the method shown, the corresponding frame duration of the compressed voice data frame is , the frame compression duration of the two is Furthermore, in order to compensate for the first time difference , then after pressing the PTT button, a total of Original speech data frames, where the symbol = represents rounding up. Then, after these M original voice data frames, it is no longer necessary to use compressed voice data frames to output to the voice output terminal. Instead, the original voice data can be directly output to the voice output terminal. Correspondingly, in step S104, after these M original voice data frames, the switch 15 is controlled to switch from connecting the first input terminal 151 and the output terminal 154 to connecting the second input terminal 152 and the output terminal 154.

[0064] Here is another way to implement it, Figure 11 As shown, the following steps are included:

[0065] In step S201, the first recognition controller 14 recognizes the presence of a valid recognition signal indicating the presence of a voice signal, records the first voice moment at which the signal appears, and at this time, the PTT control signal has not yet appeared. The first recognition controller 14 controls the filter processor 13 to perform compression processing and sends the compressed voice data to the buffer 12 for caching.

[0066] In step S202, the first recognition controller 14 first outputs the initial voice identifier, which is then output to the voice output terminal (i.e. Figure 6The control switches the third input terminal 155 to be connected to the output terminal 154), and synchronously controls the self-filtering processor 13 to compress the input voice signal and store the compressed voice data in the buffer 12. After the initial voice identifier is issued, the buffer 12 is controlled to output the compressed voice data through the switching switch to the voice output terminal (i.e. Figure 6 The first input terminal 151 is connected to the output terminal 154 by controlling the switching.

[0067] Step S203: After the compressed time length of the compressed voice data is equal to the time length of the initial voice mark, the first recognition controller 14 controls the switch to switch to the filtering processor 13 (i.e. Figure 6 The control switches the second input terminal 152 to be connected to the output terminal 154), and outputs the original voice data or the de-interferenced voice data from the filtering processor 13 to the voice output terminal through the switching switch 15.

[0068] Combine Figure 12 The speech waveform shown includes two adjacent waveforms. The first waveform represents the initial speech identifier B1 (the duration can be between 0.5 seconds and 2 seconds), and the second waveform represents the compressed speech data A2. Figure 4 To illustrate, the time length of the original voice data A1 before the compressed voice data A2 is compressed minus the time length of the compressed voice data A2 must be exactly equal to the duration of the initial voice identifier B1.

[0069] based on Figure 11 and Figure 12 The embodiment shown mainly uses the "initial voice identifier + compressed voice data" method to send audio data from the headset to the intercom. As long as the intercom detects the initial voice identifier, it can be considered that a valid speech has been made through the headset and the voice signal will be sent, instead of waiting for the PTT button to be pressed before sending the voice signal. Therefore, the problem of the initial voice being "cut off" can be solved.

[0070] If the PTT button is released before the end of speaking, Figure 13 As shown, the following steps are included:

[0071] In step S301, the first recognition controller 14 recognizes that the PTT control signal changes from a valid state to an invalid state, indicating that the PTT button is released, and also recognizes that the original voice data still exists. The first recognition controller 14 continues to control the filter processor 13 to output the original voice data to the voice output terminal.

[0072] In step S302 , the first recognition controller 14 recognizes that the original voice data ends, generates an end voice mark and outputs it to the voice output terminal.

[0073] Correspondingly, Figure 12 similar, Figure 14 The diagram schematically shows two adjacent waveforms. The first waveform represents the original voice data A1, and the second waveform represents the end voice marker B2 (the duration can be between 0.5 seconds and 2 seconds). The end voice marker B2 is added after the original voice data A1 ends.

[0074] Further, combined Figure 1 As shown, the intercom side also includes a modulator 21, a second recognition controller 22 and a power amplifier 23, wherein the second recognition controller 22 is electrically connected to the modulator 21, the power amplifier 23 and the PTT input terminal 16 and the voice output terminal 17 respectively. In combination with the above content, the second recognition controller 22 can recognize the initial voice identifier B1 and the end voice identifier B2 from the voice output terminal 17, and can also recognize the valid voice data from the voice output terminal 17 (the valid voice data includes the aforementioned compressed voice data, original voice data or descrambled voice data), and recognize the PTT control signal from the PTT input terminal.

[0075] If the second recognition controller 22 recognizes that the PTT control signal is valid, the second recognition controller 22 controls the modulator 21 and the power amplifier 23 to start operating, modulating and amplifying the valid voice data from the voice output terminal 17 for output. Alternatively, if the second recognition controller 22 recognizes that the PTT control signal is not valid, but recognizes valid voice data and also recognizes that the initial voice identifier B1 is set, the second recognition controller 22 controls the modulator 21 and the power amplifier 23 to start operating, modulating and amplifying the valid voice data from the voice output terminal 17 for output.

[0076] It should be noted that the modulator 21 here primarily performs communication modulation of valid voice data, such as performing amplitude modulation, frequency modulation, or digital modulation such as BPSK, QPSK, and QAM at the intercom's operating frequency, to obtain the corresponding modulated signal. The power amplifier 23 primarily amplifies the modulated signal produced by the modulator 21 before transmitting it through the intercom's antenna. Furthermore, this application is also applicable to intercoms based on payphones. In this case, the modulator 21 and power amplifier 23 correspond to the modulation and power amplification components used in mobile communications. These two components can be integrated into a single chip, without limitation.

[0077] In addition, when the PTT control signal is recognized to be inactive from the valid state, the PTT button is released and valid voice data still exists, the valid voice data is continued to be received until the end voice mark B2 is recognized, and then the modulator 21 and the power amplifier 23 are controlled to stop working.

[0078] Further, combined Figure 15 The second recognition controller 22 includes a second MCU 221 , an identification recognition unit 222 electrically connected to the second MCU 221 , a second PPT recognition unit 223 , a second voice recognition unit 224 , and a modulation and power amplification control unit 225 .

[0079] Identification unit 222 and Figure 1 The middle voice output terminal 17 is connected to identify the initial voice identifier B1 and the end voice identifier B2 from the voice output terminal 17, respectively for two situations, namely, the valid voice data arrives earlier than the PTT control signal and the valid voice data ends later than the PTT control signal.

[0080] The second speech recognition unit 224 is also connected to Figure 1 The voice output terminal 17 is connected to the modulator 21 and the power amplifier 23 for identifying valid voice data from the voice output terminal 17 and sending the valid voice data to the modulator 21 and the power amplifier 23 only when the valid voice data exists.

[0081] The second PPT recognition unit 223 and Figure 1 The PTT input terminal 16 in the middle is connected to identify the PTT control signal. The identification unit 222 identifies the initial voice identification B1 and the end voice identification B2, and the second PPT identification unit 223 identifies the PTT control signal, which are transmitted to the second MCU221. The second MCU221 then sends a control signal to the modulation and power amplifier control unit 225 according to these identification results to further control the Figure 1 The modulator 21 in the embodiment starts working or stops working, and controls the power amplifier 23 to start working or stop working. Here, the modulation and power amplifier control unit 225 is connected to the modulator 21 and the power amplifier 23 respectively, and is used to control the modulator 21 and the power amplifier 23 to start working and stop working.

[0082] From this we can see that Figure 15 The identification unit 222 in the main is mainly for the application scenario where the initial voice identification B1 and the end voice identification B2 need to be identified and judged. Figure 9 The application scenario of the embodiment shown does not require the use of Figure 15 The identification unit 222 in.

[0083] On the intercom side, Figure 16 A method for identifying and controlling the initial voice mark B1 is further provided, comprising the following steps:

[0084] In step S401, the second recognition controller 22 recognizes that an initial voice identifier appears from the voice output terminal 17 and that the PTT control signal from the PTT input terminal 16 is not yet valid, and determines that the valid voice data from the earphone is earlier than the PTT button is pressed.

[0085] In step S402, when valid voice data appears after the second recognition controller 22 recognizes the initial voice mark, it controls the modulator 21 and the power amplifier 23 to start working, modulating and power-amplifying the valid voice data after the initial voice mark and outputting it.

[0086] In step S403, after the second recognition controller 22 recognizes that the PTT control signal from the PTT input terminal 16 is valid and valid voice data still exists, it continues to control the modulator 21 and the power amplifier 23 to maintain the working state.

[0087] Figure 17 A method for identifying and controlling the end speech mark B2 is further provided, comprising the following steps:

[0088] Step S501, when the intercom is in the speaking state, the second recognition controller 22 recognizes that the PTT control signal from the PTT input terminal 16 switches from the valid state to the invalid state, and the PTT button is released. At the same time, it also recognizes that the valid voice data from the voice output terminal 17 continues to exist, and then continues to control the modulator 21 and the power amplifier 23 to keep working.

[0089] In step S502, the second recognition controller 22 recognizes the end voice mark after the valid voice data and also recognizes that the PTT control signal is in an invalid state, and accordingly controls the modulator 21 and the power amplifier 23 to stop working.

[0090] As can be seen, the present application discloses an earphone, earphone chip, and intercom method for an intercom. The earphone includes: a microphone, a filter processor, a buffer, a switch, and a first recognition controller, as well as a PTT input terminal and a voice output terminal. The filter processor is used to perform digital-to-analog conversion on the voice signal output by the microphone to obtain raw voice data and perform filtering processing; the first recognition controller recognizes that the raw voice data appears earlier than the PTT control signal at the PTT input terminal, then controls the raw voice data to be stored in the buffer after voice compression processing, and controls the switch to be connected to the buffer. When the PTT control signal appears, the compressed voice data in the buffer is output to the voice output terminal. Based on this application, the problem of the loss of spoken voice in the initial and final stages due to the lack of coordination between speaking and PTT button operation in the prior art can be effectively solved.

[0091] The above are merely embodiments of the present application and are not intended to limit the patent scope of the present application. Any equivalent structural transformations made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A headset for a walkie-talkie, characterized in that: include: A microphone, a filter processor, a buffer, a switch, a first recognition controller, a PTT input terminal, and a voice output terminal, wherein the switch is respectively connected to the filter processor, the buffer, the first recognition controller, and the voice output terminal, the first recognition controller is also electrically connected to the filter processor and the PTT input terminal, and the filter processor is also electrically connected to the microphone and the buffer, and is used to perform analog-to-digital conversion on the voice signal output by the microphone to obtain original voice data, and to perform filtering processing on the original voice data; The first recognition controller recognizes the original voice data generated by the filter processor and also recognizes the PTT control signal from the PTT input terminal. If the first recognition controller recognizes that the original voice data appears earlier than the PTT control signal, the first recognition controller controls the original voice data to be stored in the buffer after voice compression processing, controls the switch to be connected to the buffer, and outputs the compressed voice data in the buffer to the voice output terminal when the PTT control signal appears. The first recognition controller includes a first MCU, and a first voice recognition unit, a first PTT recognition unit and a time difference calculation unit electrically connected to the first MCU; the first voice recognition unit recognizes the original voice data from the filtering processor and then forwards the recognition result to the first MCU. After the first MCU receives the valid recognition signal from the first voice recognition unit, it records the first voice time when the valid recognition signal arrives. ; The first PTT recognition unit is also electrically connected to the PTT input terminal, recognizes the PTT control signal received by the PTT input terminal, and sends the recognition result to the first MCU. After the first MCU receives the valid PTT control signal from the first PTT recognition unit, it indicates that the PTT button has been pressed, and records the first PTT time when the valid PTT control signal arrives. ; The first MCU is electrically connected to the time difference calculation unit for calculating the first voice time and the first PTT moment are sent to the time difference calculation unit for determining the time of recognizing the first speech and the first PTT moment and calculate the time difference between these two moments, namely the first time difference .

2. The earphone for intercom according to claim 1, characterized in that: The filtering processor includes an audio amplification unit, an AD sampling unit, a compression filtering unit and a power detection unit. The audio amplification unit is used to perform distortion-free amplification on the voice signal output from the microphone to obtain a power-amplified voice signal, which then enters the AD sampling unit for analog-to-digital conversion to obtain the original voice data; the compression filtering unit is used to perform compression filtering on the original voice data to remove low-power components in the original voice data to obtain the compressed voice data.

3. The earphone for intercom according to claim 2, characterized in that: The power detection unit is used to perform power detection on the original voice data output by the AD sampling unit to detect whether there is a voice signal input.

4. The earphone for intercom according to claim 2, characterized in that: The filtering processor further includes a de-interference filtering unit configured to filter out background interference from the original voice data output by the AD sampling unit and output de-interferenced voice data.

5. The earphone for intercom according to claim 1, characterized in that: The switching switch includes a first input end, a second input end, a control end and an output end, the first input end is electrically connected to the buffer, the second input end is electrically connected to the filtering processor, the control end is electrically connected to the first recognition controller, and the output end is electrically connected to the voice output end.

6. A headphone chip, characterized in that: The interior of the earphone chip includes a buffer, a filter processor, a first identification controller and a switch, as well as a PTT input connection line and a voice output connection line; The pins of the earphone chip include a power pin, a ground pin, a voice output pin, a PTT input pin, and a pickup input pin. The power pin and the ground pin are used to connect to the positive and negative poles of direct current. The voice output pin is electrically connected to the voice output connection line, the PTT input pin is electrically connected to the PTT input connection line, and the pickup input pin is electrically connected to the filter processor for connecting to a pickup. The switching switch is respectively connected to the filter processor, the buffer, the first recognition controller and the voice output connection line, the first recognition controller is also electrically connected to the filter processor and the PTT input connection line, and the filter processor is also electrically connected to the sound pickup input pin and the buffer, and is used to perform analog-to-digital conversion on the voice signal of the sound pickup input pin to obtain original voice data, and perform filtering processing on the original voice data; The first recognition controller recognizes the original voice data generated by the filter processor and also recognizes the PTT control signal from the PTT input pin; if the first recognition controller recognizes that the original voice data appears earlier than the PTT control signal, the first recognition controller controls the original voice data to be stored in the buffer after voice compression processing, and controls the switch to be connected to the buffer; when the first recognition controller recognizes that the PTT control signal appears, the first recognition controller outputs the compressed voice data in the buffer to the voice output pin; The first recognition controller includes a first MCU, and a first voice recognition unit, a first PTT recognition unit and a time difference calculation unit electrically connected to the first MCU; the first voice recognition unit recognizes the original voice data from the filtering processor and then forwards the recognition result to the first MCU. After the first MCU receives the valid recognition signal from the first voice recognition unit, it records the first voice time when the valid recognition signal arrives. ; The first PTT recognition unit is also electrically connected to the PTT input connection line, recognizes the PTT control signal connected to the PTT input pin, and sends the recognition result to the first MCU. After the first MCU receives the valid PTT control signal from the first PTT recognition unit, it indicates that the PTT button has been pressed, and records the first PTT time when the valid PTT control signal arrives. ; The first MCU is electrically connected to the time difference calculation unit for calculating the first voice time and the first PTT moment are sent to the time difference calculation unit for determining the time of recognizing the first speech and the first PTT moment and calculate the time difference between these two moments, namely the first time difference .

7. A method for intercom, characterized in that: Using the earphone for an intercom according to claim 1 to conduct intercom speaking comprises the steps of: The first recognition controller recognizes the presence of a valid recognition signal indicating the presence of a voice signal and records the first voice moment of its presence , and the PTT control signal has not yet appeared at this time, the first recognition controller controls the filter processor to perform compression processing and sends the compressed voice data to the buffer for caching; The first recognition controller recognizes the PTT valid signal of the PTT pressing and records the first PTT moment of its occurrence , calculate the first speech moment With the first PTT moment The time difference between ; The buffer buffers the compressed voice data from the filter processor frame by frame, and the first recognition controller controls the switch to connect to the buffer, and inputs the buffered compressed voice data to the voice output terminal frame by frame; After all the compressed voice data cached in the buffer is output, the first recognition controller controls the switch to be electrically connected to the filter processor, and inputs the de-interferenced voice data or original voice data output by the filter processor to the voice output end.

8. The intercom method according to claim 7, characterized in that: The buffer caches the compressed voice data from the filter processor in frames, including: The frame length of the original speech data is , that is, the frame length of the original voice data frame is , the corresponding frame length of the compressed voice data frame is , the frame compression duration of the two is ; To compensate for the first time difference , then after pressing the PTT button, a total of Original speech data frames, where the symbol Indicates rounding up; After these M original voice data frames, the compressed voice data frames are no longer output to the voice output terminal via the buffer, but the de-interferenced voice data or original voice data output by the filter processor is directly input to the voice output terminal.

Citation Information

Patent Citations

  • Wireless earphone, talkback method based on wireless earphone and talkback system

    CN113873384A