Vehicle-mounted voice system control method

By identifying and verifying the voice commands of the owner, the problem that the on-board voice system cannot distinguish the owner from others is solved, and the safety control of special operating instructions is achieved, which improves driving safety.

CN120412593APending Publication Date: 2025-08-01FORYOU GENERAL ELECTRONICS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510517346.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing car voice system cannot distinguish between the voice commands of car owners and others, resulting in the execution of non-car owners, which poses safety risks.

Method used

By determining whether the current operation command is a special type command, and identifying whether the current voice is a system-authenticated voice, if so, the command will be executed, otherwise the execution will be refused, including preprocessing, feature signal extraction, and classifier recognition.

Benefits of technology

It realizes safety control of special operating instructions and improves driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412593A_ABST
    Figure CN120412593A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle-mounted voice system control method, which comprises the following steps of: 1, judging whether a wake-up instruction is received or not, if so, entering the next step, and if not, circularly executing the step; 2, recognizing the current voice, and extracting a current operation instruction in the current voice; 3, judging whether the current operation instruction is a special type instruction or not, if so, entering the next step, and if not, directly entering the step 5; step 4, identifying whether the current voice is a system authentication voice, if so, entering step 5, otherwise, refusing to execute the current operation instruction; and 5, executing the current operation instruction. According to the invention, the safety control of the special operation instruction is realized, and the driving safety is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human - machine interaction, and particularly to a control method for an in - vehicle voice system. Background Art

[0002] With the continuous development of intelligent vehicle technology, the in - vehicle voice system has become an indispensable part of modern intelligent cockpits, enabling the interaction between drivers and vehicles to shift from simple physical operations to more complex digital interactions, thus meeting the requirements of higher - level safety and convenience.

[0003] Current in - vehicle voice systems cannot distinguish the voices of the vehicle owner and others. As a result, after others issue commands to the in - vehicle voice system, the in - vehicle voice system will still recognize and execute the corresponding commands, which may pose certain potential safety hazards to the vehicle owner in some situations. Summary of the Invention

[0004] The present invention provides a control method for an in - vehicle voice system, aiming to solve the defects in the prior art, achieve the safe control of special operation commands, and improve driving safety.

[0005] To achieve the above - mentioned purpose, the technical solution adopted by the present invention is as follows:

[0006] The present invention provides a control method for an in - vehicle voice system, including:

[0007] Step 1: Determine whether a wake - up command is received. If yes, proceed to the next step; otherwise, loop and execute this step.

[0008] Step 2: Recognize the current voice and extract the current operation command therein.

[0009] Step 3: Determine whether the current operation command is a special - type command. If yes, proceed to the next step; otherwise, directly enter Step 5.

[0010] Step 4: Identify whether the current voice is system - authenticated voice. If yes, enter Step 5; otherwise, reject the execution of the current operation command.

[0011] Step 5: Execute the current operation command.

[0012] Further, after Step 4, it further includes:

[0013] Step 41: Determine whether the confirmation information from the vehicle owner is received within a preset time. If yes, enter Step 5; otherwise, reject the execution of the current operation command.

[0014] Specifically, step 3 includes: comparing the current operation instruction with the system preset instruction. If the current operation instruction belongs to the instruction that only the vehicle owner is allowed to issue, it is determined as a special type instruction; otherwise, it is determined as a normal type instruction.

[0015] Specifically, the recognition of whether the current voice is the system authenticated voice includes:

[0016] Step 401: Read the current voice data for preprocessing to generate several frames of initial signals. The preprocessing includes pre-emphasis, framing, and windowing.

[0017] Step 402: Read the current frame of the initial signal as the current signal to be processed.

[0018] Step 403: Determine the upper envelope signal and the lower envelope signal of the current signal to be processed.

[0019] Step 404: Generate a first feature signal according to a first preset relational expression.

[0020] Step 405: Determine whether the first feature signal meets the preset conditions. If so, use the first feature signal as the current target signal and proceed to the next step; otherwise, use the first feature signal as the new current signal to be processed and return to step 403. The preset conditions are: │a - b│≤1, where a is the number of extreme points of the first feature signal and b is the number of zero-crossing points of the first feature signal; and at any time, the mean value of the upper envelope signal and the lower envelope signal is 0.

[0021] Step 406: Generate a second feature signal according to a second preset relational expression.

[0022] Step 407: Determine whether the current target signal and the second feature signal meet the preset stop conditions. If so, save the several target signals obtained by decomposing the current signal to be processed, read the next frame of the initial signal as the current frame, and return to step 402; otherwise, use the second feature signal as the new current signal to be processed and return to step 403.

[0023] Step 408: Calculate the energy of each target signal for each frame according to a third preset relational expression.

[0024] Step 409: Calculate the total energy of each frame of the signal according to a fourth preset relational expression.

[0025] Step 410: Filter the total energy of all frames through a Mel triangular filter, and calculate the logarithmically superimposed value of the energy output by the filter according to a fifth preset relational expression.

[0026] Step 411: Perform discrete cosine transform on the energy logarithm superposition value according to the sixth preset relational expression to obtain target feature coefficients;

[0027] Step 412: Input the target feature coefficients into the pre-trained classifier for recognition, and output the recognition result to determine whether there is a characteristic sound.

[0028] Specifically, the first preset relational expression is:

[0029] h(t) = x(t) - [z1(t) + z2(t)] / 2

[0030] where h(t) represents the first feature signal, x(t) represents the current signal to be processed, z1(t) represents the upper envelope signal, and z2(t) represents the lower envelope signal.

[0031] Specifically, the second preset relational expression is:

[0032] r i (t) = x(t) - c i (t)

[0033] where r i (t) represents the second feature signal, x(t) represents the current signal to be processed, and c i (t) represents the current target signal.

[0034] Specifically, the third preset relational expression is:

[0035]

[0036] The fourth preset relational expression is:

[0037]

[0038] where E(k) represents the total energy of each frame of signal, E i (k) represents the energy of each target signal, c i (t) represents each target signal, and t represents the time parameter.

[0039] Specifically, the fifth preset relational expression is:

[0040]

[0041] where s(m) represents the energy logarithm superposition value, E(k) represents the total energy of each frame of signal, H m (k) represents the frequency response function of the triangular filter, M is the number of triangular filters, m is the filter serial number, and ln represents the logarithm operation.

[0042] Specifically, the sixth preset relational expression is:

[0043]

[0044] Among them, C(n) represents the nth target feature coefficient, s(m) represents the logarithmically superimposed energy value, L represents the order of the target feature coefficients, M is the number of triangular filters, and m is the filter serial number.

[0045] The beneficial effects of the present invention are as follows: By recognizing the current voice, if the current operation instruction is a special type of instruction, then further determine whether the current voice is a system-authenticated voice. If so, execute the current operation instruction; otherwise, reject the execution of the current operation instruction, thereby realizing the secure control of special operation instructions and improving driving safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a schematic flowchart of the control method for the in-vehicle voice system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The following specifically clarifies the embodiments of the present invention in conjunction with the drawings. The drawings are only for reference and illustration, and do not constitute a limitation on the protection scope of the present invention patent.

[0048] In the processes described in the specification, claims or drawings of the present invention, which include the serial numbers of each step (such as step 10, 20, etc.), the serial numbers are only used to distinguish each step, and the serial numbers themselves do not represent any execution order. It should be noted that the descriptions such as "first", "second", etc. in this article are only used to distinguish the described objects, etc., and do not represent a sequence, nor do they indicate that "first", "second", etc. are different types.

[0049] As Figure 1 shown, this embodiment provides a control method for an in-vehicle voice system, including:

[0050] Step 1: Determine whether a wake-up instruction is received. If so, proceed to the next step; otherwise, loop and execute this step.

[0051] Generally, in-vehicle voice systems have specific wake-up instructions. For example, the wake-up instruction for Baidu's artificial intelligence assistant is "Xiaodu Xiaodu!". When Baidu's artificial intelligence assistant receives an instruction like "Xiaodu Xiaodu!", it will enter the working state from the sleep state.

[0052] Step 2: Recognize the current voice and extract the current operation instruction therein.

[0053] This step includes voice recognition and semantic recognition, which are prior arts and will not be elaborated herein.

[0054] Step 3: Determine whether the current operation instruction is a special type of instruction. If so, proceed to the next step; otherwise, directly proceed to Step 5.

[0055] In this embodiment, Step 3 includes: comparing the current operation instruction with the system preset instructions. If the current operation instruction belongs to the instructions that are only allowed to be issued by the vehicle owner, it is determined as a special type of instruction; otherwise, it is determined as a common type of instruction.

[0056] In specific implementation, it can be pre-set in the system which instructions must be issued by the vehicle owner. For example, the instructions related to the equipment on the driver's side can be set as special type of instructions, including changing the driver's window settings, adjusting the driver's air-conditioning settings, etc.

[0057] Step 4: Identify whether the current voice is the system-certified voice. If so, proceed to Step 5; otherwise, reject the execution of the current operation instruction.

[0058] Step 5: Execute the current operation instruction.

[0059] In another embodiment of the present invention, after Step 4, it further includes:

[0060] Step 41: Determine whether the confirmation information from the vehicle owner is received within the preset time. If so, proceed to Step 5; otherwise, reject the execution of the current operation instruction.

[0061] The confirmation information includes a voice confirmation instruction or other interactive confirmation information.

[0062] In another embodiment of the present invention, the identification of whether the current voice is the system-certified voice includes:

[0063] Step 401: Read the current voice data x e (t) for preprocessing to generate several frames of initial signals x o (t), and the preprocessing includes pre-emphasis, framing, and windowing.

[0064] Step 402: Read the current frame of the initial signal x o (t) as the current signal to be processed x(t).

[0065] In this embodiment, the initial signal x o (t) is read frame by frame in sequence starting from the first frame.

[0066] Step 403: Determine the upper envelope signal z1(t) and the lower envelope signal z2(t) of the current signal to be processed x(t).

[0067] Step 404: Generate the first feature signal h(t) according to the first preset relationship.

[0068] In this embodiment, the first preset relational expression is:

[0069] h(t) = x(t) - [z1(t) + z2(t)] / 2

[0070] Where h(t) represents the first characteristic signal, x(t) represents the current signal to be processed, z1(t) represents the upper envelope signal, and z2(t) represents the lower envelope signal.

[0071] Step 405: Determine whether the first characteristic signal h(t) meets the preset condition. If so, use the first characteristic signal h(t) as the current target signal c i (t) and proceed to the next step. Otherwise, use the first characteristic signal h(t) as the new current signal to be processed and return to step 403. The preset condition is: │a - b│≤ 1, where a is the number of extreme points of the first characteristic signal h(t), and b is the number of zero-crossing points of the first characteristic signal h(t); and at any time, the mean value of the upper envelope signal z1(t) and the lower envelope signal z2(t) is 0.

[0072] Step 406: Generate a second characteristic signal r i (t).

[0073] In this embodiment, the second preset relational expression is:

[0074] r i (t) = x(t) - c i (t)

[0075] Where r i (t) represents the second characteristic signal, x(t) represents the current signal to be processed, and c i (t) represents the current target signal.

[0076] Step 407: Determine whether the current target signal c i (t) and the second characteristic signal r i (t) meet the preset stop condition. If so, save several target signals c1(t), c2(t)... c n (t) obtained by decomposing the current signal to be processed x(t), read the next frame of the initial signal x o (t) as the current frame and return to step 402. Otherwise, use the second characteristic signal r i (t) as the new current signal to be processed and return to step 403.

[0077] In this embodiment, the preset stop condition is:

[0078]

[0079] Among them, T represents the signal time length.

[0080] It is easy to understand that in this embodiment, i = 1, 2, 3... n, where n is the number of times of the loop in steps 403 - 306, and the value of n is determined by the preset stop condition in step 404.

[0081] Step 408: Calculate the energy E i (k) of each target signal in each frame.

[0082] In this embodiment, the third preset relational expression is:

[0083]

[0084] Among them, E i (k) represents the energy of each target signal, c i (t) represents each target signal, and t represents the time parameter.

[0085] Step 409: Calculate the total energy E(k) of each frame of signal according to the fourth preset relational expression.

[0086] In this embodiment, the fourth preset relational expression is:

[0087]

[0088] Among them, E(k) represents the total energy of each frame of signal, and E i (k) represents the energy of each target signal.

[0089] Step 410: Filter the total energy of all frames through a Mel triangular filter, and calculate the logarithmic superposition value s(m) of the energy output by the filter according to the fifth preset relational expression.

[0090] In this embodiment, the fifth preset relational expression is:

[0091]

[0092] Among them, s(m) represents the logarithmic superposition value of energy, E(k) represents the total energy of each frame of signal, H m (k) represents the frequency response function of the triangular filter, M is the number of triangular filters, m is the serial number of the filter, and ln represents the logarithmic operation.

[0093] Step 411: Perform discrete cosine transform on the logarithmic superposition value s(m) of energy according to the sixth preset relational expression to obtain the target feature coefficient.

[0094] In this embodiment, the sixth preset relational expression is:

[0095]

[0096] Among them, C(n) represents the nth target feature coefficient, s(m) represents the logarithmic superposition value of energy, L represents the order of the target feature coefficient, M is the number of triangular filters, and m is the filter serial number.

[0097] Step 412: Input the target feature coefficient into the pre-trained classifier for recognition, and output the recognition result to determine whether there is a characteristic sound.

[0098] The above-disclosed are only the preferred embodiments of the present invention, and the scope of the rights protection of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the scope of the patent application of the present invention still fall within the scope covered by the present invention.

Claims

1. A control method for an in-vehicle voice system, characterized in that, Including: Step 1: Determine whether a wake-up instruction is received. If yes, proceed to the next step; otherwise, loop and execute this step. Step 2: Recognize the current voice and extract the current operation instruction therein. Step 3: Determine whether the current operation instruction is a special type of instruction. If yes, proceed to the next step; otherwise, directly proceed to Step 5. Step 4: Identify whether the current voice is a system authentication voice. If yes, proceed to Step 5; otherwise, reject the execution of the current operation instruction. Step 5: Execute the current operation instruction.

2. The vehicle-mounted voice system control method according to claim 1, characterized in that, After Step 4, it further includes: Step 41: Determine whether the confirmation information from the vehicle owner is received within a preset time. If yes, proceed to Step 5; otherwise, reject the execution of the current operation instruction.

3. The vehicle-mounted voice system control method according to claim 1, wherein, Step 3 includes: Comparing the current operation instruction with the system preset instructions. If the current operation instruction belongs to the instructions only allowed to be issued by the vehicle owner, it is determined as a special type of instruction; otherwise, it is determined as a common type of instruction.

4. The characteristic sound recognition method according to claim 1, wherein Identifying whether the current voice is a system authentication voice includes: Step 401: Read the current voice data for preprocessing to generate several frames of initial signals. The preprocessing includes pre-emphasis, framing, and windowing. Step 402: Read the current frame of the initial signal as the current signal to be processed. Step 403: Determine the upper envelope signal and the lower envelope signal of the current signal to be processed. Step 404: Generate a first feature signal according to the first preset relational expression. Step 405: Determine whether the first feature signal meets the preset conditions. If yes, use the first feature signal as the current target signal and proceed to the next step; otherwise, use the first feature signal as the new current signal to be processed and return to Step 403. The preset conditions are: │a - b│≤1, where a is the number of extreme points of the first feature signal, and b is the number of zero-crossing points of the first feature signal; and at any time, the mean values of the upper envelope signal and the lower envelope signal are 0. Step 406: Generate a second feature signal according to the second preset relational expression. Step 407: Determine whether the current target signal and the second feature signal meet the preset stop conditions. If yes, save the several target signals obtained by decomposing the current signal to be processed, read the next frame of the initial signal as the current frame, and return to Step 402; otherwise, use the second feature signal as the new current signal to be processed and return to Step 403. Step 408: Calculate the energy of each target signal per frame according to the third preset relational expression. Step 409: Calculate the total energy of each frame of the signal according to the fourth preset relational expression. Step 410: Filter the total energy of all frames through a Mel triangular filter and calculate the logarithmic sum value of the energy output by the filter according to the fifth preset relational expression. Step 411: Perform discrete cosine transform on the logarithmic sum value of the energy according to the sixth preset relational expression to obtain the target feature coefficients. Step 412: Input the target feature coefficients into the pre-trained classifier for recognition and output the recognition result to determine whether there is a characteristic sound.

5. The characteristic sound recognition method according to claim 4, wherein The first preset relational expression is: h(t) = x(t) - [z1(t) + z2(t)] / 2 Among them, h(t) represents the first characteristic signal, x(t) represents the current signal to be processed, z1(t) represents the upper envelope signal, and z2(t) represents the lower envelope signal.

6. The characteristic sound recognition method according to claim 5, wherein The second preset relational expression is: r i r(t) = x(t) - c i (t) where r i (t) represents the second characteristic signal, x(t) represents the current signal to be processed, and c i (t) represents the current target signal.

7. The characteristic sound recognition method according to claim 6, wherein The third preset relational expression is: The fourth preset relational expression is: Among them, E(k) represents the total energy of each frame signal, and E i (k) represents the energy of each target signal, and c i (t) represents each target signal, and t represents the time parameter.

8. The characteristic sound recognition method according to claim 7, wherein The fifth preset relational expression is: Among them, s(m) represents the logarithmic superposition value of energy, E(k) represents the total energy of each frame of signal, H m (k) represents the frequency response function of the triangular filter, M is the number of triangular filters, m is the filter serial number, and ln represents the logarithmic operation.

9. The characteristic sound recognition method according to claim 8, wherein The sixth preset relational expression is: Among them, C(n) represents the nth target characteristic coefficient, s(m) represents the energy logarithm superposition value, L represents the order of the target characteristic coefficient, M is the number of triangular filters, and m is the filter serial number.

Citation Information

Patent Citations

  • Vehicle-mounted control device with voice recognition function

    CN108520747A

  • Identity authentication method based on voice recognition and medium

    CN115775561A

  • Voice instruction recognition method and device applied to vehicle

    CN116612754A

  • Vehicle-mounted system protection method and device

    CN117912461A

  • Empirical mode decomposition for analyzing acoustical signals

    US20030033094A1