Vehicle sound zone control method and device, storage medium, and electronic device

By analyzing the voice signals in the vehicle cabin and controlling the separation of voice zones, the problems of multi-voice zone interference and privacy leakage are solved, realizing a more intelligent and secure voice interaction system and enhancing the personalized services of the vehicle cabin.

CN116645964BActive Publication Date: 2026-04-14CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-29
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, voice interaction systems in vehicle cabins suffer from interference and privacy leaks when multiple voice zones operate the same task simultaneously, and the lack of service for voice zones without passengers results in poor flexibility.

Method used

By collecting voice signals from inside the vehicle cabin, converting them into text information, and parsing the control intent category, the system controls the opening or closing of the audio zone according to the intent category, thereby achieving audio zone separation and allowing only designated audio zones to connect to the voice interaction service.

Benefits of technology

It enhances the personalization and recognition flexibility of vehicle voice interaction, protects user privacy, and provides a smarter, more enjoyable, and safer driving experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116645964B_ABST
    Figure CN116645964B_ABST
Patent Text Reader

Abstract

The application provides a vehicle sound area control method and device, a storage medium and an electronic device, and belongs to the field of vehicle control, wherein the method comprises the following steps: collecting a first voice signal in a vehicle cabin of a target vehicle, wherein the vehicle cabin comprises multiple sound areas, and each sound area corresponds to a physical space of the vehicle cabin; converting the first voice signal into text information; analyzing a control intention category of the text information, wherein the control intention category is used to indicate the intention clarity of a control command corresponding to the text information; and performing sound area control on the target vehicle according to the control intention category. Through the embodiment of the application, the technical problem that a related art vehicle voice interaction is easily disturbed by other sound areas is solved, the individuality of vehicle voice interaction and the recognition flexibility of voice instructions are improved, the user privacy is protected, and thus a more intelligent, more pleasant and safer driving experience is brought to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle control, and more specifically, to a method and apparatus for controlling vehicle audio zones, a storage medium, and an electronic device. Background Technology

[0002] In related technologies, in-cabin voice interaction services allow passengers to navigate, order goods / services, and control the vehicle via voice commands. Voice zone separation technology divides the cabin space into several zones, allowing independent voice interaction for passengers in different zones. However, in practical applications, there are situations where only one or more designated zones are allowed for voice interaction, such as setting navigation destinations, making payments, or accessing private information. If voice commands from multiple zones simultaneously perform the same task, it can cause interference or privacy leaks. Furthermore, voice interaction services are not needed for zones without passengers, resulting in poor flexibility in vehicle voice interaction.

[0003] No efficient and accurate solution has yet been found to address the aforementioned issues in the relevant technologies. Summary of the Invention

[0004] This invention provides a method and apparatus for controlling vehicle audio zones, a storage medium, and an electronic device to solve technical problems in related technologies.

[0005] According to an embodiment of the present invention, a method for controlling vehicle audio zones is provided, comprising: acquiring a first voice signal within the vehicle cabin of a target vehicle, wherein the vehicle cabin includes multiple audio zones, each audio zone corresponding to a physical space within the vehicle cabin; converting the first voice signal into text information; parsing the control intent category of the text information, wherein the control intent category is used to indicate the clarity of the intent of the control command corresponding to the text information; and performing audio zone control on the target vehicle according to the control intent category.

[0006] Furthermore, parsing the control intent category of the text information includes: determining whether the text format of the text information is a preset explicit format; if the text format of the text information is a preset explicit format, determining that the text information is an explicit register control command; if the text format of the text information is not a preset explicit format, determining whether the text format of the text information is an implicit format; if the text format of the text information is an implicit format, determining that the text information is an implicit register control command; if the text format of the text information is not an implicit format, determining that the text information is not a register control command.

[0007] Furthermore, determining whether the text format of the text information is implicit includes: determining whether the text format of the text information is a preset implicit format; and / or, determining whether the text format of the text information is implicit using a machine learning model; and / or, determining whether the text format of the text information is implicit based on a deep neural network model.

[0008] Furthermore, controlling the audio zone of the target vehicle according to the control intent category includes: identifying the target audio zone and target operation in the text information according to the control intent category, wherein the target operation includes an on operation or an off operation; and performing the target operation on the target audio zone.

[0009] Furthermore, performing the target operation on the target audio region includes: if the target operation is an enabling operation, connecting the data channel between the target audio region and the voice interaction service; if the target operation is a disabling operation, disconnecting the data channel between the target audio region and the voice interaction service.

[0010] Furthermore, performing the target operation on the target sound region includes: if the target operation is an on operation, setting the position information and sound source enhancement control information corresponding to the physical space of the target sound region in the sound region separation module; if the target operation is a off operation, setting the position information and sound source suppression control information corresponding to the physical space of the target sound region in the sound region separation module.

[0011] Furthermore, after performing audio zone control on the target vehicle according to the control intent category, the method further includes: acquiring a second voice signal inside the vehicle cabin of the target vehicle; performing audio zone separation on the second voice signal to obtain first audio data from a first audio zone and second audio data from a second audio zone; determining that the first audio zone in the vehicle cabin is in an open state and the second audio zone in a closed state, outputting the first audio data, and disabling the output of the second audio data.

[0012] Furthermore, after performing voice zone control on the target vehicle according to the control intent category, the method further includes: acquiring a third voice signal in the vehicle cabin of the target vehicle; performing voice zone separation on the third voice signal to obtain third audio data from the third voice zone and fourth audio data from the fourth voice zone; determining that the third voice zone in the vehicle cabin is in an open state and the fourth voice zone is in a closed state, inputting the third audio data into the voice interaction service of the target vehicle, and prohibiting the input of the fourth audio data into the voice interaction service.

[0013] According to another embodiment of the present invention, a vehicle audio zone control device is provided, comprising: a first acquisition module for acquiring a first voice signal within the vehicle cabin of a target vehicle, wherein the vehicle cabin includes multiple audio zones, each audio zone corresponding to a physical space within the vehicle cabin; a conversion module for converting the first voice signal into text information; a parsing module for parsing the control intent category of the text information, wherein the control intent category is used to indicate the clarity of the intent of the control command corresponding to the text information; and a first control module for performing audio zone control on the target vehicle according to the control intent category.

[0014] Furthermore, the parsing module includes: a judgment unit, used to judge whether the text format of the text information is a preset explicit format; a first processing unit, used to determine that the text information is an explicit register control command if the text format of the text information is a preset explicit format; and to judge whether the text format of the text information is an implicit format if the text format of the text information is not a preset explicit format; and a second processing unit, used to determine that the text information is an implicit register control command if the text format of the text information is an implicit format; and to determine that the text information is not a register control command if the text format of the text information is not an implicit format.

[0015] Furthermore, the first processing unit includes: a first judgment subunit, used to judge whether the text format of the text information is a preset implicit format; and / or, a second judgment subunit, used to judge whether the text format of the text information is an implicit format through a machine learning model; and / or, a third judgment subunit, used to judge whether the text format of the text information is an implicit format based on a deep neural network model.

[0016] Furthermore, the first control module includes: a recognition unit, configured to recognize the target audio region and target operation in the text information according to the control intent category, wherein the target operation includes an on operation or an off operation; and an execution unit, configured to execute the target operation on the target audio region.

[0017] Furthermore, the execution unit includes: a connection subunit, configured to connect the data channel between the target audio region and the voice interaction service if the target operation is an open operation; and a disconnect subunit, configured to disconnect the data channel between the target audio region and the voice interaction service if the target operation is a close operation.

[0018] Furthermore, the execution unit includes: an enhancement subunit, configured to, if the target operation is an enable operation, set the position information and sound source enhancement control information corresponding to the physical space of the target sound region in the sound region separation module; and a suppression subunit, configured to, if the target operation is a disable operation, set the position information and sound source suppression control information corresponding to the physical space of the target sound region in the sound region separation module.

[0019] Furthermore, the device further includes: a second acquisition module, used to acquire a second voice signal inside the vehicle cabin of the target vehicle after the first control module performs voice zone control on the target vehicle according to the control intention category; a first separation module, used to perform voice zone separation on the second voice signal to obtain first audio data from the first voice zone and second audio data from the second voice zone; and a second control module, used to determine whether the first voice zone is in an open state and the second voice zone is in a closed state inside the vehicle cabin, output the first audio data, and prohibit the output of the second audio data.

[0020] Furthermore, the device further includes: a third acquisition module, used to acquire a third voice signal inside the vehicle cabin of the target vehicle after the first control module performs voice zone control on the target vehicle according to the control intention category; a second separation module, used to perform voice zone separation on the third voice signal to obtain third audio data from the third voice zone and fourth audio data from the fourth voice zone; and a third control module, used to determine that the third voice zone is in an open state and the fourth voice zone is in a closed state inside the vehicle cabin, input the third audio data to the voice interaction service of the target vehicle, and prohibit inputting the fourth audio data to the voice interaction service.

[0021] According to another aspect of the embodiments of this application, a storage medium is also provided, the storage medium including a stored program that executes the above steps when the program is run.

[0022] According to another aspect of the embodiments of this application, an electronic device is also provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; wherein: the memory is used to store computer programs; and the processor is used to execute the steps in the above method by running the programs stored in the memory.

[0023] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the above-described method.

[0024] This invention collects a first voice signal from the vehicle cabin of a target vehicle. The vehicle cabin includes multiple sound zones, each corresponding to a physical space within the cabin. The first voice signal is converted into text information, and the control intent category of the text information is parsed. The control intent category indicates the clarity of the control command corresponding to the text information. Sound zone control is performed on the target vehicle according to the control intent category. By parsing the control intent category of the text information and performing sound zone control on the target vehicle according to the control intent category, vehicle voice interaction based on sound zones is realized. This allows for control of the vehicle's voice interaction service according to sound zones, solving the technical problem of interference from other sound zones during vehicle voice interaction in related technologies. It improves the personalization of in-vehicle voice interaction and the flexibility of voice command recognition, protects user privacy, and ultimately brings users a more intelligent, enjoyable, and safer driving experience. Attached Figure Description

[0025] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0026] Figure 1 This is a hardware structure block diagram of an in-vehicle terminal according to an embodiment of the present invention;

[0027] Figure 2 This is a flowchart of a vehicle audio zone control method according to an embodiment of the present invention;

[0028] Figure 3 This is a flowchart illustrating the classification of speech-result text according to an embodiment of the present invention;

[0029] Figure 4 This is a flowchart of the voice interaction process before the sound zone control in an embodiment of the present invention;

[0030] Figure 5 This is a flowchart of the voice interaction process after the sound zone control in an embodiment of the present invention;

[0031] Figure 6 This is a flowchart illustrating an embodiment of the present invention;

[0032] Figure 7 This is a structural block diagram of a vehicle audio zone control device according to an embodiment of the present invention. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present application. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present application can be combined with each other.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] Example 1

[0036] The method embodiment provided in Embodiment 1 of this application can be executed in an in-vehicle terminal, in-vehicle control module, voice control module, or similar processing device. Taking its operation on an in-vehicle terminal as an example, Figure 1 This is a hardware structure block diagram of a vehicle-mounted terminal according to an embodiment of the present invention. Figure 1 As shown, the vehicle-mounted terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. Optionally, the vehicle-mounted terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned vehicle-mounted terminal. For example, the vehicle-mounted terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0037] The memory 104 can be used to store vehicle terminal programs, such as application software programs and modules, like the vehicle terminal program corresponding to a vehicle audio zone control method in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the vehicle terminal program stored in the memory 104, thereby implementing the aforementioned method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the vehicle terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0038] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the vehicle terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0039] This embodiment provides a method for controlling the sound zone of a vehicle. Figure 2 This is a flowchart of a vehicle audio zone control method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps:

[0040] Step S202: Collect the first voice signal inside the vehicle cabin of the target vehicle, wherein the vehicle cabin includes multiple sound zones, and each sound zone corresponds to a physical space in the vehicle cabin.

[0041] The sound zones in this embodiment can be divided according to factors such as vehicle space, type, number of seats, and seat layout. For example, each sound zone corresponds to one seat in the vehicle, with the longitudinal space where the seat is located forming one sound zone. For instance, a 5-seat vehicle can be divided into 5 sound zones. Alternatively, the seating arrangement can be arranged so that the front and rear rows each form a sound zone, such as a vehicle with two rows of seats being divided into 2 sound zones. Of course, the sound zones can also be divided according to the layout of the microphones and the layout of the in-vehicle human-machine interaction controls. For example, if the vehicle has 4 sets of microphones in 4 locations, then the vehicle can be divided into 4 sound zones, with the physical space of each sound zone including the location of the corresponding microphone.

[0042] In this embodiment, the first speech signal can be audio data after undergoing a region separation process, or it can be the original audio data without undergoing a region separation process. Speech recognition of the input audio can be performed by a speech recognizer on an in-vehicle mobile terminal, or by a speech recognizer on the server side of the vehicle network. The region separation in this embodiment can be a region separation process based on acoustic signal processing algorithms, or a region separation process based on neural network algorithms. By locating the sound source and determining the region from which the speech signal is emitted, region separation is performed on the original audio data, separating the mixed original audio data into multiple sub-audio data from different regions, each sub-audio corresponding to one region.

[0043] Step S204: Convert the first speech signal into text information;

[0044] Step S206: parse the control intent category of the text information, wherein the control intent category is used to indicate the clarity of the intent of the control command corresponding to the text information;

[0045] Optionally, based on the clarity of intent, text information corresponding to control commands can be divided into explicit register control commands and implicit register control commands.

[0046] Step S208: Perform sound zone control on the target vehicle according to the control intent category.

[0047] Optionally, when controlling the sound zones, you can either turn a sound zone on or off.

[0048] Through the above steps, the first voice signal inside the vehicle cabin of the target vehicle is collected. The vehicle cabin includes multiple sound zones, each corresponding to a physical space within the vehicle cabin. The first voice signal is converted into text information, and the control intent category of the text information is parsed. The control intent category is used to indicate the clarity of the control command corresponding to the text information. Sound zone control of the target vehicle is performed according to the control intent category. By parsing the control intent category of the text information and performing sound zone control of the target vehicle according to the control intent category, vehicle voice interaction based on sound zones is realized. The vehicle's voice interaction service can be controlled according to sound zones, solving the technical problem that vehicle voice interaction is easily interfered with by other sound zones. This improves the personalization of in-vehicle voice interaction and the flexibility of voice command recognition, protects user privacy, and ultimately brings users a more intelligent, enjoyable, and safer driving experience.

[0049] In one embodiment of this example, the control intent categories for parsing text information include:

[0050] S11, Determine whether the text format of the text information is a preset explicit format;

[0051] The preset explicit format in this embodiment can be defined by key fields, which can be any text field such as "audio range", "open", "on", "off", etc.

[0052] S12, if the text format of the text information is a preset explicit format, determine that the text information is an explicit audio control command; if the text format of the text information is not a preset explicit format, determine whether the text format of the text information is an implicit format.

[0053] In this embodiment, "explicit" and "implicit" refer to the degree to which the instruction executor understands the machine language instructions. "Explicit" means the instruction can be clearly and explicitly implemented. In the context of an interface, it means the implementation of the interface is clearly and explicitly specified. For other logic, "explicit" means the implementation content is clearly and explicitly specified. "Implicit" refers to an implicit implementation. In the context of an interface, as long as the method signature and return value of the implementing class are consistent with the interface definition, it is considered an implementation of the interface. There is no explicit (clear and explicit) specification, but further explicit specification is possible.

[0054] In some examples, explicit register control commands include: turn on register 2, turn off register 3, turn on registers 3 and 4, etc.

[0055] In some examples, implicit audio zone control commands include: "Open passenger-side audio zone"; "Close rear-seat audio zone"; "No one in the rear seats, enter safety mode," etc. Specifically, "Open passenger-side audio zone" means opening one or more audio zones corresponding to the passenger seat; "Close rear-seat audio zone" means closing one or more audio zones corresponding to the rear seats; "No one in the rear seats" means closing one or more audio zones corresponding to the rear seats; and "Enter safety mode" means closing all audio zones except the driver's seat, allowing only the driver to interact via voice, preventing passengers in other audio zones from accessing the owner's privacy through voice interaction.

[0056] In one example, determining whether the text format of the text information is implicit includes: determining whether the text format of the text information is a preset implicit format; and / or, determining whether the text format of the text information is implicit using a machine learning model; and / or, determining whether the text format of the text information is implicit based on a deep neural network model.

[0057] If the text information contains specific implicit fields, it is determined to be a preset implicit format. For non-preset implicit formats, an algorithmic model determines whether it is an implicit audio region control command. This can be represented as follows: the text content is input into the algorithmic model, the algorithmic model performs calculations, and the algorithmic model outputs the judgment result. By performing semantic parsing and mapping from implicit to explicit formats on the text information, if the parsing or mapping is successful, it is considered an implicit format.

[0058] S13, if the text format of the text information is implicit, determine that the text information is an implicit register control command; if the text format of the text information is not implicit, determine that the text information is not a register control command.

[0059] Figure 3 This is a flowchart illustrating the classification of speech result text according to an embodiment of the present invention, including: S21 inputting the recognition result text, S23 determining whether it is an explicit register control command, S25 determining whether it is an implicit register control command, and S26 finally executing the register control operation. When classifying the speech recognition result, determining whether it is an explicit register control command can be implemented by determining whether it is a preset explicit register control command. The determination result is one of two types: it is an explicit register control command; it is not an explicit register control command. Determining whether it is an implicit register control command can be implemented by determining whether it is a preset implicit register control command, which can be done through an algorithm model. The specific implementation of determining whether it is a preset implicit register control command can be the same as the implementation process of determining whether it is an explicit register control command, i.e., determining whether it is a preset implicit register control command. Specific implementations of determining whether it is an implicit register control command through an algorithm model can include using a machine learning model or a deep neural network model, etc. The judgment result is either one of two: it is an implicit register control command; or it is not a register control command.

[0060] In one example, controlling the audio zone of a target vehicle based on the control intent category includes: identifying the target audio zone and target operation in the text information based on the control intent category, wherein the target operation includes an on or off operation; and performing the target operation on the target audio zone.

[0061] In this embodiment, in addition to controlling the opening and closing of the audio zones, the state maintenance time of the audio zones can also be configured, such as closing for one hour or temporarily closing (the default closing time is a preset time, such as 3 minutes). The priority of each audio zone can also be configured, such as setting a primary audio zone, a secondary audio zone, etc. For example, if the priority of audio zone A is configured to be higher than that of audio zone B, then if the same control command (such as controlling the sunroof) is simultaneously acquired from audio zone A and audio zone B, the command from the higher-priority audio zone can be selected according to the priority of the audio zones and output to the vehicle's voice interaction service for execution. Alternatively, when noise reduction is performed on the acquired mixed speech during multi-person speech, audio from the lower-priority audio zone is filtered out first, while audio from the higher-priority audio zone is retained. Or, the acquisition sensitivity of the microphone in the higher-priority audio zone is increased, while the acquisition sensitivity of the microphone in the lower-priority audio zone is decreased, thereby improving the intelligence and flexibility of the multi-audio zone voice control process and enhancing the user experience.

[0062] Zone control operations include controlling one or more zones. Zone control comprises two control elements: the target zone to be controlled, such as zone 1, zones 1 and 3, etc.; and the specific control operation, such as opening a zone, closing a zone, and prioritizing a zone. A zone control operation can be expressed as opening or closing the nth zone, where n <= N, and N represents the total number of zones in the cockpit.

[0063] Optionally, performing a target operation on the target audio region includes: if the target operation is an enable operation, connecting the data channel between the target audio region and the voice interaction service; if the target operation is a disable operation, disconnecting the data channel between the target audio region and the voice interaction service.

[0064] Opening a specific audio zone involves connecting the separated audio output from that zone to the voice interaction service, allowing users within that zone to access the in-cabin voice interaction service. Alternatively, opening a specific audio zone can be achieved by setting location and control information corresponding to the physical space of that zone in the zone separation module. The zone separation module then performs sound source separation or enhancement on the corresponding physical space, allowing users within that zone to access the in-cabin voice interaction service. Closing a specific audio zone involves disconnecting the separated audio output from the voice interaction service, preventing users within that zone from accessing the in-cabin voice interaction service. Closing a specific audio zone can also be achieved by setting location and control information corresponding to the physical space of that zone in the zone separation module. The zone separation module then performs sound source suppression on the corresponding physical space, preventing users within that zone from accessing the in-cabin voice interaction service.

[0065] In this embodiment, the vehicle's cockpit voice interaction system can consist of a sound zone separation service and a voice interaction service. The sound zone separation service converts input audio data from multiple microphone channels into audio data within each sound zone; the converted audio data within each sound zone is connected to the voice interaction service through corresponding data channels, enabling voice interaction between users within each sound zone. The sound zone separation service can be based on acoustic signal processing algorithms or neural network algorithms, etc., and the voice interaction service can be offline or online.

[0066] In one implementation scenario of this embodiment, the first audio zone is an open audio zone, and the second audio zone is a closed audio zone. After controlling the audio zones of the target vehicle according to the control intent category, the method further includes: acquiring a second voice signal inside the vehicle cabin of the target vehicle; performing audio zone separation on the second voice signal to obtain first audio data from the first audio zone and second audio data from the second audio zone; determining that the first audio zone inside the vehicle cabin is open and the second audio zone is closed, outputting the first audio data, and prohibiting the output of the second audio data.

[0067] In another implementation scenario of this embodiment, after controlling the target vehicle's audio zones according to the control intent category, the method further includes: acquiring a third voice signal from the target vehicle's cabin; performing audio zone separation on the third voice signal to obtain third audio data from the third audio zone and fourth audio data from the fourth audio zone; determining that the third audio zone in the vehicle cabin is in an open state and the fourth audio zone is in a closed state; inputting the third audio data into the target vehicle's voice interaction service and prohibiting the input of the fourth audio data into the voice interaction service.

[0068] Figure 4 This is a flowchart of the voice interaction process before voice zone control in an embodiment of the present invention. Figure 5 This is a flowchart of the voice interaction process after zone control in an embodiment of the present invention. The vehicle includes N zones, namely zone 1, zone 2, ... zone N. The voice interaction service includes N voice interaction threads, each corresponding to the audio data of the N zones. Before zone control, all zones are connected to the vehicle's voice interaction service. The zone separation service separates the audio data from all voice acquisition channels (channel 1 to channel M) to obtain N audio data (audio data of zone 1 to audio data of zone N), and transmits the N audio data to the corresponding threads of the interaction service. After the zone control is applied, zone 2 is set to the off state, while the states of other zones remain unchanged. The zone separation service separates the audio data from all voice acquisition channels (channel 1 to channel M) to obtain N audio data (audio data from zone 1 to audio data from zone N). The N-1 audio data (audio data from zone 1, zone 3 to zone N) are then transmitted to the corresponding threads of the interactive service. The audio data from zone 2 cannot be transmitted to the corresponding voice interaction thread. This can be either the zone separation service not outputting the audio data from zone 2, or the zone separation service outputting the audio data from zone 2, but the data channel between the zone separation service and the voice interaction service does not input the audio data from zone 2 into the voice interaction service.

[0069] Optionally, the vehicle's voice interaction service in this embodiment can generate matching vehicle control commands based on the user's input audio data, which can be used to control vehicle components such as the power system, vehicle system, and audio-visual system, such as controlling vehicle start and stop, adjusting air conditioning, raising and lowering windows, opening and closing doors, adjusting speakers, braking, acceleration, shifting gears, and changing lanes.

[0070] Figure 6 This is a flowchart of an embodiment of the present invention, including: S61, audio input; S62, speech recognition of the input audio; S63, classification of the speech recognition result text; S64, control of the audio region based on the classification result. Figure 6 As shown, this embodiment provides a method for controlling the audio zones within a cockpit. This method allows occupants to control the state of one or more audio zones using voice. In implementation, the occupant inputs voice data; the input voice is then subjected to speech recognition to obtain recognized text; the recognized text is categorized; and based on the categorization results, control of one or more audio zones is achieved.

[0071] The audio zone control method in this embodiment can improve the controllability of audio zones in the cockpit and enhance the quality of voice interaction within the cockpit. After a person in the cockpit inputs audio data, the speech recognition process converts the spoken content in the audio data into text. The content classification process categorizes the text result of speech recognition into explicit audio zone control commands, implicit audio zone control commands, or non-audio zone control commands. When text is input into the classification process, it first determines whether it is an explicit audio zone control command. If so, the corresponding audio zone control operation is executed; otherwise, it determines whether it is an implicit audio zone control command. For the result of determining whether it is an implicit audio zone control command, if so, the corresponding operation is executed; otherwise, it indicates a non-audio zone control command, and no audio zone control operation is performed. The audio zone control process involves opening or closing the corresponding audio zone according to the explicit or implicit audio zone control command.

[0072] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0073] Example 2

[0074] This embodiment also provides a vehicle audio zone control device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0075] Figure 7 This is a structural block diagram of a vehicle audio zone control device according to an embodiment of the present invention, such as... Figure 7 As shown, the device includes:

[0076] The first acquisition module 70 is used to acquire the first voice signal inside the vehicle cabin of the target vehicle, wherein the vehicle cabin includes multiple sound zones, and each sound zone corresponds to a physical space of the vehicle cabin.

[0077] Conversion module 72 is used to convert the first voice signal into text information;

[0078] The parsing module 74 is used to parse the control intent category of the text information, wherein the control intent category is used to indicate the clarity of the intent of the control command corresponding to the text information;

[0079] The first control module 76 is used to perform audio zone control on the target vehicle according to the control intent category.

[0080] Optionally, the parsing module includes: a judgment unit, configured to judge whether the text format of the text information is a preset explicit format; a first processing unit, configured to determine that the text information is an explicit register control command if the text format of the text information is a preset explicit format; and to judge whether the text format of the text information is an implicit format if the text format of the text information is not a preset explicit format; and a second processing unit, configured to determine that the text information is an implicit register control command if the text format of the text information is an implicit format; and to determine that the text information is not a register control command if the text format of the text information is not an implicit format.

[0081] Optionally, the first processing unit includes: a first judgment subunit, used to judge whether the text format of the text information is a preset implicit format; and / or, a second judgment subunit, used to judge whether the text format of the text information is an implicit format through a machine learning model; and / or, a third judgment subunit, used to judge whether the text format of the text information is an implicit format based on a deep neural network model.

[0082] Optionally, the first control module includes: a recognition unit, configured to recognize the target audio region and target operation in the text information according to the control intent category, wherein the target operation includes an on operation or an off operation; and an execution unit, configured to execute the target operation on the target audio region.

[0083] Optionally, the execution unit includes: a connection subunit, configured to connect the data channel between the target audio region and the voice interaction service if the target operation is an open operation; and a disconnect subunit, configured to disconnect the data channel between the target audio region and the voice interaction service if the target operation is a close operation.

[0084] Optionally, the execution unit includes: an enhancement subunit, configured to, if the target operation is an enable operation, set the position information and sound source enhancement control information corresponding to the physical space of the target sound region in the sound region separation module; and a suppression subunit, configured to, if the target operation is a disable operation, set the position information and sound source suppression control information corresponding to the physical space of the target sound region in the sound region separation module.

[0085] Optionally, the device further includes: a second acquisition module, configured to acquire a second voice signal inside the vehicle cabin of the target vehicle after the first control module performs voice zone control on the target vehicle according to the control intention category; a first separation module, configured to perform voice zone separation on the second voice signal to obtain first audio data from the first voice zone and second audio data from the second voice zone; and a second control module, configured to determine that the first voice zone in the vehicle cabin is in an open state and the second voice zone in a closed state, output the first audio data, and prohibit the output of the second audio data.

[0086] Optionally, the device further includes: a third acquisition module, configured to acquire a third voice signal within the vehicle cabin of the target vehicle after the first control module performs voice zone control on the target vehicle according to the control intent category; a second separation module, configured to perform voice zone separation on the third voice signal to obtain third audio data from the third voice zone and fourth audio data from the fourth voice zone; and a third control module, configured to determine whether the third voice zone is in an open state and the fourth voice zone is in a closed state within the vehicle cabin, input the third audio data to the voice interaction service of the target vehicle, and prohibit inputting the fourth audio data to the voice interaction service.

[0087] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.

[0088] Example 3

[0089] Embodiments of the present invention also provide a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.

[0090] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:

[0091] S1, Collect the first voice signal inside the vehicle cabin of the target vehicle, wherein the vehicle cabin includes multiple sound zones, and each sound zone corresponds to a physical space of the vehicle cabin;

[0092] S2, convert the first voice signal into text information;

[0093] S3, parse the control intent category of the text information, wherein the control intent category is used to indicate the clarity of the intent of the control command corresponding to the text information;

[0094] S4, Perform audio zone control on the target vehicle according to the control intent category.

[0095] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0096] Embodiments of the present invention also provide an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.

[0097] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0098] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0099] S1, Collect the first voice signal inside the vehicle cabin of the target vehicle, wherein the vehicle cabin includes multiple sound zones, and each sound zone corresponds to a physical space of the vehicle cabin;

[0100] S2, convert the first voice signal into text information;

[0101] S3, parse the control intent category of the text information, wherein the control intent category is used to indicate the clarity of the intent of the control command corresponding to the text information;

[0102] S4, Perform audio zone control on the target vehicle according to the control intent category.

[0103] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0104] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0105] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0106] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0107] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0108] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0109] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0110] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for controlling the sound zones of a vehicle, characterized in that, include: The first voice signal inside the vehicle cabin of the target vehicle is collected, wherein the vehicle cabin includes multiple sound zones, and each sound zone corresponds to a physical space of the vehicle cabin; Convert the first speech signal into text information; The control intent category of the text information is parsed, wherein the control intent category is used to indicate the clarity of the intent of the control command corresponding to the text information; The target vehicle is controlled by its audio zone according to the control intent category. The control intent category of parsing the text information includes: determining whether the text format of the text information is a preset explicit format; if the text format of the text information is a preset explicit format, determining that the text information is an explicit audio zone control command; if the text format of the text information is not a preset explicit format, determining whether the text format of the text information is an implicit format; if the text format of the text information is an implicit format, determining that the text information is an implicit audio zone control command; if the text format of the text information is not an implicit format, determining that the text information is not an audio zone control command. Implicit audio zone control commands include: opening the passenger-side audio zone, closing the rear-side audio zone, indicating no one in the rear seats, and entering safe mode. Opening the passenger-side audio zone means opening one or more audio zones corresponding to the passenger-side position; closing the rear-side audio zone means closing one or more audio zones corresponding to the rear-side position; indicating no one in the rear seats means closing one or more audio zones corresponding to the rear-side position; and entering safe mode means closing all audio zones except the driver's seat. The method of controlling the audio zone of the target vehicle according to the control intent category includes: identifying the target audio zone and target operation in the text information according to the control intent category, wherein the target operation includes an audio zone opening operation or an audio zone closing operation, an audio zone state maintenance time, and an audio zone priority. The audio zone priority is used to increase the acquisition sensitivity of the microphone in the high-priority audio zone and decrease the acquisition sensitivity of the microphone in the low-priority audio zone; and executing the target operation on the target audio zone.

2. The method according to claim 1, characterized in that, Determining whether the text format of the text information is implicit includes: Determine whether the text format of the text information is a preset implicit format; and / or, The machine learning model is used to determine whether the text format of the text information is implicit; and / or, The deep neural network model is used to determine whether the text format of the text information is implicit.

3. The method according to claim 1, characterized in that, Performing the target operation on the target audio region includes: If the target operation is an enable operation, connect the data channel between the target audio region and the voice interaction service; If the target operation is a close operation, disconnect the data channel between the target audio region and the voice interaction service.

4. The method according to claim 1, characterized in that, Performing the target operation on the target audio region includes: If the target operation is an enable operation, the location information and sound source enhancement control information corresponding to the physical space of the target sound region are set in the sound region separation module; If the target operation is a shutdown operation, the location information and sound source suppression control information corresponding to the physical space of the target sound region are set in the sound region separation module.

5. The method according to claim 1, characterized in that, After performing audio zone control on the target vehicle according to the control intent category, the method further includes: Collect the second voice signal from inside the vehicle cabin of the target vehicle; The second speech signal is subjected to region separation to obtain first audio data from the first region and second audio data from the second region; The system determines that the first audio zone is in the open state and the second audio zone is in the closed state within the vehicle cabin, outputs the first audio data, and disables the output of the second audio data.

6. The method according to claim 1, characterized in that, After performing audio zone control on the target vehicle according to the control intent category, the method further includes: Collect third voice signals from the vehicle cabin of the target vehicle; The third speech signal is subjected to voice region separation to obtain third audio data from the third voice region and fourth audio data from the fourth voice region; The system determines that the third audio zone is in the open state and the fourth audio zone is in the closed state within the vehicle cabin, inputs the third audio data to the voice interaction service of the target vehicle, and prohibits the input of the fourth audio data to the voice interaction service.

7. A control device for vehicle audio zones, characterized in that, include: The first acquisition module is used to acquire the first voice signal inside the vehicle cabin of the target vehicle, wherein the vehicle cabin includes multiple sound zones, and each sound zone corresponds to a physical space of the vehicle cabin. A conversion module is used to convert the first speech signal into text information; A parsing module is used to parse the control intent category of the text information, wherein the control intent category is used to indicate the clarity of the intent of the control command corresponding to the text information; The first control module is used to perform audio zone control on the target vehicle according to the control intent category; The parsing module includes: a judgment unit, used to judge whether the text format of the text information is a preset explicit format; a first processing unit, used to determine that the text information is an explicit audio zone control command if the text format of the text information is a preset explicit format; and to judge whether the text format of the text information is an implicit format if the text format of the text information is not a preset explicit format; and a second processing unit, used to determine that the text information is an implicit audio zone control command if the text format of the text information is an implicit format; and to determine that the text information is not an audio zone control command if the text format of the text information is not an implicit format, wherein the implicit audio zone control command includes: opening the passenger audio zone, closing the rear audio zone, indicating no one in the rear seats, and entering safe mode, wherein opening the passenger audio zone means opening one or more audio zones corresponding to the passenger position, closing the rear audio zone means closing one or more audio zones corresponding to the rear position, indicating no one in the rear seats means closing one or more audio zones corresponding to the rear position, and entering safe mode means closing all audio zones except the driver's position; The first control module includes: a recognition unit, used to recognize the target audio region and target operation in the text information according to the control intent category, wherein the target operation includes an audio region opening operation or an audio region closing operation, an audio region state maintenance time, and an audio region priority, wherein the audio region priority is used to increase the acquisition sensitivity of the microphone in the high-priority audio region and decrease the acquisition sensitivity of the microphone in the low-priority audio region; and an execution unit, used to execute the target operation on the target audio region.

8. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method described in any one of claims 1 to 6 when it is run.

9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Vehicle-mounted audio management method and device, equipment, automobile and readable storage medium

    CN110648663A

  • Audio playing method and apparatus, and computer-readable storage medium and electronic device

    WO2023040820A1