Data processing method and related devices

By processing voices only within the acoustic ranges where the user is present, the method addresses excessive computing resource usage in intelligent vehicle cabins, ensuring efficient and accurate voice interaction.

JP2026514099APending Publication Date: 2026-05-01YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
YINWANG INTELLIGENT TECHNOLOGIES CO LTD
Filing Date
2024-03-28
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

The rapid development of intelligent cabins in vehicles leads to excessive computing resource occupation during multi-acoustic-range voice interaction, triggering protection mechanisms that limit high-load functions and impair the user experience due to delayed or failed voice activation and recognition.

Method used

Implement a data processing method that acquires user presence information across multiple acoustic ranges and processes voices only within the ranges where the user is present, reducing the number of voices to be processed and optimizing computing resource usage.

Benefits of technology

This approach ensures the smooth operation of voice interaction functions by minimizing computing resource consumption, enhancing response speed and accuracy, and maintaining the integrity of multi-acoustic-range voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026514099000001_ABST
    Figure 2026514099000001_ABST
Patent Text Reader

Abstract

To reduce the occupation of computing resources in multi-sound-range interaction, a data processing method and related devices are provided. The method includes acquiring multiple voices, wherein the multiple voices are from multiple sound ranges (S501), acquiring user information for the multiple sound ranges, wherein the user information indicates whether the user is present in the sound range (S502), and processing the multiple voices based on the user information for the multiple sound ranges (S503).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-machine interaction technology, particularly to data processing methods and related devices.

Background Art

[0002] With the development of artificial intelligence, the application of artificial intelligence technology is becoming increasingly widespread. Voice interaction functions based on artificial intelligence technology, such as voice-based question and answer, machine translation, and voice control, bring great convenience in scenarios such as users' learning, life, and work.

[0003] The application of voice interaction in intelligent electric vehicles is used as an example. In recent years, the intelligent electric vehicle industry has been developing rapidly, and the number of intelligent vehicles in use has been continuously increasing. As an important part of intelligent vehicles, the intelligent cabin is the main focus of vehicle intelligence. An intelligent vehicle can recognize the meaning of the voice emitted by a passenger in a specific acoustic area and perform corresponding responses or corresponding operations based on that meaning, such as opening and closing the vehicle window, turning on and off the multimedia, adjusting the temperature, positioning and navigation, etc., so as to improve driving safety and entertainment.

[0004] However, the enhancement of the intelligent cabin experience is increasing the challenges to the limited computing resources of the in-car infotainment system. If the computing resources occupied by the in-car infotainment system exceed the system's alarm threshold, or if the system's temperature exceeds a critical value under high load conditions, a protection mechanism is triggered, thereby limiting the use of certain high-load functions to reduce the load. This includes limiting multi-range voice interaction, one of the fundamental features of the intelligent cabin experience. Voice interaction functions are limited by issues such as delayed voice activation, delayed recognition, or complete failure. [Overview of the project]

[0005] This application provides a data processing method and related devices to solve the problem of excessive computing resource occupation in multi-range voice interaction.

[0006] A data processing method is provided according to the first aspect. The method can be applied to multi-sound-range voice interaction in means of transport, games, intelligent cinema, smart homes, or intelligent security protection scenarios. The method may be implemented by means of transport or a chip within means of transport, or by a computer, intelligent terminal device, or intelligent appliance, or a chip within a computer, intelligent terminal device, or intelligent appliance, etc. The method involves acquiring user information for multiple voices and multiple sound ranges, the multiple voices being from multiple sound ranges, and the user information including indicating whether the user is present in the sound range and processing the multiple voices based on the user information for the multiple sound ranges. For example, at least one microphone is placed in each sound range, and the multiple voices are captured by microphones in the multiple sound ranges. User information regarding whether the user is present in multiple sound ranges is acquired, and the multiple voice data from the multiple sound ranges is processed based on the user information, so that the sound ranges in which the voice data should be processed can be obtained by screening based on whether the user is present in that sound range. This can reduce the amount of voice data to be processed and reduce the computing resources required for voice processing.

[0007] In a feasible implementation, processing multiple voices based on user information across multiple acoustic ranges includes processing a portion of the voices screened from multiple voices based on user information across multiple acoustic ranges. This reduces the number of voices that need to be processed, thereby reducing computing resource usage and ensuring the proper functioning of the multi-acoustic-range voice interaction feature.

[0008] In possible implementations, processing multiple speeches based on user information across multiple acoustic ranges includes processing the speeches within the acoustic ranges in which the user resides, based on user information across multiple acoustic ranges. The speeches within the acoustic ranges in which the user resides are processed. If the number of acoustic ranges in which the user resides is less than the total number of acoustic ranges, the number of speeches that need to be processed can be reduced. Furthermore, because the distance between the microphone and the user within the same acoustic range is short, the speech captured by the microphone within the acoustic range in which the user resides has a high signal-to-noise ratio, thereby ensuring accuracy of speech recognition even when the number of speeches to be processed is reduced.

[0009] In a feasible implementation, processing multiple voices based on user information across multiple acoustic ranges includes discarding voices in acoustic ranges where no user exists, based on the user information across those ranges. Thus, since voices in acoustic ranges where no user exists are discarded, the number of voices that need to be processed is reduced, thereby decreasing the computing resources occupied by multi-acoustic-range voice interaction. Furthermore, some memory resources are freed, further reducing memory resource utilization.

[0010] In possible implementations, processing multiple speeches based on user information across multiple acoustic ranges includes processing speeches within the acoustic ranges where the user is present and speeches within the acoustic ranges where the user is not present, based on user information across multiple acoustic ranges. In this way, the accuracy of speech recognition can be further improved even when the number of speeches that need to be processed is reduced.

[0011] In a possible implementation, the method further includes obtaining computing resource usage, and if computing resource usage exceeds a threshold, processing multiple speeches based on user information across multiple acoustic ranges includes processing the speech of the portion of the acoustic range where the user resides, based on user information across multiple acoustic ranges. If computing resource usage is high, the speech of the portion of the acoustic range where the user resides is processed, thereby further reducing the number of speeches that need to be processed and further reducing the computing resource occupation of speech processing.

[0012] In a feasible implementation, processing the audio within the acoustic range in which the user resides includes processing the audio within the acoustic range in which the user resides, which is part of a target acoustic range that is part of multiple acoustic ranges. Since the number of acoustic ranges within the target acoustic range is less than the total number of acoustic ranges, the amount of audio that needs to be processed is reduced, and the computing resources required for audio processing can be reduced.

[0013] Optionally, multiple acoustic ranges are areas corresponding to multiple seats within the vehicle's cabin, and the target acoustic range includes the driver's seat area and / or passenger seat area within the areas corresponding to multiple seats. When computing resource load is high, only the voices from the driver's cabin and / or passenger seat cabin are processed, further reducing the number of voices that need to be processed and thus reducing the computing resources required for voice interaction. This ensures the normal use of driver and passenger voice interaction when computing resource load is high.

[0014] In a feasible implementation, processing audio in the portion of the acoustic range where the user resides includes processing audio in at least one acoustic range with the highest priority within the acoustic range where the user resides. When computing resources are insufficient, audio in high-priority acoustic ranges is processed preferentially, and when the number of audios to be processed is reduced and the computing resource load is lowered, the normal operation of voice interaction functions in high-priority acoustic ranges is ensured.

[0015] In possible implementations, multiple acoustic zones are areas corresponding to multiple seats within the vehicle's interior, and one acoustic zone includes areas corresponding to one or more seats.

[0016] In accordance with the second aspect, the apparatus is provided. The apparatus includes an acquisition module and a processing module. The acquisition module is configured to acquire multiple voices, which are from multiple acoustic ranges. The acquisition module is configured to acquire user information from multiple acoustic ranges, which indicates whether the user is present in an acoustic range. The processing module is configured to process the multiple voices based on the user information from multiple acoustic ranges.

[0017] In a possible implementation, the processing module is specifically configured to process a portion of the audio obtained by screening from multiple audio sources based on user information across multiple acoustic ranges.

[0018] In possible implementations, the processing module is specifically configured to process audio in the acoustic range where a user resides, based on user information across multiple acoustic ranges.

[0019] In possible implementations, the processing module is specifically configured to discard audio in acoustic ranges where the user is not present, based on user information across multiple acoustic ranges.

[0020] In a possible implementation, the processing module is specifically configured to process, based on user information across multiple acoustic ranges, the audio within the acoustic range where the user is present and a portion of the audio within the acoustic range where the user is not present, from among multiple audio.

[0021] In a possible implementation, the processing module is specifically configured to determine, based on user information from multiple acoustic ranges, whether a user is present in a target acoustic range within multiple acoustic ranges. The processing module is also specifically configured to process the audio in the target acoustic range.

[0022] In a possible implementation, the acquisition module is configured to acquire computing resource usage. The processing module is specifically configured to process the audio in the portion of the acoustic range where a user is present, based on user information across multiple acoustic ranges, when computing resource usage exceeds a threshold.

[0023] In a possible implementation, the processing module is specifically configured to process the audio of the acoustic range in which the user is located, within a target acoustic range which is part of multiple acoustic ranges.

[0024] In a possible implementation, multiple acoustic ranges are areas corresponding to multiple seats within the vehicle's cabin, and the target acoustic range includes the driver's seat area and / or passenger seat area within the areas corresponding to multiple seats.

[0025] In possible implementations, multiple acoustic ranges have a priority order, and the processing module is specifically configured to process the audio of at least one acoustic range with the highest priority among the acoustic ranges in which the user resides.

[0026] A device is provided according to the third aspect. The device includes a processor and memory. The processor is coupled to the memory and is configured to perform a data processing method relating to the first aspect or any one of the possible implementations of the first aspect, based on instructions stored in the memory.

[0027] According to a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium includes instructions, and when the computer-readable storage medium is executed on a computer, the computer can execute a data processing method according to the first aspect or any one of the possible implementations of the first aspect.

[0028] According to a fifth aspect, a computer program product including instructions is provided. When the instructions are executed by an electronic device, the electronic device can execute a data processing method according to the first aspect or any one of the possible implementations of the first aspect.

Brief Description of the Drawings

[0029] [Figure 1] It is a functional block diagram of a vehicle according to the present application. [Figure 2a] It is a diagram of acoustic region division in a vehicle according to the present application. [Figure 2b] It is another diagram of acoustic region division in a vehicle according to the present application. [Figure 3a] It is a diagram of the system architecture according to the present application. [Figure 3b] It is a diagram of another system architecture according to the present application. [Figure 4a] It is a diagram of a scenario where modules in a system according to the present application collaboratively process multiple voices. [Figure 4b] It is a diagram of another scenario where modules in a system according to the present application collaboratively process multiple voices. [Figure 4c] It is a diagram of yet another scenario where modules in a system according to the present application collaboratively process multiple voices. [Figure 5] It is a schematic flowchart of an audio processing method according to the present application. [Figure 6] It is a diagram of the structure of a device according to the present application. [Figure 7] It is a diagram of the structure of a device according to the present application.

Modes for Carrying Out the Invention

[0030] This application provides a data processing method and related devices to reduce the computing resources occupied by multi-range voice interaction.

[0031] In this application, "at least one" means one or more, and "multiple" means two or more. The term "and / or" indicates an association between related objects, indicating that three relationships exist. For example, A and / or B could mean A only, both A and B exist, and B only exists. Here, A and B only may be singular or plural. The letter " / " generally indicates an "OR" relationship between related objects. "At least one of the following items (parts)" or similar expressions indicate any combination of these items, including any single item (part) or any combination of multiple items (parts). For example, at least one item (part) of a, b, or c could mean a, b, c, ab, ac, bc, or abc. Here, a, b, and c may be singular or plural. Furthermore, in the embodiments of this application, terms such as "first" and "second" do not limit the quantity or order of execution.

[0032] In the various embodiments of the present application, unless otherwise stated or unless there is a logical contradiction, the terminology and / or descriptions in different embodiments are consistent and may be cross-referenced, and the technical features in different embodiments may be combined on the basis of their internal logical relationships to form new embodiments.

[0033] Due to the rapid development of voice processing technology, the speed and accuracy of voice processing are increasing, and as voice processing technology is applied to more electronic products, electronic products are becoming more intelligent. Voice processing technology generally includes speech preprocessing, speech recognition, etc. Speech preprocessing mainly involves processing such as noise reduction and speech amplification, and speech recognition mainly involves recognizing the semantic information carried in the preprocessed speech, thereby enabling electronic devices to make corresponding responses based on the semantic information. Currently, electronic devices with voice processing capabilities include intelligent terminal devices such as smartphones, smartwatches, smart speakers, tablet computers, or notebook computers, as well as intelligent appliances such as smart refrigerators, smart TVs, or floor cleaning robots, and even intelligent means of transport such as cars, trucks, motorcycles, buses, ships, or airplanes. Users can interact with electronic devices with voice processing capabilities by voice. For example, users can communicate with electronic devices by voice without manually entering text. As another example, users can control the power on / off, operating status, and mode of electronic devices by voice, without manual operation. This significantly streamlines the user's work and life.

[0034] To further improve the accuracy of speech recognition, in some complex or large-scale voice capture environments, such as seating areas in mobile transport vehicles or, in the case of indoor environments, the environment is divided into multiple acoustic zones. At least one microphone is placed in each acoustic zone, so that voices uttered by users at different locations within the environment can be clearly captured by at least one microphone in at least one acoustic zone. In this way, the voices captured by the microphones have a high signal-to-noise ratio, thus improving the accuracy of speech recognition performed on the voices. Multi-acoustic-zone voice interaction means that in a scenario where the environment is divided into multiple acoustic zones, microphones in multiple acoustic zones capture voices, and a device with multi-acoustic-zone voice processing capabilities further processes the voices in multiple acoustic zones, such as voice processing and speech recognition, to generate corresponding responses. Multi-acoustic-zone voice interaction can ensure that there is a response to voices uttered by users at different locations. However, in multi-acoustic-zone voice interaction, electronic devices need to process voices from multiple acoustic zones simultaneously, which occupies a large amount of computing and memory resources. If the computing and memory resources of the voice processing unit are finite, or if the load on the voice processing unit is high, it will affect the timeliness and accuracy of voice processing.

[0035] Vehicles are used as an example of a means of transportation. As vehicle intelligence and networking improve, vehicle cabins are gradually evolving into intelligent cabins with human-machine interaction as their core, and intelligent voice control within the vehicle cabin has become a mainstream requirement for current intelligent cabins. Multiple microphones may be placed within the vehicle cabin. By using multiple microphones, the vehicle can capture voice signals in the environment, recognize voice commands within the voice signals, and perform operations corresponding to the voice commands. Intelligent cabins primarily satisfy the occupant's driving and entertainment requirements, and in-vehicle infotainment must process large amounts of driving and user information, which is increasingly challenging for the limited computing resources of in-vehicle infotainment. Currently, in multi-acoustic-range (e.g., 4-acoustic-range, 5-acoustic-range, or 6-acoustic-range) voice interaction solutions for intelligent cabins, the number of voice streams that need to be processed and decoded in real time during voice preprocessing, voice activation, sound source positioning, and speech recognition, and the number of simultaneous related algorithm models, are proportional to the number of acoustic ranges. For example, if a vehicle has four acoustic ranges, the voices of all four acoustic ranges must be processed simultaneously, and if a vehicle has six acoustic ranges, the voices of all six acoustic ranges must be processed simultaneously, resulting in excessive computing resource usage across the entire voice interaction. Consequently, multi-acoustic range interaction has become one of the fundamental factors triggering high loads on in-vehicle infotainment systems. If the computing resources occupied by the in-vehicle infotainment system exceed the system's alarm threshold, or if the system's temperature exceeds a critical value under high-load conditions, a protection mechanism is triggered, thereby limiting the use of certain high-load functions to reduce the load. This includes limiting multi-acoustic range voice interaction in the intelligent cabin. The user's driving experience is significantly impaired by issues such as delayed voice activation, delayed recognition, or complete failure.This significantly degrades the user's driving experience.

[0036] To address the aforementioned technical challenges, the present invention provides the following embodiments to reduce computing resource usage in multi-range voice interaction, thereby ensuring the successful implementation of intelligent functions, including voice interaction.

[0037] The solution provided herein is applicable to multi-acoustic-range voice interaction in transportation scenarios and may be further applicable to multi-acoustic-range voice interaction in gaming, intelligent cinema, smart home, or intelligent security protection scenarios. In this application, user information is acquired for multiple acoustic ranges, indicating whether the user is present in the corresponding acoustic range, and then some of the voices in the multiple acoustic ranges are processed based on the user information to reduce the number of voices that need to be processed. This reduces the computing resources that need to be occupied for voice processing, ensuring the smooth operation of the voice interaction. Furthermore, because fewer voices need to be processed, voice processing efficiency can be improved to some extent to enhance the response speed of the voice interaction.

[0038] The following uses multi-range voice interaction in a vehicle scenario as an example for explanation. It is understandable that the solution provided herein may be applicable to other scenarios. In other words, the problem of excessive computing resource utilization in multi-range voice interaction can be solved in other scenarios based on the same principles. The architectures and service scenarios described herein are intended to more clearly illustrate the technical solution in the embodiments herein and do not constitute limitations on the technical solution provided herein. Those skilled in the art will recognize that, with the evolution of network architectures and the emergence of new service scenarios, the technical solution provided herein is applicable to similar technical challenges.

[0039] Figure 1 is a functional block diagram of the vehicle according to the present invention. The vehicle may include multiple microphones, a sensing system, and a computing platform.

[0040] The interior space (cabin) of a vehicle may be divided into multiple acoustic zones, with at least one microphone, for example, microphone 1 through microphone m (where m is a positive integer greater than or equal to 2), placed in each acoustic zone. For example, as shown in Figure 2a, the interior of a vehicle may be divided into four acoustic zones: the driver's area, the passenger's area, the second-row left area, and the second-row right area. In this case, microphones may be placed in the driver's area, the passenger's area, the second-row left area, and the second-row right area, respectively. As another example, as shown in Figure 2b, in the case of a sport utility vehicle (SUV), the interior area where the user sits may be divided into six acoustic zones: the driver's area, the passenger's area, the second-row left area, the second-row right area, the third-row left area, and the third-row right area. In this case, microphones may be placed in the driver's seat area, passenger seat area, second row left area, second row right area, third row left area, and third row right area, respectively. Here, the division into four and six acoustic frequencies is used merely as an example. Based on the type and model of the vehicle, the size of the interior space, the number of cabins, etc., the interior of the vehicle may alternatively be divided into two, five, seven, or more acoustic frequencies. Examples of each are not given here.

[0041] The sensing system is configured to detect whether a user is present in each acoustic range and to obtain user information regarding the presence of a user in each acoustic range, the user information indicating whether the user is present in the corresponding acoustic range.

[0042] The sensing system may include one or more sensors that, based on data captured by the sensors, can determine which acoustic areas a user is present in and which are not. In a possible implementation, the sensing system may include, for example, a camera device (image sensor). The camera device may capture images of the interior of the vehicle and analyze the images to determine whether a user is present in each acoustic area. In another possible implementation, the sensing system may include, for example, pressure sensors. Pressure sensors may be placed in the seats of each acoustic area. When a user sits in a seat, the pressure sensor can sense the pressure and convert the pressure signal into an electrical signal, which can then be used to determine whether a user is present in the corresponding acoustic area. In yet another possible implementation, the sensing system may include distance measuring sensors, with at least one distance measuring sensor placed in each acoustic area, and whether a user is seated in a seat can be determined by distance measurement. The sensing system may include one or more of the sensors described above. When the sensing system includes multiple sensors, whether a user is present in the corresponding acoustic area can be determined comprehensively by referring to data captured by multiple sensors, which can improve the accuracy of the acquired user information for each acoustic area. Indeed, the sensing system may further include other sensors configured to detect whether a user is present in an acoustic range, such as a temperature sensor or an infrared sensor. This is not limited here. In yet another possible implementation, the sensing system may include a graphical user interface (GUI). The GUI may be displayed on a central display screen in the vehicle or on the user's terminal device. The GUI displays each acoustic range to the user, and the user specifies the acoustic range in which they are present by voice, touch, etc.

[0043] A microphone is an acoustic sensor, and its function is to capture sound in the environment and convert that sound into electronic signals. Any device capable of capturing sound and converting the captured sound into electronic signals falls within the scope of a microphone as defined in this application. Specific implementations of microphone devices are not limited in this embodiment of the application. In addition to collecting sound emitted by sound sources within the acoustic range to which the microphone belongs, a microphone may also capture sound emitted by sound sources in other acoustic ranges. Generally, a microphone located in the same acoustic range as a user is close to the user. In other words, the distance between the microphone and the sound source is short. Therefore, the user's voice, located in the same acoustic range and captured by the microphone, has a higher signal-to-noise ratio.

[0044] The multi-range voice interaction of a vehicle may be controlled by a computing platform. The computing platform may include one or more processors, for example, processor 1 to processor n (where n is a positive integer). A processor is a circuit with signal processing capabilities. In an implementation, a processor may be a circuit with instruction read and execute capabilities, and may be, for example, a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which may be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, a processor may implement specific functions based on the logic relationships of hardware circuits. The logic relationships of hardware circuits may be fixed or reconfigurable. For example, a processor may be a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as a field programmable gate array (FPGA). In a reconfigurable hardware circuit, the process by which the processor reads a configuration document and implements the hardware circuit configuration can be understood as the process by which the processor loads instructions and implements some or all of the functions of the aforementioned unit. Furthermore, the processor may alternatively be a hardware circuit designed for artificial intelligence, and can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), or a deep learning processing unit (DPU). Furthermore, the computing platform may further include memory, which is configured to store instructions.Some or all of processors 1 through n may call and execute instructions in memory in order to implement the corresponding functions.

[0045] User information for each acoustic range acquired by the sensing system may be input to the computing platform, and audio captured by multiple microphones may also be input to the computing platform, so that the computing platform can process audio from parts of multiple acoustic ranges based on the user information for each acoustic range. Processing audio from parts of multiple acoustic ranges may include any one of the following: 1. Processing audio captured by microphones in the acoustic range where the user is present; 2. Processing audio captured by microphones in the acoustic range where the user is present and audio captured by microphones in an acoustic range where the user is not present, for example, an acoustic range adjacent to the acoustic range where the user is present and where the user is not present; or 3. Processing audio from a part of the acoustic location where the user is present, for example, processing audio from the acoustic range where the user is present within a target acoustic range, where the target acoustic range is part of multiple acoustic ranges, for example, the target acoustic range includes the driver's seat area and / or the passenger seat area; or processing audio from at least one acoustic range with the highest priority among the acoustic ranges where the user is present. In this way, since audio from multiple acoustic ranges is processed based on user information, the number of audio elements that need to be processed is reduced, thereby reducing the computing resources occupied by the computing platform.

[0046] The speech processing performed by the computing platform includes speech preprocessing and speech backend processing. Speech preprocessing includes at least one of the following: endpoint detection, amplification, filtering, noise reduction, echo cancellation, crosstalk cancellation, etc., to reduce interference in the speech such as ambient noise, echoes, and reflections. In this way, the signal-to-noise ratio of the speech is improved, a purer human voice can be obtained, and the accuracy of subsequent speech recognition can be improved. Endpoint detection involves dividing the speech into speech segments and non-speech segments, after which processing such as noise reduction and speech recognition can be performed on the speech segments, thereby reducing the amount of data that needs to be processed and improving processing efficiency. Noise estimation is performed by using the non-speech segments to suppress noise, thereby achieving a speech enhancement effect. Amplification is increasing the strength of the speech signal. Processing such as filtering, noise reduction, echo cancellation, and crosstalk cancellation suppresses interference such as ambient noise, early reflections, and multiple reflections in order to enhance the human voice signal.

[0047] The voice backend processing performs at least one of the following actions on the voice acquired by voice processing: voice detection, activation word recognition, sound source positioning, acoustic range positioning, and speech recognition, in order to recognize the meaning of the voice in the voice and determine the location of the sound source. Voice detection determines whether a human voice is present in the voice. Activation word recognition recognizes keywords in the voice, such as "hello," "hey, Celia," or "hey, X." After recognition determines that the voice contains an activation word, the speech recognition function is activated. Activation word recognition is optional. In possible implementations, speech recognition may be performed directly without speech word recognition. For example, commands in the user voice, such as "start navigation," "open the vehicle window," "increase / decrease the volume," or "increase / decrease the temperature," may be recognized directly without activation word recognition. Acoustic range positioning and sound source positioning determine the acoustic range (location) where the user who uttered the voice is located, and determine whether to appropriately perform the corresponding operation on the acoustic range or whether the corresponding acoustic range has the corresponding permission. Speech recognition, also known as automatic speech recognition (ASR), aims to convert the vocabulary text contained in a person's voice into computer-readable input such as keystrokes, binary code, or strings of characters.

[0048] Figure 3a is a diagram of the system architecture relating to the present application. Figure 3b is a diagram of another system architecture relating to the present application. As shown in Figures 3a and 3b, the system architecture mainly includes a framework comprising multiple microphones arranged in a distributed manner, a dynamic acoustic range adjustment module, a voice preprocessing algorithm module, a voice interaction algorithm module, a backend execution module, and so on. In Figure 3a, in addition to acquiring user information across multiple acoustic ranges, the dynamic acoustic range adjustment module further acquires audio captured by microphones across multiple acoustic ranges, the dynamic acoustic range adjustment module determines which audio to input to the audio preprocessing algorithm module from among the audio across multiple acoustic ranges based on the user information across multiple acoustic ranges, and the audio preprocessing algorithm module processes the audio input by the dynamic acoustic range adjustment module. In Figure 3b, the audio preprocessing algorithm module acquires audio across multiple acoustic ranges, the dynamic acoustic range adjustment module acquires user information across multiple acoustic ranges, inputs the user information across multiple acoustic ranges to the audio preprocessing algorithm module, and the audio preprocessing algorithm module determines which audio to process from among the audio based on the user information across multiple acoustic ranges. This is the difference between Figure 3a and Figure 3b. Indeed, in another implementation, the vehicle's local system architecture does not need to include a voice interaction algorithm module, and the functions of the voice interaction algorithm module may be implemented by a cloud server. After the audio is processed by the audio preprocessing algorithm module, the vehicle uploads the preprocessed audio to the cloud. After speech recognition and other processes are completed, the cloud retrieves the recognition results and then sends them to the vehicle. The recognition results are then executed by the vehicle's backend execution module. The system architecture shown in Figure 3a is used below as an example to illustrate the module's functionality. The system architecture in Figure 3b is similar, and therefore, details are not described again.

[0049] Multiple microphones are positioned in different acoustic ranges, configured to capture sound across all acoustic ranges in real time.

[0050] The dynamic acoustic range adjustment module receives audio data acquired by receiving sound from microphones in all acoustic ranges. The dynamic acoustic range adjustment module can further acquire user information, i.e., detect whether a person is present in a seat in each acoustic range within the vehicle, and determine the number and location of active acoustic ranges based on the user information. The determination results determine the number and channel of audio that needs to be processed by the downstream audio preprocessing algorithm module and voice interaction algorithm module. An active acoustic range is an acoustic range corresponding to the audio to be processed thereafter; that is, the audio of an active acoustic range among multiple acoustic ranges is the audio to be processed among multiple audio. For example, if the active acoustic range includes acoustic range 1, acoustic range 2, and acoustic range 3, the audio that needs to be processed by the audio preprocessing algorithm module and voice interaction algorithm module includes audio from three acoustic ranges: audio from acoustic range 1, audio from acoustic range 2, and audio from acoustic range 3. In this application, for the sake of brevity, audio captured by microphones in an acoustic range is usually abbreviated as acoustic range audio.

[0051] In implementation, the active acoustic range may be the acoustic range where the user is present. For example, as shown in Figure 2b, it is assumed that in acoustic ranges 1 through 6, the user is present in acoustic ranges 1, 2, and 3, and not in acoustic ranges 4, 5, and 6. In this case, the active acoustic ranges are acoustic ranges 1, 2, and 3, and the speech subsequently processed by the speech preprocessing algorithm module and the voice interaction algorithm module is the speech in acoustic ranges 1, 2, and 3. In another implementation, the active acoustic range may include the acoustic range where the user is present and a portion of the acoustic range where the user is not present, such as an acoustic range adjacent to the acoustic range where the user is present. For example, in acoustic ranges 1 to 6 shown in Figure 2b, it is assumed that users are present in acoustic ranges 1, 2, and 3, but not in acoustic ranges 4, 5, and 6. Acoustic range 4 is adjacent to acoustic range 3, and there are no obstacles such as seats obstructing acoustic ranges 4 and 3. It is also assumed that the speech emitted by the user in acoustic range 3 and captured by the microphone in acoustic range 4 has a high signal-to-noise ratio, which can improve the success rate of speech recognition. In this case, the active acoustic ranges may be acoustic ranges 1, 2, 3, and 4, and the speech subsequently processed by the speech preprocessing algorithm module and the voice interaction algorithm module is the speech in acoustic ranges 1, 2, 3, and 4. In yet another implementation, when the user is present in a target acoustic range within multiple acoustic ranges, the active acoustic range is the target acoustic range. In a vehicle scenario, the target acoustic range is, for example, the acoustic range corresponding to the driver's seat or the passenger seat. For example, in the acoustic ranges 1 through 6 shown in Figure 2b, it is assumed that users are present in acoustic ranges 1, 2, and 3, but not in acoustic ranges 4, 5, and 6, and that the target acoustic ranges are acoustic ranges 1 and 2, and that the target acoustic ranges have a high priority. In this case, the active acoustic ranges are acoustic ranges 1 and 2. In this case, the active acoustic ranges do not include acoustic range 3.The audio processed subsequently by the audio preprocessing algorithm module and the voice interaction algorithm module consists of audio from acoustic range 1 and acoustic range 2. In this way, the computing resource usage of multi-acoustic range voice interaction can be further reduced.

[0052] The speech preprocessing algorithm module preprocesses the speech in the active acoustic range input by the dynamic acoustic range adjustment module based on the location of the active acoustic range provided by the dynamic acoustic range adjustment module. Speech preprocessing includes at least one of the following: endpoint detection, speech amplification, filtering, noise reduction, echo cancellation, and crosstalk cancellation, in order to improve the signal-to-noise ratio of the speech received by the microphone corresponding to each acoustic range so that the speech quality requirements for subsequent voice interaction are satisfied. Algorithms used for voice endpoint detection include, for example, the short-time energy method, the zero crossing factor method, the cepstrum coefficient method, and the extended Gaussian mixture model. Algorithms used for noise reduction may include the least mean square (LMS) algorithm and Wiener filtering. Algorithms used for echo cancellation include, for example, the LMS algorithm or the normalized least mean square (NLMS) algorithm. Algorithms used for reflection cancellation include, for example, inverse filtering, beamforming algorithms, and deep learning algorithms.

[0053] The voice interaction algorithm module is configured to perform voice backend processing on speech in the active acoustic range. Specifically, the decoding by the voice interaction algorithm module, based on the location of the active acoustic range provided by the dynamic acoustic range adjustment module and the corresponding active acoustic range speech signal processed by the voice preprocessing algorithm module, includes trigger word recognition, sound source positioning, acoustic range locking, and speech recognition. Decoding involves performing statistical mode recognition on the feature vectors of the user's voice by using trained "acoustic models" and "language models." The function of the voice interaction algorithm module is to parse the speaker's voice command intent within the target acoustic range and determine the acoustic range in which the user who sent the corresponding voice command is located.

[0054] The backend execution module performs corresponding subsequent operations, such as responding to a voice broadcast or executing a vehicle control command, based on the decoding result of the voice dialogue algorithm module.

[0055] Signaling streams and data streams are transmitted between the dynamic acoustic range adjustment module, the voice preprocessing algorithm module, and the voice interaction algorithm module. The data stream is the active acoustic range audio, which is acquired by the dynamic acoustic range adjustment module through screening from audio in multiple acoustic ranges based on user information for each acoustic range. The signaling stream is information about the active acoustic range. The dynamic acoustic range adjustment module sends information about the active acoustic range to the voice preprocessing algorithm module and the voice interaction algorithm module. This allows the voice preprocessing algorithm module to individually determine which acoustic range audio was input by the dynamic acoustic range adjustment module based on the information about the active acoustic range, and the voice interaction algorithm module to determine which acoustic range audio was preprocessed and input by the voice preprocessing algorithm module based on user information for multiple acoustic ranges. Microphones in different acoustic ranges are placed in different positions, and the position, distance, angle, etc., of the microphones relative to the user differ. For some algorithms in the speech preprocessing algorithm module and the voice interaction algorithm module, the algorithm used may vary depending on the acoustic range of the audio. For example, when the speech preprocessing algorithm module performs noise reduction on audio in different acoustic ranges, some parameters of the noise reduction algorithm used may differ. Based on the information about the active acoustic range provided by the dynamic acoustic range adjustment module, the speech preprocessing algorithm module and the voice interaction algorithm module can accurately determine which acoustic range the input audio is from and accurately select the corresponding algorithm to be used for audio processing. Furthermore, for some algorithms, such as the echo cancellation algorithm, crosstalk algorithm, and sound source positioning algorithm, calculations need to be performed by referencing audio from multiple channels. The number of microphones, and the distance and relative position between microphones, affect these algorithms.The audio preprocessing algorithm module and the voice interaction algorithm module can accurately determine which acoustic range the input audio belongs to based on information about the active acoustic range provided by the dynamic acoustic range adjustment module, and perform corresponding algorithm adjustments, such as adjusting the number of simultaneously activated engines and ASR engines, and adjusting the number of audio channels processed by the sound source positioning algorithm.

[0056] If user information changes, for example, if the user leaves an acoustic range in which they originally resided, or if the user enters an acoustic range in which they originally did not reside, the dynamic acoustic range adjustment module may, after detecting the change in user information, promptly transmit the latest information regarding the active acoustic range to the preprocessing algorithm module and the voice interaction algorithm module. This allows the preprocessing algorithm module and the voice interaction algorithm module to adjust the corresponding algorithms accordingly. This guarantees a high success rate for speech recognition.

[0057] To facilitate understanding of the solution presented in this application, the following will be described with reference to specific scenarios. It should be understood that these scenarios are used merely as examples and should not be interpreted as limitations on this application.

[0058] As shown in Figure 4a, the active acoustic range is the acoustic range in which the user is located, and an example in which the user is located in acoustic ranges 1 and 4 is used for explanation. In this scenario, the dynamic acoustic range adjustment module detects that the user is located in acoustic ranges 1 and 4, notifies the speech preprocessing algorithm module and the voice interaction algorithm module that the active acoustic ranges are acoustic ranges 1 and 4, and sends only the audio from acoustic ranges 1 and 4 out of acoustic ranges 1 to 6 to the speech preprocessing module for processing. The speech preprocessing module receives the audio signals from the two channels and adaptively adjusts the algorithm based on the information that the two channels of audio signals are from acoustic ranges 1 and 4. After preprocessing the audio from acoustic ranges 1 and 4, the speech preprocessing algorithm module sends the preprocessed audio signals from acoustic ranges 1 and 4 to the voice interaction algorithm module. The voice interaction algorithm module receives information regarding the activation of acoustic ranges 1 and 4, as well as pre-processed audio from acoustic ranges 1 and 4. It adaptively adjusts the number of concurrently activated engines and concurrently activated ASR engines to 2, adaptively adjusts the sound source positioning algorithm for parsing the audio channels to channels 1 and 4, and finally outputs the parsing results for the target acoustic range for execution in the backend.

[0059] As shown in Figure 4b, the active acoustic range is the acoustic range in which the user is present and a portion of the acoustic range in which the user is not present. For illustrative purposes, an example is used in which the user is present in acoustic ranges 1, 2, and 3. In this scenario, the dynamic acoustic range adjustment module detects that the user is present in acoustic ranges 1, 2, and 3, determines that the active acoustic ranges are acoustic ranges 1, 2, 3, and 4, and notifies the speech preprocessing algorithm module and the voice interaction algorithm module that the active acoustic ranges are acoustic ranges 1, 2, 3, and 4. Furthermore, only the audio from acoustic ranges 1, 2, 3, and 4 out of acoustic ranges 1 through 6 is sent to the speech preprocessing module for processing. The speech preprocessing module receives the audio signals from the four channels and adaptively adjusts the algorithm based on the information that the audio signals from the four channels are from acoustic ranges 1, 2, 3, and 4. After preprocessing the audio in acoustic ranges 1, 2, 3, and 4, the audio preprocessing algorithm module sends the preprocessed audio signals from acoustic ranges 1, 2, 3, and 4 to the voice interaction algorithm module. The voice interaction algorithm module receives information regarding the activation of acoustic ranges 1, 2, 3, and 4, along with the preprocessed audio from acoustic ranges 1, 2, 3, and 4, adaptively adjusts the number of simultaneously activated engines and simultaneous ASR engines to 4, adaptively adjusts the sound source positioning algorithm for parsing the audio channels to channels 1, 2, 3, and 4, and finally outputs the parsing results for the target acoustic range for execution in the backend.

[0060] As shown in Figure 4c, the active acoustic range is the acoustic range within the target acoustic range in which the user is located, and an example in which the user is located in acoustic ranges 1, 2, and 3 is used for explanation. In this scenario, it is assumed that the target acoustic ranges are acoustic ranges 1 and 2. The dynamic acoustic range adjustment module detects that the user is located in acoustic ranges 1, 2, and 3, determines that the active acoustic ranges are acoustic ranges 1 and 2, notifies the speech preprocessing algorithm module and the voice interaction algorithm module that the active acoustic ranges are acoustic ranges 1 and 2, and sends only the audio from acoustic ranges 1 and 2 out of acoustic ranges 1 to 6 to the speech preprocessing module for processing. The speech preprocessing module receives the two channels of audio signals and adaptively adjusts the algorithm based on the information that the two channels of audio signals are from acoustic ranges 1 and 2. After preprocessing the audio in acoustic range 1 and acoustic range 2, the audio preprocessing algorithm module sends the preprocessed audio signals from acoustic range 1 and acoustic range 2 to the voice interaction algorithm module. The voice interaction algorithm module receives information regarding the activation of acoustic range 1 and acoustic range 2, as well as the preprocessed audio from acoustic range 1 and acoustic range 2, adaptively adjusts the number of simultaneous activation engines and simultaneous ASR engines to 2, adaptively adjusts the sound source positioning algorithm for parsing the audio channels to channel 1 and channel 2, and finally outputs the parsing results for the target acoustic range for execution in the backend.

[0061] In this embodiment, the dynamic acoustic range adjustment module acquires user information of the acoustic range and, based on this user information, determines which voices from among multiple voices are input for processing to the voice preprocessing algorithm module and the voice interaction algorithm module. This reduces the amount of voice that needs to be processed, thereby reducing the amount of computing resources occupied. This ensures the normal use of multi-acoustic range voice interaction.

[0062] Figure 5 is a schematic flowchart of the voice processing method according to the present invention. This embodiment may be implemented by a means of transport (e.g., a vehicle), or by the aforementioned computing platform, or by a system including a computing platform and a microphone, or by a system-on-a-chip (SOC) on the aforementioned computing platform, or by a processor within the computing platform.

[0063] S501: Multiple audio sources are acquired, and these multiple audio sources originate from multiple acoustic ranges.

[0064] Multiple sounds are captured by microphones placed in multiple acoustic ranges, and a correspondence exists between the acoustic ranges and the sounds captured by the microphones within those ranges.

[0065] S502: Obtain user information for multiple acoustic ranges, and the user information indicates whether the user is present in the acoustic range.

[0066] The primary function of a microphone within an acoustic range is to capture the user's voice within that range. The user is the service object for multi-acoustic range voice interaction, and the voice emitted by the user is the target object for capture and recognition. Generally, for voices emitted by the same user, the signal-to-noise ratio of the voice captured by a microphone in the same acoustic range as the user will be higher than the signal-to-noise ratio of the voice captured by a microphone in a different acoustic range than the user. For example, in Figure 2a, for voices emitted by the user in acoustic range 1, since the microphone in acoustic range 1 is close to the user and there are no obstacles in the sound propagation path, the signal-to-noise ratio of the user's voice in the audio captured by the microphone in acoustic range 1 will be higher than the signal-to-noise ratio of the user's voice in the audio captured by the microphones in acoustic ranges 2 through 4. Furthermore, the distance and direction of microphones in other acoustic ranges differ from that of the user who made the sound. Obstacles such as seat backs exist between some acoustic ranges and the user who made the sound, even if there are no obstacles between other acoustic ranges and the user who made the sound. As a result, the signal-to-noise ratio of the same sound captured by microphones in different acoustic ranges will differ. For example, in Figure 2a, acoustic range 3 is adjacent to acoustic range 4 and is not obstructed by a seat. Therefore, the signal-to-noise ratio of the sound made by the user in acoustic range 3 and captured by microphones in acoustic range 4 is likely to be higher than the signal-to-noise ratio of the sound made by the user in acoustic range 3 and captured by microphones in acoustic ranges 1 and 2. Thus, since user information indicating whether the user is present in an acoustic range is obtained, it is possible to determine which of the multiple sounds should be processed, ensuring the accuracy of voice interaction and reducing the computing resource occupation of multi-acoustic range voice interaction. This ensures the stability of voice interaction.

[0067] There are several ways to obtain user information across multiple acoustic ranges. For example, sensors such as pressure sensors, distance sensors, temperature sensors, or infrared sensors may be placed in the acoustic ranges, and whether a user is present in a corresponding acoustic range is determined based on the data captured by the sensors. As another example, an image of the cabin may be captured using a camera device, and the acoustic ranges in which the user is present and those in which the user is not are determined based on the image. As yet another example, the acoustic range of the vehicle cabin may be displayed using an in-vehicle display, and the user can select an acoustic range in which they are present or an acoustic range in which they are not present. As yet another example, the acoustic range of the vehicle cabin may be displayed using a GUI interface shown by a terminal device connected to the vehicle, and the user can select an acoustic range in which they are present or an acoustic range in which they are not present.

[0068] S503: Processes multiple audio based on user information across multiple acoustic ranges.

[0069] In this embodiment, audio is selected from multiple audio sources for processing based on user information across multiple acoustic ranges. Specifically, the active acoustic range is determined among the multiple acoustic ranges based on user information across multiple acoustic ranges, and then the audio in the active acoustic range is processed. In other words, portions of audio obtained by screening from multiple audio sources based on user information across multiple acoustic ranges are processed. Audio in inactive acoustic ranges may be discarded to reduce the amount of audio data that needs to be processed. This reduces the occupation of computing and memory resources.

[0070] The three cases described above, which process multiple audio based on user information across multiple acoustic ranges, will be referred to as the three processing modes below.

[0071] Mode 1: Processes only audio within the acoustic range where the user is present.

[0072] In this mode, all speech in the corresponding acoustic ranges where the user resides is processed. The active acoustic range is the one in which the user resides. For example, if the user is in one acoustic range, i.e., acoustic range 1, only the speech in acoustic range 1 is processed. If the user is in two acoustic ranges, i.e., acoustic range 1 and acoustic range 3, the speech in acoustic range 1 and acoustic range 3 is processed. The remaining ranges can be inferred by analogy. Because the signal-to-noise ratio of the user's voice captured by a microphone in the same acoustic range as the user is high, the accuracy of speech recognition can still be ensured even if the number of speeches to be processed is reduced.

[0073] Mode 2: Processes audio in the acoustic range where the user is present, and processes audio in the acoustic range where the user is not present.

[0074] In this mode, in addition to the speech in the acoustic range where the user is present, speech in the acoustic range where the user is not present is also processed, thereby reducing the number of speeches that need to be processed and ensuring the accuracy of speech recognition. In other words, the active acoustic range includes the acoustic range where the user is present, and may also include the acoustic range where the user is not present. The acoustic range within the active acoustic range where the user is not present may be an acoustic range adjacent to the acoustic range where the user is present (hereinafter referred to as the adjacent area).

[0075] Adjacent areas are acoustic areas that are adjacent to or close to each other, or acoustic areas with few or no obstacles. For example, in Figure 2a, acoustic areas 1 and 2 can be adjacent to each other, and acoustic areas 3 and 4 can be adjacent to each other. For example, in Figure 2b, acoustic areas 1 and 2 can be adjacent to each other, acoustic areas 3 and 4 can be adjacent to each other, and acoustic areas 5 and 6 can be adjacent to each other. For example, if a user is present in acoustic area 1 and no user is present in acoustic area 2, the audio in acoustic areas 1 and 2 may be processed; if a user is present in acoustic areas 1 and 3, the audio in acoustic areas 1, 2, 3, and 4 may be processed; and if a user is present in acoustic areas 1 and 5, the audio in acoustic areas 1, 2, 5, and 6 may be processed.

[0076] Optionally, for important acoustic ranges, the adjacent area range of the important acoustic range may be expanded. For example, in Figure 2a or Figure 2b, acoustic range 1 is the driver's seat area, and the adjacent area of ​​acoustic range 1 may include acoustic ranges 2 and 3, or the adjacent area of ​​acoustic range 1 may include acoustic ranges 2, 3, and 4. If the user is in acoustic range 1, the sounds of acoustic ranges 1, 2, 3, and 4 may be processed. If the user is in acoustic ranges 1 and 2, or in acoustic ranges 1 and 3, or in acoustic ranges 1 and 4, or in acoustic ranges 1, 2, and 3, or in acoustic ranges 1, 2, and 4, or in acoustic ranges 1, 3, and 4, the sounds of acoustic ranges 1, 2, 3, and 4 may be processed.

[0077] Audio from adjacent regions may be added to the processing to control the number of audio streams that need to be processed if the number of acoustic regions in which the user resides is less than the active acoustic region threshold. If the number of acoustic regions in which the user resides is equal to or greater than the active acoustic region threshold, audio from adjacent regions is not added to the processing. The active acoustic region threshold is less than the total number of acoustic regions. For example, in a 6-acoustic region voice interaction scenario, the active acoustic region threshold may be 2, 3, 4, or 5. In a 4-acoustic region voice interaction, the active acoustic region threshold may be 2 or 3. For example, in a 6-acoustic region voice interaction, the active acoustic region threshold is 4. If the number of acoustic regions in which the user resides is 3, audio from one adjacent region may be added to the processing. If the number of acoustic regions in which the user resides is 4 or 5, audio in the acoustic region in which the user resides may be processed without adding audio from adjacent regions for processing.

[0078] Optionally, if the number of acoustic ranges in which a user resides is less than the active acoustic range threshold, the number of active acoustic ranges (including the acoustic ranges in which the user resides and the portion of acoustic ranges in which the user does not reside) may be equal to the active acoustic range threshold. For example, in a 6-acoustic range voice interaction, the active acoustic range threshold is 4, and if the number of acoustic ranges in which the user resides is 1, 2, or 3, the final number of active acoustic ranges may be 4. In this way, the voice of two channels may still be reduced. During the selection of active acoustic ranges from acoustic ranges in which the user does not reside, the selection of acoustic ranges directly adjacent to the acoustic range in which the user resides, or acoustic ranges directly adjacent to the acoustic range in which the user resides and free from obstacles, is given first priority, while the selection of acoustic ranges that are not directly adjacent to the acoustic range in which the user resides but are close is given second priority. For example, in the scenario in Figure 2b, if the user resides in acoustic range 1, in addition to acoustic range 1 which is used as the active acoustic range, acoustic ranges 2, 3, and 4 are further selected as active acoustic ranges from acoustic ranges 2 to 6 in which the user does not reside. Alternatively, if a user is present in acoustic range 1 and acoustic range 6, in addition to acoustic ranges 1 and 6, which are used as active acoustic ranges, acoustic range 2, which is directly adjacent to acoustic range 1, and acoustic range 5, which is directly adjacent to acoustic range 6, are further selected as active acoustic ranges from acoustic range 2 to acoustic range 5, where no user is present.

[0079] The number of acoustic ranges k that may be selected as active acoustic ranges but do not have a user is determined by subtracting the number of acoustic ranges with users from the active acoustic range threshold. If the number of acoustic ranges l that are adjacent to acoustic ranges with users but do not have a user is greater than the number of acoustic ranges k that may be selected as active acoustic ranges but do not have a user, then k acoustic ranges may be randomly selected from the l acoustic ranges adjacent to acoustic ranges with users but no users. Alternatively, priorities may be set for multiple acoustic ranges, and the acoustic ranges adjacent to the k acoustic ranges with the highest priority among the acoustic ranges with users will be selected as active acoustic ranges. For example, in the 6-acoustic range voice interaction scenario in Figure 2b, the active acoustic range threshold is assumed to be 4, and the priority order is assumed to be acoustic range 1 > acoustic range 2 > acoustic range 3 > acoustic range 4 > acoustic range 5 > acoustic range 6. If users are present in acoustic ranges 1, 3, and 5, the number of acoustic ranges k that are not present and could be selected as active acoustic ranges is 1 (the active acoustic range threshold of 4 minus the three acoustic ranges where users are present). Acoustic range 2 is adjacent to acoustic range 1, acoustic range 4 is adjacent to acoustic range 3, and acoustic range 6 is adjacent to acoustic range 5. In other words, the number l of acoustic ranges that are not present and are adjacent to acoustic ranges where users are present is 3 (acoustic ranges 2, 4, and 6), which is greater than the number of acoustic ranges that are not present and could be selected as active acoustic ranges (1). Therefore, in addition to the three acoustic ranges used as active acoustic ranges—acoustic ranges 1, 3, and 5—one more acoustic range may be selected as an active acoustic range from the acoustic ranges where users are not present (acoustic ranges 2, 4, and 6), and since acoustic range 1 has the highest priority, acoustic range 2 may be selected as an active acoustic range.

[0080] If the number l of acoustic regions adjacent to acoustic regions where users exist, and where no users exist, is less than or equal to the number k of acoustic regions where users do not exist and which could be selected as active acoustic regions, then during the selection of active acoustic regions from acoustic regions where users do not exist, until k acoustic regions from acoustic regions where users do not exist are selected as active acoustic regions, the selection of acoustic regions directly adjacent to acoustic regions where users exist and acoustic regions directly adjacent to acoustic regions where users exist and there are no obstacles is given first priority, and the selection of acoustic regions that are not directly adjacent to but are close to acoustic regions where users exist is given second priority.

[0081] Mode 3: Processes a portion of the audio within the acoustic range where the user is present.

[0082] Mode 3 is primarily applied to scenarios where the computing platform (e.g., in-car infotainment) is hot or heavily loaded, further limiting the number of voices to be processed by processing only a portion of the voices within the acoustic range where the user is present. In this way, the computing resources occupied by voice processing are reduced, ensuring the normal use of voice interaction functions in a portion of the acoustic range even when the computing platform is heavily loaded.

[0083] In feasible implementations, processing a portion of the audio range in which the user resides may be equivalent to processing the audio within the target audio range in which the user resides. The number of target audio ranges is less than the total number of audio ranges. Target audio ranges are those with high importance and priority among all audio ranges. When computing resources on the computing platform are insufficient, priority is given to ensuring that the audio within the target audio range in which the user resides can be processed. In other words, the normal use of voice interaction functionality in the target audio range is ensured.

[0084] In a specific application scenario, such as in a vehicle, the target acoustic range is, for example, the acoustic range / multiple acoustic ranges corresponding to the driver's seat area and / or the passenger seat area, thereby ensuring that the driver's voice can be responded to preferentially. For example, the target acoustic range includes the acoustic ranges corresponding to the driver's seat area and the passenger seat area. If, based on user information, it is determined that the user is present in the acoustic ranges corresponding to the driver's seat and passenger seat, the voice in those acoustic ranges may be processed, and even if the user is in a different acoustic range, the voice in that other acoustic range will not be processed. This ensures the normal use of the voice interaction function in the acoustic ranges corresponding to the driver's seat and passenger seat by minimizing the number of voices that need to be processed.

[0085] Indeed, the target acoustic range may alternatively include only the acoustic range corresponding to the driver's seat. Alternatively, in a scenario where the vehicle has six acoustic ranges, the target acoustic range may include four acoustic ranges: the driver's seat area, the passenger seat area, the second-row left area, and the second-row right area. Alternatively, the division into target acoustic ranges may be carried out in a different way, provided that the number of target acoustic ranges is less than the total number of acoustic ranges. Examples are not listed here one by one.

[0086] In a possible implementation, processing a portion of the sound range in which the user resides may involve processing the sound from at least one sound range with the highest priority among the sound ranges in which the user resides, and discarding the sound from at least one sound range with the lowest priority among the sound ranges in which the user resides, thereby further reducing the number of sounds that need to be processed. Specifically, the sound from p sound ranges with the highest priority may be selected for processing from the sound range in which the user resides, where p is an integer greater than or equal to 1, and p is less than the number of sound ranges in which the user resides, and the sound from the remaining sound ranges in which the user resides is discarded and not processed. Alternatively, the sound from q sound ranges with the lowest priority in which the user resides is discarded and not processed, where q is an integer greater than or equal to 1, and q is less than the number of sound ranges in which the user resides, and the sound from the remaining sound ranges in which the user resides is processed.

[0087] Optionally, in scenarios where the computing platform temperature or load is high, the number of active acoustic ranges within the acoustic range where the user resides may be significantly reduced. For example, one or more acoustic ranges where the user resides that have a lower priority are excluded from the active acoustic ranges each time until the temperature or load meets the requirements. For example, in the scenario shown in Figure 2b, it is assumed that the user resides in acoustic ranges 1 through 5. If the computing resource usage of the in-car infotainment exceeds a threshold, one acoustic range is excluded from the active acoustic ranges first. For example, acoustic range 5 is excluded first, leaving acoustic ranges 1 through 4 as the remaining active acoustic ranges. In this case, the in-car infotainment processes audio from acoustic ranges 1 through 4. If the computing resource usage of the in-car infotainment is still above the threshold, one more acoustic range, for example, acoustic range 3, is excluded from the active acoustic ranges, leaving acoustic ranges 1, 2, and 4 as the remaining active acoustic ranges. In this case, the in-car infotainment processes audio from acoustic ranges 1, 2, and 4. In this case, if the computing resource usage of the in-vehicle infotainment system is below a threshold, the active acoustic ranges may remain acoustic range 1, acoustic range 2, and acoustic range 4. Indeed, if the computing resource usage of the in-vehicle infotainment system exceeds the threshold, the number of active acoustic ranges may be directly reduced to a predetermined number. This is not limited here.

[0088] If a vehicle can implement at least two of the three modes described above, the mode in which the user processes audio across multiple acoustic ranges may be determined in the following way. Here, an example is used in which the vehicle can implement the three modes described above. In practice, the three modes described above may be selected based on the temperature and / or load of the computing platform. For example, if the computing resource usage of the computing platform is below the first usage threshold, it may be determined that audio across multiple acoustic ranges will be processed in mode 2. If the computing resource usage of the computing platform is above the first usage threshold and below the second usage threshold, it may be determined that audio across multiple acoustic ranges will be processed in mode 1. If the computing resource usage of the computing platform is above the second usage threshold, it may be determined that audio across multiple acoustic ranges will be processed in mode 3. The first usage threshold is less than the second usage threshold.

[0089] Optionally, the three modes can be selected and dynamically adjusted based on the temperature or load of the computing platform. In other words, switching between the three modes may occur, allowing the number of voices to be processed to be adaptively increased or decreased. In this way, the accuracy of speech recognition can be ensured, and the computing resource utilization of multi-range voice interactions can be effectively controlled.

[0090] In another implementation, the three modes described above may be selected based on the computing platform load and the number of acoustic ranges in which the user resides. For example, if the load is below the third usage threshold and the number of acoustic ranges in which the user resides is less than the active acoustic range threshold, it is determined that audio in multiple acoustic ranges will be processed in mode 2. If the load is below the third usage threshold and the number of acoustic ranges in which the user resides is greater than or equal to the active acoustic range threshold, it is determined that audio in multiple acoustic ranges will be processed in mode 1. If the load is greater than the third usage threshold, it is determined that audio in multiple acoustic ranges will be processed in mode 3.

[0091] In an alternative implementation, the three modes described above may be selected by the user. For example, the user may be prompted to select one of the three modes using a GUI on the center console or the user's terminal device, and audio across multiple acoustic ranges may be processed in the mode selected by the user.

[0092] Indeed, one of the three modes may, as an alternative, be used by default as a mode for processing multiple audio sources.

[0093] In this embodiment, user information regarding whether the user is present in multiple acoustic ranges is acquired, and the acoustic range in which the speech should be processed is determined based on this user information, thus reducing the number of speeches that need to be processed. In this way, computing resource usage is reduced to ensure the normal operation of the multi-acoustic range voice interaction function. Furthermore, since the speech to be processed is the speech in the acoustic range in which the user is present, and the user's voice within the speech has a high signal-to-noise ratio, the accuracy of speech recognition can be guaranteed even when the number of speeches to be processed is reduced.

[0094] It can be understood that, in embodiments of the aforementioned methods, the methods and operations performed by the vehicle may, alternatively, be performed by components within the vehicle (e.g., chips, circuits, or other components). To realize the functions in the methods provided in embodiments of the present application, the vehicle may include hardware structures and / or software units, and the aforementioned functions may be realized in the form of hardware structures, software units, or a combination of hardware structures and software units. Whether any of the aforementioned functions are performed using hardware structures, software units, or a combination of hardware structures and software units depends on the specific application and design constraints of the technical solution.

[0095] Based on the same inventive concept, the present application further provides an apparatus. The apparatus may be a means of transport, an intelligent appliance, an intelligent terminal device, an Internet of Things device, etc., or a hardware module (e.g., a chip) or functional module in a means of transport, an intelligent appliance, an intelligent terminal device, or an Internet of Things device.

[0096] As shown in Figure 6, the device 600 includes an acquisition module 601 and a processing module 602. The acquisition module 601 is configured to acquire multiple audio sources, which are from multiple acoustic ranges. The acquisition module 601 is configured to acquire user information from multiple acoustic ranges, which indicates whether a user is present in an acoustic range. The processing module 602 is configured to process the multiple audio sources based on the user information from the multiple acoustic ranges.

[0097] In a possible implementation, the processing module 602 is specifically configured to process a portion of the speech obtained by screening from multiple speeches based on user information across multiple acoustic ranges.

[0098] In a possible implementation, the processing module 602 is specifically configured to process audio in the acoustic range where a user is present, based on user information in multiple acoustic ranges.

[0099] In a possible implementation, the processing module 602 is specifically configured to discard audio in acoustic ranges where no user exists, based on user information for multiple acoustic ranges.

[0100] In a possible implementation, the processing module 602 is specifically configured to process, based on user information for multiple acoustic ranges, the audio from among multiple audios that includes audio in the acoustic range where the user is present and some audio in the acoustic range where the user is not present.

[0101] In a possible implementation, the processing module 602 is configured to acquire computing resource usage. When computing resource usage exceeds a threshold, the processing module 602 is specifically configured to process the audio in the portion of the acoustic range where a user is present, based on user information for multiple acoustic ranges.

[0102] In a possible implementation, the processing module 602 is specifically configured to process the sound of the acoustic range in which the user is located, within a target acoustic range which is part of a plurality of acoustic ranges.

[0103] In a possible implementation, multiple acoustic ranges are areas corresponding to multiple seats within the vehicle's cabin, and the target acoustic range includes the driver's seat area and / or passenger seat area within the areas corresponding to multiple seats.

[0104] In possible implementations, multiple acoustic ranges have a priority order, and the processing module 602 is specifically configured to process the sound of at least one acoustic range with the highest priority among the acoustic ranges in which the user resides.

[0105] In a possible implementation, multiple acoustic zones are areas corresponding to multiple seats within the vehicle's interior, and each acoustic zone includes areas corresponding to one or more seats.

[0106] As shown in Figure 7, the present invention further provides a device, which may be a means of transport, an intelligent appliance, an intelligent terminal device, an Internet of Things device, and the like.

[0107] Device 700 includes a processor 701 and memory 702. The processor 701 is coupled to the memory 702. The processor 701 is configured to perform a data processing method according to any one of the embodiments of the above-described method based on instructions stored in the memory 702.

[0108] The present invention further provides a computer-readable storage medium that stores a computer program. When the computer program is executed by a computer, a data processing method according to one of the embodiments of the above-described method is performed.

[0109] The present invention further provides a computer program product including instructions. When the instructions are executed by an electronic device, the electronic device can perform a step of a data processing method in any one of the embodiments of the method described above.

[0110] For the convenience and in the case of a concise description, so that it may be clearly understood by those skilled in the art, the specific operating processes of the aforementioned systems, apparatus, and units should be described by referring to the corresponding processes in the embodiments of the methods described above, and further details are not described here.

[0111] It should be understood that, in some embodiments provided herein, the disclosed systems, apparatus, and methods may be implemented in other ways. For example, the embodiments of the apparatus described are merely examples. For example, the division into units is merely a logical functional division, and other divisions may be used in actual implementation. For example, multiple units or components may be coupled or integrated into other systems, or some features may be ignored or not performed. Furthermore, the mutual coupling or direct coupling or communication connection indicated or discussed may be implemented by some interface. Indirect coupling or communication connection between apparatus or units may be implemented in an electrical or other form.

[0112] Units described as separate units may or may not be physically separated, and the parts shown as units may or may not be physical units, may be located in one place, or may be distributed across multiple network units. Some or all units may be selected based on the actual requirements to achieve the objectives of the solution of the embodiment.

[0113] Furthermore, the functional units in the embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically independently, or two or more units may be integrated into a single unit. The integrated unit may be implemented in hardware form or in the form of a software functional unit.

[0114] When an integrated unit is implemented in the form of a software functional unit and sold or used as a separate product, the integrated unit may be stored on a computer-readable storage medium. Based on such understanding, all or part of the technical solutions of the present application may be implemented in the form of a software product. The computer software product is stored on a storage medium and includes several instructions that instruct a computer device (which may be a personal computer, server, network device, etc.) to perform all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage mediums include any medium capable of storing program code, such as a USB flash drive, removable hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0115] This application claims priority to Chinese Patent Application No. 202310433422.7, filed with the China National Intellectual Property Administration on April 13, 2023, with the title of the invention being "DATA PROCESSING METHOD AND RELATED DEVICE," which is incorporated herein by reference.

Claims

1. A data processing method, This involves acquiring multiple sounds, and these multiple sounds originate from multiple acoustic ranges. The process involves obtaining user information for the aforementioned multiple acoustic ranges, and the user information indicates whether or not the user is present in that acoustic range. Processing the multiple audio based on the user information of the multiple acoustic ranges. A method of having.

2. Processing the plurality of audio based on the user information of the plurality of acoustic ranges is, This includes processing a portion of the audio obtained by screening from the plurality of audio based on the user information of the plurality of acoustic ranges, The method according to claim 1.

3. Processing the plurality of audio based on the user information of the plurality of acoustic ranges is, The process includes processing the audio in the audio range where the user is present, based on the user information in the plurality of audio ranges, among the plurality of audio. The method according to claim 1 or 2.

4. Processing the plurality of audio based on the user information of the plurality of acoustic ranges is, Based on the user information for the plurality of sound ranges, the method includes discarding sound from the plurality of sounds in sound ranges where no user exists. The method according to claim 1 or 2.

5. Processing the plurality of audio based on the user information of the plurality of acoustic ranges is, The process includes processing, based on the user information of the plurality of acoustic ranges, the audio within the acoustic range where the user is present and a portion of the audio within the acoustic range where the user is not present. The method according to claim 1 or 2.

6. The method further comprises obtaining the usage status of computing resources, If the usage of the computing resources exceeds a threshold, processing the multiple audio based on the user information of the multiple acoustic ranges is: The process includes processing the audio in the portion of the acoustic range where the user is located, based on the user information of the plurality of acoustic ranges. The method according to claim 1 or 2.

7. Processing the sound in the portion of the acoustic range where the user is located is: This includes processing audio in the acoustic range where the user is located, which is part of the aforementioned plurality of acoustic ranges, within a target acoustic range that is a part of the aforementioned plurality of acoustic ranges. The method according to claim 6.

8. The plurality of acoustic areas are areas corresponding to a plurality of seats in the passenger compartment of a vehicle, and the target acoustic area includes the driver's seat area and / or passenger seat area within the areas corresponding to the plurality of seats. The method according to claim 7.

9. The aforementioned multiple acoustic ranges have a priority order, and processing the audio of the portion of the acoustic range where the user is present is: This includes processing the sound of at least one acoustic range with the highest priority among the acoustic ranges in which the user resides. The method according to claim 6.

10. It is a device, An acquisition module configured to acquire multiple audio sources, wherein the multiple audio sources are from multiple acoustic ranges, and the acquisition module, The acquisition module is configured to acquire user information for the aforementioned multiple acoustic ranges, and the user information indicates whether the user is present in the acoustic range, and the acquisition module and A processing module configured to process the plurality of audio based on the user information of the plurality of acoustic ranges. A device having.

11. The processing module is particularly configured to process a portion of the audio obtained by screening from the plurality of audio based on the user information of the plurality of acoustic ranges. The apparatus according to claim 10.

12. The processing module is particularly configured to process the audio within the audio range where the user is present, based on the user information of the plurality of audio ranges. The apparatus according to claim 10 or 11.

13. The processing module is particularly configured to discard audio from among the plurality of audio tracks that does not contain a user, based on the user information for the plurality of audio tracks. The apparatus according to claim 10 or 11.

14. The processing module is particularly configured to process, based on the user information of the plurality of acoustic ranges, the audio within the acoustic range where the user is present and the audio within a portion of the acoustic range where the user is not present among the plurality of audio. The apparatus according to claim 10 or 11.

15. The acquisition module is configured to acquire the usage of computing resources, The processing module is particularly configured to process the audio in the portion of the acoustic range where the user is located, based on the user information of the plurality of acoustic ranges, when the usage of the computing resources exceeds a threshold. The apparatus according to claim 10 or 11.

16. The processing module is specifically configured to process the sound of the acoustic range in which the user is located, which is part of the multiple acoustic ranges, within the target acoustic range. The apparatus according to claim 15.

17. The plurality of acoustic areas are areas corresponding to a plurality of seats in the passenger compartment of a vehicle, and the target acoustic area includes the driver's seat area and / or passenger seat area within the areas corresponding to the plurality of seats. The apparatus according to claim 16.

18. The aforementioned multiple acoustic ranges have a priority order. The processing module is specifically configured to process audio from at least one acoustic range with the highest priority among the acoustic ranges in which the user resides. The apparatus according to claim 15.

19. It has a processor and memory, The processor is coupled to the memory, The processor is configured to execute the data processing method described in any one of claims 1 to 9 based on instructions stored in the memory. device.

20. A computer-readable storage medium having instructions, When the computer-readable storage medium is executed on a computer, the computer can execute the data processing method described in any one of claims 1 to 9. Computer-readable storage medium.

21. A computer program product having instructions, When the aforementioned instruction is executed by the electronic device, the electronic device can perform the data processing method described in any one of claims 1 to 9. Computer program products.