Audio data processing method and apparatus, device, and storage medium
By dividing the vehicle's interior into sound zones and combining this with voiceprint feature recognition, personalized response audio data is generated, solving the problem of lack of personalization in the wake-up response of in-vehicle devices and improving the user experience and sense of technology.
Patent Information
- Application Number
- CN202111497387.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2041-12-09
AI Technical Summary
Existing in-vehicle smart devices lack personalization and a sense of technology when waking up and responding to user voice commands, resulting in a poor user experience.
By dividing the interior space of a vehicle into multiple sound zones, the sound source is identified, and personalized response audio data is generated based on the sound zone and voiceprint features, including fused keywords of sound zone features and voiceprint features, to provide differentiated response audio.
It enhances the user experience and technological feel, providing more personalized and human-like services to meet users' needs for intelligence.
Smart Images

Figure CN114120983B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of data processing, and particularly relates to the field of intelligent cockpit and Internet of Vehicles. BACKGROUND
[0002] With the development of vehicle intelligent technology, at present, almost all new vehicles launched by major vehicle manufacturers are equipped with intelligent terminals and built-in voice assistants in order to improve the sense of technology. SUMMARY
[0003] The present disclosure provides a vehicle-mounted device-based audio data processing method and device, equipment and storage medium.
[0004] According to an aspect of the present disclosure, a vehicle-mounted device-based audio data processing method is provided, comprising:
[0005] detecting that a target keyword is contained in target audio data, wherein the target keyword is used to trigger the vehicle-mounted device to enter a voice recognition state, and an internal space of a target vehicle in which the vehicle-mounted device is located is divided into at least two sound zones;
[0006] determining a target sound zone in which a sound source corresponding to the target audio data is located in the target vehicle;
[0007] determining, based on at least the target sound zone in which the sound source corresponding to the target audio data is located, reply audio data for responding to the target keyword.
[0008] According to another aspect of the present disclosure, a vehicle-mounted device-based audio data processing device is provided, comprising:
[0009] a detection unit configured to detect that a target keyword is contained in target audio data, wherein the target keyword is used to trigger the vehicle-mounted device to enter a voice recognition state, and an internal space of a target vehicle in which the vehicle-mounted device is located is divided into at least two sound zones;
[0010] a sound zone determination unit configured to determine a target sound zone in which a sound source corresponding to the target audio data is located in the target vehicle;
[0011] a reply audio determination unit configured to determine, based on at least the target sound zone in which the sound source corresponding to the target audio data is located, reply audio data for responding to the target keyword.
[0012] According to still another aspect of the present disclosure, an electronic device is provided, comprising:
[0013] at least one processor; and
[0014] a memory connected to the at least one processor in communication; wherein
[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.
[0016] According to still another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method described above.
[0017] According to still another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to the above.
[0018] According to still another aspect of the present disclosure, a vehicle-mounted device is provided, comprising:
[0019] at least one processor; and
[0020] a memory in communication connection with the at least one processor; wherein
[0021] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.
[0022] In this way, the application scheme can generate reply audio data based on the target sound area where the sound source is located, thereby increasing the experience, improving the sense of technology, and further enriching the user experience.
[0023] It should be understood that the contents described in this part are not intended to identify the key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0024] The accompanying drawings are used to better understand the present application, and do not constitute a limitation on the present disclosure. Among them:
[0025] Figure 1 is an implementation flow diagram of an audio data processing method based on a vehicle-mounted device according to an embodiment of the present disclosure;
[0026] Figure 2 is a partition result diagram in a specific example according to an embodiment of the present disclosure;
[0027] Figure 3 is a scene diagram in a specific example according to an embodiment of the present disclosure Figure 1 ;
[0028] Figure 4 is a scene diagram in a specific example according to an embodiment of the present disclosureFigure 2 ;
[0029] Figure 5 is a scenario diagram in a specific example according to an embodiment of the present disclosure Figure 3 ;
[0030] Figure 6 is a processing flow diagram in a specific example according to an embodiment of the present disclosure based on an audio data processing method of a vehicle-mounted device
[0031] Figure 7 is a structural diagram of an audio data processing apparatus based on a vehicle-mounted device according to an embodiment of the present disclosure
[0032] Figure 8 is a block diagram of an electronic device for implementing an audio data processing method based on a vehicle-mounted device according to an embodiment of the present disclosure DETAILED DESCRIPTION
[0033] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, which should be considered in a descriptive sense only. Thus, it will be apparent to one of ordinary skill in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.
[0034] The present application provides an audio data processing method based on a vehicle-mounted device. Here, the method of the present application can be specifically applied to a vehicle-mounted device arranged in a vehicle, or a server or server cluster capable of information interaction with the vehicle-mounted device. Alternatively, part of the functions are applied to the vehicle-mounted device, and the other part of the functions are applied to the server or server cluster, which is not limited by the present application.
[0035] In addition, it should be clear that the vehicle-mounted device of the present application can be integrated with functional components based on actual needs, such as integrated audio processing components (such as microphone arrays and speaker arrays), integrated touch screens, etc. At the same time, it can also be integrated with software systems and / or hardware with data processing functions, which is not limited by the present application.
[0036] The vehicle-mounted device of the present application can be understood as a broad concept, such as a device arranged in a vehicle, which can be referred to as the vehicle-mounted device of the present application.
[0037] Specifically, as shown in Figure 1 The method comprises:
[0038] Step S101: detecting that a target keyword is contained in target audio data, wherein the target keyword is used to trigger the vehicle-mounted device to enter a voice recognition state, and an internal space of a target vehicle in which the vehicle-mounted device is located is divided into at least two sound zones.
[0039] In the scheme, the target keyword can be specifically a wake-up word for the vehicle-mounted device, so as to wake up the vehicle-mounted device to perform voice recognition to obtain a user intention and perform subsequent control.
[0040] Step S102: determining a target sound zone in which a sound source corresponding to the target audio data is located in the target vehicle.
[0041] Step S103: determining, based at least on the target sound zone in which the sound source corresponding to the target audio data is located, reply audio data used to respond to the target keyword.
[0042] In the scheme, in order to distinguish the position of the sound source, a specific physical space can be divided into different sound zones. For example, as shown in Figure 2 The preset space is divided into four sound zones, namely, sound zone 1, sound zone 2, sound zone 3, and sound zone 4. At this time, a target body, such as a user, is located in the sound zone 1. When the target body outputs target audio data, the sound zone 1 in which the target body (i.e., the sound source) is located is the target sound zone.
[0043] In this way, the scheme can generate personalized reply audio data based on the target sound zone in which the sound source is located, thereby increasing the experience and improving the sense of technology, and further enriching the user experience.
[0044] In a specific example of the scheme, the reply audio data at least represents a sound zone keyword matching a sound zone feature of the target sound zone. That is, the sound zone keyword contained in the reply audio data can reflect the sound zone feature of the target sound zone, so as to further improve the sense of technology.
[0045] Here, the sound zone feature of the scheme can specifically represent a specific position of the corresponding sound zone in the internal space of the target vehicle, so as to realize personalized response based on the position.
[0046] In a specific example of the scheme, the internal space of the target vehicle is divided into a first sound zone and a second sound zone, the first sound zone matches a first sound zone keyword representing a sound zone feature of the first sound zone, the second sound zone matches a second sound zone keyword representing a sound zone feature of the second sound zone, and the first sound zone keyword is different from the second sound zone keyword.
[0047] For example, as shown in Figure 3As shown, the front row in the target vehicle is taken as the first sound area, and the second row in the target vehicle is taken as the second sound area; based on the use of the target vehicle, it is assumed that the first row, i.e., the front row, is the area where the owner of the vehicle is located, and at this time, the first sound area keyword can be specifically "owner"; after detecting that the target keyword, such as "Xiaodu Xiaodu", is contained in the target audio data in order to wake up the vehicle-mounted device, it is determined that the sound source is located in the sound area, i.e., the first sound area, and then based on the first sound area keyword of the first sound area where the sound source is located, i.e., "owner", the reply audio data at least indicating the first sound area keyword, i.e., "here, owner", is obtained.
[0048] Similarly, as shown, based on the use of the target vehicle, it is assumed that the second row is the passenger area, and at this time, the second sound area keyword can be specifically "passenger"; after detecting that the target keyword, such as "Xiaodu Xiaodu", is contained in the target audio data in order to wake up the vehicle-mounted device, it is determined that the sound source is located in the sound area, i.e., the second sound area, and then based on the second sound area keyword of the second sound area where the sound source is located, i.e., "passenger", the reply audio data at least indicating the first sound area keyword, i.e., "here, passenger", is obtained. Figure 4
[0049] In this way, the sound area keywords corresponding to different sound areas are different in the scheme of the present application, so as to differentiate the responses to audio data from different sound areas, greatly improving the user experience.
[0050] In a specific example of the scheme of the present application, the sound areas can be further divided based on the number of seats in the target vehicle, so as to further refine the user experience and further enrich and improve the user experience; specifically, the internal space of the target vehicle is divided into at least four sound areas, which are: a main driver sound area, a co-driver sound area, a main driver rear side sound area, and a co-driver rear side sound area.
[0051] The main driver sound area matches a third sound area keyword representing the sound area characteristics of the main driver sound area, the co-driver sound area matches a fourth sound area keyword representing the sound area characteristics of the co-driver sound area, the main driver rear side sound area matches a fifth sound area keyword representing the sound area characteristics of the main driver rear side sound area, and the co-driver rear side sound area matches a sixth sound area keyword representing the sound area characteristics of the co-driver rear side sound area.
[0052] Among them, the third sound area keyword, the fourth sound area keyword, the fifth sound area keyword and the sixth sound area keyword are different from each other, that is, the sound area keywords of the four sound areas are all different.
[0053] For example, as shown, Figure 5 As shown, the interior space of the target vehicle is divided into four sound zones, namely the main driver sound zone, the co-driver sound zone, the main driver rear side sound zone, and the co-driver rear side sound zone. At this time, different sound zone keywords can be set for different sound zones based on the physical positions of the four sound zones. For example, the sound zone keyword of the co-driver sound zone is set as "co-driver". After detecting that the target keyword, such as "Xiaodu Xiaodu", is contained in the target audio data to wake up the vehicle-mounted device, it is determined that the sound source is in the co-driver sound zone. Then, based on the sound zone keyword of the co-driver sound zone where the sound source is located, that is, "co-driver", the reply audio data at least representing the sound zone keyword of the co-driver sound zone, that is, "here, co-driver passenger", is obtained.
[0054] Alternatively, in order to further distinguish the main driver from other passengers, the third sound zone keyword of the main driver sound zone can be different from the sound zone keywords of the other three sound zones. Whether the sound zone keywords of the other three sound zones are the same or not can be not limited, that is, they can be the same, different, or partially the same and partially different, and the like. In this way, the main driver is further highlighted and the user experience is refined. That is, the third sound zone keyword is different from the fourth sound zone keyword, the fifth sound zone keyword, and the sixth sound zone keyword, and the fourth sound zone keyword, the fifth sound zone keyword, and the sixth sound zone keyword are the same or not.
[0055] It should be noted that based on the above specific example, the reply audio data not only represents the sound zone feature of the sound zone, but also contains other greetings. The greetings can be set based on actual needs, and the present application scheme does not limit this.
[0056] In a specific example of the present application scheme, the above-mentioned at least based on the target audio data corresponding to the sound source in the target sound zone to determine the reply audio data for responding to the target keyword includes: at least based on the target audio data corresponding to the sound source in the target sound zone, determining the sound zone keyword matched with the sound zone feature of the target sound zone; determining the reply audio data for responding to the target keyword, wherein the reply audio data at least represents the sound zone keyword matched with the sound zone feature of the target sound zone. For example, a mapping table can be pre-set, which records the mapping relationship between the sound zone and the sound zone keyword. Then, based on the mapping table, the sound zone keyword corresponding to the target sound zone can be obtained after determining that the target audio data corresponds to the sound source in the target sound zone. This way is simple and efficient.
[0057] In a specific example of the solution, in order to further refine the user experience, enrich and improve the user experience, and provide more personalized and differentiated services, the solution can also determine the reply audio data based on the voiceprint feature. Specifically, the voiceprint feature of the target audio data is obtained, and then the reply audio data for responding to the target keyword is determined based on the target sound area where the sound source corresponding to the target audio data is located, which can include: determining the reply audio data for responding to the target keyword based on the target sound area where the sound source corresponding to the target audio data is located and the voiceprint feature of the target audio data. That is, in the process of determining the reply audio data, not only the dimension of the sound area is considered, but also the dimension of the voiceprint feature is considered, so as to provide more personalized and differentiated services and further improve the user experience and the sense of technology.
[0058] It should be noted that the user can choose whether to perform voiceprint recognition based on actual needs, or both voiceprint recognition and sound area recognition are functions of the vehicle-mounted device, but whether to need and which function to need are determined by the user based on personal preferences. Alternatively, the voiceprint recognition and sound area recognition functions are set before the device is shipped without manual intervention by the user. The solution does not limit this, and the actual needs of the final product can be determined.
[0059] In a specific example of the solution, the reply audio data at least represents a fusion keyword matching the sound area feature of the target sound area and the voiceprint feature of the target audio data. That is, the sound area feature of the target sound area and the voiceprint feature of the target audio data are considered based on the fusion keyword, so as to improve the feasibility of the solution.
[0060] In a specific example of the solution, the fusion keyword includes a sound area keyword matching the sound area feature of the target sound area and a type keyword matching the voiceprint feature of the target audio data. In other words, the fusion keyword includes two types of keywords, one type of keyword representing the sound area feature of the sound area keyword, and the other type of keyword representing the voiceprint feature of the type keyword. Here, the sound area keyword representing the sound area feature can refer to the description above, which is not repeated here. For the type keyword, it can correspond to gender, age, etc. That is, the solution classifies the sound source based on the voiceprint feature, such as male, female, adult, or child.
[0061] For example, when the target audio data corresponds to a female based on the voiceprint feature analysis, the type keyword can be specifically "goddess"; for another example, when the target audio data corresponds to a child based on the voiceprint feature analysis, the type keyword can be specifically "child". Correspondingly, in combination with the above sound area keyword, the fusion keyword can be specifically "main driver goddess", and the reply audio data can be specifically "yes, main driver goddess". Or, the fusion keyword can be specifically "child passenger", and the reply audio data can be specifically "yes, child passenger". In this way, more personalized services are further provided, and user experience is improved.
[0062] It should be noted that in actual application, the fusion keywords corresponding to the same sound area but different voiceprint features (corresponding to different categories) are different; the fusion keywords corresponding to the same voiceprint feature (corresponding to the same category, for example, both are female) but different sound areas are different; and the fusion keywords corresponding to the same sound area and the same voiceprint feature (corresponding to the same category) are the same.
[0063] In this way, the sound area feature and the voiceprint feature are embodied by the fusion keyword, and user experience is further improved. Moreover, the scheme is simple and feasible, and lays a foundation for engineering promotion and application.
[0064] In a specific example of the scheme, considering that there can be multiple audios in actual scenarios, in order to meet the scene requirements, the detection of the target audio data containing the target keyword can specifically include: obtaining multiple audio data from the internal space of the target vehicle; for example, there are multiple people speaking in the target vehicle, and the target audio data containing the target keyword can be identified from the multiple audio data. In other words, whether the multiple audio data contains the target keyword for waking up the vehicle-mounted device is identified, and if so, the target audio data corresponding to the target keyword is identified. In this way, actual requirements are met, user experience is further improved, and processing accuracy is also improved.
[0065] In a specific example of the scheme, after the reply audio data is determined, the reply audio data can be output, and the reply audio data is played to prompt the user that the vehicle-mounted device enters the voice recognition state and can be controlled subsequently, such as playing a song, opening a map and displaying on a vehicle display screen (i.e., a central control screen of the vehicle), or starting a seat heating function, etc. In this way, the user's demand for intelligence is met, and user experience is also improved.
[0066] In another specific example, outputting the reply audio data can specifically be displaying the reply audio data, in other words, the reply audio data can be output in the form of audio playing or in the form of screen display (such as display on a vehicle display screen, etc.), or one of the two, so as to meet different needs of users.
[0067] In this way, the scheme can generate reply audio data based on the target sound area where the sound source is located, or can also generate reply audio data based on the voiceprint features of the sound source, thereby increasing the experience, improving the sense of technology, further enriching the user experience, and meeting the user's demand for intelligence.
[0068] The scheme will be further described in detail below in combination with specific examples. Specifically, the examples aim to propose a more personalized voice assistant wake-up response language scheme based on a vehicle device, such as a vehicle voice device. After the vehicle is woken up, the vehicle voice device can respond in a personalized manner according to the sound area where the wake-up sound source is located, for example, "here, main driver" or "here, rear passenger". Or, the vehicle voice device can also provide a personalized response after combining the sound area and the voiceprint features of the sound source, for example, "here, main driver goddess" or "here, rear child". In this way, compared with the existing default or random response mode, the scheme has a stronger interactive experience, can provide a sense of technology, and meets the user's personalized needs.
[0069] Specifically, as shown in Figure 6 After obtaining the target audio data containing the wake-up word, the vehicle device is woken up. At this time, the wake-up engine of the vehicle device feeds back the sound area where the target audio data is located or feeds back the sound area and the voiceprint features. Here, the function can be selected by the user. The processing result (such as the sound area where the target audio data is located or the sound area and the voiceprint features) is sent to the logical processing module of the vehicle device to obtain a welcome language such as "hello, main driver" or "hello, main driver goddess", and then sent to the broadcast engine of the vehicle device through the logical processing template of the vehicle device to be played by the broadcast engine; and sent to the UI (User Interface) display template of the vehicle device for display.
[0070] In practical applications, the voiceprint recognition and sound area recognition functions can be configured as user-selectable functions, and then the user's selection result is used to determine which function to execute. Similarly, the playing and screen display functions can also be configured as user-selectable functions, which are not limited by the examples.
[0071] In this way, the scheme can generate reply audio data based on the target sound area where the sound source is located, or can also generate reply audio data based on the voiceprint features of the sound source, thereby increasing the experience, improving the sense of technology, further enriching the user experience, and meeting the user's demand for intelligence.
[0072] The scheme also provides an audio data processing device based on a vehicle-mounted device. It should be noted that the device of the scheme can be integrated into the vehicle-mounted device, or integrated into a server or server cluster in communication with the vehicle-mounted device, or part of the functions are in the vehicle-mounted device and the other part of the functions are in the server or server cluster, and the scheme does not limit this.
[0073] Specifically, as shown in Figure 7 , comprising:
[0074] The detection unit 701 is configured to detect that the target audio data contains a target keyword, wherein the target keyword is used to trigger the vehicle-mounted device to enter a voice recognition state, and the internal space of the target vehicle where the vehicle-mounted device is located is divided into at least two sound areas.
[0075] The sound area determination unit 702 is configured to determine a target sound area where a sound source corresponding to the target audio data is located in the target vehicle.
[0076] The reply audio determination unit 703 is configured to determine, based on at least the target sound area where the sound source corresponding to the target audio data is located, reply audio data for responding to the target keyword.
[0077] In a specific example of the scheme, the reply audio data at least represents a sound area keyword matching the sound area features of the target sound area.
[0078] In a specific example of the scheme, the internal space of the target vehicle is divided into a first sound area and a second sound area, the first sound area matches a first sound area keyword representing the sound area features of the first sound area, the second sound area matches a second sound area keyword representing the sound area features of the second sound area, and the first sound area keyword is different from the second sound area keyword.
[0079] In a specific example of the scheme, the internal space of the target vehicle is divided into at least four sound areas, which are: a main driver sound area, a co-driver sound area, a main driver rear sound area, and a co-driver rear sound area.
[0080] The main driver audio zone matches a third audio zone keyword representing an audio zone feature of the main driver audio zone, the co-driver audio zone matches a fourth audio zone keyword representing an audio zone feature of the co-driver audio zone, the main driver rear-side audio zone matches a fifth audio zone keyword representing an audio zone feature of the main driver rear-side audio zone, and the co-driver rear-side audio zone matches a sixth audio zone keyword representing an audio zone feature of the co-driver rear-side audio zone;
[0081] The third audio zone keyword, the fourth audio zone keyword, the fifth audio zone keyword, and the sixth audio zone keyword are different from each other; or,
[0082] The third audio zone keyword is different from the fourth audio zone keyword, the fifth audio zone keyword, and the sixth audio zone keyword, and the fourth audio zone keyword, the fifth audio zone keyword, and the sixth audio zone keyword are the same or different.
[0083] In a specific example of the application scheme, the reply audio determination unit is specifically configured to determine an audio zone keyword matching an audio zone feature of a target audio zone in which a sound source corresponding to the target audio data is located based on at least the target audio data, and determine reply audio data for responding to the target keyword, wherein the reply audio data at least represents the audio zone keyword matching the audio zone feature of the target audio zone.
[0084] In a specific example of the application scheme, further comprising: a voiceprint detection unit; wherein,
[0085] The voiceprint detection unit is configured to obtain a voiceprint feature of the target audio data.
[0086] The reply audio determination unit is specifically configured to determine reply audio data for responding to the target keyword based on a target audio zone in which a sound source corresponding to the target audio data is located and a voiceprint feature of the target audio data.
[0087] In a specific example of the application scheme, the reply audio data at least represents a fusion keyword matching both the audio zone feature of the target audio zone and the voiceprint feature of the target audio data.
[0088] In a specific example of the application scheme, the fusion keyword includes an audio zone keyword matching the audio zone feature of the target audio zone and a type keyword matching the voiceprint feature of the target audio data.
[0089] In a specific example of the application scheme, the detection unit is specifically configured to obtain a plurality of audio data from an internal space of the target vehicle, and identify target audio data containing the target keyword from the plurality of audio data.
[0090] In a specific example of the solution of the present application, further comprising:
[0091] The output unit is configured to output the reply audio data to prompt the user to enter a voice recognition state.
[0092] The specific functions of each unit in the above device can refer to the description of the above method, which will not be repeated here.
[0093] In the technical solution of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0094] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a vehicle-mounted device, a readable storage medium and a computer program product.
[0095] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit implementations of the present disclosure described and / or claimed in this document.
[0096] It should be noted that the specific structure of the vehicle-mounted device can refer to the electronic device described below, and will not be repeated here.
[0097] As shown in Figure 8 The device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0098] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0099] The computing unit 801 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs various methods and processes described above, such as the in-vehicle device-based audio data processing method. For example, in some embodiments, the in-vehicle device-based audio data processing method can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the in-vehicle device-based audio data processing method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the in-vehicle device-based audio data processing method by any other appropriate means, such as by means of firmware.
[0100] Various implementations of the systems and techniques described above herein can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0101] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.
[0102] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0103] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0104] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0105] The computer system can include clients and servers. This relationship can be. The servers are typically remote from the clients with the interactions between them occurring over a communication network. The relationship between a client and a server is one of client-server. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0106] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure are achieved, and the present disclosure is not limited herein.
[0107] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.
Claims
1. A vehicle-mounted device-based audio data processing method, comprising: detecting that a target keyword is contained in target audio data, wherein the target keyword is used to trigger the vehicle-mounted device to enter a voice recognition state, and an internal space of a target vehicle in which the vehicle-mounted device is located is divided into at least two sound zones; determining a target sound zone in which a sound source corresponding to the target audio data is located in the target vehicle; determining, based at least on the target sound zone in which the sound source corresponding to the target audio data is located, reply audio data for responding to the target keyword; the reply audio data at least represents a fusion keyword matching a sound zone feature of the target sound zone and a voiceprint feature of the target audio data, the fusion keyword containing a sound zone keyword matching the sound zone feature of the target sound zone and a type keyword matching the voiceprint feature of the target audio data, so as to differentiate responses to audio data in different sound zones through the reply audio data; wherein, in a case where the internal space of the target vehicle is divided into at least four sound zones, the four sound zones are a main driver sound zone, a co-driver sound zone, a main driver rear side sound zone, and a co-driver rear side sound zone; the main driver sound zone matches a third sound zone keyword representing a sound zone feature of the main driver sound zone, the co-driver sound zone matches a fourth sound zone keyword representing a sound zone feature of the co-driver sound zone, the main driver rear side sound zone matches a fifth sound zone keyword representing a sound zone feature of the main driver rear side sound zone, and the co-driver rear side sound zone matches a sixth sound zone keyword representing a sound zone feature of the co-driver rear side sound zone; wherein the third sound zone keyword, the fourth sound zone keyword, the fifth sound zone keyword, and the sixth sound zone keyword are different from each other; or the third sound zone keyword is different from the fourth sound zone keyword, the fifth sound zone keyword, and the sixth sound zone keyword, and the fourth sound zone keyword, the fifth sound zone keyword, and the sixth sound zone keyword are the same or different.
2. The method of claim 1, wherein, in a case where the internal space of the target vehicle is divided into a first sound zone and a second sound zone, the first sound zone matches a first sound zone keyword representing a sound zone feature of the first sound zone, and the second sound zone matches a second sound zone keyword representing a sound zone feature of the second sound zone, and the first sound zone keyword is different from the second sound zone keyword.
3. The method of claim 1 or 2, wherein, The determination of the reply audio data for responding to the target keyword based at least on the target sound zone in which the sound source corresponding to the target audio data is located comprises: determining, based at least on the target sound zone in which the sound source corresponding to the target audio data is located, a sound zone keyword matching a sound zone feature of the target sound zone; determining the reply audio data for responding to the target keyword, wherein the reply audio data at least represents the sound zone keyword matching the sound zone feature of the target sound zone.
4. The method of claim 1, further comprising: obtaining a voiceprint feature of the target audio data; wherein the determination of the reply audio data for responding to the target keyword based at least on the target sound zone in which the sound source corresponding to the target audio data is located comprises: determine, based on the target audio data corresponding to a target sound source and a voiceprint feature of the target audio data, reply audio data for responding to the target keyword.
5. The method of claim 1 or 2 or 4, wherein, The detection of the target keyword in the target audio data includes: obtaining a plurality of audio data from an internal space of the target vehicle; identifying target audio data containing the target keyword from the plurality of audio data.
6. The method of claim 1 or 2 or 4, further comprising: outputting the reply audio data to prompt a user to enter a voice recognition state of the vehicle-mounted device.
7. An audio data processing apparatus based on a vehicle-mounted device, comprising: a detection unit configured to detect a target keyword in target audio data, wherein the target keyword is used to trigger the vehicle-mounted device to enter a voice recognition state, and an internal space of a target vehicle in which the vehicle-mounted device is located is divided into at least two sound zones; a sound zone determination unit configured to determine a target sound zone in which a sound source corresponding to the target audio data is located in the target vehicle; a reply audio determination unit configured to determine, based on at least the target sound zone in which the sound source corresponding to the target audio data is located, reply audio data for responding to the target keyword; the reply audio data at least represents a fusion keyword matching the sound zone feature of the target sound zone and the voiceprint feature of the target audio data, the fusion keyword includes a sound zone keyword matching the sound zone feature of the target sound zone, and a type keyword matching the voiceprint feature of the target audio data, so as to differentiate the audio data of different sound zones through the reply audio data; wherein, in the case that the internal space of the target vehicle is divided into at least four sound zones, the four sound zones are: a main driver sound zone, a co-driver sound zone, a main driver rear side sound zone, and a co-driver rear side sound zone; the main driver sound zone matches a third sound zone keyword representing the sound zone feature of the main driver sound zone, the co-driver sound zone matches a fourth sound zone keyword representing the sound zone feature of the co-driver sound zone, the main driver rear side sound zone matches a fifth sound zone keyword representing the sound zone feature of the main driver rear side sound zone, and the co-driver rear side sound zone matches a sixth sound zone keyword representing the sound zone feature of the co-driver rear side sound zone; wherein, the third sound zone keyword, the fourth sound zone keyword, the fifth sound zone keyword, and the sixth sound zone keyword are different from each other; or the third sound zone keyword is different from the fourth sound zone keyword, the fifth sound zone keyword, and the sixth sound zone keyword, and the fourth sound zone keyword, the fifth sound zone keyword, and the sixth sound zone keyword are the same or different.
8. The apparatus of claim 7, wherein, In the case that the internal space of the target vehicle is divided into a first sound zone and a second sound zone, the first sound zone matches a first sound zone keyword representing the sound zone feature of the first sound zone, and the second sound zone matches a second sound zone keyword representing the sound zone feature of the second sound zone, and the first sound zone keyword is different from the second sound zone keyword.
9. The apparatus of claim 7 or 8, wherein, The reply audio determination unit is specifically configured to determine an audio area keyword matched with an audio area characteristic of the target audio area based on at least the target audio data corresponding to a target sound source in the target audio area, and determine reply audio data for responding to the target keyword, wherein the reply audio data at least represents the audio area keyword matched with the audio area characteristic of the target audio area.
10. The apparatus of claim 7, further comprising: The voiceprint detection unit is configured to obtain a voiceprint feature of the target audio data. The voiceprint detection unit is configured to obtain a voiceprint feature of the target audio data. The reply audio determination unit is specifically configured to determine an audio area keyword matched with an audio area characteristic of the target audio area based on at least the target audio data corresponding to a target sound source in the target audio area, and determine reply audio data for responding to the target keyword, wherein the reply audio data at least represents the audio area keyword matched with the audio area characteristic of the target audio area.
11. The apparatus of claim 7 or 8 or 10, wherein, The detection unit is specifically configured to obtain a plurality of audio data from an internal space of the target vehicle, and identify target audio data containing the target keyword from the plurality of audio data.
12. The apparatus of claim 7 or 8 or 10, further comprising: The output unit is configured to output the reply audio data to prompt the user to enter a voice recognition state of the vehicle-mounted device.
13. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
14. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-6.
15. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-6.
16. An in-vehicle device characterized by comprising: comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
Citation Information
Patent Citations
Vehicle-mounted robot control method and device, vehicle, electronic equipment and medium
CN112026790A
Multi-voice-register voice wake-up method and device, multi-voice-register voice recognition method and device, equipment and storage medium
CN113380247A