Speech Information Recognition Method, Apparatus, Electronic Device, and Storage Medium
By combining sound and visual information to determine the location of the target person in the vehicle, and controlling the voice assistant to only receive voice in the target sound area, the problem of inaccurate recognition of the midrange area in the prior art is solved, and the smoothness of voice interaction is improved.
Patent Information
- Application Number
- CN202311041528.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-08-17
AI Technical Summary
Inaccurate midrange recognition of prior art leads to unsmooth voice interaction, especially when discussing between noisy environments and passengers, which affects the accuracy of rear voice recognition.
By combining the sound dimensions and visual dimensions, the location of the target person in the car is determined. The specific steps include: receiving voice wake-up instructions, determining the initial position of the target person through the sound information; acquiring image information in the vehicle, determining the second position of the target person through the image information; determining the target sound area based on the two; controlling the voice assistant to receive and recognize only the voice information in the target sound area.
It improves the accuracy of the target sound area, reduces recognition errors, ensures that the voice assistant can accurately receive and recognize the target person's instructions, and improves the smoothness of voice interaction in the car.
Smart Images

Figure CN117153160B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technology, and in particular, to a method, device, electronic device, and storage medium for identifying voice information. Background Art
[0002] Intelligent cockpit voice interaction mainly includes four parts: wake-up, listening, understanding, and broadcasting. When waking up the in-vehicle voice assistant, it is also necessary to locate the sound source. Currently, the voice zone recognition technology on the market mainly uses independent pick-up microphones installed beside each seat, and determines the position of the sound source by comparing the sound intensity of the sound received by each pick-up microphone. However, this method of determining the position of the sound source by the sound volume is not applicable in some scenarios. For example, when a passenger in the left rear position turns to the right and issues a wake-up word to wake up the voice assistant, the vehicle computer may determine that the passenger in the right rear wakes up the voice assistant, and thus choose not to accept the subsequent instructions issued by the passenger in the left rear, resulting in incorrect recognition of the rear voice zone. In addition, when the noise inside and outside the vehicle is relatively large, the discussion and echo among passengers will affect the pick-up quality of voice interaction and the accuracy of rear voice zone recognition.
[0003] Therefore, there is a problem in the prior art that inaccurate voice zone recognition leads to unsmooth voice interaction. Summary of the Invention
[0004] In view of this, embodiments of this application provide a method, device, electronic device, and storage medium for identifying voice information to solve the problem in the prior art that inaccurate voice zone recognition leads to unsmooth voice interaction.
[0005] In a first aspect of the embodiments of this application, a method for identifying voice information is provided, including:
[0006] When receiving a voice wake-up instruction for a voice assistant, determining first position information of a first target person in a vehicle according to the voice wake-up instruction, where the first target person is the person who issues the voice wake-up instruction;
[0007] Obtaining image information in the vehicle and determining second position information of the first target person in the vehicle according to the image information;
[0008] Determining a first target voice zone of the first target person in the vehicle according to the first position information and the second position information;
[0009] Controlling the voice assistant to only receive and recognize voice information in the first target voice zone according to the first target voice zone.
[0010] In a second aspect of the embodiments of the present application, there is provided an apparatus for recognizing voice information, including:
[0011] A first determination module, configured to determine first position information of a first target person in a vehicle according to the voice wake-up instruction when receiving the voice wake-up instruction for a voice assistant, where the first target person is the person who issues the voice wake-up instruction;
[0012] A second determination module, configured to obtain image information in the vehicle and determine second position information of the first target person in the vehicle according to the image information;
[0013] A third determination module, configured to determine a first target sound area of the first target person in the vehicle according to the first position information and the second position information;
[0014] A voice recognition module, configured to control the voice assistant to only receive and recognize voice information within the first target sound area according to the first target sound area.
[0015] In a third aspect of the embodiments of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the steps of the above method when executing the computer program.
[0016] In a fourth aspect of the embodiments of the present application, there is provided a readable storage medium storing a computer program, where the computer program implements the steps of the above method when executed by a processor.
[0017] The beneficial effects of the embodiments of the present application compared with the prior art are:
[0018] In this embodiment, when a voice wake-up instruction for the voice assistant is received, the first position information of the first target person who issues the voice wake-up instruction in the vehicle is determined according to the voice wake-up instruction, and the second position information of the first target person in the vehicle is determined according to the image information. The first target sound area of the first target person in the vehicle is determined according to the first position information and the second position information, and then the voice assistant is controlled to only receive and recognize the voice information in the first target sound area. It realizes jointly determining the first target sound area of the first target person in the vehicle by combining the sound dimension and the visual dimension, improves the accuracy of the determined first target sound area, reduces the occurrence of scenarios where the target sound area is misrecognized. In addition, controlling the voice assistant to only receive and recognize the voice information in the first target sound area avoids the interference of the sounds in other sound areas on the voice recognition in the first target sound area, ensures that a series of subsequent instructions issued by the first target person who issues the voice wake-up instruction can be accurately received and recognized by the voice assistant, improves the smoothness of voice interaction among the people in the vehicle, and solves the problem of unsmooth voice interaction caused by inaccurate sound area recognition in the prior art. Brief Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic flowchart of a method for recognizing voice information provided by an embodiment of the present application;
[0021] Figure 2 It is a schematic flowchart of another method for recognizing voice information provided by an embodiment of the present application;
[0022] Figure 3 It is a schematic structural diagram of a device for recognizing voice information provided by an embodiment of the present application;
[0023] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Description of the Embodiments
[0024] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0025] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.
[0026] In addition, it should be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the presence of additional identical elements in the process, method, article or device including the elements.
[0027] A method and device for recognizing voice information according to an embodiment of this application will be described in detail below with reference to the accompanying drawings.
[0028] Figure 1 It is a schematic flowchart of a method for recognizing voice information provided by an embodiment of this application. As Figure 1 shown, the method for recognizing voice information includes:
[0029] Step 101, when receiving a voice wake-up instruction for a voice assistant, determine the first position information of the first target person in the vehicle according to the voice wake-up instruction.
[0030] The first target person is the person who issues the voice wake-up instruction.
[0031] Specifically, the voice wake-up instruction is generally a pre-set voice instruction, such as "Hello, Xiao Zhu" or "Hi, Xiao Zhu", etc.
[0032] When a person in the vehicle makes a sound, the pick-up microphone in the vehicle can receive the sound information and calculate the position of the person making the sound in the vehicle. In this embodiment, if it is detected that the user issues a voice wake-up instruction, the vehicle can receive the voice wake-up instruction and determine the first position information of the first target person who issues the voice wake-up instruction in the vehicle according to the voice wake-up instruction.
[0033] It should be noted that since the front row seats of the vehicle only include the driver's seat and the passenger seat, it is not easy to have misidentification when recognizing the sound area only through voice information. Therefore, in this embodiment, the process described in this embodiment can be executed after it is determined that there are people in the rear row seats of the vehicle.
[0034] In this embodiment, the first position information of the first target person who issues the voice wake-up instruction in the vehicle is determined through voice information, realizing the recognition of the position of the first target person from the voice dimension.
[0035] Step 102, obtain the image information in the vehicle, and determine the second position information of the first target person in the vehicle according to the image information.
[0036] Specifically, in this embodiment, an Occupancy Monitoring System (OMS) camera can be arranged at the roof of the vehicle on the same plane as the front row seats of the vehicle, and the image information in the vehicle is obtained through the OMS camera. Generally, OMS cameras are set in the front row of the vehicle, but the OMS cameras located in the front row of the vehicle may be blocked by the front row seats or the passengers in the passenger seat and the driver in the driver's seat. Therefore, by arranging an OMS camera at the roof of the vehicle on the same plane as the front row seats, it is ensured that the image information of the passengers in the rear row can be obtained by the vehicle, so that the position information of each person can be determined according to the image information. Of course, it should be noted that this embodiment does not specifically limit the installation position of the camera, as long as the captured image information includes the information of all the people in the vehicle.
[0037] It should be noted that when the voice pickup microphone in the vehicle monitors that a passenger issues a voice wake-up instruction, the vehicle will record the time period when the voice wake-up instruction is issued. And since the vehicle has been detecting the behavior of the passengers through the OMS camera, the image data corresponding to the time period when the first target person issues the voice wake-up instruction, that is, the image information, can be retrieved. The image information includes the face images and position information of all the people in the vehicle when the first target person issues the voice wake-up instruction, so that the second position information of the first target person in the vehicle can be determined according to the image information.
[0038] By obtaining the image information in the vehicle, the position of the first target person who issues the voice wake-up instruction is determined from the visual dimension.
[0039] Step 103, determine the first target sound area of the first target person in the vehicle according to the first position information and the second position information.
[0040] Specifically, the sound zone refers to a preset sound source area inside the vehicle, and the goal of in-vehicle sound source localization is to accurately distinguish the sound zones. For example, in one example, the sound zones inside the vehicle can be divided into 4 along the central axis of the vehicle, namely the area where the driver's seat is located, the area where the co-driver's seat is located, the area of the left half of the rear seats in the vehicle, and the area of the right half of the rear seats in the vehicle. Another example, in one example, if there are three rows of seats in the vehicle, the sound zones inside the vehicle can be divided into 6 along the central axis of the vehicle, and each sound zone will not be described one by one here.
[0041] In this embodiment, the first target sound zone of the first target person is determined according to the first position information and the second position information. Since the first position information determines the position of the first target person from the sound dimension, and the second position information determines the position of the first target person from the visual dimension, the first position information and the second position information can be combined to jointly determine the first target sound zone of the first target person inside the vehicle. And compared with determining the position of the first target person only from the sound dimension, the accuracy of the determined first target sound zone is improved.
[0042] Step 104, according to the first target sound zone, control the voice assistant to only receive and recognize the voice information within the first target sound zone.
[0043] Specifically, after determining the first target sound zone, the voice assistant can be controlled to only receive the voice information within the first target sound zone and perform recognition. This avoids the interference of the sounds in other sound zones on the voice recognition within the first target sound zone, ensuring that a series of subsequent instructions issued by the first target person who issues the voice wake-up command can be accurately received and recognized by the voice assistant.
[0044] In this way, this embodiment realizes jointly determining the first target sound zone of the first target person inside the vehicle by combining the sound dimension and the visual dimension, improves the accuracy of the first target sound zone, reduces the occurrence of scenarios where the target sound zone is misrecognized. In addition, controlling the voice assistant to only receive and recognize the voice information within the first target sound zone avoids the interference of the sounds in other sound zones on the voice recognition within the first target sound zone, ensures that a series of subsequent instructions issued by the first target person who issues the voice wake-up command can be accurately received and recognized by the voice assistant, improves the smoothness of voice interaction among the people inside the vehicle, and solves the problem in the prior art that inaccurate sound zone recognition leads to unsmooth voice interaction.
[0045] In some embodiments, the image information includes the face images of each person inside the vehicle and the seating information of each person; the determining the second position information of the first target person inside the vehicle according to the image information includes:
[0046] Based on the face images of the respective persons, obtain the facial features of the respective persons, where the facial features include facial features and mouth features; determine the first target person according to the facial features of the respective persons and the preset facial features of the persons stored in advance when outputting the voice wake-up instruction; determine the second position information of the first target person in the vehicle according to the seat information of the respective persons.
[0047] Specifically, when a person speaks, they will mobilize the facial muscles, and the facial features and mouth shapes are different when uttering different words. For example, when a person says "Xiaodu" and "Xiaoan", there are obvious differences in the facial features and mouth features. In this embodiment, the preset facial features of the persons when outputting the voice wake-up instruction can be stored in advance in the preset database, and the preset facial features include facial features and mouth features, to be used as a basis for determining whether there is a person issuing the voice wake-up instruction.
[0048] It should be noted that the voice wake-up instruction can be the default wake-up word of the voice assistant or a user-defined wake-up word, and this embodiment does not limit this.
[0049] In addition, the obtained image information includes the face images of the respective persons in the vehicle. At this time, the facial features corresponding to the respective persons can be analyzed from the face images, and the facial features corresponding to the respective persons are compared with the preset facial features of the persons stored in advance when outputting the voice wake-up instruction to determine which person is the first target person issuing the voice wake-up instruction.
[0050] It should also be noted that in order to ensure the accuracy of determining the first target person, the preset facial features can be the facial features of the same person when outputting the voice wake-up instruction, that is, the facial features corresponding to the respective persons are compared with the preset facial features of this person when outputting the voice wake-up instruction to determine which person is the first target person issuing the voice wake-up instruction.
[0051] In addition, the image information also includes the seating positions of the respective persons, which enables the seating position of the first target person to be found through the seating position information of the respective persons, so as to determine the second position information of the first target person in the vehicle.
[0052] In this way, determining the first target person through the facial features of the person ensures the accuracy of identifying the first target person.
[0053] In some embodiments, when determining the first target person according to the facial features of each person and the preset facial features of the person stored in advance when the voice wake-up instruction is output, the similarity between the facial features of each person and the preset facial features can be determined; in the case where at least one target similarity is greater than a preset value, the person corresponding to each target similarity is determined as the first target person.
[0054] Specifically, when calculating the similarity between the facial features of each person and the preset facial features, it can be calculated by means of cosine distance or Euclidean distance.
[0055] In addition, the preset value can be determined according to the accuracy requirement for the sound area. For example, if the accuracy requirement for the sound area is relatively high, the value of the preset value can be set larger. For example, the preset value can be 95%; if the accuracy requirement for the sound area is not very high, the value of the preset value can be set smaller. For example, the preset value can be set to 85%.
[0056] If only one similarity in the calculated similarities is greater than the set preset value, the person corresponding to this similarity can be determined as the first target person; if at least two similarities in the calculated similarities are greater than the set preset value, the persons corresponding to these at least two similarities can be determined as the first target persons.
[0057] In this way, by calculating the similarity between the facial features of each person and the preset facial features to determine the first target person, it is realized to determine the first target person according to the accuracy requirement for the sound area, that is, to determine the first target sound area, meeting the user's requirements for voice interaction.
[0058] In addition, in some embodiments, determining the first target sound area of the first target person in the vehicle according to the first position information and the second position information includes:
[0059] If the first position information is the same as the second position information, the sound area corresponding to the first position information or the second position information is determined as the first target sound area;
[0060] If the first position information is different from the second position information, the sound area corresponding to the second position information is determined as the first target sound area.
[0061] Specifically, if the first position information and the second position information are the same, for example, the first position information determined through the sound dimension indicates that the first target person is located at the left rear position of the vehicle, and the second position information determined through the visual dimension also indicates that the first target person is located at the left rear position of the vehicle, then the sound area corresponding to the first position information or the second position information can be determined as the first target sound area.
[0062] In addition, if the first position information and the second position information are different, for example, the first position information determined through the sound dimension indicates that the first target person is located in the left rear position of the vehicle, and the second position information determined through the visual dimension indicates that the first target person is located in the right rear position of the vehicle, it is considered that the position determined through the sound dimension is inaccurate. In this case, the position determined through the visual dimension shall prevail, that is, the sound zone corresponding to the second position information is the first target sound zone.
[0063] In this way, by distinguishing whether the first position information is the same as the second position information, the first target sound zone where the first target person is located is finally determined, ensuring the accuracy of the first target sound zone.
[0064] In addition, since the person interacting with the voice assistant in the vehicle may change, this embodiment also needs to detect other persons interacting with the voice assistant in real time. Specifically, in some embodiments, after controlling the voice assistant to only receive and recognize the voice information within the first target sound zone according to the first target sound zone, any one or more of the following may be performed:
[0065] First, obtain the real-time voice in the vehicle, and detect whether a preset command word for the voice assistant is received according to the real-time voice; if it is detected that a preset command word for the voice assistant is received, determine the second target sound zone where the second target person who issued the preset command word is located; control the voice assistant to receive and recognize the voice information within the second target sound zone.
[0066] Specifically, the preset command words for the voice assistant may include "automatically turn on the air conditioner", "automatically open the window", etc.
[0067] When determining the second target sound zone where the second target person who issued the preset command word is located, the third position information of the second target person in the vehicle may be determined according to the preset command word, and the fourth position information of the second target person in the vehicle may be determined according to the image information, and then the second target sound zone of the second target person in the vehicle is jointly determined according to the third position information and the fourth position information.
[0068] In this way, after this embodiment controls the voice assistant to only receive and recognize the voice information within the first target sound zone, it can also monitor and obtain the real-time voice in the vehicle in real time. If it is detected that the real-time voice includes a preset command word for the voice assistant, it determines the second target sound zone where the second target person who issued the preset command word is located, and controls the voice assistant to receive and recognize the voice information within the second target sound zone, enabling other persons in the vehicle to join the voice interaction with the voice assistant at any time, and avoiding the situation where other persons have a need to interact with the voice assistant but the voice assistant fails to receive the voice information.
[0069] Second, update the image information to obtain real-time image information; when it is determined that there is a third target person in the vehicle and a preset command word is output according to the real-time image information, determine the third target sound area where the third target person is located; control the voice assistant to receive and recognize the voice information in the third target sound area.
[0070] Specifically, in this embodiment, the acquired image information can be updated in real time to obtain real-time image information, that is, the real-time image information includes the face images of the people at the current moment.
[0071] When it is determined that there is a third target person in the vehicle and a preset command word is output according to the real-time image information, the facial features corresponding to the face image at the current moment can be compared with the preset facial features of the person when the preset command word is output, and the similarity between the facial features corresponding to the face image at the current moment and the preset facial features when the preset command word is output can be calculated to determine the third target person.
[0072] After determining the third target person, the third target sound area where the third target person is located can be determined according to the seating position information of each person included in the image information, and the voice assistant can be controlled to receive and recognize the voice information in the third target sound area, realizing that other people can interact with the voice assistant by voice at any time.
[0073] In this way, after this embodiment controls the voice assistant to only receive and recognize the voice information in the first target sound area, when it is determined that there is a third target person in the vehicle and a preset command word is output according to the real-time image information, the third target sound area where the third target person is located is determined, and the voice assistant is controlled to receive and recognize the voice information in the third target sound area, enabling other people in the vehicle to join the voice interaction with the voice assistant at any time, and avoiding the situation where other people have the need to interact with the voice assistant but the voice assistant cannot receive the voice information.
[0074] Third, if it is detected that no voice information in the first target sound area is received within a preset time period, control the voice assistant to no longer receive the voice information in the first target sound area.
[0075] Specifically, the preset time period can be set according to actual needs, for example, it can be set to 3 minutes, 5 minutes, etc.
[0076] If it is detected that no voice information in the first target sound area is received within the preset time period, that is, no control command from the first target person is received, the voice assistant can be controlled to no longer receive the voice information in the first target sound area, thereby avoiding the problem that the voice assistant has been able to receive the voice information in the first target sound area and affecting the voice recognition of other sound areas.
[0077] Next, throughFigure 2 An embodiment of the present application will be described. As Figure 2 shown, the recognition process of voice information includes:
[0078] First, the rear-row passengers are recognized by the camera. Specifically, the in-vehicle OMS camera retrieves the real-time video image information inside the vehicle, recognizes the rear-row passengers of the vehicle, outputs the quantity and position, and saves the recognition result at this time. There are three results output by the OMS camera for recognizing the rear-row passengers. Result one is that there is only one person in the rear row. Result two is that there are people on both the left and right in the rear row. Result three is that there is no one in the rear row.
[0079] If there is only one person in the rear row, it is monitored whether a voice wake-up command for the voice assistant is received. If the voice wake-up command is received, the position of the target person who issued the voice wake-up command is determined, that is, the position result is output through the microphone. The time period when the voice wake-up command is received is recorded, and the image information captured by the OMS camera within this time period is called, and then the target sound area where the target person is located is determined according to the position result determined by the image information.
[0080] If there are people on both the left and right in the rear row, the voice wake-up command for the voice assistant is monitored through the microphone. If the voice wake-up command is received, the position result is output through the microphone. Then, the image information is captured by the OMS camera. It is detected whether the position result output by the microphone is consistent with the position result determined by the image information. If they are inconsistent, the position result determined by the image information shall prevail. If they are consistent, the target sound area where the target person is located is determined according to the same position result.
[0081] All the above optional technical solutions can be combined arbitrarily to form the optional embodiments of the present application, which will not be elaborated one by one here.
[0082] The following is the device embodiment of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.
[0083] Figure 3 is a schematic diagram of a voice information recognition device provided by an embodiment of the present application. As Figure 3 shown, the vehicle rear-row sound area recognition device includes:
[0084] A first determination module 301, configured to determine the first position information of a first target person in the vehicle according to the voice wake-up command when receiving the voice wake-up command for the voice assistant, where the first target person is the person who issued the voice wake-up command;
[0085] A second determination module 302, configured to obtain the image information inside the vehicle and determine the second position information of the first target person in the vehicle according to the image information;
[0086] A third determination module 303, configured to determine a first target sound area of the first target person in the vehicle according to the first position information and the second position information;
[0087] A voice recognition module 304, configured to control the voice assistant to only receive and recognize voice information within the first target sound area according to the first target sound area.
[0088] In some embodiments, the image information includes face images of each person in the vehicle and the seating information of each person;
[0089] The second determination module is specifically configured to obtain facial features of each person according to the face images of each person, where the facial features include face features and mouth features; determine the first target person according to the facial features of each person and preset facial features of the person when outputting the voice wake-up instruction stored in advance; and determine second position information of the first target person in the vehicle according to the seating information of each person.
[0090] In some embodiments, the second determination module is specifically configured to determine the similarity between the facial features of each person and the preset facial features; and when there is at least one target similarity greater than a preset value, determine each person corresponding to the target similarity as the first target person.
[0091] In some embodiments, the third determination module is specifically configured to, if the first position information is the same as the second position information, determine the sound area corresponding to the first position information or the second position information as the first target sound area; if the first position information is different from the second position information, determine the sound area corresponding to the second position information as the first target sound area.
[0092] In some embodiments, the voice recognition module is further configured to obtain real-time voice in the vehicle, and detect whether a preset instruction word for the voice assistant is received according to the real-time voice; if it is detected that the preset instruction word for the voice assistant is received, determine a second target sound area where the second target person who issues the preset instruction word is located; and control the voice assistant to receive and recognize voice information within the second target sound area.
[0093] In some embodiments, the voice recognition module is further configured to update the image information to obtain real-time image information; when it is determined according to the real-time image information that there is a third target person in the vehicle who outputs a preset instruction word, determine a third target sound area where the third target person is located; and control the voice assistant to receive and recognize voice information within the third target sound area.
[0094] In some embodiments, the voice recognition module is further configured to, if it is detected that no voice information in the first target sound area is received within a preset time period, control the voice assistant to no longer receive voice information in the first target sound area.
[0095] It should be understood that the sequence numbers of the steps in the above embodiments do not represent the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0096] Figure 4 is a schematic diagram of the electronic device 4 provided by the embodiments of the present application. As Figure 4 shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above various method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above device embodiments are implemented.
[0097] The electronic device 4 may be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 4 may include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 merely examples of the electronic device 4 do not constitute a limitation to the electronic device 4, and may include more or fewer components than those shown in the figure, or different components.
[0098] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0099] The memory 402 can be an internal storage unit of the electronic device 4. For example, it can be the hard disk or memory of the electronic device 4. The memory 402 can also be an external storage device of the electronic device 4. For example, it can be a plug-in hard disk equipped on the electronic device 4, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory 402 can also include both the internal storage unit and the external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.
[0100] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0101] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0102] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for recognizing voice information, characterized in that, it includes: When receiving a voice wake-up instruction for a voice assistant, determining first position information of a first target person in the vehicle according to the voice wake-up instruction, where the first target person is the person who issues the voice wake-up instruction; Obtaining image information in the vehicle and determining second position information of the first target person in the vehicle according to the image information; Determining a first target sound area of the first target person in the vehicle according to the first position information and the second position information; Controlling the voice assistant to only receive and recognize voice information within the first target sound area according to the first target sound area; After controlling the voice assistant to only receive and recognize voice information within the first target sound area according to the first target sound area, it further includes: Updating the image information to obtain real-time image information; When determining that there is a third target person in the vehicle outputting a preset command word according to the real-time image information, determining a third target sound area where the third target person is located; Controlling the voice assistant to receive and recognize voice information within the third target sound area.
2. The method for recognizing voice information according to claim 1, characterized in that, the image information includes face images of each person in the vehicle and the seating information of each person; The determining the second position information of the first target person in the vehicle according to the image information includes: Obtaining facial features of each person according to the face images of each person, where the facial features include facial features and mouth features; Determining the first target person according to the facial features of each person and preset facial features of the person when outputting the voice wake-up instruction stored in advance; Determining the second position information of the first target person in the vehicle according to the seating information of each person.
3. The method for recognizing voice information according to claim 2, characterized in that, The determining the first target person according to the facial features of each person and preset facial features of the person when outputting the voice wake-up instruction stored in advance includes: Determining the similarity between the facial features of each person and the preset facial features; When there is at least one target similarity greater than a preset value, determining each person corresponding to each target similarity as the first target person.
4. The method for recognizing voice information according to claim 1, characterized in that, The determining the first target sound area of the first target person in the vehicle according to the first position information and the second position information includes: If the first position information is the same as the second position information, determining the sound area corresponding to the first position information or the second position information as the first target sound area; If the first position information is different from the second position information, determining the sound area corresponding to the second position information as the first target sound area.
5. The method for recognizing voice information according to claim 1, characterized in that, After controlling the voice assistant to only receive and recognize voice information within the first target voice range according to the first target voice range, the method further includes: Obtaining real-time voice inside the vehicle, and detecting whether a preset command word for the voice assistant is received according to the real-time voice; If it is detected that the preset command word for the voice assistant is received, determining a second target voice range where a second target person who issues the preset command word is located; Controlling the voice assistant to receive and recognize voice information within the second target voice range.
6. The method for recognizing voice information according to claim 1, wherein, After controlling the voice assistant to only receive and recognize voice information within the first target voice range according to the first target voice range, the method further includes: If it is detected that no voice information within the first target voice range is received within a preset time period, controlling the voice assistant to no longer receive voice information within the first target voice range.
7. A device for recognizing voice information, wherein, it includes: A first determination module, configured to determine first position information of a first target person inside the vehicle according to a voice wake-up instruction for the voice assistant when the voice wake-up instruction for the voice assistant is received, where the first target person is the person who issues the voice wake-up instruction; A second determination module, configured to obtain image information inside the vehicle, and determine second position information of the first target person inside the vehicle according to the image information; A third determination module, configured to determine a first target voice range of the first target person inside the vehicle according to the first position information and the second position information; A voice recognition module, configured to control the voice assistant to only receive and recognize voice information within the first target voice range according to the first target voice range; The voice recognition module is further configured to update the image information to obtain real-time image information; when it is determined that a third target person inside the vehicle outputs a preset command word according to the real-time image information, determining a third target voice range where the third target person is located; controlling the voice assistant to receive and recognize voice information within the third target voice range.
8. An electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Vehicle-mounted terminal equipment, vehicle-mounted interaction system and interaction method
CN109941231A
Man-machine conversation method and device, robot, computer equipment and storage medium
CN112309395A