Voice control method and device, equipment, storage medium and vehicle
By combining sound source positioning and in-vehicle image recognition technology, accurately positioning the issuer and target equipment of in-vehicle voice commands, the problem of inaccurate execution of voice commands in the prior art is solved, and the accuracy and user experience of voice interaction are improved.
Patent Information
- Application Number
- CN202311465463.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-06
AI Technical Summary
In the case of multiple windows and multiple seats in the vehicle, it is difficult for the prior art to accurately locate the voice area where the user sending the voice command is located, resulting in the device that ultimately executes the voice command is not the device that the user wants to control, and the user experience is poor.
By obtaining the sound source positioning results in the target sound area and the in-car image at the target moment, image recognition processing is performed to obtain the position information of the personnel in the vehicle, and the target positioning results are used to determine the target position closest to the sound source, and the target device corresponding to the target position is controlled to execute voice commands.
It realizes that when the user issues voice commands from the boundary or transverse zone, the target device that the user wants to control is accurately positioned, improve the accuracy of voice interaction and improve the user experience.
Smart Images

Figure CN119943041A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of vehicle technology, and in particular to a voice control method, device, equipment, storage medium and vehicle. Background Art
[0002] With the development of voice technology, more and more interactive scenes in life can be completed by voice. In the car environment, the driver or passengers can control the in-car devices through voice commands, such as opening windows, adjusting seats, etc., but there are multiple windows and multiple seats in the car, and it is necessary to determine which in-car device the user wants to control. The existing technology often uses the sound zone pickup separation algorithm to determine the sound zone where the user who issued the voice command is located, and then determines the in-car device that the user wants to control based on the sound zone. However, the range of activities of people in the car is large, and the voice command may be issued from the boundary of the sound zone, or across the sound zone. At this time, it is difficult to accurately locate the sound zone, resulting in the device that finally executes the voice command is not the device that the user wants to control, and the user experience is poor. Summary of the invention
[0003] In order to solve the above technical problems, the present disclosure provides a voice control method, device, equipment, storage medium and vehicle.
[0004] A first aspect of an embodiment of the present disclosure provides a voice control method, which is applicable to a vehicle, and includes:
[0005] Acquire a sound source localization result in a target sound zone, where the target sound zone is a sound zone where a voice command is detected at a target time;
[0006] Obtain the in-vehicle image collected at the target time;
[0007] Performing image recognition processing on the in-vehicle image to obtain position information of the person in the vehicle, and determining a target position closest to the sound source based on the position information of the person in the vehicle and the sound source localization result, wherein the target position is the position information of the person who issued the voice command;
[0008] The target device corresponding to the target position is controlled to execute the operation corresponding to the voice instruction.
[0009] A second aspect of an embodiment of the present disclosure provides a voice control device, which is applicable to a vehicle, and includes:
[0010] A sound source localization module, used to obtain a sound source localization result in a target sound zone, where the target sound zone is a sound zone where a voice command is detected at a target time;
[0011] An image acquisition module, used to acquire the in-vehicle image captured at the target moment;
[0012] a position determination module, configured to perform image recognition processing on the in-vehicle image to obtain position information of the occupants in the vehicle, and determine a target position closest to the sound source based on the position information of the occupants in the vehicle and the sound source localization result, wherein the target position is the position information of the person who issued the voice command;
[0013] The device control module is used to control the target device corresponding to the target position to execute the operation corresponding to the voice instruction.
[0014] A third aspect of an embodiment of the present disclosure provides a computer device, including a memory and a processor, and a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, the voice control method as described in the first aspect above is implemented.
[0015] A fourth aspect of an embodiment of the present disclosure provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the voice control method of the first aspect described above is implemented.
[0016] The fifth aspect of an embodiment of the present disclosure provides a vehicle, comprising a memory and a processor, and a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, the voice control method as described in the first aspect above is implemented.
[0017] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:
[0018] In the voice control method, apparatus, device, storage medium and vehicle provided in the embodiments of the present disclosure, by obtaining the sound source positioning result in the target sound zone, the target sound zone is the sound zone where the voice command is detected at the target time, obtaining the in-vehicle image collected at the target time, performing image recognition processing on the in-vehicle image, obtaining the person position information of the people in the vehicle, and determining the target position closest to the sound source based on the person position information of the people in the vehicle and the sound source positioning result, the target position is the position information of the person who issued the voice command, and controlling the target device corresponding to the target position to perform the operation corresponding to the voice command. The target position of the user who issued the voice command in the vehicle can be accurately located in combination with the position of the sound source and the in-vehicle image, so as to determine the target device that the user wants to control according to the target position, and further, when the user issues a voice command from the boundary of the sound zone or across the sound zone, the target device that the user wants to control can also be accurately located, thereby improving the accuracy of voice interaction and enhancing user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0021] Figure 1 is a schematic diagram of a person in a vehicle issuing a voice command provided by an embodiment of the present disclosure;
[0022] Figure 2 is a schematic diagram of another embodiment of the present disclosure in which a person in a vehicle issues a voice command;
[0023] Figure 3 is a flow chart of a voice control method provided by an embodiment of the present disclosure;
[0024] Figure 4 is a flow chart of a method for determining a target pitch range provided by an embodiment of the present disclosure;
[0025] Figure 5 is a flow chart of a method for determining a sound source localization result provided by an embodiment of the present disclosure;
[0026] Figure 6 is a flow chart of a method for determining a target position provided by an embodiment of the present disclosure;
[0027] Figure 7 is a flow chart of another method for determining a target position provided by an embodiment of the present disclosure;
[0028] Figure 8 is a flow chart of a method for controlling a target device provided by an embodiment of the present disclosure;
[0029] Fig. 9 is a flowchart of another method for controlling a target device provided by an embodiment of the present disclosure;
[0030] Fig.10 is a flow chart of another voice control method provided by an embodiment of the present disclosure;
[0031] Fig.11 is a structural schematic diagram of a voice control device provided by an embodiment of the present disclosure;
[0032] Fig.12 It is a structural diagram of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0034] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0035] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0036] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0037] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0038] Figure 1 is a schematic diagram of a vehicle occupant issuing a voice command provided by an embodiment of the present disclosure, such as Figure 1 As shown, the interior of the car includes sound zone A and sound zone B. There is a boundary area at the connection between sound zone A and sound zone B. The user originally in sound zone A moves to the boundary area to issue a voice command. At this time, it is difficult to determine which sound zone the voice command issued by user A belongs to by using the traditional sound zone pickup separation algorithm.
[0039] Figure 2 FIG. 1 is another schematic diagram of a person in a vehicle issuing a voice command provided by an embodiment of the present disclosure, such as Figure 2As shown, there are four sound zones in the car, namely, sound zone 1, sound zone 2, sound zone 3, and sound zone 4. There is a occupant in each of sound zone 1, sound zone 2, and sound zone 3. The occupant in sound zone 3 issues a voice command, and the voice control device determines the acoustic positioning coordinates according to the voice command, that is, the sound source positioning result in S301. At this time, sound zone 4 is the target sound zone in S301, and then obtains the feature position information and the person position information including the visual positioning coordinates of the occupant in sound zone 3 according to the in-car image. Based on the sound source positioning result, the feature position information, and the person position information, it is determined that the occupant in the car who is closest to the acoustic positioning coordinates in sound zone 4 is the occupant in sound zone 3, and then it is determined that the occupant in the car who issues the voice command is the occupant in sound zone 3.
[0040] Figure 3 is a flow chart of a voice control method provided by an embodiment of the present disclosure, and the method can be executed by a voice control device. Figure 3 As shown, the voice control method provided in this embodiment includes the following steps:
[0041] S301: Acquire a sound source localization result in a target sound zone, where the target sound zone is a sound zone where a voice command is detected at a target time.
[0042] The target moment in the embodiment of the present disclosure may be understood as the moment when the voice command is detected.
[0043] The sound zone in the embodiment of the present disclosure may be understood as an area for collecting sounds inside the vehicle obtained by pre-dividing the space inside the vehicle, and the target sound zone may be understood as a sound zone where voice commands are collected.
[0044] In the disclosed embodiment, the voice control device can monitor the sounds inside the car in real time through a microphone, monitor the voice command at the target time, determine the target sound area of the monitored voice command, determine the position of the sound source of the voice command based on the time difference and amplitude difference when the voice command reaches each microphone, and obtain the sound source positioning result.
[0045] In an exemplary implementation of the disclosed embodiment, when multiple microphones are provided in a vehicle, the voice control device can, after one or more microphones detect a voice command, determine the microphone closest to the sound source of the voice command based on the loudness and time of the voice commands detected by each microphone, and then determine the target sound area where the voice command is detected based on the position of the microphone in the vehicle, determine the position of the sound source of the voice command based on the time difference and amplitude difference when the voice command reaches each microphone, and obtain a sound source localization result.
[0046] S302: Acquire the in-vehicle image collected at the target time.
[0047] In the disclosed embodiment, the voice control device may acquire the in-vehicle image captured by the image capture device at the target time after determining that the voice command is detected at the target time.
[0048] In an exemplary implementation of the disclosed embodiment, the image acquisition device can acquire in-vehicle images in real time and record the acquisition time of each image, so that after detecting a voice command, the in-vehicle image acquired at the target time is selected from multiple in-vehicle images acquired in advance.
[0049] In another exemplary implementation of the disclosed embodiment, the voice control device may send an image acquisition instruction to the image acquisition device after detecting the voice instruction, so that the image acquisition device acquires the image inside the vehicle after receiving the image acquisition instruction.
[0050] In another exemplary implementation of the disclosed embodiment, when there are multiple image acquisition devices in a vehicle for acquiring images of different areas, the voice control device can obtain the in-vehicle image acquired by a target image acquisition device at a target time, and the image acquisition area of the target image acquisition device includes the position of the sound source.
[0051] In another exemplary implementation of the embodiment of the present disclosure, the image acquisition device may be a time-of-flight camera or a binocular camera.
[0052] S303, performing image recognition processing on the in-car image to obtain the position information of the occupants in the car, and determining the target position closest to the sound source based on the position information of the occupants in the car and the sound source positioning result, where the target position is the position information of the person who issued the voice command.
[0053] In the disclosed embodiment, the voice control device can perform image recognition processing on the image inside the vehicle after obtaining the image inside the vehicle, thereby identifying the person position information of the people inside the vehicle, and based on the identified position information of each person and the sound source positioning result, determine the person position information closest to the sound source from the person position information. The person position information is the person position information of the person who issued the voice command, and the person position information is determined as the target position.
[0054] In an exemplary implementation of the disclosed embodiment, the voice control device can perform image recognition processing on the in-vehicle image after obtaining it, determine the pixel position corresponding to the person position of each person contained in the in-vehicle image, and then based on the internal parameters and external parameters of the image acquisition device that captures the in-vehicle image acquired in advance, convert the pixel position corresponding to the two-dimensional person position of the person in the vehicle into a three-dimensional spatial position, and determine it as the person position information, and then based on the person position information and the sound source localization result, calculate the distance between the person position of each person in the vehicle and the sound source, and determine the person position information with the shortest distance as the target position.
[0055] S304: Control the target device corresponding to the target position to execute the operation corresponding to the voice command.
[0056] The target device in the embodiment of the present disclosure can be understood as an in-vehicle device controlled by a voice command target. For example, the in-vehicle device may include windows, seats, air conditioners, vehicle screens, audio systems, etc., which are not limited here.
[0057] In the embodiment of the present disclosure, after determining the target position, the voice control device can determine the target device controlled by the voice command corresponding to the target position from the in-vehicle devices, and control the target device to perform the operation corresponding to the voice command.
[0058] In an exemplary implementation of the disclosed embodiment, the voice control device may determine a control instruction based on the voice instruction, and send the control instruction to a target device corresponding to the target location, so that the target device performs the operation corresponding to the voice instruction after receiving the control instruction.
[0059] The disclosed embodiment obtains a sound source positioning result within a target sound zone, where the target sound zone is a sound zone where a voice command is detected at a target moment, obtains an in-vehicle image collected at the target moment, performs image recognition processing on the in-vehicle image, obtains position information of people in the vehicle, and determines a target position closest to the sound source based on the position information of the people in the vehicle and the sound source positioning result, where the target position is the position information of the person who issues the voice command, controls a target device corresponding to the target position to perform an operation corresponding to the voice command, and can accurately locate the target position of a user who issues a voice command in the vehicle in combination with the position of the sound source and the in-vehicle image, thereby determining a target device that the user wants to control based on the target position, and furthermore, when the user issues a voice command from a sound zone boundary or across sound zones, the target device that the user wants to control can also be accurately located, thereby improving the accuracy of voice interaction and enhancing user experience.
[0060] Figure 4 is a flow chart of a method for determining a target pitch range provided by an embodiment of the present disclosure. Figure 4 As shown, based on the above embodiment, the target sound range can be determined by the following method.
[0061] S401, performing multi-channel separation on the speech signal collected at the target time.
[0062] In the embodiment of the present disclosure, after determining that the microphone detects a voice signal at a target time, the voice control device may perform multi-channel separation on all voice signals detected at the target time to obtain a multi-channel signal.
[0063] S402: Perform voice activity detection on the separated multi-channel signal.
[0064] In the disclosed embodiment, after obtaining the multi-channel signal, the voice control device may perform voice activity detection on the multi-channel signal to determine whether the voice signal of each channel contains voice activity, and filter out the voice signal containing voice activity.
[0065] S403: Based on the voice activity detection result, determine the sound zone where the voice command is detected as the target sound zone.
[0066] In the disclosed embodiment, the voice control device can determine the channel to which the voice signal belongs after screening out the voice signal containing voice activity, and then determine the sound zone where the voice signal containing voice activity is located based on the correspondence between the channel and the sound zone.
[0067] The disclosed embodiment performs multi-channel separation on the voice signal collected at the target moment, performs voice activity detection on the separated multi-channel signal, and based on the voice activity detection result, determines the sound area where the voice command is detected as the target sound area, thereby accurately locating the target sound area where the voice command is located and improving the accuracy of sound source positioning.
[0068] Figure 5 is a flow chart of a method for determining a sound source localization result provided by an embodiment of the present disclosure, such as Figure 5 As shown, based on the above embodiment, the sound source localization result can be determined by the following method.
[0069] S501: Perform spatial spectrum estimation based on the speech signal collected by the microphone array to determine the sound source localization result.
[0070] In the disclosed embodiment, the voice control device can utilize all voice signals collected by the microphone array composed of all microphones installed in the vehicle to perform spatial spectrum estimation on the entire vehicle interior space and determine the position of the sound source of each voice signal. Specifically, the voice control device can utilize the time difference and amplitude difference from the sound source to each microphone to estimate the distance difference from the sound source to each microphone, and then infer the spatial position of the sound source through the geometric relationship between the microphones to obtain the sound source localization result.
[0071] S502: Filter the sound source localization results located in the target sound area from the sound source localization results.
[0072] In the embodiment of the present disclosure, after obtaining multiple sound source localization results, the voice control device can filter out the sound source localization results located in the target sound zone based on the sound zones where the respective sound source localization results are located.
[0073] The embodiment of the present disclosure determines the sound source localization result by performing spatial spectrum estimation based on the speech signal collected by the microphone array, and filters the sound source localization results located in the target sound area from the sound source localization results. The sound source localization results of the voice commands located in the target sound area can be filtered out, thereby further improving the accuracy of sound source localization.
[0074] Figure 6 is a flow chart of a method for determining a target position provided by an embodiment of the present disclosure, such as Figure 6 As shown, based on the above embodiment, the target position can be determined by the following method.
[0075] S601: Determine the number of people in the vehicle.
[0076] In the disclosed embodiment, the voice control device can determine the number of people in the vehicle. Specifically, after obtaining the image of the vehicle, the voice control device can perform human body recognition processing on the image of the vehicle and determine the number of people in the vehicle based on the recognition result. It can also determine the number of people in the vehicle based on data collected by a seat pressure sensor, which is not limited here.
[0077] In an exemplary implementation of the disclosed embodiment, the voice control device may input the in-vehicle image into a pre-trained human body detection model, and the human body detection model may annotate the human bodies in the in-vehicle image, thereby determining the number of people in the in-vehicle image.
[0078] S602: In response to the number of people in the vehicle being greater than one, determining a target position closest to the sound source based on the position information of the people in the vehicle and the sound source localization result.
[0079] In the embodiment of the present disclosure, the voice control device can perform human body recognition on the image inside the car, and after obtaining the number of people in the car, determine whether the number of people in the car is greater than one. If it is greater, execute the step of S303 to determine the target position closest to the sound source based on the person position information of the people in the car and the sound source positioning result.
[0080] S603: In response to the number of people in the vehicle being equal to one, determining the person's position information as a target position.
[0081] In the embodiment of the present disclosure, the voice control device can determine that the voice command is issued by the only person in the car when it is determined that the number of people in the car is equal to one. At this time, the position of the only person in the car can be determined as the target position, and the person's position information is determined as the target position, and then the target device corresponding to the target position is controlled to perform the operation corresponding to the voice command.
[0082] In another exemplary implementation of the disclosed embodiment, if the number of people in the vehicle is zero, the voice command is not processed.
[0083] The disclosed embodiment determines the number of people in the car, and in response to the number of people in the car being greater than one, determines the target position closest to the sound source based on the person position information of the people in the car and the sound source positioning result, and in response to the number of people in the car being equal to one, determines the person position information as the target position. When there are multiple people in the car, the sender of the voice command closest to the sound source can be filtered according to the distance between the person position information and the sound source. Otherwise, the sender of the voice command is directly determined according to the person position information, thereby reducing the overall workload and reducing the occupancy rate of computing resources.
[0084] Figure 7 is a flowchart of another method for determining a target location provided by an embodiment of the present disclosure. Figure 7 As shown, based on the above embodiment, the target position can be determined by the following method, wherein the personnel position information is seat position information.
[0085] S701. Perform image recognition processing on the in-vehicle image to determine characteristic position information of the occupants of the vehicle, where there is a corresponding relationship between the characteristic position information and the seat position information.
[0086] The seat position information in the embodiment of the present disclosure may be understood as the position information of the part where a person contacts the seat, and the feature position information may be understood as the facial feature position information used to characterize the voice position of the person in the vehicle. For example, the feature position information may be the mouth position information, and there is a corresponding relationship between the feature position information and the seat position information of the same person.
[0087] In the disclosed embodiment, the voice control device can, after obtaining the position information of the person, specifically the seat position information of the part of the seat where the person is in contact, further perform image recognition processing on the image inside the vehicle to identify the characteristic position information of the person inside the vehicle. There is a corresponding relationship between the characteristic position information and the seat position information of the same person.
[0088] In an exemplary implementation of the disclosed embodiment, the voice control device can perform image recognition processing on the in-vehicle image, determine the pixel coordinates of the facial position of each person contained in the in-vehicle image, and then based on the internal and external parameters of the image acquisition device that captures the in-vehicle image acquired in advance, convert the two-dimensional pixel coordinates into three-dimensional spatial coordinates to obtain the characteristic position information of the people in the vehicle, and at the same time establish a corresponding relationship between the characteristic position information and seat position information of the same person in the vehicle.
[0089] In another exemplary implementation of the disclosed embodiment, after obtaining the in-vehicle image, the voice control device may identify the characteristic position information of the occupant of the vehicle while identifying the position information of the occupant of the vehicle.
[0090] S702: Determine the target characteristic position information closest to the sound source from the characteristic position information of the people in the vehicle.
[0091] In the disclosed embodiment, after obtaining the characteristic position information of the occupant of the vehicle, the voice control device can determine the target characteristic position information closest to the sound source from each characteristic position information.
[0092] In an exemplary implementation of the disclosed embodiment, the voice control device can obtain the position of the sound source and the three-dimensional coordinates of the characteristic positions of each person in the vehicle, and then calculate the Euclidean distance between the position of the sound source and each characteristic position based on the three-dimensional coordinates, and determine the characteristic position information corresponding to the characteristic position with the shortest distance as the target characteristic position information.
[0093] S703: Determine the seat position information corresponding to the target feature position information as the target position.
[0094] In the disclosed embodiment, after determining the target feature position information, the voice control device may determine the seat position information corresponding to the target feature position information as the target position.
[0095] The disclosed embodiment determines characteristic position information of occupants of the vehicle by performing image recognition processing on images within the vehicle. There is a corresponding relationship between the characteristic position information and the seat position information. The target characteristic position information closest to the sound source is determined from the characteristic position information of the occupants of the vehicle, and the seat position information corresponding to the target characteristic position information is determined as the target position. When the occupants of the vehicle have a large range of activities and issue voice commands in other sound zones other than the sound zone to which their sitting positions belong, the sitting positions are still used as the basis, and the in-vehicle devices that the user wants to control are determined to be the in-vehicle devices around the sitting position, thereby further improving the accuracy of voice interaction and enhancing the user experience.
[0096] Figure 8 is a flow chart of a method for controlling a target device provided by an embodiment of the present disclosure, such as Figure 8 As shown, based on the above embodiment, the target device can be controlled by the following method.
[0097] S801: Identify the device type of the in-vehicle device controlled by the voice command target.
[0098] In the disclosed embodiment, the voice control device may identify the voice command after determining the target position of the person in the vehicle who issues the voice command, and determine the type of the in-vehicle device that is targeted for control by the voice command.
[0099] In an exemplary implementation of the disclosed embodiment, the voice control device may convert the voice command into text information after receiving the voice command. Specifically, the voice control device may process the obtained voice command through a pre-trained text conversion model, eliminate the silent period contained in the voice command, perform frame processing, extract sound features from the obtained audio, and then obtain text information through an acoustic model and a language model. Then, in the converted text information, it is detected whether the keywords in a preset keyword library are included to obtain the target keywords included in the text information. Based on the correspondence between the pre-acquired keywords and the device types, the device type corresponding to the target keyword is determined, and it is determined as the device type of the in-vehicle device targeted for control by the voice command.
[0100] S802: Based on the pre-acquired location information of the in-vehicle device, determine a target device that matches the device type at the target location.
[0101] In the disclosed embodiment, the voice control device can obtain the location information of each in-vehicle device in advance, and then after identifying the device type of the in-vehicle device to be controlled by the voice command, determine the device location of each in-vehicle device of this device type, and then determine the device location that matches the target location where the person is located among multiple device locations, and determine the in-vehicle device at this device location as the target device.
[0102] In an exemplary implementation of the disclosed embodiment, the voice control device may determine the sound zone to which the target position belongs after determining the target position. For example, the sound zone may include a left front sound zone, a right front sound zone, a left rear sound zone, and a right rear sound zone. Further, after determining the sound zone to which the target position belongs, an in-vehicle device that matches the device type in the sound zone is determined as the target device. For example, when the device type is determined to be a window and the target position is located in the left rear sound zone in the vehicle, the left rear window is determined as the target device.
[0103] S803: Send a control instruction to the target device, so that the target device performs an operation corresponding to the voice instruction after receiving the control instruction.
[0104] The control instructions in the embodiments of the present disclosure may be understood as control instructions for controlling an in-vehicle device to perform a corresponding operation. For example, the control instructions may include specific operations that the target device needs to perform.
[0105] In the embodiment of the present disclosure, the voice control apparatus may send a control instruction to the target device after determining the target device, so that the target device performs an operation corresponding to the voice instruction after receiving the control instruction.
[0106] In an exemplary implementation of the embodiment of the present disclosure, the voice control device may send a control instruction to the vehicle computer, so that the vehicle computer controls the target device to perform an operation corresponding to the voice instruction after receiving the control instruction.
[0107] The disclosed embodiment identifies the device type of the in-vehicle device that is controlled by the voice command, determines the target device that matches the voice control device type at the voice control target location based on the pre-acquired location information of the in-vehicle device, sends a control command to the target device so that the target device executes the operation corresponding to the voice command after receiving the control command, and can accurately send the control command to the in-vehicle device that is controlled by the target, thereby realizing the user's voice control of the in-vehicle device.
[0108] Fig. 9 is a flowchart of another method for controlling a target device provided by an embodiment of the present disclosure. Fig. 9 As shown, based on the above embodiment, the target device can be controlled by the following method.
[0109] S901: Perform personnel recognition on the in-vehicle image to obtain personnel information of the in-vehicle personnel.
[0110] In the disclosed embodiment, after obtaining the in-vehicle image, the voice control device can perform personnel recognition on the in-vehicle image to determine the personnel information of each person in the vehicle.
[0111] In an exemplary implementation of the disclosed embodiment, after obtaining the in-vehicle image, the voice control device can recognize the facial image contained in the in-vehicle image, extract facial features, and match them with the facial features stored in a preset database to determine the identity ID of each person in the vehicle, and determine the identity ID as the personal information of the person in the vehicle.
[0112] In another exemplary implementation of the disclosed embodiment, after obtaining the in-car image, the voice control device can identify the age of each occupant contained in the in-car image, determine whether each occupant is an adult or a minor, and determine the determination result as the personnel information.
[0113] S902: Determine target personnel information of a target person located at a target location based on the personnel information.
[0114] In the disclosed embodiment, after determining the personal information of each person in the vehicle, the voice control device may determine the person in the vehicle at the target position as the target person, and determine the personal information of the target person as the target person information.
[0115] S903: Based on the target person information and the voice command, control the target device corresponding to the target location to perform the operation corresponding to the voice command in a manner preferred by the target person.
[0116] In the embodiment of the present disclosure, the voice control device can determine that the voice command is a voice command issued by the target person after determining the target person information of the target person located at the target location, and can determine the device control method preferred by the target person based on the target person information, and then control the target device corresponding to the target location based on the device control method to perform the operation corresponding to the voice command in the method preferred by the target person.
[0117] In an exemplary implementation of the disclosed embodiment, when the target person information is the target person information determined based on the identity ID, the target person information may include the control mode that the target person prefers, such as the seat heating temperature, the window opening, etc., which is pre-entered. When the target person information is an adult or a minor, the target person information may include the control authority of the target person over the in-vehicle equipment, such as an adult can fully open the window, while a minor can only open the window by one-third. The voice control device may determine the specific control mode of the target person over the target device based on the target person information, and then control the target device corresponding to the target position based on the control mode to perform the operation corresponding to the voice command according to the control mode corresponding to the target person.
[0118] The disclosed embodiment obtains personal information of the people in the vehicle by performing personnel recognition on the in-vehicle image, determines the target personal information of the target person at the target position based on the personal information, controls the target device corresponding to the target position to execute the operation corresponding to the voice command in a manner preferred by the target person based on the target personal information and the voice command, and can control the in-vehicle devices based on the personal information, thereby realizing personalized control of the in-vehicle devices and further improving the user experience.
[0119] Fig.10 is a flow chart of another voice control method provided by an embodiment of the present disclosure. Fig.10 As shown, after detecting voice activity, the acoustic module sends a command to the visual module. The visual module performs coordinate system conversion according to the occupant in the collected in-car image, and converts the feature position information and occupant position information of the occupant in the camera coordinate system into visual coordinates in the vehicle coordinate system. At the same time, the acoustic module determines the acoustic coordinates, that is, the sound source positioning result, according to the voice command, and then sends the acoustic coordinates and the voice command audio to the multi-mode fusion decision module. The multi-mode fusion decision module determines the occupant who issued the voice command based on the visual coordinates and the acoustic coordinates, and then sends the occupant information of the occupant together with the voice command audio to the wake-up / recognition module, which performs subsequent interactive control operations.
[0120] Fig.11 is a structural diagram of a voice control device provided by an embodiment of the present disclosure, such as Fig.11As shown, the voice control device 1100 includes: a sound source localization module 1110, an image acquisition module 1120, a position determination module 1130, and a device control module 1140, wherein the sound source localization module 1110 is used to obtain a sound source localization result in a target sound zone, and the target sound zone is a sound zone where a voice command is detected at a target time; the image acquisition module 1120 is used to obtain an in-vehicle image acquired at a target time; the position determination module 1130 is used to perform image recognition processing on the in-vehicle image to obtain the position information of the occupants in the vehicle, and determine the target position closest to the sound source based on the position information of the occupants in the vehicle and the sound source localization result, and the target position is the position information of the person who issued the voice command; the device control module 1140 is used to control the target device corresponding to the target position to perform the operation corresponding to the voice command.
[0121] Optionally, the voice control device 1100 also includes: a separation module for performing multi-channel separation on the voice signal collected at the target moment; a detection module for performing voice activity detection on the separated multi-channel signal; and a sound zone determination module for determining the sound zone in which the voice command is detected as the target sound zone based on the voice activity detection result.
[0122] Optionally, the sound source localization module 1110 includes: a first determination unit, used to perform spatial spectrum estimation based on the speech signal collected by the microphone array to determine the sound source localization result; and a screening unit, used to screen the sound source localization results located in the target sound area from the sound source localization results.
[0123] Optionally, the voice control device 1100 also includes: a quantity determination module, used to determine the number of people in the car; the position determination module 1130, including: a first determination module, used to determine the target position closest to the sound source based on the person position information of the people in the car and the sound source positioning result in response to the number of people in the car being greater than one; a second determination module, used to determine the person position information as the target position in response to the number of people in the car being equal to one.
[0124] Optionally, the personnel position information is seat position information, and the position determination module 1130 includes: an image recognition unit, used to perform image recognition processing on the in-vehicle image to determine characteristic position information of the personnel in the vehicle, and there is a corresponding relationship between the characteristic position information and the seat position information; a second determination unit, used to determine the target characteristic position information closest to the sound source from the characteristic position information of the personnel in the vehicle; and a third determination unit, used to determine the seat position information corresponding to the target characteristic position information as the target position.
[0125] Optionally, the device control module 1140 includes: a voice recognition unit, used to identify the device type of the in-vehicle device that is controlled by the voice command; a fourth determination unit, used to determine the target device that matches the device type at the target location based on pre-acquired location information of the in-vehicle device; and a sending unit, used to send a control command to the target device, so that the target device performs the operation corresponding to the voice command after receiving the control command.
[0126] Optionally, the voice control device 1100 also includes: a personnel identification module, which is used to perform personnel identification on the in-vehicle image to obtain personnel information of the personnel in the vehicle; an information determination module, which is used to determine the target personnel information of the target person located at the target position based on the personnel information; the device control module 1140, which is specifically used to control the target device corresponding to the target position to perform the operation corresponding to the voice command in a manner preferred by the target person based on the target personnel information and the voice command.
[0127] The voice control device provided in this embodiment can execute the method described in any of the above embodiments. Its execution method and beneficial effects are similar and will not be described in detail here.
[0128] Fig.12 It is a structural diagram of a computer device provided in an embodiment of the present disclosure.
[0129] like Fig.12 As shown, the computer device may include a processor 1210 and a memory 1220 storing computer program instructions.
[0130] Specifically, the processor 1210 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0131] The memory 1220 may include a large capacity memory for information or instructions. By way of example and not limitation, the memory 1220 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a tape, or a universal serial bus (USB) drive or a combination of two or more of these. Where appropriate, the memory 1220 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 1220 may be inside or outside the integrated gateway device. In a specific embodiment, the memory 1220 is a non-volatile solid-state memory. In a specific embodiment, the memory 1220 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (Electrically Erasable Programmable ROM, EEPROM), an electrically alterable ROM (EAROM) or flash memory, or a combination of two or more of these.
[0132] The processor 1210 reads and executes the computer program instructions stored in the memory 1220 to perform the steps of the voice control method provided in the embodiment of the present disclosure.
[0133] In one example, the computer device may further include a transceiver 1230 and a bus 1240. Fig.12 As shown, the processor 1210, the memory 1220 and the transceiver 1230 are connected via a bus 1240 and communicate with each other.
[0134] The bus 1240 includes hardware, software, or both. For example, but not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a Memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, the bus 1240 may include one or more buses. Although embodiments of the present application describe and illustrate a particular bus, the present application contemplates any suitable bus or interconnect.
[0135] The embodiments of the present disclosure also provide a computer-readable storage medium, which may store a computer program. When the computer program is executed by a processor, the processor implements the voice control method provided by the embodiments of the present disclosure.
[0136] The above-mentioned storage medium may, for example, include a memory 1220 of computer program instructions, and the above-mentioned instructions may be executed by the processor 1210 of the voice control device to complete the voice control method provided by the embodiment of the present disclosure. Optionally, the storage medium may be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium may be a ROM, a random access memory (Random Access Memory, RAM), a compact disc read-only memory (Compact Disc ROM, CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc. The above-mentioned computer program may be written in any combination of one or more programming languages to perform the program code for the operation of the embodiment of the present disclosure, and the programming language includes an object-oriented programming language, such as Java, C++, etc., and also includes a conventional procedural programming language, such as "C" language or similar programming language. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device, partially on the remote computing device, or completely on the remote computing device or server.
[0137] An embodiment of the present disclosure also provides a vehicle, which includes a memory, a processor, and a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, the various processes and effects in the above-mentioned embodiments of the present disclosure can be implemented, which will not be elaborated here.
[0138] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A voice control method, characterized in that: include: Acquire a sound source localization result in a target sound zone, where the target sound zone is a sound zone where a voice command is detected at a target time; Obtain the in-vehicle image collected at the target time; Performing image recognition processing on the in-vehicle image to obtain position information of the person in the vehicle, and determining a target position closest to the sound source based on the position information of the person in the vehicle and the sound source localization result, wherein the target position is the position information of the person who issued the voice command; The target device corresponding to the target position is controlled to execute the operation corresponding to the voice instruction.
2. The method according to claim 1, characterized in that Before obtaining the sound source localization result in the target sound area, the method further includes: Perform multi-channel separation on the speech signal collected at the target time; Performing voice activity detection on the separated multi-channel signals; Based on the voice activity detection result, the sound zone where the voice instruction is detected is determined as the target sound zone.
3. The method according to claim 1, characterized in that The obtaining of the sound source localization result in the target sound area includes: Perform spatial spectrum estimation based on the speech signal collected by the microphone array to determine the sound source localization result; The sound source localization results located in the target sound area are selected from the sound source localization results.
4. The method according to claim 1, characterized in that: Before determining the target position closest to the sound source based on the position information of the occupant in the vehicle and the sound source localization result, the method further includes: Determine the number of people in the vehicle; In response to the number of people in the vehicle being greater than one, determining a target position closest to the sound source based on the person position information of the people in the vehicle and the sound source localization result; In response to the number of people in the vehicle being equal to one, the person position information is determined as the target position.
5. The method according to claim 1, characterized in that The personnel position information is seat position information; The determining the target position closest to the sound source based on the position information of the person in the vehicle and the sound source positioning result includes: Performing image recognition processing on the in-vehicle image to determine characteristic position information of the person in the vehicle, wherein there is a corresponding relationship between the characteristic position information and the seat position information; Determine the target characteristic position information closest to the sound source from the characteristic position information of the person in the vehicle; The seat position information corresponding to the target feature position information is determined as the target position.
6. The method according to claim 1, characterized in that The controlling the target device corresponding to the target position to perform the operation corresponding to the voice command includes: Identify the device type of the in-vehicle device that is controlled by the voice command target; Based on the pre-acquired position information of the in-vehicle device, determining a target device matching the device type at the target location; A control instruction is sent to the target device, so that the target device performs the operation corresponding to the voice instruction after receiving the control instruction.
7. The method according to claim 1, characterized in that After acquiring the in-vehicle image captured at the target time, the method further includes: Performing personnel recognition on the in-vehicle image to obtain personnel information of the in-vehicle personnel; Based on the personnel information, determining target personnel information of a target person located at the target location; The controlling the target device corresponding to the target position to execute the operation corresponding to the voice command includes: Based on the target person information and the voice command, the target device corresponding to the target location is controlled to perform the operation corresponding to the voice command in a manner preferred by the target person.
8. A voice control device, characterized in that: The device is applicable to a vehicle, and comprises: A sound source localization module, used to obtain a sound source localization result in a target sound zone, where the target sound zone is a sound zone where a voice command is detected at a target time; An image acquisition module, used to acquire the in-vehicle image captured at the target moment; a position determination module, configured to perform image recognition processing on the in-vehicle image to obtain position information of the occupants in the vehicle, and determine a target position closest to the sound source based on the position information of the occupants in the vehicle and the sound source localization result, wherein the target position is the position information of the person who issued the voice command; The device control module is used to control the target device corresponding to the target position to execute the operation corresponding to the voice instruction.
9. A computer device, characterized in that: include: Memory; processor; and a computer program; wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A vehicle, characterized in that: include: Memory; processor; and a computer program; wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
Sound source positioning method, controller and vehicle
CN121687046A