A method and related device for voice command control in a vehicle cabin
By obtaining lip movement information in the cabin and matching voice commands, the problem of identifying the location of voice commands in multiple scenarios in the cabin is solved, and refined control and authority judgment are achieved, which improves the accuracy and safety of vehicle operations.
Patent Information
- Application Number
- CN202010631879.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-03
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2040-07-03
AI Technical Summary
In the case of multiple people in the cabin, it is difficult for the prior art to accurately identify the location of the voice command to be emitted, resulting in the inability to achieve refined voice control and prevent misoperation.
By obtaining the lip movement information of members in the car to match the voice commands, the feature matching model is used to identify the specific location of the voice commands, and the refined voice control and authority judgment are achieved.
It improves the accuracy of execution of voice commands in the cabin, reduces the occurrence of erroneous operations, and ensures the safety and refined control of vehicle operations.
Smart Images

Figure CN113963692B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction, and in particular to a voice command control method and related devices in a vehicle cabin. Background Art
[0002] With the development of driverless technology, the intelligence of vehicles is getting higher and higher. Intelligent voice control in the vehicle cabin has also become a mainstream requirement of the current intelligent cockpit. More and more vehicles are beginning to have voice interaction functions and can execute corresponding functions based on the voice commands of vehicle occupants.
[0003] Since the vehicle cabin is also a multi-person space, different vehicle occupants may have different control requirements. Therefore, when executing voice commands, for some special commands, the in-vehicle intelligent devices need to obtain the specific location of the command issuer to determine how to execute the corresponding commands. In a multi-person scenario in the vehicle cabin, some vehicle commands need to operate on specific positions in the vehicle cabin, such as adjusting the air volume of a specific air outlet or adjusting the volume of a specific speaker. At this time, when receiving a voice command to reduce the air volume of the air outlet, if the vehicle intelligent device cannot identify the specific location of the current command issuer, it is often impossible to perform refined control adjustments in the vehicle cabin based on the voice; some vehicles can perform driving control of the vehicle through voice, such as autonomous driving and automatic parking. However, when there are multiple occupants in the vehicle, such as children, how to determine when driving-related voice commands can be executed and when they cannot be executed, and confirm the identity and authority of the command issuer to prevent misoperation of the vehicle is also one of the problems that need to be solved in the field of in-vehicle human-computer interaction control in autonomous driving. Summary of the Invention
[0004] Embodiments of the present invention provide a voice command control method and related devices to improve the execution accuracy of voice commands in a multi-person scenario in the vehicle and reduce the situations of misoperation or misrecognition.
[0005] In a first aspect, an embodiment of the present invention provides a method for in-vehicle voice control, including: obtaining a first type of instruction and lip movement information of in-vehicle members at N positions in the vehicle cabin during a target time period, where the first type of instruction is obtained based on target audio data collected in the vehicle cabin, the lip movement information of the in-vehicle members is obtained when the first type of instruction is recognized from the target audio data, and the target time period is the time period corresponding to the first type of instruction in the audio data; matching the first type of instruction with the lip movement information of the in-vehicle members at the N positions in the vehicle cabin, and obtaining a target position according to the matching result between the lip movement information of the in-vehicle members at the N positions and the first type of instruction, where the target position is the position where the in-vehicle member whose lip movement information matches the first type of instruction is located; sending indication information for executing the first type of instruction for the target position.
[0006] The embodiment of the present invention can be applied to refined voice command control in the vehicle cabin. When it is recognized that the voice command is an instruction that needs to further determine the position before execution, such as an operation instruction for a device in the vehicle cabin, such as playing a video, speaker adjustment, air conditioner adjustment, seat adjustment, etc., which requires determining the specific position targeted by the instruction for targeted local operations. By obtaining the lip movement information of members at each position in the vehicle and the specific instruction information recognized, it can be determined which position the member who issued the instruction is located at, so as to be able to perform operation control on a specific position area in a targeted manner. Since the above voice control method involves the processing and analysis of video data and the matching degree of its local features with the instruction information, it can occur locally, that is, the vehicle or the intelligent device on the vehicle executes the above method, or the above video data processing and matching actions can also be executed in the cloud.
[0007] The method of the embodiment of the present invention is applicable to vehicle scenarios with any number of people, especially applicable to scenarios where there are multiple members in the vehicle and multiple people are speaking at the same time. At this time, by combining the matching of lip movement information and instruction information with the in-vehicle position distribution information, the position information of the member who issued the voice instruction can be accurately located, and then the position area targeted by the instruction execution can be determined. In the method described in the embodiment of the present invention, N does not necessarily represent all the members in the vehicle, and it can be all the members in the vehicle or some members.
[0008] In a possible implementation, the obtaining of the first type of instruction and the lip movement information of in-vehicle members at N positions in the vehicle cabin is specifically as follows: Obtain the target audio data in the vehicle cabin; when it is recognized that the target audio data includes the first type of instruction, obtain the in-vehicle image data; extract the lip movement information of in-vehicle members at N positions in the vehicle cabin from the in-vehicle image data. In the embodiments of the present invention, the extraction of the lip movement information of the members at the N positions in the vehicle cabin from the in-vehicle image data may specifically be based on a face recognition algorithm to recognize the N face regions in the video data and extract the lip movement video or sample the video frame sequence in each of the N face regions; determine the lip movement information of the N members based on the lip movement video or video frame sequence in each face region. The image data is usually obtained by one or more cameras in the vehicle cabin, and there can be various types of cameras.
[0009] At the same time, since there are usually multiple cameras in the vehicle cabin environment, some cameras can obtain video data from different perspectives by changing the shooting angle. Therefore, the image data mentioned in the embodiments of the present invention can be the image data obtained by one camera from one angle. However, when shooting from some angles, there may be occlusion situations between members at different positions in the vehicle. Therefore, the image data in the embodiments of the present invention can also be the image data from different perspectives of the same camera, or the image data from multiple cameras, as well as other combinations as mentioned above.
[0010] In a possible implementation, the extraction of the lip movement information of in-vehicle members at N positions in the vehicle cabin from the in-vehicle image data is specifically as follows: When the number of recognized in-vehicle members is greater than 1, extract the lip movement information of in-vehicle members at N positions in the vehicle cabin from the in-vehicle image data. In the specific implementation process of the embodiments of the present invention, to avoid waste of computing resources, after recognizing the first instruction type, the number of people in the vehicle cabin can be judged according to the obtained in-vehicle image data. When there is only one person in the vehicle cabin, there is no need to extract the lip movement information. Only when the number of people is more than 1, the lip movement information of the people at the occupied positions is extracted.
[0011] In a possible implementation, the matching of the first type of instruction and the lip movement information of the in-vehicle members at N positions in the vehicle cabin, and obtaining the target position according to the matching result between the lip movement information of the in-vehicle members at the N positions and the first type of instruction are specifically as follows: according to the first type of instruction and the lip movement information of the N in-vehicle members in the vehicle cabin, obtain the matching degree between the lip movement information of each of the N in-vehicle members at the N positions and the instruction information respectively, where N is an integer greater than 1; take the position where the in-vehicle member corresponding to the lip movement information with the highest matching degree is located as the target position.
[0012] There are various ways to obtain the matching degree. Generally, it can be obtained through a target matching model. The target feature matching model is a feature matching model trained with the lip movement information of the training user and one or more voice information (the voice information can be a voice waveform sequence or the text information corresponding to the voice) as the input, and the matching degrees between the lip movement information of the training user and the M voice information as the M labels.
[0013] In the way of training the model in the embodiments of the present invention, according to the different training sample data, the inference method of the model will also be different. When multiple groups of lip movement information are used as the input, such as 5 groups of lip movement information, then each of the 5 groups of lip movement information and each of the M voice information are constructed into M groups of samples as the input, and the matching degrees between the 5 groups of lip movement information and the voice information in the samples are used as the output, and the label information is the target output result during training. The model trained in this way also takes 5 groups of lip movement information and the target instruction information as the input during inference. When the number of in-vehicle personnel is less than 5, the lip information of the vacant positions can be input with default values, such as all 0 sequences, and the matching degree with the instruction information is output. It is also possible to use a single group of lip movement information and voice information as a sample, and the matching label is the target training result for training. In this way, when the trained model is performing inference, it is necessary to take the lip movement instructions of each in-vehicle member and the instruction information as the input, and obtain the matching degrees of multiple position members respectively.
[0014] In a possible implementation, the first type of instruction is a voice waveform sequence extracted from the audio data or text instruction information recognized according to the audio data.
[0015] In a possible implementation, the lip movement information of the in-vehicle members at N positions in the vehicle cabin is an image sequence of the lip movements of the in-vehicle members at the N positions in the vehicle cabin during the target time period.
[0016] In an embodiment of the present invention, the member who issues the voice related to the instruction is determined by the matching degree between the obtained voice instruction and the lip information of the vehicle occupants at each position. The matching degree between the instruction information and the lip movement information of the member is obtained through a matching model. Depending on the different matching models, the instruction information to be obtained is also different. According to the needs of model training, different lip movement information can be extracted from the video of lip movement. For example, it can be an image sequence of lip movement within the target time period or a vector parameter representing the temporal change of the distance between the upper and lower lips. The first type of instruction can also have various forms, which can be a voice waveform sequence or text information corresponding to the instruction.
[0017] In a possible implementation manner, the in-vehicle audio data is obtained from the audio data collected by a microphone in a specified position area inside the vehicle cabin.
[0018] In a possible implementation manner, the in-vehicle audio data is obtained based on target audio data selected from the audio data collected by multiple microphones in the vehicle cabin.
[0019] There are usually multiple microphones in the vehicle cabin, which are arranged at different positions in the vehicle cabin. Therefore, for the collection of in-vehicle audio data to obtain the best audio effect, the in-vehicle audio data mentioned in the embodiments of the present invention can be obtained after comprehensive processing of multiple audio data; it can also be the audio data with the optimal parameters selected according to a preset rule after comprehensively comparing the audio data collected by multiple microphones in the vehicle cabin, that is, the audio data with the best voice quality recorded; or the audio data collected by the microphone at the instruction position, such as the audio data collected by the microphone arranged in the central position area inside the vehicle.
[0020] In a possible implementation manner, the target feature matching model includes a first model, a second model, and a third model; inputting the instruction information and the lip movement information of the vehicle occupants at N positions in the vehicle cabin into the target feature matching model to obtain the matching degrees between the lip movement information of the vehicle occupants at N positions in the vehicle cabin and the instruction information respectively includes: inputting the instruction information into the first model to obtain voice features, where the voice features are K-dimensional voice features and K is an integer greater than 0; inputting the lip movement information of the vehicle occupants at N positions in the vehicle cabin into the second model to obtain N image sequence features, and each of the N image sequence features is a K-dimensional image sequence feature; inputting the voice features and the N image sequence features into the third model to obtain the matching degrees between the N image sequence features and the voice features respectively.
[0021] In a possible implementation, the target feature matching model includes a first model and a second model; the step of inputting the instruction information and the lip movement information of the in-vehicle members at N positions in the vehicle cabin into the target feature matching model to obtain the matching degrees between the lip movement information of the in-vehicle members at N positions in the vehicle cabin and the instruction information respectively includes: inputting the audio data into the first model to obtain the corresponding instruction information, inputting the lip movement information of the in-vehicle members at N positions in the vehicle cabin into the second model, where the lip movement information of each in-vehicle member corresponds to a set of image sequence features, and inputting the N image sequence features into the second model simultaneously or separately to obtain the instruction information corresponding to each lip movement information; based on the recognition results of the two models, determine the target position member that issues the instruction.
[0022] In one implementation of the model in the embodiment of the present invention, the instruction information output by the first model may be the identifier and matching degree corresponding to the instruction. The model selects the instruction identifier corresponding to the instruction with the highest matching degree. The identifier may be an instruction code or the text feature corresponding to the instruction. The above judgment is based on the output results of the first model and the second model. There are various judgment rules. For example, identify the position member that is the same as the instruction information output by the first model. If there are multiple members at the identified positions that are the same as the instruction information identified by the first model, then compare the matching degrees output by the second model and select the target position with the higher matching degree to execute the instruction, or select the instruction information with the highest matching degree identified in the second model and compare it with the instruction information of the first model. If the instruction information is the same, then determine the position corresponding to the instruction information with the highest matching degree in the second model as the target position.
[0023] In a possible implementation, generate the correspondence between the lip movement information of the in-vehicle members at N positions and the N positions; the step of using the position where the in-vehicle member corresponding to the lip movement information with the highest matching degree is located as the target position includes: obtaining the target lip movement information with the highest matching degree; determining the position corresponding to the target lip movement information as the target position according to the correspondence between the lip movement information of the in-vehicle members at N positions and the N positions.
[0024] The generation of the correspondence between the lip movement information of the in-vehicle members at N positions and the N positions may be to obtain the member relationship at each position through image acquisition in the vehicle cabin, and then correspond the lip movement information of each member with the position relationship. The image data may be the same image from which the lip information is extracted or an independent acquisition process.
[0025] In a possible implementation, a correspondence between the lip movement information of the in-vehicle members at the N positions and the identities of the in-vehicle members at the N positions is generated; the position where the in-vehicle member corresponding to the lip movement information with the highest matching degree is located is used as the target position, including: obtaining the target lip movement information with the highest matching degree; determining the target in-vehicle member according to the correspondence between the lip movement information of the in-vehicle members at the N positions and the identities of the in-vehicle members at the N positions; determining the position information of the target in-vehicle member as the target position, and the position information of the target in-vehicle member is determined according to the sensor data in the vehicle.
[0026] In the embodiments of the present invention, the first type of instruction mentioned may be an in-vehicle control instruction, which is mainly applicable to an instruction interaction scenario where it is necessary to identify position information in the vehicle cabin to determine the target area for executing the instruction. Since sometimes the in-vehicle user will give clear position information when executing a voice instruction, such as closing the window of the right rear seat, then in a possible implementation, this type of instruction with clear position information may be considered not to belong to the first type of instruction, and the first type of instruction may be an instruction that needs to distinguish position areas for execution but does not contain position area information in the instruction.
[0027] In a second aspect, the embodiments of the present invention further provide a method for in-vehicle voice control, including obtaining a second type of instruction and the lip movement information of an in-vehicle member at a first position in the vehicle, where the second type of instruction is obtained according to the target audio data collected in the vehicle cabin, and the lip movement information of the in-vehicle member at the first position is obtained when the second type of instruction is recognized from the target audio data, and the target time period is the time period corresponding to the second type of instruction in the audio data; matching the second type of instruction and the lip movement information of the in-vehicle member at the first position to obtain a matching result; when it is determined according to the matching result that the second type of instruction and the lip movement information of the in-vehicle member at the first position are matched, sending an indication information for instructing to execute the second type of instruction.
[0028] Embodiments of the present invention can be applied to refined voice command control inside a vehicle cabin. When it is recognized that the voice command is an instruction that can only be executed after determining the identity or permission of the command issuer, such as a driving control instruction of the vehicle, or other instructions involving privacy information operations, such as switching the driving mode, controlling the driving direction, viewing historical call data, etc., usually based on the special environment of the vehicle, we often default that the member in the driver's seat has the highest permission. Therefore, when the vehicle obtains a relevant instruction that requires confirmation of the issuer's identity, it determines whether the instruction is issued by the member in the driver's seat by obtaining the lip movement information of the member in a specific position (such as the driver's seat) and the specific instruction information recognized, so as to be able to perform operation control targeted at a specific position area. Since the above voice control method involves the processing and analysis of video data and the matching degree of its local features with the instruction information, it can occur locally, that is, the vehicle or an intelligent device on the vehicle executes the above method, or the above video data processing and matching actions can also be executed in the cloud. In addition to collecting the lip movement information of the member in the driver's seat, it can also be preset or manually set that members in one or more positions have the control permission for a certain type of instruction. When such an instruction is recognized, the lip movement information of the person in the corresponding position is obtained for matching analysis.
[0029] In a possible implementation manner, it can also be to identify the position of a specific user (such as the vehicle owner) in the current vehicle environment through face recognition technology, and use the position where the specific user is located as the first position. Usually, it is default to obtain the lip movement information of the member in the driver's seat.
[0030] The method of the embodiments of the present invention is applicable to vehicle scenarios with any number of people, especially applicable to scenarios where there are multiple members in the vehicle and multiple people are speaking at the same time. At this time, by combining the matching of lip movement information and instruction information with the in-vehicle position distribution information, it can accurately determine whether the voice instruction is issued by a member in a specific position, and then determine whether to execute the instruction. The matching method can be to obtain multiple members in the vehicle (including the member in the first position) for matching, and determine whether the matching degree of the member in the first position is the highest, or it can also be to only obtain the member in the first position for matching, and when the matching degree reaches the threshold, it is determined as a match and the instruction can be executed.
[0031] In an embodiment where a second type of instruction is judged, the situation where a specific position usually needs to be judged is a scenario where only members in a specific position have the permission to execute such instructions. For example, the second type of instruction is usually an instruction for vehicle control. To prevent misoperation, such instructions are usually set so that only members in a specific position, such as the driver's seat, have the ability to control the vehicle's driving through voice. The usual implementation method is to obtain the lip movement information of the user in the driver's seat and match it with the second instruction information. When the structure is matched, it is judged that the second type of instruction is issued by the user in the driver's seat, and then the second type of instruction is executed. Because it is usually default that the user in the driver's seat has the permission to control the vehicle. It is also possible to manually set that passengers in other positions have voice control permissions, then the first position is still other positions inside the vehicle.
[0032] It is also possible to set according to needs that certain instructions can only be performed by members in a specific position. Then, the execution rules of such instructions can also refer to the execution method of the second type of instruction to judge whether to execute.
[0033] In the embodiment of the present invention, the obtaining of the second type of instruction and the lip movement information of the in-vehicle member at the first position inside the vehicle can specifically be: obtaining the target audio data inside the vehicle cabin; when it is recognized that the target audio data includes the second type of instruction, obtaining the in-vehicle image data; extracting the lip movement information of the in-vehicle member at the first position from the in-vehicle image data.
[0034] In the implementation process of the above solution, the matching of the second type of instruction and the lip movement information of the in-vehicle member at the first position to obtain a matching result is specifically: determining the matching result of the second type of instruction and the lip movement information of the in-vehicle member at the first position according to the matching degree between the second type of instruction and the lip movement information of the in-vehicle member at the first position and a preset threshold.
[0035] The second type of instruction in the embodiment of the present invention can be a voice waveform sequence extracted from the audio data or text instruction information recognized according to the audio data. The lip movement information of the in-vehicle member is an image sequence of the lip movement of the in-vehicle member during the target time period.
[0036] In the implementation process, the embodiments of the present invention further include: when the audio data includes the second type of instruction, acquiring the image data of the in-vehicle members at N other positions in the vehicle; extracting the lip movement information of the in-vehicle members at the N other positions in the vehicle from the image data of the in-vehicle members at the N other positions in the vehicle during the target time period; matching the second type of instruction with the lip movement information of the in-vehicle member at the first position in the vehicle to obtain a matching result, specifically: matching the second type of instruction with the lip movement information of the in-vehicle member at the first position in the vehicle and the lip movement information of the in-vehicle members at the N positions to obtain the matching degrees between the lip movement information of N + 1 in-vehicle members and the second type of instruction respectively, and acquiring the lip movement information with the highest matching degree; when it is determined according to the matching result that the second type of instruction and the lip movement information of the in-vehicle member at the first position in the vehicle are matched, sending an indication information indicating the execution of the second type of instruction, specifically: when the lip movement information with the highest matching degree is the lip movement information of the in-vehicle member at the first position in the vehicle, sending the indication information indicating the execution of the second type of instruction.
[0037] In the embodiments of the present invention, the lip movement information of the in-vehicle members at the in-vehicle positions is extracted from the video data of the in-vehicle members at the in-vehicle positions. The specific method is to identify multiple face regions in the video data based on the face recognition algorithm, and extract the lip movement videos in each face region of the multiple face regions; and determine the lip movement information corresponding to each face based on the lip movement videos in each face region.
[0038] In a possible implementation manner, the in-vehicle audio data is obtained according to the data collected by multiple microphones in the vehicle, or the in-vehicle audio data is obtained according to the audio data collected by the microphones in the specified position area in the vehicle.
[0039] In a third aspect, an embodiment of the present invention provides a voice command control device, including a processor; the processor is configured to: obtain a first type of command and lip movement information of in-vehicle members at N positions in the vehicle cabin during a target time period, where the first type of command is obtained according to target audio data collected in the vehicle cabin, the lip movement information of the in-vehicle members is obtained when the first type of command is recognized from the target audio data, and the target time period is the time period corresponding to the first type of command in the audio data; match the first type of command with the lip movement information of the in-vehicle members at the N positions in the vehicle cabin, and obtain a target position according to the matching result between the lip movement information of the in-vehicle members at the N positions and the first type of command, where the target position is the position where the in-vehicle member whose lip movement information matches the first type of command according to the matching result is located; send indication information indicating that the first type of command is to be executed for the target position.
[0040] The embodiment of the present invention can be applied to refined voice command control in the vehicle cabin. When it is recognized that the voice command is a command that needs to further determine the position before execution, such as an operation command for a device in the vehicle cabin, such as playing a video, speaker adjustment, air conditioner adjustment, seat adjustment, etc., which are commands that need to determine the specific position targeted by the command for targeted local operations, by obtaining the lip movement information of the members at each position in the vehicle and the specific command information recognized, it is possible to determine which position member issued the command, so as to be able to perform operation control for a specific position area in a targeted manner. The above voice control device needs to process video data and analyze the matching degree between its local features and command information. Therefore, the above device can be a local device, such as an intelligent in-vehicle device, or an in-vehicle processor chip, or it can be an in-vehicle system including a microphone and a camera, or an intelligent vehicle. At the same time, according to different implementation methods of the solution, it can also be a cloud server, which performs the above video data processing and matching actions after obtaining the data of the in-vehicle camera and the in-vehicle speaker.
[0041] In a possible implementation, the processor is configured to obtain the target audio data in the vehicle cabin; when it is recognized that the target audio data includes a first type of instruction, obtain the image data in the vehicle cabin; and extract the lip movement information of the vehicle occupants at N positions in the vehicle cabin from the image data in the vehicle cabin. In the embodiments of the present invention, the processor extracts the lip movement information of the occupants at the N positions in the vehicle cabin from the image data in the vehicle cabin. Specifically, the processor may, based on a face recognition algorithm, recognize the N face regions in the video data, and extract the lip movement video or sample the video frame sequence in each of the N face regions; and determine the lip movement information of the N occupants based on the lip movement video or the video frame sequence in each of the N face regions. The image data is usually obtained by one or more cameras in the vehicle, and there can be various types of the cameras.
[0042] In a possible implementation, the processor is further configured to, when it is recognized that the number of vehicle occupants is greater than 1, extract the lip movement information of the vehicle occupants at N positions in the vehicle cabin from the image data in the vehicle cabin. In the specific implementation process of the embodiments of the present invention, in order to avoid wasting computing resources, after recognizing the first instruction type, the processor will determine the number of people in the vehicle cabin according to the obtained image data in the vehicle cabin. When there is only one person in the vehicle cabin, there is no need to extract the lip movement information. Only when the number of people is more than 1, the lip movement information of the people at the occupied positions is extracted.
[0043] In a possible implementation, the processor matches the first type of instruction with the lip movement information of the vehicle occupants at N positions in the vehicle cabin, and obtains the target position according to the matching result between the lip movement information of the vehicle occupants at the N positions and the first type of instruction. Specifically: the processor obtains the matching degree between the lip movement information of each of the vehicle occupants at the N positions and the instruction information according to the first type of instruction and the lip movement information of the N vehicle occupants in the vehicle cabin, where N is an integer greater than 1; and takes the position where the vehicle occupant corresponding to the lip movement information with the highest matching degree is located as the target position.
[0044] Among them, the processor can obtain the matching degree through different target matching models. The target feature matching model is a feature matching model trained with the lip movement information of the trained user and one or more voice information (the voice information can be a voice waveform sequence or the text information corresponding to the voice) as inputs, and with the matching degrees between the lip movement information of the trained user and the M voice information as M labels. The training of the model is usually performed on the cloud side and is independent of the use of the model. That is, the model is trained on one device, and after training is completed, it is sent to the device that needs to perform instruction matching for running and using.
[0045] In the method for training the model in the embodiments of the present invention, according to different training sample data, the inference method of the model will also be different. When multiple sets of lip movement information are used as input, such as 5 sets of lip movement information, then each of the 5 sets of lip movement information is combined with each of the M sets of speech information to form M sets of samples as input. The matching degrees between the 5 sets of lip movement information and the speech information within the samples are used as output, and the label information is the target output result during training. The model trained in this way also uses 5 sets of lip movement information and target instruction information as input during inference. When the number of people in the vehicle is less than 5, the lip information of the vacant positions can be input with default values, such as an all-0 sequence, and the matching degree between the output and the instruction information is obtained. It is also possible to use a single set of lip movement information and speech information as a sample, with the matching label as the target training result for training. When the model trained in this way is used for inference, the lip movement instructions and instruction information of each vehicle occupant need to be used as input to obtain the matching degrees of multiple position members respectively.
[0046] In a possible implementation manner, the processor is further configured to generate the correspondence between the lip movement information of the vehicle occupants at the N positions and the N positions; the processor determines the position where the vehicle occupant corresponding to the lip movement information with the highest matching degree is located as the target position, including: the processor obtains the target lip movement information with the highest matching degree; and determines the position corresponding to the target lip movement information as the target position according to the correspondence between the lip movement information of the vehicle occupants at the N positions and the N positions.
[0047] Regarding the correspondence between the lip movement information of the vehicle occupants and their positions, the processor can identify the member relationships at each position through the images collected inside the vehicle, and then correspond the lip movement information of each member with the position relationship. The image data can be the same image from which the lip information is extracted or an independent collection process.
[0048] In a possible implementation manner, the processor generates the correspondence between the lip movement information of the vehicle occupants at the N positions and the identities of the vehicle occupants at the N positions; the processor determines the position where the vehicle occupant corresponding to the lip movement information with the highest matching degree is located as the target position, including: the processor obtains the target lip movement information with the highest matching degree; determines the target vehicle occupant according to the correspondence between the lip movement information of the vehicle occupants at the N positions and the identities of the vehicle occupants at the N positions; and determines the position information of the target vehicle occupant as the target position, where the position information of the target vehicle occupant is determined according to the sensor data inside the vehicle.
[0049] In the embodiments of the present invention, the first type of instruction mentioned may be an in-vehicle control instruction, which is mainly applicable to instruction interaction scenarios in the vehicle cabin where it is necessary to identify position information to determine the target area for executing the instruction. Sometimes, when a vehicle user executes a voice instruction, they will give clear position information, such as closing the window of the right rear seat. In this case, the processor can directly identify the target position of the instruction. Therefore, in a possible implementation manner, when the processor identifies that the obtained instruction is an instruction with clear position information of this type, the processor will determine that this instruction does not belong to the first type of instruction. The first type of instruction may be an instruction that needs to distinguish position areas for execution but does not contain position area information in the instruction.
[0050] Fourthly, embodiments of the present invention provide a voice instruction control device, including a processor; the processor is configured to: obtain a second type of instruction and lip movement information of a vehicle occupant at a first position inside the vehicle, where the second type of instruction is obtained according to target audio data collected inside the vehicle cabin, and the lip movement information of the vehicle occupant at the first position is obtained when the second type of instruction is recognized from the target audio data, and the target time period is the time period corresponding to the second type of instruction in the audio data; match the second type of instruction and the lip movement information of the vehicle occupant at the first position to obtain a matching result; when it is determined according to the matching result that the second type of instruction and the lip movement information of the vehicle occupant at the first position are matched, send an indication information indicating to execute the second type of instruction.
[0051] Embodiments of the present invention can be applied to refined voice instruction control in the vehicle cabin. When a voice instruction is recognized as an instruction that needs to determine the identity or permission of the instruction issuer before execution, such as a vehicle driving control instruction, or other instructions involving privacy information operations, such as switching the driving mode, controlling the driving direction, viewing historical call data, etc., usually based on the special environment of the vehicle, we often default that the occupant in the driver's seat has the highest permission. Therefore, when the vehicle obtains relevant instructions that need to confirm the identity of the issuer, by obtaining the lip movement information of the occupant at a specific position (such as the driver's seat) and the specific instruction information recognized, it is determined whether the instruction is issued by the occupant in the driver's seat, so as to be able to perform operation control on a specific position area in a targeted manner. Since the above voice control device involves the processing and analysis of video data and the matching degree of its local features with instruction information, it can be a local device, such as an in-vehicle intelligent device, an in-vehicle intelligent chip, an in-vehicle system including a camera and a microphone, or an intelligent vehicle, or it can be a cloud-side cloud processor. In addition to collecting the lip movement information of the occupant in the driver's seat, the processor can also collect the lip movement information of the occupants at one or more positions according to system presets or manual settings for targeted collection. When such an instruction is recognized, the lip movement information of the personnel at the corresponding position is obtained for matching analysis.
[0052] In a possible implementation, the first position is the driver's seat.
[0053] In a possible implementation, the second type of instruction is a vehicle driving control instruction.
[0054] The device according to the embodiment of the present invention can be used in a vehicle scenario with any number of people, especially applicable to a scenario where there are multiple members in the vehicle and multiple people are speaking simultaneously, for instruction recognition and execution judgment. At this time, the processor can accurately determine whether a voice instruction is issued by a member at a specific position by combining the matching situation between the lip movement information and the instruction information to obtain the in-vehicle position distribution information, and then determine whether to execute the instruction. The matching method can be to obtain multiple members in the vehicle (including the member at the first position) for matching to determine whether the member at the first position has the highest matching degree, or to only obtain the member at the first position for matching, and when the matching degree reaches the threshold, it is determined as a match and the instruction can be executed.
[0055] In an embodiment where the processor needs to judge a specific position for the second type of instruction, the situation where the processor needs to judge a specific position is for a scenario where only members at a specific position have the permission to execute such instructions. For example, the second type of instruction is usually a vehicle control instruction, and such instructions are usually set so that only members at a specific position such as the driver's seat have the ability to control the vehicle driving through voice to prevent misoperation. The general implementation method is that the processor obtains the lip movement information of the user in the driver's seat and the second instruction information for matching. When the structure is a match, it is judged that the second type of instruction is issued by the user in the driver's seat, and thus the second type of instruction is executed. Because usually, it is default that the user in the driver's seat has the control permission of the vehicle. It is also possible to manually set that passengers in other positions have voice control permissions, then the first position is still other positions in the vehicle.
[0056] In a possible implementation, the processor obtains the second type of instruction and the lip movement information of the in-vehicle member at the first position in the vehicle, specifically: the processor obtains the target audio data in the vehicle cabin; when it is recognized that the target audio data includes the second type of instruction, it obtains the image data in the vehicle cabin; and extracts the lip movement information of the in-vehicle member at the first position from the image data in the vehicle cabin.
[0057] During the implementation process of the above solution, the processor matches the second type of instruction and the lip movement information of the in-vehicle member at the first position to obtain a matching result, specifically: the processor determines the matching result of the second type of instruction and the lip movement information of the in-vehicle member at the first position according to the matching degree between the second type of instruction and the lip movement information of the in-vehicle member at the first position and a preset threshold.
[0058] The second type of instruction in the embodiments of the present invention may be a speech waveform sequence extracted from the audio data or text instruction information recognized from the audio data. The lip movement information of the in-vehicle member is an image sequence of the lip movement of the in-vehicle member during the target time period.
[0059] In a possible implementation manner, the processor is further configured to, when the audio data includes the second type of instruction, obtain the image data of the in-vehicle members at other N positions in the vehicle; extract the lip movement information of the in-vehicle members at other N positions in the vehicle from the image data of the in-vehicle members at other N positions in the vehicle during the target time period; the processor matches the second type of instruction with the lip movement information of the in-vehicle member at the first position in the vehicle to obtain a matching result, specifically, the processor matches the second type of instruction with the lip movement information of the in-vehicle member at the first position in the vehicle and the lip movement information of the in-vehicle members at the N positions to obtain the matching degrees between the lip movement information of N + 1 in-vehicle members and the second type of instruction respectively, and obtains the lip movement information with the highest matching degree; when the processor determines that the second type of instruction and the lip movement information of the in-vehicle member at the first position in the vehicle are matched according to the matching result, it sends an instruction information indicating the execution of the second type of instruction, specifically: when the lip movement information with the highest matching degree is the lip movement information of the in-vehicle member at the first position in the vehicle, the processor sends an instruction information indicating the execution of the second type of instruction.
[0060] In the embodiments of the present invention, to extract the lip movement information of the members at the in-vehicle positions from the video data of the members at the in-vehicle positions, the specific method is to identify multiple face regions in the video data based on a face recognition algorithm, and extract the lip movement videos in each face region among the multiple face regions; and determine the lip movement information corresponding to each face based on the lip movement videos in each face region.
[0061] In a fifth aspect, an embodiment of the present invention provides a chip system, which includes at least one processor, a memory, and an interface circuit. The memory, the interface circuit, and the at least one processor are interconnected through a line. Instructions are stored in the at least one memory; the instructions are executed by the processor to implement the method according to any one of the first and second aspects.
[0062] In a sixth aspect, an embodiment of the present invention provides a computer-readable storage medium, which is used to store program codes, and the program codes include the method for executing any one of the first and second aspects.
[0063] In a seventh aspect, an embodiment of the present invention provides a computer program, the computer program includes instructions, and when the computer program is executed, it is used to implement the method of any one of the first and second aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the following will describe the drawings required to be used in the embodiments of the present application or the background art.
[0065] Figure 1 FIG. is a schematic diagram of a scenario of multi-person interaction in a vehicle provided by an embodiment of the present invention.
[0066] Figure 2 FIG. is a schematic diagram of a scenario of multi-person interaction in a vehicle provided by an embodiment of the present invention.
[0067] Figure 3 An embodiment of the present invention provides a system architecture 100.
[0068] Figure 4 FIG. is a schematic diagram of a convolutional neural network provided by an embodiment of the present invention.
[0069] Figure 5 FIG. is a flowchart of a method for training a neural network provided by an embodiment of the present invention.
[0070] Figure 6 FIG. is an example diagram of a sound waveform provided by an embodiment of the present invention.
[0071] Figure 7A FIG. is a method for matching voice commands provided by an embodiment of the present invention.
[0072] Figure 7B FIG. is a method for matching voice commands provided by an embodiment of the present invention.
[0073] Figure 8 FIG. is a schematic diagram of a cloud interaction scenario provided by an embodiment of the present invention.
[0074] Figure 9 FIG. is a flowchart of a method according to an embodiment of the present invention.
[0075] Figure 10 FIG. is a flowchart of a method according to an embodiment of the present invention.
[0076] Figure 11 FIG. is a flowchart of a method according to an embodiment of the present invention.
[0077] Figure 12 FIG. is a flowchart of a method according to an embodiment of the present invention.
[0078] Figure 13It is a schematic structural diagram of an instruction control device provided by an embodiment of the present invention.
[0079] Figure 14 It is a schematic structural diagram of a training device for a neural network provided by an embodiment of the present invention.
[0080] Figure 15 It is another instruction control system provided by an embodiment of the present invention. Detailed implementation manners
[0081] Next, the embodiments of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0082] The terms "first", "second", "third", and "fourth" in the description and claims of this application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include unlisted steps or units, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.
[0083] Referring to "embodiments" herein means that specific features, structures, or characteristics described in conjunction with the embodiments may be included in at least one embodiment of this application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0084] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, an application running on a computing device and the computing device can both be components. One or more components can reside in a process and / or an execution thread, and the components can be located on one computer and / or distributed between two or more computers. In addition, these components can execute from various computer-readable media storing various data structures. The components can communicate, for example, according to signals having one or more data packets (such as data from two components interacting with each other between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems through signals) through local and / or remote processes.
[0085] First, some terms in this application are explained to facilitate understanding by those skilled in the art.
[0086] (1) Bitmap: Also known as raster graphics or dot matrix graphics, it is an image represented by a pixel array (pixel-array / dot-matrix). According to the bit depth, bitmaps can be divided into 1-bit, 4-bit, 8-bit, 16-bit, 24-bit, and 32-bit images, etc. The more bits of information used by each pixel, the more colors are available, the more realistic the color representation is, and the larger the corresponding data volume. For example, a pixel bitmap with a bit depth of 1 has only two possible values (black and white), so it is also called a binary bitmap. An image with a bit depth of 8 has 2^8 (i.e., 256) possible values. A grayscale image with a bit depth of 8 has 256 possible gray values. An RGB image consists of three color channels. Each channel in an 8-bit / channel RGB image has 256 possible values, which means that the image has more than 16 million possible color values. Sometimes an RGB image with 8 bits per channel (bpc) is called a 24-bit image (8 bits x 3 channels = 24 bits of data per pixel). [2] A bitmap represented by 24-bit RGB combined data bits is usually called a true color bitmap.
[0087] (2) Automatic Speech Recognition (ASR), also known as automatic speech recognition, aims to convert the lexical content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.
[0088] (3) Voiceprint is the sound wave spectrum carrying speech information displayed by electroacoustic instruments and is a biometric feature composed of more than a hundred characteristic dimensions such as wavelength, frequency, and intensity. Voiceprint recognition aims to identify an unknown voice by analyzing the characteristics of one or more speech signals. Simply put, it is a technology for identifying whether a certain sentence is spoken by a certain person. The identity of the speaker can be determined through the voiceprint, and then a targeted response can be made.
[0089] (4) Mel Frequency Cepstrum Coefficient (MFCC) In the field of sound processing, Mel-Frequency Cepstrum is a linear transformation of the logarithmic energy spectrum based on the non-linear Mel scale of sound frequency. Mel Frequency Cepstrum Coefficient (MFCC) is widely used in the function of speech recognition.
[0090] (5) Multi-way cross-Entropy Loss. Cross-entropy describes the distance between two probability distributions. The smaller the cross-entropy, the closer the two distributions are.
[0091] (6) Neural network
[0092] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes x s and the intercept 1 as inputs. The output of this operation unit can be:
[0093]
[0094] where s = 1, 2,..., n, n is a natural number greater than 1, Ws is the weight of xs, b is the bias of the neural unit, and f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many such single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.
[0095] (7) Deep neural network
[0096] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with many hidden layers. Here, "many" does not have a specific measurement standard. Dividing DNN according to the positions of different layers, the neural network inside DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i + 1-th layer. Although DNN seems very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: where is the input vector, is the output vector, b is the offset vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer simply performs such a simple operation on the input vector to obtain the output vector Since the DNN has many layers, the number of coefficients W and bias vectors b is also very large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscripts correspond to the index 2 of the output third layer and the index 4 of the input second layer. In summary: The coefficient from the kth neuron in the (L-1)th layer to the jth neuron in the Lth layer is defined as It should be noted that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is also a process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0097] (8) Convolutional Neural Network
[0098] A convolutional neural network (CNN, convolutional neuron network) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as using a trainable filter to convolve with an input image or a convolutional feature plane (featuremap). A convolutional layer refers to the neuron layer in a convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can only be connected to some neighboring layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weight here is the convolutional kernel. Sharing weights can be understood as a way of extracting image information that is independent of position. The underlying principle here is that the statistical information of a certain part of an image is the same as that of other parts. That is to say, the image information learned in one part can also be used in another part. Therefore, for all positions on the image, the same learned image information can be used. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels there are, the richer the image information reflected by the convolution operation.
[0099] The convolutional kernel can be initialized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolutional kernel can learn reasonable weights. In addition, the direct benefit brought by sharing weights is to reduce the connections between layers of the convolutional neural network and at the same time reduce the risk of overfitting.
[0100] (9) Loss function
[0101] During the process of training a deep neural network, since it is desired that the output of the deep neural network is as close as possible to the value that is truly wanted to be predicted, the weight vectors of each layer of the neural network can be updated by comparing the predicted value of the current network with the truly desired target value and then according to the difference between the two (of course, there is usually a process before the first update, that is, configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is high, the weight vector is adjusted to make it predict lower, and continuous adjustment is made until the deep neural network can predict the truly desired target value or a value very close to the truly desired target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function (loss function) or objective function (objective function), and they are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0102] (10) Backpropagation algorithm
[0103] The convolutional neural network can use the error backpropagation (BP) algorithm to correct the size of the parameters in the convolutional neural network during the training process, so that the reconstruction error loss of the convolutional neural network becomes smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the error loss information is propagated backward to update the parameters in the convolutional neural network, so that the error loss converges. The backpropagation algorithm is a backward propagation movement dominated by the error loss, aiming to obtain the optimal parameters of the convolutional neural network, such as the weight matrix.
[0104] (11) Pixel value
[0105] The pixel value of an image can be an RGB (Red, Green, Blue) color value, and the pixel value can be a long integer representing the color. For example, the pixel value is 256*Red + 100*Green + 76*Blue, where Blue represents the blue component, Green represents the green component, and Red represents the red component. Among the respective color components, the smaller the value, the lower the brightness, and the larger the value, the higher the brightness. For a grayscale image, the pixel value can be a grayscale value.
[0106] First, to facilitate the understanding of the embodiments of the present invention, the technical problems to be specifically solved by this application are further analyzed and proposed. In the prior art, there are various implementation methods for identifying the location of the issuer of a detected voice command in a scenario with multiple people in the vehicle. For example, it can be achieved through voiceprint recognition and / or sound source localization in the vehicle cabin. The following are two commonly used solutions listed by way of example. Among them,
[0107] Solution 1:
[0108] A voiceprint is the sound wave spectrum carrying speech information displayed by electroacoustic instruments, and it is a biometric feature composed of more than a hundred characteristic dimensions such as wavelength, frequency, and intensity. Voiceprint recognition aims to identify unknown sounds through the feature analysis of one or more voice signals. Simply put, it is a technology for identifying whether a certain sentence is spoken by a certain person. The speaker's identity can be determined through the voiceprint, and then a targeted response can be made. It is mainly divided into two stages: the registration stage and the verification stage. Among them, in the registration stage: according to the voiceprint characteristics of the speaker's voice, a corresponding voiceprint model is established; in the verification stage: receive the speaker's voice, extract its voiceprint characteristics and match them with the registered voiceprint model. If the match is successful, it proves that it is the original registered speaker.
[0109] Solution 2:
[0110] Sound source localization technology is a technology that uses acoustic and electronic devices to receive target sound field information to determine the position of the target sound source. The sound source localization of a microphone array refers to picking up the sound source signal with the microphone array, analyzing and processing multiple channels of sound signals, and determining one or more sound source planes or spatial coordinates in the spatial domain, that is, obtaining the position of the sound source. Further control the beam of the microphone array to align with the speaker.
[0111] Disadvantages of Solution 1 and Solution 2 applied to the analysis of the position of the command issuer in the vehicle cabin:
[0112] For the application of voiceprint recognition, first, the voiceprint information of the passengers needs to be stored in advance. If a person has not undergone voiceprint recognition and recording, they cannot be recognized. At the same time, the voice of the same person is variable, and the vehicle cabin is a multi-person environment. When multiple people speak simultaneously, it is difficult to extract voiceprint characteristics, or environmental noise will also interfere with the recognition.
[0113] For sound source localization technology, since the vehicle cabin is a relatively narrow and crowded space, especially the space distance between the rear passengers is very close, and the members will also shake or tilt their bodies when speaking. The above factors will all lead to a decrease in the accuracy of sound source localization. At the same time, there are often multiple people speaking simultaneously in the vehicle cabin, which will also affect the accuracy of sound source localization.
[0114] In summary, if the above two solutions are applied to the recognition of the position of the voice command issuer in the vehicle, especially to the recognition of the position of the command issuer in the in-vehicle scenario where multiple people speak simultaneously, it will be impossible to accurately identify which in-vehicle member at which position the collected command is from. Therefore, it is impossible to achieve a more accurate and effective human-machine interaction. Therefore, the technical problems to be solved by this application include the following aspects: in the case of multiple users in the vehicle cabin, when a specific type of command is collected, how to accurately determine the specific position of the voice issuer and issue corresponding commands accordingly.
[0115] The voice matching method provided by the embodiments of this application can be applied to the human-machine interaction scenario of intelligent vehicles. The following are exemplary lists of the human-machine interaction scenarios to which the voice command control method in this application can be applied, which can include the following two scenarios.
[0116] In-vehicle interaction scenario one:
[0117] Usually, there are multiple speakers distributed in the vehicle, which are distributed at different positions in the vehicle cabin. The multiple speakers can provide music with different volume levels for passengers in different areas of the vehicle according to the needs of passengers and the driver. For example, if passenger A wants to rest, a quiet environment is needed, so the volume of the speaker in the area where he is located can be adjusted to the lowest. And if passenger B needs to listen to music normally, the speaker in the area where he is located can be set to a normal size. Or the multiple speakers can also provide different audio playback contents for users in different areas. For example, if there are children sitting in the back row, fairy tales can be selected for playback in the back row, while the driver and the co-pilot in the front row want to listen to pop music, then pop music can be played on the speakers in the front row area.
[0118] The embodiments of the present invention can provide a method for a member in the vehicle cabin to control the speaker in the area where the voice command issuer is located by voice. For example, as Figure 1 shown, when there are four people, A, B, C, and D, in the vehicle, sitting in the driver's seat, co-pilot seat, left rear row, and right rear row respectively. At this time, member D says: "Turn the volume down". At this time, as Figure 7A shown, the audio command information and video information of multiple in-vehicle members can be obtained through the cameras and microphones in the vehicle cabin respectively. By feature matching of the lip movement information and command information of the in-vehicle members, the member who is speaking is determined, and based on the command and position issued by the member who is speaking, the speaker at the position of the speaking member is controlled. If the speaking member is identified as the passenger in the right rear, the speaker in the right rear area is turned down.
[0119] If the voice command of member C is: "Play a song **** for me", the audio command information and video information of the in-vehicle members can also be obtained through the camera and microphone in the vehicle cabin respectively. By processing and analyzing the lip movement information and voice information of in-vehicle members A, B, C, and D, it is determined that the speaking member is member C, and based on the command issued by the speaking member C and the position where the member is located, the speaker at the position of the speaking member is controlled. If it is recognized that the speaking member C is the passenger in the left rear, the speaker in the left rear is controlled to play the song ****.
[0120] Similar application scenarios can also include that there are multiple air vents distributed in the vehicle, which are distributed at different positions in the vehicle cabin. The multiple air vents can provide different air volumes for the passengers in different areas of the vehicle according to the needs of the passengers and the driver, realizing differential adjustment of the temperature in the local area. For example, passenger A feels a little cold, so the air volume in the area where he is located can be increased, while passenger B feels cold, and the air outlet direction of the air vent in the area where he is located can be adjusted by command so that the air does not directly blow on people or the air volume can be reduced. Or when the seats in the vehicle can independently adjust the angle and height respectively, the passengers in different areas will also adjust the various parameters of the seats according to their own needs. In the above scenarios, the voice recognition method of the embodiments of the present invention can be used for convenient control. Similarly, the audio command information and video information of the in-vehicle members can be obtained through the camera and microphone in the vehicle cabin respectively. By processing and analyzing the lip movement information and voice information of the in-vehicle members, the speaking member is determined, and based on the command issued by the speaking member and the position where the member is located, the air outlet direction or air volume of the air vent at the position of the speaking member is controlled, or the backrest angle of the seat or the height, front and back of the seat is controlled.
[0121] In-vehicle interaction scenario two:
[0122] In addition to the control of the in-vehicle settings mentioned in the previous scenarios for the voice command control in the vehicle cabin, since the control of some in-vehicle facilities needs to distinguish the target area where the specific command is implemented, it is necessary to identify which position the member who issued the voice command is. In addition to the above scenarios, when the driver wants to execute the driving control of the vehicle, he can also choose to control the driving of the vehicle by voice. In this voice command interaction scenario, it is also necessary to identify whether the current vehicle control command is issued by the member in the driver's seat.
[0123] Therefore, the embodiments of the present invention can provide a method for identifying the authority of voice commands for the driving control of the vehicle. For example, when there are multiple people in the vehicle, if a voice command related to vehicle driving control is received at this time, such as, "Switch to the autopilot mode", and the vehicle system defaults that only the member in the driver's seat has the execution authority for such commands, at this time the vehicle needs to Figure 2 obtain the lip movement information of the member in the driver's seat as shown, and asFigure 7B Match the obtained lip movement information and voice command information of the driver's seat member to obtain the matching degree with the command information, so as to judge whether the voice command is issued by the driver's seat member, and thus judge whether to execute the command.
[0124] Specifically, to judge whether the voice command is issued by the driver's seat member, it can also be to obtain the lip movement information of multiple members in the vehicle, analyze the matching degree between it and the command information, and see whether the matching degree of the lip movement information of the member in the driver's position is the highest, and then judge whether to execute the command.
[0125] It can be understood that Figure 1 、 Figure 2 The application scenarios in the vehicle cabin are only several exemplary implementation manners in the embodiments of the present invention. When the embodiments of the present invention are specifically implemented, there can be various flexible implementation manners. For example, for scenario one, it is not necessary to obtain the lip movement information of all members in the vehicle. It may only obtain the lip movement information of some members according to the specific command type. For example, when only the front row seats are adjustable and a seat adjustment command is detected, only the lip movement information of the front row members is obtained. For scenario two, it is not necessarily to obtain the lip movement information of the driver's seat member. When the vehicle defaults that the owner has the operation authority for the recognized command, obtain the position where the owner is located, extract the lip movement information of the owner, and judge whether the command is issued by the owner.
[0126] Since the matching of the command information and the lip movement information of the vehicle members can be carried out by using the method of model training to obtain a model and output the corresponding matching degree by inputting the lip movement information and the command information, the method provided by this application will be described below from the model training side and the model application side:
[0127] Any training method of a neural network provided by this application involves the fusion processing of computer audition and vision, and can be specifically applied to data processing methods such as data training, machine learning, and deep learning. It performs symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on training data (such as the lip movement information of training users and M voice messages in this application), and finally obtains a trained target feature matching model. Moreover, any voice matching method provided by this application can use the above-mentioned trained target feature matching model to input input data (such as the voice message to be recognized and the lip movement information of N users in this application) into the trained target feature matching model to obtain output data (such as the matching degrees between the lip movement information of N users and the voice message to be recognized in this application). It should be noted that a training method of a neural network and a voice matching method provided by an embodiment of this application are inventions generated based on the same concept, and can also be understood as two parts of a system or two stages of an overall process: such as a model training stage and a model application stage.
[0128] See the appendix Figure 3 , Figure 3 FIG. is a system architecture 100 provided by an embodiment of the present invention. As shown in the system architecture 100, the data acquisition device 160 is used to acquire training data. In this application, the data acquisition device 160 may include a microphone and a camera. The training data (i.e., the input data on the model training side) in the embodiment of the present invention may include: video sample data and voice sample data, that is, the lip movement information of training users and M voice messages in the embodiment of the present invention respectively. Among them, the M voice messages may include voice messages matching the lip movement information of the training users. For example, the video sample data is a lip movement image sequence of a certain training user when uttering the voice: "The weather is particularly good today. Where shall we go to play?" And the voice sample data is a voice waveform sequence (as a voice positive sample) containing the above-mentioned training user uttering "The weather is particularly good today. Where shall we go to play?" and (M - 1) other voice waveform sequences (as voice negative samples). And the above video sample data and audio sample data can be acquired by the data acquisition device 160 or downloaded from the cloud. Figure 3 This is just an exemplary architecture and is not limited thereto. Further, the data acquisition device 160 stores the training data in the database 130, and the training device 120 trains a target feature matching model / rule 101 based on the training data maintained in the database 130 (the target feature matching model 101 here is the target feature matching model in the embodiment of the present invention. For example, it is a model trained through the above training stage and can be a neural network model for feature matching between voice and lip movement trajectories).
[0129] The following will describe in more detail how the training device 120 obtains the target feature matching model / rule 101 based on the training data. The target feature matching model / rule 101 can be used to implement any one of the voice matching methods provided in the embodiments of the present invention. That is, the audio data and image data obtained by the data acquisition device 160 are input into the target feature matching model / rule 101 after relevant preprocessing, and the matching degrees / confidences between the image sequence features of the lip movements of multiple users and the voice features to be recognized can be obtained. The target feature matching model / rule 101 in the embodiments of the present invention can specifically be a spatio-temporal convolutional network (STCNN). In the embodiments provided in this application, the spatio-temporal convolutional network can be obtained by training a convolutional neural network. It should be noted that in actual applications, the training data maintained in the database 130 may not necessarily all come from the acquisition of the data acquisition device 160, and it may also be received from other devices. Additionally, it should be noted that the training device 120 does not necessarily completely train the target feature matching model / rule 101 based on the training data maintained in the database 130, and it may also obtain training data from the cloud or other places for model training. The above descriptions should not be regarded as limitations on the embodiments of the present invention.
[0130] As Figure 3 shown, the target feature matching model / rule 101 is obtained by training the training device 120. The target feature matching model / rule 101 can be referred to as an audiovisual cross convolutional neural network (V&A CrossCNN) / spatio-temporal convolutional network in the embodiments of the present invention. Specifically, the target feature matching model provided in the embodiments of the present invention may include: a first model, a second model, and a third model. The first model is used to extract voice features, the second model is used to extract the image sequence features of the lip movements of multiple users (N users in this application), and the third model is used to calculate the matching degrees / confidences between the above voice features and the image sequence features of N users. In the target feature matching model provided in the embodiments of the present invention, the first model, the second model, and the third model can all be convolutional neural networks, that is, it can be understood that the target feature matching model / rule 101 itself can be regarded as an overall spatio-temporal convolutional network, and this spatio-temporal convolutional network contains multiple independent networks, such as the above first model, second model, and third model.
[0131] In addition to the above model training and execution methods, the embodiments of the present invention can also be implemented through other model training and execution solutions.
[0132] Similar to the source of sample data collection in the training method described above, the training data (i.e., the input data on the model training side) in the embodiments of the present invention may include: video sample data and voice sample data, which are respectively the lip movement information and M voice information of the training users in the embodiments of the present invention. Among them, the lip movement information includes the lip movement information corresponding to various voice command statements of different users, and the voice information includes the voice command statements issued by different users. Optionally, it may also include some negative samples, that is, the lip movement information corresponding to statements that are not voice commands, and the voice information that is not a voice command. Here, the voice command refers to the voice information that the vehicle-mounted system can recognize and respond to, which can be a keyword or a complete sentence. The above video sample data and audio sample data can be collected by the data collection device 160, downloaded from the cloud, or provided by a third-party data holder. Further, the data collection device 160 stores the training data in the database 130, and the training device 120 trains the target feature matching model / rule 101 based on the training data maintained in the database 130 (the target feature matching model 101 here is the target feature matching model in the embodiments of the present invention. For example, it is a neural network model that can be used for feature matching between voice and lip movement trajectories and is trained through the above training stage).
[0133] Next, it will be described in more detail how the training device 120 obtains the target feature matching model / rule 101 based on the training data. The target feature matching model / rule 101 can be used to implement any one of the voice matching methods provided in the embodiments of the present invention, that is, by preprocessing the audio data and image data obtained by the data collection device 160 and inputting them into the target feature matching model / rule 101, the matching degrees / confidences between the image sequence features of the lip movements of multiple users and the voice features to be recognized can be obtained. The target feature matching model / rule 101 in the embodiments of the present invention can specifically be a convolutional network (CNN) in the embodiments provided in the present application. It should be noted that in actual applications, the training data maintained in the database 130 may not necessarily all come from the collection of the data collection device 160, and it may also be received from other devices. Additionally, it should be noted that the training device 120 may not necessarily train the target feature matching model / rule 101 completely based on the training data maintained in the database 130, and it may also obtain training data from the cloud or other places for model training. The above description should not be used as a limitation to the embodiments of the present invention.
[0134] Such as Figure 3As shown, the target feature matching model / rule 101 is trained according to the training device 120. Specifically, the target feature matching model provided in the embodiments of the present invention may include: a first model and a second model. The first model is used to perform matching recognition of voice commands to identify the command information corresponding to the voice commands, which may specifically be a command identifier or the text feature of the command. The second model is used to respectively identify the corresponding relationship between the voice commands corresponding to each lip movement information based on the image sequence features of N users. For example, it can match and output the identifier of the corresponding command and its matching degree. Finally, it outputs the target user who issues the voice command according to the command identifier corresponding to the voice command and the voice identifier corresponding to each user's lip movement information and its matching degree. In the target feature matching model provided in the embodiments of the present invention, the first model and the second model may be CNN, RNN, DBN, DNN, etc.
[0135] The training of the first model uses the voice command as the input and the command identifier corresponding to the voice command (the manifestation form of the identifier may be a code) as the label for training. The second model uses the lip movement information of the user as the input (the lip movement information may specifically be the lip movement image sequence feature, such as the opening and closing amplitude of the lip sampled according to time as a vector sequence), and the command identifier corresponding to the lip movement information and its matching degree as the output. The command identifier may be the code corresponding to the command, and the matching degree may be the output matching value. A judgment of whether to match is made according to the matching value. For example, a value greater than 0.5 is a match, and a value less than 0.5 is a non-match.
[0136] The target feature matching model / rule 101 trained according to the training device 120 can be applied to different systems or devices, such as applied to Figure 4 the execution device 110 shown. The execution device 110 may be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, augmented reality (AR) / virtual reality (VR), a smart wearable device, a smart robot, a vehicle terminal, a smart cockpit environment, etc., or it may also be a server or the cloud, etc. In the appendix Figure 4Among them, the execution device 110 is configured with an I / O interface 112 for data interaction with external devices. Users can input data to the I / O interface 112 through the client device 140 (the client device in this application may also include data acquisition devices such as microphones and cameras). The input data (i.e., the input data on the model application side) may include, in the embodiments of the present invention: the voice information to be recognized and the lip movement information of N users, that is, the voice waveform sequence within the target time period in the embodiments of the present invention and the lip movement information of each of the N users. The lip movement information of each user includes the image sequence of the user's lip movement within the target time period. For example, if it is currently necessary to recognize which person in a group of people spoke the voice information "What's the weather like tomorrow and where is suitable for traveling", then the voice waveform sequence corresponding to "What's the weather like tomorrow and where is suitable for traveling" and the image sequences of the lip movements of all people present are used as input data. It can be understood that the input data here can be input by users or provided by relevant databases, depending on different application scenarios, and the embodiments of the present invention do not make specific limitations on this.
[0137] In an embodiment of the present invention, the client device 140 and the execution device 110 may be on the same device, and the data acquisition device 160, the database 130, and the training device 120 may also be on the same device as the execution device 110 and the client device 140. Taking the execution entity in this application as a robot as an example, after the robot extracts the collected audio data and image data through the client device 140 (including a microphone, a camera, and a processor) to obtain the voice information to be recognized and the lip movement information of N users, the execution device 110 inside the robot can further perform feature matching between the extracted voice information and lip movement information, and finally output the result to the client device 140. The processor in the client device 140 analyzes to obtain the target user to whom the voice information to be recognized belongs among the N users. Moreover, the devices on the model training side (the data acquisition device 160, the database 130, and the training device 120) may be inside the robot or in the cloud. When they are inside the robot, it can be considered that the robot has the function of realizing model training or model update and optimization. At this time, the robot has both the functions of the model training side and the model application side. When in the cloud, it can be considered that the robot side only has the function of the model application side. Optionally, the client device 140 and the execution device 110 may also not be on the same device, that is, the collection of audio data and image data, and the extraction of the voice information to be recognized and the lip movement information of N users may be performed by the client device 140 (such as a smart phone, a smart robot, etc.), while the process of performing feature matching between the voice information to be recognized and the lip movement information of N users may be performed by the execution device 110 (such as a cloud server, a server, etc.). Or, optionally, the collection of audio data and image data is performed by the client device 140, while the extraction of the voice information to be recognized and the lip movement information of N users, and the process of performing feature matching between the voice information to be recognized and the lip movement information of N users are all completed by the execution device 110.
[0138] In the appendix Figure 3In the case shown, the user can manually input data, and this manual input can be operated through the interface provided by the I / O interface 112. In another case, the client device 140 can automatically send input data to the I / O interface 112. If the client device 140 is required to automatically send input data and user authorization is needed, the user can set the corresponding permissions in the client device 140. The user can view the results output by the execution device 110 on the client device 140, and the specific presentation forms can be display, sound, actions, etc. The client device 140 can also be used as a data collection end (such as a microphone or a camera) to collect the input data input to the I / O interface 112 and the output results of the output I / O interface 112 as new sample data, and store them in the database 130. Of course, it is also possible not to collect data through the client device 140, but directly store the input data input to the I / O interface 112 and the output results of the output I / O interface 112 as new sample data in the database 130.
[0139] The preprocessing module 113 is used to preprocess the input data received by the I / O interface 112 (such as the voice data). In the embodiment of the present invention, the preprocessing module 113 can be used to preprocess the voice data, for example, extract the voice information to be recognized from the voice data.
[0140] The preprocessing module 114 is used to preprocess the input data received by the I / O interface 112, such as (the image data). In the embodiment of the present invention, the preprocessing module 114 can be used to preprocess the image data, for example, extract the lip movement information of N users corresponding to the above-mentioned voice information to be recognized from the image data.
[0141] When the execution device 110 preprocesses the input data or the calculation module 111 of the execution device 110 performs calculations and other related processing, the execution device 110 can call the data, code, etc. in the data storage system 150 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing in the data storage system 150. Finally, the I / O interface 112 returns the output results, such as the matching degrees between the lip movement information of N users in the embodiment of the present invention and the voice information to be recognized, or the target user ID of the highest matching degree among them, to the client device 140. The client device 140 then determines the user information of the target user based on the above-mentioned matching degrees, and generates a control instruction matching the user information based on the user information.
[0142] It should be noted that the training device 120 can generate corresponding target feature matching models / rules 101 based on different targets or different tasks and different training data. The corresponding target feature matching models / rules 101 can be used to achieve the above targets or complete the above tasks, so as to provide the required results for users.
[0143] It should be noted that the attached Figure 4 is only a schematic diagram of a system architecture provided by an embodiment of the present invention. The positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in the attached Figure 4 , the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110.
[0144] Based on the above introduction of the system architecture, the neural network model involved in the model training side and the model application side in the embodiments of the present invention is the convolutional neural network. The convolutional neural network CNN is a deep neural network with a convolutional structure and is a deep learning architecture. The deep learning architecture refers to performing multiple levels of learning at different abstraction levels through machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network, and each neuron in the feed-forward artificial neural network responds to the overlapping regions in the input image.
[0145] As Figure 4 shown, Figure 4 is a schematic diagram of a convolutional neural network provided by an embodiment of the present invention. The convolutional neural network (CNN) 200 can include an input layer 210, a convolutional layer / pooling layer 220, and a neural network layer 230, where the pooling layer is optional.
[0146] Convolutional layer / pooling layer 220:
[0147] As Figure 4 shown, the convolutional layer / pooling layer 120 can include layers such as example 221 - 226. In one implementation, layer 221 is a convolutional layer, layer 222 is a pooling layer, layer 223 is a convolutional layer, layer 224 is a pooling layer, 225 is a convolutional layer, and 226 is a pooling layer; in another implementation, 221 and 222 are convolutional layers, 223 is a pooling layer, 224 and 225 are convolutional layers, and 226 is a pooling layer. That is, the output of the convolutional layer can be used as the input of the subsequent pooling layer or as the input of another convolutional layer to continue the convolutional operation.
[0148] Convolutional layer:
[0149] Taking the convolutional layer 221 as an example, the convolutional layer 221 can include many convolutional operators, which are also called kernels. Their role in image processing is equivalent to a filter that extracts specific information from the input image matrix. Essentially, a convolutional operator can be a weight matrix, which is usually predefined. During the convolution operation on the image, the weight matrix usually processes the input image pixel by pixel (or two pixels by two pixels... depending on the value of the stride) along the horizontal direction, thus completing the work of extracting specific features from the image. The size of the weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as that of the input image. During the convolution operation, the weight matrix extends to the entire depth of the input image. Therefore, convolving with a single weight matrix will produce a convolved output with a single depth dimension. However, in most cases, instead of using a single weight matrix, multiple weight matrices with the same dimension are applied. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image. Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract image edge information, another weight matrix is used to extract specific colors in the image, and yet another weight matrix is used to blur the unwanted noise in the image... These multiple weight matrices have the same dimension, and the feature maps extracted by these multiple weight matrices with the same dimension also have the same dimension. Then, the multiple feature maps with the same dimension that are extracted are merged to form the output of the convolution operation.
[0150] The weight values in these weight matrices need to be obtained through a large amount of training in practical applications. Each weight matrix formed by the weight values obtained through training can extract information from the input image, thus helping the convolutional neural network 200 to make correct predictions.
[0151] When the convolutional neural network 200 has multiple convolutional layers, the initial convolutional layer (such as 221) often extracts more general features, which can also be called low-level features; as the depth of the convolutional neural network 200 increases, the later convolutional layers (such as 226) extract more and more complex features, such as high-level semantic features. The higher the semantic features, the more suitable they are for the problem to be solved.
[0152] Pooling layer:
[0153] Since it is often necessary to reduce the number of training parameters, a pooling layer is often introduced periodically after the convolutional layer, that is, as Figure 4Each of the layers 221 - 226 shown in 220 can be a convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers. During image processing, the sole purpose of the pooling layer is to reduce the spatial size of the image. The pooling layer can include an average pooling operator and / or a max pooling operator for sampling the input image to obtain a smaller-sized image. The average pooling operator can calculate the average value of pixel values in the image within a specific range. The max pooling operator can take the pixel with the maximum value within a specific range as the result of max pooling. Additionally, just as the size of the weight matrix in the convolutional layer should be related to the image size, the operators in the pooling layer should also be related to the image size. The size of the image output after processing by the pooling layer can be smaller than the size of the image input to the pooling layer. Each pixel point in the image output by the pooling layer represents the average value or the maximum value of the corresponding sub-region of the image input to the pooling layer.
[0154] Neural network layer 130:
[0155] After being processed by the convolutional / pooling layer 220, the convolutional neural network 100 is still not sufficient to output the required output information. As mentioned before, the convolutional / pooling layer 220 only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other relevant information), the convolutional neural network 200 needs to use the neural network layer 230 to generate an output of one or a set of required classes. Therefore, the neural network layer 230 can include multiple hidden layers (such as Figure 4 231, 232 to 23n shown) and an output layer 240. The parameters contained in the multiple hidden layers can be pre-trained according to the relevant training data of the specific task type. For example, the task type can include image recognition, image classification, image super-resolution reconstruction, and so on...
[0156] After the multiple hidden layers in the neural network layer 230, that is, the last layer of the entire convolutional neural network 200 is the output layer 240. The output layer 240 has a loss function similar to categorical cross-entropy, specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network 200 ( Figure 4 The propagation from 210 to 240 in is the forward propagation) is completed, the backpropagation ( Figure 4 The propagation from 240 to 210 in is the backpropagation) will start to update the weight values and biases of the aforementioned layers to reduce the loss of the convolutional neural network 200 and the error between the result output by the convolutional neural network 200 through the output layer and the ideal result.
[0157] It should be noted that as Figure 4The convolutional neural network 200 shown is only an example of a convolutional neural network. In specific applications, the convolutional neural network can also exist in the form of other network models. For example, multiple convolutional layers / pooling layers are parallel, and the features extracted respectively are all input to the fully neural network layer 230 for processing.
[0158] The normalization layer in this application, as a functional layer of the CNN, can in principle be performed after any layer or before any layer in the above CNN, and uses the feature matrix output by the previous layer as the input, and its output can also be used as the input of any functional layer in the CNN. However, in actual CNN applications, the normalization layer is generally performed after the convolutional layer and uses the feature matrix output by the previous convolutional layer as the input matrix.
[0159] Based on the above Figure 3 and Figure 4 Regarding the description of the system architecture 100 and the related functions of the convolutional neural network 200, the following describes the embodiments of the training method and speech matching method of the neural network provided in this application from the model training side and the model application side in combination with the above application scenarios, system architecture, structure of the convolutional neural network, and structure of the neural network processor, and specifically analyzes and solves the technical problems proposed in this application.
[0160] See s , Figure 5 is a schematic flowchart of a training method of a neural network provided by an embodiment of the present invention. This method can be applied to the application scenarios and system architectures described in the above Figure 5 , Figure 1 and can be specifically applied to the training device 120 in the above Figure 2 . The following combines the attached Figure 3 Taking the execution subject as the training device 120 in the above Figure 5 or a device including the training device 120 as an example for description. This method may include the following steps S701 - step S702.
[0161] S701: Obtain training samples, where the training samples include the lip movement information of the training user and M instruction information.
[0162] Specifically, for example, the lip movement information of the training user is the lip movement information corresponding to the voice message sent by user Xiaofang: "Hello, my name is Xiaofang. I'm from Hunan, China. What about you?" That is, the lip movement video or the image sequence of continuous lip movement, or the vector parameters formed by the distances between the upper and lower lips that can reflect the opening and closing movement of the lips in chronological order. Then, the M instruction messages include the waveform sequence or text information of the above "turn up the air conditioner temperature" instruction message as an instruction sample, and other instruction messages, such as voice messages like "lower the seat back angle a bit", "open the window", "turn off the music", etc. as negative samples. Optionally, the M instruction messages include the instruction messages that match the lip movement information of the training user and (M - 1) instruction messages that do not match the lip movement information of the training user. For example, the above lip movement information is the image sequence of continuous lip movement (i.e., the video of the pronunciation mouth shape) corresponding to user A when sending the instruction message "turn up the air conditioner temperature", and the above M instruction messages include the voice waveform sequence of the above voice positive sample and the voice waveform sequences of M - 1 negative samples. It can be understood that the above M instruction messages can also include multiple positive samples and negative samples, that is, the specific quantities of positive samples and negative samples are not limited, as long as they are all included.
[0163] S702: Using the lip movement information of the training user and the M voice messages as training inputs, and using the matching degrees between the lip movement information of the training user and the M voice messages as M labels, train the initialized neural network to obtain a target feature matching model.
[0164] Specifically, for example, the label between the lip movement information of the above training user and the positive sample instruction message "turn up the air conditioner temperature" is "matching degree = 1", while the labels between the lip movement information of the above training user and other negative sample instruction messages such as "lower the seat back angle a bit", "open the window", "turn off the music" are "matching degree = 0.2", "matching degree = 0", "matching degree = 0", etc., which will not be elaborated here. That is, through the above training inputs and the pre - set labels, the initialized neural network model can be trained to obtain the target feature matching model required in this application. This target feature matching model can be used to match the matching relationship between the instruction information to be recognized and the lip movement information of multiple users, and is used to implement any one of the voice matching methods in this application.
[0165] In a possible implementation, the lip movement information of the training user and the M instruction information are used as training inputs, and the matching degrees between the lip movement information of the training user and the M instruction information are used as M labels to train the initialized neural network to obtain a target feature matching model, including: inputting the lip movement information of the training user and the M instruction information into the initialized neural network, and calculating the matching degrees between the M instruction information and the lip movement information of the training user respectively; comparing the calculated matching degrees between the M instruction information and the lip movement information of the training user with the M labels, and training the initialized neural network to obtain a target feature matching model.
[0166] In a possible implementation, the target feature matching model includes a first model, a second model and a third model; the step of inputting the lip movement information of the training user and the M instruction information into the initialized neural network and calculating the matching degrees between the M instruction information and the lip movement information of the training user respectively includes: inputting the M instruction information into the first model to obtain M voice features, each of the M voice features is a K-dimensional voice feature, and K is an integer greater than 0; inputting the lip movement information of the training user into the second model to obtain the image sequence feature of the training user, and the image sequence feature of the training user is a K-dimensional image sequence feature; inputting the M voice features and the image sequence feature of the training user into the third model, and calculating the matching degrees between the M voice features and the image sequence feature of the training user respectively.
[0167] Regarding how to specifically train the initialized neural network model into the target feature matching model in the present application, it will be described together in the method embodiment on the model application side corresponding to FIG. 7 later, and will not be elaborated here.
[0168] In the embodiment of the present invention, by using the lip movement information of a certain training user, as well as the matching instruction information and multiple non-matching instruction information as the input of the initialized neural network, and using the actual matching degrees between the above M instruction information and the lip movement information of the training user as labels to train the above initialized neural network model to obtain a target feature matching model. For example, the matching degree corresponding to a complete match is the label 1, and the matching degree corresponding to a non-match is the label 0. When the matching degrees between the lip movement information of the training user calculated by the trained initialized neural network and the M instruction information are closer to the M labels, the trained initialized neural network is closer to the target feature matching model.
[0169] See Figure 3 ,Figure 9 FIG. Figure 9 is a schematic flowchart of another voice command control method provided by an embodiment of the present invention, which is mainly applicable to the scenario of voice interaction control of in-vehicle devices by in-vehicle members. Usually, in the scenario where there are multiple in-vehicle members in the vehicle, when the in-vehicle device receives a voice command for controlling the in-vehicle device, and when the in-vehicle device needs to determine which in-vehicle member at which position issues the command and perform response control for a specific position area, this solution can be used to accurately identify which member at which position issues the voice command. This method can be applied to the application scenarios and system architectures in the vehicle cabin, and specifically can be applied to the above-mentioned Figure 9 customer device 140 and execution device 110. It can be understood that both the customer device 140 and the execution device 110 can be set in the vehicle. The following combines the attached Figure 3 Taking an intelligent vehicle as the execution subject as an example for description. This method may include the following steps S1601 - step S1605
[0170] Step S1601: Obtain in-vehicle audio data.
[0171] Specifically, obtain the in-vehicle audio data collected by the in-vehicle microphone. The audio data includes the ambient sound in the vehicle, such as the music of the speaker, the noise of the engine and air conditioner, the sound outside the vehicle, and other ambient sounds, as well as the voice commands issued by the user.
[0172] Usually, there is a microphone array in the vehicle cabin of an intelligent vehicle, that is, there are multiple microphones distributed at different positions in the vehicle cabin. Therefore, when there is a microphone array in the vehicle, at this time, step S1601 can specifically be:
[0173] S1601 a: Obtain the audio data collected by multiple in-vehicle microphones.
[0174] In the scenario of human-machine interaction, the microphone array will collect audio data, or the in-vehicle microphone array is in a real-time audio data collection state after the vehicle is started, or after a specific operation is performed by an in-vehicle member, such as the owner, for example, after the audio collection function is turned on, the microphone array enters the audio collection state. The way for the microphone array to collect audio data is that multiple microphones collect audio data at different positions in the vehicle cabin respectively.
[0175] S1601 b: Obtain target audio data based on the audio data collected by multiple microphones.
[0176] An in-vehicle microphone array usually has multiple microphones set at different positions inside the vehicle. Therefore, when acquiring audio data in the in-vehicle environment, there are multiple audio sources to choose from. Since the effects of audio data collected at different positions are different. For example, when the person issuing the voice command is sitting in the back row of the vehicle, and the passenger in the front passenger seat is listening to music, and the speaker in the passenger seat is playing the song at this time, then the audio data collected at the passenger seat will have a relatively loud music sound due to the speaker in the passenger seat, and the command information of the rear row passenger is relatively small. While the speaker in the rear row will relatively collect a relatively clear voice signal only accompanied by a small music sound. At this time, when acquiring audio data, usually after preprocessing the audio data collected by each microphone, through analysis and comparison, the target audio data is selected. For example, because the environmental noise, music sound, and the frequency band where the voice command is located are different, the preprocessing can be to perform filtering processing on the audio data collected by multiple microphones, and select the audio signal with the strongest voice signal after filtering processing as the target audio signal.
[0177] Here, it can also be to use other existing preprocessing methods to determine which microphone collects the audio data with the best signal quality of the voice command-related signal, and select this audio signal as the target audio signal. The selected target audio signal can be the original audio signal collected by the microphone, or the audio signal after preprocessing.
[0178] Step S1602: When it is recognized that the audio data includes the first type of command information, acquire image data.
[0179] There are various ways to recognize whether the audio data obtained in step S1601 includes the first type of command information. For example, it can be based on an RNN model to perform semantic recognition of audio information, and then based on the recognized text information, perform command content recognition, and judge the command type according to the command content, or directly judge the command type according to the feature information in the text information, such as keywords. There are already various specific solutions for command recognition based on voice in the prior art, and they will not be listed one by one here. The audio data used for model input can be the audio data after preprocessing such as environmental noise filtering of the collected audio data, or can be directly input based on the collected audio data. It can also be other voice recognition methods in the prior art to judge whether there is command information included.
[0180] The first type of command information in this embodiment refers to the command information that the in-vehicle device can receive and recognize, and needs to perform corresponding operation responses on the position area by judging the position where the command initiator is located, that is, usually the regulation commands for the in-vehicle facilities, such as the air-conditioning adjustment command in the vehicle cabin, sound adjustment, and commands related to audio content selection adjustment.
[0181] The instruction information can be a sequence of speech waveforms corresponding to the instruction time period in the audio data, or a sequence of text features of the text information extracted from the audio data during the target time period. When introducing the model in this article, the speech information to be recognized will be mentioned. In essence, it is also a sequence of speech waveforms during the corresponding time period when the speech instruction is issued. Therefore Figure 9 when the instruction information mentioned in [reference] is in the form of a sequence of speech waveforms, it is also a kind of speech information to be recognized.
[0182] Obtaining in-vehicle audio image data is the same as performing audio acquisition with a microphone. It can start automatically for real-time acquisition after the vehicle starts, or the real-time acquisition function can be enabled according to the user's instructions, or the acquisition of image data can be enabled simultaneously when the default audio acquisition starts. Usually, multiple cameras are installed on the vehicle, and different types of cameras are also set, such as monocular cameras, binocular cameras, TOF cameras, infrared cameras, etc. In this solution, the deployment location, number, and type of cameras for acquiring in-vehicle image data are not limited. Those skilled in the art can make corresponding selection and deployment according to the needs of specific solution implementation. The microphone in step S1601 can be an independently set microphone or a microphone integrated in the camera.
[0183] The image data can be that the in-vehicle processing system acquires image data through a camera while acquiring voice data through a microphone. That is, the above-mentioned audio data and image data are the original audio data and image data within a certain time period, that is, the audio data source and the image data source. Optionally, the audio data and the image data are acquired within the same time period for the same scene.
[0184] Since there are usually more than two rows of seats in the vehicle cabin, when acquiring image data from a single camera, there is often an occlusion situation between members. Therefore, in order to clearly acquire the lip movement information of each member, it is often necessary to acquire image data through multiple cameras at different positions in the vehicle cabin. However, the number of audio data sources and the number of image data sources do not necessarily need to match. For example, audio data can be acquired through microphones installed at various positions in the vehicle, and image data of the vehicle cabin can be acquired through a global camera in the vehicle, or audio data can be acquired through a specified microphone, and image data of the vehicle cabin can be acquired through cameras at multiple positions in the vehicle.
[0185] Step S1603: Extract the lip movement information of the members at N positions in the vehicle from the image data.
[0186] Specifically, based on the collected in-vehicle video information, the position distribution of in-vehicle members is determined, and the lip movement information of the members in each position is extracted, and the lip movement information carries a corresponding position identifier.
[0187] The lip movement information of each of the multiple in-vehicle members includes an image sequence of the corresponding user's lip movement during the corresponding target time period, and the target time period is the time period corresponding to the instruction information in the audio. That is, the lip videos of each member extracted from the original image data, that is, the image sequence of continuous lip movements, contain the continuous lip shape change characteristics of the corresponding member. For example, the format of each frame of the image data collected by the camera is a 24-bit BMP bitmap. Among them, the BMP image file (Bitmap-File) format is the image file storage format adopted by Windows, and the 24-bit image uses 3 bytes to save the color value. Each byte represents a color, arranged in red (R), green (G), and blue (B), and the RGB color image is converted into a grayscale image. The intelligent device obtains at least one face region from the image data collected by the above camera based on the face recognition algorithm, and further, taking each face region as a unit, assigns a face ID to each face region (different from the robot or smart speaker scenario, the face ID is used to correspond to the position in the vehicle), and extracts the video sequence stream of the mouth region, where the frame rate of the video is 30f / s (frame rate (Framerate) = number of frames (Frames) / time (Time), unit: frames per second (f / s, frames per second, fps)). 9 consecutive image frames form a 0.3-second video stream. The 9-frame image data (video speed 30fps) is concatenated (concat) into a cube with a size of 9×60×100, where 9 represents the number of frames of time information (temporal feature). Each channel is a 60×100 grayscale image (2d spatial feature) of the oral cavity region. Taking these N image sequences of lip movements within the corresponding 0.3s as the input of video features, where 0.3s is the target time period.
[0188] For the specific method of extracting the lip movement information of multiple members from the image data, refer to the corresponding technical solution description in the prior embodiment of the present invention.
[0189] Step S1604: Input the instruction information and the lip movement information of the N members into the target feature matching model to obtain the matching degrees between the lip movement information of the N members and the instruction information respectively.
[0190] Specifically, the instruction information and the image sequences of the lip movements of each of the N members during the target time period are respectively used as the input of the audio feature and the input of the video feature, and are input into the target feature matching model, and the matching degrees between the instruction information and the lip movement features of the N members are respectively calculated. The matching degree can specifically be a value greater than or equal to 0 and less than or equal to 1.
[0191] The instruction information here can be, for example, Figure 9 as shown Figure 6 a sound waveform example diagram provided by an embodiment of the present invention, or in the form of identification information of an instruction, such as in the form of a sequence number, or in the form of an instruction statement, etc. In a possible implementation manner, assuming that there are multiple users speaking simultaneously in the audio data, then at this time, it is necessary to determine which user emits a certain segment of speech information among them, and it is necessary to first identify and extract the target speech information in the audio data, that is, the above-mentioned speech information to be recognized. Or, assuming that the audio data includes multiple segments of speech information spoken by a certain user, and the intelligent device only needs to recognize a certain segment of speech information among them, then this segment of speech information is the speech information to be recognized. For example, the intelligent device extracts audio features from the audio data obtained by the microphone array in S801. The specific method can use Mel Frequency Cepstral Coefficients for speech feature extraction. Mel Frequency Cepstral Coefficients (MFCC) are used to extract 40-dimensional features from data with a frame length of 20 ms. There is no overlap between frames (non-overlapping), and every 15 frames (corresponding to an audio segment of 0.3 seconds) are concatenated (concat) into a cube with a dimension of 15×40×3 (where 15 is the temporal feature, and 40×3 is the 2D spatial feature). The speech waveform sequence within this 0.3 s is used as the input of the audio feature, and 0.3 s is the target time period. In addition to the above methods, there are also other methods in the prior art that can be used to separate the target statement in a segment of speech.
[0192] Regarding the specific implementation of the target feature matching model, it will be specifically introduced below. For the model structure, refer to the subsequent description of FIG. 7 and the Figure 6 description of model training and acquisition in the previous text.
[0193] Step S1605: Determine which position area to execute the instruction corresponding to the instruction information according to the matching degree.
[0194] Since the matching degree is usually in the form of a numerical value, the corresponding determination strategy for S1605 can be to determine the position in the vehicle where the member corresponding to the lip movement information of the member with the highest matching degree is located as the target area for executing the instruction information, and execute the
[0195] For example, when the instruction is to turn down the air conditioner, only the operation of turning down the temperature or the air volume of the air outlet is performed on the target area.
[0196] In addition, S1604 - S1605 can also be:
[0197] S1604: Input the instruction information and the lip movement information of one of the members into the target feature matching model to obtain the matching degree between the lip movement information of the member and the instruction information.
[0198] Step S1605: When the matching degree is greater than the known threshold value, execute the instruction corresponding to the instruction information in the position area where the member is located.
[0199] If the matching degree is less than the known threshold value, continue to judge the matching degree between the lip movement information of the member at another position in the vehicle and the instruction information according to a certain rule until the lip movement information with a matching degree greater than the known threshold value is obtained, or all the members in the vehicle are matched, and then the matching process ends.
[0200] In addition to identifying the position of the instruction initiator in the vehicle cabin as in the above embodiments and performing corresponding operations for specific positions, there are also scenarios where it is necessary to judge the identity of the instruction initiator. For example, when a voice instruction related to vehicle control is recognized, it is necessary to judge whether it is an instruction issued by the driver, so as to judge whether the instruction can be executed. For such scenarios, the specific implementation method is as follows:
[0201] See Figure 3 , Figure 10 is a schematic flowchart of another voice matching method provided by an embodiment of the present invention, which is mainly applicable to the scenario of performing operation control on the vehicle based on voice instructions for driving. Since there are usually multiple members in the vehicle, it is generally considered that only the driver has the permission to perform voice operation control on the vehicle driving. To avoid misoperation and misrecognition, when the in-vehicle device receives a voice instruction for vehicle driving control, it is necessary to judge whether it is a voice instruction issued by the driver, and then judge whether to execute the vehicle driving instruction based on the recognition result. This method may include the following steps S1701 - step S1705.
[0202] Step S1701: Obtain in-vehicle audio data.
[0203] The specific implementation of S1701 is the same as that of S1601.
[0204] Step S1702: When it is recognized that the audio data includes the second type of instruction information, obtain image data.
[0205] The second type of instruction information in S1702 mainly refers to instruction information related to vehicle form control, such as vehicle turning, accelerating, starting, switching of driving modes, etc. When this type of instruction information is recognized, it is necessary to obtain the image data of the member in the driver's seat.
[0206] For the specific instruction recognition method and image data acquisition method, refer to S1602.
[0207] Step S1703: Extract the lip movement information of the first position member from the image data. For how to extract the lip movement information and how to identify the lip movement information, refer to S1603.
[0208] Step S1704: Input the instruction information and the lip movement information of the first position member into the target feature matching model to obtain the matching degrees between the lip movement information of the driver's seat member and the instruction information respectively.
[0209] Step S1705: Determine whether to execute the instruction corresponding to the instruction information according to the matching degree.
[0210] There are multiple judgment methods for S1705. Since the matching degree is usually in the form of a numerical value, S1705 can judge whether to execute the instruction information according to whether the matching degree is higher than a preset threshold. That is, it can be that when the matching degree is greater than the preset threshold, it is considered that the instruction is issued by the first position member, and then the vehicle form control instruction is executed. Otherwise, the instruction is not executed.
[0211] S1705 can also be to judge whether the matching degree of the lip information of the first position member is the highest among the matching degrees of the lip movement information and instruction information of all in-vehicle members. If this is the case, then in S1703, in addition to extracting the lip movement information of the first position member, it is also necessary to extract the lip movement information of other in-vehicle members. Similarly, in S1704, in addition to inputting the instruction information and the lip movement information of the first position member into the target feature matching model, it is also necessary to input the lip information of other members into the target feature matching model to obtain the corresponding matching degrees.
[0212] When the solution is specifically implemented, the first position in the above embodiments is usually the driver's seat. For example, the initial setting of the in-vehicle control system can default that the driver's seat member has the permission to control the vehicle driving operation by voice, or it can be changed based on the user's manual setting according to the specific position distribution situation during each ride. For example, it is set that both the driver's seat and the co-driver's seat have the vehicle driving control permission. In this case, the first position is the driver's seat and the co-driver's seat.
[0213] Alternatively, when the embodiment of the present invention is specifically implemented, during vehicle initialization, according to vehicle prompts or the owner's initiative, the image information and permission information of the members who will use the vehicle in the family are input on the vehicle. At this time, when the solution of the embodiment of the present invention is specifically implemented, before the vehicle starts or after it starts, the in-vehicle camera obtains the position information of the registered members with driving control authority. Then, when a vehicle control-related instruction is recognized, it is determined whether the voice instruction is issued by the member at the position based on the lip movement information of the member at the position with control authority.
[0214] In addition to determining whether the vehicle driving control type voice instruction is issued by the driver's seat member, the embodiment of the present invention can also be applied to the judgment of whether other types of instructions can be executed. For example, for the call function, it can be manually or by default set in the vehicle that only the owner or driver can execute voice control. The above embodiments are only specific examples and do not limit the specific instruction type or specific fixed position.
[0215] Through the above two in-vehicle interaction embodiments, it is possible to implement instruction operations specifically by determining which member in the vehicle seat issues the instruction, providing more precise in-vehicle interaction control for users.
[0216] For voice control in vehicle driving and operation, it can well prevent misoperations and misidentifications, ensuring that only the driver can perform corresponding vehicle driving control, providing the safety of vehicle driving control.
[0217] The embodiment of the present invention also provides another voice instruction control method, which can be applied to the in-vehicle application scenarios and system architectures described above Figure 10 、 Figure 1 and can be specifically applied to the execution device 110 described above Figure 2 It can be understood that at this time, the client device 140 and the execution device 110 may not be on the same physical device. As shown in Figure 3 shown, Figure 8 FIG. is a system architecture diagram of a voice instruction control system provided by an embodiment of the present invention. In this system, for example, it includes an intelligent vehicle 800, which serves as a collection device for audio data and image data, and further can also serve as an extraction device for the to-be-recognized instruction information and the lip information of N users. The matching between the extracted to-be-recognized instruction information and the lip information of N users can be executed on the server / service device / service apparatus / cloud service device 801 where the execution device 110 is located. Optionally, the extraction of the to-be-recognized instruction information and the lip information of N users can also be executed on the device side where the execution device 110 is located. The embodiment of the present invention does not make specific limitations in this regard. The following takes the inclusion of Figure 8Taking the cloud service device 801 in it as an example for description. The method is as follows Figure 8 shown, and may include the following steps S1001 - step S1003.
[0218] Step S1001: Obtain instruction information and lip movement information of N in - vehicle members located in the vehicle cabin;
[0219] In the above - mentioned steps, the instruction information is obtained according to the audio data collected in the vehicle cabin, and the lip movement information of the in - vehicle members is obtained when it is determined that the instruction corresponding to the instruction information is a first - type instruction. The lip movement information includes an image sequence of the lip movement of the in - vehicle member located at the first position in the vehicle cabin within a target time period, and the target time period is the time period corresponding to the instruction in the audio data.
[0220] Step S1002: Input the instruction information and the lip movement information of N in - vehicle members located in the vehicle cabin into a target feature matching model to obtain the matching degrees between the lip movement information of the in - vehicle members at the N positions and the instruction information respectively;
[0221] Step S1003: Take the position where the member corresponding to the lip movement information of the user with the highest matching degree is located as the target position for executing the instruction corresponding to the instruction information.
[0222] In addition, there is also a cloud solution as Figure 11 shown, which needs to identify the permissions of the member who specifically issues the instruction, so as to judge the target execution area of the instruction:
[0223] Step S1021: Obtain instruction information and lip movement information of the in - vehicle member located at the first position in the vehicle;
[0224] In the above - mentioned steps, the instruction information is obtained according to the audio data collected in the vehicle cabin, and the lip movement information of the in - vehicle member located at the first position in the vehicle is obtained when it is identified that the instruction corresponding to the instruction information is a second - type instruction. The lip movement information includes an image sequence of the lip movement of the in - vehicle member located at the first position in the vehicle cabin within a target time period, and the target time period is the time period corresponding to the instruction in the audio data.
[0225] Step S1022: Input the instruction information and the lip movement information of the in - vehicle member located at the first position in the vehicle into a target feature matching model to obtain the first matching degree between the lip movement information of the in - vehicle member located at the first position in the vehicle and the instruction information;
[0226] Step S1023: Determine whether to execute the instruction corresponding to the instruction information according to the first matching degree.
[0227] In a possible implementation, the target feature matching model includes a first model, a second model, and a third model;
[0228] Inputting the speech information to be recognized and the lip movement information of the N users into the target feature matching model to obtain the matching degrees between the lip movement information of the N users and the speech information to be recognized respectively, includes:
[0229] Inputting the speech information to be recognized into the first model to obtain speech features, where the speech features are K-dimensional speech features, and K is an integer greater than 0;
[0230] Inputting the lip movement information of the N users into the second model to obtain N image sequence features, and each of the N image sequence features is a K-dimensional image sequence feature;
[0231] Inputting the speech features and the N image sequence features into the third model to obtain the matching degrees between the N image sequence features and the speech features respectively.
[0232] In a possible implementation, the target feature matching model is a feature matching model trained with the lip movement information of training users and M instruction information as inputs and with the matching degrees between the lip movement information of the training users and the M speech information as M labels.
[0233] In a possible implementation, the method further includes:
[0234] Determining the user information of the target user, where the user information includes one or more of personal attribute information, facial expression information corresponding to the speech information to be recognized, and environmental information corresponding to the speech information to be recognized;
[0235] Generating a control instruction matching the user information based on the user information.
[0236] In a possible implementation, the method further includes: extracting the lip movement information of N users from image data; further, the extracting the lip movement information of N users from the image data includes:
[0237] Based on a face recognition algorithm, recognizing N face regions in the image data and extracting the lip movement videos in each of the N face regions;
[0238] Determining the lip movement information of the N users based on the lip movement videos in each face region.
[0239] In a possible implementation, the method further includes: extracting voice information to be recognized from the audio data; further, the extracting voice information to be recognized from the audio data includes:
[0240] Based on a spectrum recognition algorithm, recognizing audio data of different spectra in the audio data, and recognizing the audio data of the target spectrum as the voice information to be recognized.
[0241] It should be noted that for the method flow executed by the cloud service device described in the embodiments of the present invention, reference can be made to the relevant method embodiments described above Figure 12 and will not be elaborated here.
[0242] Please refer to Figures 9 - 12 , Figure 13 which is a schematic structural diagram of an intelligent device provided by an embodiment of the present invention, Figure 13 and is a schematic functional principle diagram of an intelligent device provided by an embodiment of the present invention. The intelligent device can be an in-vehicle device, an in-vehicle system, or an intelligent vehicle. The intelligent device 40 may include a processor 401, and a microphone 402 and a camera 403 coupled to the processor 401. When it is an intelligent vehicle or an in-vehicle voice processing system, the microphone 402 and the camera 403 are usually multiple, such as corresponding to Figure 13 the application scenario, where
[0243] the microphone 402 is used to collect audio data;
[0244] the camera 403 is used to collect image data, and the audio data and the image data are collected for the same scenario;
[0245] The processor 401 obtains the in-vehicle audio data. When it recognizes that the in-vehicle audio data includes a first type of instruction, it obtains the in-vehicle image data; extracts the lip movement information of the vehicle occupants at N positions in the vehicle compartment from the in-vehicle image data; and is used to input the instruction information corresponding to the first type of instruction and the lip movement information of the vehicle occupants at N positions in the vehicle compartment into a target feature matching model to obtain the matching degrees between the lip movement information of the vehicle occupants at the N positions and the instruction information respectively; and takes the position where the member corresponding to the lip movement information of the user with the highest matching degree is located as the target position for executing the instruction corresponding to the instruction information.
[0246] Such as corresponding to Figure 12 the application scenario, the microphone 402 is used to collect audio data;
[0247] the camera 403 is used to collect image data, and the audio data and the image data are collected for the same scenario;
[0248] The processor 401 obtains the in-vehicle audio data. When it recognizes that the in-vehicle audio data includes a second type of instruction, it obtains the in-vehicle image data, and obtains first image data from the in-vehicle image data. The first image data is image data including an in-vehicle member located at a first position inside the vehicle, and extracts the lip movement information of the in-vehicle member located at the first position inside the vehicle from the first image data; and is used to input the instruction information corresponding to the second type of instruction and the lip movement information of the in-vehicle member located at the first position inside the vehicle into a target feature matching model, to obtain a first matching degree between the lip movement information of the in-vehicle member located at the first position inside the vehicle and the instruction information, and determines whether to execute the instruction corresponding to the instruction information according to the first matching degree.
[0249] In a possible implementation manner, the voice information to be recognized includes a voice waveform sequence within a target time period; and the lip movement information of each of the N users among the lip movement information of the N users includes an image sequence of the corresponding user's lip movement within the target time period.
[0250] In a possible implementation manner, the processor 401 is specifically configured to: input the voice information to be recognized and the lip movement information of the N users into a target feature matching model, to obtain matching degrees between the lip movement information of the N users and the voice information to be recognized respectively; and determine the user corresponding to the lip movement information of the user with the highest matching degree as the target user to which the voice information to be recognized belongs.
[0251] In a possible implementation manner, the target feature matching model includes a first model, a second model, and a third model; the processor 401 is specifically configured to: input the voice information to be recognized into the first model to obtain voice features, where the voice features are K-dimensional voice features, and K is an integer greater than 0; input the lip movement information of the N users into the second model to obtain N image sequence features, and each of the N image sequence features is a K-dimensional image sequence feature; and input the voice features and the N image sequence features into the third model to obtain matching degrees between the N image sequence features and the voice features respectively.
[0252] In a possible implementation manner, the target feature matching model is a feature matching model trained with the lip movement information of training users and M pieces of voice information as inputs, and with the matching degrees between the lip movement information of the training users and the M pieces of voice information as M labels, where the M pieces of voice information include the voice information that matches the lip movement information of the training users.
[0253] In a possible implementation, the processor 401 is further configured to: determine user information of the target user, where the user information includes one or more of person attribute information, facial expression information corresponding to the speech information to be recognized, and environmental information corresponding to the speech information to be recognized; and generate a control instruction matching the user information based on the user information.
[0254] In a possible implementation, the processor 401 is specifically configured to: based on a face recognition algorithm, recognize N face regions in the image data, and extract lip movement videos in each of the N face regions; and determine lip movement information of the N users based on the lip movement videos in each of the face regions.
[0255] In a possible implementation, the processor 401 is specifically configured to: based on a spectrum recognition algorithm, recognize audio data of different spectra in the audio data, and recognize the audio data of the target spectrum as the speech information to be recognized.
[0256] It should be noted that for the functions of the relevant modules in the intelligent device 40 described in the embodiments of the present invention, reference may be made to the relevant method embodiments described above Figure 12 and will not be elaborated herein.
[0257] Please refer to Figures 9 - 12 , Figure 14 which is a schematic structural diagram of a training device for a neural network provided by an embodiment of the present invention, Figure 14 and is a schematic functional principle diagram of an intelligent device provided by an embodiment of the present invention. The model trained by the training device for the neural network can be used in in-vehicle devices, vehicle-mounted systems, intelligent vehicles, cloud servers, etc. The training device 60 for the neural network may include an acquisition unit 601 and a training unit 602; where,
[0258] The acquisition unit 601 is configured to acquire training samples, where the training samples include lip movement information of a training user and M instruction information; optionally, the M instruction information includes instruction information matching the lip movement information of the training user and (M - 1) instruction information not matching the lip movement information of the training user;
[0259] The training unit 602 is configured to use the lip movement information of the training user and the M instruction information as training inputs, and use the matching degrees between the lip movement information of the training user and the M instruction information as M labels to train an initialized neural network to obtain a target feature matching model.
[0260] In a possible implementation, the lip movement information of the training user includes a sequence of lip movement images of the training user, and the M instruction messages include a sequence of speech waveforms matching the sequence of lip movement images of the training user and (M - 1) sequences of speech waveforms not matching the sequence of lip movement images of the training user.
[0261] In a possible implementation, the training unit 602 is specifically configured to:
[0262] Input the lip movement information of the training user and the M instruction messages into the initialized neural network, and calculate the matching degrees between the M instruction messages and the lip movement information of the training user respectively;
[0263] Compare the calculated matching degrees between the M instruction messages and the lip movement information of the training user respectively with the M labels, and train the initialized neural network to obtain a target feature matching model.
[0264] In a possible implementation, the target feature matching model includes a first model, a second model, and a third model; the training unit 602 is specifically configured to:
[0265] Input the M instruction messages into the first model to obtain M speech features, and each of the M speech features is a K-dimensional speech feature, where K is an integer greater than 0;
[0266] Input the lip movement information of the training user into the second model to obtain the image sequence features of the training user, and the image sequence features of the training user are K-dimensional image sequence features;
[0267] Input the M speech features and the image sequence features of the training user into the third model, and calculate the matching degrees between the M speech features and the image sequence features of the training user respectively;
[0268] Compare the calculated matching degrees between the M speech features and the image sequence features of the training user respectively with the M labels, and train the initialized neural network to obtain a target feature matching model.
[0269] Please refer to Figure 14 , Figure 15 which is the system structure diagram provided by an embodiment of the present invention, including a schematic structural diagram of an intelligent device 70 and a server device 80. The intelligent device may be an intelligent vehicle. The intelligent device 70 may include a processor 701, and a microphone 702 and a camera 703 coupled to the processor 701; where
[0270] A microphone 702 for collecting audio data;
[0271] A camera 703 for collecting image data;
[0272] A processor 701 for obtaining audio data and image data;
[0273] Extract the speech information to be recognized from the audio data, where the speech information to be recognized includes a speech waveform sequence within a target time period;
[0274] Extract the lip movement information of N users from the image data, where the lip movement information of each of the N users includes an image sequence of the corresponding user's lip movement within the target time period, and N is an integer greater than 1;
[0275] When applied to an intelligent vehicle or an in-vehicle voice interaction system, the processor 701 is used to obtain audio data. When the audio data includes a target instruction, obtain in-vehicle image data; extract the lip movement information of in-vehicle members at N positions in the vehicle cabin from the in-vehicle image data. Here, the lip movement information of the in-vehicle members can be sent to the service device, or the collected in-vehicle image information can be sent to the service device, and the service device extracts the lip movement information.
[0276] Or the processor 701 is used to obtain in-vehicle audio data; when it is recognized that the audio data includes a second type of instruction, obtain first image data, where the first image data is image data including an in-vehicle member located at a first position in the vehicle; extract the lip movement information of the in-vehicle member located at the first position in the vehicle from the first image data.
[0277] It should be noted that for the functions of the relevant modules in the intelligent device 70 described in the embodiments of the present invention, reference can be made to the relevant method embodiments described above Figure 15 and will not be elaborated here.
[0278] Figures 9 - 12 Figure 15 Provide a schematic structural diagram of a service device, and the service device can be a server, a cloud server, etc. The service device 80 may include a processor; optionally, the processor may be composed of a neural network processor and a processor 802 coupled to the neural network processor, or directly composed of a processor; where
[0279] Corresponding to the in-vehicle implementation scenario, the neural network processor 801 is used for:
[0280] Input the instruction information corresponding to the first type of instruction and the lip movement information of the in-vehicle members at N positions in the vehicle cabin into the target feature matching model to obtain the matching degrees between the lip movement information of the in-vehicle members at the N positions and the instruction information respectively; take the position where the member corresponding to the lip movement information of the user with the highest matching degree is located as the target position for executing the instruction corresponding to the instruction information.
[0281] Or it is used for: inputting the instruction information corresponding to the second type of instruction and the lip movement information of the in-vehicle member located at the first position in the vehicle into the target feature matching model to obtain the first matching degree between the lip movement information of the in-vehicle member located at the first position and the instruction information; determining whether to execute the instruction corresponding to the instruction information according to the first matching degree.
[0282] In a possible implementation manner, the target feature matching model includes a first model, a second model, and a third model; the processor 802 is specifically configured to: input the voice information or instruction information to be recognized into the first model to obtain voice features, where the voice features are K-dimensional voice features, and K is an integer greater than 0; input the lip movement information of the N users into the second model to obtain N image sequence features, and each of the N image sequence features is a K-dimensional image sequence feature; input the voice features and the N image sequence features into the third model to obtain the matching degrees between the N image sequence features and the voice features respectively.
[0283] In a possible implementation manner, the target feature matching model is a feature matching model trained with the lip movement information of the training users and M pieces of voice information as inputs and the matching degrees between the lip movement information of the training users and the M pieces of voice information as M labels.
[0284] In a possible implementation manner, the server further includes a processor 802; the processor 802 is configured to: determine the user information of the target user, where the user information includes one or more of personal attribute information, facial expression information corresponding to the voice information to be recognized, and environmental information corresponding to the voice information to be recognized; generate a control instruction matching the user information based on the user information.
[0285] In a possible implementation manner, the server further includes a processor 802; the processor 802 is further configured to: based on the face recognition algorithm, recognize N face regions in the image data, and extract the lip movement videos in each of the N face regions; determine the lip movement information of the N users based on the lip movement videos in each of the face regions.
[0286] In a possible implementation, the server further includes a processor 802; the processor 802 is further configured to: based on a spectrum recognition algorithm, recognize audio data of different spectrums in the audio data, and recognize the audio data of the target spectrum as the speech information to be recognized.
[0287] An embodiment of the present invention further provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, it includes some or all of the steps described in any one of the foregoing method embodiments.
[0288] An embodiment of the present invention further provides a computer program, the computer program includes instructions, and when the computer program is executed by a computer, the computer can execute some or all of the steps described in any one of the foregoing method embodiments.
[0289] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0290] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be adopted in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0291] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in an electrical or other form.
[0292] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0293] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0294] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc., specifically, the processor in the computer device) to execute all or part of the steps of the above-mentioned methods in each embodiment of the present application. Among them, the aforementioned storage medium may include: USB flash drives, mobile hard disks, magnetic disks, optical disks, read-only memory (ROM), or random access memory (RAM), etc., various media that can store program codes.
[0295] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A voice command control method, characterized in that, Including: Obtain a first type of instruction and lip movement information of in-vehicle members at N positions in the vehicle cabin during a target time period. The first type of instruction is obtained according to target audio data collected in the vehicle cabin. The lip movement information of the in-vehicle members is extracted from in-vehicle image data when the first type of instruction is recognized in the target audio data. The target time period is the time period corresponding to the first type of instruction in the audio data. Match the first type of instruction with the lip movement information of the in-vehicle members at N positions in the vehicle cabin, and obtain a target position according to the matching result between the lip movement information of the in-vehicle members at the N positions and the first type of instruction. The target position is the position where the in-vehicle member whose lip movement information matches the first type of instruction as indicated by the matching result is located. Send indication information instructing to execute the first type of instruction for the target position. The step of matching the first type of instruction with the lip movement information of the in-vehicle members at N positions in the vehicle cabin and obtaining a target position according to the matching result between the lip movement information of the in-vehicle members at the N positions and the first type of instruction is specifically: according to the first type of instruction and the lip movement information of the in-vehicle members at N positions in the vehicle cabin, obtain the matching degree between the lip movement information of the in-vehicle members at each position and the instruction information respectively; take the position where the in-vehicle member corresponding to the lip movement information with the highest matching degree is located as the target position.
2. The method according to claim 1, wherein: The step of obtaining a first type of instruction and lip movement information of in-vehicle members at N positions in the vehicle cabin is specifically: Obtain the target audio data in the vehicle cabin; When it is recognized that the target audio data includes a first type of instruction, obtain in-vehicle image data; Extract the lip movement information of the in-vehicle members at N positions in the vehicle cabin from the in-vehicle image data.
3. The method according to claim 2, wherein: The step of extracting the lip movement information of the in-vehicle members at N positions in the vehicle cabin from the in-vehicle image data is specifically: When the number of recognized in-vehicle members is greater than 1, extract the lip movement information of the in-vehicle members at N positions in the vehicle cabin from the in-vehicle image data.
4. The method according to any one of claims 1-3, characterized in that: N is an integer greater than 1.
5. The method according to claim 1, wherein The first type of instruction is a voice waveform sequence extracted from the audio data or text instruction information recognized according to the audio data.
6. The method according to claim 1, wherein The lip movement information of the in-vehicle members at N positions in the vehicle cabin is an image sequence of the lip movements of the in-vehicle members at the N positions in the vehicle cabin during the target time period.
7. The method according to any one of claims 1-3, 5-6, characterized in that, The method further includes: Generate the correspondence between the lip movement information of the in-vehicle members at the N positions and the N positions; The step of taking the position where the in-vehicle member corresponding to the lip movement information with the highest matching degree is located as the target position includes: Obtain the target lip movement information with the highest matching degree; Determine the position corresponding to the target lip movement information as the target position according to the correspondence between the lip movement information of the in-vehicle members at the N positions and the N positions.
8. The method according to any one of claims 1-3, 5-6, characterized in that, The method further includes: generating a correspondence between the lip movement information of the in-vehicle members at the N positions and the identities of the in-vehicle members at the N positions; The step of using the position where the in-vehicle member corresponding to the lip movement information with the highest matching degree is located as the target position includes: obtaining the target lip movement information with the highest matching degree; determining the target in-vehicle member according to the correspondence between the lip movement information of the in-vehicle members at the N positions and the identities of the in-vehicle members at the N positions; determining the position information of the target in-vehicle member as the target position, where the position information of the target in-vehicle member is determined according to the sensor data in the vehicle.
9. The method according to any one of claims 1-3, 5-6, wherein the in-vehicle audio data is obtained according to the data collected by multiple microphones in the vehicle cabin, or the in-vehicle audio data is obtained according to the audio data collected by the microphones in a specified position area in the vehicle cabin.
10. According to the method described in any one of claims 1-3, 5-6, it is characterized in that The first type of instruction is an in-vehicle control instruction.
11. A voice command control method, characterized in that, including: obtaining a second type of instruction and the lip movement information of the in-vehicle member at the first position in the vehicle, where the second type of instruction is obtained according to the target audio data collected in the vehicle cabin, the in-vehicle member at the first position has the control authority for the second type of instruction, the lip movement information of the in-vehicle member at the first position is extracted from the in-vehicle image data when the second type of instruction is recognized from the target audio data, and the target time period is the time period corresponding to the second type of instruction in the audio data; matching the second type of instruction and the lip movement information of the in-vehicle member at the first position to obtain a matching result; when it is determined according to the matching result that the second type of instruction and the lip movement information of the in-vehicle member at the first position are matched, sending an indication information for instructing to execute the second type of instruction.
12. The method according to claim 11, wherein The first position is the driver's seat.
13. The method according to any one of claims 11-12, wherein: The step of obtaining the second type of instruction and the lip movement information of the in-vehicle member at the first position in the vehicle specifically is: obtaining the target audio data in the vehicle cabin; when it is recognized that the target audio data includes a second type of instruction, obtaining the in-vehicle image data; extracting the lip movement information of the in-vehicle member at the first position from the in-vehicle image data.
14. The method according to any one of claims 11-12, wherein: The step of matching the second type of instruction and the lip movement information of the in-vehicle member at the first position to obtain a matching result specifically is: determining the matching result of the second type of instruction and the lip movement information of the in-vehicle member at the first position according to the matching degree between the second type of instruction and the lip movement information of the in-vehicle member at the first position and a preset threshold.
15. The method according to claim 11, wherein The second type of instruction is a speech waveform sequence extracted from the audio data or a text instruction information recognized according to the audio data.
16. The method according to claim 11, wherein The lip movement information of the in-vehicle member at the first position is an image sequence of the lip movement of the in-vehicle member at the first position during the target time period.
17. The method according to any one of claims 11-12, 15-16, characterized in that The method further includes: When the second type of instruction is included in the audio data, obtain the image data of the vehicle occupants at N other positions in the vehicle; Extract the lip movement information of the vehicle occupants at N other positions in the vehicle from the image data of the vehicle occupants at N other positions in the vehicle during the target time period; Match the second type of instruction with the lip movement information of the vehicle occupant at the first position in the vehicle to obtain a matching result, specifically: Match the second type of instruction with the lip movement information of the vehicle occupant at the first position in the vehicle and the lip movement information of the vehicle occupants at the N positions to obtain the matching degrees between the lip movement information of N + 1 vehicle occupants and the second type of instruction respectively, and obtain the lip movement information with the highest matching degree; When it is determined according to the matching result that the second type of instruction and the lip movement information of the vehicle occupant at the first position in the vehicle are matched, send an indication information indicating the execution of the second type of instruction, specifically: When the lip movement information with the highest matching degree is the lip movement information of the vehicle occupant at the first position in the vehicle, send an indication information indicating the execution of the second type of instruction.
18. The method according to any one of claims 11-12, 15-16, characterized in that The in-vehicle audio data is obtained based on the data collected by multiple microphones in the vehicle cabin, or The in-vehicle audio data is obtained based on the audio data collected by the microphones in the specified position area in the vehicle cabin.
19. The method according to any one of claims 11-12, 15-16, characterized in that, The second type of instruction is a vehicle driving control instruction.
20. A voice command control device, characterized in that, Including a processor; the processor is used for: Obtain the first type of instruction and the lip movement information of the vehicle occupants at N positions in the vehicle cabin during the target time period, where the first type of instruction is obtained based on the target audio data collected in the vehicle cabin, and the lip movement information of the vehicle occupants is extracted from the in-vehicle image data when the first type of instruction is recognized from the target audio data, and the target time period is the time period corresponding to the first type of instruction in the audio data; Match the first type of instruction with the lip movement information of the vehicle occupants at N positions in the vehicle cabin, and obtain the target position according to the matching result between the lip movement information of the vehicle occupants at the N positions and the first type of instruction, where the target position is the position where the vehicle occupant whose lip movement information is indicated to match the first type of instruction by the matching result is located; Send an indication information indicating the execution of the first type of instruction for the target position; The processor matches the first type of instruction with the lip movement information of the vehicle occupants at N positions in the vehicle cabin, and obtains the target position according to the matching result between the lip movement information of the vehicle occupants at the N positions and the first type of instruction, specifically: the processor obtains the matching degrees between the lip movement information of each vehicle occupant at each position and the instruction information according to the first type of instruction and the lip movement information of N vehicle occupants in the vehicle cabin; the position where the vehicle occupant corresponding to the lip movement information with the highest matching degree is located is used as the target position.
21. The device according to claim 20, wherein The processor is configured to obtain the target audio data in the vehicle cabin; when it is recognized that the target audio data includes a first type of instruction, obtain the image data in the vehicle cabin; and extract the lip movement information of the vehicle occupants at N positions in the vehicle cabin from the image data in the vehicle cabin.
22. The device according to claim 20, wherein: The processor is further configured to, when it is recognized that there are more than one vehicle occupants, extract the lip movement information of the vehicle occupants at N positions in the vehicle cabin from the image data in the vehicle cabin.
23. The device according to any one of claims 20-22, characterized in that: N is an integer greater than 1.
24. The device according to any one of claims 20-22, wherein The processor is further configured to generate a correspondence between the lip movement information of the vehicle occupants at the N positions and the N positions; The processor determines the position where the vehicle occupant corresponding to the lip movement information with the highest matching degree is located as the target position, including: The processor obtains the target lip movement information with the highest matching degree; and determines the position corresponding to the target lip movement information as the target position according to the correspondence between the lip movement information of the vehicle occupants at the N positions and the N positions.
25. The device according to any one of claims 20-22, wherein The processor generates a correspondence between the lip movement information of the vehicle occupants at the N positions and the identities of the vehicle occupants at the N positions; The processor determines the position where the vehicle occupant corresponding to the lip movement information with the highest matching degree is located as the target position, including: The processor obtains the target lip movement information with the highest matching degree; determines the target vehicle occupant according to the correspondence between the lip movement information of the vehicle occupants at the N positions and the identities of the vehicle occupants at the N positions; and determines the position information of the target vehicle occupant as the target position, where the position information of the target vehicle occupant is determined according to the sensor data in the vehicle.
26. A voice command control device, characterized in that, Comprising a processor; the processor is configured to: Obtain a second type of instruction and the lip movement information of the vehicle occupant at the first position in the vehicle. The second type of instruction is obtained according to the target audio data collected in the vehicle cabin. The vehicle occupant at the first position has the control authority for the second type of instruction. The lip movement information of the vehicle occupant at the first position is extracted from the image data in the vehicle cabin when the second type of instruction is recognized in the target audio data. The target time period is the time period corresponding to the second type of instruction in the audio data; Match the second type of instruction and the lip movement information of the vehicle occupant at the first position to obtain a matching result; When it is determined according to the matching result that the second type of instruction and the lip movement information of the vehicle occupant at the first position are matched, send an indication message indicating the execution of the second type of instruction.
27. The device according to claim 26, wherein, The first position is the driver's seat.
28. The device according to any one of claims 26-27, wherein The processor obtains the second type of instruction and the lip movement information of the vehicle occupant at the first position in the vehicle, specifically: The processor obtains the target audio data in the vehicle cabin; when it is recognized that the target audio data includes a second type of instruction, it obtains the image data in the vehicle cabin; and extracts the lip movement information of the vehicle occupant at the first position from the image data in the vehicle cabin.
29. The device according to any one of claims 26-27, wherein: The processor matches the second type of instruction with the lip movement information of the vehicle occupant at the first position to obtain a matching result, specifically: The processor determines the matching result of the second type of instruction and the lip movement information of the vehicle occupant at the first position according to the matching degree between the second type of instruction and the lip movement information of the vehicle occupant at the first position and a preset threshold.
30. The device according to claim 26, wherein, The second type of instruction is a speech waveform sequence extracted from the audio data or text instruction information recognized according to the audio data.
31. The device according to claim 26, wherein, The lip movement information of the vehicle occupant at the first position is an image sequence of the lip movement of the vehicle occupant at the first position during the target time period.
32. The device according to any one of claims 26-27, 30-31, wherein The processor is further configured to, when the audio data includes the second type of instruction, obtain the image data of the vehicle occupants at the other N positions in the vehicle; extract the lip movement information of the vehicle occupants at the other N positions in the vehicle from the image data of the vehicle occupants at the other N positions in the vehicle during the target time period; The processor matches the second type of instruction with the lip movement information of the vehicle occupant at the first position in the vehicle to obtain a matching result, specifically: The processor matches the second type of instruction with the lip movement information of the vehicle occupant at the first position in the vehicle and the lip movement information of the vehicle occupants at the N positions to obtain the matching degrees between the lip movement information of the N+1 vehicle occupants and the second type of instruction respectively, and obtains the lip movement information with the highest matching degree; When the processor determines that the second type of instruction and the lip movement information of the vehicle occupant at the first position in the vehicle are matched according to the matching result, it sends an indication information for instructing to execute the second type of instruction, specifically: When the lip movement information with the highest matching degree is the lip movement information of the vehicle occupant at the first position in the vehicle, the processor sends an indication information for instructing to execute the second type of instruction.
33. The device according to any one of claims 26-27, 30-31, characterized in that The second type of instruction is a vehicle driving control instruction.
34. A chip system, characterized in that, The chip system includes at least one processor, a memory and an interface circuit. The memory, the interface circuit and the at least one processor are interconnected by lines. Instructions are stored in the at least one memory; when the instructions are executed by the processor, the method according to any one of claims 1-19 is implemented.
35. A computer-readable storage medium, characterized in that, The computer-readable medium is used to store program code, and the program code includes instructions for executing the method according to any one of claims 1-19.
36. A computer program, characterized in that, The computer program includes instructions, and when the computer program is executed, the method according to any one of claims 1-19 is implemented.
Citation Information
Patent Citations
Regulating system and method for devices in vehicle
CN105667433A
Vehicle-borne voice control method, device and equipment
CN105867179A
Cited By
Speech instruction control method in vehicle cabin and related device
EP4682846A2