Voice instruction control method and device of vehicle, electronic equipment, vehicle and storage medium
By collecting image data and audio data in the vehicle, and using lip movement information to determine the target occupant of the voice command, the position detection error caused by swaying of the occupant in the vehicle is solved, and the accuracy of voice control is improved.
Patent Information
- Application Number
- CN202510725128.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-26
AI Technical Summary
In the prior art, the position detection error occurs due to shaking of the occupant in the vehicle, which leads to the problem of voice control errors.
By collecting image data inside the vehicle, extracting lip movement information, combining voice commands in the audio data, the target occupant in the vehicle that issues the voice command, and controlling their corresponding cockpit area to perform operations.
Improves the accuracy of voice control and reduces position detection errors due to occupant shaking.
Smart Images

Figure CN120544574A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle technology, and specifically to a method, device, electronic device, vehicle, and storage medium for controlling a vehicle by voice command. Background Art
[0002] Vehicles typically include onboard intelligent voice interaction devices. Based on speech recognition, natural language processing, and artificial intelligence algorithms, these devices enable voice interaction between vehicle occupants and the vehicle. They can also execute corresponding actions based on the occupants' voice commands, thereby enhancing the occupants' experience. For example, if a occupant issues a voice command to "open the first window," the vehicle will execute the "open the first window" action; if a occupant issues a voice command to "open the second window," the vehicle will execute the "open the second window" action.
[0003] Furthermore, to simplify voice commands, the location of the target occupant (the occupant issuing the voice command) can be combined with the voice command. The vehicle then performs the corresponding operation based on the target occupant's location and the voice command. For example, if the first occupant in the vehicle needs to open the first window at their location, they only need to issue the voice command "open window." Based on the first occupant's location and the voice command "open window," the vehicle can determine that the first window needs to be opened and then execute the operation to open the first window. It can be understood that compared to the voice command "open first window," the voice command "open window" is more concise and easier for the user to operate.
[0004] In the related art, directional beam picking can be used to collect audio data. This collected audio data is then detected using directional sound source detection technology to determine the position of the vehicle occupants. However, if the occupants' movement is excessive, this may lead to errors in the occupant's position detection, which in turn may cause voice control errors. For example, when the first occupant in the vehicle needs to open the first window in their location, they issue a voice command of "open the window." However, when the first occupant issues the voice command of "open the window," their body sways and they lean toward the position of the second occupant in the vehicle. Using the position detection scheme described in the related art, the vehicle mistakenly identifies the occupant who issued the voice command of "open the window" as the second occupant in the vehicle, and then executes the operation of opening the second window (i.e., the window in the location of the second occupant), resulting in voice control errors.
[0005] It should be pointed out that the information disclosed in the background technology section of this application is only intended to deepen the understanding of the general background technology of this application, and should not be regarded as an admission or any form of implication that the information constitutes prior art already known to those skilled in the art. Summary of the Invention
[0006] The present application provides a vehicle voice command control method, device, electronic device, vehicle and storage medium to help solve the problem in the prior art of detecting the position of the occupant in the vehicle who issues the voice command, which may lead to position detection errors and then voice control errors due to excessive shaking of the occupant in the vehicle.
[0007] In a first aspect, an embodiment of the present application provides a method for controlling a vehicle via voice commands, comprising: collecting first image data and first audio data of an interior of a vehicle, wherein the first audio data includes a voice command; extracting lip movement information of each vehicle occupant from the first image data; extracting voice instructions from the first audio data; determining, based on the lip movement information of each of the vehicle occupants, a target vehicle occupant who issues the voice command among the vehicle occupants; The cockpit area corresponding to the occupant in the target vehicle is controlled to perform an operation corresponding to the voice command.
[0008] In a possible implementation, determining, based on the lip movement information of each of the vehicle occupants, a target vehicle occupant who issues the voice command, includes: According to the lip movement information of each of the vehicle occupants and the voice command, a target vehicle occupant who issues the voice command is determined among the vehicle occupants, and the lip movement information of the target vehicle occupant matches the voice command.
[0009] In one possible implementation, collecting the first image data and the first audio data of the interior of the vehicle includes: collecting second image data of the interior of the vehicle; If it is determined based on the second image data that there is a passenger inside the vehicle, first image data and first audio data of the interior of the vehicle are collected.
[0010] In one possible implementation, collecting the first image data and the first audio data of the interior of the vehicle includes: collecting second audio data inside the vehicle; If it is determined based on the second audio data that there is a passenger inside the vehicle, first image data and first audio data of the interior of the vehicle are collected.
[0011] In a possible implementation, after collecting the first image data of the interior of the vehicle, the method further includes: A cabin area corresponding to each vehicle occupant is determined based on the first image data.
[0012] In a possible implementation, after determining, based on the lip movement information of each of the vehicle occupants, a target vehicle occupant who issues the voice command, the method further includes: A cabin area corresponding to the occupant in the target vehicle is determined according to the first image data.
[0013] In a second aspect, an embodiment of the present application provides a vehicle voice command control device, comprising: a data acquisition module, configured to acquire first image data and first audio data from inside the vehicle, wherein the first audio data includes a voice command; a lip movement information extraction module, configured to extract lip movement information of each vehicle occupant from the first image data; A voice command extraction module, configured to extract voice commands from the first audio data; a target vehicle occupant determination module, configured to determine a target vehicle occupant who issues the voice command among the vehicle occupants based on lip movement information of each vehicle occupant; The voice command execution module is used to control the cabin area corresponding to the occupant in the target vehicle to perform an operation corresponding to the voice command.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, including: processor; Memory; and a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, the electronic device executes the method described in any one of the first aspects.
[0015] In a fourth aspect, an embodiment of the present application provides a vehicle, comprising: A controller is configured to execute the method according to any one of the first aspects.
[0016] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of the first aspects is implemented.
[0017] In this embodiment of the present application, the occupant issuing the voice command (i.e., the target occupant) is identified using lip movement information in the image data, and the target occupant's position can be determined. Because the target occupant's position is detected using image data, it is generally not affected by the target occupant's movement, thereby improving detection accuracy and, in turn, voice control accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 A schematic diagram of an application scenario provided for an embodiment of the present application.
[0020] Figure 2 A schematic structural diagram of a vehicle provided in an embodiment of the present application.
[0021] Figure 3 A flow chart of a method for controlling a vehicle by voice commands provided in an embodiment of the present application.
[0022] Figure 4 A flowchart of another method for controlling a vehicle by voice commands provided in an embodiment of the present application.
[0023] Figure 5 A flowchart of another method for controlling a vehicle by voice commands provided in an embodiment of the present application.
[0024] Figure 6 A schematic structural diagram of a voice command control device for a vehicle provided in an embodiment of the present application.
[0025] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0026] Figure 8 A schematic structural diagram of another vehicle provided in an embodiment of the present application. DETAILED DESCRIPTION
[0027] In order to better understand the technical solution of the present application, the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0028] It should be clear that the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0029] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "an", "the" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0030] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. Furthermore, the character " / " in this document generally indicates an "or" relationship between the associated objects.
[0031] See also Figure 1 , is a schematic diagram of an application scenario provided by an embodiment of the present application. Figure 1 As shown, vehicle 100 generally includes an in-vehicle intelligent voice interaction device 101. Based on voice recognition, natural language processing, and artificial intelligence algorithms, in-vehicle intelligent voice interaction device 101 can realize voice interaction between vehicle occupants and vehicle 100, and can execute corresponding operations based on the voice commands of the vehicle occupants, thereby improving the experience of the vehicle occupants.
[0032] It is understood that the space inside the vehicle 100 can generally be divided into a plurality of cabin areas ("cabin areas" may also be referred to as "locations" elsewhere in this document). For example, Figure 2 In the vehicle, from front to back, the first-row cabin area 201 and the second-row cabin area 202 are sequentially included along the vehicle's travel direction. The first-row cabin area 201 includes two cabin areas, namely, a first-row left cabin area 2011 and a first-row right cabin area 2012; the second-row cabin area 202 includes three cabin areas, namely, a second-row left cabin area 2021, a second-row middle cabin area 2022, and a second-row right cabin area 2023.
[0033] In actual application, when the passenger in the car issues a voice command to "open the first row left cabin area window", the vehicle executes the operation of "opening the first row left cabin area window"; when the passenger in the car issues a voice command to "open the first row right cabin area window", the vehicle executes the operation of "opening the first row right cabin area window".
[0034] Furthermore, to simplify voice commands, the location of the target occupant (the one issuing the voice command) can be combined with the voice command. The vehicle then uses the target occupant's location and the voice command to execute the corresponding operation. For example, if a occupant in the first-row, left-side cabin area wants to open the first-row left-side cabin window, they only need to issue the voice command "open window." Based on the occupant's location (i.e., the first-row left-side cabin area) and the voice command "open window," the vehicle can determine that the first-row left-side cabin window needs to be opened and then execute the operation to open the first-row left-side cabin window. It can be understood that compared to the voice command "open first-row left-side cabin window," the voice command "open window" is more concise and easier for the user to operate.
[0035] It should be pointed out that Figure 2 The vehicle shown in the figure is only an exemplary illustration. The vehicle may also include other numbers of cabin areas, for example, three-row cabin areas, five-row cabin areas, etc., and the embodiments of the present application do not impose specific limitations on this.
[0036] In the related art, it is possible to collect audio data by picking up sound in a directional beam area; then, the collected audio data is detected based on directional sound source detection technology to determine the position of the occupants in the vehicle. However, when the occupants in the vehicle shake too much, it may cause errors in the position detection of the occupants in the vehicle, which in turn may cause errors in voice control. For example, when the occupant in the first row, left cabin area needs to open the window in the first row, left cabin area, he or she issues a voice command of "open the window". However, when the occupant issues the voice command of "open the window", due to the shaking of his or her body, he or she leans towards the position of the occupant in the first row, right cabin area. Through the above-mentioned position detection scheme in the related art, the vehicle mistakenly judges the occupant in the vehicle who issued the voice command of "open the window" as a occupant in the first row, right cabin area, and then executes the operation of opening the window in the first row, right cabin area, resulting in voice control errors.
[0037] To address the above issues, in an embodiment of the present application, lip movement information in image data is used to identify the vehicle occupant issuing the voice command (i.e., the target occupant), and thus the target occupant's position. Because the target occupant's position is detected using image data, it is generally not affected by the target occupant's movement, thereby improving detection accuracy and, in turn, voice control accuracy.
[0038] Specifically, a detailed description is given below with reference to the accompanying drawings and specific embodiments.
[0039] See also Figure 3 , is a flow chart of a method for controlling a vehicle by voice command provided in an embodiment of the present application. This method can be applied to Figure 1 The in-vehicle intelligent voice interaction device shown in Figure 3 As shown, it specifically includes the following steps.
[0040] Step S301: collecting first image data and first audio data inside the vehicle.
[0041] In practical applications, when capturing images of a vehicle's interior, the captured image data should generally include the entire vehicle cabin area. If there are passengers in a certain cabin area, the captured image data generally includes the passenger's portrait information.
[0042] Illustratively, when there are passengers in the first-row left cabin area, the collected image data generally includes portrait information of the passengers in the first-row left cabin area; when there are passengers in both the first-row left cabin area and the first-row right cabin area, the collected image data generally includes portrait information of the passengers in the first-row left cabin area and the first-row right cabin area; when there are passengers in the first-row left cabin area, the first-row right cabin area and the second-row left cabin area, the collected image data generally includes portrait information of the passengers in the first-row left cabin area, the first-row right cabin area and the second-row left cabin area.
[0043] It is understood that portrait information typically includes lip movement information. In this embodiment of the present application, after obtaining image data, the lip movement information of the vehicle occupants can be extracted from the image data. The specific content of "lip movement information" is described in detail below and will not be repeated here. To facilitate differentiation from other image data described below, in this embodiment of the present application, the image data used to extract lip movement information is referred to as "first image data."
[0044] In the embodiment of the present application, the image data of the interior of the vehicle can be collected by the vehicle-mounted camera. It should be noted that the number of the vehicle-mounted cameras may be one or more, and the embodiment of the present application does not impose a specific limitation on this.
[0045] In practical applications, when a vehicle occupant needs to control the vehicle, they may issue voice commands. Understandably, the audio data collected in this situation typically includes voice commands. Of course, the audio data may also contain other audio information besides voice commands. For example, information unrelated to voice commands or noise emitted by the vehicle occupant may also be present.
[0046] In the embodiment of the present application, after obtaining the audio data, the voice command issued by the vehicle occupant can be extracted from the audio data. In order to distinguish it from other audio data in the following text, in the embodiment of the present application, the audio data used to extract the voice command is referred to as "first audio data".
[0047] In the embodiment of the present application, the audio data inside the vehicle can be collected by the vehicle-mounted microphone. It should be noted that the number of the vehicle-mounted microphone may be one or more, and the embodiment of the present application does not impose a specific limitation on this.
[0048] In practical applications, as described above, the location of the vehicle occupant issuing the voice command can be combined with the voice command to execute the corresponding operation. It is understood that when there are occupants inside the vehicle, the occupant's location and voice command can be used to execute the corresponding operation. However, when there are no occupants inside the vehicle, it is obviously not necessary to determine the occupant's location and voice command to execute the corresponding operation, and therefore, there is no need to collect image data and audio data. In this case, if the vehicle continuously collects image data and audio data, it will increase the vehicle's performance overhead.
[0049] To address the above problem, it is possible to first determine whether there is an occupant in the vehicle. If there is an occupant in the vehicle, the first image data and the first audio data are collected; if there is no occupant in the vehicle, the first image data and the first audio data are not collected.
[0050] In practical applications, there may be multiple "schemes for determining whether there are passengers inside a vehicle." For ease of understanding, three "schemes for determining whether there are passengers inside a vehicle" are provided in the embodiments of this application and are described below.
[0051] The first solution: determining whether there are passengers inside the vehicle through image data. To distinguish it from the "first image data" mentioned above, in this embodiment of the application, the image data used to determine whether there are passengers inside the vehicle is referred to as "second image data."
[0052] In the embodiment of the present application, step S301 specifically includes: collecting second image data of the interior of the vehicle; if it is determined that there is a passenger inside the vehicle based on the second image data, collecting first image data and first audio data of the interior of the vehicle.
[0053] It is understood that when there is a passenger inside the vehicle, the second image data typically includes portrait information. Therefore, the presence of a passenger inside the vehicle can be determined by the presence of portrait information in the second image data. Specifically, after obtaining the second image data, the second image data is analyzed to determine whether there is portrait information in the second image data. If the second image data contains portrait information, it is determined that there is a passenger inside the vehicle; if the second image data does not contain portrait information, it is determined that there is no passenger inside the vehicle.
[0054] In the embodiment of the present application, the second image data can accurately record the entire cabin area inside the vehicle, and thus can more accurately determine whether there are passengers inside the vehicle.
[0055] The second solution: determining whether there are passengers inside the vehicle through audio data. To distinguish it from the "first audio data" mentioned above, in this embodiment of the application, the audio data used to determine whether there are passengers inside the vehicle is referred to as "second audio data."
[0056] In an embodiment of the present application, step S301 specifically includes: collecting second audio data from the interior of the vehicle; if it is determined based on the second audio data that there is a passenger inside the vehicle, collecting first image data and first audio data from the interior of the vehicle.
[0057] It will be appreciated that when a passenger is present in the vehicle and the passenger is speaking, the second audio data typically includes voice information of the passenger's speech. Therefore, the presence of speech information in the second audio data can be used to determine whether a passenger is present in the vehicle. Specifically, after obtaining the second audio data, the second audio data is analyzed to determine whether speech information is present. If speech information is present in the second audio data, it is determined that a passenger is present in the vehicle; if speech information is absent in the second audio data, it is determined that no passenger is present in the vehicle.
[0058] In the embodiments of the present application, the presence of a passenger in the vehicle can be determined based on the second audio data. However, in some application scenarios, when there are passengers in the vehicle but no one is speaking, determining the presence of a passenger based on the second audio data may lead to a false assumption that no passengers are present, resulting in a false detection. Because the second image data in the first solution accurately captures the entire cabin area, the first solution can more accurately determine the presence of a passenger in the vehicle compared to the second solution, improving accuracy to a certain extent.
[0059] The third solution is to determine whether there are passengers inside the vehicle through image data and audio data.
[0060] In the embodiment of the present application, step S301 specifically includes: collecting second image data and second audio data of the vehicle interior; and if it is determined based on the second image data and second audio data that there is an occupant in the vehicle interior, collecting first image data and first audio data of the vehicle interior. The specific details of the examples in the implementation of this application can be found in the description of the above-illustrated embodiment, and for the sake of brevity, they are not further described.
[0061] In the embodiment of the present application, the first image data and the first audio data need only be collected when it is determined that there is a passenger inside the vehicle. It is understood that only when there is a passenger inside the vehicle can the passenger issue a voice command, and thus it is necessary to determine the passenger's position and voice command. Therefore, the first image data and the first audio data need only be collected when there is a passenger inside the vehicle, which reduces the data processing volume to a certain extent and saves vehicle performance overhead.
[0062] Step S302: extracting lip movement information of each vehicle occupant from the first image data.
[0063] In the embodiment of the present application, as described above, the collected first image data generally includes portrait information of the occupant in the vehicle, and the portrait information of the occupant in the vehicle generally includes lip movement information of the occupant in the vehicle.
[0064] In practical applications, whether a vehicle occupant is speaking is often determined by their lip movements. It is understood that the lip movement information of a vehicle occupant speaking is different from that of a vehicle occupant not speaking. Therefore, whether a vehicle occupant is speaking can be determined based on their lip movement information.
[0065] In the embodiment of the present application, since there may be multiple occupants in the vehicle, before determining whether the occupant is speaking based on the lip movement information of the occupant, it is necessary to extract the lip movement information of each occupant from the first image data. Specifically, the lip movement information of the occupant can be extracted from the first image data using an image algorithm.
[0066] It should be pointed out that extracting the lip movement information of the occupants in the vehicle from the first image data through an image algorithm is only an exemplary description. Those skilled in the art can extract the lip movement information of the occupants in the vehicle from the first image data through other methods, and the embodiments of the present application do not impose specific restrictions on this.
[0067] Step S303: extracting voice instructions from the first audio data.
[0068] In practical applications, as described above, the collected first audio data may include not only voice commands but also other audio information, such as information sent by passengers in the vehicle that is unrelated to the voice commands, or noise.
[0069] In an embodiment of the present application, the vehicle needs to perform an operation corresponding to the voice command based on the voice command. Therefore, before performing the operation corresponding to the voice command, the voice command can be extracted from the first audio data. Specifically, the voice command can be extracted from the first audio data using a voice command extraction algorithm.
[0070] It should be pointed out that extracting voice commands from the first audio data through a voice command extraction algorithm is only an exemplary description. Those skilled in the art can extract the lip movement information of the vehicle occupants from the first audio data through other methods, and the embodiments of the present application do not impose specific restrictions on this.
[0071] Step S304: Determine a target vehicle occupant who issues a voice command among the vehicle occupants based on the lip movement information of each vehicle occupant.
[0072] In practical applications, since the occupant issuing the voice command is usually the one currently speaking, the speaking occupant can be used as the occupant issuing the voice command. Furthermore, lip movement information can be used to determine whether each occupant is currently speaking, thereby determining the target occupant issuing the voice command.
[0073] For example, Figure 2 In the vehicle shown, when the passenger in the first row left cabin area is speaking, the lip movement information of the passenger in the first row left cabin area can be used to determine that the passenger in the first row left cabin area is speaking, and further, the passenger in the first row left cabin area can be determined as the target passenger for issuing the voice command.
[0074] However, in some application scenarios, when there may be multiple occupants in the vehicle speaking at the same time, but only one occupant issues a voice command, the speaking occupant is determined to be the target occupant issuing the voice command based solely on the lip movement information of the occupant. The occupant who does not speak may be mistakenly identified as the target occupant issuing the voice command, resulting in an error in identifying the target occupant and, in turn, an error in voice control.
[0075] For example, Figure 2 In the vehicle shown, when the occupants in the first, right, and second row left cabin areas speak simultaneously, and only the first row left cabin occupant issues a voice command, the lip movement information from these occupants identifies the first row left cabin area, first row right cabin area, and second row left cabin area as the target occupant issuing the voice command. Theoretically, only the first row left cabin area occupant is the target occupant, but the above method for determining the target occupant will mistakenly identify the first row right cabin area and second row left cabin area as the target occupant. Clearly, identifying the first row right cabin area and second row left cabin area as the target occupant is a misidentification.
[0076] To address the above problem, the target vehicle occupant who issues the voice command can be determined by judging whether the lip movement information of the vehicle occupant matches the voice command.
[0077] In an embodiment of the present application, step S304 specifically includes: determining the target occupant in the car who issues the voice command among the occupants in the car based on the lip movement information and voice command of each occupant in the car, and the lip movement information of the target occupant in the car matches the voice command.
[0078] It can be understood that if the lip movement information of the vehicle occupant matches the voice command, it is determined that the vehicle occupant is the target vehicle occupant who issued the voice command; if the lip movement information of the vehicle occupant does not match the voice command, it is determined that the vehicle occupant is not the target vehicle occupant who issued the voice command.
[0079] In this embodiment of the present application, when multiple vehicle occupants are speaking simultaneously, the target occupant issuing the voice command can be identified from among the multiple speaking occupants by determining whether the lip movement information of each occupant matches the voice command. Since the occupant whose lip movement information matches the voice command is the target occupant, it is possible to avoid misidentifying other occupants who are speaking but have not issued a voice command as the target occupant, thereby improving the accuracy of target occupant identification and, in turn, voice control.
[0080] In practical applications, as described above, the vehicle performs corresponding operations based on the position of the occupants and voice commands. Therefore, it is necessary to determine the position of the occupants inside the vehicle.
[0081] In the embodiments of the present application, two "solutions for determining the position of the occupants inside the vehicle" are provided, which are described below respectively.
[0082] Solution 1: After collecting the first image data, the cabin area corresponding to each vehicle occupant can be determined based on the first image data.
[0083] As will be appreciated, as described above, the first image data typically includes the portrait information of the vehicle occupants, indicating the presence of occupants within the vehicle. Therefore, after acquiring the first image data, the position of the occupants (i.e., the cabin area corresponding to the occupants) can be determined based on the first image data. Specifically, after obtaining the first image data, the portrait position relationship of the occupants' portrait information in the first image data can be calculated to determine the cabin area corresponding to the occupants.
[0084] In this embodiment of the present application, the position of each vehicle occupant is determined after obtaining the first image data. Therefore, once the target vehicle occupant issuing the voice command is determined, the cabin zone corresponding to the target vehicle occupant is determined. Further analysis of the first image data to determine the cabin zone corresponding to the target vehicle occupant is unnecessary, eliminating the need for redundant data processing, thereby reducing the data processing load to a certain extent.
[0085] See also Figure 4 , is a flow chart of another method for controlling a vehicle via voice commands provided by an embodiment of the present application. Figure 4 As shown, after step S301, step 400 is further included. The specific contents involved in the embodiment of the present application can be found in the description of the above method embodiment, and for the sake of brevity, no further details will be given.
[0086] It should be pointed out that Figure 4 The embodiment shown is only an exemplary description, and step 400 may also be performed after step 302 or step S303. This embodiment of the present application does not impose any specific limitation on this.
[0087] Solution 2: After determining the target occupant who issues the voice command among the occupants based on the lip movement information of each occupant, the cabin area corresponding to the target occupant can be determined based on the first image data.
[0088] It is understood that after the target vehicle occupant issuing the voice command is determined, only the position of the target vehicle occupant needs to be determined. Therefore, the cabin area corresponding to the target vehicle occupant can be determined by analyzing only the image data related to the target vehicle occupant in the first image data.
[0089] In an embodiment of the present application, determining the position of the target occupant only requires analyzing the image data related to the target vehicle occupant, and there is no need to process image data unrelated to the target vehicle occupant, which reduces the data processing volume to a certain extent and saves the vehicle's performance overhead.
[0090] See also Figure 5 , is a flow chart of another method for controlling a vehicle via voice commands provided by an embodiment of the present application. Figure 5 As shown, after step S304, step 500 is further included. The specific contents involved in the embodiment of the present application can be found in the description of the above method embodiment, and for the sake of brevity, no further details will be given.
[0091] Step S305: Control the cabin area corresponding to the occupant in the target vehicle to perform an operation corresponding to the voice command.
[0092] In practical applications, corresponding operations can be performed based on the position of the occupants in the vehicle and the voice command. It is understood that based on the position of the occupants in the vehicle, the cabin area where the voice command needs to be executed can be determined, and then the operation corresponding to the voice command can be performed in the cabin area where the voice command needs to be executed.
[0093] In the embodiment of the present application, determining the cabin zone corresponding to the target vehicle occupant is equivalent to determining the cabin zone where the voice command needs to be executed. Therefore, the cabin zone can be controlled to execute the operation corresponding to the voice command.
[0094] In this embodiment of the present application, the occupant issuing the voice command (i.e., the target occupant) is identified using lip movement information in the image data, and the target occupant's position can be determined. Because the target occupant's position is detected using image data, it is generally not affected by the target occupant's movement, thereby improving detection accuracy and, in turn, voice control accuracy.
[0095] Corresponding to the above method embodiment, the present application also provides a voice command control device for a vehicle.
[0096] See also Figure 6 , is a structural diagram of a vehicle voice command control device provided by an embodiment of the present application. Figure 6 As shown, the vehicle voice command control device 600 includes: a data acquisition module 601, a lip movement information extraction module 602, a voice command extraction module 603, a target vehicle occupant determination module 604 and a voice command execution module 605.
[0097] Specifically, the data acquisition module 601 is used to collect first image data and first audio data of the interior of the vehicle, where the first audio data includes voice commands; the lip movement information extraction module 602 is used to extract the lip movement information of each vehicle occupant from the first image data; the voice command extraction module 603 is used to extract voice commands from the first audio data; the target vehicle occupant determination module 604 is used to determine the target vehicle occupant who issues the voice command among the vehicle occupants based on the lip movement information of each vehicle occupant; and the voice command execution module 605 is used to control the cabin area corresponding to the target vehicle occupant to perform an operation corresponding to the voice command.
[0098] The specific contents involved in the embodiments of this application can be found in the description of the above method embodiments. For the sake of brevity, they will not be repeated here.
[0099] Corresponding to the above embodiment, an embodiment of the present application further provides an electronic device.
[0100] See also Figure 7, is a structural diagram of an electronic device provided in an embodiment of the present application. Figure 7 As shown, the electronic device 700 may include: a processor 701, a memory 702, and a communication unit 703. These components communicate via one or more buses. Those skilled in the art will appreciate that the electronic device structure shown in the figure does not limit the embodiments of the present application. It may be a bus structure or a star structure, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0101] The communication unit 703 is used to establish a communication channel so that the electronic device can communicate with other devices.
[0102] The processor 701 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. It runs or executes software programs and / or modules stored in the memory 702, and calls data stored in the memory to perform various functions of the electronic device and / or process data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor 701 can include only a central processing unit (CPU). In the embodiment of the present application, the CPU can be a single computing core or multiple computing cores.
[0103] The memory 702 is used to store execution instructions of the processor 701. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0104] When the execution instructions in the memory 702 are executed by the processor 701 , the electronic device 700 is enabled to execute part or all of the steps in the above method embodiments.
[0105] The specific contents involved in the embodiments of this application can be found in the description of the above method embodiments. For the sake of brevity, they will not be repeated here.
[0106] Corresponding to the above method embodiment, the present application also provides a vehicle.
[0107] See also Figure 8 , is a structural diagram of another vehicle provided in an embodiment of the present application. Figure 8 As shown, the vehicle 800 includes a controller 801 , which is configured to execute the method described in the above method embodiment.
[0108] The specific contents involved in the embodiments of this application can be found in the description of the above method embodiments. For the sake of brevity, they will not be repeated here.
[0109] Corresponding to the above embodiments, embodiments of the present application also provide a computer-readable storage medium, wherein the computer-readable storage medium may store a program. When the program is executed, the device containing the computer-readable storage medium may be controlled to perform some or all of the steps in the above method embodiments. In a specific implementation, the computer-readable storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM). In the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean that A exists alone, A and B exist simultaneously, or B exists alone. A and B can be singular or plural. The character " / " generally indicates that the associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0110] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0111] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0112] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0113] The above description is merely a specific embodiment of the present application. Any person skilled in the art may easily conceive of variations or substitutions within the technical scope disclosed in this application, and such variations or substitutions shall be within the scope of protection of this application. The scope of protection of this application shall be subject to the scope of protection of the claims.
Claims
1. A method for controlling a vehicle by voice command, characterized in that: include: collecting first image data and first audio data of an interior of a vehicle, wherein the first audio data includes a voice command; extracting lip movement information of each vehicle occupant from the first image data; extracting voice instructions from the first audio data; determining, based on the lip movement information of each of the vehicle occupants, a target vehicle occupant who issues the voice command among the vehicle occupants; The cockpit area corresponding to the occupant in the target vehicle is controlled to perform an operation corresponding to the voice command.
2. The method according to claim 1, characterized in that The step of determining, based on the lip movement information of each of the vehicle occupants, a target vehicle occupant who issues the voice command, among the vehicle occupants, includes: According to the lip movement information of each of the vehicle occupants and the voice command, a target vehicle occupant who issues the voice command is determined among the vehicle occupants, and the lip movement information of the target vehicle occupant matches the voice command.
3. The method according to claim 1, characterized in that The collecting of the first image data and the first audio data inside the vehicle includes: collecting second image data of the interior of the vehicle; If it is determined based on the second image data that there is a passenger inside the vehicle, first image data and first audio data of the interior of the vehicle are collected.
4. The method according to claim 1, wherein The collecting of the first image data and the first audio data inside the vehicle includes: collecting second audio data inside the vehicle; If it is determined based on the second audio data that there is a passenger inside the vehicle, first image data and first audio data of the interior of the vehicle are collected.
5. The method according to claim 1, characterized in that After collecting the first image data of the interior of the vehicle, the method further includes: A cabin area corresponding to each vehicle occupant is determined based on the first image data.
6. The method according to claim 1, characterized in that After determining, based on the lip movement information of each of the vehicle occupants, a target vehicle occupant who issues the voice command among the vehicle occupants, the method further includes: A cabin area corresponding to the occupant in the target vehicle is determined according to the first image data.
7. A voice command control device for a vehicle, characterized in that: include: a data acquisition module, configured to acquire first image data and first audio data from inside the vehicle, wherein the first audio data includes a voice command; a lip movement information extraction module, configured to extract lip movement information of each vehicle occupant from the first image data; A voice command extraction module, configured to extract voice commands from the first audio data; a target vehicle occupant determination module, configured to determine a target vehicle occupant who issues the voice command among the vehicle occupants based on lip movement information of each vehicle occupant; The voice command execution module is used to control the cabin area corresponding to the occupant in the target vehicle to perform an operation corresponding to the voice command.
8. An electronic device, characterized in that: include: processor; Memory; and a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, causes the electronic device to perform the method according to any one of claims 1 to 6.
9. A vehicle, characterized in that: include: A controller configured to execute the method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.