Voice instruction execution method and device, vehicle and storage medium
By acquiring the location of the sound source and images of people in the vehicle, and combining historical commands with real-time sound signals to generate target commands, the problem of discontinuous voice command reception caused by changes in user location is solved, achieving efficient and accurate voice command recognition and improving driving safety.
Patent Information
- Application Number
- CN202411181778.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2025-11-04
AI Technical Summary
In existing technologies, the vehicle's voice command receiving unit cannot continuously acquire control commands after the user's location changes, affecting user experience and driving safety.
By acquiring target sound signals and images of people in real-world scenarios, the location of the sound source is determined and the target person is matched. The target command is generated by combining historical command text and real-time sound signals. The beamforming parameters of the receiver array are adjusted to capture and recognize the user's voice commands.
It enables real-time tracking of user position changes inside the vehicle and accurate generation of commands, improving the accuracy and efficiency of command recognition, and enhancing driving safety and user experience.
Smart Images

Figure CN120895029A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicles, and more particularly to voice command execution methods, devices, vehicles, and storage media. Background Technology
[0002] With the continuous development of automotive technology, more and more residents are choosing vehicles as their daily mode of transportation. In related technologies, when control commands are given to a vehicle via voice, the vehicle's command receiving unit can only acquire voice signals from a specific location area. However, when the location of the person issuing the command changes, it becomes impossible to continuously acquire control commands from the same person, thus affecting the user experience. Summary of the Invention
[0003] This application provides a voice command execution method, apparatus, vehicle, and storage medium, aiming to perform real-time dynamic control of the vehicle, thereby improving driving safety. The technical solution is as follows:
[0004] In a first aspect, this application provides a voice command execution method, including: obtaining the sound source location corresponding to the target sound signal in a real scene; obtaining the target person located at the sound source location; generating the target person's target command based on the position matching relationship between the sound source location and the recorded location of the target person, combined with the target person's historical command text and the target sound signal, and executing the target command.
[0005] In the above technical solution, by acquiring the target sound signal and images of each person in the real-world scene, determining the location of the sound source corresponding to the target sound signal, and further acquiring the target person located at that sound source location, the sound source and the person are accurately matched. By combining the positional matching relationship between the sound source location and the recorded location of the target person, as well as the target person's historical command text and the current target sound signal, the target person's target command can be accurately generated, thereby improving the accuracy and efficiency of command recognition.
[0006] Combining the first aspect and the above implementation methods, in some possible implementation methods, obtaining the target personnel located at the sound source location includes: obtaining a panoramic image of the real scene; identifying personnel images in the panoramic image of the scene; determining the regional position of each personnel in the real scene based on the position of the personnel images in the panoramic image of the scene; obtaining the target personnel image of the target personnel located at the sound source location according to the relationship between the regional position and the sound source location; and determining the target personnel located at the sound source location based on the target personnel image.
[0007] In the above technical solution, by acquiring a panoramic image of the real scene and identifying the images of each person in it, a comprehensive visual monitoring of the entire scene is achieved. This helps to track the movement of people in the real scene in real time, thereby enabling accurate tracking of the location of the target person and improving operational efficiency.
[0008] Combining the first aspect and the above implementation methods, in some possible implementation methods, based on the positional matching relationship between the sound source location and the recorded location of the target person, combined with the target person's historical instruction text and target sound signal, a target instruction for the target person is generated and executed. This includes: obtaining the target person's facial features in the target person's image and obtaining the recorded location corresponding to the facial features; if the positional matching relationship between the sound source location and the target person's recorded location is the same, then the target person's target instruction is generated and executed by combining the target person's historical instruction text and target sound signal; if the positional matching relationship between the sound source location and the target person's recorded location is different, then the recorded location is updated using the sound source location, and the target person's target instruction is generated and executed by combining the target person's historical instruction text and target sound signal.
[0009] Combining the first aspect and the above implementation methods, in some possible implementation methods, the target personnel's target instructions are generated by combining the target personnel's historical instruction text and target voice signal, including: obtaining the target personnel's historical instruction text based on the recording location; obtaining the target voice text of the target voice signal; obtaining the target instruction text based on the target voice text and historical instruction text; and performing instruction matching on the target instruction text to generate the target personnel's target instructions.
[0010] In the above technical solution, facial recognition technology is used to capture the facial features of the target person in the image and obtain the corresponding recorded location, thereby improving the accuracy and efficiency of recognition. When the area location is completely consistent with the recorded location, the target command is generated by timely combining the target person's historical command text and target voice signal and integrating the information, thereby achieving high precision in triggering control commands. By real-time determination of the target person's location information, it is confirmed that the target person's location change information can be obtained. On this basis, the target person's historical command text and voice signal are combined again to generate new target commands. Through an adaptive update mechanism, the system can flexibly respond to environmental changes, thereby improving the intelligence level of command generation and enhancing the pertinence and effectiveness of vehicle task execution.
[0011] In combination with the first aspect and the above implementation methods, in some possible implementation methods, after acquiring the target person image located at the sound source location, the method further includes: acquiring the target person image, determining the face region of the target person in the target person image, and adjusting the beamforming parameters of the receiving array corresponding to the sound source location based on the face region.
[0012] Combining the first aspect and the above implementation methods, in some possible implementation methods, determining the face region in the target person image and adjusting the beamforming parameters of the receiving array corresponding to the sound source location based on the face region includes: performing facial feature analysis on the target person image to determine the face region of the target person and the location information of the face region in the real scene; determining the phase adjustment parameters and amplitude adjustment parameters of the receiving array based on the location information; and generating beamforming parameters based on the phase adjustment parameters and amplitude adjustment parameters.
[0013] In the above technical solution, by acquiring the scene area location of each person's image and matching the scene area location with the sound source location, the target person image containing the target person can be accurately determined. Based on this, the face region of the target person is further determined, and the beamforming parameters of the receiving array corresponding to the sound source location are adjusted based on the face region. This optimizes the sound reception for the target person, ensuring the clarity and quality of the sound source signal. By performing facial feature analysis on the target person image, the face region of the target person and its specific location information in the real scene are accurately located. Based on the location information of the face region, the phase adjustment parameters and amplitude adjustment parameters required by the receiving array are further calculated. Finally, by combining the phase adjustment parameters and amplitude adjustment parameters, the final beamforming parameters are generated and applied to the receiving array to achieve efficient sound source capture and high-quality audio signal acquisition. This enables accurate recognition of voice commands issued by specific users, improving recognition accuracy.
[0014] Secondly, this application provides a voice command execution method, comprising: in response to the voice interaction function corresponding to the sound source location of the target person being enabled, acquiring the target person; determining the target area location of the target person in the real scene based on the target person's image; if the target area location is different from the recorded location of the target person, adjusting the beamforming parameters of the receiving array corresponding to the sound source location based on the target area location; controlling the receiving array to call the beamforming parameters to acquire the target person's target sound signal, combining the target person's historical command text and the target sound signal to generate the target person's target command, and executing the target command.
[0015] In the above technical solution, after the voice interaction function corresponding to the sound source location of the target person is activated, the image of the target person is captured, and then image recognition technology is used to determine the specific location of the target person in the real scene. If the target person is detected to have moved to a new area, the beamforming parameters of the receiving array are automatically adjusted to ensure that the sound signal after the target person's position change can be acquired. By combining the target person's historical command text and the real-time captured target sound signal, the target command is generated and executed, thereby improving the execution efficiency of the voice interaction function.
[0016] Thirdly, this application provides a voice command execution device, comprising:
[0017] The sound source location acquisition unit is used to acquire the sound source location corresponding to the target sound signal in the real scene;
[0018] Personnel image acquisition unit, used to acquire target personnel located at the sound source location;
[0019] The instruction generation unit is used to generate the target instruction for the target person based on the position matching relationship between the sound source location and the recorded location of the target person, combined with the target person's historical instruction text and target sound signal, and then execute the target instruction.
[0020] Fourthly, this application provides a vehicle, the vehicle including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, it implements the voice command execution method as described above.
[0021] Fifthly, this application provides a computer-readable storage medium storing a computer program, which, when executed, implements the voice command execution method as described above. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a system architecture diagram of a voice command execution method provided in an embodiment of this application.
[0024] Figure 2 This is a flowchart illustrating a voice command execution method provided in an embodiment of this application;
[0025] Figure 3This is a schematic diagram of a scenario for a voice command execution method provided in an embodiment of this application;
[0026] Figure 4 This is a flowchart illustrating a voice command execution method provided in an embodiment of this application;
[0027] Figure 5 This is a flowchart illustrating a voice command execution method provided in an embodiment of this application;
[0028] Figure 6 This is a flowchart illustrating a voice command execution method provided in an embodiment of this application;
[0029] Figure 7 This is a schematic diagram of a scenario for a voice command execution method provided in an embodiment of this application;
[0030] Figure 8 This is a flowchart illustrating a voice command execution method provided in an embodiment of this application;
[0031] Figure 9 This is a schematic diagram of the structure of a voice command execution device provided in an embodiment of this application;
[0032] Figure 10 This is a schematic diagram of the structure of a voice command execution device provided in an embodiment of this application;
[0033] Figure 11 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application. Detailed Implementation
[0034] To make the features and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0035] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0036] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0037] Among related technologies, traditional multi-zone voice control technology has significant limitations. Primarily, it can only target voice commands to specific locations or directions, resulting in limited coverage within the cabin and an inability to fully meet the diverse needs of users in the vehicle. Furthermore, this type of technology lacks the ability to continuously track the user's location, making it difficult to achieve seamless and continuous communication with constantly moving users. This significantly impacts the intelligence level and user experience of in-vehicle voice interaction systems.
[0038] To improve the user experience, this application provides a voice command execution method, wherein the execution subject of this method is a vehicle. Detailed descriptions are provided below. It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments. Please refer to... Figure 1 , Figure 1 This is a system architecture diagram of a voice command execution method provided in an embodiment of this application. The specific flow of the voice command execution method is as follows:
[0039] Figure 1 This is a system schematic diagram of a voice command execution method provided in an embodiment of this application.
[0040] It should be noted that the real-world scenarios in this solution include, but are not limited to, vehicle driving scenarios, meeting scenarios, training scenarios, and teaching scenarios. To clearly illustrate this solution, the following explanation will preferably use a vehicle driving scenario.
[0041] It should be noted that this solution includes, but is not limited to, image acquisition devices, audio acquisition devices, feature processing modules, instruction generation modules, and instruction execution modules. The image acquisition devices are used to acquire panoramic images of the real-world scene and images of individuals within the scene.
[0042] The type of audio acquisition device can be an omnidirectional microphone or a directional microphone, without specific limitations. The audio acquisition device is used to collect sound signals in real-world scenes, including human voices and ambient sounds. The number and placement of the audio acquisition devices can be determined according to actual needs. Multiple audio acquisition devices are arranged in a real-world scene according to array design principles (such as uniform distribution or specific geometry) and data is associated using wired or wireless connections to form a sound receiving array in the real-world scene.
[0043] The feature processing unit is used to extract audio features from the sound signals acquired by the audio acquisition device to determine the target sound signal from multiple sound signals, and to determine the sound source location features and audio spectrum of the target sound signal. The audio spectrum is used to distinguish the sound signals of each person. The feature processing unit is further used to extract the person image of each person and the regional position of the person image in the panoramic image of the scene. The feature processing module also generates a record location table corresponding to each person. The record location table is used to record the location information when the location information of the target person changes in the real scene.
[0044] The instruction generation module is used to translate the audio signal processed by the feature processing module into text to obtain audio text. After further text correction, the audio text is matched with a preset instruction set to determine the target instruction corresponding to the audio signal. At the same time, the instruction generation module is also used to adjust the parameters of the audio array in the real scene based on the human portrait image obtained from the feature processing module to generate beamforming parameters. The beamforming parameters are used to dynamically enhance the audio array, including suppressing interfering audio signals other than the target audio signal and enhancing the channel characteristics of the target audio signal.
[0045] The instruction generation module is used to execute the target instruction generated in the instruction generation module, and at the same time, display the response data after the target instruction is executed through the display device.
[0046] based on Figure 1 The system architecture diagram shown below will be used in conjunction with... Figures 2-6 This application provides a detailed description of a voice command execution method according to an embodiment.
[0047] Based on the above, this application proposes a voice command execution method. Please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart illustrating a voice command execution method provided in an embodiment of this application. Figure 2 As shown, the method in this application embodiment may include the following steps S101-S103.
[0048] S101, Obtain the location of the sound source corresponding to the target sound signal in the real scene.
[0049] In this embodiment of the application, in order to ensure that the user's privacy is not illegally obtained, the sound array composed of multiple audio acquisition devices is controlled to acquire sound signals in the real scene when the user is authorized. The sound signals include the sound signals of people in the real scene and the environmental sound signals. At the same time, the image acquisition device is controlled to acquire images of people in the real scene when authorized and save them to a preset image database.
[0050] Specifically, when the receiver array is detected to be in an authorized state, the receiver array's sound signal acquisition function is activated. At the same time, when the image acquisition device is detected to be in an authorized state, the image acquisition device's image capture function is activated to obtain a panoramic image of the scene in the real scene. Furthermore, the panoramic image of the scene is analyzed for human features to obtain the images of all the people in the real scene.
[0051] Audio feature analysis is performed on the sound signal to filter out ambient noise and interference audio. Feature vectors are then extracted from the filtered sound signal to extract feature vectors that can characterize the unique properties of the sound signal, such as Mel-frequency cepstral coefficients (MFCC) and linear predictive coding (LPC).
[0052] After extracting the feature vector of the sound signal, the feature vector is further mapped using a pre-trained sound processing model. The feature vector is then analyzed for pronunciation frequency to obtain the pronunciation features of the sound signal. These pronunciation features are then matched with a preset pronunciation feature library. The pronunciation feature library stores the audio features corresponding to each instruction keyword. If the pronunciation features of the sound signal match the audio features corresponding to a certain instruction keyword in the preset pronunciation feature library, then the sound signal is identified as the target sound signal. For example, in a vehicle driving scenario, the target sound signal refers to the instructions issued by the driver or passengers to control various functions of the vehicle, such as "turn on the interior lights," "start the navigation system," "adjust the air conditioning temperature," and "play relaxing music."
[0053] The image acquisition device acquires a preset number of images of each person at intervals. Based on the three-dimensional geometric information of the images acquired at multiple different times, a depth image is generated for each person. The depth image of each person is then registered with the panoramic image of the real scene to match the pixels of the depth image with the pixels of the panoramic image. The depth information is usually represented as the distance value of each pixel, reflecting the distance information of each pixel in the target person image from the image acquisition device.
[0054] By using the intrinsic and extrinsic parameter matrices of the image acquisition device, the pixel coordinates in the depth image are transformed into three-dimensional coordinates. The intrinsic parameter matrix primarily describes the internal characteristics of the image acquisition device, such as focal length, principal point coordinates, and distortion coefficients, while the extrinsic parameter matrix describes the position and orientation of the image acquisition device relative to the real-world scene coordinate system. A preset algorithm then projects the three-dimensional coordinates back onto the two-dimensional plane of the panoramic scene image, thereby obtaining the regional positions of each person's image within the real-world scene.
[0055] S102, Obtain the target personnel located at the sound source location.
[0056] Specifically, since the sound receiving array is set up in the real scene according to the preset array distribution pattern, the sound receiving array can calculate the arrival time of the sound signal and the sound propagation amplitude when the sound signal reaches each array area according to the time difference of arrival algorithm, thereby determining the sound source location of the sound signal. Therefore, by obtaining the area location of the personnel image of each person in the above S101 in the panoramic image of the scene, and by analyzing whether there is a personnel image area location in the sound source location that falls into the location information corresponding to the sound source location, the personnel image that falls into the sound source location is determined as the target personnel.
[0057] S103: Based on the position matching relationship between the sound source location and the recorded location of the target person, combined with the target person's historical instruction text and target sound signal, generate the target person's target instruction and execute the target instruction.
[0058] In this embodiment of the application, a location information record table is generated for each person based on the facial features of the person in the real scene. The location information record table records the historical location of each person and the instruction text information generated by each person when in the historical location.
[0059] It should be noted that the number of historical location areas in the location information record table can be single or multiple. When the number of historical location areas is single, it indicates that the target person's location has not changed in the real-world scenario. When the number of historical location areas is multiple, it indicates that the target person's location has changed in the real-world scenario. The recording time of the historical location area is determined according to the actual application scenario. For example, in a vehicle driving scenario, the recording starts when the vehicle is powered on and ends when it is powered off. In a meeting scenario, the recording starts when the host triggers the "start meeting" command and ends when the host triggers the "end meeting" command.
[0060] Specifically, at preset time intervals, the location of the target person's sound source at the current moment is matched with the most recently recorded historical area location in the location information record table, i.e., the recorded location, to determine whether the target person has moved. If the target person has moved, the historical command text corresponding to the historical area location recorded in the location information record table is contextually associated with the target sound text corresponding to the target sound signal at the current moment, thereby forming a complete information chain of target sound text. When associating the target sound text with the historical command text, all recorded historical command texts in the location information record table can be obtained, or a custom number of historical command texts can be obtained. The target sound text is then matched with a preset command set to determine the target command.
[0061] The target instruction is sent to the instruction execution module. The instruction execution module will output the response data after executing the target instruction through the output component. The output component may include an audio playback component and a display component. The response data may include voice data and text data.
[0062] Please refer to the following: Figure 3 , Figure 3 This is a schematic diagram of a scenario for a voice command execution method provided in an embodiment of this application, such as... Figure 3 As shown, taking a vehicle driving scenario as an example, there are multiple seats inside the vehicle cabin, and an audio acquisition device is installed in the area above each seat. By linking the audio acquisition devices in multiple seats, a sound array is formed inside the vehicle cabin. At the same time, an image acquisition device is installed inside the vehicle cabin. This image acquisition device is used to acquire a panoramic image of the vehicle cabin and to acquire the image of each person from the panoramic image of the cabin.
[0063] As shown above, by acquiring the target sound signal and images of individuals in the real-world scene, determining the location of the sound source corresponding to the target sound signal, and further identifying the target person located at that sound source location, precise matching between the sound source and the person can be achieved. By combining the positional matching relationship between the sound source location and the recorded location of the target person, as well as the target person's historical command text and the current target sound signal, the target person's target command can be accurately generated, thereby improving the accuracy and efficiency of command recognition.
[0064] Because real-world scenarios involve the movement and change of people, it is necessary to determine the location of each person within the real-world scene based on panoramic images to accommodate these changes. Please see [link / reference]. Figure 4 , Figure 4 This is a flowchart illustrating a voice command execution method provided in an embodiment of this application. Figure 4 As shown, the method in this application embodiment may include the following steps S201-S204.
[0065] S201, acquire a panoramic image of the real-world scene.
[0066] S202, Identify people images in a panoramic scene image.
[0067] Specifically, in S201-S202, a panoramic image of the real scene is acquired through an image acquisition device. A preset human image analysis algorithm is used to analyze and process the human images in the captured panoramic image. The preset human image analysis algorithm can identify and segment different regions in the image, especially human body contours and facial features. It also identifies the images of each person by combining a deep learning model and separates them from the panoramic image of the scene.
[0068] For example, taking a vehicle driving scenario, there are three people in the vehicle cabin, namely person A, person B, and person C. At this time, a panoramic image of the vehicle cabin is acquired through an image acquisition device. The panoramic image of the cabin includes person A, person B, and person C. Then, after performing facial feature detection on the panoramic image of the cabin through a preset facial analysis algorithm, the facial images of person A, person B, and person C are obtained respectively.
[0069] S203, Based on the position of the personnel images in the panoramic image of the scene, determine the regional position of each person in the real scene.
[0070] Specifically, multiple images of different individuals are acquired at preset time intervals, and depth images corresponding to each individual image are generated based on the three-dimensional geometric information of the images of different individuals.
[0071] The depth images corresponding to each person are registered with the panoramic images of the real scene to match the pixels in the depth images with the pixels in the panoramic images. The depth information is usually represented by the distance value of each pixel, reflecting the distance information of each pixel in the person image to the image acquisition device.
[0072] By using the intrinsic and extrinsic parameter matrices of the image acquisition device, the pixel coordinates in the depth image are transformed into three-dimensional coordinates. The intrinsic parameter matrix primarily describes the internal characteristics of the image acquisition device, such as focal length, principal point coordinates, and distortion coefficients, while the extrinsic parameter matrix describes the position and orientation of the image acquisition device relative to the real-world scene coordinate system. A preset algorithm then projects the three-dimensional coordinates back onto the two-dimensional plane of the panoramic scene image, thereby obtaining the regional positions of each person's image within the real-world scene.
[0073] S204. Based on the relationship between the location of the area and the location of the sound source, obtain the target person image located at the sound source location.
[0074] S205, Determine the target person located at the sound source location based on the target person image.
[0075] Specifically, in S204-S205, the position coordinates of each person image in the panoramic image of the scene are obtained, and it is detected whether there is a position coordinate region of a person image that overlaps with the coordinate region of the sound source. The person corresponding to the overlapping person image is determined as the target person image.
[0076] As shown above, by acquiring a panoramic image of the real-world scene and identifying the images of each person, comprehensive visual monitoring of the entire scene is achieved, which helps to understand the distribution of people within the scene in real time. Furthermore, based on the regional position of each person's image in the panoramic scene image, the system can accurately track the location of specific individuals, thereby improving operational efficiency.
[0077] Because the location information of each person in a real-world scenario changes dynamically, dynamic location tracking of each person is necessary to ensure the accuracy of the instructions issued by each person. Please see [link / reference]. Figure 5 , Figure 5 This is a flowchart illustrating a voice command execution method provided in an embodiment of this application. Figure 5 As shown, the method in this application embodiment may include the following steps S301-S307.
[0078] S301, based on the position matching relationship between the sound source location and the recorded location of the target person, combined with the target person's historical instruction text and target sound signal, generates the target person's target instruction and executes the target instruction.
[0079] For details on the specific implementation of S301, please refer to step S103 above, which will not be elaborated here.
[0080] S302, Obtain the facial features of the target person from the target person image, and obtain the recording position corresponding to the facial features.
[0081] Specifically, the system acquires the target person's image and performs facial feature analysis on the image to obtain the target person's facial features, including analyzing the size, distribution, and outline of the facial features. Based on the analyzed facial feature data, a unique location information record table for the target person is generated. Simultaneously with acquiring the facial features, the system obtains the regional position of the target person's image relative to the panoramic image of the scene and saves the regional position to the location record table.
[0082] For example, after performing facial feature analysis on the target person, a location information record table with the number 001 is generated. If the current location of the target person is the coordinate value A, then the coordinate value A is written into the location information record table.
[0083] S303, if the location of the sound source matches the location of the target person's recorded location, then the target person's target instruction is generated by combining the target person's historical instruction text and the target sound signal, and the target instruction is executed.
[0084] Specifically, the location of the sound source in the panoramic image of the target person at the current moment is determined. At the same time, the most recently recorded historical location in the location information record table corresponding to the target person is used as the recording location. If the sound source location at the current moment is the same as the recorded location, it is considered that the target person has not changed location. At the same time, the historical command text of the target person in the historical location is obtained from the location information record table. The historical command text is obtained by analyzing and recording the historical sound signals of the target person.
[0085] The target audio signal is parsed to generate target audio text. The target audio text is then context-dependently correlated with historical instruction text to obtain the target audio text. This target audio text is then matched against a preset instruction set to obtain the target instruction for the target person. This target instruction is sent to the instruction execution module, which outputs the response data after executing the instruction through an output component. This output component may include an audio playback component and a display component, and the response data may include both voice and text data.
[0086] S304 If the position matching relationship between the sound source location and the recorded position of the target person is not the same, the recorded position is updated using the sound source location, and the target person's target command is generated and executed by combining the target person's historical command text and target sound signal.
[0087] If the current location of the sound source is different from the recorded location, it is assumed that the target person's location has changed. At this time, the current location of the target person's sound source is written into the target person's corresponding location record table. At the same time, the historical command text recorded in the target person's historical location is obtained from the location information record table. The historical command text is the text data obtained after analyzing the target person's historical sound signals, and is recorded as historical command text.
[0088] The target audio signal is parsed to generate target audio text. The target audio text is then context-dependently correlated with historical command text to obtain the target audio text. This target audio text is then matched against a preset command set to obtain the target command for the target person. This target command is sent to the command execution module, which displays the response data after executing the command via a display device. The response data includes both audio and text data.
[0089] S305, based on the recorded location, retrieve the historical instruction text of the target personnel.
[0090] S306, acquire the target audio text of the target audio signal, and obtain the target instruction text based on the target audio text and the historical instruction text.
[0091] Specifically, in S305-S306, the most recent historical location of the target personnel is recorded in the location record table. At the same time, the historical command text of the target personnel when they were in the historical location is recorded in the location record table. That is, the historical command text corresponding to the target personnel when they were in the historical location is obtained by converting the sound signal generated by the target personnel when they were in the historical location into audio-to-text.
[0092] It should be noted that by analyzing keywords in the voice or text data input by the target personnel, it is possible to identify their intentions and determine the corresponding control command type. For example, when the target personnel's voice text is "check the weather," the keywords "check" and "weather" will be identified, thus determining that this is a query command and executing the relevant operation of checking the weather; while when the target voice text is "turn on the air conditioner," the keywords "turn on" and "air conditioner" will be identified, thus determining that this is a device control command and executing the relevant operation of turning on the air conditioner.
[0093] For example, the obtained historical command text is "Query Kaifeng City Weather", and the target audio text is "Query Gulou District Weather". The keyword in the target audio text is "Query, Weather", which indicates that the command type corresponding to the target audio text is a query command. At this time, the target command text obtained by contextually associating the historical command text with the target audio text is "Query Kaifeng City Gulou District Weather".
[0094] S307, perform instruction matching on the target instruction text to generate target instructions for the target personnel.
[0095] Specifically, the target instruction text is matched with a preset instruction set to generate the target instruction for the target person. For example, the target instruction text "query weather in Gulou District, Kaifeng City" is subjected to keyword combination probability analysis. The keyword combination with the highest probability is "query weather". Therefore, the target instruction is a weather query instruction. At the same time, "Gulou District, Kaifeng City" is written as a query condition for the weather query instruction to ensure that the correct regional weather can be queried.
[0096] As shown above, by employing facial recognition technology to capture the facial features of the target person in the image and obtaining the corresponding recording location, the accuracy and efficiency of recognition are improved. When the sound source location is completely consistent with the recording location, the target command is generated by associating the target person's historical command text with the target sound signal, thereby achieving high precision in triggering control commands. By determining the target person's location information in real time, it is confirmed that the target person's location change information can be obtained. Based on this, the target person's historical command text and the target sound signal are semantically associated to generate the target command, thereby enhancing the targeting and effectiveness of the vehicle's task execution and providing users with a convenient and safe driving experience.
[0097] To enhance the driver's vehicle usage experience. Please see [link / reference]. Figure 6 , Figure 6 This is a flowchart illustrating a voice command execution method provided in an embodiment of this application. Figure 6 As shown, the method in this application embodiment may include the following steps S401-S404.
[0098] Please refer to the following: Figure 7 , Figure 7 This is a schematic diagram illustrating a scenario of a voice command execution method provided in an embodiment of this application. For example... Figure 7 As shown, the spatial coordinate range of the panoramic scene image is (300-300). Within this spatial coordinate range, the location coordinates of the sound source are (100,100) to (270-270). There are three personnel images: personnel image A, personnel image B, and personnel image C. By further matching the coordinate regions corresponding to the scene regions of the three personnel images with the coordinate regions of the sound source, it can be determined that personnel image A coincides with the sound source location, and the personnel corresponding to personnel image A is identified as the target personnel.
[0099] S401, acquire the target person image, determine the face region of the target person in the target person image, and adjust the beamforming parameters of the receiving array corresponding to the sound source position based on the face region.
[0100] Specifically, a pre-defined face detection algorithm is used to locate the face of the target person in the image, i.e., the face region. The beamforming parameters of the receiver array corresponding to the sound source location are then adjusted. Beamforming is a method of amplifying a sound signal in a specific direction by controlling the signal phase of multiple receiver arrays. In this scenario, the beamforming parameters of the receiver array can be adjusted in real time according to the positional changes of the target person's face region to ensure the directivity and focus of the beamforming.
[0101] S402, Perform facial feature analysis on the target person image to determine the target person's face region and the location information of the face region in the real scene.
[0102] Specifically, a preset face detection algorithm is used to determine the facial region of the target person from the target person image, and then the coordinates of the facial region in the real scene are obtained.
[0103] S403 determines the phase adjustment parameters and amplitude adjustment parameters of the radio array based on the position information.
[0104] S404 generates beamforming parameters based on phase adjustment parameters and amplitude adjustment parameters.
[0105] Specifically, facial recognition technology is used to identify the face region of the target person in an image, and the two-dimensional or three-dimensional coordinates of the face region are obtained to determine its location. Then, based on the location information of the face region, the phase adjustment parameters and amplitude adjustment parameters of each microphone in the sound array are calculated. The phase adjustment parameters control the relative time delay of the sound signals received by each audio acquisition device. By adjusting the phase parameters, the target sound signal from the direction of the sound source can be strengthened in the time domain, while noise and interference from other directions are weakened. The amplitude adjustment parameters are used to adjust the intensity of the sound signals received by each audio acquisition device to further optimize the beamforming effect.
[0106] After determining the phase adjustment and amplitude adjustment parameters, the beamforming parameters are adjusted. These beamforming parameters will be used to control the operating status of each audio acquisition device in the receiver array, such as adjusting the power of audio acquisition devices in specific areas to enhance the sound signal at the sound source location.
[0107] For example, suppose a linear microphone array consists of four audio acquisition devices, with the target sound signal located directly in front of the array. To form a beam pointing towards the target sound signal, the phase and amplitude of each audio acquisition device need to be adjusted. Assuming the target sound signal is located in the array region at 0 degrees, the phase of each audio acquisition device can be adjusted so that sound waves from the 0-degree direction are superimposed in phase in the time domain, thus enhancing the sound signal in that direction. By adjusting the gain of each audio acquisition device, we can control the degree of beam focusing. For example, audio acquisition devices farther from the target sound signal can have their gain reduced to decrease sidelobe effects, thereby reducing interference from sound signals at non-source locations.
[0108] As shown above, by acquiring the scene area location of each person's image and matching the scene area location with the sound source location, the target person image containing the target person can be accurately determined. Based on this, the target person's face region is further determined, and the beamforming parameters of the receiving array corresponding to the sound source location are adjusted based on the face region. This optimizes the sound reception for the target person, ensuring the clarity and quality of the sound source signal. By performing facial feature analysis on the target person image, the face region of the target person and its specific location information in the real scene are accurately located. Based on the location information of the face region, the phase adjustment parameters and amplitude adjustment parameters required by the receiving array are further calculated. Finally, by combining the phase adjustment parameters and amplitude adjustment parameters, the final beamforming parameters are generated and applied to the receiving array to achieve efficient sound source capture and high-quality audio signal acquisition. This enables accurate recognition of voice commands issued by specific users, improving recognition accuracy.
[0109] based on Figure 1 The system architecture diagram shown below will be used in conjunction with... Figure 8 This application provides a detailed description of a voice command execution method according to an embodiment.
[0110] To improve the continuity of user experience when using voice interaction functions in real-world scenarios, this application proposes a voice command execution method. Please refer to... Figure 8 , Figure 8 This is a flowchart illustrating a voice command execution method provided in an embodiment of this application. Figure 8 As shown, the method in this application embodiment may include the following steps S501-S504.
[0111] S501, in response to the voice interaction function corresponding to the sound source location of the target person being turned on, acquire the target person image.
[0112] In this embodiment, the target person can be any person in a real-world scenario, and the number of target persons can be single or multiple, without specific limitations. The voice interaction function is activated when the target person is in the area corresponding to the sound source location by triggering a keyword to enable the voice interaction function. The specific trigger keyword is determined according to the actual scenario and is not specifically limited here.
[0113] Specifically, when the target person is located in the area corresponding to the sound source location, the trigger keyword is issued verbally. For example, the trigger keyword is "turn on voice function". At this time, the voice interaction function will be turned on, and after the voice interaction function is turned on, the image acquisition device is controlled to acquire the image of the target person whose sound source location is aligned with the target person.
[0114] S502, Determine the location of the target person in the target area in the real scene based on the target person image.
[0115] Specifically, a panoramic image of the real scene is acquired through an image acquisition device, and multiple images of the target person in the panoramic image are acquired at preset time intervals. Based on the three-dimensional geometric information of the multiple target person images, a depth image corresponding to the target person image is generated.
[0116] The depth image corresponding to the target person is registered with the panoramic image of the real scene to match the pixels in the depth image with the pixels in the panoramic image. The depth information is usually represented by the distance value of each pixel, reflecting the distance information of each pixel in the target person image to the image acquisition device.
[0117] By using the intrinsic and extrinsic parameter matrices of the image acquisition device, the pixel coordinates in the depth image are transformed into three-dimensional coordinates. The intrinsic parameter matrix mainly describes the internal characteristics of the image acquisition device, such as focal length, principal point coordinates, and distortion coefficients, while the extrinsic parameter matrix describes the position and orientation of the image acquisition device relative to the coordinate system of the real scene. A preset algorithm then projects the three-dimensional coordinates back onto the two-dimensional plane of the panoramic image of the scene, thereby obtaining the target area location of the person in the real scene.
[0118] S503, if the location of the target area is different from the recording location of the target personnel, the beamforming parameters of the receiving array corresponding to the sound source location are adjusted based on the location of the target area.
[0119] Specifically, if the target area location is different from the recorded location, it is assumed that the target person's current location has changed. At this time, the target area location is written into the target person's corresponding location record table. That is, the timestamp of the target area location acquisition is used as the storage time point and written into the location record table together with the target area location. Among them, the target area location with the timestamp closest to the current time is the most recent location information of the target person.
[0120] The system acquires an image of the target person corresponding to the target area location, and uses a preset face detection algorithm to locate the face region within the image. It then adjusts the beamforming parameters of the receiving array corresponding to the sound source location so that the array can receive the target sound signal emitted when the person is in the target area. Beamforming is a method of amplifying a sound signal in a specific direction by controlling the signal phase of multiple receiving arrays. In this scenario, the beamforming parameters of the receiving array can be adjusted in real time according to the positional changes of the target person's face region to ensure the directivity and focus of the beamforming.
[0121] It should be noted that the location information record table is a location information record table generated for each person based on the facial features of the person in the real scene. The location information record table records the historical location of each person and the historical command text generated by each person when in the historical location. The number of historical locations can be single or multiple. When the number of historical locations is single, it indicates that the target person has not changed location. When the number of historical locations is multiple, it indicates that the target person has changed location.
[0122] Optionally, if the target area location is the same as the recorded location of the target person, the historical instruction text generated by the target person when in the historical area location is obtained from the location record table corresponding to the target person. The target sound text is generated by context association between the historical instruction text and the target sound text corresponding to the target sound signal generated by the target person when in the target area location. The target sound text is then matched with a preset instruction set to obtain the target instruction of the target person.
[0123] Optionally, the target user can select the method to turn off the voice interaction function according to actual needs. For example, in the case of a vehicle in motion, when the vehicle switches from an active state to a power-off state, it means that the voice interaction function is turned off. Or, the target user can trigger a voice command to turn off the voice interaction function, such as issuing the voice command "turn off voice interaction function". The specific method of turning off the voice interaction function will be determined according to the actual scenario and will not be specifically limited here.
[0124] S504 controls the radio array to call beamforming parameters to obtain the target person's voice signal, combines the target person's historical command text and the target voice signal to generate the target person's target command, and executes the target command.
[0125] Specifically, the target sound signal emitted by the target person when at the target area location is acquired through a radio array. At this time, since the target person's location has changed, the target area location of the target person is written into the target person's corresponding location record table. At the same time, the historical command text recorded by the target person's historical area location is obtained from the location information record table. The historical command text is text data obtained after analyzing the target person's historical sound signal.
[0126] The target audio signal is parsed to generate target audio text. The target audio text is then context-dependently correlated with historical instruction text to obtain the target audio text. This target audio text is then matched against a preset instruction set to obtain the target instruction for the target person. This target instruction is sent to the instruction execution module, which outputs the response data after executing the instruction through an output component. This output component may include an audio playback component and a display component, and the response data may include both voice and text data.
[0127] As shown above, by capturing the target person's image after the voice interaction function corresponding to the sound source location of the target person is activated, image recognition technology is used to determine the target person's specific location in the real scene. If the target person is detected to have moved to a new area, the beamforming parameters of the receiver array are automatically adjusted to ensure that the sound signal after the target person's location change can be acquired. By combining the target person's historical command text with the real-time captured target sound signal, the target command is generated and executed, thereby improving the execution efficiency of the voice interaction function.
[0128] based on Figure 1 The system architecture will be discussed below. Figure 9 This application provides a detailed description of the voice command execution device provided in its embodiments. It should be noted that... Figure 9 The voice command execution device in the present application is used to execute the present application. Figures 2-6 The methods shown in the embodiments are for illustrative purposes only, illustrating the parts relevant to the embodiments of this application. For specific technical details not disclosed, please refer to this application. Figures 2-6 In the embodiment shown, the voice command execution device 500 may include a sound source location acquisition unit 501, a personnel image acquisition unit 502, and a command generation unit 503, as detailed below:
[0129] The sound source location acquisition unit 501 is used to acquire the sound source location corresponding to the target sound signal in the real scene;
[0130] Personnel image acquisition unit 502 is used to acquire target personnel images of target personnel located at the sound source location;
[0131] The instruction generation unit 503 is used to generate the target instruction of the target person based on the position matching relationship between the sound source position and the recorded position of the target person, combined with the historical instruction text of the target person and the target sound signal, and to execute the target instruction.
[0132] In some embodiments, the sound source location acquisition unit 501 further includes a panoramic image acquisition unit, a recognition unit, a location determination unit, and a location relationship matching unit.
[0133] A panoramic image acquisition unit is used to acquire panoramic images of real-world scenes.
[0134] The recognition unit is used to identify people images in a panoramic scene image;
[0135] The location determination unit is used to determine the regional location of each person in the real scene based on the position of the person's image in the panoramic image of the scene;
[0136] The location relationship matching unit is used to obtain the target person image located at the sound source location based on the relationship between the area location and the sound source location.
[0137] In some embodiments, the instruction generation unit 503 further includes a human face feature acquisition unit, a first determination unit, and a second determination unit.
[0138] The facial feature acquisition unit is used to acquire the facial features of the target person in the target person image and to acquire the recording position corresponding to the facial features;
[0139] The first determination unit is used to generate the target personnel's target command and execute the target command if the position matching relationship between the sound source position and the recorded position of the target personnel is the same.
[0140] The second determination unit is used to update the recorded position by using the sound source position if the position matching relationship between the sound source position and the recorded position of the target person is not the same, and to generate the target person's target instruction by combining the target person's historical instruction text and target sound signal, and then execute the target instruction.
[0141] In some embodiments, the personnel image acquisition unit 502 further includes an instruction text acquisition unit, a target instruction text determination unit, and a target instruction generation unit.
[0142] The instruction text acquisition unit is used to acquire the historical instruction text of the target personnel based on the recorded location;
[0143] The target instruction text determination unit is used to acquire the target sound text of the target sound signal, and obtain the target instruction text based on the target sound text and historical instruction text;
[0144] The target instruction generation unit is used to perform instruction matching on the target instruction text to generate target instructions for the target personnel.
[0145] In some embodiments, the personnel image acquisition unit 502 further includes an adjustment unit.
[0146] The first adjustment unit is used to acquire the target person's image, determine the target person's face region in the target person's image, and adjust the beamforming parameters of the receiving array corresponding to the sound source location based on the face region.
[0147] In some embodiments, the personnel image acquisition unit 502 further includes a feature analysis unit, a parameter determination unit, and a generation unit.
[0148] The feature analysis unit is used to perform facial feature analysis on the target person image to determine the target person's face region and the location information of the face region in the real scene;
[0149] The parameter determination unit is used to determine the phase adjustment parameters and amplitude adjustment parameters of the radio array based on the position information;
[0150] The generation unit is used to generate beamforming parameters based on phase adjustment parameters and amplitude adjustment parameters.
[0151] In this embodiment, by acquiring target sound signals and images of individuals in a real-world scene, determining the location of the sound source corresponding to the target sound signal, and further acquiring the target individual located at that sound source location, precise matching between the sound source and the individual is achieved. By combining the positional matching relationship between the sound source location and the recorded location of the target individual, as well as the target individual's historical command text and the current target sound signal, the target individual's target command can be accurately generated, thereby improving the accuracy and efficiency of command recognition.
[0152] based on Figure 1 The system architecture will be discussed below. Figure 10 This application provides a detailed description of the voice command execution device provided in its embodiments. It should be noted that... Figure 10 The voice command execution device in the present application is used to execute the present application. Figure 8 The methods shown in the embodiments are for illustrative purposes only, illustrating the parts relevant to the embodiments of this application. For specific technical details not disclosed, please refer to this application. Figure 8 In the embodiment shown, the voice command execution device 600 may include a response unit 601, a region location acquisition unit 602, a parameter adjustment unit 603, and a command execution unit 604, as detailed below:
[0153] The response unit 601 is used to acquire the target person in response to the voice interaction function corresponding to the sound source location of the target person being turned on.
[0154] The region location acquisition unit 602 is used to determine the target region location of the target person in a real scene based on the target person image;
[0155] The parameter adjustment unit 603 is used to adjust the beamforming parameters of the receiver array corresponding to the sound source location based on the target area location if the target area location is different from the recording location of the target personnel.
[0156] The instruction execution unit 604 is used to control the radio array to call beamforming parameters to obtain the target person's target voice signal, combine the target person's historical instruction text and the target voice signal to generate the target person's target instruction, and execute the target instruction.
[0157] In this embodiment, after the voice interaction function corresponding to the sound source location of the target person is activated, an image of the target person is captured, and then image recognition technology is used to determine the specific location of the target person in the real scene. If the target person is detected to have moved to a new area, the beamforming parameters of the receiving array are automatically adjusted to ensure that the sound signal after the target person's position change can be acquired. By combining the target person's historical command text and the real-time captured target sound signal, the target command is generated and executed, thereby improving the execution efficiency of the voice interaction function.
[0158] Furthermore, the voice command execution device provided in the above embodiments and the voice command execution method embodiment belong to the same concept, and the implementation process can be found in the method embodiment, which will not be repeated here.
[0159] The sequence numbers of the embodiments described above are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0160] Please see Figure 11 This provides a structural schematic diagram of a vehicle according to an embodiment of this application. Figure 11 As shown, the vehicle 700 includes a processor 701 and a memory 702. The processor 701 and the memory 702 are electrically connected.
[0161] The processor 701 is the control center of the vehicle 700 and may include one or more processing cores. The processor 701 connects to various parts of the vehicle via various interfaces and lines, executing various vehicle functions and processing data by running or calling computer programs stored in the memory 702 and calling data stored in the memory 702, thereby performing overall vehicle control. Optionally, the processor 701 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 701 may integrate one or more of the following: CPU, Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user page, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 701 and may be implemented separately through a communication chip.
[0162] The memory 702 can be used to store software programs and modules. The processor 701 executes various functional applications and data processing by running the computer programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, computer programs required for at least one function, etc.; the data storage area may store data created based on the use of the vehicle, etc.
[0163] Furthermore, memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory 702 may also include a memory controller to provide processor 701 with access to memory 702.
[0164] In this embodiment, the processor 701 in the vehicle 700 loads the instructions corresponding to the processes of one or more computer programs into the memory 702 according to the following steps, and the processor 701 runs the computer programs stored in the memory 702 to realize various functions, as follows:
[0165] Obtain the location of the sound source corresponding to the target sound signal in the real scene; obtain the target person located at the sound source location; based on the location matching relationship between the sound source location and the recorded location of the target person, combined with the target person's historical instruction text and the target sound signal, generate the target person's target instruction and execute the target instruction.
[0166] Optionally, the processor 701, in acquiring the target person located at the sound source location, specifically performs the following: acquiring a panoramic image of the real scene; identifying the person image in the panoramic image; determining the regional position of each person in the real scene based on the position of the person image in the panoramic image; obtaining the target person image of the target person located at the sound source location based on the relationship between the regional position and the sound source location; and determining the target person located at the sound source location based on the target person image.
[0167] Optionally, the processor 701, in executing the position matching relationship between the sound source location and the recorded location of the target person, combines the target person's historical instruction text and the target sound signal to generate the target person's target instruction and execute the target instruction. Specifically, it performs the following: obtaining the target person's facial features from the target person's image and obtaining the recorded location corresponding to the facial features; if the position matching relationship between the sound source location and the target person's recorded location is the same, then combining the target person's historical instruction text and the target sound signal to generate the target person's target instruction and execute the target instruction; if the position matching relationship between the sound source location and the target person's recorded location is different, then using the sound source location to update the recorded location, and combining the target person's historical instruction text and the target sound signal to generate the target person's target instruction and execute the target instruction.
[0168] Optionally, the processor 701, in executing the process of combining the target person's historical instruction text and target voice signal to generate the target person's target instruction, specifically performs the following steps: based on the recorded location, obtain the target person's historical instruction text; obtain the target voice text from the target voice signal; based on the target voice text and the historical instruction text, obtain the target instruction text; and perform instruction matching on the target instruction text to generate the target person's target instruction.
[0169] Optionally, after the processor 701 executes the process of acquiring the target person image located at the sound source location, it specifically performs the following: acquiring the target person image, determining the face region of the target person in the target person image, and adjusting the beamforming parameters of the receiving array corresponding to the sound source location based on the face region.
[0170] Optionally, the processor 701 performs the following steps: determining the face region in the target person image and adjusting the beamforming parameters of the receiving array corresponding to the sound source location based on the face region. Specifically, this involves: performing facial feature analysis on the target person image to determine the face region of the target person and its location information in the real scene; determining the phase adjustment parameters and amplitude adjustment parameters of the receiving array based on the location information; and generating beamforming parameters based on the phase adjustment parameters and amplitude adjustment parameters.
[0171] In this embodiment, by acquiring target sound signals and images of individuals in a real-world scene, comprehensive environmental perception is achieved, which helps improve the comprehensiveness and accuracy of command recognition. By determining the location of the sound source corresponding to the target sound signal and further acquiring the image of the target individual located at that sound source location, precise location of the sound source and the target individual is achieved. By combining the positional matching relationship between the sound source location and the recorded location of the target individual, as well as the target individual's historical command text and the current target sound signal, the target individual's target command can be accurately generated, thereby improving the accuracy and efficiency of command recognition.
[0172] Optionally, the processor 701 is also used to specifically perform: in response to the voice interaction function corresponding to the sound source location of the target person being enabled, acquiring the target person; determining the target area location of the target person in the real scene based on the target person image; if the target area location is different from the recorded location of the target person, adjusting the beamforming parameters of the receiving array corresponding to the sound source location based on the target area location; controlling the receiving array to call the beamforming parameters to acquire the target person's target sound signal, combining the target person's historical instruction text and the target sound signal to generate the target person's target instruction, and executing the target instruction.
[0173] In this embodiment, after the voice interaction function corresponding to the sound source location of the target person is activated, an image of the target person is captured, and then image recognition technology is used to determine the specific location of the target person in the real scene. If the target person is detected to have moved to a new area, the beamforming parameters of the receiving array are automatically adjusted to ensure that the sound signal after the target person's position change can be acquired. By combining the target person's historical command text and the real-time captured target sound signal, the target command is generated and executed, thereby improving the execution efficiency of the voice interaction function.
[0174] In addition, the device provided in this application embodiment may specifically be a chip, component or module. The chip may include a connected processor and a memory. The memory is used to store instructions. When the processor calls and executes the instructions, the chip can execute a voice command execution method provided in the above embodiment.
[0175] This application also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the above-described related method steps to implement a voice command execution method provided in the above embodiments.
[0176] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement a voice command execution method provided in the above embodiments.
[0177] In this application, the apparatus, computer-readable storage medium, computer program product or chip provided in the embodiments are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0178] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0179] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the related couplings or direct couplings or communication connections shown or discussed may be through some interfaces; indirect couplings or communication connections between apparatuses or units may be electrical, mechanical, or other forms.
[0180] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for executing voice commands, characterized in that, include: Obtain the location of the sound source corresponding to the target sound signal in a real-world scene; Identify the target person located at the sound source location; Based on the location matching relationship between the sound source location and the recorded location of the target person, combined with the target person's historical instruction text and the target sound signal, the target person's target instruction is generated and executed.
2. The method according to claim 1, characterized in that, The acquisition of the target person located at the sound source location includes: Obtain a panoramic image of the real-world scene; Identify people images in the panoramic image of the scene; Based on the position of the personnel images in the panoramic image of the scene, determine the location of each person in the real scene. Based on the relationship between the location of the area and the location of the sound source, an image of the target person located at the location of the sound source is obtained; The target person is identified at the location of the sound source based on the image of the target person.
3. The method according to claim 1, wherein generating the target instruction of the target person based on the positional matching relationship between the sound source location and the recorded location of the target person, combined with the target person's historical instruction text and the target sound signal, and executing the target instruction, comprises: Obtain the facial features of the target person from the target person image, and obtain the recording location corresponding to the facial features; If the location of the sound source matches the location of the recorded position of the target person, then the target person's target instruction is generated by combining the target person's historical instruction text and the target sound signal, and the target instruction is executed. If the location of the sound source is not the same as the location of the recorded position of the target person, the recorded position is updated using the sound source location, and the target person's target instruction is generated by combining the target person's historical instruction text and the target sound signal, and the target instruction is executed.
4. The method according to claim 1, characterized in that, The step of combining the target person's historical command text and the target's voice signal to generate the target person's target command includes: Based on the recorded location, obtain the target person's historical instruction text; Obtain the target audio text of the target audio signal, and obtain the target instruction text based on the target audio text and the historical instruction text; The target instruction text is matched to generate the target instruction for the target person.
5. The method according to claim 1, after acquiring the target person image located at the sound source location, the method further includes: The facial region of the target person is determined in the image of the target person, and the beamforming parameters of the receiving array corresponding to the sound source location are adjusted based on the facial region.
6. The method according to claim 5, characterized in that, The step of determining a face region in the target person image and adjusting the beamforming parameters of the receiving array corresponding to the sound source location based on the face region includes: Facial feature analysis is performed on the target person image to determine the face region of the target person and the location information of the face region in the real scene; Based on the location information, the phase adjustment parameters and amplitude adjustment parameters of the radio array are determined; The beamforming parameters are generated based on the phase adjustment parameters and the amplitude adjustment parameters.
7. A method for executing voice commands, characterized in that, The method includes: In response to the voice interaction function corresponding to the sound source location of the target person being enabled, the target person is obtained; Based on the image of the target person, determine the location of the target area in the real-world scene; If the location of the target area is different from the recording location of the target person, the beamforming parameters of the receiving array corresponding to the sound source location are adjusted based on the location of the target area. The receiver array is controlled to call the beamforming parameters to obtain the target person's voice signal. The target person's historical command text and the target voice signal are combined to generate the target person's target command and execute the target command.
8. A voice command execution device, characterized in that, include: The sound source location acquisition unit is used to acquire the sound source location corresponding to the target sound signal in the real scene; A personnel image acquisition unit is used to acquire a target personnel image of a target personnel located at the sound source location; The instruction generation unit is used to generate the target instruction for the target person based on the position matching relationship between the sound source location and the recorded location of the target person, combined with the target person's historical instruction text and target sound signal, and then execute the target instruction.
9. A vehicle, characterized in that, The vehicles include: Memory, used to store executable program code; A processor is configured to call and run the executable program code from the memory, causing the vehicle to perform the voice command execution method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the voice command execution method as described in any one of claims 1 to 7.