Vehicle-mounted screen-based control method and device for control, vehicle, and storage medium
By combining gesture recognition and voice control, the problem of users being unable to physically touch the screen is solved, resulting in a more efficient in-vehicle screen operation experience.
Patent Information
- Application Number
- PCT/CN2025/097808
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-30
- Filing Date
- 2025-05-28
- Publication Date
- 2025-12-04
AI Technical Summary
In existing technologies, the control of in-vehicle screens mainly relies on click and drag operations, which leads to a poor user experience in some situations (such as when rear-seat users cannot access the screen while the vehicle is in motion).
Gesture recognition technology determines whether the position of the gesture on the in-vehicle screen belongs to the target display area, and combines it with voice commands to control the corresponding controls, simplifying the complexity of voice commands.
It improves the convenience and user experience of in-vehicle screen control, especially in situations where it is inconvenient to touch the screen, simplifying operation complexity and improving human-computer interaction efficiency.
Smart Images

Figure CN2025097808_04122025_PF_FP_ABST
Abstract
Description
Control methods, devices, vehicles, and storage media based on in-vehicle screens.
[0001] This application claims Chinese Patent Application No. 2024106879105, filed on May 30, 2024, entitled "Control Method, Apparatus, Vehicle and Storage Medium Based on In-Vehicle Screen", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of vehicle technology, specifically to a control method, device, vehicle, and storage medium based on an in-vehicle screen. Background Technology
[0003] With the development of vehicle technology, there are more and more in-vehicle screens available in vehicles. In some embodiments, in addition to the traditional central control screen, some vehicles also provide passenger entertainment screens and rear entertainment screens, allowing users to control the vehicle or enjoy entertainment through these in-vehicle screens. Summary of the Invention
[0004] This application provides a control method, device, vehicle, and storage medium based on an in-vehicle screen. The technical solution is as follows:
[0005] On the one hand, a control method based on an in-vehicle screen is provided, the method comprising:
[0006] When the in-vehicle screen is activated in the vehicle and there is a gesture operation on the in-vehicle screen, it is determined whether the position indicated by the gesture operation on the in-vehicle screen belongs to the target display area, the target display area including at least one control displayed on the in-vehicle screen;
[0007] When the position indicated by the gesture operation on the vehicle screen belongs to the target display area, at least one preset instruction text corresponding to the target display area is loaded. The preset instruction text is instruction text used to control the controls within the target display area.
[0008] When a voice command is detected and the voice command is a target voice command, the control targeted by the voice command in the target display area is controlled based on the preset command text. The target voice command is a voice command that matches any of the at least one preset command text.
[0009] On the one hand, a control device based on an in-vehicle screen is provided, the device comprising:
[0010] The determination module is used to determine whether the position indicated by the gesture operation on the vehicle screen belongs to a target display area when the vehicle screen is started and there is a gesture operation on the vehicle screen. The target display area includes at least one control displayed on the vehicle screen.
[0011] A loading module is used to load at least one preset instruction text corresponding to the target display area when the position indicated by the gesture operation on the vehicle screen belongs to the target display area. The preset instruction text is instruction text for controlling the controls within the target display area.
[0012] A control module is configured to, when a voice command is detected and the voice command is a target voice command, control the control targeted by the voice command within the target display area based on the preset command text, wherein the target voice command is a voice command that matches any of the at least one preset command text.
[0013] On one hand, a vehicle is provided, the vehicle including one or more processors and one or more memories, the one or more memories storing at least one piece of program code, the program code being loaded and executed by the one or more processors to implement the operations performed by the control method based on the in-vehicle screen.
[0014] On one hand, a computer-readable storage medium is provided, wherein at least one piece of program code is stored in the computer-readable storage medium, the program code being loaded and executed by a processor to implement the operations performed by the control control method based on the vehicle screen. Attached Figure Description
[0015] Figure 1 is a schematic diagram of the implementation environment of a control method based on an in-vehicle screen provided in an embodiment of this application;
[0016] Figure 2 is a flowchart of a control method based on an in-vehicle screen provided in an embodiment of this application;
[0017] Figure 3 is a flowchart of another control method based on an in-vehicle screen provided in an embodiment of this application;
[0018] Figure 4 is a flowchart of another control method based on an in-vehicle screen provided in an embodiment of this application;
[0019] Figure 5 is a flowchart of another control method based on an in-vehicle screen provided in an embodiment of this application;
[0020] Figure 6 is a schematic diagram of the structure of a control device based on an in-vehicle screen provided in an embodiment of this application;
[0021] Figure 7 is a structural schematic diagram of a vehicle provided in an embodiment of this application. Detailed Implementation
[0022] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0023] In the following text, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features reflected. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0024] In order to illustrate the technical solutions provided in the embodiments of this application, the terms involved in the embodiments of this application will be introduced below.
[0025] In-vehicle screens: These are electronic screens installed inside a car to display information. These screens have a wide range of uses, from basic driving information display to multimedia entertainment, navigation, vehicle settings, and safety assistance systems.
[0026] Rear-seat entertainment screen: Also known as a rear-seat entertainment system, a rear-seat entertainment screen is an independent LCD display designed for rear passengers. It allows rear passengers to watch videos, movies, listen to music, and other multimedia content during the journey, thus enhancing the riding experience. Generally, rear-seat entertainment screens are quite feature-rich, including basic audio and video playback functions, as well as communication, office work, internet access, and even external game consoles. Some models also support independent control of the rear-seat LCD screen, allowing passengers to choose content according to their personal preferences for a personalized entertainment experience.
[0027] Rear ceiling-mounted screen: A rear ceiling-mounted screen is an entertainment device installed on the rear roof of a vehicle. It is typically used to provide audio-visual entertainment, interactive games, and other functions to enhance the riding experience for rear passengers. These devices usually support multiple input methods, such as USB (Universal Serial Bus) and HDMI (High Definition Multimedia Interface), and can be operated via touch screen or remote control.
[0028] Passenger entertainment screen: Designed to provide passengers with a diverse entertainment experience. It offers various functions such as movie playback, music listening, and e-book reading, enriching passengers' time on long journeys. Some passenger entertainment screens also support interconnection with other vehicle systems, such as navigation and voice recognition, allowing passengers to operate the entertainment screen and increasing the interactivity of the journey.
[0029] Speech recognition (SR) is the automated process of converting speech signals into text or commands. With the rapid development of deep learning technology, speech recognition has made significant progress, especially in accuracy and robustness.
[0030] Gesture recognition: Gesture recognition is a topic in computer science and language technology, aiming to recognize human gestures through mathematical algorithms.
[0031] Air gestures: a technology that controls electronic devices through non-contact hand gestures.
[0032] In related technologies, the control of in-vehicle screens usually relies on click and drag operations on the screen. In other words, users need to touch the screen to control it.
[0033] However, in some situations, it is inconvenient for users to touch the in-vehicle screen. In some embodiments, rear-seat users will fasten their seat belts while the vehicle is in motion to ensure safety. At this time, it is inconvenient for rear-seat users to touch the rear entertainment screen. Therefore, more diverse methods are needed to facilitate users to control the in-vehicle screen.
[0034] After introducing the terms involved in the embodiments of this application, the implementation environment of the embodiments of this application will be described below. Referring to Figure 1, the implementation environment of the control method based on the vehicle screen provided in the embodiments of this application includes the vehicle terminal 101, the vehicle screen 102, and the gesture recognition device 103.
[0035] The vehicle-mounted terminal 101 is a terminal installed in a vehicle to control different vehicle components. It is an intelligent device integrated into the vehicle, serving as a front-end facility for the vehicle monitoring and management system. Its core function is to collect, process, transmit, and interact with vehicle data, thereby supporting vehicle management, driving decisions, and intelligent services. The vehicle-mounted terminal includes a processor and memory; for example, it may be the vehicle's infotainment system. In this embodiment, the vehicle-mounted terminal 101 can perform gesture recognition based on the acquired data. Vehicle components are also referred to as functional components within the vehicle, such as in-vehicle screens, gesture recognition devices, in-vehicle air conditioning, and vehicle doors. In this embodiment, the vehicle-mounted terminal 101 can control the content displayed on the in-vehicle screen 102. In-vehicle applications can run on the vehicle-mounted terminal 101, and these applications are displayed on the in-vehicle screen 102.
[0036] The in-vehicle screen 102 includes a passenger entertainment screen and / or a rear entertainment screen. The passenger can use the passenger entertainment screen for entertainment, and the rear passengers can use the rear entertainment screen for entertainment.
[0037] The gesture recognition device 103 is connected to the vehicle terminal 101 and can be used to control the vehicle screen 102. The gesture recognition device 103 is a vehicle component used for gesture recognition. In this embodiment, the gesture recognition device 103 is mainly used to recognize air gestures. The gesture recognition device 103 has image acquisition capabilities. For example, the gesture recognition device 103 is a camera or image acquisition component inside the vehicle. The gesture recognition device 103 sends the acquired images inside the vehicle to the vehicle terminal 101, and the vehicle terminal 101 completes the final gesture recognition.
[0038] After introducing the implementation environment of the embodiments of this application, the technical solution provided by the embodiments of this application will be introduced below. Referring to Figure 2, taking the vehicle terminal as the execution subject as an example, the method includes the following steps.
[0039] 201. When the in-vehicle screen is activated in the vehicle and there is a gesture operation on the in-vehicle screen, the in-vehicle terminal determines whether the position indicated by the gesture operation on the in-vehicle screen belongs to the target display area, the target display area including at least one control displayed on the in-vehicle screen.
[0040] The in-vehicle screens include a passenger entertainment screen and / or a rear entertainment screen. The passenger entertainment screen is installed in front of the passenger seat, and the rear entertainment screen is installed in front of the rear seats. In some embodiments, it is installed behind the front seats or on the roof of the vehicle (also called a rear ceiling-mounted screen).
[0041] When the in-vehicle screen is activated, it sends a activation message to the in-vehicle terminal. Upon receiving this message, the in-vehicle terminal confirms that the screen is activated. Furthermore, after the screen is activated, the in-vehicle terminal maintains a communication connection with it. If this connection is interrupted, it indicates that the screen is off or malfunctioning. The presence of gesture controls for the screen means the user is controlling it via gestures. These gestures can be air gestures, meaning contactless operations where the user doesn't need to directly touch the device. Examples include waving, swiping, clenching a fist, giving a thumbs-up, and making an OK sign. Different gestures correspond to different screens; for instance, waving and swiping might correspond to the rear ceiling-mounted screen, while clenching a fist and giving a thumbs-up might correspond to the passenger-side screen.
[0042] The location or type of the gesture operation within the vehicle can be used to determine whether the gesture is targeted at the in-vehicle screen. In some embodiments, if the in-vehicle screen is activated and the gesture is executed in the corresponding in-vehicle space or of a preset type, the gesture is targeted at the in-vehicle screen. If the in-vehicle screen is activated but the gesture is not executed in the corresponding in-vehicle space or of a preset type, the gesture is not targeted at the in-vehicle screen. The type of gesture is used to distinguish different gestures; for example, waving, swiping, clenching a fist, giving a thumbs up, and making an OK sign are different types of gestures.
[0043] The target display area is a specific display area on the interface of the in-vehicle screen. The target display area includes at least one control; a control is a component used to create a graphical user interface (GUI), enabling the user to interact with the application. The user can control the software's behavior by clicking, dragging, or similar actions. Controls include sliders, buttons, etc., with sliders used to adjust volume, brightness, etc. This at least one control includes application launch controls or function control controls.
[0044] Determining whether the location indicated by the gesture on the in-vehicle screen belongs to the target display area is crucial for activating at least one control corresponding to that target display area. This allows for subsequent control of the at least one control based on voice commands. The location indicated by the gesture on the in-vehicle screen maps the execution location of the gesture inside the vehicle to the location on the in-vehicle screen. Activating a control triggers the injection of preset command text for that control, which can then be used to control the control.
[0045] 202. When the position indicated by the gesture operation on the vehicle screen belongs to the target display area, the vehicle terminal loads at least one preset instruction text corresponding to the target display area. The preset instruction text is an instruction text used to control the controls within the target display area.
[0046] Voice commands are used to control controls within the target display area to perform corresponding functions. In some embodiments, when the control is an application launch control, the target voice command can control the application launch control to launch the corresponding application. When the control is a function control control, the target voice command can control the function control control to perform the corresponding function. Loading at least one preset command text corresponding to the target display area refers to pre-loading the command text corresponding to the control within the target display area; this process is also called preset command text injection. In this embodiment, the preset command text is a short command text, which is either a verb or a noun, rather than a combination of both. That is, when the preset command text is a verb, it includes words like "open," "start," "uninstall," "raise," and "lower"; when the preset command text is a noun, it includes words like "number 1," "first," and "A." The injected preset command text differs for different display areas, allowing for personalized control of different display areas. In other words, when switching between different target display areas via gestures, the corresponding preset command text changes with the target display area.
[0047] 203. When a voice command is detected and the voice command is a target voice command, the vehicle terminal controls the control targeted by the voice command in the target display area based on the preset command text. The target voice command is a voice command that matches any of the preset command texts in the at least one preset command text.
[0048] In this context, since the target display area may include multiple controls, the voice command is used to control one of these controls, which is the control targeted by the voice command. When the target display area includes multiple controls, the voice command carries the name of the targeted control. Controlling the specified control can be achieved directly using the name, without needing to combine it with other content. For example, if the target display area includes launch controls for a social media application, an audio playback application, and a video playback application, and these are all application launch controls, the preset command text corresponding to the target display area can be A, B, and C. A indicates the launch control for the social media application, B indicates the launch control for the audio playback application, and C controls the launch control for the video playback application. In this case, the target voice command carries the name "A," meaning that the user only needs to say "A" for the social media application to be launched. In related technologies, if a user wants to launch a social media application using a voice command, they need to at least say "launch social media application," which involves both the action "launch" and the name "social media application," making it significantly more complex than the technical solution provided in this application. Furthermore, when the target display area includes a control, the voice command does not need to carry the noun of the target control. Control of the control can be achieved using the verb carried by the voice command. For example, if the target display area includes a volume control, the target voice command carries the noun of the specified action; that is, the user only needs to say "turn up" to increase the volume. In related technologies, if the user wants to adjust the volume using a voice command, they need to say at least "turn up the volume," which involves both the action "turn up" and the object "volume," making it significantly more complex than the "turn up" in this application. Also, if the target display area includes a brightness control used to control the screen brightness, the target voice command carries the noun of the specified action; that is, the user only needs to say "turn up" to increase the volume. It can be seen that the same target voice command can control both volume and brightness controls of different controls, which is completely impossible in related technologies.
[0049] As can be seen from the above description, the technical solution combining gesture operation and voice command provided in this application embodiment can greatly simplify the complexity of voice command and improve the efficiency of human-computer interaction while maintaining accuracy, thereby improving the user experience.
[0050] The technical solution provided in this application uses gesture operation to determine the target control, and combines it with concise words for voice control, which reduces the complexity of operation and improves the convenience of controlling the vehicle screen.
[0051] Referring to Figure 3, taking the vehicle-mounted terminal as the executing entity as an example, the method includes the following steps.
[0052] 301. When the in-vehicle screen is turned on in the vehicle, the in-vehicle terminal determines whether there is a gesture operation targeting the in-vehicle screen.
[0053] In some embodiments, when the in-vehicle screen is activated, the in-vehicle terminal determines whether there is a gesture operation in the vehicle interior space corresponding to the in-vehicle screen. If a gesture operation exists in the vehicle interior space corresponding to the in-vehicle screen, the in-vehicle terminal determines the type of the gesture operation. Based on the type of the gesture operation, the in-vehicle terminal determines whether the gesture operation is a gesture operation targeting the in-vehicle screen.
[0054] The in-vehicle space corresponding to the in-vehicle screen refers to the space inside the vehicle that can be controlled via gesture operation. The in-vehicle space corresponding to the in-vehicle screen is configured by technicians according to actual conditions, and this application embodiment does not limit this. The type of gesture operation is determined by the hand shape corresponding to the gesture operation.
[0055] In this implementation, it is determined whether there is a gesture operation in the vehicle interior space corresponding to the vehicle screen. If a gesture operation exists in the vehicle interior space, it is then determined whether the gesture operation is for the vehicle screen based on the type of gesture operation, which has a high accuracy.
[0056] To provide a clearer explanation of the above embodiments, the following description will be divided into three parts.
[0057] Part 1: When the in-vehicle screen is activated, the in-vehicle terminal determines whether there is a gesture operation in the in-vehicle space corresponding to the in-vehicle screen.
[0058] In some embodiments, when the in-vehicle screen is activated, the in-vehicle terminal acquires in-vehicle images captured by the gesture recognition device of the in-vehicle screen. The in-vehicle terminal performs target detection on the in-vehicle images to determine whether there is a gesture operation in the in-vehicle space corresponding to the in-vehicle screen.
[0059] The in-vehicle space covered by the in-vehicle image acquisition range of the gesture recognition device is the same as the in-vehicle space corresponding to the in-vehicle screen.
[0060] In some embodiments, when the in-vehicle screen is activated, the in-vehicle terminal acquires in-vehicle images captured by the gesture recognition device of the in-vehicle screen. The in-vehicle terminal inputs the in-vehicle image into a target detection model, and the target detection model extracts features from the in-vehicle image to obtain in-vehicle image features. Based on the in-vehicle image features, the in-vehicle terminal uses the target detection model to classify the in-vehicle image and determine whether there is a gesture operation in the in-vehicle image.
[0061] The object detection model is a binary classification model, outputting results indicating whether a gesture operation exists or not. A gesture operation refers to a specific hand gesture, such as waving, swiping, clenching a fist, giving a thumbs up, or making an OK sign; these are considered specific gestures and are therefore recognized as gesture operations. Conversely, gestures like making a heart shape with your hands, flipping your palm, clapping, or crossing your arms are not considered specific gestures and will not be recognized. When training this object detection model, positive and negative sample sets are used. Both sets include multiple in-vehicle images. Images in the positive sample set contain gesture operations (i.e., specific hand gestures). Images in the negative sample set do not contain gesture operations (i.e., no specific hand gestures or no passengers). By iteratively training the initial detection model using the positive and negative sample sets, the object detection model can be obtained. A contrastive loss function can be used as the loss function for training this initial model.
[0062] In some embodiments, when the in-vehicle screen is activated, the in-vehicle terminal acquires in-vehicle images captured by the gesture recognition device of the in-vehicle screen. The in-vehicle terminal inputs the in-vehicle image into a target detection model, and performs multiple convolutions on the in-vehicle image through the target detection model to obtain in-vehicle image features. The in-vehicle terminal then uses the target detection model to perform fully connected and normalized processing on the in-vehicle image features to obtain a classification value corresponding to the in-vehicle image. If the classification value is greater than or equal to a classification value threshold, it is determined that a gesture operation exists in the in-vehicle image; if the classification value is less than the classification value threshold, it is determined that no gesture operation exists in the in-vehicle image.
[0063] Part Two: When there is a gesture operation in the in-vehicle space corresponding to the in-vehicle screen, the in-vehicle terminal determines the type of the gesture operation.
[0064] In some embodiments, when there is a gesture operation in the in-vehicle space corresponding to the in-vehicle screen, the in-vehicle terminal determines the type of the gesture operation based on the image area corresponding to the gesture operation in the in-vehicle image.
[0065] The in-vehicle image was captured by the gesture recognition device on the vehicle's screen.
[0066] In some embodiments, when a gesture operation occurs in the vehicle interior space corresponding to the in-vehicle screen, the in-vehicle terminal performs key point detection on the image area to obtain multiple gesture key points. Based on these multiple gesture key points, the in-vehicle terminal determines the type of the gesture operation.
[0067] In some embodiments, when a gesture operation occurs in the vehicle interior space corresponding to the in-vehicle screen, the in-vehicle terminal inputs the image region into a keypoint detection model. The keypoint detection model extracts features from the image region to obtain its features. Based on these features, the in-vehicle terminal determines multiple gesture keypoints within the image region. The in-vehicle terminal then matches the shape formed by these multiple gesture keypoints with shapes corresponding to multiple candidate types to determine the type of the gesture operation.
[0068] The gesture operation type is used to distinguish different gesture operations. For example, waving, swiping, clenching a fist, giving a thumbs up, and making an OK sign are different types of gesture operations. Different gesture operations correspond to different in-vehicle screens. For example, waving and swiping may correspond to the rear ceiling screen, while clenching a fist and giving a thumbs up may correspond to the passenger screen.
[0069] Part Three: Based on the type of the gesture operation, the vehicle terminal determines whether the gesture operation is a gesture operation targeting the vehicle screen.
[0070] In some embodiments, if the type of the gesture operation is a preset type, the gesture operation is determined to be a gesture operation for the vehicle screen. If the type of the gesture operation is not the preset type, the gesture operation is determined to be a gesture operation for the vehicle screen.
[0071] The preset type is set by technicians according to the actual situation, and this application embodiment does not limit it. For example, if the recognized gesture operation type is waving, and the preset type is also waving, then the gesture operation is a gesture operation for the vehicle screen. If the recognized gesture operation type is clenching a fist, and the preset type is also waving, then the gesture operation is not a gesture operation for the vehicle screen.
[0072] This application also provides another implementation of step 301 described above.
[0073] In some embodiments, when the in-vehicle screen is activated, the in-vehicle terminal acquires an image of the vehicle's interior. Based on this image, the in-vehicle terminal determines whether there is a gesture operation targeting the in-vehicle screen.
[0074] The in-vehicle image is an image covering the interior space of the vehicle.
[0075] In this implementation, it is possible to determine whether there is a gesture operation on the in-vehicle screen directly using the in-vehicle image, which is highly efficient.
[0076] In some embodiments, when the in-vehicle screen is activated, the in-vehicle terminal acquires an in-vehicle image via an in-vehicle image acquisition device. The in-vehicle terminal inputs the in-vehicle image into a classification model, which extracts features from the image to obtain in-vehicle image features. Based on these features, the in-vehicle terminal uses the classification model to classify the in-vehicle image and determine whether there is a gesture operation targeting the in-vehicle screen.
[0077] The classification model is a binary classification model, and the output includes whether a gesture operation is performed on the in-vehicle screen or not. In some embodiments, the classification model is trained based on multiple sample in-vehicle images and the corresponding annotations for each sample in-vehicle image. The annotations for the sample in-vehicle images are used to indicate whether a gesture operation is performed on the in-vehicle screen in the sample in-vehicle images.
[0078] In some embodiments, when the in-vehicle screen is activated, the in-vehicle terminal acquires in-vehicle images captured by an in-vehicle image acquisition device. The in-vehicle terminal inputs these images into a classification model, which performs multiple convolutions on the images to obtain in-vehicle image features. The in-vehicle terminal then uses the classification model to perform fully connected and normalized processing on these features to obtain a classification value for the in-vehicle image. If the classification value is greater than or equal to a classification value threshold, it is determined that a gesture operation has been performed on the in-vehicle screen; if the classification value is less than the classification value threshold, it is determined that no gesture operation has been performed on the in-vehicle screen.
[0079] This application also provides another implementation of step 301 described above.
[0080] In some embodiments, when the in-vehicle screen is activated, the in-vehicle terminal determines whether there is a gesture operation in the vehicle interior space corresponding to the in-vehicle screen. If a gesture operation exists in the vehicle interior space corresponding to the in-vehicle screen, the in-vehicle terminal determines that a gesture operation exists for the in-vehicle screen. If no gesture operation exists in the vehicle interior space corresponding to the in-vehicle screen, the in-vehicle terminal determines that no gesture operation exists for the in-vehicle screen.
[0081] The method for determining whether a gesture operation exists belongs to the same inventive concept as the first implementation method. The implementation process is described in the relevant description of the first implementation method and will not be repeated here.
[0082] 302. In the event of a gesture operation on the vehicle screen, the vehicle terminal determines whether the location indicated by the gesture operation on the vehicle screen belongs to a target display area, the target display area including at least one control displayed on the vehicle screen.
[0083] The presence of gesture operations on the in-vehicle screen indicates that the user is controlling the screen through gestures. Referring to Figure 4, when the target display area is the in-vehicle application display area 400 on the desktop of the in-vehicle terminal, this in-vehicle application display area 400 includes multiple application launch controls 401-406. Similarly, referring to Figure 5, when the target display area is the application control area 500 on the application interface displayed on the in-vehicle terminal, this application control area 500 includes a function control control 501, which is used to implement a specific function.
[0084] In some embodiments, when a gesture operation is performed on the vehicle screen, the vehicle terminal determines the spatial location and direction of the gesture operation. Based on the spatial location and direction of the gesture operation, the vehicle terminal determines the location indicated by the gesture operation on the vehicle screen. Based on the screen coordinates of the location and the screen coordinates of the target display area, the vehicle terminal determines whether the location belongs to the target display area.
[0085] To provide a clearer explanation of the above embodiments, the following description will be divided into three parts.
[0086] Part 1: When there is a gesture operation on the vehicle screen, the vehicle terminal determines the spatial location and direction of the gesture operation.
[0087] In some embodiments, when a gesture operation is performed on the vehicle screen, the vehicle terminal acquires a three-dimensional image of the gesture operation. Based on the three-dimensional image, the vehicle terminal determines the spatial location and gesture posture of the gesture operation. Based on the gesture posture, the vehicle terminal determines the direction of the gesture operation.
[0088] The three-dimensional image of the gesture operation is obtained by fusing the in-vehicle image collected by the gesture recognition device with the in-vehicle point cloud. The fusion method can be front fusion or back fusion, and this application embodiment does not limit it.
[0089] In some embodiments, when a gesture operation is performed on the vehicle screen, the vehicle terminal acquires a three-dimensional image of the gesture operation. Based on the three-dimensional image, the vehicle terminal determines multiple spatial key points of the gesture operation. Based on the multiple spatial key points, the vehicle terminal determines the spatial position and gesture posture of the gesture operation. The vehicle terminal determines the direction corresponding to the gesture posture of the gesture operation as the pointing direction of the gesture operation.
[0090] In some embodiments, when a gesture operation is performed on the vehicle screen, the vehicle terminal acquires a 3D image of the gesture operation. The vehicle terminal performs keypoint detection on the 3D image to obtain a plurality of spatial keypoints. Based on the spatial coordinates of each spatial keypoint, the vehicle terminal determines the spatial position of the gesture operation. Based on the relative positional relationship of the plurality of spatial keypoints, the vehicle terminal determines the gesture posture of the gesture operation. The vehicle terminal determines the direction corresponding to the gesture posture of the gesture operation as the direction of the gesture operation.
[0091] Among them, the gesture posture includes the pitch angle and roll angle of the gesture action. The direction corresponding to the gesture posture is determined based on the pitch angle and roll angle. That is, the pitch angle and roll angle are used to determine the angle between the gesture operation and the horizontal plane. This angle can be regarded as the direction corresponding to the gesture posture.
[0092] The above process will be further explained below.
[0093] In some embodiments, when a gesture operation is performed on the vehicle screen, the vehicle terminal acquires a 3D image of the gesture operation. The vehicle terminal inputs this 3D image into a keypoint detection model, which performs keypoint detection on the 3D image to obtain multiple spatial keypoints. The vehicle terminal determines the spatial position of the gesture operation by averaging the spatial coordinates of each spatial keypoint. Based on the relative positional relationships of the multiple spatial keypoints, the vehicle terminal queries multiple candidate gesture postures to obtain the gesture posture of the gesture operation. The vehicle terminal determines the direction corresponding to the gesture posture of the gesture operation as the pointing direction of the gesture operation.
[0094] The keypoint detection model is trained based on multiple sample 3D images and the corresponding labeled spatial keypoints of each sample 3D image, and has the ability to determine spatial keypoints in 3D images. The relative positional relationships of multiple spatial keypoints corresponding to different candidate gesture poses are different. Therefore, by utilizing the relative positional relationships of multiple spatial keypoints, the gesture pose of the gesture operation can be found from the candidate gesture poses.
[0095] Part Two: The vehicle terminal determines the location indicated by the gesture operation on the vehicle screen based on the spatial location and direction of the gesture operation.
[0096] In some embodiments, the in-vehicle terminal determines the position transformation function corresponding to the gesture operation based on the direction of the gesture operation. The in-vehicle terminal substitutes the spatial position of the gesture operation into the position transformation function to obtain the position indicated by the gesture operation on the in-vehicle screen.
[0097] The position transformation function reflects the correspondence between the spatial position of the gesture operation and its position on the vehicle screen; that is, it is a function that transforms three-dimensional spatial coordinates into two-dimensional screen coordinates. This position transformation function is calibrated by technicians according to actual conditions, and this application embodiment does not limit it. Different directions correspond to different position transformation functions to accommodate different users pointing differently when performing the same gesture. The direction of the gesture operation refers to the direction of the middle finger, index finger, or thumb. When at least two of the middle finger, index finger, or thumb are simultaneously recognized, the direction of the gesture operation is determined according to the weight of the middle finger, index finger, or thumb. For example, if the weight of the middle finger is higher than that of the index finger, and the weight of the index finger is higher than that of the thumb, then as long as the direction of the middle finger is recognized, the direction of the middle finger is determined as the direction of the gesture operation. If the directions of the index finger and thumb are recognized, then the direction of the index finger is determined as the direction of the gesture operation. Of course, the above weight setting is only an example. In actual use, the weight can be set according to needs, and this application embodiment does not limit it. This position transformation function is used to convert three-dimensional spatial coordinates into two-dimensional screen coordinates. This process is actually a coordinate system transformation process. The position transformation function is equivalent to a coordinate system transformation function. For example, the above process involves the transformation from a spatial plane coordinate system to a camera coordinate system to a screen coordinate system. The direction of the gesture operation affects the result of the coordinate transformation. For example, if a user performs a leftward gesture and a rightward gesture at the same position, the corresponding position on the vehicle screen will be different. Combining the direction of the gesture operation to determine the position transformation function is to eliminate the influence of the direction on the position transformation. In other words, the position transformation function corresponding to the direction already carries the influence of the direction on the position transformation. This influence does not need to be determined in real time. It is only necessary to substitute the three-dimensional spatial coordinates of the gesture operation into the position transformation function to obtain the corresponding two-dimensional screen coordinates. This can reduce the amount of computation required for real-time calculation and achieve efficient coordinate transformation.
[0098] In this implementation, the position transformation function corresponding to the gesture operation is determined by the direction of the gesture operation, and the position transformation function is used to realize the conversion from the spatial position of the gesture operation to the indicated position on the vehicle screen, which is highly efficient.
[0099] In some embodiments, the vehicle terminal queries multiple candidate position transformation functions based on the direction of the gesture to obtain the position transformation function corresponding to the gesture. Different candidate position transformation functions correspond to different pointing ranges. The vehicle terminal substitutes the three-dimensional coordinates of the spatial position of the gesture into the position transformation function and performs a linear transformation on the three-dimensional coordinates through the position transformation function to obtain the position indicated by the gesture on the vehicle screen.
[0100] It should be noted that the influence of gesture pointing on coordinate system transformation stems from the difference in projection from three-dimensional space to two-dimensional screen. In other words, the spatial pointing of the gesture changes the spatial mapping relationship between the three-dimensional coordinate system and the screen plane.
[0101] Different directions determine different operation plane normal vectors. For example, a horizontal index finger establishes the XY operation plane, while a vertical downward finger establishes the XZ operation plane, resulting in different screen coordinates for the same spatial location projected onto different planes. Different directions correspond to different perspective projection parameters. For example, a straight forward finger uses orthographic projection, while an oblique finger uses perspective projection compensation, resulting in significant differences in their projection matrices. This difference is reflected in the differences in the aforementioned position transformation functions. The direction of the finger carries the user's operational intent. A rightward waving gesture maps to operations on the right side of the screen, while the same spatial location might map to controls at the top of the screen when pointing upward. Therefore, the same location may correspond to different screen positions depending on the direction of the finger.
[0102] Part Three: The vehicle terminal determines whether the location belongs to the target display area based on the screen coordinates of the location and the screen coordinates of the target display area.
[0103] The screen coordinates of the target display area include the screen coordinates of the boundary of the target display area.
[0104] In some embodiments, the vehicle terminal compares the screen coordinates of the location with the screen coordinates of the target display area to determine whether the location belongs to the target display area.
[0105] In some embodiments, the vehicle terminal compares the screen coordinates of the location on the vehicle screen with the screen coordinates of the target display area to determine whether the screen coordinates of the location fall within the area enclosed by the screen coordinates of the target display area. If the screen coordinates of the location fall within the area enclosed by the screen coordinates of the target display area, the vehicle terminal determines that the location belongs to the target display area of the vehicle screen. If the screen coordinates of the location do not fall within the area enclosed by the screen coordinates of the target display area, the vehicle terminal determines that the location does not belong to the target display area of the vehicle screen.
[0106] In addition to determining the position indicated by the gesture operation on the vehicle screen through the above-described embodiments, this application also provides the following method for determining the position indicated by the gesture operation on the vehicle screen.
[0107] In some embodiments, the in-vehicle terminal determines the palm region from the image region corresponding to the gesture operation in the in-vehicle image. The in-vehicle terminal determines the position indicated by the gesture operation on the in-vehicle screen based on the center point of the palm region.
[0108] In this implementation, gesture operations are performed by the user using their hands, which include the fingers and palm. The palm area is the region of the hand excluding the fingers. The center point of the palm area is located at the center of the palm. In this implementation, the position of the center point of the palm area is associated with the position indicated on the in-vehicle screen. If the center point of the palm area changes, the position indicated by the gesture on the in-vehicle screen will also change. For the user, adjusting the position indicated by the gesture on the in-vehicle screen is achieved by moving the palm. In this implementation, the gesture operation for the in-vehicle screen is the gesture of pointing to "5". This allows for direct recognition of the palm area, and the center point of the palm area is used to control the position indicated by the gesture on the in-vehicle screen.
[0109] To provide a clearer explanation of the above embodiments, the following description will be divided into several parts.
[0110] Part 1: The vehicle terminal determines the palm region from the image region corresponding to the gesture operation in the in-vehicle image.
[0111] In some embodiments, the vehicle terminal inputs the image region into a target detection model, and performs target detection on the image region through the target detection model to obtain the palm region in the image region.
[0112] The object detection model used is from related technologies such as the YOLO series, EfficientDet, RetinaNet, DETR, Faster R-CNN, and SSD, and this application does not limit the specific model used. The object detection model is trained using multiple sample images and their corresponding annotation information. The sample images are images of hand gestures, and the annotation information indicates the location of the palm region. After training, inputting the image into the object detection model directly yields the palm region.
[0113] Part Two: The vehicle terminal determines the position indicated by the gesture operation on the vehicle screen based on the center point of the palm area.
[0114] In some embodiments, the vehicle terminal maps the three-dimensional coordinates of the center point of the palm area onto the vehicle screen to obtain two-dimensional screen coordinates, which are used to represent the position indicated by the gesture operation on the vehicle screen.
[0115] In the above embodiment, the three-dimensional coordinates of the center point include coordinate values on three coordinate axes: x, y, and z. The plane formed by the x-axis and y-axis is parallel to the vehicle screen, and the y-coordinate represents the distance between the gesture operation and the vehicle screen. The process of mapping the three-dimensional coordinates to two-dimensional screen coordinates involves substituting x1 and y1 from the three-dimensional coordinates into a coordinate transformation function to directly obtain the corresponding x2 and y2. x2 and y2 are the two-dimensional screen coordinates corresponding to the three-dimensional coordinates. This is equivalent to "projecting" the palm along the y-axis onto the vehicle screen to obtain the two-dimensional screen coordinates. The form of the coordinate transformation function is (x2, y2) = (x2, y2)·(k1, k2). T Where k1 and k2 are coordinate transformation coefficients, which are set by technicians according to the actual situation, and this application embodiment does not limit them.
[0116] In addition, the vehicle-mounted terminal can also perform the following steps.
[0117] In some embodiments, after determining the location indicated by the gesture operation on the in-vehicle screen, the in-vehicle terminal displays a cursor at that location, thereby visualizing the location through the cursor and facilitating user control using gesture operations.
[0118] The cursor can be displayed in the form of a finger, an arrow, etc., and users can adjust the cursor display according to their preferences. This application embodiment does not limit this. When a gesture operation is detected, the cursor displayed on the vehicle terminal will also move with the gesture operation.
[0119] 303. When the position indicated by the gesture operation on the vehicle screen belongs to the target display area, the vehicle terminal loads at least one preset instruction text corresponding to the target display area, and the preset instruction text is instruction text used to control the controls within the target display area.
[0120] The preset instruction text is used to control the controls within the target display area.
[0121] In some embodiments, when the location indicated by the gesture on the vehicle screen belongs to the target display area, the vehicle terminal recalls instruction text based on controls within the target display area to obtain at least one preset instruction text corresponding to the target display area. The vehicle terminal then loads the at least one preset instruction text corresponding to the target display area.
[0122] The instruction text recall is to find the preset instruction text that corresponds to the target display area from multiple preset instruction texts. The preset instruction text corresponding to the target display area is the preset instruction text corresponding to the control within the target display area. The correspondence between the control and the preset instruction text is set by the technician according to the actual situation, and this application embodiment does not limit it.
[0123] To provide a clearer explanation of the above embodiments, the following description will be divided into two parts.
[0124] Part 1: The vehicle terminal recalls instruction text based on controls within the target display area to obtain at least one preset instruction text corresponding to the target display area.
[0125] In some embodiments, when the control in the target display area is an application launch control, the vehicle terminal recalls the application based on the application identifier corresponding to the control in the target display area to obtain at least one preset instruction text corresponding to the target display area. The at least one preset instruction text includes at least one of the following: application name, application name abbreviation, application number, and application alias.
[0126] Since the control is an application launch control, its function is to launch the application, and the control corresponds to an application identifier. The application name is the full name of the application indicated by the application identifier, which in some embodiments is "XX Music"; the application name abbreviation is a simplified version of the full name of the application, which in some embodiments is "XX"; the application number describes the position of the control corresponding to the application within the target display area, and can be 1 if there is only one control within the target display area; the application alias is another name for the application, which in some embodiments might be referred to as "YY" by some users, so "YY" is the application alias for "XX Music". It should be noted that the application name, application name abbreviation, application number, and application alias are configured in advance by technical personnel according to actual conditions, and this application embodiment does not limit this. Furthermore, the number of controls within the target display area can be one or more; for ease of explanation, the following example assumes therein is only one control within the target display area.
[0127] In some embodiments, when the control in the target display area is an application launch control, the vehicle terminal queries multiple preset instruction texts based on the application identifier corresponding to the control in the target display area, and determines at least one preset instruction text corresponding to the application identifier from the multiple preset instruction texts.
[0128] In some embodiments, when the control in the target display area is a function control control, the vehicle terminal recalls at least one preset instruction text corresponding to the target display area based on the function identifier corresponding to the control in the target display area. The function identifier is used to indicate the function provided by the corresponding control, and the at least one preset instruction text includes function control terms.
[0129] Among them, the function control words can be "increase" and "decrease", that is, there are only actions, but no corresponding verbs for the execution of actions.
[0130] In some embodiments, when the control in the target display area is a function control control, the vehicle terminal queries multiple preset instruction texts based on the function identifier corresponding to the control in the target display area, and determines at least one preset instruction text corresponding to the function identifier from the multiple preset instruction texts.
[0131] Part Two: The vehicle-mounted terminal loads at least one preset instruction text corresponding to the target display area.
[0132] Among them, the at least one preset instruction text is also referred to as a hot word, and the process of loading the at least one preset instruction text is also referred to as the hot word injection process.
[0133] In some embodiments, the in-vehicle terminal adds at least one preset instruction text to the cache for subsequent matching with voice commands.
[0134] 304. The vehicle terminal determines whether a voice command is present.
[0135] In some embodiments, the in-vehicle terminal performs voice activity detection on audio frames captured by the in-vehicle microphone to determine whether a voice command exists.
[0136] Among them, Voice Activity Detection (VAD) can identify speech signals in audio frames. These audio frames are acquired in real time.
[0137] In this implementation, by detecting voice activity in the audio frames captured by the in-vehicle microphone, it is possible to identify whether there are voice commands, with high accuracy.
[0138] In some embodiments, the in-vehicle terminal extracts features from audio frames captured by the in-vehicle microphone to obtain audio frame features. The in-vehicle terminal classifies the audio frames based on these features to determine whether each audio frame is a speech frame. If the audio frame is a speech frame, the terminal classifies the N consecutive audio frames captured after it to determine whether each of the N audio frames is a speech frame, where N is a positive integer. If a speech frame is present among the N audio frames, a voice command is determined to exist. If no speech frame is present among the N audio frames, a voice command is determined to exist. Alternatively, if the number of speech frames among the N audio frames is greater than or equal to M, a voice command is determined to exist. If the number of speech frames among the N audio frames is less than M, a voice command is determined to exist, where M is a positive integer less than N.
[0139] In this embodiment, N and M are set by technicians according to actual conditions, and this application does not limit them. In some embodiments, audio frame features include time-domain features and frequency-domain features. In some embodiments, audio frame features include energy fluctuations, zero-crossing rate, and spectral composition, etc. In the above embodiments, voice commands can be regarded as a set of multiple consecutive voice frames.
[0140] In some embodiments, the location of the in-vehicle microphone in the above embodiments is related to the type of in-vehicle screen. When the in-vehicle screen is a passenger-side entertainment screen, the in-vehicle microphone is located near the passenger seat. When the in-vehicle screen is a rear-seat entertainment screen, the in-vehicle microphone is located near the rear seats. This method can, to some extent, eliminate interference from voice commands from users not using the in-vehicle screen.
[0141] 305. In the presence of a voice command, the vehicle terminal determines whether the voice command is a target voice command, wherein the target voice command is a voice command that matches any of the at least one preset command texts.
[0142] Among them, a voice command that matches a preset command text means that the command text corresponding to the voice command matches the preset command text, and the matching includes being the same or similar.
[0143] In some embodiments, when a voice command is present, the vehicle terminal performs voice recognition on the voice command to obtain the command text corresponding to the voice command. Based on the command text and at least one preset command text corresponding to the target display area, the vehicle terminal determines whether the voice command is the target voice command.
[0144] In this implementation, the voice command can be determined to be the target voice command by using the command text corresponding to the voice command, which is highly efficient.
[0145] To provide a clearer explanation of the above embodiments, the following description will be divided into two parts.
[0146] Part 1: When a voice command is available, the vehicle terminal performs voice recognition on the voice command to obtain the corresponding command text.
[0147] In some embodiments, when a voice command is available, the in-vehicle terminal inputs the voice command into a voice recognition model, encodes the voice command using the voice recognition model, and obtains the voice command features. The in-vehicle terminal then decodes the voice command features using the voice recognition model to obtain the corresponding command text.
[0148] The speech recognition model has the ability to convert speech to text. It is trained under supervision based on multiple sample speech commands and the labeled command text corresponding to each sample speech command.
[0149] In some embodiments, when a voice command is present, the vehicle terminal inputs the voice command into a speech recognition model. The speech recognition model performs frame segmentation and windowing on the voice command to obtain multiple command audio frames. The vehicle terminal uses the speech recognition model to perform time-frequency transformation on the multiple command audio frames to obtain the spectrum of each command audio frame. The vehicle terminal uses the speech recognition model to embed and encode the spectrum of each command audio frame and the position of each command audio frame within the multiple command audio frames, obtaining audio frame embedding features and position embedding features for each command audio frame. The vehicle terminal uses the speech recognition model to concatenate the audio frame embedding features and position embedding features of each command audio frame to obtain the embedding features of each command audio frame. The vehicle terminal uses the speech recognition model to encode the embedding features of each command audio frame based on an attention mechanism, obtaining the audio frame features of each command audio frame. The vehicle terminal uses the speech recognition model to fuse the audio frame features of each command audio frame to obtain the speech command features of the voice command. The vehicle terminal uses a speech recognition model and an attention mechanism to perform multiple rounds of iterative decoding on the features of the voice command to obtain the corresponding command text.
[0150] Part Two: The vehicle-mounted terminal determines whether the voice command is the target voice command based on the command text and at least one preset command text corresponding to the target display area.
[0151] In some embodiments, if a preset instruction text corresponding to the target instruction text exists in at least one preset instruction text corresponding to the target display area, the vehicle terminal determines that the voice instruction is the target voice instruction. If a preset instruction text corresponding to the target instruction text does not exist in at least one preset instruction text corresponding to the target display area, the vehicle terminal determines that the voice instruction is not the target voice instruction.
[0152] In this context, the preset instruction text corresponding to the instruction text in at least one preset instruction text refers to a preset instruction text whose semantic similarity to the instruction text is greater than or equal to a similarity threshold, or it refers to a preset instruction text that is identical to the instruction text in at least one preset instruction text. The similarity threshold is set by a technician according to the actual situation, and this application embodiment does not limit this. In some embodiments, when at least one control corresponding to the target display area is an application launch control, the preset instruction text corresponding to the instruction text in at least one preset instruction text refers to a preset instruction text that is identical to the instruction text in at least one preset instruction text. When at least one control corresponding to the target display area is a function control control, the preset instruction text corresponding to the instruction text in at least one preset instruction text refers to a preset instruction text whose semantic similarity to the instruction text is greater than or equal to a similarity threshold.
[0153] The following describes the method for determining the semantic similarity between at least one preset instruction text and the instruction text.
[0154] In some embodiments, the vehicle terminal extracts semantic features from the instruction text to obtain its semantic features. Based on the semantic features of the instruction text and the semantic features of at least one preset instruction text, the vehicle terminal determines the semantic similarity between the instruction text and the instruction text. When the semantic features are semantic feature vectors, the similarity is represented by the cosine similarity or cosine distance between the semantic feature vectors.
[0155] Optionally, the vehicle terminal can also perform the following steps.
[0156] In some embodiments, when the location indicated by the gesture operation on the vehicle screen belongs to the target display area, the vehicle terminal highlights the controls within the target display area based on the control type of the controls within the target display area.
[0157] The control types include application launch controls and function control controls. Highlighting controls is to remind users which controls are active, so that users can control the controls using corresponding voice commands.
[0158] The above implementation method is illustrated below with two examples.
[0159] Example 1: When the gesture operation indicates a location within the target display area on the vehicle screen and the control type within the target display area is an application launch control, the vehicle terminal highlights or hovers the application launch control, and / or displays the corresponding application number on the application launch control.
[0160] Highlighting includes increasing the brightness of the application launch control or displaying it in a striking color. Floating display refers to creating a "floating" effect by adding a shadow below the application launch control, thus distinguishing it from other controls. The application number of the application launch control can be determined based on its arrangement order. In some embodiments, the application launch controls are numbered sequentially from left to right to obtain their respective application numbers.
[0161] Example 2: When the gesture operation indicates a location within the target display area on the vehicle screen, and the control type within the target display area is a function control control, the vehicle terminal highlights or displays the function control control in a floating manner.
[0162] 306. When the voice command is a target voice command, the vehicle terminal controls the control targeted by the voice command within the target display area based on the preset command text.
[0163] Since the target display area may include multiple controls, the voice command is used to control one of these controls, which is the control targeted by the voice command. When the target display area includes multiple controls, the voice command carries the control identifier of the targeted control. Conversely, when the target display area includes only one control, the voice command does not need to carry the control identifier of the targeted control. In some embodiments, if the control type within the target display area is an application launch control, the target display area may include one or more application launch controls. If the control type within the target display area is a function control control, the target display area includes one function control control. This approach facilitates user control of controls via voice commands.
[0164] In some embodiments, when a voice command is detected, the voice command is a target voice command, and the control type of the control in the target display area is an application launch control, the vehicle terminal launches the application corresponding to the preset command text.
[0165] The application corresponding to the preset command text can be launched by the vehicle terminal by simulating a click operation on the control targeted by the voice command, or by launching the application according to the application identifier corresponding to the control targeted by the voice command. This application embodiment does not limit this.
[0166] In some embodiments, when a voice command is detected, and the voice command is a target voice command, and the control type of the control within the target display area is an application launch control, taking the preset command text corresponding to the target display area as including "Application A", "Application B", and "Application C" and / or "Application ID of Application A", "Application ID of Application B", and "Application ID of Application C" as an example, if the vehicle terminal determines that the preset command text corresponding to the voice command is "Application A", then the vehicle terminal launches "Application A". In this implementation, the combination of gesture operation and voice command enables quick launch of the application.
[0167] In some embodiments, when a voice command is detected, the voice command is a target voice command, and the control type of the control in the target display area is a function control control, the vehicle terminal adjusts the function corresponding to the control in the target display area in accordance with the preset command text indication.
[0168] In some embodiments, when a voice command is detected, the voice command is a target voice command, and the control type of the control within the target display area is a function control control, taking the function control control used to control the volume as an example, and the preset command text corresponding to the target display area includes "increase" and "decrease," if the vehicle terminal determines that the preset command text corresponding to the voice command is "decrease," then the vehicle terminal decreases the audio playback volume. In this implementation, the combination of gesture operation and voice command enables rapid adjustment of functions.
[0169] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0170] Figure 6 is a schematic diagram of a control device based on an in-vehicle screen provided in an embodiment of this application. Referring to Figure 6, the device includes: a determination module 601, a loading module 602, and a control module 603.
[0171] The determination module 601 is used to determine whether the position indicated by the gesture operation on the vehicle screen belongs to a target display area when the vehicle screen is started and there is a gesture operation on the vehicle screen. The target display area includes at least one control displayed on the vehicle screen.
[0172] The loading module 602 is used to load at least one preset instruction text corresponding to the target display area when the position indicated by the gesture operation on the vehicle screen belongs to the target display area. The preset instruction text is instruction text used to control the controls within the target display area.
[0173] The control module 603 is used to control the control targeted by the voice command in the target display area based on the preset command text when a voice command is detected and the voice command is a target voice command. The target voice command is a voice command that matches any of the preset command texts in the at least one preset command text.
[0174] In some embodiments, the determining module 601 is configured to determine the spatial location and direction of a gesture operation when the in-vehicle screen is activated and a gesture operation is performed on the in-vehicle screen. Based on the spatial location and direction of the gesture operation, the position indicated by the gesture operation on the in-vehicle screen is determined. Based on the screen coordinates of the position and the screen coordinates of the target display area, it is determined whether the position belongs to the target display area.
[0175] In some embodiments, the determining module 601 is configured to acquire a three-dimensional image of a gesture operation when the in-vehicle screen is activated and a gesture operation is performed on the in-vehicle screen. Based on the three-dimensional image, the spatial position and gesture posture of the gesture operation are determined. Based on the gesture posture of the gesture operation, the direction of the gesture operation is determined.
[0176] In some embodiments, the loading module 602 is configured to, when the location indicated by the gesture operation on the vehicle screen belongs to the target display area, recall instruction text based on controls within the target display area to obtain at least one preset instruction text corresponding to the target display area. The at least one preset instruction text corresponding to the target display area is then loaded.
[0177] In some embodiments, the loading module 602 is configured to, when the control in the target display area is an application launch control, recall based on the application identifier corresponding to the control in the target display area to obtain at least one preset instruction text corresponding to the target display area. The at least one preset instruction text includes at least one of the following: application name, application name abbreviation, application number, and application alias. When the control in the target display area is a function control control, recall based on the function identifier corresponding to the control in the target display area to obtain at least one preset instruction text corresponding to the target display area. The function identifier is used to indicate the function provided by the corresponding control, and the at least one preset instruction text includes function control terms.
[0178] In some embodiments, the device further includes:
[0179] The detection module is used to detect the presence of a voice command. If a voice command is present, it performs speech recognition on the voice command to obtain the corresponding command text. Based on the command text and at least one preset command text corresponding to the target display area, it determines whether the voice command is the target voice command.
[0180] In some embodiments, the detection module is configured to determine that the voice instruction is the target voice instruction if a corresponding preset instruction text exists in at least one preset instruction text corresponding to the target display area; and to determine that the voice instruction is not the target voice instruction if a corresponding preset instruction text does not exist in at least one preset instruction text corresponding to the target display area.
[0181] In some embodiments, the device further includes:
[0182] The display module is used to highlight controls within the target display area based on their control types when the location indicated by the gesture operation on the vehicle screen is within the target display area.
[0183] In some embodiments, the display module is configured to, when the control type of the control within the target display area is an application launch control, highlight or hover the application launch control, and / or display the corresponding application number on the application launch control. When the control type of the control within the target display area is a function control control, highlight or hover the function control control.
[0184] In some embodiments, the control module 603 is configured to launch the application corresponding to the preset instruction text when a voice instruction is detected, the voice instruction is a target voice instruction, and the control type of the control in the target display area is an application launch control. When a voice instruction is detected, the voice instruction is a target voice instruction, and the control type of the control in the target display area is a function control control, the module adjusts the function corresponding to the control in the target display area according to the method indicated by the preset instruction text.
[0185] It should be noted that the above embodiments of the vehicle-mounted screen-based control device are only illustrative examples of the above-described functional module divisions when controlling controls based on the vehicle-mounted screen. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the vehicle-mounted screen-based control device and the vehicle-mounted screen-based control method embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0186] The technical solution provided in this application, when the in-vehicle screen is activated and a gesture operation is performed on it, indicates that the user is controlling the screen via gesture. It determines whether the location indicated by the gesture on the screen belongs to a target display area, which includes at least one control displayed on the screen. If the location indicated by the gesture belongs to the target display area, at least one preset instruction text corresponding to that target display area is loaded. This preset instruction text can control the controls within the target display area. When a voice command is detected and is a target voice command, the control corresponding to that voice command within the target display area is controlled based on the preset instruction text. The target voice command is a voice command that matches any preset instruction text. This combination of gesture operation and voice command allows the user to control the in-vehicle screen without physical contact, improving the convenience of screen control.
[0187] This application also provides a vehicle, and Figure 7 is a structural schematic diagram of a vehicle provided in this application embodiment.
[0188] Typically, vehicle 700 includes one or more processors 701 and one or more memories 702.
[0189] Processor 701 may include one or more processing cores, such as a 4-core processor or a 7-core processor in some embodiments. Processor 701 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 701 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 701 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 701 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0190] The memory 702 may include one or more computer-readable storage media, which may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, and in some embodiments, one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 is used to store at least one computer program, which is executed by the processor 701 to implement the control control method based on an in-vehicle screen provided in the method embodiments of this application.
[0191] Those skilled in the art will understand that the structure shown in FIG7 does not constitute a limitation on vehicle 700, and may include more or fewer components than shown, or combine certain components, or employ different component arrangements.
[0192] In addition, the device provided in the embodiments of this application may specifically be a chip, component or module. The chip may include a connected processor and a memory. The memory is used to store instructions. When the processor calls and executes the instructions, the chip can execute the control method based on the vehicle screen provided in the above embodiments.
[0193] This embodiment also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the above-described related method steps to implement the control method based on an in-vehicle screen provided in the above embodiment.
[0194] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement the control method based on an in-vehicle screen provided in the above embodiment.
[0195] In this embodiment, the device, computer-readable storage medium, computer program product, or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0196] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0197] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0198] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A control method based on an in-vehicle screen, the method comprising: When the in-vehicle screen is activated in the vehicle and there is a gesture operation on the in-vehicle screen, it is determined whether the position indicated by the gesture operation on the in-vehicle screen belongs to the target display area, the target display area including at least one control displayed on the in-vehicle screen; When the position indicated by the gesture operation on the vehicle screen belongs to the target display area, at least one preset instruction text corresponding to the target display area is loaded. The preset instruction text is instruction text used to control the controls within the target display area. When a voice command is detected and the voice command is a target voice command, the control targeted by the voice command in the target display area is controlled based on the preset command text. The target voice command is a voice command that matches any of the at least one preset command text.
2. The method according to claim 1, wherein, When the in-vehicle screen is activated and a gesture operation is performed on the in-vehicle screen, determining whether the location indicated by the gesture operation on the in-vehicle screen belongs to the target display area includes: When the in-vehicle screen is activated and there is a gesture operation on the in-vehicle screen, determine the spatial location and direction of the gesture operation; Based on the spatial location and direction of the gesture operation, determine the position indicated by the gesture operation on the vehicle screen; Based on the screen coordinates of the location and the screen coordinates of the target display area, determine whether the location belongs to the target display area.
3. The method according to claim 2, wherein, When the in-vehicle screen is activated and a gesture operation is performed on the in-vehicle screen, determining the spatial location and direction of the gesture operation includes: When the in-vehicle screen is activated in the vehicle and there is a gesture operation on the in-vehicle screen, a three-dimensional image of the gesture operation is acquired. Based on the three-dimensional image, the spatial location and gesture posture of the gesture operation are determined; Based on the gesture posture of the gesture operation, the direction of the gesture operation is determined.
4. The method according to claim 3, wherein, Determining the spatial location and gesture posture of the gesture operation based on the three-dimensional image includes: Based on the three-dimensional image, multiple spatial key points of the gesture operation are determined; Based on the multiple spatial key points, the spatial position and gesture posture of the gesture operation are determined.
5. The method according to claim 4, wherein, Determining the spatial location and gesture posture of the gesture operation based on the multiple spatial key points includes: The spatial position of the gesture operation is determined based on the spatial coordinates of each key spatial point. The gesture posture of the gesture operation is determined based on the relative positional relationship of the multiple spatial key points.
6. The method according to claim 1, wherein, When the location indicated by the gesture operation on the vehicle screen belongs to the target display area, loading at least one preset instruction text corresponding to the target display area includes: When the location indicated by the gesture on the vehicle screen belongs to the target display area, the command text is recalled based on the controls within the target display area to obtain at least one preset command text corresponding to the target display area. Load at least one preset instruction text corresponding to the target display area.
7. The method according to claim 6, wherein, The step of recalling instruction text based on controls within the target display area to obtain at least one preset instruction text corresponding to the target display area includes: If the control in the target display area is an application launch control, a recall is performed based on the application identifier corresponding to the control in the target display area to obtain at least one preset instruction text corresponding to the target display area. The at least one preset instruction text includes at least one of the following: application name, application name abbreviation, application number, and application alias. When the control in the target display area is a function control control, a recall is performed based on the function identifier corresponding to the control in the target display area to at least one preset instruction text corresponding to the target display area. The function identifier is used to indicate the function provided by the corresponding control, and the at least one preset instruction text includes function control vocabulary.
8. The method according to claim 1, wherein, Before controlling the control targeted by the voice command within the target display area based on the preset command text when a voice command is detected and the voice command is a target voice command, the method further includes: Detect the presence of voice commands; In the presence of a voice command, the voice command is recognized to obtain the command text corresponding to the voice command; Based on the instruction text and at least one preset instruction text corresponding to the target display area, determine whether the voice instruction is a target voice instruction.
9. The method according to claim 8, wherein, The step of determining whether the voice command is a target voice command based on the command text and at least one preset command text corresponding to the target display area includes: If a preset instruction text corresponding to the instruction text exists in at least one preset instruction text corresponding to the target display area, the voice instruction is determined to be the target voice instruction; If no preset instruction text corresponding to the instruction text exists in at least one preset instruction text corresponding to the target display area, it is determined that the voice instruction is not the target voice instruction.
10. The method according to claim 1, wherein, The method further includes: When the location indicated by the gesture operation on the vehicle screen belongs to the target display area, the controls within the target display area are highlighted based on their control types.
11. The method according to claim 10, wherein, The step of highlighting controls within the target display area based on their control types includes: If the control type of the control in the target display area is an application launch control, the application launch control is highlighted or floated, and / or the corresponding application number is displayed on the application launch control; If the control type within the target display area is a function control control, then the function control control is highlighted or displayed in a floating manner.
12. The method according to claim 1, wherein, When a voice command is detected and the voice command is a target voice command, the control of the control targeted by the voice command within the target display area based on the preset command text includes: If a voice command is detected, the voice command is a target voice command, and the control type of the control in the target display area is an application launch control, the application corresponding to the preset command text is launched; Upon detecting a voice command, wherein the voice command is a target voice command and the control type of the control in the target display area is a function control control, the function corresponding to the control in the target display area is adjusted according to the preset command text indication.
13. The method according to claim 1, wherein, The method further includes: When the in-vehicle screen is activated in the vehicle, determine whether there is a gesture operation in the in-vehicle space corresponding to the in-vehicle screen; If a gesture operation exists in the in-vehicle space corresponding to the in-vehicle screen, determine the type of the gesture operation; Based on the type of the gesture operation, determine whether the gesture operation is a gesture operation targeting the vehicle screen.
14. The method according to claim 13, wherein, When the in-vehicle screen is activated in the vehicle, determining whether there is a gesture operation in the in-vehicle space corresponding to the in-vehicle screen includes: When the in-vehicle screen is activated in the vehicle, the in-vehicle images captured by the gesture recognition device of the in-vehicle screen are obtained; Target detection is performed on the in-vehicle image to determine whether there is a gesture operation in the in-vehicle space corresponding to the in-vehicle screen.
15. The method according to claim 14, wherein, When a gesture operation exists in the in-vehicle space corresponding to the in-vehicle screen, determining the type of the gesture operation includes: When a gesture operation exists in the in-vehicle space corresponding to the in-vehicle screen, the type of the gesture operation is determined based on the image area corresponding to the gesture operation in the in-vehicle image.
16. The method according to claim 13, wherein, Determining whether a gesture operation is a gesture operation targeting the in-vehicle screen based on the type of the gesture operation includes: If the type of the gesture operation is a preset type, then the gesture operation is determined to be a gesture operation for the vehicle screen. If the type of the gesture operation is not the preset type, it is determined that the gesture operation is not a gesture operation for the vehicle screen.
17. A control device based on an in-vehicle screen, the device comprising: The determination module is used to determine whether the position indicated by the gesture operation on the vehicle screen belongs to a target display area when the vehicle screen is started and there is a gesture operation on the vehicle screen. The target display area includes at least one control displayed on the vehicle screen. A loading module is used to load at least one preset instruction text corresponding to the target display area when the position indicated by the gesture operation on the vehicle screen belongs to the target display area. The preset instruction text is instruction text for controlling the controls within the target display area. A control module is configured to, when a voice command is detected and the voice command is a target voice command, control the control targeted by the voice command within the target display area based on the preset command text, wherein the target voice command is a voice command that matches any of the at least one preset command text.
18. A vehicle, the vehicle comprising: Memory, used to store executable program code; A processor is configured to call and run the executable program code from the memory, causing the vehicle to perform the control method based on the in-vehicle screen as described in any one of claims 1 to 16.
19. A computer-readable storage medium storing at least one piece of program code, the program code being loaded and executed by a processor to implement the control method based on an in-vehicle screen as described in any one of claims 1-16.
Citation Information
Patent Citations
Air gesture operation method and apparatus
CN106055098A
Voice control method and device, electronic equipment and storage medium
CN115457951A
Remote screen control method, vehicle and computer readable storage medium
CN115617232A
Voice control method and device, equipment and medium
CN115810354A
Control control method based on vehicle-mounted screen and vehicle
CN118519530A