Control method of target vehicle, server, device, and storage medium

By using cloud servers and large language models to identify the position of controls on the vehicle display screen, the problem of cumbersome and inefficient control recognition in existing technologies has been solved, achieving efficient and accurate control control and improving the user experience.

CN116958569BActive Publication Date: 2026-04-28APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD
Filing Date
2023-07-12
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies are cumbersome and inefficient in complex scenarios, making it difficult to efficiently identify and control controls on vehicle display screens.

Method used

Based on the display screen image features of the target vehicle, the system uses a large language model to identify the control position information via a cloud server, and then uses the target audio to achieve automated control of the controls.

Benefits of technology

It improves the recognition capability and accuracy of "what you see is what you hear", enhances the user experience, simplifies the control recognition process, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958569B_ABST
    Figure CN116958569B_ABST
Patent Text Reader

Abstract

The present disclosure provides a target vehicle control method, a server, a device and a storage medium, relates to the technical field of data processing, and particularly relates to the technical field of car-machine interaction, artificial intelligence, cloud computing and big data. The specific implementation scheme is: in the case that control data for a target vehicle is detected, target intention features of the control data are obtained based on the control data and target image features corresponding to the target vehicle; the control data is used to indicate a control operation on a target control in a display screen of the target vehicle; the target image features represent image features of an interface image displayed by the display screen; target position information of a target control pointed to by the target intention features is output; and the target position information represents position information of the target control in the interface image displayed by the display screen.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to the fields of vehicle-to-everything (V2X) interaction, artificial intelligence, cloud computing and big data. Background Technology

[0002] With the advancement of voice recognition technology, voice recognition can be implemented in more and more scenarios. For example, in driving scenarios, voice recognition can meet user needs, thereby improving both the level of intelligence and driving safety. Summary of the Invention

[0003] This disclosure provides a method for controlling a target vehicle, a server, a device, and a storage medium.

[0004] According to one aspect of this disclosure, a method for controlling a target vehicle is provided, comprising:

[0005] Upon detecting control data targeting a target vehicle, a target intent feature of the control data is obtained based on the control data and the target image features corresponding to the target vehicle; wherein, the control data is used to instruct control operations on target controls on the display screen of the target vehicle; and the target image features characterize the image features of the interface image displayed on the display screen.

[0006] Output the target position information of the target control pointed to by the target intent feature; wherein, the target position information represents the position information of the target control in the interface image displayed on the display screen.

[0007] According to another aspect of this disclosure, a method for controlling a target vehicle is provided, comprising:

[0008] In response to the target audio, control data is sent; wherein the control data is obtained based on the target audio and is used to instruct control operations to be performed on the target controls on the display screen of the target vehicle.

[0009] The target position information of the target control pointed to by the target intent feature of the control data is obtained; wherein, the target position information represents the position information of the target control in the interface image displayed on the display screen;

[0010] Based on the target position information of the target control, control operations are performed on the target control.

[0011] According to another aspect of this disclosure, a cloud server is provided, comprising:

[0012] The processing unit is configured to, upon detecting control data for a target vehicle, obtain target intent features of the control data based on the control data and target image features corresponding to the target vehicle; wherein the control data is used to instruct control operations on target controls on the display screen of the target vehicle; and the target image features characterize the image features of the interface image displayed on the display screen.

[0013] The output unit is used to output the target position information of the target control pointed to by the target intent feature; wherein the target position information represents the position information of the target control in the interface image displayed on the display screen.

[0014] According to another aspect of this disclosure, an in-vehicle device is provided, comprising:

[0015] A sending unit is configured to send control data in response to target audio; wherein the control data is obtained based on the target audio and is used to instruct control operations to be performed on target controls on the display screen of the target vehicle.

[0016] The acquisition unit is used to acquire the target position information of the target control pointed to by the target intent feature of the control data; wherein, the target position information represents the position information of the target control in the interface image displayed on the display screen;

[0017] The control unit is used to perform control operations on the target control based on the target position information of the target control.

[0018] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0019] At least one processor; and

[0020] The memory is communicatively connected to the at least one processor; wherein,

[0021] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.

[0022] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.

[0023] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.

[0024] In this way, the disclosed solution can obtain the target intent features corresponding to the control data of the target vehicle based on the image features of the interface image displayed on the target vehicle's display screen (that is, the target image features corresponding to the target vehicle), and then obtain the target position information of the target control to be controlled. This lays the foundation for subsequent automated control operations on the target control in the target vehicle's display screen. At the same time, it also effectively improves the recognition capability and recognition accuracy of "what you see is what you can say", thus laying the foundation for improving the user experience.

[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0026] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0027] Figure 1 This is an illustrative flow diagram of a control method for a target vehicle applied to a cloud server according to an embodiment of this application. Figure 1 ;

[0028] Figure 2 This is a schematic flowchart of a target vehicle control method applied in a target vehicle according to an embodiment of this application;

[0029] Figure 3 This is a schematic diagram of the coordinate system constructed for the target position information transformation of the target control according to an embodiment of this application;

[0030] Figure 4 This is an illustrative flow diagram of a control method for a target vehicle applied to a cloud server according to an embodiment of this application. Figure 2 ;

[0031] Figure 5 This is a schematic diagram of a control method for a target vehicle according to an embodiment of this application. Figure 2 ;

[0032] Figure 6(a) is a schematic diagram of a control method for a target vehicle according to an embodiment of this application. Figure 3 ;

[0033] Figure 6(b) is a schematic diagram of a control method for a target vehicle according to an embodiment of this application. Figure 4 ;

[0034] Figure 7 This is an illustrative flow diagram of a control method for a target vehicle applied to a cloud server according to an embodiment of this application. Figure 3 ;

[0035] Figure 8 This is a schematic diagram of the implementation process of the target vehicle control method according to an embodiment of this application in a specific embodiment;

[0036] Figure 9 This is a schematic diagram of the structure of a cloud server according to an embodiment of this disclosure;

[0037] Figure 10 This is a schematic diagram of the structure of the vehicle-mounted device according to an embodiment of the present disclosure;

[0038] Figure 11 This is a block diagram of an electronic device used to implement the control method for the target vehicle in the embodiments of this disclosure. Detailed Implementation

[0039] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0040] In this document, the term "and / or" merely describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The term "at least one" in this document indicates any combination of at least two of a plurality of elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this document refer to and distinguish between multiple similar technical terms, not to restrict the order or to limit there to only two. For example, "first feature" and "second feature" refer to two categories / two features; the first feature can be one or more, and the second feature can also be one or more.

[0041] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can still be practiced even without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0042] Existing "what you see can be said" technologies typically employ the following three methods:

[0043] Registration-based: This method pre-registers the text corresponding to a control. When the corresponding text is detected (e.g., when user-inputted speech containing text), the control can be identified. However, registration-based methods can only pre-register certain controls; unregistered controls cannot be recognized.

[0044] Accessibility Assistance: This method extracts the text corresponding to controls using accessibility assistance. If the text corresponding to a control is detected (for example, if the user's voice input contains text), the control can be identified. However, accessibility assistance can only extract text from native controls (such as pre-installed controls, which can be collectively referred to as native controls). It cannot recognize non-native controls such as web pages, nor can it recognize other non-text controls.

[0045] OCR (Optical Character Recognition) method: Images (such as control icons) are pre-trained. If a pre-trained image is detected, the corresponding control can be identified. However, OCR can only recognize pre-trained icons; it cannot recognize untrained icons.

[0046] It is foreseeable that as the complexity of scenarios increases, the process of "what you see can be said" will become cumbersome and inefficient. Therefore, in order to improve the accuracy of "what you see can be said" recognition, a highly practical implementation method is urgently needed.

[0047] Based on this, the present disclosure provides a control method for a target vehicle that significantly improves recognition capabilities and efficiency while achieving the "what you see is what you speak" function. Furthermore, the present disclosure requires no modifications to the target vehicle itself, thus enhancing the user experience without significantly increasing costs.

[0048] Figure 1 This is an illustrative flow diagram of a control method for a target vehicle applied to a cloud server according to an embodiment of this application. Figure 1 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. For example, it can be applied to cloud servers that interact with target vehicles.

[0049] Furthermore, the method includes at least a portion of the following: For example... Figure 1 As shown, it includes:

[0050] Step S101: When control data for a target vehicle is detected, the target intent feature of the control data is obtained based on the control data and the target image features corresponding to the target vehicle.

[0051] Here, the control data is used to instruct control operations on the target controls in the display screen of the target vehicle; the target image features characterize the image features of the interface image displayed on the display screen.

[0052] Step S102: Output the target location information of the target control pointed to by the target intent feature.

[0053] Here, the target location information refers to the position information of the target control in the interface image displayed on the display screen.

[0054] In this way, the present solution can obtain the target intent features corresponding to the control data based on the image features of the interface image displayed on the target vehicle's display screen (that is, the target image features corresponding to the target vehicle), and then obtain the target position information of the target control. This lays the foundation for subsequent automated control operations on the target control in the target vehicle's display screen. At the same time, it also effectively improves the recognition capability and recognition accuracy of "what you see is what you can say", thus laying the foundation for improving the user experience.

[0055] In a specific example, after obtaining the target location information of the target control, the cloud server can also send the target location information of the target control pointed to by the target intent feature. This facilitates the receiving end, such as the target vehicle, to realize the function of "what you see is what you can say", thereby improving the intelligence level of the target vehicle and further enhancing the user experience.

[0056] Figure 2 This is a schematic flowchart illustrating a control method for a target vehicle applied in a target vehicle according to an embodiment of this application. The method can optionally be applied to electronic devices, such as personal computers, servers, server clusters, etc. For example, it can be applied to a target vehicle, and more specifically, to in-vehicle equipment installed in the target vehicle.

[0057] Furthermore, the method includes at least a portion of the following: For example... Figure 2 As shown, it includes:

[0058] Step S201: In response to the target audio, send control data.

[0059] Here, the control data is obtained based on the target audio and is used to instruct the target controls on the display screen of the target vehicle to perform control operations.

[0060] Furthermore, in one specific example, the control data may specifically be the target audio; in this case, the target vehicle directly sends the target audio as the control data to the cloud server. Alternatively, in another example, the control data may include text data obtained after text recognition based on the target audio; in this case, the target vehicle sends the text data obtained after text recognition of the target audio to the cloud server, thereby utilizing the cloud server for control recognition.

[0061] In a specific example, control data is sent only if the target audio meets the intent conditions, such as the see-is-say intent.

[0062] Step S202: Obtain the target position information of the target control pointed to by the target intent feature of the control data.

[0063] Here, the target location information refers to the position information of the target control in the interface image displayed on the display screen.

[0064] For example, the target vehicle obtains the target location information of the target control from the cloud server, or the target vehicle receives the target location information sent by the cloud server.

[0065] Step S203: Based on the target position information of the target control, perform control operations on the target control.

[0066] For example, the target vehicle can control the target control based on the target location information sent by the cloud server, thus realizing the principle of "what you see is what you get".

[0067] Here, the control operation can be to process the target control accordingly, such as to start, close, click or switch the target control on the display screen of the target vehicle. This disclosure does not limit the specific content of the control operation.

[0068] In this way, the disclosed solution can realize the control operation of the target control based on the target audio, which can greatly improve the recognition capability and accuracy of "what you see can be spoken", enhance the intelligence level of the target vehicle, and further improve the user experience.

[0069] Furthermore, in a specific example, the position information can be converted in the target vehicle in the following way, which facilitates determining the specific position of the target control on the target vehicle's display screen, laying the foundation for accurate control operations. Specifically, the control operation on the target control based on the target position information of the target control (e.g., step S203 above) specifically includes:

[0070] Based on the target position information of the target control (e.g., the image coordinate information of the target control in the current interface image displayed on the display screen), the screen position information of the target control on the display screen (e.g., the screen coordinate information of the target control on the display screen) is obtained.

[0071] Based on the screen position information of the target control on the display screen, control operations are performed on the target control.

[0072] In a specific example, the target vehicle takes a screenshot of the current interface of the display screen to obtain an image of the current interface. In practical applications, to conserve network resources, the target vehicle can compress the screenshot image, for example, by proportional compression, and then upload the compressed image (e.g., the bitmap data of the compressed image) to a cloud server. Furthermore, the target position information of the target control output by the cloud server can be the image coordinate information of the target control within the current interface image (i.e., the compressed image), for example... Figure 3 As shown, in the cloud server, a coordinate system is established with a specific location of the received current interface image (i.e., the compressed current interface image), such as a vertex of a corner, as the origin. X-axis and Y-axis are established along the width and length directions of the compressed current interface image, respectively. The image coordinates of the target control in the compressed current interface image can then be denoted as...<x,y> The cloud server will display the image coordinates of the target control within the compressed current interface image.<x,y> The image coordinates are sent to the target vehicle; correspondingly, the target vehicle receives the image coordinates sent by the cloud server.<x,y> Then, the following coordinate transformation is performed to obtain the screen coordinates of the target control on the display screen:

[0073] For the received image coordinate points<x,y> Processing is performed to obtain the intermediate coordinate point.<dx,dy> Where dx = x / the width of the current interface image (i.e., the compressed current interface image) received by the cloud server, for example dx = x / 700, dy = y / the length of the current interface image (i.e., the compressed current interface image) received by the cloud server, for example dy = y / 500.

[0074] Based on the length and width of the display interface on the screen, the intermediate coordinate point<dx,dy> Processing is performed to obtain the screen coordinates.<sx,sy> Where sx = dx × the width of the display screen, for example, sx = dx × 1920; sy = dy × the length (i.e., height) of the display screen, for example, sy = dy × 1080.

[0075] In a specific example of the scheme disclosed herein, a target model, such as a large language model, can be used for control recognition; specifically, when control data for a target vehicle is detected, the control data is input into the target model.

[0076] Here, the target model is used to determine the target image features corresponding to the target vehicle, and obtain the target intent features based on the control data and the target image features, and then output the target position information of the target control pointed to by the target intent features.

[0077] In a specific example, the control method of the target vehicle, such as Figure 4 As shown, it specifically includes:

[0078] Step S401: If control data for the target vehicle is detected, the control data is input to the target model.

[0079] Here, the target model is used to determine the target image features corresponding to the target vehicle, and to obtain the target intent features based on the control data and the target image features.

[0080] Furthermore, in a specific example, the target model is also used to label the bitmap data of the interface image displayed on the target vehicle's screen, and extract the image features of the labeled bitmap data to obtain the target image features corresponding to the target vehicle. This allows the target model to extract the target image features corresponding to the target vehicle from its pre-processed data when receiving control data from the target vehicle, thereby providing data support for rapid response to control data.

[0081] It should be noted that bitmap data can specifically refer to the bit array of the bitmap image of the interface image displayed on the screen.

[0082] Step S402: Output the target location information of the target control pointed to by the target intent feature.

[0083] Here, the target position information refers to the position information of the target control within the interface image displayed on the screen. Further, the target position information refers to the position information of the target control within the bitmap data of the interface image marked by the target model.

[0084] In this way, the present invention utilizes the target model to obtain the target intent features corresponding to the control data, and then obtains the target position information of the target control. This lays the foundation for subsequent automated control operations on the target control in the display screen of the target vehicle. At the same time, it also effectively improves the recognition capability and recognition accuracy of "what you see is what you can say", thus laying the foundation for improving the user experience.

[0085] In a specific example, Figure 5 This is a scenario illustration of the control method for the target vehicle in this disclosure. Figure 2 This scenario involves a target vehicle 501 and a cloud server (or server cluster) 502. Specifically, the target vehicle 501 sends control data to the cloud server 502, which then inputs the control data into a target model to obtain the target location information of the target control pointed to by the target intent feature. The cloud server 502 then sends the target location information of the target control pointed to by the target intent feature back to the target vehicle 501. Based on the target location information of the target control, the target vehicle 501 can perform control operations on the target control, thus quickly realizing what you see is what you can say, thereby effectively improving the user experience.

[0086] In a specific example of the disclosed solution, the bitmap data of the interface image at a specific moment can be input into the target model in the following manner. This facilitates the target model to pre-process the bitmap data of the interface image, laying the foundation for subsequent rapid response control data and improving response efficiency. Specifically, the method further includes:

[0087] Step S400-1: Obtain the bitmap data of the interface image displayed on the display screen at the first moment.

[0088] Step S400-2: Compare the bitmap data of the interface image displayed on the screen at the first moment with the historical bitmap data marked by the target model.

[0089] For example, in one instance, if it is determined that the bitmap data of the interface image displayed on the screen at the second moment is the latest bitmap data of the interface image, the bitmap data of the interface image displayed on the screen at the second moment can be input into the target model. At this time, the target model marks the bitmap data of the interface image displayed on the screen at the second moment as the bitmap data of the interface image displayed on the screen of the target vehicle, and then extracts the image features of the marked bitmap data to obtain the target image features corresponding to the target vehicle. This provides data support for subsequent rapid response control data. Here, the second moment is at least the previous moment of the first moment (i.e., the historical moment corresponding to the first moment). Based on this, the bitmap data of the interface image displayed on the screen at the second moment marked by the target model is historical bitmap data used for comparison with the bitmap data at the first moment.

[0090] Step S400-3: If the bitmap data of the interface image displayed on the screen at the first moment is different from the historical bitmap data marked by the target model, the bitmap data of the interface image displayed on the screen at the first moment is input into the target model.

[0091] Furthermore, in one example, after inputting the bitmap data of the interface image displayed on the screen at the first moment into the target model, the target model is also used to update the marked historical bitmap data and update it to the bitmap data of the interface image displayed on the screen at the first moment, thereby extracting the image features of the newly marked bitmap data to obtain the new target image features corresponding to the target vehicle. This ensures that the obtained target image features are the image features of the latest bitmap data of the interface image, laying the foundation for effectively improving the response efficiency of control data and improving the response accuracy.

[0092] In a specific example, steps S400-1 to S400-3 described above may be performed before the control data is input to the target model, such as step S401 described above.

[0093] It should be noted that in this example, if the bitmap data of the interface image obtained by the cloud server at the first moment is different from the historical bitmap data marked by the target model, it indicates that the interface displayed on the target vehicle's screen has changed. In this case, the latest interface image after the change, that is, the bitmap data of the interface image displayed on the screen at the first moment, can be input into the target model. This allows the target model to use the latest interface image updated in real time to recognize controls, thereby improving recognition accuracy and laying a foundation for improving user experience.

[0094] For example, in a specific example, as shown in Figure 6(a), the target vehicle uploads the interface image (e.g., bitmap data of the interface image) displayed on the screen at the second moment to the cloud server. If the cloud server determines that the interface image displayed on the screen at the second moment is different from the interface image displayed on the screen at other historical moments (e.g., at least the moment before the second moment), for example, different from the historical bitmap data marked by the target model, it indicates that the interface image displayed on the screen at the second moment is the latest interface image of the target vehicle's screen. At this time, the cloud server can input the bitmap data of the interface image displayed on the screen at the second moment into the target model. Correspondingly, the target model marks the bitmap data of the interface image displayed on the screen at the second moment as the bitmap data of the interface image displayed on the screen, and then extracts the image features of the marked bitmap data to obtain the target image features corresponding to the target vehicle.

[0095] Further, as shown in Figure 6(b), the target vehicle's display interface changes from the interface at the second moment to the interface at the first moment. At this time, the target vehicle can take a screenshot of the current interface of the display screen to obtain the interface image at the first moment and upload it to the cloud server. The cloud server compares the interface image displayed on the display screen at the first moment with the interface image displayed on the display screen at the second moment, that is, it compares it with the historical interface images marked by the target model. For example, the cloud server compares the bitmap data of the interface image at the first moment with the bitmap data of the interface image at the second moment (that is, the historical bitmap data marked by the target model). If it is determined that the two are different, the acquired bitmap data of the interface image at the first moment, that is, the bitmap data of the latest changed interface image, is input into the target model. In this way, it is convenient to quickly use the bitmap data of the interface image at the first moment to perform control recognition when control data is detected, thus further improving the recognition efficiency.

[0096] Furthermore, in a specific example, after inputting the bitmap data of the interface image displayed on the screen at the first moment into the target model, the target model is further used to update the marked interface image displayed on the screen, so as to update the bitmap data of the interface image displayed on the screen at the first moment to the bitmap data of the interface image displayed on the screen. This facilitates the target model to use the updated latest interface image for control recognition, thereby improving recognition accuracy and laying a foundation for improving the user experience.

[0097] Furthermore, the target model can also extract features from the bitmap data of the interface image displayed on the updated display screen, that is, the bitmap data of the interface image at the first moment, to obtain new target image features. This facilitates the rapid use of the updated new target image features for control recognition when control data is detected, thereby further improving recognition efficiency.

[0098] In a specific example of the scheme disclosed herein, the time when the control data is input to the target model is after the first time moment, that is, at least the next time moment after the first time moment.

[0099] In other words, the data required for control recognition by the target model in this disclosed solution, such as control data and bitmap data of the interface image displayed on the screen, can be input to the target model at different times. For example, in one example, the target vehicle can take a screenshot of the current interface of the target vehicle's display screen at a preset interval and upload it to the cloud server. This allows the cloud server to input the latest bitmap data of the interface image into the target model, enabling the target model to mark the latest bitmap data of the target vehicle's display screen and perform feature extraction in advance. Consequently, when control data targeting the target vehicle is detected, the target model can quickly call the relevant features of the marked interface image of the target vehicle's display screen, i.e., the target image features corresponding to the target vehicle, to perform control recognition. This improves the response speed and lays the foundation for further improving the user experience.

[0100] Figure 7 This is an illustrative flow diagram of a control method for a target vehicle applied to a cloud server according to an embodiment of this application. Figure 3 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. For example, it can be applied to cloud servers.

[0101] Furthermore, the method includes at least a portion of the following: For example... Figure 7 As shown, it includes:

[0102] Step S701: Obtain the bitmap data of the interface image displayed on the display screen at the first moment.

[0103] In a specific example of the disclosed solution, the target vehicle can upload interface images at different times to a cloud server. Specifically, the target vehicle obtains interface images of its own display screen at different times and sends the interface images of the target vehicle's display screen at different times. In this way, the cloud server can receive interface images at different times, thereby laying the foundation for improving the recognition accuracy by performing control recognition based on the latest interface images.

[0104] It should be noted that, in a specific example, the target vehicle can take screenshots of the display screen at preset time intervals to obtain the interface image at the current moment, and then upload the interface image at the current moment to the cloud server. In this way, the cloud server can obtain interface images at different times. Alternatively, the target vehicle can record the display screen and upload the obtained video data to the cloud server. In this case, the cloud server can obtain interface images at different times based on the video data.

[0105] Furthermore, in another specific example, the data uploaded by the target vehicle to the cloud server can specifically be interface images, or it can be processed interface images, such as bitmap data of the interface images. That is, the target vehicle acquires interface images of its own display screen at different times; and sends the bitmap data of the interface images of the target vehicle's display screen at different times. This allows the cloud server to receive the bitmap data of the interface images at different times, thus laying the foundation for improving recognition accuracy by performing control recognition based on the latest interface images. At the same time, since the transmitted data is bitmap data, rather than the image data itself, the amount of data transmitted is effectively reduced, thereby laying the foundation for further improving the user experience.

[0106] Step S702: Compare the bitmap data of the interface image displayed on the display screen at the first moment with the historical bitmap data marked by the target model.

[0107] In a specific example, two bitmap data can be compared in the following way: specifically, the comparison of the bitmap data of the interface image displayed on the screen at the first moment with the historical bitmap data marked by the target model (e.g., step S702) as described above includes:

[0108] Step S702-1: Obtain the total coordinate information of the bitmap data of the interface image displayed on the display screen at the first moment.

[0109] Step S702-2: Compare the total coordinate information of the bitmap data of the interface image at the first moment with the total coordinate information of the historical bitmap data marked by the target model.

[0110] It should be noted that, in one example, the total coordinate information of the historical bitmap data marked by the target model can be obtained in advance. For example, in one example, the total coordinate information of the historical bitmap data can be calculated when the historical bitmap data is obtained from the cloud server; or, in another example, the total coordinate information of the historical bitmap data is calculated by comparing it with the total coordinate information of other interface images.

[0111] Thus, this disclosure provides a specific implementation scheme based on bitmap data comparison. For example, it provides a specific scheme for comparing bitmap data of interface images at different times. This scheme is simple, easy to implement, and has low computational cost. This lays the foundation for greatly improving the recognition efficiency of "what you see is what you can say", and in turn, lays the foundation for improving the user experience.

[0112] Step S703: Based on the comparison result, determine whether the bitmap data of the interface image displayed on the display screen at the first moment is the same as the historical bitmap data marked by the target model; if they are the same, proceed to step S704; otherwise, proceed to step S705.

[0113] Step S704: If it is determined that the bitmap data of the interface image displayed on the display screen at the first moment is the same as the historical bitmap data marked by the target model, the bitmap data of the interface image displayed on the display screen at the first moment is masked, and the process ends.

[0114] Here, "masking" means that the bitmap data of the interface image displayed on the screen at the first moment does not need to be input into the target model. For example, the bitmap data of the interface image displayed on the screen at the first moment can be directly deleted, or the bitmap data of the interface image displayed on the screen at the first moment can be ignored. The present disclosure does not limit the specific processing method of "bitmap data of the interface image displayed on the screen at the first moment", as long as the bitmap data of the interface image displayed on the screen at the first moment does not need to be input into the target model.

[0115] Step S705: If it is determined that the bitmap data of the interface image displayed on the screen at the first moment is different from the historical bitmap data marked by the target model, the bitmap data of the interface image displayed on the screen at the first moment is input into the target model. Proceed to step S706.

[0116] Here, the processing procedure after the bitmap data of the interface image displayed on the screen at the first moment is input into the target model can be referred to the above description, and will not be repeated here.

[0117] Step S706: If control data for the target vehicle is detected, the control data is input to the target model.

[0118] The target model is used to determine the target image features corresponding to the target vehicle, and to obtain the target intent features based on the control data and the target image features.

[0119] Step S707: Output the target location information of the target control pointed to by the target intent feature.

[0120] Here, the target location information refers to the position information of the target control in the interface image displayed on the display screen.

[0121] In summary, the proposed solution can acquire the latest updated interface image, enabling the target model to use the bitmap data of the latest updated interface image to perform control recognition and output the target position information of the target control. This allows for the execution of control operations on the target control in the display screen of the target vehicle. At the same time, it effectively improves the recognition capability and accuracy of "what you see is what you get," thus laying the foundation for improving the user experience.

[0122] The following provides a more detailed explanation of this disclosure with specific examples; specifically, this disclosure proposes a "what you see can be spoken" implementation scheme based on accurate non-textual (e.g., target audio) recognition using a large language model, such as... Figure 8 As shown, the specific steps are as follows:

[0123] Step S801: The target vehicle (e.g., the vehicle-mounted equipment in the target vehicle, such as the vehicle computer) takes screenshots of the display interface of the display screen at preset time intervals and uploads the obtained interface images to the cloud server.

[0124] For example, the target vehicle uploads the bitmap data of the interface image to a remote server.

[0125] Step S802: After receiving a new interface image, the cloud server compares the new interface image with historical interface images. For example, it compares the bitmap data of the new interface image with the historical bitmap data of the historical interface images marked by the large language model. After determining that the interface image has been updated, the bitmap data of the new interface image is input into the large language model.

[0126] Here, the large language model can relabel the received new interface image as the interface image displayed on the target vehicle's screen. For example, it can label the bitmap data of the received new interface image as the bitmap data of the interface image displayed on the target vehicle's screen. Furthermore, the large language model can also extract features from the bitmap data of the latest labeled interface image displayed on the target vehicle's screen, thus obtaining the target image features of the target vehicle and laying the foundation for subsequent accurate recognition.

[0127] Step S803: Target vehicle detects target audio, for example, the target audio "Click Bull Demon King" is detected.

[0128] Here, the target vehicle detects the target audio and sends the target audio to the cloud server; the target vehicle uploads the target audio to the cloud server, and the cloud server converts the target audio into text data, or the target vehicle converts the target audio into text data and uploads the text data to the cloud server.

[0129] Step S804: The target vehicle recognizes the target audio, such as "Click on the Bull Demon King", which is the "what you see can be said" intent.

[0130] It should be noted here that "Bull Demon King" is the name of the target control.

[0131] Step S805: After the target vehicle converts the target audio "Click on the Bull Demon King" into text data, it uploads it to the cloud server.

[0132] Step S806: The cloud server inputs the text data into the large language model.

[0133] Step S807: The cloud server uses the large language model to identify the image coordinates of the target control (such as "Bull Demon King") in the interface image (such as the bitmap data of the interface image) marked by the large language model.<x,y> .

[0134] Step S808: The cloud server will assign the image coordinates of the target control to the interface image (e.g., bitmap data of the interface image) marked by the large language model.<x,y> Send to the target vehicle.

[0135] Step S809: Target vehicle coordinates the received image points<x,y> The target control (such as "Bull Demon King") is processed (for example, processed according to the coordinate point transformation method described above) to obtain the screen coordinates of the target control on the display screen.<sx,sy> .

[0136] Step S810: Target vehicle based on screen coordinates<sx,sy> Click on "Bull Demon King" to complete what you see.

[0137] Thus, the "what you see can be spoken" function, which utilizes a large language model in a cloud server, can easily and efficiently achieve the function compared to existing solutions. It also supports a wider range of recognition and provides more accurate recognition results, thereby effectively improving the user experience.

[0138] This disclosure also provides a cloud server, such as Figure 9 As shown, it includes:

[0139] The processing unit 901 is configured to, upon detecting control data for a target vehicle, obtain target intent features of the control data based on the control data and target image features corresponding to the target vehicle; wherein the control data is used to instruct control operations to be performed on target controls on the display screen of the target vehicle; and the target image features characterize the image features of the interface image displayed on the display screen.

[0140] The output unit 902 is used to output the target position information of the target control pointed to by the target intent feature; wherein the target position information represents the position information of the target control in the interface image displayed on the display screen.

[0141] In a specific example of the scheme disclosed herein, the processing unit 901 is specifically used to input the control data into the target model; wherein, the target model is used to determine the target image features corresponding to the target vehicle, and to obtain target intent features based on the control data and the target image features.

[0142] In a specific example of the scheme disclosed herein, the target model is further used to mark the bitmap data of the interface image displayed on the display screen of the target vehicle, and extract the image features of the marked bitmap data to obtain the target image features corresponding to the target vehicle.

[0143] In a specific example of the scheme disclosed herein, the processing unit 901 is further configured to:

[0144] Obtain the bitmap data of the interface image displayed on the screen at the first moment;

[0145] The bitmap data of the interface image displayed on the screen at the first moment is compared with the historical bitmap data marked by the target model;

[0146] If it is determined that the bitmap data of the interface image displayed on the screen at the first moment is different from the historical bitmap data marked by the target model, the bitmap data of the interface image displayed on the screen at the first moment is input into the target model; wherein, the target model is used to update the marked historical bitmap data and update it to the bitmap data of the interface image displayed on the screen at the first moment.

[0147] In a specific example of the scheme disclosed herein, the processing unit 901 is further configured to:

[0148] If it is determined that the bitmap data of the interface image displayed on the screen at the first moment is the same as the historical bitmap data marked by the target model, the bitmap data of the interface image displayed on the screen at the first moment is masked.

[0149] In a specific example of the disclosed solution, the processing unit 901 is specifically used for:

[0150] Obtain the total coordinate information of the bitmap data of the interface image displayed on the screen at the first moment;

[0151] The total coordinate information of the bitmap data of the interface image at the first moment is compared with the total coordinate information of the historical bitmap data marked by the target model.

[0152] For a description of the specific functions and examples of each unit of the apparatus in this disclosure embodiment, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.

[0153] This disclosure also provides an in-vehicle device, such as Figure 10 As shown, it includes:

[0154] The sending unit 1001 is used to send control data in response to the target audio; wherein the control data is obtained based on the target audio and is used to instruct the target control on the display screen of the target vehicle to perform control operations.

[0155] The acquisition unit 1002 is used to acquire the target position information of the target control pointed to by the target intent feature of the control data; wherein, the target position information represents the position information of the target control in the interface image displayed on the display screen;

[0156] The control unit 1003 is used to perform control operations on the target control based on the target position information of the target control.

[0157] In a specific example of the disclosed scheme, wherein,

[0158] The acquisition unit 1002 is also used to acquire the interface images of the target vehicle's display screen at different times;

[0159] The sending unit 1001 is also used to send the interface images of the target vehicle's display screen at different times.

[0160] In a specific example of the scheme disclosed herein, the sending unit 1001 is specifically used for:

[0161] Send bitmap data of the interface images of the target vehicle's display screen at different times.

[0162] In a specific example of the disclosed solution, the control unit 1003 is specifically used for:

[0163] Based on the target position information of the target control, the screen position information of the target control in the display screen is obtained;

[0164] Based on the screen position information of the target control on the display screen, control operations are performed on the target control.

[0165] For a description of the specific functions and examples of each unit of the apparatus in this disclosure embodiment, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.

[0166] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0167] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0168] Figure 11 A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0169] like Figure 11 As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1102 or a computer program loaded from storage unit 1108 into random access memory (RAM) 1103. The RAM 1103 may also store various programs and data required for the operation of device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.

[0170] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0171] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as the control method for a target vehicle. For example, in some embodiments, the control method for a target vehicle may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the control method for a target vehicle described above may be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to execute the control method of the target vehicle by any other suitable means (e.g., by means of firmware).

[0172] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0173] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0174] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0175] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0176] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0177] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0178] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0179] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for controlling a target vehicle, applied to a cloud server, the method comprising: Upon detecting control data targeting a target vehicle, a target intent feature of the control data is obtained based on the control data and the target image features corresponding to the target vehicle; wherein, the control data is used to instruct control operations on target controls on the display screen of the target vehicle; and the target image features characterize the image features of the interface image displayed on the display screen. Output the target position information of the target control pointed to by the target intent feature; wherein, the target position information represents the position information of the target control in the interface image displayed on the display screen; The step of obtaining the target intent feature of the control data based on the control data and the target image features corresponding to the target vehicle includes: The control data is input into the target model to obtain the target intent features of the control data; The target model is used to pre-label the bitmap data of the interface image displayed on the display screen of the target vehicle, and extract the image features of the labeled bitmap data to obtain the target image features corresponding to the target vehicle. The target model is also used to, after acquiring the control data, first determine the target image features corresponding to the target vehicle, and then obtain the target intent features based on the control data and the target image features.

2. The method according to claim 1, further comprising: Obtain the bitmap data of the interface image displayed on the screen at the first moment; The bitmap data of the interface image displayed on the screen at the first moment is compared with the historical bitmap data marked by the target model; If it is determined that the bitmap data of the interface image displayed on the screen at the first moment is different from the historical bitmap data marked by the target model, the bitmap data of the interface image displayed on the screen at the first moment is input into the target model; wherein, the target model is used to update the marked historical bitmap data and update it to the bitmap data of the interface image displayed on the screen at the first moment.

3. The method according to claim 2, further comprising: If it is determined that the bitmap data of the interface image displayed on the screen at the first moment is the same as the historical bitmap data marked by the target model, the bitmap data of the interface image displayed on the screen at the first moment is masked.

4. The method according to claim 2 or 3, wherein, The step of comparing the bitmap data of the interface image displayed on the screen at the first moment with the historical bitmap data marked by the target model includes: Obtain the total coordinate information of the bitmap data of the interface image displayed on the screen at the first moment; The total coordinate information of the bitmap data of the interface image at the first moment is compared with the total coordinate information of the historical bitmap data marked by the target model.

5. A method for controlling a target vehicle, the method comprising: In response to the target audio, control data is sent; wherein the control data is obtained based on the target audio and is used to instruct control operations to be performed on the target controls on the display screen of the target vehicle. The target position information of the target control pointed to by the target intent feature of the control data is obtained; wherein the target position information represents the position information of the target control in the interface image displayed on the display screen; wherein the target intent feature is obtained based on the method of any one of claims 1 to 4; Based on the target position information of the target control, control operations are performed on the target control.

6. The method according to claim 5, further comprising: Acquire the interface images of the target vehicle's display screen at different times; Send the interface images of the target vehicle's display screen at different times.

7. The method according to claim 6, wherein, The step of sending the interface images of the target vehicle's display screen at different times includes: Send bitmap data of the interface images of the target vehicle's display screen at different times.

8. The method according to any one of claims 5-7, wherein, The control operation based on the target position information of the target control includes: Based on the target position information of the target control, the screen position information of the target control in the display screen is obtained; Based on the screen position information of the target control on the display screen, control operations are performed on the target control.

9. A cloud server, comprising: A processing unit is configured to, upon detecting control data for a target vehicle, obtain target intent features of the control data based on the control data and target image features corresponding to the target vehicle; wherein the control data is used to instruct control operations on target controls on the display screen of the target vehicle; the target image features characterize the image features of the interface image displayed on the display screen; specifically, the processing unit is configured to input the control data into a target model to obtain the target intent features of the control data; wherein the target model is configured to pre-label the bitmap data of the interface image displayed on the display screen of the target vehicle, and extract the image features of the labeled bitmap data to obtain the target image features corresponding to the target vehicle; the target model is configured to, after obtaining the control data, first determine the target image features corresponding to the target vehicle, and then obtain the target intent features based on the control data and the target image features; The output unit is used to output the target position information of the target control pointed to by the target intent feature; wherein the target position information represents the position information of the target control in the interface image displayed on the display screen.

10. The cloud server according to claim 9, wherein, The processing unit is further configured to: Obtain the bitmap data of the interface image displayed on the screen at the first moment; The bitmap data of the interface image displayed on the screen at the first moment is compared with the historical bitmap data marked by the target model; If it is determined that the bitmap data of the interface image displayed on the screen at the first moment is different from the historical bitmap data marked by the target model, the bitmap data of the interface image displayed on the screen at the first moment is input into the target model; wherein, the target model is used to update the marked historical bitmap data and update it to the bitmap data of the interface image displayed on the screen at the first moment.

11. The cloud server according to claim 10, wherein the processing unit is further configured to: If it is determined that the bitmap data of the interface image displayed on the screen at the first moment is the same as the historical bitmap data marked by the target model, the bitmap data of the interface image displayed on the screen at the first moment is masked.

12. The cloud server according to claim 10 or 11, wherein, The processing unit is specifically used for: Obtain the total coordinate information of the bitmap data of the interface image displayed on the screen at the first moment; The total coordinate information of the bitmap data of the interface image at the first moment is compared with the total coordinate information of the historical bitmap data marked by the target model.

13. A vehicle-mounted device, comprising: A sending unit is configured to send control data in response to target audio; wherein the control data is obtained based on the target audio and is used to instruct control operations to be performed on target controls on the display screen of the target vehicle. An acquisition unit is configured to acquire target position information of a target control pointed to by the target intent feature of the control data; wherein the target position information represents the position information of the target control in the interface image displayed on the display screen; wherein the target intent feature is obtained based on the method of any one of claims 1 to 4; The control unit is used to perform control operations on the target control based on the target position information of the target control.

14. The vehicle-mounted device according to claim 13, wherein, The acquisition unit is also used to acquire the interface images of the target vehicle's display screen at different times; The sending unit is also used to send the interface images of the target vehicle's display screen at different times.

15. The vehicle-mounted device according to claim 14, wherein, The transmitting unit is specifically used for: Send bitmap data of the interface images of the target vehicle's display screen at different times.

16. The vehicle-mounted device according to any one of claims 13-15, wherein, The control unit is specifically used for: Based on the target position information of the target control, the screen position information of the target control in the display screen is obtained; Based on the screen position information of the target control on the display screen, control operations are performed on the target control.

17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.

19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Voice interaction method and device and storage medium

    CN116382614A