Vehicle machine cross-device communication control system and method based on AI self-learning
Through the cross-device communication control system of vehicle-machine computers based on AI self-learning, the compatibility and single function of mobile phones and vehicle-machine interconnection is solved, and full-function communication control and dynamic adaptation interface changes are realized, improving user experience.
Patent Information
- Application Number
- CN202510364124.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, there are problems such as limited compatibility, single functions, and rigid interaction between mobile phones and car computers, and full-function communication control cannot be achieved, especially in intelligent connected cars, which cannot safely and conveniently operate the software ecological resources on the mobile phone.
The cross-device communication control system of the vehicle-machine computer based on AI self-learning is adopted, including screen projection module, AI visual analysis module, simulation operation module and voice interaction module. Through the application interface after AI self-learning mobile phone screen projection, an operation logic tree is generated to achieve full-function communication operations.
It realizes compatibility with any mobile phone model and APP that supports screen projection and touch anti-control, supports dynamic learning interface changes, can complete complex operations, improve user experience, and achieve seamless cross-device communication control.
Smart Images

Figure CN120335744A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of locomotive communication, and particularly relates to a vehicle-mounted cross-device communication control system and method based on AI self-learning. Background Art
[0002] Currently, on intelligent connected vehicles, the mobile phone can be conveniently screen-cast to the vehicle-mounted device, but it is not possible to safely and conveniently operate the software ecological resources on the mobile phone on the vehicle-mounted device.
[0003] In the prior art, the interconnection between the mobile phone and the vehicle-mounted device is mainly achieved through protocol adaptation (such as HiCar, CarLife, CarPlay and other protocols), but there are the following problems:
[0004] 1. Compatibility is limited, and it is necessary to rely on specific mobile phone brands, operating systems or communication protocols, and it cannot support all models of mobile phones and APPs; 2. The functions are single: only basic functions (such as navigation, music playback) can be realized, and it is impossible to deeply operate social software (such as sending WeChat messages, voice calls); 3. The interaction is rigid: it is necessary to preset fixed instructions or interface templates, and it is impossible to dynamically learn the interface changes of the new version of the APP.
[0005] For example, the Chinese invention patent with the publication number CN108712521A discloses a vehicle-mounted device interconnection system based on protocol adaptation, but its functions are limited by specific protocols and mobile phone brands, and full-function communication control cannot be achieved. Summary of the Invention
[0006] In order to solve the above problems, the present invention proposes a vehicle-mounted cross-device communication control system and method based on AI self-learning, a vehicle-mounted device control system that does not need to rely on specific protocols or mobile phone models, and realizes full-function communication operations by AI self-learning the application interface after the mobile phone is screen-cast.
[0007] To achieve the above object, the technical solution adopted by the present invention is:
[0008] A vehicle-mounted cross-device communication control system based on AI self-learning, including;
[0009] A screen-casting module, configured to establish a communication connection between the mobile phone and the vehicle-mounted device, and project the mobile phone screen content onto the vehicle-mounted device large screen in real time for display, and support touch reverse control of the mobile phone;
[0010] An AI vision analysis module, using large model technology, semantically analyzes the elements on the screen-casting interface, generates an interface element attribute set including element names, functions, position coordinates, and touch click operation sequence attributes, and constructs an operation logic tree according to the touch click operation sequence obtained by the analysis;
[0011] The simulation operation module generates a simulated click signal sequence according to the operation logic tree generated by the AI vision analysis module, and under the monitoring and guidance of the AI vision analysis module, simulates human operations to perform corresponding touch click operations in the projection area;
[0012] And the voice interaction module is configured to receive user voice instructions, convert the user voice instructions into an operation sequence, and call the simulation operation module and the AI vision analysis module to execute the tasks corresponding to the user instructions.
[0013] Furthermore, the projection module includes:
[0014] A communication interface unit to achieve a wired or wireless communication connection between the mobile phone and the in-vehicle computer;
[0015] And a projection display unit that transmits and displays the mobile phone screen content on the in-vehicle computer screen in real time, supports reverse touch control of the mobile phone, and simultaneously maintains the synchronous update of the projection content and the mobile phone screen content.
[0016] Furthermore, the AI vision analysis module includes:
[0017] An image acquisition unit to obtain the interface image from the projection module;
[0018] A semantic parsing unit that uses a large model to perform in-depth semantic understanding of the projection interface and extracts various attributes of the interface elements;
[0019] And a logic tree construction unit that constructs an operation logic tree using machine learning algorithms according to the touch click operation sequence extracted by the semantic parsing unit to optimize the operation path and decision-making process.
[0020] Furthermore, the simulation operation module includes:
[0021] A signal generation unit that generates simulated click signals according to the operation logic tree;
[0022] And an operation execution unit that, under the real-time monitoring of the AI vision analysis module, simulates human touch click operations in the projection area according to the generated click signal sequence.
[0023] Furthermore, the voice interaction module includes:
[0024] A voice recognition unit for recognizing and parsing user voice instructions;
[0025] An instruction conversion unit that converts the recognized voice instructions into an operable task sequence;
[0026] And a task scheduling unit that coordinates and calls the simulation operation module and the AI vision analysis module according to the converted task sequence to complete the tasks required by the user instructions.
[0027] On the other hand, the present invention also provides a vehicle-mounted device cross-device communication control method based on AI self-learning, including the steps of:
[0028] In the first step, self-learning is performed on the mobile phone screen-cast to the vehicle-mounted device; the vehicle-mounted device calls the AI module to automatically perform simulated click operations, perform self-learning on the screen content of the mobile phone screen-cast to the vehicle-mounted device and the specified APP, perform semantic parsing on each displayed interface element, understand all functions of the mobile phone and all functions of the APP, and deeply learn the general operation sequence of a function and the touch-screen operation method of each operation;
[0029] In the second step, the vehicle-mounted device helps the user achieve the intention according to the results of self-learning; when the user operates the mobile phone by voice on the vehicle-mounted device, the vehicle-mounted device AI automatically parses the user's voice intention, forms a simulated click operation sequence through the learned general operation sequence, and uses the simulated click method to perform simulated human operations in the mobile phone screen area until the user's intention is completed;
[0030] In the third step, the vehicle-mounted device AI self-evolves; during the simulated click operation process, the operation results are monitored in real time to achieve automatic error correction and self-learning of the error correction experience; or when the user's mobile phone APP is updated or the operations are different due to other reasons, the vehicle-mounted device AI automatically identifies the different operation steps and performs a self-learning correction.
[0031] Beneficial effects of adopting this technical solution:
[0032] The present invention analyzes the interface elements of the mobile phone screen-cast through AI vision analysis, generates an operation logic tree, and combines the simulated operation module to achieve full-function communication control. The present invention realizes full-function communication control after the mobile phone screen-cast through the simulated click technology, is compatible with any mobile phone models and APPs that support screen-casting and touch reverse control, supports dynamic learning of interface changes, and can be widely applied to the scenarios of intelligent vehicle and mobile phone interconnection, or other similar fields where screen-casting can be reversely controlled.
[0033] The present invention has protocol independence: it replaces traditional protocol adaptation through AI vision analysis and is compatible with any mobile phone models and APPs that support screen-casting and touch reverse control.
[0034] The present invention can perform dynamic self-learning: it uses a large model to parse interface elements in real time and adapts to APP version updates.
[0035] The present invention can cover all functions: it supports complex operations (such as group chat @ someone, sending location, switching voice / text input).
[0036] The present invention can improve the user experience: the user does not need to repeat operations, and seamless cross-device communication control is realized. Description of the Drawings
[0037] Figure 1 Schematic diagram of the structure of a vehicle-mounted cross-device communication control system based on AI self-learning of the present invention;
[0038] Figure 2 Schematic diagram of interface element recognition of the AI vision analysis module in the embodiment of the present invention;
[0039] Figure 3 Schematic diagram of signal generation and execution of the simulation operation module in the embodiment of the present invention. Detailed implementation manners
[0040] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings.
[0041] In this embodiment, as shown in Figure 1 the present invention proposes a vehicle-mounted cross-device communication control system based on AI self-learning, including; a screen mirroring module, an AI vision analysis module, a voice interaction module, and a simulation operation module.
[0042] (1) The screen mirroring module is configured to establish a communication connection between the mobile phone and the vehicle-mounted device, and project the mobile phone screen content onto the vehicle-mounted large screen in real time, and support touch reverse control of the mobile phone.
[0043] The screen mirroring module includes:
[0044] A communication interface unit to realize wired or wireless communication connection between the mobile phone and the vehicle-mounted device;
[0045] and a screen mirroring display unit to transmit and display the mobile phone screen content on the vehicle-mounted large screen in real time, support touch reverse control of the mobile phone, and keep the screen mirroring content synchronized with the mobile phone screen content.
[0046] The screen mirroring module is connected to the mobile phone by wireless or wired means, and transmits the mobile phone screen image to the vehicle-mounted large screen in real time.
[0047] (2) The AI vision analysis module, as shown in Figure 2 uses large model technology to semantically analyze the elements on the screen mirroring interface, generate an interface element attribute set including element name, function, position coordinates, and touch click operation sequence attributes, and construct an operation logic tree according to the parsed touch click operation sequence.
[0048] The AI vision analysis module includes: an image acquisition unit, a semantic analysis unit, and a logic tree construction unit.
[0049] The image acquisition unit obtains the interface image from the screen mirroring module.
[0050] The semantic parsing unit uses a large model to perform in-depth semantic understanding of the screen mirroring interface and extract various attributes of the interface elements. The semantic parsing unit uses the large model to identify interface elements (such as buttons, input boxes, menus) and touch click paths, and obtains an attribute of the interface element, which includes the name, function, position coordinates, touch click operation sequence, etc. of the element.
[0051] And the logic tree construction unit constructs an operation logic tree using machine learning algorithms based on the touch click operation sequence extracted by the semantic parsing unit to optimize the operation path and decision-making process.
[0052] (3) The simulation operation module, as Figure 3 shown, generates a simulated click signal sequence according to the operation logic tree generated by the AI vision analysis module, and under the monitoring and guidance of the AI vision analysis module, simulates human operations and performs corresponding touch click operations in the screen mirroring area.
[0053] The simulation operation module includes:
[0054] The signal generation unit generates simulated click signals according to the operation logic tree;
[0055] And the operation execution unit, under the real-time monitoring of the AI vision analysis module, simulates human touch click operations in the screen mirroring area according to the generated click signal sequence.
[0056] According to the output of the AI vision analysis module, a set of simulated click command sequences for completing user instructions is generated, and human operations are simulated in the mobile phone screen mirroring area of the vehicle using simulated clicks. The AI will monitor the operation results in real time and correct them when differences occur.
[0057] (4) The voice interaction module is configured to receive user voice instructions, convert the user voice instructions into an operation sequence, and call the simulation operation module and the AI vision analysis module to execute the tasks corresponding to the user instructions.
[0058] The voice interaction module includes:
[0059] The voice recognition unit is used to recognize and parse user voice instructions;
[0060] The instruction conversion unit converts the recognized voice instructions into an operable task sequence;
[0061] And the task scheduling unit coordinates and calls the simulation operation module and the AI vision analysis module according to the converted task sequence to complete the tasks required by the user instructions.
[0062] To cooperate with the implementation of the system of the present invention, based on the same inventive concept, the present invention also provides a vehicle-mounted device cross-device communication control method based on AI self-learning, including the steps:
[0063] In the first step, self-learning is performed on the mobile phone projected onto the vehicle-mounted device; the vehicle-mounted device calls the AI module to automatically perform simulated click operations, perform self-learning on the content of the mobile phone screen projected onto the vehicle-mounted device and the specified APP, perform semantic parsing on each displayed interface element, understand all functions of the mobile phone and all functions of the APP, and deeply learn the general operation sequence for implementing a function and the touch screen operation method for each operation.
[0064] In the second step, the vehicle-mounted device helps the user achieve the intention according to the result of self-learning; when the user operates the mobile phone by voice on the vehicle-mounted device, the vehicle-mounted device AI automatically parses the user's voice intention, forms a simulated click operation sequence through the learned general operation sequence, and uses the simulated click method to perform simulated human operations in the mobile phone screen area until the user's intention is completed.
[0065] In the third step, the vehicle-mounted device AI self-evolves; during the simulated click operation process, it monitors the operation result in real time to achieve automatic error correction and self-learns the error correction experience; or when the mobile phone APP is updated or the operation is different due to other reasons, the vehicle-mounted device AI automatically identifies the different operation steps and performs a self-learning correction.
Specific Embodiment 1
[0067] A vehicle-mounted device cross-device communication control system based on AI self-learning includes: a screen projection module, an AI visual analysis module, a voice interaction module, and a simulated operation module.
[0068] The screen projection module projects the content of the mobile phone screen onto the large screen S1 of the vehicle-mounted device in real time through the HiCar protocol.
[0069] The AI visual analysis module obtains the mobile phone screen projection interface from the screen projection module, identifies interface elements (such as the "send" button in the WeChat chat window) through the visual large model of Mianbi Intelligence, and generates an operation logic tree.
[0070] The voice interaction module converts the user's voice instruction (such as "Send a WeChat message to Li Si saying that I will arrive home at 7 pm tonight") into an operation sequence and calls the simulated operation module to execute.
[0071] The simulated operation module generates simulated click signals according to the operation logic tree, and the AI assistant executes the simulated click command sequence, clicks on the mobile phone screen projection interface area, and notifies the user after all instructions are completed.
Specific Embodiment 2
[0073] The AI visual analysis module supports dynamic learning of APP interface changes.
[0074] For example:
[0075] 1. The AI vision analysis module obtains the mobile phone screen mirroring interface from the screen mirroring module and monitors the operation results in real time. If an unexpected result occurs during the execution of the simulated click process, such as an advertisement popping up and blocking the simulated click position, the AI vision analysis module automatically identifies it, calls the simulated operation module to correct the error until the task is completed, and records and learns from the errors at the same time.
[0076] 2. After the WeChat update, the position of the "Send" button changes. The system automatically corrects the operation logic tree by comparing the historical interface data to ensure the operation accuracy.
Specific Embodiment 3
[0078] The simulated operation module supports the generation of various simulated click signals, including clicks, swipes, long presses, etc. For example, for the user voice command "Swipe to the next message of Zhang San", the simulated operation module generates a signal sequence containing a swipe click and swipes up a message in Zhang San's chat dialog box area on the mobile phone screen mirroring interface.
Specific Embodiment 4
[0080] The voice interaction module supports multi-round conversations and context understanding. For example, the user first says "Open WeChat", and then says "Send a message to Zhang San: Have a meeting tomorrow", and the system automatically executes operations such as opening WeChat, searching for contacts, inputting text, and clicking to send.
[0081] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A vehicle-mounted device cross-device communication control system based on AI self-learning, characterized in that including; a screen mirroring module configured to establish a communication connection between a mobile phone and a vehicle head unit, and to project the mobile phone screen content onto the large screen of the vehicle head unit in real time for display, supporting touch reverse control of the mobile phone; an AI vision analysis module that uses large model technology to semantically analyze the elements on the screen mirroring interface, generates a set of interface element attributes including element name, function, position coordinates, and touch click operation sequence attributes, and constructs an operation logic tree based on the parsed touch click operation sequence; a simulation operation module that generates a sequence of simulated click signals according to the operation logic tree generated by the AI vision analysis module, and under the monitoring and guidance of the AI vision analysis module, simulates human operations to perform corresponding touch click operations in the screen mirroring area; and a voice interaction module configured to receive user voice commands, convert the user voice commands into an operation sequence, and call the simulation operation module and the AI vision analysis module to execute the tasks corresponding to the user commands.
2. The vehicle-mounted device cross-device communication control system and method based on AI self-learning according to claim 1, characterized in that The screen mirroring module includes: a communication interface unit that realizes a wired or wireless communication connection between the mobile phone and the vehicle head unit; and a screen mirroring display unit that transmits and displays the mobile phone screen content on the large screen of the vehicle head unit in real time, supports touch reverse control of the mobile phone, and simultaneously maintains the synchronous update of the screen mirroring content and the mobile phone screen content.
3. A vehicle-mounted device cross-device communication control system and method based on AI self-learning according to claim 1, characterized in that The AI vision analysis module includes: an image acquisition unit that obtains the interface image from the screen mirroring module; a semantic analysis unit that uses a large model to perform in-depth semantic understanding of the screen mirroring interface and extracts various attributes of the interface elements; and a logic tree construction unit that constructs an operation logic tree using machine learning algorithms based on the touch click operation sequence extracted by the semantic analysis unit to optimize the operation path and decision-making process.
4. A vehicle-mounted device cross-device communication control system and method based on AI self-learning according to claim 1, characterized in that, The simulation operation module includes: a signal generation unit that generates simulated click signals according to the operation logic tree; and an operation execution unit that, under the real-time monitoring of the AI vision analysis module, simulates human touch click operations in the screen mirroring area according to the generated sequence of click signals.
5. A vehicle-mounted device cross-device communication control system and method based on AI self-learning according to claim 1, characterized in that, The voice interaction module includes: a voice recognition unit for recognizing and parsing user voice commands; a command conversion unit that converts the recognized voice commands into an operable task sequence; and a task scheduling unit that coordinates and calls the simulation operation module and the AI vision analysis module according to the converted task sequence to complete the tasks required by the user commands.
6. A vehicle-mounted device cross-device communication control method based on AI self-learning, characterized in that, including the steps: The first step is to perform self-learning on the mobile phone projected onto the vehicle head unit; the vehicle head unit calls the AI module to automatically perform simulated click operations to perform self-learning on the mobile phone screen content and the specified APP projected onto the vehicle head unit, semantically analyze each displayed interface element, understand all the functions of the mobile phone and all the functions of the APP, and deeply learn the general operation sequence to implement a function and the touch screen operation method for each operation; The second step is for the vehicle head unit to help the user achieve the intention according to the results of self-learning; When the user operates the mobile phone by voice on the vehicle head unit, the vehicle head unit AI automatically analyzes the user's voice intention, forms a simulated click operation sequence through the learned general operation sequence, and uses the simulated click method to perform simulated human operations in the mobile phone screen area until the user's intention is completed; Step 3: The in-vehicle AI self-evolves; during the simulated click operation process, it monitors the operation results in real time to achieve automatic error correction and self-learns from the error correction experience; or when the user's mobile phone APP is updated or the operations are different due to other reasons, the in-vehicle AI automatically identifies the different operation steps and makes a self-learning correction once.
Citation Information
Patent Citations
System, method and equipment for configuring node information of equipment, and readable storage medium
CN108712521A
Cited By
Voice interaction method and device, and storage medium
CN121708916A