Intelligent glasses control method and device, computing equipment and system
By converting and encrypting voice commands in AR glasses and combining them with cloud-based multimodal large-scale model analysis interface screenshots, the response delay and privacy risks of AR glasses are resolved, the intelligence and environmental adaptability are improved, and more efficient user intent understanding and operation are achieved.
Patent Information
- Application Number
- CN202511341972.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Existing AR glasses control solutions have problems such as high response delay, high privacy risk, low intelligence and poor environmental adaptability, especially the insufficient level of interactive intelligence in complex scenarios.
Smart glasses convert voice commands into target text commands, encrypt them, and send them to the cloud server. The cloud-based multimodal large model analyzes the target interface screenshots and text commands to determine the operation steps required for the user's intention. The results are then encrypted and fed back to the glasses for execution, reducing network latency and protecting privacy.
It reduces the interaction delay between AR glasses and cloud servers, improves intelligence and environmental adaptability, enhances the understanding of user intentions and the accuracy of operations, and reduces the risk of privacy leakage.
Smart Images

Figure CN120832683A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video processing, in particular to a control method and device of intelligent glasses, a computing device and system. BACKGROUND
[0002] With the increasing richness of functions of AR (Augmented Reality, augmented reality) wearable devices and the growing demand of users, how to realize the automatic control of AR intelligent glasses through intelligent means has become a hot issue in the industry. For example, existing AR glasses usually execute operations through fixed voice instructions or simple gestures, and some remote collaboration applications and translation functions also begin to appear on AR platforms. However, the existing AR glasses control scheme still has many deficiencies in actual application. For example, many existing schemes need to upload the voice instructions or camera pictures collected by the AR glasses to the cloud server for processing, and the network round-trip communication and cloud computing often cause significant response delay, which cannot meet the real-time requirements of AR scenarios. The existing technology often needs to upload the sensitive data such as the voice of the user and the seen pictures to the cloud for analysis and processing, and once the transmission and storage are improper, it may cause the risk of privacy leakage. The current interaction mode of many AR devices is based on a pre-set fixed instruction set or simple logic, and the device cannot truly "understand" the user's intention and the content of the environment. The traditional voice assistant usually only executes limited functions according to the voice keywords, and lacks the use of camera visual information, so the decision-making intelligence level is limited. The system is difficult to process the free expression instructions and the associated understanding of the visual scene in complex scenarios, and the intelligent degree of interaction is not high. The existing scheme is often customized for specific applications or situations, and lacks universality. Therefore, how to reduce the response delay of the AR glasses, protect the privacy of the AR glasses users, improve the intelligent degree of the AR glasses, and improve the environmental adaptability of the AR glasses are problems to be solved. SUMMARY
[0003] Therefore, it is necessary to provide a control method, device, computing device and system of intelligent glasses capable of reducing the response delay of the AR glasses, protecting the privacy of the AR glasses users, improving the intelligent degree of the AR glasses, and improving the environmental adaptability of the AR glasses in view of the above technical problems.
[0004] In a first aspect, the present application provides a control method of intelligent glasses, the control method of intelligent glasses being realized by an intelligent glasses control system, the intelligent glasses control system comprising an intelligent glasses and a cloud server, and the method comprising:
[0005] The intelligent glasses acquire a voice instruction issued by a user, and convert the voice instruction into a target text instruction;
[0006] The smart glasses determine a target interface screenshot, encrypt the target interface screenshot and the target text instruction, determine encrypted transmission information, and send the encrypted transmission information to a cloud server;
[0007] The cloud server decrypts the encrypted transmission information, determines the target interface screenshot and the target text instruction, and determines operation steps required to achieve a user intent according to the target interface screenshot and the target text instruction through a cloud multi-modal large model;
[0008] The cloud server assembles the operation steps into a standardized instruction sequence and sends the standardized instruction sequence to the smart glasses;
[0009] The smart glasses determine instruction operation information according to the standardized instruction sequence and execute the instruction operation information.
[0010] In one of the embodiments, the training method of the cloud multi-modal large model comprises:
[0011] The cloud server combines a pre-trained visual language model, a text perceiver, a graph perceiver, and a space perceiver to determine a multi-modal fusion model;
[0012] The multi-modal fusion model is trained using a sample interface screenshot of the smart glasses and a sample description text corresponding to the sample interface screenshot to determine a first to-be-trained multi-modal model;
[0013] The first to-be-trained multi-modal model is trained according to the sample interface screenshot, sample annotation information of the sample interface screenshot, and a sample interactive element corresponding to the sample interface screenshot to determine a second to-be-trained multi-modal model; the sample annotation information includes a sample prompt word, a sample icon, and a sample virtual control in the sample interface screenshot;
[0014] The second to-be-trained multi-modal model is trained according to the sample interface screenshot, a sample instruction sequence, and a sample text instruction corresponding to the sample interface screenshot to determine a cloud multi-modal large model.
[0015] In one of the embodiments, the second to-be-trained multi-modal model is trained according to the sample interface screenshot, a sample instruction sequence, and a sample text instruction corresponding to the sample interface screenshot to determine a cloud multi-modal large model, which comprises:
[0016] The second to-be-trained multi-modal model is trained according to the sample interface screenshot, a sample instruction sequence, and a sample text instruction corresponding to the sample interface screenshot to determine a candidate multi-modal large model;
[0017] The candidate multi-modal large model is reinforced training by a group relative strategy optimization algorithm to determine the cloud multi-modal large model.
[0018] In one of the embodiments, the cloud multi-modal large model is determined by reinforcing training the candidate multi-modal large model by a group relative strategy optimization algorithm, including:
[0019] The model test task is executed by the candidate multi-modal large model, a reward function is determined according to a task execution result of the candidate multi-modal large model executing the model test task by using the group relative strategy optimization algorithm;
[0020] A model reward score is determined based on the reward function by the group relative strategy optimization algorithm, and a model parameter of the candidate multi-modal large model is optimized based on the model reward score to determine the cloud multi-modal large model.
[0021] In one of the embodiments, the operation steps required to achieve the user's intention are determined by the cloud multi-modal large model based on the target interface screenshot and the target text instruction, including:
[0022] The target interface screenshot and the target text instruction are input into the cloud multi-modal large model, and a target prompt word of the target interface screenshot is extracted by a text perceiver in the cloud multi-modal large model.
[0023] A target icon and a target virtual control in the target interface screenshot are determined by a graphic perceiver in the cloud multi-modal large model.
[0024] A target spatial relationship between the target icon and the target virtual control is determined by a spatial perceiver in the cloud multi-modal large model.
[0025] The target prompt word, the target icon and the target virtual control are taken as target interactive elements, and the operation steps required to achieve the user's intention are determined by a visual language model in the cloud multi-modal large model based on the target spatial relationship, the target interactive elements and the target text instruction.
[0026] In one of the embodiments, the control method of the smart glasses further includes:
[0027] The smart glasses determine instruction correction information by a dynamic programming algorithm based on the target text instruction, historical task execution data of a subtask that has been executed in the instruction operation information execution process, and a real-time interface screenshot of the smart glasses after executing the subtask when executing the instruction operation information.
[0028] From the standardized instruction sequence, determine the to-be-adjusted instruction sequence corresponding to the unexecuted subtask in the instruction operation information;
[0029] Based on the instruction correction information, adjust the to-be-adjusted instruction sequence.
[0030] In one of the embodiments, the control method of the smart glasses further comprises:
[0031] After the smart glasses execute the instruction operation information, the smart glasses generate operation result feedback information, and send the operation result feedback information, intermediate data generated when the instruction operation information is executed, operation interface screenshots corresponding to the intermediate data, and task completion interface screenshots of the smart glasses after the execution of the instruction operation information to the cloud server;
[0032] The cloud server re-trains the cloud multi-modal large model according to the operation result feedback information, the intermediate data, the operation interface screenshots, and the task completion interface screenshots.
[0033] In a second aspect, the application further provides a control device of smart glasses, the device comprising:
[0034] The voice conversion module is deployed on the smart glasses, and is configured to obtain voice instructions issued by a user and convert the voice instructions into target text instructions;
[0035] The interface screenshot acquisition module is deployed on the smart glasses, and is configured to determine target interface screenshots, encrypt the target interface screenshots and the target text instructions, determine encrypted transmission information, and send the encrypted transmission information to the cloud server;
[0036] The cloud large model inference module is deployed on the cloud server, and is configured to decrypt the encrypted transmission information, determine the target interface screenshots and the target text instructions, and determine operation steps required to achieve a user's intention according to the target interface screenshots and the target text instructions through a cloud multi-modal large model;
[0037] The action instruction analysis module is deployed on the cloud server, and is configured to assemble the operation steps into a standardized instruction sequence, and send the standardized instruction sequence to the smart glasses;
[0038] The instruction execution module is deployed on the smart glasses, and is configured to determine instruction operation information according to the standardized instruction sequence, and execute the instruction operation information.
[0039] In a third aspect, the application further provides a computing device comprising a memory and a processor, the memory storing a computer program, when the computing device is a smart glasses, the following steps are implemented: obtaining a voice instruction issued by a user, and converting the voice instruction into a target text instruction; determining a target interface screenshot, and encrypting the target interface screenshot and the target text instruction, determining encrypted transmission information, and sending the encrypted transmission information to a cloud server; determining instruction operation information according to a standardized instruction sequence fed back by the cloud server, and executing the instruction operation information;
[0040] When the computing device is a cloud server, the following steps are implemented: decrypting the encrypted transmission information, determining the target interface screenshot and the target text instruction, and determining operation steps required to achieve a user's intention according to the target interface screenshot and the target text instruction through a cloud multi-modal large model; assembling the operation steps into a standardized instruction sequence, and sending the standardized instruction sequence to the smart glasses.
[0041] In a fourth aspect, the application further provides a control system of smart glasses, the control system of smart glasses comprising smart glasses and a cloud server.
[0042] The control method, device, computing device and system of the above-mentioned smart glasses are as follows: the smart glasses obtain the voice commands issued by the user and convert the voice commands into target text commands; the smart glasses determine the target interface screenshot, encrypt the target interface screenshot and the target text command, determine the encrypted transmission information, and send the encrypted transmission information to the cloud server; the cloud server decrypts the encrypted transmission information, determines the target interface screenshot and the target text command, and determines the operation steps required to realize the user's intention based on the target interface screenshot and the target text command through the cloud multimodal large model; the cloud server assembles the operation steps into a standardized instruction sequence and sends the standardized instruction sequence to the smart glasses; the smart glasses determine the instruction operation information based on the standardized instruction sequence and execute the instruction operation information. This solves the problems of high response delay, high privacy risk, low intelligence and poor adaptability of the current AR device automation control method. The above solution builds an intelligent system that closely collaborates with an AR glasses device and a cloud server. The AR glasses use voice recognition technology to rapidly process user voice commands, converting them into target text commands and uploading them to the cloud server. This eliminates the need to upload audio to the cloud server, reducing interaction latency between the AR glasses and the cloud server. The target text commands and screenshots of the target interface are encrypted and sent to the cloud server, protecting user privacy. Both the target text commands and screenshots serve as input data for a multimodal large-scale model, which intuitively reflects the real-time environment and virtual interface state of the AR glasses during command execution. The cloud-based multimodal large-scale model deployed on the cloud server analyzes the screenshots and text commands to determine the steps required to fulfill the user's intent, enhancing the intelligent control of the AR glasses. By determining the steps required to fulfill the user's intent using the multimodal large-scale model, the target interface screenshots and text commands can be analyzed based on the large-scale model's extensive training data, thereby improving the environmental adaptability of the AR glasses control method. By assembling the steps into a standardized command sequence and feeding it back to the AR glasses, interaction latency between the cloud server and the AR glasses can be further reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 1 is a flow chart of a method for controlling smart glasses according to an embodiment;
[0044] Figure 2 1 is a flow chart of a method for training a large multimodal model in the cloud in one embodiment;
[0045] Figure 3 A schematic flow chart of a method for determining an operation step in one embodiment;
[0046] Figure 4 is a structural block diagram of a control device for smart glasses in one embodiment;
[0047] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0049] In one embodiment, Figure 1 As shown, a control method for smart glasses is provided. The control method for smart glasses is implemented by a smart glasses control system, which includes smart glasses and a cloud server. In this embodiment, the method includes the following steps:
[0050] S110. The smart glasses obtain a voice command issued by the user and convert the voice command into a target text command.
[0051] Smart glasses, also known as AR glasses, are used as intelligent terminal devices. AR glasses are wearable computing devices equipped with a camera, microphone, display module, and local computing chip. AR glasses can overlay digital content on the user's real-world view, creating an interactive experience that blends virtual and real life. The display module can be a micro-projector or optical waveguide lens.
[0052] Specifically, the user's voice instructions input through the microphone of the AR glasses are obtained, and the voice instructions are converted into instructions in text format, namely, target text instructions, through voice recognition and audio coding technology. The method of converting voice instructions into target text instructions may include: voice instruction preprocessing, that is, performing noise reduction processing and rebound elimination processing on the voice instructions to improve the audio quality of the voice instructions, thereby improving the accuracy of voice recognition. Feature extraction is performed on the preprocessed voice instructions to determine the voice features. The voice features can be the Mel-frequency cepstral coefficients of the preprocessed voice instructions, or the linear prediction coefficients of the preprocessed voice instructions. The target text instructions are determined according to the voice features through a speech recognition algorithm. The speech recognition algorithm can be a pretrained deep learning algorithm or a pretrained convolutional neural network.
[0053] For example, the user says to the glasses "Connect to my supervisor to request remote assistance", and the AR glasses can recognize the voice command issued by the user as the text "Connect to my supervisor for remote assistance" through the voice recognition module.
[0054] It should be noted that converting voice commands into target text commands and then transmitting them to the cloud server can improve the efficiency of command transmission and reduce the interaction delay between smart glasses and the cloud server.
[0055] In S120, the smart glasses determine the target interface screenshot, encrypt the target interface screenshot and the target text instruction, determine the encrypted transmission information, and send the encrypted transmission information to the cloud server.
[0056] Specifically, the AR glasses capture the environment picture around the user through the camera, combine the environment picture with the virtual interface elements of the AR glasses, and determine the target interface screenshot. The target interface screenshot and the target text instruction are encrypted, the encrypted transmission information is determined, and the encrypted transmission information is sent to the cloud server.
[0057] Illustratively, the key server can generate a key pair according to the protocol of the AR glasses and the cloud server in advance, distribute the public key in the key pair to the AR glasses, and assign the private key in the key pair to the cloud server. The AR glasses encrypt the target interface screenshot and the target text instruction using the public key to determine the encrypted transmission information. After the cloud server obtains the encrypted transmission information, the cloud server can decrypt the encrypted transmission information using the private key to obtain the target interface screenshot and the target text instruction.
[0058] Illustratively, the AR glasses capture the picture image seen by the user through the AR glasses as the target interface screenshot. If the AR glasses are in the function menu virtual interface when the screenshot is taken, and the virtual interface contains a remote collaboration application icon, or the user can see the user interface prompt area, then the target interface screenshot contains the user interface prompt area, the remote collaboration application icon and the function menu virtual interface, and the target interface screenshot contains the virtual control layout in the function menu virtual interface and the key information of the scene around the user.
[0059] It should be noted that the encryption cost of directly encrypting the voice instruction is high, so when directly controlling the AR glasses based on the voice instruction, the voice instruction is not encrypted, but is directly sent to the cloud server, which has a risk of privacy leakage. The encryption cost of encrypting the target text instruction is less than that of directly encrypting the voice instruction, and the encryption time of encrypting the target text instruction is also less than that of directly encrypting the voice instruction. Therefore, encrypting and transmitting the target text instruction can further reduce the interaction delay between the smart glasses and the cloud server.
[0060] In S130, the cloud server decrypts the encrypted transmission information to determine the target interface screenshot and the target text instruction, and determines the operation steps needed to achieve the user's intention through the cloud multi-modal large model according to the target interface screenshot and the target text instruction.
[0061] The operation step required to realize the user's intention refers to an operation step determined to be completed by the AR glasses according to the target text instruction and the target interface screenshot. For example, if the target text instruction is clicking, the operation step required to realize the user's intention is to click the virtual control corresponding to the clicking operation on the virtual interface of the AR glasses once.
[0062] In S140, the cloud server assembles the operation steps into a standardized instruction sequence and sends the standardized instruction sequence to the smart glasses.
[0063] The standardized instruction sequence is a standardized JSON instruction sequence. The standardized instruction sequence defines a series of standardized operations to be performed in sequence, such as clicking a virtual button, sliding a view, inputting text, waiting, and information prompting. A normalized relative coordinate system ranging from 0 to 1000 is used to represent the target position in the virtual interface of the AR glasses or the user's field of view that needs to be operated.
[0064] Specifically, after determining the operation steps, the cloud server converts the data corresponding to the operation steps into a JSON-formatted string, thereby determining the standardized instruction sequence. The cloud server then sends the standardized instruction sequence to the smart glasses.
[0065] In S150, the smart glasses determine instruction operation information according to the standardized instruction sequence and execute the instruction operation information.
[0066] The instruction operation information includes the instruction operation to be executed and the target position corresponding to the instruction operation to be executed. The target position is located in the virtual interface of the AR glasses.
[0067] Specifically, the smart glasses are deployed with an action instruction parsing module for parsing the JSON-formatted standardized instruction sequence. The action instruction parsing module of the smart glasses determines the instruction operation information according to the standardized instruction sequence and calls the virtual control interface or the graphical interaction control layer of the AR glasses system in sequence to execute the corresponding instruction operation information, thereby converting the JSON-formatted standardized instruction sequence into actual interface interaction behavior.
[0068] It should be noted that the virtual control interface is a technology used to define and operate areas in the interface that cannot be recognized through conventional methods in automated testing. By encapsulating specific interface areas as virtual controls, operations such as clicking, text recognition, and image comparison can be achieved, which is suitable for scenarios such as custom buttons and complex interface elements. The graphical interaction control layer is the controller layer in the MVC (Model-View-Controller) architecture, responsible for handling the interaction logic between the user and the graphical interface. It receives user operation instructions and calls the model layer to process data, and finally updates the view layer to display the results.
[0069] The control method of the smart glasses includes the following steps: the smart glasses acquire a voice instruction issued by a user and convert the voice instruction into a target text instruction; the smart glasses determine a target interface screenshot and encrypt the target interface screenshot and the target text instruction, determine encrypted transmission information, and send the encrypted transmission information to a cloud server; the cloud server decrypts the encrypted transmission information, determines the target interface screenshot and the target text instruction, and determines an operation step required to achieve a user intent according to the target interface screenshot and the target text instruction through a cloud multimodal large model; the cloud server assembles the operation step into a standardized instruction sequence and sends the standardized instruction sequence to the smart glasses; and the smart glasses determine instruction operation information according to the standardized instruction sequence and execute the instruction operation information. The current AR device automatic control method has the problems of high response delay, high privacy risk, low intelligent degree, and poor adaptability. The above scheme constructs an intelligent agent system in which an AR glasses device and a cloud server are closely coordinated. The AR glasses quickly process a voice instruction issued by a user through voice recognition technology, convert the voice instruction into a target text instruction, and upload the target text instruction to the cloud server without uploading audio to the cloud server, which can reduce the interaction delay between the AR glasses and the cloud server. The target text instruction and the target interface screenshot are encrypted and sent to the cloud server, which can protect the user privacy. The target text instruction and the target interface screenshot are used as input data of the multimodal large model, which can intuitively reflect the real-time environment and the virtual interface state of the AR glasses when the instruction is executed. The cloud multimodal large model deployed on the cloud server analyzes the target interface screenshot and the target text instruction to determine the operation step required to achieve the user intent, which can improve the intelligent degree of the AR glasses control method. The operation step required to achieve the user intent is determined through the multimodal large model, which can analyze the target interface screenshot and the target text instruction based on the massive training data of the large model, thereby improving the environmental adaptability of the AR glasses control method. The operation step is assembled into a standardized instruction sequence and then fed back to the AR glasses, which can further reduce the interaction delay between the cloud server and the AR glasses.
[0070] In one embodiment, as shown in FIG. 1, Figure 2 The training method of the cloud multimodal large model includes the following steps:
[0071] In S210, the cloud server combines a pre-trained visual language model, a text perceiver, a graph perceiver, and a space perceiver to determine a multimodal fusion model.
[0072] The visual language model includes a large language model and a visual encoder. The large language model is a model trained based on deep learning technology, which learns the rules of human language by analyzing massive amounts of text data, and can generate or understand natural language text. The visual encoder is a core component that converts images or videos into numerical data understandable by machines, mainly used to connect image and text understanding, enabling the language model to recognize topics, colors, locations, and other features in images, and perform complex reasoning and interaction. The text sensor is a sensor model used in the field of artificial intelligence to process text data, and is one of the sensory organs of artificial intelligence. It can understand and process text information, extract key features by analyzing text content, and provide data support for subsequent decision-making or tasks. The graph sensor is a basic binary classification model in the field of machine learning, which simulates the structure of biological neurons to make linear classification decisions. It classifies input features by adjusting weights and thresholds, and is mainly used to handle linearly separable problems. The spatial sensor is a fusion technology combining spatial perception and sensor concepts, mainly referring to a technology component that acquires environmental data through sensors and converts it into usable information. The spatial sensor needs to have three-dimensional spatial environment perception capability, including positioning its own location, recognizing surrounding objects, constructing an environmental model, and planning a path of action. Its core is to acquire environmental data through sensors and generate executable instructions through algorithm processing. The constructed multi-modal fusion model has the ability to process target interface screenshots and target text instructions.
[0073] S220, using a sample interface screenshot of smart glasses and a sample description text corresponding to the sample interface screenshot, model training is performed on the multi-modal fusion model to determine a first to-be-trained multi-modal model.
[0074] The sample interface screenshot can be a virtual interface screenshot generated by the AR glasses during historical operation. The sample description text corresponding to the sample interface screenshot can be picture description information added by a person according to the sample interface screenshot. For example, the sample description text can include UI element description and interface structure description of the virtual interface in the sample interface screenshot. The first to-be-trained multi-modal model can determine the image description text corresponding to the input image according to the input image. The UI element refers to a visual element used in interface design, for example, it can be a five-point star representing collection in the virtual interface.
[0075] It should be noted that the sample interface screenshot of the smart glasses and the sample description text corresponding to the sample interface screenshot are used to train the multi-modal fusion model, and the trained multi-modal fusion model is used as the first to-be-trained multi-modal model. The visual language model in the multi-modal fusion model can be trained, so that the first to-be-trained multi-modal model has the ability to extract visual content in the graphical user interface, and at the same time, the first to-be-trained multi-modal model has the ability to understand the semantics of the image.
[0076] S230, according to the sample interface screenshot, the sample annotation information of the sample interface screenshot, and the sample interaction element corresponding to the sample interface screenshot, the model training is performed on the to-be-trained multi-modal model to determine a second to-be-trained multi-modal model.
[0077] The sample annotation information includes a sample prompt word, a sample icon, a sample virtual control, position coordinates of the sample icon, position coordinates of the sample virtual control, and position coordinates of the sample prompt word in the sample interface screenshot. The sample interaction element refers to an operation function in the sample interface screenshot, for example, the sample interaction element includes but is not limited to: up sliding, down sliding, and clicking. The sample prompt word refers to a prompt word in the sample interface screenshot, which is an instruction carrier used when a user interacts with an augmented reality system, and guides the system to feed back operation instructions or information by structuring the task requirements. The sample icon can include a UI element in the sample interface screenshot, and can also include a road icon in the environment around the user. The sample prompt word in the sample interface screenshot can be recognized by an optical character recognition technology.
[0078] Specifically, the first to-be-trained multi-modal model is trained according to the sample interface screenshot of the intelligent glasses and the sample annotation information of the sample interface screenshot, so that the trained first to-be-trained multi-modal model has the ability to locate and recognize each virtual control, icon, and prompt word in the virtual interface of the AR glasses. The trained first to-be-trained multi-modal model is further trained according to the sample interface screenshot and the sample interaction element corresponding to the sample interface screenshot to determine the second to-be-trained multi-modal model, so that the second to-be-trained multi-modal model has the recognition ability of the elements in the graphical user interface of the AR glasses.
[0079] S240, according to the sample interface screenshot, the sample instruction sequence, and the sample text instruction corresponding to the sample interface screenshot, the model training is performed on the second to-be-trained multi-modal model to determine a cloud multi-modal large model.
[0080] Specifically, according to the sample interface screenshot, the sample instruction sequence, and the sample text instruction corresponding to the sample interface screenshot, the second to-be-trained multi-modal model is subjected to supervised model training, wherein the sample instruction sequence is the supervision data. The sample instruction sequence can be a standardized instruction in JSON format generated in the historical AR glasses control process collected manually, and the sample text instruction can be a text instruction converted from the user instruction issued by the user in the historical AR glasses control process collected manually. The sample interface screenshot, the sample instruction sequence, and the sample text instruction corresponding to the sample interface screenshot can be taken as a sample data set, the sample data set is divided into a model training data set and a model test data set, the model training data set is used to train the second to-be-trained multi-modal model, the trained second to-be-trained multi-modal model is subjected to model test, a prediction instruction sequence output by the trained second to-be-trained multi-modal model in the model test process and the sample instruction sequence in the model test data set are used to calculate a model loss function, and the model loss function can be a cross-entropy loss function. Whether the second to-be-trained multi-modal model is trained is determined according to the model loss function, if the training is completed, the trained second to-be-trained multi-modal model is taken as the cloud multi-modal large model, and if the training is not completed, the second to-be-trained multi-modal model is subjected to iterative training to update the model parameters of the second to-be-trained multi-modal model.
[0081] The above scheme provides a method for training a multi-modal fusion model and determining a cloud multi-modal large model, which can enable the cloud multi-modal large model obtained through training to have the ability to locate and identify various virtual controls, icons, and prompt words in the virtual interface of AR glasses, the ability to identify elements in the graphical user interface of AR glasses, the ability to analyze complex virtual interfaces containing real scenes, and the ability to plan a task based on the virtual interface of AR glasses and target text instructions, thereby providing accurate semantic description information for subsequent decision-making.
[0082] In one embodiment, the model training of the second to-be-trained multi-modal model and the determination of the cloud multi-modal large model according to the sample interface screenshot, the sample instruction sequence, and the sample text instruction corresponding to the sample interface screenshot include:
[0083] The model training of the second to-be-trained multi-modal model and the determination of the cloud multi-modal large model according to the sample interface screenshot, the sample instruction sequence, and the sample text instruction corresponding to the sample interface screenshot include:
[0084] The group relative strategy optimization is a reinforcement learning algorithm aiming to improve the performance of a large language model in a complex reasoning task.
[0085] The scheme can further improve the decision robustness and long-range planning capability of the cloud multi-modal large model in unknown scenarios.
[0086] In one embodiment, the cloud multi-modal large model is determined by reinforcing training of the candidate multi-modal large model through a group relative strategy optimization algorithm, comprising:
[0087] The model test task is performed by the candidate multi-modal large model, and the reward function is determined according to the task execution result of the candidate multi-modal large model performing the model test task by using the group relative strategy optimization algorithm. The model reward score is determined based on the reward function by using the group relative strategy optimization algorithm, and the model parameters of the candidate multi-modal large model are optimized based on the model reward score to determine the cloud multi-modal large model.
[0088] Specifically, the candidate multi-modal large model is determined by performing model training on the second to-be-trained multi-modal model according to the sample interface screenshot, the sample instruction sequence, and the sample text instruction corresponding to the sample interface screenshot. The test task of automatically controlling the AR glasses is performed by using the candidate multi-modal large model, and in the execution process of the test task, the candidate multi-modal large model can generate a plurality of different test action sequences according to the test interface screenshot and the test text instruction corresponding to the test task. The test action sequence is selected according to the reward function fed back after each test action sequence is executed by using the group relative strategy optimization algorithm, and the test action sequence with a reward function greater than a preset function threshold is determined as a strategy feedback sequence, and the candidate multi-modal large model is trained by using the strategy feedback sequence, the test interface screenshot corresponding to the strategy feedback sequence, and the test text instruction to determine the cloud multi-modal large model.
[0089] The reward function can include an action format reward function, an action type reward function, an action parameter reward function, and a task completion reward function. For example, the determination method of the action format reward function is that if the output JSON format test instruction sequence does not conform to the specification, -1 reward is immediately assigned, and if it conforms to the specification, +1 reward is assigned. The determination method of the action type reward function is that if the execution operation falls within the true value box when the test instruction is executed, +1 reward is assigned, and if the execution operation does not fall within the true value box, -1 reward is assigned. The test instruction refers to the instruction operation information determined by the AR glasses according to the test instruction sequence. The determination method of the action parameter reward function is that if the execution operation correctly hits the target control when the test instruction is executed, +1 reward is assigned, and if the execution operation fails to hit the target control when the test instruction is executed, -1 reward is assigned. The determination method of the task completion reward function is that if the instruction execution result is correct after the test instruction is executed, +1 reward is assigned, and if the instruction execution result is incorrect after the test instruction is executed, -1 reward is assigned.
[0090] The above scheme not only consolidates the correct operation mode learned previously, but also learns to improve the strategy through "trial and error-feedback-adjustment" when facing complex or never-seen interface conditions, showing stronger adaptive decision-making ability.
[0091] In one embodiment, as shown in Figure 3 The cloud multi-modal large model determines the operation steps required to achieve the user's intention according to the target interface screenshot and the target text instruction, including:
[0092] S310, input the target interface screenshot and the target text instruction into the cloud multi-modal large model, and extract the target prompt word of the target interface screenshot through the text perceiver in the cloud multi-modal large model.
[0093] S320, determine the target icon and the target virtual control in the target interface screenshot through the graph perceiver in the cloud multi-modal large model.
[0094] S330, determine the target spatial relationship between the target icon and the target virtual control through the spatial perceiver in the cloud multi-modal large model.
[0095] S340, take the target prompt word, the target icon and the target virtual control as the target interaction element, and determine the operation steps required to achieve the user's intention according to the target spatial relationship, the target interaction element and the target text instruction through the visual language model in the cloud multi-modal large model.
[0096] The above scheme uses the reasoning ability of the visual language model to generate operation steps required to achieve the user's intention according to the target spatial relationship, the target interaction element and the target text instruction. It can improve the reliability of the determined operation steps.
[0097] In one embodiment, the control method of the smart glasses further includes:
[0098] When executing the instruction operation information, the smart glasses determine the instruction correction information according to the target task instruction, the historical task execution data of the sub-tasks that have been executed in the execution process of the instruction operation information, and the real-time interface screenshot of the smart glasses after executing the sub-tasks through the dynamic programming algorithm; determine the to-be-adjusted instruction sequence corresponding to the unexecuted sub-tasks in the instruction operation information from the standardized instruction sequence; and adjust the to-be-adjusted instruction sequence based on the instruction correction information.
[0099] It can be understood that the standardized instruction sequence refers to a sequence composed of multiple instructions, and each instruction in the standardized instruction sequence can be regarded as a subtask of the standardized instruction sequence. The historical execution data refers to the data generated in the task execution process of the subtasks in the standardized instruction sequence that has been executed, which can be recorded in a log file and obtained by reading the log file. The real-time interface screenshot of the smart glasses after executing the subtask refers to the screenshot of the virtual interface of the smart glasses obtained after executing each subtask, and the execution of the subtask can be understood through the real-time interface screenshot. The dynamic programming algorithm can generate instruction correction information according to the target task instruction, the historical task execution data, and the real-time interface screenshot.
[0100] The above scheme can dynamically correct the standardized instruction sequence according to the execution result of the instruction operation information when the instruction operation information is executed, thereby improving the adaptability of the method and the completion rate of the complex task.
[0101] In one embodiment, the control method of the smart glasses further comprises:
[0102] The smart glasses generate operation result feedback information after the execution of the instruction operation information ends, and send the operation result feedback information, intermediate data generated when the instruction operation information is executed, an operation interface screenshot corresponding to the intermediate data, and a task completion interface screenshot of the smart glasses after the execution of the instruction operation information ends to the cloud server; and the cloud server re-trains the cloud multi-modal large model according to the operation result feedback information, the intermediate data, the operation interface screenshot, and the task completion interface screenshot.
[0103] For example, the feedback mode of the operation result feedback information can be to broadcast "You have been connected to the supervisor" through voice, or to display a notification in the field of view to inform the user that the task is successfully completed. The feedback mode of the operation result feedback information can be to broadcast "You have been connected to the supervisor" through voice, or to display a notification in the field of view to inform the user that the task is successfully completed. If the execution result of a step in the operation execution process fails to achieve the expected effect, for example, the target application fails to start successfully or the specified contact person is not found, the task result feedback module can capture abnormal information and timely feedback the reason for the operation execution failure to the user through voice or text, and give corresponding suggestion operations. It can also re-capture the virtual interface screenshot of the AR glasses and send the virtual state screenshot and the target text instruction to the cloud server again, and the cloud server re-determines the operation steps through the cloud multi-modal large model to ensure that the task can be completed smoothly.
[0104] The above scheme can feed back the operation result to the user after the instruction operation information execution is completed, and feed back the intermediate data, operation interface screenshot and task completion interface screenshot according to the operation result, so as to retrain the cloud multi-modal large model, continuously optimize the cloud multi-modal large model, and ensure that the task can be completed smoothly.
[0105] For example, on the basis of the above embodiment, the control method of the intelligent glasses comprises:
[0106] The voice instruction input by the user based on the microphone of the AR glasses is acquired, and the voice instruction is converted into a text format instruction, i.e., a target text instruction, through voice recognition and audio coding technology. The AR glasses acquire the environment picture around the user through the camera, combine the environment picture with the virtual interface element of the AR glasses, and determine a target interface screenshot. The target interface screenshot and the target text instruction are encrypted, encrypted transmission information is determined, and the encrypted transmission information is sent to the cloud server.
[0107] The cloud server combines the pre-trained visual language model, the text perceiver, the graph perceiver and the space perceiver to determine a multi-modal fusion model. The sample interface screenshot of the intelligent glasses and the sample description text corresponding to the sample interface screenshot are used to train the multi-modal fusion model to determine a first to-be-trained multi-modal model. The first to-be-trained multi-modal model is trained according to the sample interface screenshot of the intelligent glasses and the sample annotation information of the sample interface screenshot, so that the trained first to-be-trained multi-modal model has the ability to locate and identify each virtual control, icon and prompt word in the virtual interface of the AR glasses. The trained first to-be-trained multi-modal model is further trained according to the sample interface screenshot and the sample interaction element corresponding to the sample interface screenshot to determine a second to-be-trained multi-modal model.
[0108] According to the sample interface screenshot, the sample instruction sequence, and the sample text instruction corresponding to the sample interface screenshot, the second to-be-trained multi-modal model is trained to determine a candidate multi-modal large model. The candidate multi-modal large model is used to perform a test task of automatic control of AR glasses. During the execution of the test task, the candidate multi-modal large model can generate a plurality of different test action sequences according to a test interface screenshot and a test text instruction corresponding to the test task, select the test action sequences according to the reward function feedback after the execution of each test action sequence by a group relative strategy optimization algorithm, determine a test action sequence with a reward function greater than a preset function threshold as a strategy feedback sequence, and train the candidate multi-modal large model by using the strategy feedback sequence, the test interface screenshot corresponding to the strategy feedback sequence, and the test text instruction to determine a cloud multi-modal large model.
[0109] The target interface screenshot and the target text instruction are input into the cloud multi-modal large model, the target prompt word of the target interface screenshot is extracted by a text perceiver in the cloud multi-modal large model, the target icon and the target virtual control in the target interface screenshot are determined by a graphic perceiver in the cloud multi-modal large model, and the target spatial relationship between the target icon and the target virtual control is determined by a spatial perceiver in the cloud multi-modal large model. The target prompt word, the target icon, and the target virtual control are taken as target interactive elements, and an operation step required to realize the user's intention is determined by a visual language model in the cloud multi-modal large model according to the target spatial relationship, the target interactive elements, and the target text instruction.
[0110] The intelligent glasses determine instruction correction information by a dynamic programming algorithm according to the target task instruction, historical task execution data of a subtask that has been executed in the execution process of the instruction operation information, and a real-time interface screenshot of the intelligent glasses after the execution of the subtask when the instruction operation information is executed.
[0111] After the execution of the instruction operation information ends, the intelligent glasses generate operation result feedback information, and send the operation result feedback information, intermediate data generated when the instruction operation information is executed, an operation interface screenshot corresponding to the intermediate data, and a task completion interface screenshot of the intelligent glasses after the execution of the instruction operation information to a cloud server. The cloud server re-trains the cloud multi-modal large model according to the operation result feedback information, the intermediate data, the operation interface screenshot, and the task completion interface screenshot.
[0112] For example, if the target text instruction is a click instruction, the AR glasses' action instruction parsing module will locate the corresponding virtual control position in the current AR glasses' virtual interface based on the provided target identifier or coordinates, and simulate a selection or click operation; if the target text instruction is an input text instruction, the AR glasses can automatically inject the corresponding characters into the designated text input area; if the target text instruction is a wait instruction, the AR glasses can pause for a specified period of time; if the target text instruction is a prompt instruction, it can trigger the AR glasses' voice broadcast or interface message display function. In this way, the high-level actions described in the standardized instruction sequence in JSON format are implemented one by one as specific control instructions for the AR glasses.
[0113] For example, the AR glasses call an action instruction parsing module to sequentially parse and execute the instruction information in the standardized instruction sequence. The action instruction parsing module reads each instruction information in the standardized instruction sequence, locates the corresponding element in the AR glasses' virtual interface based on the control identifier or coordinate information in the read instruction information, and simulates the execution of the corresponding instruction operation information through the AR glasses' virtual control interface.
[0114] For example, the standardized instruction sequence determined by the cloud server may be: [
[0116] {"action": "click", "target": "Remote collaboration application icon", "x": 500, "y":
[0117] 200},
[0118] {"action": "wait", "duration": 3000},
[0119] {"action": "click", "target": "Call Supervisor Button", "x": 300, "y":
[0120] 150} ]
[0122] The three-step operation described in the above standardized instruction sequence needs to be executed in sequence: first, click the remote collaboration application icon, the coordinates of which are approximately 500, 200 in the field of view; second, wait for 3 seconds to start the application; third, click the "Call Supervisor" button, the coordinates of which are approximately 300, 150 in the field of view. Each instruction contains an action type and a description or location parameter of the target element. In actual application, the standardized instruction sequence can also contain more auxiliary fields, such as "target" for indicating the text label or ID of the UI (User Interface, user interface) control, "type" for indicating the control type, or "speakText" for indicating the content that needs to be spoken.
[0123] The action instruction analysis module called by the AR glasses locates the remote collaboration application icon in the virtual interface of the AR glasses according to the coordinates and performs "click"; then reads the "wait" instruction to pause for about 3 seconds; and then locates the "Call Supervisor" button and performs "click". Through the system interface of the AR glasses, the above operations are all automatically completed on the device side without the need for manual intervention. From the user's perspective, the virtual interface of the AR glasses will appear in sequence the effects of the application being opened, the remote call being dialed out, etc. The entire execution process takes place on the AR glasses device, and the actual interactive operation is simulated by using the graphical interaction control interface.
[0124] In the control method of the smart glasses, the smart glasses acquire a voice instruction issued by a user, and convert the voice instruction into a target text instruction; the smart glasses determine a target interface screenshot, and encrypt the target interface screenshot and the target text instruction, determine encrypted transmission information, and send the encrypted transmission information to a cloud server; the cloud server decrypts the encrypted transmission information, determines the target interface screenshot and the target text instruction, and determines an operation step required to achieve a user intention according to the target interface screenshot and the target text instruction through a cloud multi-modal large model; the cloud server assembles the operation step into a standardized instruction sequence, and sends the standardized instruction sequence to the smart glasses; and the smart glasses determine instruction operation information according to the standardized instruction sequence, and execute the instruction operation information. The problems of high response delay, high privacy risk, low intelligent degree and poor adaptability of the current AR device automatic control method are solved. The above scheme constructs a smart agent system in which an AR glasses device and a cloud server are closely coordinated. The AR glasses quickly process a voice instruction issued by a user through voice recognition technology, convert the voice instruction into a target text instruction, and upload the target text instruction to the cloud server, without uploading audio to the cloud server, so that the interaction delay between the AR glasses and the cloud server can be reduced. The target text instruction and the target interface screenshot are encrypted and sent to the cloud server, so that the user privacy can be protected. Meanwhile, the target text instruction and the target interface screenshot are used as input data of a multi-modal large model, so that the real-time environment and the virtual interface state of the AR glasses during instruction execution can be intuitively reflected. The cloud multi-modal large model deployed on the cloud server analyzes the target interface screenshot and the target text instruction, and determines an operation step required to achieve a user intention, so that the intelligent degree of the control of the AR glasses can be improved. The operation step required to achieve a user intention is determined through the multi-modal large model, so that the target interface screenshot and the target text instruction can be analyzed based on the massive training data of the large model, thereby improving the environmental adaptability of the AR glasses control method. The operation step is assembled into a standardized instruction sequence and then fed back to the AR glasses, so that the interaction delay between the cloud server and the AR glasses can be further reduced.
[0125] Based on the same inventive concept, the embodiments of the present application also provide a control device of smart glasses for implementing the control method of the smart glasses described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, and therefore the specific limitations in one or more control device embodiments of the smart glasses provided below can be referred to the limitations of the control method of the smart glasses described above, which will not be described herein again.
[0126] In one embodiment, as shown in Figure 4 a control device of smart glasses is provided, comprising a voice conversion module 401, an interface screenshot acquisition module 402, a cloud large model inference module 403, an action instruction analysis module 404 and an instruction execution module 405, wherein:
[0127] The voice conversion module 401 is arranged in the smart glasses, and is configured to acquire a voice instruction issued by a user and convert the voice instruction into a target text instruction.
[0128] The interface screenshot acquisition module 402 is arranged in the smart glasses, and is configured to determine a target interface screenshot, encrypt the target interface screenshot and the target text instruction, determine encrypted transmission information, and send the encrypted transmission information to a cloud server.
[0129] The cloud large model inference module 403 is arranged in the cloud server, and is configured to decrypt the encrypted transmission information, determine the target interface screenshot and the target text instruction, and determine an operation step required to achieve a user intention according to the target interface screenshot and the target text instruction by using a cloud multi-modal large model.
[0130] The action instruction analysis module 404 is arranged in the cloud server, and is configured to assemble the operation step into a standardized instruction sequence, and send the standardized instruction sequence to the smart glasses.
[0131] The instruction execution module 405 is arranged in the smart glasses, and is configured to determine instruction operation information according to the standardized instruction sequence, and execute the instruction operation information.
[0132] In an example, the control device of the smart glasses further includes:
[0133] The model training module is arranged in the cloud server, and is configured to combine a pre-trained visual language model, a text perceiver, a graph perceiver and a space perceiver to determine a multi-modal fusion model; use a sample interface screenshot of the smart glasses and a sample description text corresponding to the sample interface screenshot to perform model training on the multi-modal fusion model to determine a first to-be-trained multi-modal model; use the sample interface screenshot, sample annotation information of the sample interface screenshot and a sample interactive element corresponding to the sample interface screenshot to perform model training on the first to-be-trained multi-modal model to determine a second to-be-trained multi-modal model; the sample annotation information includes a sample prompt word, a sample icon and a sample virtual control in the sample interface screenshot; use the sample interface screenshot, a sample instruction sequence and a sample text instruction corresponding to the sample interface screenshot to perform model training on the second to-be-trained multi-modal model to determine a cloud multi-modal large model.
[0134] Further, the model training module is further configured to:
[0135] According to the sample interface screenshot, the sample instruction sequence, and the sample text instruction corresponding to the sample interface screenshot, the second to-be-trained multi-modal model is trained to determine a candidate multi-modal large model; and the candidate multi-modal large model is trained by a group relative strategy optimization algorithm to determine a cloud multi-modal large model.
[0136] Further, the model training module is further specifically configured to:
[0137] The candidate multi-modal large model is used to perform a model test task, a reward function is determined according to a task execution result of the candidate multi-modal large model performing the model test task by using the group relative strategy optimization algorithm; a model reward score is determined based on the reward function by using the group relative strategy optimization algorithm, and model parameters of the candidate multi-modal large model are optimized based on the model reward score to determine the cloud multi-modal large model.
[0138] For example, the cloud large model inference module 403 is specifically configured to:
[0139] The target interface screenshot and the target text instruction are input into the cloud multi-modal large model, and a target prompt word of the target interface screenshot is extracted by a text perceiver in the cloud multi-modal large model.
[0140] A target icon and a target virtual control in the target interface screenshot are determined by a graphic perceiver in the cloud multi-modal large model.
[0141] A target spatial relationship between the target icon and the target virtual control is determined by a spatial perceiver in the cloud multi-modal large model.
[0142] The target prompt word, the target icon, and the target virtual control are taken as target interactive elements, and an operation step required to realize a user's intention is determined by a visual language model in the cloud multi-modal large model according to the target spatial relationship, the target interactive elements, and the target text instruction.
[0143] For example, the control device of the smart glasses further includes:
[0144] An instruction planning module is deployed on the smart glasses and is configured to, when the instruction operation information is executed, determine instruction correction information by a dynamic programming algorithm according to the target text instruction, historical task execution data of a subtask that has been executed in an execution process of the instruction operation information, and a real-time interface screenshot of the smart glasses after the subtask is executed; determine, from the standardized instruction sequence, a to-be-adjusted instruction sequence corresponding to a subtask that is not executed in the instruction operation information; and adjust the to-be-adjusted instruction sequence based on the instruction correction information.
[0145] Exemplarily, the control device of the smart glasses further comprises:
[0146] a feedback information determination module deployed in the smart glasses, configured to generate operation result feedback information after the execution of the instruction operation information ends, and send the operation result feedback information, intermediate data generated when the instruction operation information is executed, operation interface screenshots corresponding to the intermediate data, and task completion interface screenshots of the smart glasses after the execution of the instruction operation information ends to the cloud server;
[0147] a model retraining module deployed in the cloud server, configured to perform model retraining on the cloud multi-modal large model according to the operation result feedback information, the intermediate data, the operation interface screenshots, and the task completion interface screenshots.
[0148] The modules in the control device of the smart glasses can be all or partially implemented by software, hardware, or a combination thereof. The modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in the computer device in software form, so as to be called and executed by the processor.
[0149] In one embodiment, a computer device is provided, comprising a memory and a processor, the memory storing a computer program, when the computer device is a smart glasses, the following steps are implemented: obtaining a voice instruction issued by a user, and converting the voice instruction into a target text instruction; determining a target interface screenshot, and encrypting the target interface screenshot and the target text instruction to determine encrypted transmission information, and sending the encrypted transmission information to a cloud server; determining instruction operation information according to a standardized instruction sequence fed back by the cloud server, and executing the instruction operation information;
[0150] when the computer device is a cloud server, the following steps are executed: decrypting the encrypted transmission information, determining the target interface screenshot and the target text instruction, and determining operation steps required to achieve the user's intention according to the target interface screenshot and the target text instruction through a cloud multi-modal large model; assembling the operation steps into a standardized instruction sequence, and sending the standardized instruction sequence to the smart glasses.
[0151] Exemplarily, in one embodiment, the computer device can be a terminal, and its internal structure diagram can be as shown in Figure 5As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be achieved via Wi-Fi, mobile cellular networks, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a positioning method. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0152] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0153] In one embodiment, a control system for smart glasses is provided. The control system for smart glasses includes a smart glasses terminal and a cloud server, wherein:
[0154] The smart glasses are used to receive voice commands issued by the user and convert the voice commands into target text commands; determine the target interface screenshot, encrypt the target interface screenshot and the target text command, determine the encrypted transmission information, and send the encrypted transmission information to the cloud server; obtain the standardized command sequence sent by the cloud server, determine the command operation information based on the standardized command sequence, and execute the command operation information;
[0155] The cloud server is configured to obtain the encrypted transmission information sent by the smart glasses, decrypt the encrypted transmission information, determine the target interface screenshot and the target text instruction, and determine the operation steps required to achieve the user's intention according to the target interface screenshot and the target text instruction through a cloud multi-modal large model.
[0156] In one embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the following steps:
[0157] Step one, the smart glasses obtain the voice instruction issued by the user, and convert the voice instruction into a target text instruction;
[0158] Step two, the smart glasses determine a target interface screenshot, encrypt the target interface screenshot and the target text instruction, determine encrypted transmission information, and send the encrypted transmission information to the cloud server;
[0159] Step three, the cloud server decrypts the encrypted transmission information, determines the target interface screenshot and the target text instruction, and determines the operation steps required to achieve the user's intention according to the target interface screenshot and the target text instruction through a cloud multi-modal large model;
[0160] Step four, the cloud server assembles the operation steps into a standardized instruction sequence, and sends the standardized instruction sequence to the smart glasses;
[0161] Step five, the smart glasses determine instruction operation information according to the standardized instruction sequence, and execute the instruction operation information.
[0162] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant national and regional laws, regulations and standards.
[0163] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0164] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0165] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A control method of smart glasses, characterized by, The control method of the smart glasses is realized by a smart glasses control system, the smart glasses control system comprising the smart glasses and a cloud server, and the control method of the smart glasses comprising: The smart glasses acquire a voice instruction issued by a user and convert the voice instruction into a target text instruction; The smart glasses determine a target interface screenshot, encrypt the target interface screenshot and the target text instruction, determine encrypted transmission information, and send the encrypted transmission information to the cloud server; The cloud server decrypts the encrypted transmission information, determines the target interface screenshot and the target text instruction, and determines operation steps required to achieve a user intent according to the target interface screenshot and the target text instruction through a cloud multi-modal large model; The cloud server assembles the operation steps into a standardized instruction sequence and sends the standardized instruction sequence to the smart glasses; The smart glasses determine instruction operation information according to the standardized instruction sequence and execute the instruction operation information.
2. The method of claim 1, wherein, The training method of the cloud multi-modal large model comprises: The cloud server combines a pre-trained visual language model, a text perceiver, a graph perceiver, and a spatial perceiver to determine a multi-modal fusion model; Sample interface screenshots of the smart glasses and sample description texts corresponding to the sample interface screenshots are used to train the multi-modal fusion model to determine a first to-be-trained multi-modal model; According to sample interface screenshots, sample annotation information of the sample interface screenshots, and sample interactive elements corresponding to the sample interface screenshots, the to-be-trained multi-modal model is trained to determine a second to-be-trained multi-modal model; the sample annotation information comprises sample prompt words, sample icons, and sample virtual controls in the sample interface screenshots; According to the sample interface screenshots, sample instruction sequences, and sample text instructions corresponding to the sample interface screenshots, the second to-be-trained multi-modal model is trained to determine a cloud multi-modal large model.
3. The method of claim 2, wherein, According to the sample interface screenshots, sample instruction sequences, and sample text instructions corresponding to the sample interface screenshots, the second to-be-trained multi-modal model is trained to determine a cloud multi-modal large model, comprising: According to the sample interface screenshots, sample instruction sequences, and sample text instructions corresponding to the sample interface screenshots, the second to-be-trained multi-modal model is trained to determine a candidate multi-modal large model; The candidate multi-modal large model is reinforced trained through a group relative strategy optimization algorithm to determine the cloud multi-modal large model.
4. The method of claim 3, wherein, The candidate multi-modal large model is reinforced trained through a group relative strategy optimization algorithm to determine the cloud multi-modal large model, comprising: A model test task is executed through the candidate multi-modal large model, a group relative strategy optimization algorithm is used, a reward function is determined according to a task execution result of the candidate multi-modal large model executing the model test task; A model reward score is determined based on the reward function through the group relative strategy optimization algorithm, model parameters of the candidate multi-modal large model are optimized based on the model reward score, and the cloud multi-modal large model is determined.
5. The method of claim 1, wherein, The cloud multi-modal large model is used to determine operation steps required to achieve the user intent according to the target interface screenshot and the target text instruction, including: The target interface screenshot and the target text instruction are input into the cloud multi-modal large model, the target prompt word of the target interface screenshot is extracted through a text perceiver in the cloud multi-modal large model; The target icon and the target virtual control in the target interface screenshot are determined through a graph perceiver in the cloud multi-modal large model; The target spatial relationship between the target icon and the target virtual control is determined through a spatial perceiver in the cloud multi-modal large model; The target prompt word, the target icon and the target virtual control are taken as target interaction elements, and the operation steps required to achieve the user intent are determined through a visual language model in the cloud multi-modal large model according to the target spatial relationship, the target interaction elements and the target text instruction.
6. The method of claim 1, wherein, Further comprising: The intelligent glasses determine instruction correction information through a dynamic programming algorithm according to the target text instruction, historical task execution data of a subtask that has been executed in the instruction operation information execution process, and a real-time interface screenshot of the intelligent glasses after the subtask is executed when the instruction operation information is executed; From the standardized instruction sequence, a to-be-adjusted instruction sequence corresponding to a subtask that has not been executed in the instruction operation information is determined; The to-be-adjusted instruction sequence is adjusted based on the instruction correction information.
7. The method of claim 1, wherein, Further comprising: The intelligent glasses generate operation result feedback information after the instruction operation information execution ends, and send the operation result feedback information, intermediate data generated when the user operation is executed, an operation interface screenshot corresponding to the intermediate data, and a task completion interface screenshot of the intelligent glasses after the instruction operation information execution ends to a cloud server; The cloud server re-trains the cloud multi-modal large model according to the operation result feedback information, the intermediate data, the operation interface screenshot and the task completion interface screenshot.
8. A control device of smart glasses, characterized by, The control device of the intelligent glasses comprises: A voice conversion module deployed on the intelligent glasses, used to acquire a voice instruction issued by a user and convert the voice instruction into a target text instruction; An interface screenshot acquisition module deployed on the intelligent glasses, used to determine a target interface screenshot, encrypt the target interface screenshot and the target text instruction, determine encrypted transmission information, and send the encrypted transmission information to a cloud server; A cloud large model inference module deployed on the cloud server, used to decrypt the encrypted transmission information, determine the target interface screenshot and the target text instruction, and determine operation steps required to achieve the user intent through a cloud multi-modal large model according to the target interface screenshot and the target text instruction; An action instruction analysis module deployed on the cloud server, used to assemble the operation steps into a standardized instruction sequence and send the standardized instruction sequence to the intelligent glasses; An instruction execution module is arranged in the smart glasses, configured to determine instruction operation information according to the standardized instruction sequence, and execute the instruction operation information.
9. A computing device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the computing device is the smart glasses, the following steps are implemented: obtaining a voice instruction issued by a user, and converting the voice instruction into a target text instruction; determining a target interface screenshot, and encrypting the target interface screenshot and the target text instruction; determining encrypted transmission information, and sending the encrypted transmission information to a cloud server; determining instruction operation information according to a standardized instruction sequence fed back by the cloud server, and executing the instruction operation information; When the computing device is the cloud server, the following steps are implemented: decrypting the encrypted transmission information, determining the target interface screenshot and the target text instruction, and determining operation steps required for realizing a user's intention according to the target interface screenshot and the target text instruction through a cloud multi-modal large model; assembling the operation steps into a standardized instruction sequence, and sending the standardized instruction sequence to the smart glasses.
10. A control system for smart glasses, characterized in that, The control system of the smart glasses comprises the smart glasses and the cloud server according to claim 9.
Citation Information
Patent Citations
Image processing method and device, equipment and medium
CN117711001A
Multi-mode large model and augmented reality integrated daily activity auxiliary system for old people
CN118824486A
Visual assisting system and method for visually impaired person based on multi-mode large model, intelligent glasses and blind assisting application software
CN119473005A
Information search system based on large language model, intelligent glasses and information search method
CN119577112A
Multi-modal model visual perception ability enhancement method and device, and medium
CN119809925A
Cited By
Intelligent interaction method, device and system
CN122086223A