System and method for identity recognition by means of smart glasses, and smart mobile terminal

By using smart glasses to capture faces and matching them with the address book using a large language model, a voice reply is generated, solving the problem that smart glasses cannot recognize identities in complex social scenarios and realizing intelligent identity verification.

WO2026040652A1PCT designated stage Publication Date: 2026-02-26SOLOS TECH SHENZHEN LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/106233
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-21
Filing Date
2025-06-30
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Existing smart glasses have limited functionality and cannot effectively identify acquaintances in complex social scenarios, thus lacking a high level of intelligence.

Method used

By capturing facial images with smart glasses and combining them with a large language model and the photo and contact list of a smart mobile terminal, facial matching is performed to generate a voice response and achieve identity recognition.

Benefits of technology

It enriches the functions of smart glasses, helps users identify people they don't remember, improves the level of intelligence, and provides convenient identity verification services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025106233_26022026_PF_FP_ABST
    Figure CN2025106233_26022026_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for identity recognition by means of smart glasses, and a smart mobile terminal, wherein the system comprises: smart glasses, a smart mobile terminal connected to the smart glasses, and a cloud server configured with a large language model; the smart glasses control a camera to photograph a person in front of a user, and send a resulting captured image and a first voice instruction of the user to the smart mobile terminal; the smart mobile terminal performs face matching between the image and a photo of a contact in an album and address book, and inputs into the large language model contact data information and first prompt information of a target contact corresponding to a matched photo; and the large language model generates a first voice response according to the contact data information and the first prompt information, and sends, by means of the smart mobile terminal, the first voice response to the smart glasses for playback. The described system and method for identity recognition by means of the smart glasses, and the smart mobile terminal, expand the functionality of smart glasses, and improve the ease of use by users.
Need to check novelty before this filing date? Find Prior Art

Description

System, method and smart mobile terminal for recognizing identity through smart glasses

[0001] The present application claims priority to the Chinese patent application No. CN 2024111582654, filed on August 21, 2024, entitled "System, method and smart mobile terminal for recognizing identity through smart glasses", the entire content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] Embodiments of the present application relate to the field of wearable devices, in particular to a system, method and smart mobile terminal for recognizing identity through smart glasses. BACKGROUND

[0003] With the development of computer technology, smart glasses are becoming more and more popular. The existing smart glasses usually have the functions of listening to music and making / calling phone, and the functions are relatively simple and the degree of intelligence is relatively low.

[0004] Smart glasses implement some simple face recognition technology, but in some complex application scenarios, for example, as people's social circle expands, they meet more and more people, including business partners, classmates in training activities and social objects based on relatives. However, sometimes for people who do not contact frequently, they may not remember the name when they meet again. There is currently no technology for recognizing the identity of people known based on smart glasses in the above-mentioned scenarios, and the functions of smart glasses are not rich enough. TECHNICAL PROBLEM

[0005] Embodiments of the present application provide a system, method and smart mobile terminal for recognizing identity through smart glasses, which can capture the image of a person through smart glasses and recognize the person in the address book in combination with a large language model. TECHNICAL SOLUTION

[0006] In one aspect, the embodiments of the present application provide a system for recognizing identity through smart glasses, the system comprising:

[0007] a smart glasses, a smart mobile terminal connected to the smart glasses, and a cloud server configured with a large language model;

[0008] The smart glasses are configured to control the camera of the smart glasses to capture a person in front of a user, and send the captured image and a first voice instruction of the user to the smart mobile terminal, wherein the smart mobile terminal stores a photo address book, and the photo address book includes photos of contacts and contact information corresponding to the photos;

[0009] The intelligent mobile terminal is configured to perform face matching between the image and the photos of the contacts in the album address book, and input the contact information of the target contact corresponding to the matched photo, a first text instruction converted from the first voice instruction, and first prompt information generated based on the first voice instruction into a large language model;

[0010] The large language model is configured to generate a first text reply to the first voice instruction based on the contact information of the target contact and the first prompt information, and return the first text reply to the intelligent mobile terminal;

[0011] The intelligent mobile terminal is configured to send a first voice reply converted from the first text reply to the intelligent glasses;

[0012] The intelligent glasses are configured to play the first voice reply.

[0013] The embodiment of the present application also provides a method for identifying an identity through intelligent glasses, comprising:

[0014] The intelligent glasses control a camera of the intelligent glasses to capture a person in front of a user, and the intelligent glasses are configured with a large language model and store an album address book, the album address book including photos of contacts and contact information corresponding to the photos;

[0015] The intelligent glasses perform face matching between the image and the photos of the contacts in the album address book based on a first voice instruction of the user, and input the contact information of a target contact corresponding to the matched photo, a first text instruction converted from the first voice instruction, and first prompt information generated based on the first voice instruction into a large language model;

[0016] The large language model generates a first text reply to the first voice instruction based on the contact information of the target contact and the first prompt information, and the intelligent glasses convert the first text reply into a first voice reply and play the first voice reply.

[0017] The embodiment of the present application also provides an intelligent mobile terminal connected with intelligent glasses, the intelligent mobile terminal being configured to acquire an image sent by the intelligent glasses and a first voice instruction of a user wearing the intelligent glasses to the intelligent glasses, the image being a photo of a person in front of the user captured by the intelligent glasses, the intelligent mobile terminal being configured with a large language model and storing an album address book, the album address book including photos of contacts and contact information corresponding to the photos;

[0018] The intelligent mobile terminal is further configured to perform face matching on the image and photos of contacts in the album address book, and input contact information of a target contact corresponding to a matched photo, a first text instruction converted from the first voice instruction, and first prompt information generated based on the first voice instruction of the user into a large language model.

[0019] The large language model generates a first text reply to the first voice instruction according to the contact information of the target contact and the first prompt information.

[0020] The intelligent mobile terminal obtains the first text reply and converts it into a first voice reply, and sends the first voice reply to the intelligent glasses, so that the intelligent glasses play the first voice reply. Advantages

[0021] As can be seen from the above embodiments of the present application, the intelligent glasses capture a person in front of a user through a camera, and send an image captured and a first voice instruction of the user to an intelligent mobile terminal. The intelligent mobile terminal performs face matching on the image and photos of contacts in an album address book stored locally in the intelligent mobile terminal, inputs contact information of a target contact matched and prompt information generated based on the first voice instruction into a large language model of a cloud server, and the large language model generates a voice reply for replying to the voice instruction according to the contact information and the prompt information, and sends the voice reply to the intelligent mobile terminal. The intelligent mobile terminal sends the voice reply to the intelligent glasses for playing, so that the identity of a person can be determined through the intelligent glasses, the user is helped to get rid of the difficulty of not remembering the identity of a person, the function of the intelligent glasses is enriched, and the intelligent degree of the intelligent glasses and a system for identifying the identity through the intelligent glasses is improved. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application.

[0023] FIG. 1 is a structural schematic diagram of a system for identifying an identity through intelligent glasses according to an embodiment of the present application;

[0024] FIG. 2 is a hardware structural schematic diagram of intelligent glasses according to an embodiment of the present application;

[0025] FIG. 3 is a hardware structural schematic diagram of intelligent glasses according to another embodiment of the present application;

[0026] FIG. 4 is an architectural schematic diagram of a system for identifying an identity through intelligent glasses according to an embodiment of the present application; FIG. 5 is a flowchart of a method for identifying an identity through intelligent glasses according to an embodiment of the present application;

[0027] FIG. 5 is a schematic diagram of an implementation flow of the method for identifying identity through smart glasses according to an embodiment of the present application. Embodiments of the present application

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0029] Referring to FIG. 1, the present application provides a system for identifying identity through smart glasses, which comprises:

[0030] The smart glasses 100, the smart mobile terminal 200 connected with the smart glasses 100, and the cloud server 300 configured with a large language model (LLM);

[0031] The smart glasses 100 and the smart mobile terminal 200 are connected through wireless or wired mode, and the cloud server 300 is wirelessly connected with the smart mobile terminal 200.

[0032] The smart glasses 100 are provided with a camera, a microphone, and are internally provided with a nine-axis sensor. When the smart glasses 100 detect that a preset virtual button is pressed, or, detect that a user uses a preset wake-up word, or, detect a preset action of the user's head, a shooting mode is started to control the camera to shoot a person in front of the user.

[0033] The preset virtual button can be set on the touch operation interface of the smart glasses 100.

[0034] The preset wake-up word can be, for example, “open the shooting mode”.

[0035] The preset action can be, for example, shaking or swinging of the user's head according to a preset rule, which can be recognized by the nine-axis sensor internally provided in the smart glasses 100.

[0036] The starting of the shooting mode means that the camera of the smart glasses 100 is automatically controlled to shoot the person in front of the user. The smart glasses 100 send the image obtained by shooting to the smart mobile terminal 200 together with the first voice instruction of the user. The image at least includes a face image of the person, and can further include a half-body image or a full-body image of the person.

[0037] The first voice instruction is acquired through a microphone of the smart glasses 100. The user can issue the first voice instruction before, after or simultaneously with the camera of the smart glasses 100 taking a photo.

[0038] The first voice instruction contains information requesting confirmation of the identity of the person in front of the user. For example, the user issues the first voice instruction "Tell me the name of the person who just entered the hotel lobby. I think I know him."

[0039] The smart mobile terminal 200 is provided with a voice-to-text module that can convert received voice information into text information, i.e., convert the first voice instruction into a first text instruction.

[0040] The voice-to-text module can convert the voice instruction acquired by the smart mobile terminal 200, specifically the voice instruction acquired by the APP, into text and generate the prompt information. If the content of the voice instruction is not clear enough, a more accurate prompt information is generated by querying historical records to make the LLM more accurately execute the voice instruction. The smart mobile terminal 200 stores a photo address book that includes photos of contacts and contact information corresponding to the photos, i.e., each piece of contact information in the photo address book includes at least one facial photo of a contact and contact information corresponding to the facial photo, such as the name, occupation, identity profile, contact number, social account, social account link and email of the contact.

[0041] The smart mobile terminal 200 is built-in with an APP with a "contact image matching" function for facial matching of the image taken and the photos of contacts in the photo address book. When the APP is running, the permission to access only the photo address book of the smart mobile terminal 200 can be set, while access to other photos on the smart mobile terminal 200 is prohibited, and the image taken by the camera of the smart glasses 100 is not stored and transmitted, thereby improving the security of the system.

[0042] Specifically, the facial features in the image are matched with the facial features in the photos of each contact, and the contact corresponding to the photo with the highest matching degree is confirmed as the target contact, i.e., the person in front of the user is determined as the target contact, and the contact information of the target contact is taken as the contact information of the person. One-time identification of multiple people is supported.

[0043] The smart glasses 100 can control the camera to continuously take pictures of the person in front of the user and continuously send the taken pictures to the smart mobile terminal 200. Correspondingly, the smart mobile terminal 200 continuously performs face detection on the acquired pictures to detect whether there is a face in the picture; if a face is detected, the acquired picture is continuously matched with the photos of the contacts in the album address book; if no face is detected, the continuous processing of the picture is abandoned.

[0044] The smart mobile terminal 200 inputs the contact information of the target contact corresponding to the matched photo, the first text instruction converted from the first voice instruction, and the first prompt information generated based on the first voice instruction into the LLM.

[0045] The LLM is a deep learning model trained using a large amount of text data, and can generate natural language text or understand the meaning of language text. The large language model can process various natural language tasks such as text classification, question answering, and dialogue.

[0046] The first prompt information includes application example data learned by the LLM. The application example data can be historical data. After the LLM learns the application example, it can learn a data analysis and processing method similar to the content indicated by the first voice instruction. Based on the first voice instruction and the contact information of the target contact, a first text reply can be generated. For example, for the first voice instruction "Tell me the name of the person who just entered the hotel lobby. I think I know him" issued by the user, a first text reply "His name is Chen Dawen, and he is the general manager of the Ecological Circle Company" is generated.

[0047] The LLM returns the first text reply to the smart mobile terminal 200, and the smart mobile terminal 200 sends a first voice reply converted from the first text reply to the smart glasses 100. The smart glasses 100 can play the first voice reply through a loudspeaker.

[0048] The smart mobile terminal 200 is also provided with a text-to-speech module that can convert the received text format voice reply into an audio format voice reply.

[0049] The text-to-speech module and the speech-to-text module can also be provided on the cloud server 300.

[0050] Referring to FIG. 2 and FIG. 3, the smart glasses 100 comprises a frame 101, at least one temple 102, at least one microphone 103, at least one speaker 104, at least one processor 105, at least one memory 106, and at least one camera 107. The camera 107 can be disposed on the left side of the frame 101 as shown in FIG. 3, or on the right side of the frame 101, or in the middle of the frame 101, which is not limited in the present application. FIG. 2 is only the best example, and in actual application, the smart glasses 100 can have fewer or more components than shown in FIG. 2.

[0051] The frame 101 can be a front frame with lenses (e.g., sunglass lenses, transparent lenses, or corrective lenses). The at least one temple 102 can include a left temple and a right temple.

[0052] The temple 102 is connected to the frame 101, and the processor 105 is electrically connected to the microphone 103, the speaker 104, the memory 106, and the camera 107. The microphone 103, the speaker 104, the processor 105, the memory 106, and the camera 107 are disposed on at least one temple 102 and / or the frame 101. Preferably, the temple 102 is detachably connected to the frame 101.

[0053] The smart glasses 100 further comprises a wireless communication component 108 electrically connected to the processor 105, which includes a wireless signal transceiver and its surrounding circuit, and can be disposed in the inner cavity of the frame 101 and / or at least one temple 102. The wireless signal transceiver can use at least one of the following protocols for data transmission: WiFi (Wireless Fidelity) protocol, NFC (Near Field Communication) protocol, ZigBee, UWB (Ultra Wideband), RFID (Radio Frequency Identification) protocol, and cellular mobile communication (Cellular Mobile Communication) protocol (such as 3G / 4G / 5G, etc.).

[0054] Optionally, in other embodiments of the present application, the smart glasses 100 further comprises a Bluetooth component 109 electrically connected to the processor 105.

[0055] The smart glasses 100 further comprises a battery 110 for providing power support for the various electronic components of the smart glasses 100, such as the microphone 103, the speaker 104, the processor 105, the memory 106, the camera 107, and other components.

[0056] The various electronic components of the smart glasses can be connected through a bus.

[0057] It should be noted that the relationship between each component of the smart glasses described above can be a replacement relationship or a superimposed relationship. That is, all components in the above embodiment can be installed on a smart glasses, or a part of the components can be selectively installed according to requirements. When it is a replacement relationship, the smart glasses are further provided with a connection interface of an external device, which can be at least one of a PS / 2 interface, a serial interface, a parallel interface, an IEEE 1394 interface, a USB (Universal Serial Bus) interface, and the like. The function of the replaced component can be realized by the external device connected to the connection interface, such as an external speaker, an external sensor, and the like.

[0058] Further, the processor 105 includes a central processing unit (CPU) and a DSP (Digital Signal Processor). The DSP is used to process voice data acquired by the microphone 103. The CPU is preferably an MCU (Microcontroller Unit).

[0059] The memory 106 is a non-transitory memory, which specifically can include a RAM (Random Access Memory) and a flash memory component, and has one or more programs stored therein, which includes a plurality of instructions. The plurality of instructions are used to:

[0060] The first voice instruction of the user is acquired by the microphone 103, the image is captured by the at least one camera 107, and the image is sent to the smart mobile terminal 200 through the wireless communication component 108 or the Bluetooth component 109. The first voice reply sent by the smart mobile terminal 200 is acquired, and the first voice reply is played through the speaker 104.

[0061] The smart mobile terminal 200 receives the image captured by the camera sent by the smart glasses 100, performs face matching on the image and the photos of the contacts in the album address book saved locally in the smart mobile terminal 200, and inputs the contact information of the target contact corresponding to the matched photo, the first text instruction converted from the first voice instruction, and the first prompt information generated based on the first voice instruction to the LLM on the cloud server 300. The LLM generates a first text reply to the first voice instruction according to the contact information of the target contact and the first prompt information, and sends the first text reply to the smart mobile terminal 200.

[0062] The smart mobile terminal 200 can include, but is not limited to, a smart phone, a tablet computer, a cellular phone, a personal digital assistant (PDA), and other wireless communication devices. The smart mobile terminal 200 is also installed with an Android, IOS, or other operating system.

[0063] The cloud server 300 can be a single server or a distributed server cluster composed of multiple servers.

[0064] The cloud server 300 is configured with an LLM, which can include, but is not limited to, a generative artificial intelligence large language model (GAILLM) or a multimodal large language model (MLLM).

[0065] The generative artificial intelligence large language model can include, but is not limited to, ChatGPT of Open AI, Bard of Google, and other models with similar functions. The multimodal large language model can include, but is not limited to, BLIP-2, LLaVA, MiniGPT-4, mPLUG-Owl, LLaMA-Adapter-v2, Otter, Multimodal-GPT, InstructBLIP, VisualGLM-6B, PandaGPT, LaVIN, and other models with similar functions. The LLM can include multiple parts, which are trained based on different samples to reply to different task requests from the smart mobile terminal 200, such as generating natural language voice replies, searching emails, searching messages, searching photos, and calling network services and other third-party software development kits (SDKs) to execute tasks provided by the SDKs.

[0066] The embodiment provides a system for identifying an identity through smart glasses, the smart glasses capture a person in front of a user through a camera, and send an image captured to a smart mobile terminal based on a first voice instruction of the user, the smart mobile terminal performs face matching on the image and photos of contacts in a local album address book of the smart mobile terminal, inputs contact information of a target contact matched, text instructions converted from the voice instruction, and prompt information generated based on the voice instruction into a large language model of a cloud server, the large language model generates text reply used for replying to the voice instruction based on the contact information and the prompt information, and sends the text reply to the smart mobile terminal, the smart mobile terminal converts the text reply into a voice reply, and sends the voice reply to the smart glasses for playing, so that the identity of the person can be identified through the smart glasses, the user is helped to get rid of the difficulty of forgetting the identity of the person, the function of the smart glasses is enriched, and the intelligent degree of the smart glasses and the system for identifying the identity through the smart glasses is improved.

[0067] In another embodiment, the user can continue to inquire based on the first voice instruction, and issue a second voice instruction.

[0068] The user issues the second voice instruction, the second voice instruction can be a further instruction based on the first voice reply, for example, the user issues the second voice instruction "when did we meet last time? What is it about?"

[0069] The smart glasses 100 acquire the second voice instruction through a microphone, and send the second voice instruction to the smart mobile terminal 200, the smart mobile terminal 200 generates second prompt information based on the second voice instruction, and the smart mobile terminal 200 inputs second text instructions converted from the second voice instruction, the second prompt information, and the latest time contact information of the user into the LLM;

[0070] The second prompt information contains application example data corresponding to the second voice instruction for learning of the LLM, after the LLM learns the application example, a data analysis processing mode similar to content indicated by the second voice instruction can be learned, and a second text reply can be generated based on the second voice instruction and information of the target contact in the latest time contact information of the user;

[0071] The LLM specifically searches information related to the target contact in the latest time contact information of the user according to the indication of the second voice instruction;

[0072] The latest time contact information includes the latest information generated when the user contacts other people, such as an email, a message, and a photo.

[0073] The LLM generates a second text reply to the second voice instruction according to the searched information related to the target contact and the second prompt information, for example, generates a second text reply "you had a business meeting in Japan three years ago, and discussed potential partnership" to the second voice instruction "when did we last meet? What was it about?" issued by the user;

[0074] The LLM returns the second text reply to the smart mobile terminal 200, and the smart mobile terminal 200 sends a second voice reply converted from the second text reply to the smart glasses 100, and the smart glasses 100 play the second voice reply through a loudspeaker.

[0075] The embodiment provides a system for identifying identity through smart glasses, and the system can obtain the latest relevant information of a target contact in front of a user and the user through a smart mobile terminal controlling a large language model in contact information of the user in a recent time according to a voice instruction of the user through the smart glasses, thereby further enriching the function of confirming the identity of the other party through the smart glasses, providing higher convenience for the user, and improving the intelligent degree of the smart glasses and the system for identifying identity through the smart glasses.

[0076] In another embodiment, the user can continue to inquire on the basis of the first voice instruction or the second voice instruction, and issue a third voice instruction.

[0077] The user issues a third voice instruction, which can be a further instruction on the basis of the first voice reply or a further instruction on the basis of the second voice reply, for example, the user issues a third voice instruction "can you tell me what he has been doing recently?".

[0078] The smart glasses 100 acquire the third voice instruction through a microphone, and send the third voice instruction to the smart mobile terminal 200, and the smart mobile terminal 200 generates a third prompt information based on the third voice instruction, and inputs a third text instruction converted from the third voice instruction and the third prompt information into the LLM;

[0079] The third prompt information contains application example data corresponding to the third voice instruction for the LLM to learn, and after the LLM learns the application example, the LLM can search the latest network information of the target contact based on the third voice instruction and the contact information of the target contact, and generate a third voice reply;

[0080] The LLM searches the latest network information of the target contact according to the contact information of the target contact according to the indication of the third voice instruction;

[0081] The latest network information includes: the latest time from now, the information of the person including name, company, etc. used in social media (such as Facebook) or network search.

[0082] The LLM generates a third text reply to the third voice instruction according to the searched latest network information and the third prompt information, for example, for the third voice instruction "Can you tell me what he has recently?" issued by the user, a third text reply "His company acquired another company NovaNest at a price of 10 billion US dollars" is generated;

[0083] The LLM returns the third text reply to the smart mobile terminal 200, and the smart mobile terminal 200 sends a third voice reply converted from the third text reply to the smart glasses 100, and the smart glasses 100 play the third voice reply through the speaker.

[0084] The embodiment provides a system for identifying identity through smart glasses, and the system further enriches the function of confirming the identity of the other party through smart glasses by the smart glasses according to the voice instruction of the user, the smart mobile terminal controls the large language model to search the latest network information of the target contact on the network, provides higher convenience for the user, and improves the intelligent degree of the smart glasses and the system for identifying identity through smart glasses.

[0085] In another embodiment, based on the first voice instruction, if the smart mobile terminal 200 does not match the face of the person in the image in the photos of the contacts in the album address book, the social information of the user is obtained, and the target contact is matched in the social information.

[0086] The social information includes social images and social contact information corresponding to the social images, the social images include photos of each social platform, and the social contact information includes contact information of the social contact, such as name, occupation, identity profile, contact number, social account, social account link and email, etc.

[0087] The smart mobile terminal 200 performs face matching on the images obtained by the camera and each social image of the user, if a social image is matched, the social contact corresponding to the matched social image is taken as a target social contact, and the target social contact is the person in front of the user, which is not a contact in the album address book of the user, but a contact of a certain contact of the user in the social image, for example, a friend B of a colleague A of the user.

[0088] The social contact information of the target social contact corresponding to the matched social image, the first text instruction converted from the first voice instruction, and the first prompt information generated based on the first voice instruction are input into the LLM;

[0089] The LLM is configured to generate a fourth text reply to the first voice instruction according to the social contact information of the target social contact and the first prompt information, and return the fourth text reply to the intelligent mobile terminal 200.

[0090] For example, the first voice instruction is "Tell me the name of the person who just entered the hotel lobby, I think I know him", and the fourth text reply can be "His name is Zhang A, who is a colleague of your friend Wang B".

[0091] The intelligent mobile terminal 200 sends a fourth voice reply converted from the fourth text reply to the intelligent glasses 100, and the intelligent glasses 100 play the fourth voice reply.

[0092] The embodiment provides a system for identifying an identity through intelligent glasses. Based on a voice instruction for identifying an identity of a person in front of a user issued by the user, if the intelligent mobile terminal fails to match a face in an image of the person captured by the intelligent glasses in a photo album address book, the intelligent mobile terminal matches the face in the image from a social image of the user, so as to realize indirect identity recognition of a social personnel of the user, further enrich the functionality of confirming an identity of a person through the intelligent glasses, provide higher convenience for the user, and improve the intelligent degree of the intelligent glasses and the system for identifying an identity through the intelligent glasses.

[0093] In another embodiment, based on the first voice instruction, if the intelligent mobile terminal 200 fails to match the face in the image in the photos of the contacts in the photo album address book, the intelligent mobile terminal 200 allows the LLM to perform face matching through a third-party tool.

[0094] If the intelligent mobile terminal 200 fails to match the face in the image in the photos of the contacts in the photo album address book, the intelligent mobile terminal 200 generates fourth prompt information, and sends the image, the photo album address book, the first text instruction, and the fourth prompt information to the LLM.

[0095] The fourth prompt information includes application example data corresponding to the first voice instruction for learning by the LLM. After learning the application example, the LLM can send the photos of the contacts in the photo album address book and the image to the third-party tool for face matching.

[0096] Specifically, the LLM sends the image and the photo album address book to the third-party tool for face matching according to the fourth prompt information, and obtains a matching result of the third-party tool.

[0097] The third-party tool can be another face matching model configured on the cloud server 300. Different matching algorithms have different accuracies and different requirements for image clarity. If the smart mobile terminal 200 fails to match the result, the third-party tool can find a target contact person in the album address book who matches the face of the image.

[0098] Further, if the third-party tool finds a target contact person in the album address book whose contact photo matches the face of the image, the matching result is returned to the LLM.

[0099] The LLM generates the first text reply to the first voice command according to the contact information of the target contact person and the first prompt information, and sends the first text reply to the smart mobile terminal 200.

[0100] The smart mobile terminal 200 sends the first voice reply converted from the first text reply to the smart glasses 100, and the smart glasses 100 play the first voice reply through a loudspeaker.

[0101] The system for identifying identity through smart glasses provided in the embodiment is based on the voice command for identifying the identity of the person in front of the user issued by the user. If the smart mobile terminal fails to match the face in the image of the person photographed by the smart glasses in the album address book, the large language model is controlled to perform face matching through a third-party tool, which further enriches the functionality of confirming the identity of the person through the smart glasses, provides higher convenience for the user, and improves the intelligent degree of the smart glasses and the system for identifying identity through smart glasses.

[0102] Further, if the smart mobile terminal 200 fails to find a target contact person in the album address book whose contact photo matches the face of the image, or if the third-party tool fails to find a target contact person in the album address book whose contact photo matches the face of the image, or if the smart mobile terminal fails to find a target social contact person in the social information whose social image matches the face of the image, the smart mobile terminal 200 issues a voice prompt to prompt the user to establish a new contact information based on the image photographed by the camera.

[0103] For example, the content of the voice prompt is: He is not in the address book. Do you want to add him to the address book?

[0104] The system for identifying identity through the smart glasses provided in the embodiment is based on the voice instruction for identifying the identity of the person in front of the user issued by the user, and if the smart mobile terminal and the third-party tool cannot match the face in the image of the person photographed by the smart glasses in the album address book or social information, the user is reminded to add the person in front to the album address book, which further enriches the function of confirming the identity of the person through the smart glasses for the user, provides higher convenience for the user, and improves the intelligent degree of the smart glasses and the system for identifying identity through the smart glasses.

[0105] Referring to FIG. 4, FIG. 4 is a schematic diagram of the architecture of the system for identifying identity through the smart glasses. The voice instruction issued by the user is sent to the LLM through the smart terminal, and the LLM returns the voice reply based on each voice instruction and sends it to the smart glasses through the smart terminal for playing.

[0106] The voice-to-text module and the text-to-speech module in FIG. 4 are arranged on the cloud server, the voice-to-text module converts the voice instruction into text and sends it to the large language model, and each reply generated by the large language model is converted into an audio file that can be played on the smart glasses through the text-to-speech module.

[0107] It should be noted that the LLM can also be arranged on the smart mobile terminal 200, and then the system for identifying identity through the smart glasses includes the smart glasses and the smart mobile terminal connected with the smart glasses. The data transmission between the smart mobile terminal and the LLM is changed from external transmission to internal transmission, and their functions are the same as those in the above embodiments.

[0108] The embodiment of the application also provides a method for identifying identity through the smart glasses, which realizes the confirmation of the identity of the person in front of the user through the system for identifying identity through the smart glasses as described above. The application scenarios of the method include the smart glasses. The APP and the LLM of the "contact image matching" function in the above embodiments are arranged on the smart glasses. The smart glasses in the embodiment of the application are the smart glasses 100 in the above embodiments.

[0109] Referring to FIG. 5, FIG. 5 is a schematic diagram of the implementation flow of the method for identifying identity through the smart glasses, which includes:

[0110] S501, the smart glasses photograph the person in front of the user to obtain an image;

[0111] The smart glasses control the camera of the smart glasses to photograph the person in front of the user;

[0112] The smart glasses are configured with a large language model and store an album address book, and the album address book includes the photos of the contacts and the contact information corresponding to the photos.

[0113] S502, the smart glasses perform face matching on the image and the photos of the contacts in the album address book based on the first voice instruction, and input the contact information of the target contact corresponding to the matched photo, the first text instruction converted from the first voice instruction, and the first prompt information generated based on the first voice instruction into a large language model.

[0114] S503, the large language model generates a first text reply to the first voice instruction according to the contact information of the target contact and the first prompt information, the smart glasses convert the first text reply into a first voice reply and play the first voice reply through a loudspeaker.

[0115] The specific contents in the embodiments of the present application are described in the foregoing embodiments.

[0116] In the embodiments of the present application, the smart glasses capture a person in front of the user through a camera, the smart glasses perform face matching on the image and the photos of the contacts in the album address book stored locally in the smart glasses, input the contact information of the target contact matched and the prompt information generated based on the voice instruction into a large language model of a cloud server, the large language model generates a voice reply to reply to the voice instruction according to the contact information and the prompt information, and sends the voice reply to the smart glasses for playing, so that the identity of a person can be determined through the smart glasses, helping the user to get rid of the difficulty of not remembering the identity of the person, enriching the function of the smart glasses, and improving the intelligent degree of the smart glasses and the system for identifying the identity through the smart glasses.

[0117] Further, the smart glasses obtain a second voice instruction of the user, and input the recent contact information, a second text instruction converted from the second voice instruction, and a second prompt information generated based on the second voice instruction into the large language model.

[0118] The large language model searches for information related to the target contact in the recent contact information of the user according to the indication of the second voice instruction.

[0119] According to the searched information related to the target contact and the second prompt information, a second text reply to the second voice instruction is generated; the smart glasses convert the second text reply into a second voice reply, and play the second voice reply.

[0120] Further, the smart glasses obtain a third voice instruction of the user, and input a third text instruction converted from the third voice instruction and a third prompt information generated based on the third voice instruction into the large language model.

[0121] The large language model searches for the latest network information of the target contact according to the contact information of the target contact according to the indication of the third voice instruction;

[0122] According to the searched latest network information and the third prompt information, a third text reply to the third voice instruction is generated; the smart glasses convert the third text reply into a third voice reply and play the third voice reply.

[0123] Further, if the smart glasses do not match the face of the person in the image in the photos of the contacts in the album address book, the social information of the user is obtained, the social information including a social image and social contact information corresponding to the social image;

[0124] The social image and the image taken by the camera are face-matched;

[0125] The social contact information of the target social contact corresponding to the matched social image, the first text instruction converted from the first voice instruction, and the first prompt information generated based on the first voice instruction are input into the large language model;

[0126] The large language model generates a fourth text reply to the first voice instruction according to the social contact information of the target social contact and the first prompt information; the smart glasses convert the fourth text reply into a fourth voice reply and play the fourth voice reply.

[0127] Further, if the smart glasses do not match the face of the person in the image in the photos of the contacts in the album address book, a fourth prompt information is generated, and the image, the album address book and the fourth prompt information are sent to the large language model;

[0128] The large language model sends the image and the album address book to a third-party tool for face matching according to the fourth prompt information, and obtains the matching result of the third-party tool.

[0129] Further, if the third-party tool finds a target contact in the album address book whose contact photo matches the face of the image,

[0130] The large language model is used to generate the first text reply to the first voice instruction according to the contact information of the target contact and the first prompt information; the smart glasses convert the first text reply into a first voice reply and play the first voice reply.

[0131] Further, the smart glasses are configured to issue a voice prompt to prompt the user to establish a new contact information based on the image captured by the camera, if no target contact person is found in the album address book, or if no target contact person is found in the album address book by the third-party tool, or if no target social contact person is found in the social information.

[0132] Further, the smart glasses are configured to start the shooting mode when detecting that a preset virtual button is pressed, or detecting that the user uses a preset wake-up word, or detecting a preset action of the user.

[0133] Further, the smart glasses are configured to control the camera of the smart glasses to continuously shoot the person in front of the user, and continuously perform face detection on the captured image; if a face is detected, continuously perform face matching between the captured image and the photos of the contacts in the album address book; if no face is detected, abandon the continuous processing of the image.

[0134] For other technical details of the present embodiment, please refer to the description in the foregoing embodiments.

[0135] The present embodiment also provides a smart mobile terminal connected with the smart glasses. The APP and LLM of the "contact image matching" function in the foregoing embodiments are arranged on the smart mobile terminal. The smart mobile terminal and the smart glasses in the present embodiment are respectively the smart mobile terminal 200 and the smart glasses 100 in the foregoing embodiments.

[0136] The smart mobile terminal is configured to acquire an image and a first voice instruction issued by the user to the smart glasses, the image being a photo of a person in front of the user wearing the smart glasses captured by the smart glasses;

[0137] The smart mobile terminal is configured to acquire an image and a first voice instruction issued by the user to the smart glasses, the image being a photo of a person in front of the user wearing the smart glasses captured by the smart glasses;

[0138] The smart mobile terminal is configured to acquire an image and a first voice instruction issued by the user to the smart glasses, the image being a photo of a person in front of the user wearing the smart glasses captured by the smart glasses;

[0139] The intelligent mobile terminal obtains the first text reply and converts it into a first voice reply, and sends the first voice reply to the intelligent glasses, so that the intelligent glasses play the first voice reply.

[0140] Further, the intelligent mobile terminal is configured to obtain a second voice instruction of the user, and input recent time contact information, a second text instruction converted from the second voice instruction, and second prompt information generated based on the second voice instruction and the recent time contact information into the large language model, so that the large language model is configured to search for information related to the target contact in the recent time contact information of the user according to an indication of the second voice instruction, and generate a second text reply for the second voice instruction according to the searched information related to the target contact and the second prompt information.

[0141] The intelligent mobile terminal obtains the second text reply and converts it into a second voice reply, and sends the second voice reply to the intelligent glasses, so that the intelligent glasses play the second voice reply.

[0142] Further, the intelligent mobile terminal is configured to obtain a third voice instruction of the user, and input a third text instruction converted from the third voice instruction and third prompt information generated based on the third voice instruction into the large language model, so that the large language model is configured to search for the latest network information of the target contact according to an indication of the third voice instruction and according to the contact information of the target contact, and generate a third text reply for the third voice instruction according to the searched latest network information and the third prompt information.

[0143] The intelligent mobile terminal obtains the third text reply and converts it into a third voice reply, and sends the third voice reply to the intelligent glasses, so that the intelligent glasses play the third voice reply.

[0144] Further, the intelligent mobile terminal is further configured to obtain social information of the user if the face in the image is not matched in the photos of the contacts in the album address book, the social information including a social image and social contact information corresponding to the social image.

[0145] The social image and the image captured by the camera are face-matched.

[0146] The social contact information of the target social contact corresponding to the matched social image, a first text instruction converted from the first voice instruction, and first prompt information generated based on the first voice instruction are input into the large language model, so that the large language model is configured to generate a fourth text reply for the first voice instruction according to the social contact information of the target social contact and the first prompt information.

[0147] The intelligent mobile terminal obtains the fourth text reply and converts it into a fourth voice reply, and sends the fourth voice reply to the intelligent glasses, so that the intelligent glasses play the fourth voice reply.

[0148] Further, the intelligent mobile terminal is further configured to generate a fourth prompt information if the face in the image is not matched with the photos of the contacts in the album address book, and send the image, the album address book, the first text instruction, and the fourth prompt information to the large language model. The large language model sends the image and the album address book to a third-party tool for face matching according to the fourth prompt information, and obtains a matching result of the third-party tool.

[0149] Further, if the third-party tool finds a target contact in the album address book whose photo is matched with the face in the image, the intelligent mobile terminal is further configured to generate the first prompt information based on the first voice instruction, and input the first prompt information into the large language model, so that the large language model generates the first text reply for the first voice instruction according to the contact information of the target contact and the first prompt information.

[0150] The intelligent mobile terminal obtains the first text reply and converts it into a first voice reply, and sends the first voice reply to the intelligent glasses for playing.

[0151] Further, if the target contact whose photo is matched with the face in the image is not found in the album address book, or if the third-party tool does not find the target contact whose photo is matched with the face in the image in the album address book, or if the target social contact whose social image is matched with the face in the image is not found in the social information, the intelligent mobile terminal is further configured to issue a voice prompt to prompt the user to establish a new contact information based on the image captured by the camera.

[0152] In the embodiments of the present application, the intelligent mobile terminal performs face matching on the image of the person in front of the user captured by the intelligent glasses and the photos of the contacts in the album address book stored locally in the intelligent mobile terminal, inputs the contact information of the target contact matched and the prompt information generated based on the voice instruction into the large language model of the cloud server, so that the large language model generates a voice reply for replying to the voice instruction according to the contact information and the prompt information, obtains the voice reply and sends the voice reply to the intelligent glasses for playing. Thus, the identity of a person can be determined through the intelligent mobile terminal, the user is helped to get rid of the difficulty of not remembering the identity of a person, and the convenience of the user in confirming the identity of the person in front of him through the intelligent glasses is improved.

[0153] The specific contents in the embodiments of the present application are described in the foregoing embodiments.

[0154] It should be noted that, for the foregoing method embodiments, the purposes of brief description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0155] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0156] The above is the description of the system and method for identifying identity through smart glasses provided by the present application. For those skilled in the art, according to the idea of the embodiments of the present application, there will be changes in specific implementation and application range. In summary, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A system for identifying identity through smart glasses, characterized in that, The system comprises: smart glasses, a smart mobile terminal connected with the smart glasses, and a cloud server configured with a large language model; The smart glasses are used to control the camera of the smart glasses to take a photo of a person in front of a user, and send the taken photo and a first voice instruction of the user to the smart mobile terminal, wherein the smart mobile terminal stores a photo address book, and the photo address book includes photos of contacts and contact information of the contacts corresponding to the photos; The smart mobile terminal is used to perform face matching between the photo and the photos of the contacts in the photo address book, and input contact information of a target contact corresponding to the matched photo, a first text instruction converted from the first voice instruction, and a first prompt information generated based on the first voice instruction into the large language model; The large language model is used to generate a first text reply to the first voice instruction according to the contact information of the target contact and the first prompt information, and return the first text reply to the smart mobile terminal; The smart mobile terminal is used to send a first voice reply converted from the first text reply to the smart glasses; The smart glasses are used to play the first voice reply.

2. The system of claim 1, wherein, The smart glasses are used to obtain a second voice instruction of the user, and send the second voice instruction to the smart mobile terminal; The smart mobile terminal is used to generate a second prompt information based on the second voice instruction, and input recent contact information, a second text instruction converted from the second voice instruction, and the second prompt information into the large language model; The large language model is used to search information related to the target contact in the recent contact information of the user according to an indication of the second voice instruction; generate a second text reply to the second voice instruction according to the searched information related to the target contact and the second prompt information, and return the second text reply to the smart mobile terminal; The smart mobile terminal is used to send a second voice reply converted from the second text reply to the smart glasses; The smart glasses are used to play the second voice reply.

3. The system of claim 1, wherein, The smart glasses are used to obtain a third voice instruction of the user, and send the third voice instruction to the smart mobile terminal; The smart mobile terminal is used to generate a third prompt information based on the third voice instruction, and input a third text instruction converted from the third voice instruction and the third prompt information into the large language model; The large language model is used to search latest network information of the target contact according to the contact information of the target contact according to an indication of the third voice instruction; generate a third text reply to the third voice instruction according to the searched latest network information and the third prompt information, and return the third text reply to the smart mobile terminal; The smart mobile terminal is used to send a third voice reply converted from the third text reply to the smart glasses; The smart glasses are configured to play the third voice reply.

4. The system of claim 1, wherein, The smart mobile terminal is configured to, if a face in the image is not matched in the photos of the contacts in the album address book, acquire social information of the user, the social information including a social image and social contact information corresponding to the social image; perform face matching on the social image and the image captured by the camera; and input the social contact information of a target social contact corresponding to the matched social image, the first text instruction converted from the first voice instruction, and the first prompt information generated based on the first voice instruction into the large language model; The large language model is configured to generate a fourth text reply to the first voice instruction according to the social contact information of the target social contact and the first prompt information, and return the fourth text reply to the smart mobile terminal; The smart mobile terminal is configured to send a fourth voice reply converted from the fourth text reply to the smart glasses; The smart glasses are configured to play the fourth voice reply.

5. The system of claim 1, wherein, The smart mobile terminal is configured to, if a face in the image is not matched in the photos of the contacts in the album address book, generate a fourth prompt information, and send the image, the album address book, the first text instruction, and the fourth prompt information to the large language model; The large language model is configured to, according to the fourth prompt information, send the image and the album address book to a third-party tool for face matching, and acquire a matching result of the third-party tool.

6. The system of claim 5, wherein, The large language model is configured to, if the third-party tool finds a target contact in the album address book whose contact photo is matched with the face in the image, generate the first text reply to the first voice instruction according to contact information of the target contact and the first prompt information, and return the first text reply to the smart mobile terminal; The smart mobile terminal is configured to send a first voice reply converted from the first text reply to the smart glasses; The smart glasses are configured to play the first voice reply.

7. The system of claim 6, wherein, The smart mobile terminal is configured to, if a target contact whose contact photo is matched with the face in the image is not found in the album address book, or if the third-party tool does not find a target contact whose contact photo is matched with the face in the image in the album address book, or if the smart mobile terminal does not find a target social contact whose social image is matched with the face in the image in the social information, issue a voice prompt to prompt the user to establish a new contact information based on the image captured by the camera.

8. The system of claim 1, wherein, The smart glasses are configured to detect that a preset virtual button is pressed, or detect that the user uses a preset wake-up word, or detect a preset action of the head of the user, and start a shooting mode to control the camera to shoot a person in front of the user.

9. The system of claim 1, wherein, The smart glasses are used to control the camera of the smart glasses to continuously take pictures of the people in front of the user, and continuously send the taken pictures to the smart mobile terminal; The smart mobile terminal is used to continuously perform face detection on the acquired pictures; If a face is detected, the acquired pictures are continuously matched with the photos of the contacts in the album address book; If no face is detected, the continuous processing of the pictures is abandoned.

10. A method for identifying identity through smart glasses, characterized in that, The method comprises: The smart glasses control the camera of the smart glasses to take pictures of the people in front of the user, and the smart glasses are configured with a large language model and store an album address book, which includes photos of contacts and contact information of the contacts corresponding to the photos; The smart glasses perform face matching between the pictures and the photos of the contacts in the album address book based on a first voice instruction of the user, and input the contact information of a target contact corresponding to the matched photo, a first text instruction converted from the first voice instruction, and first prompt information generated based on the first voice instruction of the user into the large language model; The large language model generates a first text reply to the first voice instruction according to the contact information of the target contact and the first prompt information; the smart glasses convert the first text reply into a first voice reply and play the first voice reply.

11. The method of claim 10, wherein, The smart glasses acquire a second voice instruction of the user, and input recent contact information, a second text instruction converted from the second voice instruction, and second prompt information generated based on the second voice instruction into the large language model; The large language model searches for information related to the target contact in the recent contact information of the user according to the indication of the second prompt information; and generates a second text reply to the second voice instruction according to the searched information related to the target contact and the second prompt information; The smart glasses convert the second text reply into a second voice reply and play the second voice reply.

12. The method of claim 10, wherein, The smart glasses acquire a third voice instruction of the user, and input a third text instruction converted from the third voice instruction and third prompt information generated based on the third voice instruction into the large language model; The large language model searches for the latest network information of the target contact according to the contact information of the target contact according to the indication of the third prompt information; and generates a third text reply to the third voice instruction according to the searched latest network information and the third prompt information; the smart glasses convert the third text reply into a third voice reply and play the third voice reply.

13. The method of claim 10, wherein, If the smart glasses do not match the face in the pictures with the people in the pictures in the album address book, the social information of the user is acquired, and the social information includes social images and social contact information corresponding to the social images; Face matching is performed between the social images and the pictures taken by the camera; The social contact information of the target social contact corresponding to the matched social image, the first text instruction converted from the first voice instruction, and the first prompt information generated based on the first voice instruction are input into the large language model; The large language model generates a fourth text reply to the first voice instruction according to the social contact information of the target social contact and the first prompt information; the smart glasses convert the fourth text reply into a fourth voice reply and play the fourth voice reply.

14. The method of claim 10, wherein, If the smart glasses do not match the face in the image in the photos of contacts in the album address book, a fourth prompt information is generated, and the image, the album address book, the first text instruction, and the fourth prompt information are sent to the large language model; The large language model sends the image and the album address book to a third-party tool for face matching according to the fourth prompt information, and obtains the matching result of the third-party tool.

15. The method of claim 14, wherein, If the third-party tool finds a target contact in the album address book whose contact photo matches the face in the image, the large language model is used to generate the first text reply to the first voice instruction according to the contact information of the target contact and the first prompt information; the smart glasses convert the first text reply into a first voice reply and play the first voice reply.

16. The method of claim 15, wherein, The smart glasses are used to issue a voice prompt to prompt the user to establish a new contact information based on the image captured by the camera if no target contact whose contact photo matches the face in the image is found in the album address book, or if the third-party tool does not find a target contact whose contact photo matches the face in the image in the album address book, or if no target social contact whose social image matches the face in the image is found in the social information.

17. A smart mobile terminal connected with smart glasses, characterized in that, The smart mobile terminal is used to obtain an image captured by smart glasses and a first voice instruction issued by a user wearing the smart glasses, the image being a photo of a person in front of the user captured by the smart glasses, the smart mobile terminal being configured with a large language model and storing an album address book, the album address book including photos of contacts and contact information corresponding to the photos; The smart mobile terminal is also used to perform face matching on the image and the photos of contacts in the album address book, and input the contact information of a target contact corresponding to the matched photo, a first text instruction converted from the first voice instruction, and a first prompt information generated based on the first voice instruction of the user into the large language model; The large language model generates a first text reply to the first voice instruction according to the contact information of the target contact and the first prompt information; The smart mobile terminal obtains the first text reply and converts it into a first voice reply, and sends the first voice reply to the smart glasses to make the smart glasses play the first voice reply. 18.The intelligent mobile terminal of claim 17, wherein, The intelligent mobile terminal is configured to acquire a second voice instruction of the user, and input recent time contact information, a second text instruction converted from the second voice instruction, and second prompt information generated based on the second voice instruction into the large language model; The large language model is configured to search for information related to the target contact in the contact information of the user at the recent time according to the indication of the second prompt information, and generate a second text reply to the second voice instruction according to the searched information related to the target contact and the second prompt information; The intelligent mobile terminal acquires the second text reply and converts it into a second voice reply, and sends the second voice reply to the intelligent glasses, so that the intelligent glasses play the second voice reply. 19.The intelligent mobile terminal of claim 17, wherein, The intelligent mobile terminal is configured to acquire a third voice instruction of the user, and input a third text instruction converted from the third voice instruction and third prompt information generated based on the third voice instruction into the large language model; The large language model is configured to search for the latest network information of the target contact according to the contact information of the target contact according to the indication of the third prompt information, and generate a third text reply to the third voice instruction according to the searched latest network information and the third prompt information; The intelligent mobile terminal acquires the third text reply and converts it into a third voice reply, and sends the third voice reply to the intelligent glasses, so that the intelligent glasses play the third voice reply. 20.The intelligent mobile terminal of claim 10, wherein, The intelligent mobile terminal is further configured to acquire social information of the user if the face in the image is not matched in the photos of the contacts in the album address book, wherein the social information includes a social image and social contact information corresponding to the social image; The social image and the image captured by the camera are matched in face; The social contact information of a target social contact corresponding to the matched social image, a first text instruction converted from the first voice instruction, and the first prompt information generated based on the first voice instruction are input into the large language model; The large language model is configured to generate a fourth text reply to the first voice instruction according to the social contact information of the target social contact and the first prompt information; The intelligent mobile terminal acquires the fourth text reply and converts it into a fourth voice reply, and sends the fourth voice reply to the intelligent glasses, so that the intelligent glasses play the fourth voice reply.

Citation Information

Patent Citations

  • Intelligent glasses

    CN104880835A

  • Wearable intelligent device and identity identification method and system based on wearable intelligent device

    CN108776798A

  • Information interaction and / or social contact method based on face recognition and intelligent glasses system thereof

    CN112231613A

  • Intelligent glasses, intelligent glasses system and intelligent glasses interaction method

    CN112433372A

  • Information prompting method based on intelligent glasses, intelligent glasses and storage medium

    CN118151881A