Intelligent home assistant system implementation method based on large language model

By adopting large language models and vector databases in the smart home assistant system, the limitations of existing systems in voice interaction and personalized services are solved, and more efficient, accurate and personalized smart home services are achieved.

CN119988326APending Publication Date: 2025-05-13WUHAN UNIV OF TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510041436.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing smart home assistant system has limitations in voice interaction, information understanding and personalized services. It is impossible to accurately understand users' voice commands or provide personalized services, resulting in a decline in user experience.

Method used

The intelligent home assistant system implementation method based on a large language model is adopted. By obtaining user information and smart home settings files, it is vectorized and stored in a vector database. Combining speech recognition and large language model to generate return results, the analysis results are replies content and executable code to realize personalized services.

Benefits of technology

It improves the interaction efficiency and speech recognition accuracy of smart home systems, enhances the personalization level of services, and improves user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988326A_ABST
    Figure CN119988326A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent home assistant system implementation method based on a large language model, and the method comprises the following steps: S1) obtaining information of a user, carrying out the vectorization of an intelligent home setting function file and an intelligent home personalized setting file of the user through a word embedding model, and converting the vectorization into vector representation; s2) carrying out user identity recognition, and reading corresponding personalized settings from the vector database; s3) receiving a voice instruction input of a user, and converting the voice input of the user into a recognized text through voice recognition; s4, the text recognition result is combined with the cue word to be sent to the large language model.S5, the large language model generates a return result according to the received text recognition result and the cue word, analyzes the result generated by the large language model and divides the result into reply content and executable cods.By means of the method, the personalized experience of the smart home can be improved, and the user satisfaction degree is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to artificial intelligence technology, and in particular to a method for realizing an intelligent home assistant system based on a large language model. Background Art

[0002] Currently, the popularity of smart home devices and systems has led to an increase in the demand for smart home assistants. However, traditional smart assistant systems have some limitations in voice interaction, information understanding, and personalized services. For example, existing systems may not be able to accurately understand users' voice commands or provide personalized services, resulting in a decline in user experience.

[0003] In this context, further improvement and innovation of the smart home assistant system becomes crucial. It is necessary to design a smart home assistant implementation method based on a large language model to improve interaction efficiency, speech recognition accuracy, and service personalization. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a method for implementing an intelligent home assistant system based on a large language model in view of the defects in the prior art.

[0005] The technical solution adopted by the present invention to solve the technical problem is: a method for implementing an intelligent home assistant system based on a large language model, comprising the following steps: S1) obtaining user information, smart home setting function files, and user smart home personalized setting files, vectorizing them through a word embedding model, converting them into vector representations, and storing them in a vector database; Among them, the user's information includes the user's ID and user identification information; The smart home setting function file includes the state quantity and parameter setting range of each smart home device; The user's smart home personalized settings are the user's preference settings according to the scenario, including the status and parameter settings of each smart home device; S2) Perform user identification to determine the identity of the current user. By identifying the user, the system will read the corresponding personalized settings from the vector database; S3) receiving a user's voice command input, and converting the user's voice input into recognized text through offline speech recognition; S4) sending the text recognition result combined with the prompt word to the large language model; The process of generating the prompt word is as follows: 4.1) Send the text instruction to the word embedding model, the model returns the corresponding vector, and sends the vector to the vector database for similarity search, and saves the search results of the top n similarities with the smart home setting function file vector and the top k similarities with the personalized setting vector; 4.2) Define prompt words, which should include the following: The top n search results of smart home setting function file vector similarity are used as constraints; The top k search results of personalized setting vector similarity are used as reference; S5) The large language model generates a return result based on the received text recognition result combined with the prompt word, parses the result generated by the large language model, and divides it into reply content and executable code. The reply content is converted into voice form for playback, and the executable code will be executed by the smart home system.

[0006] According to the above scheme, in step 1), the vector data converted from the smart home setting function file is named API, and the vector data converted from the smart home personalized setting file is named mode.

[0007] According to the above solution, in step 1), the user identity identification information is voiceprint information.

[0008] According to the above scheme, in step 2), voiceprint recognition is as follows: When the user issues a command, the user's voice is recorded as an audio clip and saved; Then the audio file is sent to the voiceprint recognition model. The model will first encode the audio, and the encoding result will be calculated through the previously trained neural network. The result of the calculation will be decoded again. The model will return the voiceprint information result, and the result will be compared with the voiceprint information saved in the database. When the confidence level is greater than the set threshold, the identity of the user is determined.

[0009] According to the above scheme, in step 4.1), the top three search results of API similarity and the top two search results of mode similarity are saved.

[0010] According to the above scheme, in step 4.2), the prompt words include the following: "Please combine with the following API: {api}", "Use this as a reference: {mode}".

[0011] According to the above solution, in step 4.2), the prompt words also include prompt words used to make the return result of the large language model a structured result.

[0012] According to the above scheme, in step 5), the large language model generates a return result based on the received text recognition result combined with the prompt word and the user's historical conversation data.

[0013] The beneficial effects produced by the present invention are: 1. Through the smart home setting function file, the present invention enables the system to more accurately understand the user's needs and intentions, thereby providing services that better meet the user's expectations, and combining the user's smart home personalized settings to enhance the personalized experience and increase user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which: Figure 1 It is a method flow chart of an embodiment of the present invention. DETAILED DESCRIPTION

[0015] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0016] like Figure 1 As shown, a method for implementing an intelligent home assistant system based on a large language model includes the following steps: S100) obtaining user information, a smart home setting function file, and a user's smart home personalized setting file, vectorizing them through a word embedding model, converting them into vector representations, and storing them in a vector database; Among them, the user's information includes the user's ID and voiceprint information. The smart home setting function file includes the state quantity and parameter setting range of each smart home device; The user's smart home personalized settings are the user's preference settings according to the scenario, including the status and parameter settings of each smart home device; The vector data converted from the smart home setting function file is named API, and the vector data converted from the smart home personalized setting file is named mode.

[0017] First, we open a blank file and name it API.md. Write various smart home call functions in it, such as: "Smart lighting: api.Light(state) - controls the light on or off, where state can be "on" or "off"." Save the file.

[0018] Create another file and name it mode.md, and write the user's default home scene settings in it, such as: "###Sleep mode Smart Lights: Turn off lights Smart door lock: lock the door Intelligent temperature control system: adjust the temperature to 18 Smart Camera: Open Smart Socket: Off Smart curtains: (0)" Save the file after writing.

[0019] After getting the above information, we need to split it before vector conversion for easy query. First, for the API.md file, we observe that each API is written in one line as much as possible, so we can split the "\n" character. "\n" represents a line break in the file, so we can split the API content in this way. As for the mode.md file, we split it according to the three "#"s before each mode. Through this special symbol design, we can ensure that the settings corresponding to each mode are split together.

[0020] After the information is segmented, we need to use the word embedding model to convert this information into vector form. First, create a loop. Each loop feeds the previously segmented information into the word embedding model and saves the returned vector. Then create a vector database and add the vector just returned to the database. After the loop is completed, create a new VectorStore folder to store the vector database just now. Name the vector database API and mode according to the content.

[0021] To save the user identification information, first record an audio file of the user, then upload the user's audio to the system, decode the audio and perform feature extraction to obtain the user's voiceprint information, upload the user's identification information and save it.

[0022] S200) Perform voiceprint recognition to determine the identity of the current user. By identifying the user, the system will read the corresponding personalized settings from the vector database; When the system starts running, the microphone will be turned on to record audio. When the user gives a command, the user's voice will be recorded as an audio clip, and the recorded audio will be saved in one place.

[0023] Then send this audio file into the voiceprint recognition model. The model will first encode the audio, and the encoding result will be calculated through the previously trained neural network. The result of the calculation will be decoded, and the model will return the voiceprint information. We use the result to compare with the voiceprint information obtained in the previous step. When the confidence level is greater than 0.6, we can determine the identity of the user.

[0024] After obtaining the user identity, the system will search for the previously set vector database under the user's directory and load the found user vector database into the system.

[0025] S300) receiving a user's voice command input, and converting the user's voice input into recognized text through offline voice recognition; The user command audio is sent to the speech recognition model to obtain the text information corresponding to the voice command. This information may be erroneous, so the model is used to correct it and the corrected result is returned.

[0026] The returned results are then processed in some way, such as removing punctuation, converting between simplified and traditional Chinese, and other operations, so that these instructions can be sent to the large language model later.

[0027] S400) sending the text recognition result combined with the prompt word to the large language model; The process of generating the prompt word is as follows: The text instruction is sent to the word embedding model, and the model returns the corresponding vector. The vector is sent to the vector database for similarity search, and the vector is compared with the top n search results of the smart home setting function file vector similarity and the top k search results of the personalized setting vector similarity; Define the prompt word, which needs to include the following: The top n search results of smart home setting function file vector similarity are used as constraints; The top k search results of personalized setting vector similarity are used as reference; In this embodiment, the prompt word needs to include the following content: "Please combine the following API: {api}", "Use this as a reference: {mode}", "Please strictly return according to the following format: AI reply: ```Response Your response after completing the orders given by humans.

[0028] ``` Python code: ```Python Python code that uses the API to complete human commands ```", The rest of the prompt word can be defined by yourself.

[0029] The first two prompts are to concatenate the results obtained from the vector database to achieve personalized services. The last prompt is to make the return of the large language model structured, which is convenient for our subsequent response analysis.

[0030] S500) The large language model generates a return result based on the received text recognition result combined with the prompt word, parses the result generated by the large language model, and divides it into reply content and executable code. The reply content is converted into voice form for playback, and the executable code will be executed by the smart home system.

[0031] The response results of the large language model are also combined with the user's historical conversation data.

[0032] Through the above steps, we have obtained the response results of the large language model. We need to parse it. Thanks to the prompt words in the previous step, we can now use simple regular expressions to parse the results. For the reply content of the large language model, we can use the regular expression (```response(.*?)```) to parse it, and for the executable code replied by the large language model, we can use the regular expression (```python(.*?)```) to parse it.

[0033] We directly send the parsed reply content into the speech synthesis model. The model encodes the text reply content and then sends it to the trained neural network model. The model returns a piece of audio encoding information, and then combines this encoding information with the existing timbre to finally get the content in voice form. We then open the player and play this audio to achieve the effect of voice reply.

[0034] For the parsed executable code, we create a code executor and send the executable code to the executor line by line. The executor will execute the code in sequence and finally achieve the effect of controlling the smart home.

[0035] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all these improvements and changes should fall within the scope of protection of the appended claims of the present invention.

Claims

1. A method for implementing an intelligent home assistant system based on a large language model, characterized in that: The following steps are involved: S1) obtaining user information, smart home setting function files, and user smart home personalized setting files, vectorizing them through a word embedding model, converting them into vector representations, and storing them in a vector database; Among them, the user's information includes the user's ID and user identification information; The smart home setting function file includes the state quantity and parameter setting range of each smart home device; The user's smart home personalized settings are the user's preference settings according to the scenario, including the status and parameter settings of each smart home device; S2) Perform user identification to determine the identity of the current user. By identifying the user, the system will read the corresponding personalized settings from the vector database; S3) receiving a user's voice command input, and converting the user's voice input into recognized text through voice recognition; S4) sending the text recognition result combined with the prompt word to the large language model; The process of generating the prompt word is as follows: 4.1) Send the text instruction to the word embedding model, the model returns the corresponding vector, and sends the vector to the vector database for similarity search, and saves the search results of the top n similarities with the smart home setting function file vector and the top k similarities with the personalized setting vector; 4.2) Define prompt words, which need to include the following: The top n search results of smart home setting function file vector similarity are used as constraints; The top k search results of personalized setting vector similarity are used as reference; S5) The large language model generates a return result based on the received text recognition result combined with the prompt word, parses the result generated by the large language model, and divides it into reply content and executable code. The reply content is converted into voice form for playback, and the executable code will be executed by the smart home system.

2. The method for implementing the intelligent home assistant system based on a large language model according to claim 1, characterized in that: In the step 1), the vector data converted from the smart home setting function file is named API, and the vector data converted from the smart home personalized setting file is named mode.

3. The method for implementing the intelligent home assistant system based on a large language model according to claim 1, characterized in that: In the step 1), the user identification information is voiceprint information.

4. The method for implementing the intelligent home assistant system based on a large language model according to claim 3, characterized in that: In step 2), voiceprint recognition is as follows: When the user issues a command, the user's voice is recorded as an audio clip and saved; Then the audio file is sent to the voiceprint recognition model. The model will first encode the audio, and the encoding result will be calculated through the previously trained neural network. The result of the calculation will be decoded again. The model will return the voiceprint information result, and the result will be compared with the voiceprint information saved in the database. When the confidence level is greater than the set threshold, the identity of the user is determined.

5. The method for implementing the intelligent home assistant system based on a large language model according to claim 2, characterized in that: In the step 4.1), the top three search results with API similarity and the top two search results with mode similarity are saved.

6. The method for implementing the intelligent home assistant system based on a large language model according to claim 5, characterized in that: In step 4.2), the prompt words include the following: "Please combine the following API: {api}", "Use this as a reference: {mode}".

7. The method for implementing the intelligent home assistant system based on a large language model according to claim 1, characterized in that: In the step 4.2), the prompt words also include prompt words used to make the return result of the large language model a structured result.

8. The method for implementing the intelligent home assistant system based on a large language model according to claim 1, characterized in that: In step 5), the large language model generates a return result based on the received text recognition result combined with the prompt word and the user's historical conversation data.

Citation Information

Patent Citations

  • Enterprise internal search engine method based on large language model

    CN116775853A

  • Intelligent family medical equipment control system and method based on large language model

    CN117334191A

  • Method for generating smart home control scene based on large language model

    CN117687314A

  • Voice interaction method and device based on large language model and intelligent voice equipment

    CN118212925A

  • Smart home system voice control method, device and equipment and smart home system

    CN118248145A