Atmography machine control method and device, electronic equipment and storage medium

By displaying the main item on the main interface of the learning machine and collecting voice question data, and using a large language model to generate answer data, the problem of the small display interface making it impossible to set interactive buttons is solved, thus satisfying user interaction needs.

CN120928972APending Publication Date: 2025-11-11SHENZHEN LUKA DR TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510900550.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing learning machines have a small display interface and cannot be equipped with interactive buttons, thus failing to meet users' interactive needs.

Method used

By displaying the main object on the main interface of the learning machine and turning on the microphone through the voice input button, the system collects the user's voice question data, generates answer data using a large language model, and outputs the answer data through voice.

Benefits of technology

This achieves the goal of meeting user interaction needs without adding physical buttons, thus improving the interactive capabilities of the learning machine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120928972A_ABST
    Figure CN120928972A_ABST
Patent Text Reader

Abstract

The invention provides a control method of an apex learning machine, and the method comprises the steps: displaying a main body article on a main interface when the main body article is determined, and turning on a microphone of the apex learning machine through triggering a voice input key in the main interface; collecting voice questioning data of a user through a microphone; based on the voice question data and the main body article, generating corresponding answer data through a preset big language model; and controlling the auction learning machine to output the answer data. When the main body article is determined, the main body article is displayed on the main interface, the voice input key in the main interface is triggered, the microphone of the shooting learning machine is turned on, the voice question data of the user is collected through the microphone, and the corresponding answer data is generated through the preset large language model by using the voice question data and the main body article. The problems that an existing auction learning machine is small in display interface, interaction keys cannot be arranged, and the user interaction requirement cannot be met are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and in particular to a control method, device, electronic device, and storage medium for a learning machine. Background Technology

[0002] Learning cameras are used to photograph objects, identify them, and output relevant information. However, due to their size, learning cameras cannot have a large display interface or many interactive buttons, thus failing to meet user interaction needs. Therefore, there is an urgent need for a control method for learning cameras with interactive features to solve the problems of small display interfaces and limited interactive buttons in existing learning cameras. Summary of the Invention

[0003] This invention provides a control method for a learning machine, aiming to solve the problem that existing learning machines have small display interfaces and cannot set interactive buttons, thus failing to meet user interaction needs. This invention displays the main object on the main interface when it is identified, and activates the microphone of the learning machine by triggering the voice input button on the main interface. The microphone collects the user's voice question data, and using the voice question data and the main object, generates corresponding answer data through a preset large language model. The learning machine is then controlled to output the answer data, thus solving the problem of existing learning machines having small display interfaces and not being able to set interactive buttons, thus failing to meet user interaction needs.

[0004] In a first aspect, embodiments of the present invention provide a control method for a learning machine, the method comprising the following steps:

[0005] Once the main item is identified, it is displayed on the main interface, and the microphone of the learning machine is turned on by triggering the voice input button on the main interface.

[0006] The microphone is used to collect the user's voice questions.

[0007] Based on the voice question data and the main item, corresponding answer data is generated through a preset large language model;

[0008] The camera is controlled to output the answer data.

[0009] Optionally, determining the main item includes:

[0010] Acquire image data;

[0011] The image data is processed to identify the main object, thus determining the main object in the image data.

[0012] Optionally, the step of generating corresponding answer data based on the voice question data and the main item using a preset large language model includes:

[0013] Based on the main items, the target knowledge base is determined;

[0014] Based on the target knowledge base and the voice question data, corresponding answer data is generated through a preset large language model.

[0015] Optionally, determining the target knowledge base based on the main item includes:

[0016] Based on the main item, determine the item type of the main item;

[0017] Based on the correspondence between item types and knowledge bases, the target knowledge base is determined, with different item types corresponding to different knowledge bases.

[0018] Optionally, the step of generating corresponding answer data based on the target knowledge base and the voice question data using a preset large language model includes:

[0019] Based on the target knowledge base and the voice question data, a preset prompt word for a large language model is constructed. The prompt word is used to instruct the preset large language model to call the target knowledge base to perform question-and-answer processing on the voice question data.

[0020] The prompt words are input into the preset large language model for question-and-answer processing to obtain the answer data.

[0021] Optionally, controlling the learning machine to output the answer data includes:

[0022] Based on the response data, the target selected timbre is selected in the learning machine;

[0023] The learning machine is controlled to output the response data in speech according to the selected tone of the target.

[0024] Optionally, after controlling the learning machine to select a timbre based on the target and output the response data as speech, the method further includes:

[0025] The learning machine is controlled to generate a corresponding temporary interface image based on the answer data for display.

[0026] Secondly, embodiments of the present invention provide a control device for a learning machine, the control device comprising:

[0027] The determination module is used to display the main item on the main interface when the main item is determined, and to turn on the microphone of the learning machine by triggering the voice input button on the main interface;

[0028] The acquisition module is used to acquire the user's voice question data through the microphone;

[0029] The generation module is used to generate corresponding answer data based on the voice question data and the main item, using a preset large language model.

[0030] The control module is used to control the learning machine to output the answer data.

[0031] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the control method of the learning machine provided in the embodiments of the present invention.

[0032] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the control method of the learning machine provided in the embodiments of the present invention.

[0033] In this embodiment of the invention, when the main item is identified, it is displayed on the main interface, and the microphone of the learning machine is activated by triggering the voice input button on the main interface. The microphone collects the user's voice question data. Based on the voice question data and the main item, corresponding answer data is generated using a preset large language model. The learning machine is then controlled to output the answer data. This invention solves the problem of existing learning machines having small display interfaces and lacking interactive buttons, thus failing to meet user interaction needs. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a flowchart of a control method for a learning machine provided in an embodiment of the present invention;

[0036] Figure 2This is a schematic diagram of the structure of a control device for a learning machine provided in an embodiment of the present invention;

[0037] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] like Figure 1 As shown, Figure 1 This is a flowchart of a control method for a learning machine provided in an embodiment of the present invention. The control method for the learning machine includes the following steps:

[0040] 101. When the main item is identified, display the main item on the main interface and turn on the microphone of the learning machine by triggering the voice input button on the main interface.

[0041] In this embodiment of the invention, the control method of the learning machine described above can be applied to a server, and the server and the learning machine communicate with each other. The learning machine includes a voice acquisition device and an image acquisition device. The voice acquisition device is used to acquire voice data, and the voice acquisition device can be a microphone, etc.; the image acquisition device is used to acquire image data, and the image acquisition device can be a camera, etc.

[0042] The aforementioned main items can be understood as items that need to be identified or explained by the learning machine, such as roses and apples.

[0043] The main interface described above can be understood as the main display screen of the learning machine.

[0044] The aforementioned voice input button can be understood as a button or function that allows users to input information via voice.

[0045] The microphone described above is an energy conversion device that converts sound signals into electrical signals, used to collect voice data.

[0046] It should be noted that once the main item is identified, it will be displayed on the main interface. Users can activate the microphone of the learning machine by triggering the voice input button on the main interface for voice interaction.

[0047] 102. Collect users' voice questions via microphone.

[0048] In this embodiment of the invention, the aforementioned user can be understood as a user who interacts with the learning machine.

[0049] The aforementioned voice question data is collected through the microphone of the learning machine. The voice question data can be the user's request or question to the learning machine.

[0050] 103. Based on voice question data and the main item, generate corresponding answer data through a preset large language model.

[0051] In this embodiment of the invention, prompt words for a preset large language model can be constructed based on the knowledge base corresponding to the subject item and the voice question data. These prompt words are then input into the preset large language model for question-and-answer processing to obtain question-and-answer data. The prompt words instruct the large language model to call the knowledge base corresponding to the subject item to process the voice question data. The preset large language model can be a large language model (LLM) built based on deep learning or machine learning, such as GPT, BERT, or LLaMA. The question-and-answer processing can be understood as the process of finding the most relevant answer to the voice question data from the knowledge base corresponding to the subject item.

[0052] The above answer data can be understood as the most relevant answer data obtained by analyzing the main items and the user's voice question data.

[0053] In one possible implementation, for example, if the main item is a rose and the user's voice question is "Why are roses red?", the knowledge base corresponding to the main item includes relevant knowledge about roses. A prompt word for the large language model is constructed using the user's voice question data and the knowledge base corresponding to the main item. The prompt word could be, "Please use the information from the knowledge base corresponding to the main item to answer 'Why are roses red?'". Then, the prompt word "Please use the information from the knowledge base corresponding to the main item to answer 'Why are roses red?'" is input into the large language model. The large language model will, based on "Please use the information from the knowledge base corresponding to the main item to answer 'Why are roses red?'", obtain the answer data: "The reason roses are red is because rose petals contain a substance called anthocyanin. Anthocyanin exhibits different colors depending on the pH level. In an acidic environment, anthocyanin appears red, therefore the roses appear red."

[0054] 104. Control the learning machine to output the answer data.

[0055] In this embodiment of the invention, the above output can be understood as the output format of the learning machine, which is voice output. The above voice output can be understood as the learning machine converting the response data into audible voice information and outputting it.

[0056] Specifically, the learning machine can be controlled to convert the answer data into audible voice information for output.

[0057] In this embodiment of the invention, when the main item is identified, it is displayed on the main interface, and the microphone of the learning machine is activated by triggering the voice input button on the main interface. The microphone collects the user's voice question data. Based on the voice question data and the main item, corresponding answer data is generated using a preset large language model. The learning machine is then controlled to output the answer data. This invention solves the problem of existing learning machines having small display interfaces and lacking interactive buttons, thus failing to meet user interaction needs.

[0058] It is understood that in the specific implementation of this application, data related to main items, voice data, question data, knowledge data, answer data, etc. are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required. Furthermore, the collection, use and processing of related data, as well as the training, deployment and invocation of algorithm models, must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0059] Optionally, in the step of determining the main object, image data can be acquired; the main object can be identified from the image data.

[0060] In this embodiment of the invention, the above-mentioned image data is acquired by the image acquisition device of the learning machine.

[0061] The aforementioned main image item recognition can be understood as the process of identifying the main object in image data. Specifically, image recognition technology can be used to identify the main object in image data. The image recognition model described above is trained using sample image data and corresponding main object annotation data from those sample images. This untrained image recognition model can be based on deep learning or machine learning, such as ResNet or AlexNet. Specifically, the untrained image recognition model is trained using sample image data and corresponding main object annotation data from those sample images. During training, the parameters of the image recognition model are adjusted using a minimum loss function to obtain a trained image recognition model. The main object annotation data can be information that classifies and describes the main object, such as flowers, fruits, or daily necessities. The training can be supervised training, which uses a set of known labeled data to train the model. By optimizing the model parameters, the model can predict the label of new data or make decisions based on the characteristics of existing data. The loss function is used to evaluate and optimize model performance by comparing the model's predicted value with the true value. The loss function mentioned above can be the mean squared error loss function, the cross-entropy loss function, etc.

[0062] The aforementioned main items can be understood as items that need to be identified or explained by the learning machine, such as roses and apples.

[0063] Optionally, in the step of generating corresponding answer data based on voice question data and subject items using a preset large language model, the target knowledge base can be determined based on the subject items; and the corresponding answer data can be generated based on the target knowledge base and voice question data using a preset large language model.

[0064] In this embodiment of the invention, the aforementioned target knowledge base can be a knowledge base corresponding to the main item. The aforementioned knowledge base is a pre-set knowledge base of the system, used for storing and managing knowledge, including knowledge bases for flowers, fruits, green plants, daily necessities, and school supplies, etc.

[0065] In one possible implementation, for example, when the main item is a rose, the knowledge base corresponding to the rose can be determined to be a flower knowledge base, and the flower knowledge base can be used as the target knowledge base; when the main item is an apple, the knowledge base corresponding to the apple can be determined to be a fruit knowledge base, and the fruit knowledge base can be used as the target knowledge base, and so on.

[0066] The aforementioned voice question data is collected through the microphone of the learning machine. The voice question data can be the user's request or question to the learning machine.

[0067] Furthermore, based on the target knowledge base and the voice question data, pre-defined prompt words for a large language model can be constructed. These prompt words are then input into the pre-defined large language model for question-and-answer processing to obtain question-and-answer data. The aforementioned prompt words instruct the large language model to access the knowledge base corresponding to the subject item to process the voice question data. The pre-defined large language model can be a large language model (LLM) built based on deep learning or machine learning, such as GPT, BERT, or LLaMA. The question-and-answer processing can be understood as the process of finding the most relevant answer to the voice question data from the knowledge base corresponding to the subject item.

[0068] The above answer data can be understood as the most relevant answer data obtained by analyzing the main items and the user's voice question data.

[0069] In another possible implementation, for example, if the main item is an apple and the user's voice question is "Are apples sweet?", the target knowledge base includes relevant knowledge about apples. A prompt word for the large language model is constructed using the user's voice question data and the target knowledge base. The prompt word could be, "Please use information from the target knowledge base to answer 'Are apples sweet?'". Then, the prompt word "Please use information from the target knowledge base to answer 'Are apples sweet?'" is input into the large language model. Based on this prompt, the large language model will obtain the answer data: "Apples are sweet, accompanied by a fresh and pleasant fruity aroma. The sweetness of apples comes from the sugars inside the apple, including glucose and fructose."

[0070] It should be noted that when the user's voice question is about a background object, the learning machine can add a background object image, identify the type of the background object, determine the knowledge base corresponding to the type of object, and generate corresponding question and answer data based on the corresponding knowledge base and the voice question data through a preset large language model.

[0071] Optionally, in the step of determining the target knowledge base based on the main item, the item type of the main item can be determined based on the main item; the target knowledge base can be determined based on the correspondence between the item type and the knowledge base, with different item types corresponding to different knowledge bases.

[0072] In this embodiment of the invention, the aforementioned main item may be an item that needs to be identified or explained by the learning machine.

[0073] The above item types can be understood as the classification types of the main items.

[0074] Furthermore, the item type of the main item can be determined based on the main item. For example, if the main item is a rose, and roses belong to fresh flowers, then the item type of roses is determined to be fresh flowers; if the main item is an apple, and apples belong to fruit, then the item type of apples is determined to be fruit, and so on.

[0075] The above-mentioned correspondence between item types and knowledge bases is a pre-set system mapping, with different item types corresponding to different knowledge bases. The aforementioned knowledge bases are pre-set system knowledge bases used for storing and managing knowledge, including knowledge bases for flowers, fruits, plants, and daily necessities, etc. For example, the correspondence between item types and knowledge bases could be: flowers correspond to the flower knowledge base, fruits to the fruit knowledge base, plants to the plant knowledge base, and so on.

[0076] The aforementioned target knowledge base is a knowledge base corresponding to the item type of the main item.

[0077] In one possible implementation, for example, if the main item is a rose, and roses belong to the category of fresh flowers, then the item type of rose is determined to be the category of fresh flowers. In the correspondence between item types and knowledge bases, the knowledge base corresponding to roses is determined to be the category of fresh flowers, and the knowledge base of fresh flowers is used as the target knowledge base. If the main item is an apple, and apples belong to the category of fruits, then the item type of apple is determined to be the category of fruits. In the correspondence between item types and knowledge bases, the knowledge base corresponding to apples is determined to be the category of fruits, and the knowledge base of fruits is used as the target knowledge base, and so on.

[0078] Optionally, in the step of generating corresponding answer data based on the target knowledge base and voice question data through a preset large language model, prompt words for the preset large language model can be constructed according to the target knowledge base and voice question data; the prompt words are then input into the preset large language model for question-and-answer processing to obtain the answer data.

[0079] In this embodiment of the invention, the aforementioned prompt words are used to instruct a preset large language model to call the target knowledge base to perform question-and-answer processing on the voice question data.

[0080] The aforementioned target knowledge base can be understood as a knowledge base corresponding to the item type of the main item.

[0081] The aforementioned pre-defined large language model can be a large language model built based on deep learning or machine learning, such as GPT, BERT, LLaMA, etc.

[0082] The above question-and-answer processing can be understood as the process of finding the most relevant answer to the voice question data from the target knowledge base.

[0083] The above answer data can be understood as the most relevant answer data obtained by analyzing voice question data.

[0084] Specifically, based on the target knowledge base and the current dialogue data, a pre-defined large language model can be constructed with prompt words. These prompt words are then input into the pre-defined large language model for question-and-answer processing to obtain the answer data.

[0085] In one possible implementation, for example, if the main item is a rose and the user's voice question is "Why are roses red?", the knowledge base corresponding to the main item includes relevant knowledge about roses. A prompt word for the large language model is constructed using the user's voice question data and the knowledge base corresponding to the main item. The prompt word could be, "Please use the information from the knowledge base corresponding to the main item to answer 'Why are roses red?'". Then, the prompt word "Please use the information from the knowledge base corresponding to the main item to answer 'Why are roses red?'" is input into the large language model. The large language model will, based on "Please use the information from the knowledge base corresponding to the main item to answer 'Why are roses red?'", obtain the answer data: "The reason roses are red is because rose petals contain a substance called anthocyanin. Anthocyanin exhibits different colors depending on the pH level. In an acidic environment, anthocyanin appears red, therefore the roses appear red."

[0086] In another possible implementation, for example, if the main item is an apple and the user's voice question is "Are apples sweet?", the target knowledge base includes relevant knowledge about apples. A prompt word for the large language model is constructed using the user's voice question data and the target knowledge base. The prompt word could be, "Please use information from the target knowledge base to answer 'Are apples sweet?'". Then, the prompt word "Please use information from the target knowledge base to answer 'Are apples sweet?'" is input into the large language model. Based on this prompt, the large language model will obtain the answer data: "Apples are sweet, accompanied by a fresh and pleasant fruity aroma. The sweetness of apples comes from the sugars inside the apple, including glucose and fructose."

[0087] Optionally, in the step of controlling the learning machine to output the answer data, a target selected tone can be selected in the learning machine based on the answer data; the learning machine can then be controlled to output the answer data in speech according to the target selected tone.

[0088] In this embodiment of the invention, the above-mentioned answer data is the most relevant answer data obtained by analyzing the subject item and voice question data.

[0089] The aforementioned timbres are preset by the system, including virtual character timbres, children's timbres, etc.

[0090] The selected timbre can be the timbre chosen by the user in the learning machine.

[0091] The aforementioned voice output can be the output format of the learning machine. The aforementioned voice output can be understood as the learning machine converting the answer data into audible voice information and outputting it.

[0092] In one possible implementation, for example, when the user selects a virtual character's voice as the target voice in the learning machine, the learning machine is controlled to output the answer data as speech based on the virtual character's voice; when the user selects a child's voice as the target voice in the learning machine, the learning machine is controlled to output the answer data as speech based on the child's voice.

[0093] It should be noted that when receiving the response data, the user can choose not to select a tone, and then control the learning machine to output the response data in speech according to the system's default tone.

[0094] Optionally, after controlling the learning machine to output the answer data in voice according to the selected tone of the target, the learning machine can also be controlled to generate a corresponding temporary interface image for display based on the answer data.

[0095] In this embodiment of the invention, the aforementioned temporary interface diagram can be understood as a temporary display interface. Specifically, the aforementioned interface diagram may be a text diagram corresponding to the answer data.

[0096] It should be noted that this invention allows the learning machine to simultaneously display a text image corresponding to the answer data on its main interface while the machine is playing the answer data. This invention controls the learning machine to generate a temporary text image based on the answer data and present it to the user, allowing the user to intuitively view the answer data and improving the user experience.

[0097] like Figure 2 As shown, an embodiment of the present invention provides a control device for a learning machine, the control device for the learning machine comprising:

[0098] The determination module 201 is used to display the main item on the main interface when the main item is determined, and to turn on the microphone of the learning machine by triggering the voice input button on the main interface;

[0099] The acquisition module 202 is used to acquire the user's voice question data through the microphone;

[0100] The generation module 203 is used to generate corresponding answer data based on the voice question data and the main item, using a preset large language model.

[0101] The control module 204 is used to control the learning machine to output the answer data.

[0102] Optionally, the determining module 201 is further configured to acquire image data; perform subject object recognition on the image data, and determine the subject object in the image data.

[0103] Optionally, the generation module 203 is further configured to determine a target knowledge base based on the main item; and generate corresponding answer data based on the target knowledge base and the voice question data through a preset large language model.

[0104] Optionally, the generation module 203 is further configured to determine the item type of the main item based on the main item; and to determine the target knowledge base based on the correspondence between the item type and the knowledge base, with different item types corresponding to different knowledge bases.

[0105] Optionally, the generation module 203 is further configured to construct prompt words for a preset large language model based on the target knowledge base and the voice question data. The prompt words are used to instruct the preset large language model to call the target knowledge base to perform question-and-answer processing on the voice question data. The prompt words are then input into the preset large language model for question-and-answer processing to obtain the answer data.

[0106] Optionally, the control module 204 is further configured to select a target selected tone in the learning machine based on the answer data; and control the learning machine to output the answer data in speech according to the target selected tone.

[0107] Optionally, the device is also used to control the learning machine to generate a corresponding temporary interface image for display based on the answer data.

[0108] like Figure 3 As shown, this embodiment of the invention also provides an electronic device, including a processor, which can execute any of the above-described control methods for a learning machine.

[0109] Specifically, it includes a processor 301 and a memory 302, as well as a computer program stored in the memory 302 and capable of running on the processor 301, which executes the control method for the learning machine, wherein:

[0110] The processor 301 executes the calculator program containing the control method of the learning machine stored in the memory 302, and performs the following steps:

[0111] Once the main item is identified, it is displayed on the main interface, and the microphone of the learning machine is turned on by triggering the voice input button on the main interface.

[0112] The microphone is used to collect the user's voice questions.

[0113] Based on the voice question data and the main item, corresponding answer data is generated through a preset large language model;

[0114] The camera is controlled to output the answer data.

[0115] Optionally, the determination of the principal article performed by processor 301 includes:

[0116] Acquire image data;

[0117] The image data is processed to identify the main object, thus determining the main object in the image data.

[0118] Optionally, the processor 301 executes the step of generating corresponding answer data based on the voice question data and the main item using a preset large language model, including:

[0119] Based on the main items, the target knowledge base is determined;

[0120] Based on the target knowledge base and the voice question data, corresponding answer data is generated through a preset large language model.

[0121] Optionally, the process executed by processor 301 to determine the target knowledge base based on the subject item includes:

[0122] Based on the main item, determine the item type of the main item;

[0123] Based on the correspondence between item types and knowledge bases, the target knowledge base is determined, with different item types corresponding to different knowledge bases.

[0124] Optionally, the step of processor 301 generating corresponding answer data based on the target knowledge base and the voice question data using a preset large language model includes:

[0125] Based on the target knowledge base and the voice question data, a preset prompt word for a large language model is constructed. The prompt word is used to instruct the preset large language model to call the target knowledge base to perform question-and-answer processing on the voice question data.

[0126] The prompt words are input into the preset large language model for question-and-answer processing to obtain the answer data.

[0127] Optionally, the processor 301's control of the learning machine to output the answer data includes:

[0128] Based on the response data, the target selected timbre is selected in the learning machine;

[0129] The learning machine is controlled to output the response data in speech according to the selected tone of the target.

[0130] Optionally, after controlling the learning machine to select a timbre based on the target and output the response data as speech, the method executed by the processor 301 further includes:

[0131] The learning machine is controlled to generate a corresponding temporary interface image based on the answer data for display.

[0132] This invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the control method for the learning machine provided in this invention and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0133] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0134] The above description discloses only preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.

Claims

1. A control method for a learning machine, characterized in that, The method includes the following steps: Once the main item is identified, it is displayed on the main interface, and the microphone of the learning machine is turned on by triggering the voice input button on the main interface. The microphone is used to collect the user's voice questions. Based on the voice question data and the main item, corresponding answer data is generated through a preset large language model; The camera is controlled to output the answer data.

2. The control method for the learning machine as described in claim 1, characterized in that, The identified main items include: Acquire image data; The image data is processed to identify the main object, thus determining the main object in the image data.

3. The control method for the learning machine as described in claim 1, characterized in that, The step of generating corresponding answer data based on the voice question data and the main item using a preset large language model includes: Based on the main items, the target knowledge base is determined; Based on the target knowledge base and the voice question data, corresponding answer data is generated through a preset large language model.

4. The control method for the learning machine as described in claim 3, characterized in that, The process of determining the target knowledge base based on the main item includes: Based on the main item, determine the item type of the main item; Based on the correspondence between item types and knowledge bases, the target knowledge base is determined, with different item types corresponding to different knowledge bases.

5. The control method for the learning machine as described in claim 3, characterized in that, The step of generating corresponding answer data based on the target knowledge base and the voice question data through a preset large language model includes: Based on the target knowledge base and the voice question data, a preset prompt word for a large language model is constructed. The prompt word is used to instruct the preset large language model to call the target knowledge base to perform question-and-answer processing on the voice question data. The prompt words are input into the preset large language model for question-and-answer processing to obtain the answer data.

6. The control method for the learning machine as described in claim 1, characterized in that, The control of the learning machine to output the answer data includes: Based on the response data, the target selected timbre is selected in the learning machine; The learning machine is controlled to output the response data in speech according to the selected tone of the target.

7. The control method for the learning machine as described in claim 6, characterized in that, After controlling the learning machine to output the response data as speech based on the selected timbre according to the target, the method further includes: The learning machine is controlled to generate a corresponding temporary interface image based on the answer data for display.

8. A control device for a learning machine, characterized in that, The control device for the learning machine includes: The determination module is used to display the main item on the main interface when the main item is determined, and to turn on the microphone of the learning machine by triggering the voice input button on the main interface; The acquisition module is used to acquire the user's voice question data through the microphone; The generation module is used to generate corresponding answer data based on the voice question data and the main item, using a preset large language model. The control module is used to control the learning machine to output the answer data.

9. An electronic device, characterized in that, include: The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the control method for the learning machine as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the control method for the learning machine as described in any one of claims 1 to 7.