Information processing method and device and electronic equipment

By using the first language model in the electronic device to process the information to be processed directly and sending the information that cannot be processed directly to the second language model that cannot be processed directly to the cloud server for processing, the problem of inaccurate complex information processing results is solved, and the rapid and accurate processing of multimodal information is achieved.

CN120336737APending Publication Date: 2025-07-18GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410064900.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, when processing complex pending information based on a large language model, there is a problem of inaccurate processing results, especially when the pending information contains speech or image information, or when the logical relationship between multiple subtasks cannot be understood.

Method used

The first language model is used to confirm whether the pending information can be processed directly. If it can be processed directly, the result will be generated in the electronic device. If it cannot be processed directly, it will be sent to the cloud server and processed by the second language model.

Benefits of technology

The accuracy of the processing results of multimodal pending information is improved, the limitations on the complexity of information form and content are reduced, and rapid and accurate information processing is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336737A_ABST
    Figure CN120336737A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an information processing method and device and electronic equipment. The method comprises the following steps: acquiring to-be-processed information; if it is confirmed that the to-be-processed information is information capable of being directly processed based on the first large language model, obtaining a processing result based on the to-be-processed information and the first large language model; and if the to-be-processed information is confirmed to be information which cannot be directly processed based on the first large language model, sending the to-be-processed information to a cloud server, so that the cloud server obtains a processing result based on the to-be-processed information and a second large language model. By means of the mode, the limitation on the complexity of the form and content of the obtained to-be-processed information can be reduced, and then the accuracy of the processing result of the complex to-be-processed information is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more particularly, to an information processing method, apparatus, and electronic device. Background Art

[0002] With the development of artificial intelligence technology, people have begun to rely on artificial intelligence technology to complete some daily tasks. For example, a user can input information to be processed into an electronic device (such as a computer, mobile phone, tablet, etc.) to quickly obtain a processing result corresponding to the information to be processed. In related methods, a processing result corresponding to the information to be processed can be obtained based on the information to be processed and a large language model. However, in related methods, there is still a problem that the processing results for complex information to be processed are inaccurate. Summary of the Invention

[0003] In view of the above problems, this application proposes an information processing method, apparatus, and electronic device to improve the above problems.

[0004] In a first aspect, this application provides an information processing method applied to an electronic device. The method includes: obtaining information to be processed, where the information to be processed includes one or more of text information, voice information, and image information; if it is confirmed based on a first large language model that the information to be processed is directly processable information, obtaining a processing result based on the information to be processed and the first large language model; if it is confirmed based on the first large language model that the information to be processed is not directly processable information, sending the information to be processed to a cloud server so that the cloud server obtains the processing result based on the information to be processed and a second large language model.

[0005] In a second aspect, this application provides an information processing method applied to a cloud server. The method includes: in response to receiving information to be processed from an electronic device, obtaining the processing result based on the information to be processed and a second large language model, where the information to be processed is sent to the cloud server when the electronic device confirms based on a first large language model that the information to be processed is not directly processable information.

[0006] In a third aspect, the present application provides an information processing device that runs on an electronic device. The device includes: an information acquisition unit configured to acquire information to be processed, where the information to be processed includes one or more of text information, voice information, and image information; a processing result generation unit configured to, if it is confirmed based on a first large language model that the information to be processed is directly processable information, obtain a processing result based on the information to be processed and the first large language model; and if it is confirmed based on the first large language model that the information to be processed is not directly processable information, send the information to be processed to a cloud server so that the cloud server obtains the processing result based on the information to be processed and a second large language model.

[0007] In a fourth aspect, the present application provides an information processing device that runs on a cloud server. The device includes: a processing result generation unit configured to, in response to receiving information to be processed from an electronic device, obtain the processing result based on the information to be processed and a second large language model, where the information to be processed is sent to the cloud server when the electronic device confirms based on a first large language model that the information to be processed is not directly processable information.

[0008] In a fifth aspect, the present application provides an electronic device including one or more processors and a memory; one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the above method.

[0009] In a sixth aspect, the present application provides a cloud server including one or more processors and a memory; one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the above method.

[0010] In a seventh aspect, the present application provides a computer-readable storage medium storing program code, where the above method is executed when the program code runs.

[0011] An information processing method, apparatus, electronic device, and storage medium provided by this application. After obtaining information to be processed including one or more of text information, voice information, and image information, if it is confirmed based on the first large language model that the information to be processed is directly processable information, a processing result is obtained based on the information to be processed and the first large language model; if it is confirmed based on the first large language model that the information to be processed is not directly processable information, the information to be processed is sent to a cloud server so that the cloud server obtains the processing result based on the information to be processed and the second large language model. By the above method, it is possible to obtain multi-modal information to be processed. When it is confirmed based on the first large language model in the electronic device that the multi-modal information to be processed is directly processable information, a processing result can be directly obtained. At the same time, when it is confirmed based on the first large language model in the electronic device that the multi-modal information to be processed is not directly processable information, a corresponding processing result can also be obtained based on the second large language model in the cloud, thereby reducing the restrictions on the form and content complexity of the obtained information to be processed, and further improving the accuracy of the processing result of complex information to be processed. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 The flowchart of an information processing method proposed in an embodiment of this application is shown;

[0014] Figure 2 The schematic diagram of obtaining information to be processed proposed in an embodiment of this application is shown;

[0015] Figure 3 The schematic diagram of generating a processing result proposed in an embodiment of this application is shown;

[0016] Figure 4 The flowchart of an information processing method proposed in another embodiment of this application is shown;

[0017] Figure 5 The schematic diagram of the business process of an information processing method proposed in an embodiment of this application is shown;

[0018] Figure 6 The structural block diagram of an information processing apparatus proposed in an embodiment of this application is shown;

[0019] Figure 7The structural block diagram of another information processing device proposed by an embodiment of the present application is shown;

[0020] Figure 8 The structural block diagram of an electronic device proposed by the present application is shown;

[0021] Figure 9 It is a storage unit for storing or carrying program codes for implementing the information processing method according to the embodiments of the present application. Detailed implementation manners

[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0023] With the development of technology, people begin to rely on artificial intelligence technology to complete some daily tasks. For example, a user can input information to be processed into an electronic device (such as a computer, a mobile phone, a tablet, etc.) to quickly obtain a processing result corresponding to the information to be processed. In the related art, a processing result corresponding to the information to be processed can be obtained based on the information to be processed and a large language model.

[0024] The inventors found in relevant research that there are still problems with inaccurate processing results for complex information to be processed in the related art. For example, when the information to be processed contains voice or image information, the large language may not be able to understand the content of the voice or image, resulting in inaccurate processing results. Another example is that when the information to be processed contains multiple subtasks and there is an execution order between the multiple subtasks, the large language model may not be able to understand the logical relationship between the multiple tasks, resulting in inaccurate processing results.

[0025] Therefore, the inventors propose an information processing method, device, and electronic device in this application. After obtaining the information to be processed including one or more of text information, voice information, and image information, if it is confirmed based on the first large language model that the information to be processed is directly processable information, a processing result is obtained based on the information to be processed and the first large language model; if it is confirmed based on the first large language model that the information to be processed is not directly processable information, the information to be processed is sent to the cloud server so that the cloud server obtains the processing result based on the information to be processed and the second large language model. Through the above method, it is possible to obtain multi-modal information to be processed. When it is confirmed based on the first large language model in the electronic device that the multi-modal information to be processed is directly processable information, the processing result can be directly obtained. At the same time, when it is confirmed based on the first large language model in the electronic device that the multi-modal information to be processed is not directly processable information, the corresponding processing result can also be obtained based on the second large language model in the cloud, thereby reducing the limitation on the complexity of the form and content of the obtained information to be processed, and further improving the accuracy of the processing result of complex information to be processed.

[0026] To better understand the solutions of the embodiments of this application, the following first explains the technical terms used in the embodiments of this application.

[0027] LLM (Large Language Model): It can refer to a deep learning model trained using a large amount of text data. The LLM can generate natural language text or understand the meaning of language text, and is usually used to process various natural language tasks, such as text classification, question answering, dialogue, etc. Exemplarily, the large language model can be the ChatGPT (Chat Generative Pre-trained Transformer) model.

[0028] Gorilla LLM (Large Language Model Linking Massive APIs): It can refer to a fine-tuning model based on LLaMA (Large Language and Vision Assistant), aiming to connect large language models with various services and applications provided through APIs (Application Programming Interfaces).

[0029] ICL (In-Context Learning): It can refer to a conditional probability distribution model that uses a well-trained language model to estimate under the condition of given examples. ICL can give a "prompt" to the language model, and the prompt can be a list composed of input-output pairs. An input-output pair can be an example, and this example can be used to describe a task. At the end of the prompt, there can be a test input, and the language model is required to predict the corresponding output only conditioned on the prompt.

[0030] COT (Chain-of-Throught): It can refer to a series of intermediate reasoning steps.

[0031] The following will specifically describe the embodiments of the present application in conjunction with the accompanying drawings.

[0032] Please refer to Figure 1 , an information processing method provided by an embodiment of the present application, which is applied to an electronic device. The method includes:

[0033] S110: Obtain the information to be processed, where the information to be processed includes one or more of text information, voice information, and image information.

[0034] Among them, the information to be processed can refer to the reference basis for generating the processing result, and the intention of the user can be included in the information to be processed. The intention of the user can refer to the plan for the user to achieve a certain purpose. For example, when the information to be processed is "What is the mechanism of mannitol in reducing intracranial pressure", the intention of the user can be to understand the mechanism of mannitol in reducing intracranial pressure. Another example, when the information to be processed is "What is 114 multiplied by 896", the intention of the user can be to obtain the calculation result of 114×896.

[0035] As a way, the user can directly input the information to be processed into the electronic device.

[0036] Optionally, the user can directly input the information to be processed into the electronic device by means of voice interaction. Exemplarily, the electronic device can be provided with a voice assistant. The user can call the name of the voice assistant, and after getting the response of the voice assistant, the user can directly tell the voice assistant the information to be processed by voice.

[0037] Optionally, the user can directly input the information to be processed into the electronic device by means of text input. Exemplarily, as Figure 2 shown in the left figure, the screen of the electronic device can display cards such as intelligent Q&A and map navigation. The user can click on the intelligent Q&A card to enter Figure 2 the information input interface to be processed shown in the right figure, and directly input the information to be processed into this interface by means of text input.

[0038] Optionally, the user can take a photo through an image acquisition device (such as a camera) of the electronic device or perform a QR code scan, etc., and then use the taken picture or the QR code scan result, etc. as the information to be processed.

[0039] As another way, the user can indirectly input the information to be processed to the electronic device through other devices. Exemplarily, the user can wear a VR (Virtual Reality) device (such as a VR glasses, etc.), and this VR device can make the content displayed on the electronic device three-dimensional. The user can achieve the selection operation of the content in the electronic device through head movement, etc. When the user selects a card representing intelligent Q&A, the information input interface to be processed can be entered to input the information to be processed by means of gestures, etc.

[0040] S120: If it is confirmed based on the first large language model that the information to be processed is directly processable information, based on the information to be processed and the first large language model, obtain the processing result.

[0041] Among them, the first large language model can be the large language model Gorilla LLM linked to the Speech to Text (STT) application programming interface and the Optical Character Recognition (OCR) application programming interface.

[0042] As a way, in the case where the information to be processed includes voice information and / or image information, if it is confirmed based on the first large language model that the information to be processed is directly processable information, the sub-models corresponding to the speech to text application programming interface and / or the optical character recognition application programming interface can be called based on the first large language model to obtain the text information corresponding to the voice information and / or image information in the information to be processed; generate a second prompt word based on the text information; based on the second prompt word and the first large language model, obtain the processing result.

[0043] Among them, the second prompt word can refer to the prompt word for obtaining the processing result. The second prompt word can be set with a prompt word template, and the prompt word template can include a task description part, a task example part, and an information to be processed part.

[0044] Among them, the task description part can be used to inform the first large language model of the content in the prompt that can be used as the basis for answer reasoning. Exemplarily, the task description part can be "When generating the processing result, it is necessary to refer to the task example part and the information to be processed part." The task example part can be used to enable the first large language model to learn the logical relationship between the input and the output. Exemplarily, in the scenario of mathematical calculation, the task example part can be "789×345 = 272205". The information to be processed part can be used to fill in the information to be processed. Exemplarily, in the scenario of mathematical calculation, the task example part can be "69×52 =?"

[0045] As a way, when the information to be processed only includes text information, if it is confirmed based on the first large language model that the information to be processed is directly processable information, the second prompt can be directly generated based on the information to be processed; based on the second prompt and the first large language model, the processing result is obtained.

[0046] Optionally, as Figure 3 shown, after generating the second prompt, the second prompt can be vectorized by text encoding through a Tokenizer to obtain the vector corresponding to the second prompt, and the vector is input into the first large language model to obtain the processing result.

[0047] Optionally, the first prompt can be obtained based on the information to be processed; based on the first prompt and the first large language model, the perplexity corresponding to the information to be processed is obtained. The perplexity can represent the understanding degree of the large language model for the information to be processed; if the perplexity is less than the perplexity threshold, it is confirmed that the information to be processed is directly processable information; if the perplexity is greater than or equal to the perplexity threshold, it is confirmed that the information to be processed is not directly processable information.

[0048] Among them, the first prompt can refer to the prompt used to obtain the perplexity.

[0049] Optionally, the second prompt can be the same as the first prompt. That is to say, the first prompt can be generated in the same way as the second prompt, and then based on the first prompt and the first large language model, the perplexity is obtained. When it is confirmed based on the perplexity that the information to be processed is directly processable information, the first large language model can use the first prompt as the second prompt and then output the processing result.

[0050] In the embodiments of the present application, the to-be-processed information of multiple modalities can be converted into the to-be-processed information of the text modality, so that the processing results of the to-be-processed information of multiple modalities can be obtained, and the flexibility and real-time performance of the information processing method provided by the present application can be improved. For example, in the data calculation scenario, it can help users quickly obtain calculation results through multiple input forms (text, voice, image). Compared with traditional calculators that can only obtain answers by the user clicking on the screen to input specific numbers for calculation, the information processing method provided by the present application supports users to calculate while speaking and calculate while scanning. After the chat is completed or the photo is taken, the desired answer can be obtained, saving the cumbersome processes such as taking notes or memos and manually calculating step by step in the prior art. At the same time, it can also support problems that cannot be calculated by calculators such as logical reasoning.

[0051] S130: If it is confirmed based on the first large language model that the to-be-processed information is information that cannot be directly processed, send the to-be-processed information to the cloud server, so that the cloud server obtains the processing result based on the to-be-processed information and the second large language model.

[0052] Among them, the second large language model can be the large language model Gorilla LLM linked to a large number of application programming interfaces (such as more than 1,600 APIs). The application programming interfaces of the second large language model can include third-party interfaces that require network communication links. For example, when the to-be-processed information is "What's the weather like in XXX place today", it can be linked to a third-party service provider related to weather query through the corresponding API.

[0053] As a way, if it is confirmed based on the first large language model that the to-be-processed information is information that cannot be directly processed, the to-be-processed information can be sent to the cloud server, so that the cloud server obtains the processing result based on the to-be-processed information and the second large language model.

[0054] Optionally, the application corresponding to the processing result can be automatically executed based on the processing result. That is to say, when the processing result contains services that the electronic device can provide, the electronic device can ask the user whether to provide the corresponding service in the form of text or voice, etc. If the user's feedback information is to confirm the provision of the corresponding service, the application that provides the corresponding service can be determined, and the determined application is used as the application corresponding to the processing result, so that the application corresponding to the processing result can be automatically executed.

[0055] Exemplarily, when the processing result contains "Taking a plane to XXX place on XXX date and returning from XXX place on XXX date", it can ask the user whether to book a ticket. If the user's feedback information is yes, the ticket booking APP can be automatically executed and the ticket booking operation can be performed.

[0056] Optionally, if there are multiple application programs in the electronic device that support providing the same service included in the processing result, the execution results of the multiple application programs can be compared, and based on the comparison result, the final execution result can be obtained. Among them, the selection rule for obtaining the final execution result based on the comparison result can be to select the execution result that best conforms to the user's interests. For example, the rule of the best price, the rule of the least waiting time, the principle closest to the user's habits, the principle closest to the user's preferences, etc.

[0057] In the embodiment of the present application, by automatically executing the application program corresponding to the processing result based on the processing result, the reasoning and decision-making ability and the inductive and summarizing ability of the first large language model / second large language model can be used to help the user automatically complete some things, so that the application program no longer waits passively for the user to input instructions, but automatically executes the relevant application program according to the user's intentions, habits, preferences, etc., thereby improving the intelligence level and user experience of the information processing method proposed in the present application.

[0058] An information processing method provided in this embodiment, after obtaining the information to be processed including one or more of text information, voice information, and image information, if it is confirmed based on the first large language model that the information to be processed is directly processable information, based on the information to be processed and the first large language model, a processing result is obtained; if it is confirmed based on the first large language model that the information to be processed is not directly processable information, the information to be processed is sent to the cloud server, so that the cloud server based on the information to be processed and the second large language model, obtains the processing result. Through the above method, multi-modal information to be processed can be obtained. When it is confirmed based on the first large language model in the electronic device that the multi-modal information to be processed is directly processable information, a processing result can be directly obtained. At the same time, when it is confirmed based on the first large language model in the electronic device that the multi-modal information to be processed is not directly processable information, a corresponding processing result can also be obtained based on the second large language model in the cloud, thereby reducing the limitation on the complexity of the form and content of the obtained information to be processed, and further improving the accuracy of the processing result of complex information to be processed.

[0059] Please refer to Figure 4 , an information processing method provided in the embodiment of the present application, which is applied to a cloud server, and the method includes:

[0060] S210: In response to receiving the information to be processed from the electronic device, based on the information to be processed and the second large language model, obtain the processing result, where the information to be processed is sent to the cloud server when the electronic device confirms based on the first large language model that the information to be processed is not directly processable information.

[0061] As a way, in response to receiving the information to be processed from an electronic device, if it is confirmed based on the second large language model that the information to be processed does not contain multiple subtasks, based on the information to be processed and the second large language model, obtain and send the processing result to the electronic device, and the processing result may include the intermediate process obtained based on the chain of thought. If it is confirmed based on the second large language model that the information to be processed contains multiple subtasks, the sub-models corresponding to the multiple subtasks can be confirmed based on the second large language model; based on the multiple subtasks and the sub-models corresponding to the multiple subtasks respectively, obtain the reference results corresponding to the multiple subtasks respectively; based on the reference results corresponding to the multiple subtasks respectively and the second large language model, obtain and send the processing result to the electronic device.

[0062] In the embodiments of the present application, the second large language model can confirm whether the information to be processed contains multiple subtasks and the sub-model corresponding to each subtask through its own reasoning and analysis ability. The sub-model corresponding to each subtask can be a large language model containing the corresponding API. This reasoning and analysis ability is obtained during the training process. Exemplarily, when the information to be processed is "There are 23 apples in the cafeteria. If they use 20 for lunch and buy 6 more, how many apples do they have?", it can be confirmed that the information to be processed does not contain multiple subtasks; when the information to be processed is "Going on a business trip to City A tomorrow for two days. What's the recent weather in City A? What are the recommended foods? How to get there fastest? Which hotel to stay in when arriving in City A is the closest to XXX mall?", it can be confirmed that the information to be processed contains multiple subtasks, and the multiple subtasks can be respectively: weather query, food recommendation, transportation recommendation, accommodation recommendation.

[0063] Optionally, in the case of obtaining the processing result based on the information to be processed and the second large language model, a prompt can be composed based on the information to be processed and the task example introducing the chain of thought, and the processing result can be obtained based on the composed prompt and the second large language model. Or, prompts corresponding to the multiple subtasks can be composed based on the text information corresponding to the multiple subtasks in the information to be processed and the task examples of the chain of thought corresponding to the multiple subtasks respectively, and the reference results corresponding to the multiple subtasks can be obtained based on the prompts corresponding to the multiple subtasks respectively and the sub-models; a prompt is composed based on the reference results corresponding to the multiple subtasks respectively and the task example introducing the chain of thought, and the processing result is obtained based on this prompt and the second large language model.

[0064] Exemplarily, when the information to be processed is "There are 23 apples in the cafeteria. If they use 20 for lunch and buy 6 more, how many apples do they have?", the task example introducing the chain of thought can be "Roger has 5 tennis balls. He buys two more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now? Roger started with 5 balls. 2 cans of 3 tennis balls each is 6 tennis balls. 5 + 6 = 11, the answer is 11."

[0065] In the embodiments of the present application, by introducing a chain of thought, the second large language model can achieve better results in reasoning tasks such as arithmetic reasoning, common sense reasoning, and symbolic reasoning, so as to obtain a more accurate or more user-expected processing result.

[0066] Optionally, in the process of the second large language model obtaining a processing result, multiple candidate processing results can be obtained based on different calculation or logical thinking methods, and then the candidate processing result with the largest proportion among the multiple candidate processing results is used as the processing result. Exemplarily, the multiple candidate processing results can be A, B, C, A, A, where the proportion of A is 3 / 5, and the proportions of B and C are both 1 / 5, then A can be used as the processing result.

[0067] In the embodiments of the present application, by obtaining multiple candidate processing results based on different calculation or logical thinking methods and then obtaining a processing result based on the multiple candidate processing results, the accuracy of the processing result can be improved.

[0068] Optionally, in the process of calling multiple sub-models, the text information corresponding to each of the multiple sub-tasks in the information to be processed can be used as a parameter to fill in the corresponding API function to obtain the corresponding reference result.

[0069] An information processing method provided in this embodiment enables, through the above method, the acquisition of multi-modal information to be processed. When it is confirmed by the first large language model in the electronic device that the multi-modal information to be processed is directly processable information, the processing result can be directly obtained. At the same time, when it is confirmed by the first large language model in the electronic device that the multi-modal information to be processed is not directly processable information, the corresponding processing result can also be obtained based on the second large language model in the cloud, thereby reducing the limitation on the complexity of the form and content of the acquired information to be processed, and further improving the accuracy of the processing result of complex information to be processed. Moreover, in this embodiment, when the first large language model confirms that the information to be processed is not directly processable information and the information to be processed contains multiple sub-tasks, the processing result can be obtained based on the multiple sub-tasks and the second large language model. Among them, since the second large language model can call more API sub-models than the first large language model, and each API sub-model can possess knowledge in related fields, the accuracy of the reference result corresponding to each sub-task can be improved, and thus the accuracy of the processing result can be improved without the need to retrain the second large language model. Furthermore, through the combination of the device and the cloud, the user can quickly obtain the processing result corresponding to simple information to be processed, and can obtain the processing result corresponding to complex information to be processed without the user's awareness (without the need to input other information to the electronic device).

[0070] To better understand the solution in this application, the following introduces the business process of the information processing method proposed in this application.

[0071] Please refer to Figure 5 , the to-be-processed information can be obtained based on step S1, and it is confirmed whether the to-be-processed information is directly processable information based on step S2. If so, step S3 is executed; if not, steps S4 and S5 can be sequentially executed, and when the to-be-processed information contains multiple subtasks, step S6 is executed; when the to-be-processed information does not contain subtasks, step S7 is executed.

[0072] Please refer to Figure 6 , an information processing device 600 provided by this application runs on an electronic device, and the device 600 includes:

[0073] An information acquisition unit 610, configured to acquire to-be-processed information, where the to-be-processed information includes one or more of text information, voice information, and image information;

[0074] A processing result generation unit 620, configured to, if it is confirmed based on the first large language model that the to-be-processed information is directly processable information, obtain a processing result based on the to-be-processed information and the first large language model; if it is confirmed based on the first large language model that the to-be-processed information is not directly processable information, send the to-be-processed information to a cloud server, so that the cloud server obtains the processing result based on the to-be-processed information and the second large language model.

[0075] As a way, the processing result generation unit 620 is specifically configured to obtain a first prompt word based on the to-be-processed information; obtain a perplexity corresponding to the to-be-processed information based on the first prompt word and the first large language model, where the perplexity represents the degree of understanding of the to-be-processed information by the large language model; if the perplexity is less than a perplexity threshold, confirm that the to-be-processed information is the directly processable information.

[0076] As a way, the to-be-processed information includes voice information and / or image information, and the processing result generation unit 620 is specifically configured to, if it is confirmed based on the first large language model that the to-be-processed information is directly processable information, call a sub-model corresponding to a voice-to-text application programming interface and / or an optical character recognition application programming interface based on the first large language model, to obtain text information corresponding to the voice information and / or image information in the to-be-processed information; generate a second prompt word based on the text information; obtain the processing result based on the second prompt word and the first large language model.

[0077] As a way, the processing result generation unit 620 is specifically configured to automatically execute an application program corresponding to the processing result based on the processing result.

[0078] Please refer to Figure 7 , an information processing device 800 provided by the present application runs on a cloud server. The device 800 includes:

[0079] A processing result generation unit 810 is configured to, in response to receiving information to be processed from an electronic device, obtain the processing result based on the information to be processed and a second large language model. Wherein, the information to be processed is sent to the cloud server when the electronic device determines that the information to be processed is information that cannot be directly processed based on a first large language model.

[0080] As a manner, the processing result generation unit 810 is specifically configured to, in response to receiving the information to be processed from the electronic device, if it is determined based on the second large language model that the information to be processed does not include multiple subtasks, obtain and send the processing result to the electronic device based on the information to be processed and the second large language model. The processing result includes an intermediate process obtained based on the chain of thought.

[0081] Optionally, the processing result generation unit 810 is specifically configured to, if it is determined based on the second large language model that the information to be processed includes multiple subtasks, determine sub-models corresponding to the multiple subtasks based on the second large language model; obtain reference results corresponding to the multiple subtasks based on the multiple subtasks and the sub-models corresponding to the multiple subtasks; obtain and send the processing result to the electronic device based on the reference results corresponding to the multiple subtasks and the second large language model.

[0082] Next, an electronic device provided by the present application will be described in conjunction with Figure 8 the following.

[0083] Please refer to Figure 8 , based on the above information processing method and device, another electronic device 100 that can execute the foregoing information processing method is further provided in an embodiment of the present application. The electronic device 100 includes a processor 102 and a memory 104. Wherein, a program that can execute the content in the foregoing embodiment is stored in the memory 104, and the processor 102 can execute the program stored in the memory 104.

[0084] Among them, the processor 102 may include one or more processing cores. The processor 102 connects various parts within the entire electronic device 100 through various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 104, and by invoking the data stored in the memory 104, it performs various functions of the electronic device 100 and processes data. Optionally, the processor 102 may be implemented in at least one hardware form of a network processor (Neural network Processing Unit, NPU), a digital signal processing (Digital Signal Processing, DSP), a field-programmable gate array (Field-Programmable Gate Array, FPGA), or a programmable logic array (Programmable Logic Array, PLA). The processor 102 may integrate one or a combination of several of a central processing unit (Central Processing Unit, CPU), a graphics processing unit (Graphics Processing Unit, GPU), a neural network processing unit (Neural networkProcessing Unit, NPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the displayed content; the NPU is responsible for processing multimedia data such as videos and images; the modem is responsible for processing wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 102 and may be implemented separately through a communication chip.

[0085] The memory 104 may include a random access memory (Random Access Memory, RAM), and may also include a read-only memory and a double data rate synchronous dynamic random access memory (Double DataRate, DDR). The memory 104 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store the data created during the use of the electronic device 100 (such as a phone book, audio and video data, chat record data, etc.).

[0086] Please refer to Figure 9, which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable storage medium 1000, and the program code can be called by a processor to execute the method described in the above method embodiment.

[0087] The computer-readable storage medium 1000 can be an electronic memory such as a flash memory, EEPROM (electrically erasable programmable read-only memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 1000 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 1000 has a storage space for the program code 1010 that executes any method step in the above method. These program codes can be read from or written into one or more computer program products. The program code 1010 can be compressed in an appropriate form, for example.

[0088] In summary, for an information processing method, device, and electronic device provided by the present application, after obtaining the to-be-processed information including one or more of text information, voice information, and image information, if it is confirmed based on the first large language model that the to-be-processed information is directly processable information, a processing result is obtained based on the to-be-processed information and the first large language model; if it is confirmed based on the first large language model that the to-be-processed information is not directly processable information, the to-be-processed information is sent to a cloud server so that the cloud server obtains the processing result based on the to-be-processed information and the second large language model. By the above method, it is possible to obtain multi-modal to-be-processed information. When it is confirmed based on the first large language model in the electronic device that the multi-modal to-be-processed information is directly processable information, a processing result can be directly obtained. At the same time, when it is confirmed based on the first large language model in the electronic device that the multi-modal to-be-processed information is not directly processable information, a corresponding processing result can also be obtained based on the second large language model in the cloud, thereby reducing the limitation on the complexity of the form and content of the obtained to-be-processed information, and further improving the accuracy of the processing result of complex to-be-processed information.

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An information processing method, characterized in that, Applied to an electronic device, the method includes: Obtain the information to be processed, where the information to be processed includes one or more of text information, voice information, and image information; If it is confirmed based on the first large language model that the information to be processed is directly processable information, obtain a processing result based on the information to be processed and the first large language model; If it is confirmed based on the first large language model that the information to be processed is not directly processable information, send the information to be processed to a cloud server so that the cloud server obtains the processing result based on the information to be processed and a second large language model.

2. The method according to claim 1, wherein Before the step of "if it is confirmed based on the first large language model that the information to be processed is directly processable information, obtain a processing result based on the information to be processed and the first large language model", it further includes: Obtain a first prompt word based on the information to be processed; Obtain the perplexity corresponding to the information to be processed based on the first prompt word and the first large language model, where the perplexity represents the degree of understanding of the information to be processed by the large language model; If the perplexity is less than a perplexity threshold, confirm that the information to be processed is the directly processable information.

3. The method according to claim 1, characterized in that When the information to be processed includes voice information and / or image information, and the step of "if it is confirmed based on the first large language model that the information to be processed is directly processable information, obtain a processing result based on the information to be processed and the first large language model" includes: If it is confirmed based on the first large language model that the information to be processed is directly processable information, call a sub-model corresponding to a speech-to-text application programming interface and / or an optical character recognition application programming interface based on the first large language model to obtain text information corresponding to the voice information and / or image information in the information to be processed; Generate a second prompt word based on the text information; Obtain the processing result based on the second prompt word and the first large language model.

4. The method according to any one of claims 1-3, characterized in that, The method further includes: Automatically execute the application program corresponding to the processing result based on the processing result.

5. An information processing method, characterized in that, Applied to a cloud server, the method includes: In response to receiving the information to be processed from an electronic device, obtain the processing result based on the information to be processed and a second large language model, where the information to be processed is sent to the cloud server when the electronic device confirms based on the first large language model that the information to be processed is not directly processable information.

6. The method according to claim 5, characterized in that The step of "in response to receiving the information to be processed from an electronic device, obtain the processing result based on the information to be processed and a second large language model" includes: In response to receiving the information to be processed from the electronic device, if it is confirmed based on the second large language model that the information to be processed does not include multiple subtasks, obtain and send the processing result to the electronic device based on the information to be processed and the second large language model, where the processing result includes an intermediate process obtained based on a chain of thought.

7. The method according to claim 6, wherein The method further includes: If it is confirmed based on the second large language model that the information to be processed includes multiple subtasks, confirm the sub-models corresponding to the multiple subtasks based on the second large language model; Obtain the reference results corresponding to the multiple subtasks based on the multiple subtasks and the sub-models respectively corresponding to the multiple subtasks; Based on the reference results corresponding to the multiple subtasks and the second large language model, obtain and send the processing result to the electronic device.

8. An information processing apparatus, characterized in that, Running on an electronic device, the device includes: An information acquisition unit, configured to acquire information to be processed, where the information to be processed includes one or more of text information, voice information, and image information; A processing result generation unit, configured to, if it is confirmed based on the first large language model that the information to be processed is directly processable information, obtain a processing result based on the information to be processed and the first large language model; if it is confirmed based on the first large language model that the information to be processed is not directly processable information, send the information to be processed to a cloud server so that the cloud server obtains the processing result based on the information to be processed and the second large language model.

9. An information processing apparatus, characterized in that, Running on a cloud server, the device includes: A processing result generation unit, configured to, in response to receiving the information to be processed from an electronic device, obtain the processing result based on the information to be processed and the second large language model, where the information to be processed is sent to the cloud server when the electronic device confirms based on the first large language model that the information to be processed is not directly processable information.

10. An electronic device, characterized in that, Comprising one or more processors and a memory; One or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1-4.

11. A cloud server, characterized in that, Comprising one or more processors and a memory; One or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 5-7.

12. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, where the method according to any one of claims 1-4 or the method according to any one of claims 5-7 is executed when the program code runs.