Information processing method

By obtaining the correlation between problem information and environmental information, and using sensing elements to generate prompt information and input it into the artificial intelligence model, the problem that the artificial intelligence system in the existing technology fails to combine real-time environmental information is solved, and a more accurate and intelligent response is achieved.

CN120688648APending Publication Date: 2025-09-23LENOVO (BEIJING) LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510703666.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In existing technologies, artificial intelligence systems fail to effectively incorporate real-time environmental information when generating answers, resulting in the answers lacking adaptability to actual scenarios, affecting the accuracy and practicality of information processing.

Method used

By obtaining the association between problem information and current environmental information, using sensing elements to obtain environmental information, and fusing it with problem information to generate prompt information, which is input into the artificial intelligence model to generate a response that is closer to the real scene.

Benefits of technology

It improves the environmental perception capability of the artificial intelligence system and its practicality in different application scenarios, and enhances the accuracy and intelligence level of answering questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688648A_ABST
    Figure CN120688648A_ABST
Patent Text Reader

Abstract

The invention discloses an information processing method. The method comprises the following steps: acquiring question information; the problem information is determined to meet a first condition, and the first condition represents that the problem information is associated with current environment information; obtaining at least one kind of environmental information associated with the problem information, wherein the at least one kind of environmental information is obtained based on the environmental information of the electronic equipment sensed by a sensing element; generating prompt information based on the at least one kind of environment information and the problem information; the prompt information is input into an artificial intelligence model, first generation information is obtained, and the first generation information is used for responding to the question information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence, and are related to but not limited to an information processing method. Background Art

[0002] In modern smart devices, information processing technologies are widely used in areas such as human-computer interaction, environmental perception, and response decision-making. By integrating sensor data and user input, electronic devices can more accurately understand the current situation and respond accordingly. With the development of artificial intelligence models, the ability to generate context-based information is continuously improving, making personalized services and intelligent responses possible.

[0003] Current AI applications primarily focus on providing answers to specific questions. This involves obtaining user information and responding based on pre-set rules or general models. For example, some systems utilize built-in natural language processing models to analyze the content of questions and generate semantically appropriate responses. However, these approaches often overlook the specific environment in which the device is located, resulting in responses that lack adaptability to real-world scenarios.

[0004] Because real-time environmental information is not incorporated into the generation of responses, truly context-aware responses are difficult to achieve, which in turn affects the accuracy and practicality of information processing. How to effectively obtain real-time scene information and improve the accuracy of responses to questions has become a pressing technical challenge. Summary of the Invention

[0005] In view of this, an embodiment of the present application provides an information processing method.

[0006] The technical solution of the embodiment of the present application is implemented as follows:

[0007] In a first aspect, an embodiment of the present application provides an information processing method, including:

[0008] Get problem information;

[0009] Determining that the problem information satisfies a first condition, wherein the first condition indicates that the problem information is associated with current environment information;

[0010] Acquire at least one type of environmental information associated with the problem information, where the at least one type of environmental information is obtained based on environmental information of the electronic device being sensed by a sensing element;

[0011] generating prompt information based on at least one of the environmental information and the question information;

[0012] The prompt information is input into the artificial intelligence model to obtain first generated information, wherein the first generated information is used to respond to the question information.

[0013] In a second aspect, an embodiment of the present application provides an information processing device, including:

[0014] A first acquisition module is used to obtain problem information;

[0015] a determination module, configured to determine that the problem information satisfies a first condition, wherein the first condition indicates that the problem information is associated with current environment information;

[0016] A second acquisition module is configured to acquire at least one type of environmental information associated with the problem information, wherein the at least one type of environmental information is obtained based on environmental information of the electronic device being sensed by a sensing element;

[0017] A first generating module, configured to generate prompt information based on at least one piece of environmental information and question information;

[0018] The first information processing module is used to input the prompt information into the artificial intelligence model to obtain first generated information, wherein the first generated information is used to respond to the question information.

[0019] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor obtains problem information when executing the program; determines that the problem information satisfies a first condition, wherein the first condition indicates that the problem information is associated with current environmental information; obtains at least one environmental information associated with the problem information, wherein the at least one environmental information is obtained based on environmental information of the electronic device sensed by a sensing element; generates prompt information based on the at least one environmental information and the problem information; and inputs the prompt information into an artificial intelligence model to obtain first generated information, wherein the first generated information is used to respond to the problem information.

[0020] In a fourth aspect, an embodiment of the present application provides a storage medium storing executable instructions for obtaining problem information when executed by a processor; determining that the problem information satisfies a first condition, wherein the first condition indicates that the problem information is associated with current environmental information; obtaining at least one type of environmental information associated with the problem information, wherein the at least one type of environmental information is obtained based on environmental information of the electronic device sensed by a sensing element; generating prompt information based on the at least one type of environmental information and the problem information; and inputting the prompt information into an artificial intelligence model to obtain first generated information, wherein the first generated information is used to respond to the problem information.

[0021] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising a computer program or instructions, which, when executed by a processor, realizes obtaining problem information; determining that the problem information satisfies a first condition, wherein the first condition indicates that the problem information is associated with current environmental information; obtaining at least one environmental information associated with the problem information, wherein the at least one environmental information is obtained based on environmental information of the electronic device sensed by a sensing element; generating prompt information based on the at least one environmental information and the problem information; inputting the prompt information into an artificial intelligence model to obtain first generated information, wherein the first generated information is used to respond to the problem information. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A schematic diagram of an implementation flow of an information processing method provided in an embodiment of the present application;

[0023] Figure 2 A schematic diagram of a process for determining that question information meets the first condition provided in an embodiment of the present application;

[0024] Figure 3 A schematic diagram of an implementation flow of generating prompt information provided in an embodiment of the present application;

[0025] Figure 4 A schematic diagram of the processing flow for achieving real-time perception by the AI ​​system provided in an embodiment of the present application;

[0026] Figure 5 A schematic diagram of the structure of an information processing device provided in an embodiment of the present application;

[0027] Figure 6 A schematic diagram of a hardware entity of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] To make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the specific technical solutions of the embodiments of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.

[0029] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0030] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0032] The present application provides an information processing method, which is applied to electronic devices such as Figure 1 As shown, the method includes:

[0033] Step S110: Obtaining problem information;

[0034] Here, question information refers to the question that a user inputs or the system receives, requiring an answer. It is typically presented in the form of natural language text and can be input in either text or voice. There are no restrictions on the input format of question information. Examples of question information include: "Please count the number of people in the classroom," "What color lipstick is the lady across from you wearing?" This question information serves as the basis for triggering subsequent processing flows.

[0035] During implementation, question information can be provided by the user through voice, keyboard input, image recognition, or automatically generated by the system based on context. For example, in an educational setting, a teacher could issue a headcount instruction via voice; in a social setting, a user could submit a text request to identify a lipstick color.

[0036] In actual implementation, electronic devices can receive user question information through various methods such as voice recognition modules, text input interfaces, image recognition modules, etc., and convert them into structured text data as a basis for subsequent judgment and processing.

[0037] Step S120: Determine whether the problem information satisfies a first condition, wherein the first condition indicates that the problem information is associated with current environment information;

[0038] Here, the first condition may refer to whether the question information is directly or indirectly related to the current physical environment. If it is determined that the question information meets the first condition, it can be determined that the question information is related to the current environment information. The first condition can be preset.

[0039] During implementation, if the problem information involves the perception, judgment, or operation of real-world scenarios, it is considered to meet the first condition. For example, counting the number of people in a classroom requires relying on image information captured by a camera; identifying lipstick shades requires relying on color information in an image; and recommending music that suits the current mood may require combining multiple environmental information such as sound, lighting, and motion status.

[0040] In actual implementation, the system can use a preset keyword matching mechanism, semantic analysis model, or rule engine to determine whether the question information meets the first condition. For example, if the question information contains words such as "identify," "inventory," "recommend," and "current," the system can preliminarily determine that the question needs to be handled in conjunction with the context.

[0041] For example, if the question information is "counting the number of people attending the meeting on site", it can be determined that the question information meets the first condition, that is, based on the keywords "site" and "count", it is determined that the question information is associated with the information of the people currently attending the meeting on site.

[0042] When the question information is "get the lipstick color of the lady opposite", it can be determined that the question information meets the first condition, that is, based on the keyword "lady opposite", it is determined that the question information is associated with the information of the lady opposite.

[0043] Step S130: Acquire at least one type of environmental information associated with the problem information, wherein the at least one type of environmental information is obtained based on environmental information of the electronic device being sensed by a sensing element;

[0044] Here, sensing elements refer to various hardware modules embedded in electronic devices for sensing the external environment, such as cameras, microphones, accelerometers, gyroscopes, infrared sensors, light sensors, etc. These components can obtain environmental data in real time and serve as an important source of environmental information in the present invention. Different problem information may require calling different types of sensing elements. For example, an audio acquisition device can be used to obtain sound information; a camera can collect pictures or videos; a time of flight (TOF) sensor can collect depth information; a heart rate sensor can collect heartbeat information; a visual sensor (EVS) can collect motion information; and a brightness / temperature sensor can collect brightness / temperature.

[0045] Environmental information refers to information related to the current physical environment collected by various sensing elements on electronic devices (such as cameras, microphones, depth sensors, EVS sensors, temperature sensors, etc.).

[0046] During implementation, at least one type of environmental information associated with the problem information can be obtained based on the environmental information of the electronic device sensed by the sensing element. This information can include images, audio, light intensity, temperature, depth data, motion status, and other information, which is used to supplement the content of the problem information and enable it to have the ability to perceive the real scene. For example, a camera provided on the electronic device can be used to capture image or video information of the environment in which the electronic device is located to obtain image or video information associated with the problem information.

[0047] For example, the wide-angle camera device can be activated to capture photos of the conference site to obtain photos of the conference site associated with the question information "Count the number of people attending the meeting on site"; in the scenario of identifying lipstick color, the system can call the image sensor to obtain the facial image of the target person; the heart rate sensor can be activated to capture heartbeats to obtain user heartbeat information associated with the question information "Recommend suitable songs based on my current state."

[0048] In actual implementation, the system can dynamically select and combine multiple sensing elements based on the content and type of the query information, thereby achieving comprehensive perception of complex environments. For example, in a personalized music recommendation scenario, the system can simultaneously use the camera to obtain the user's facial expression information, the microphone to obtain the surrounding sound information, and the light sensor to obtain the ambient brightness information to comprehensively judge the user's emotional state and preferences.

[0049] Step S140: Generate prompt information based on the at least one piece of environmental information and the problem information;

[0050] Here, prompt information is enhanced input information generated by integrating relevant context information with the question information. This prompt information, as input to the AI ​​model, can guide the model to generate more accurate and context-aware responses based on the relationship between the question and the context. Prompt information can be structured text, multimodal data, or a combination thereof.

[0051] During the implementation process, the environmental information can be converted into text descriptions and combined with the problem information to generate prompt information.

[0052] For example, in the scenario of counting the number of people in the classroom, the system can fuse the question information "Please count the number of people in the classroom" with the classroom image obtained from the camera to generate a prompt message such as "Please count the number of people in the classroom according to the picture below"; in the scenario of identifying lipstick colors, the system can fuse the question information "Identify the lipstick color of the lady opposite" with the facial image obtained from the camera to generate a prompt message such as "Please identify the lipstick color of the lady in the picture below".

[0053] In actual implementation, the system can use natural language processing technology, image description generation technology or multimodal fusion technology to intelligently splice and integrate problem information with environmental information to generate prompt information that conforms to the input format of the artificial intelligence model.

[0054] Step S150: input the prompt information into the artificial intelligence model to obtain first generated information, wherein the first generated information is used to respond to the question information.

[0055] Here, artificial intelligence models generally refer to machine learning models used to handle natural language understanding and generation tasks, such as large language models based on the Transformer architecture. The model can generate answers based on the input prompt information, that is, the first generated information, thereby realizing intelligent response to questions. The prompt information is the question information input by the user that has been preprocessed, that is, the input information combined with environmental information. For example, in the scenario of counting the number of students in the classroom, the artificial intelligence model can output that there are 32 students in the classroom; in the scenario of identifying the color of lipstick, the artificial intelligence model can output that the lipstick color used by the other lady is XX.

[0056] In some embodiments, the system can also automatically generate some guiding text based on the prompt information to help the artificial intelligence model better understand the user's intentions.

[0057] The processed prompt information is input into a pre-trained artificial intelligence model. The artificial intelligence model analyzes and processes the input prompt information and uses the knowledge and patterns it has learned to generate first generated information related to the question information.

[0058] Here, the AI ​​model can use the generated first information as output to directly or indirectly respond to the user's question information. The AI ​​model can also execute specific applications based on the prompt information. For example, if the first generated information is "Recommended purchase link for lipstick color XX," it can link to a shopping platform for purchasing that lipstick color.

[0059] In actual implementation, AI models can be deployed on cloud servers or run on local devices, depending on the application scenario and the availability of computing resources. Regardless of the deployment method, the core function of the AI ​​model is to generate first-hand information that meets the requirements of the problem based on the prompt information, thereby achieving an intelligent response to the problem.

[0060] In this embodiment, the system first determines whether the question information meets the conditions related to the current environment. If so, it uses sensing elements to obtain the corresponding environmental information. This information is then integrated with the question information to generate prompt information, which is then input into the artificial intelligence model to generate a response that is more realistic. This not only improves the artificial intelligence system's environmental perception capabilities, but also enhances its practicality and intelligence in different application scenarios, effectively improving the accuracy of responses to questions.

[0061] In some embodiments, as Figure 2 As shown, in the above step S120, "determining that the question information satisfies the first condition" can be achieved by the following steps:

[0062] Step S210: parsing the question information to obtain at least one target information in the question information that matches preset matching information;

[0063] Here, preset matching information refers to a set of key information types pre-set by the system for identifying specific semantics or scenarios, including object information, orientation information, time information, and location information. These information types can exist independently or be used in combination to determine whether the current question requires the combination of real-time perception data to generate an answer. For example, the preset matching information can be "mine", "now", "work environment", etc. When there is target information in the question information that matches the preset matching information, and there is no contextual description of the target information, it can indicate the need to obtain current environmental information and turn on the camera or other perception devices.

[0064] Target information is the specific content extracted from the question that matches the preset matching information. For example, if the question involves counting the number of people in the classroom at 3 p.m., then "classroom" is the target information, "3 p.m. today" is the time information, and "classroom" is the location information. The system uses natural language processing technology to perform word segmentation, syntactic analysis, and semantic understanding on the question to extract this target information. It then compares this information with the preset matching information to determine whether the question requires real-world context to answer.

[0065] The parsing process can be performed by an AI model, which can be deployed on a local device (such as a mobile phone or tablet) or run on a cloud server. Regardless of the method, the parsing results will influence whether the sensor module is subsequently used for data collection. This approach enables the system to make dynamic decisions, intelligently determining whether an answer requires incorporating external environmental information based on the question content.

[0066] Step S220: Determine whether the first condition is met based on at least one target information;

[0067] The preset matching information includes at least one of the following: object information, orientation information, time information and location information.

[0068] When the system's parsed target information contains any of the preset matching information, it determines that the question meets the first condition, triggering the subsequent sensor call process. For example, if the question contains information such as lipstick color, woman opposite, and current time, the corresponding object information, direction information, and time information will be used by the system to determine whether to activate the camera, audio, and time synchronization modules to obtain the corresponding real-time scene data.

[0069] In the pre-set matching information, object information refers to information related to a specific individual, object, or entity, such as a person's name and identity, or an item's name and model. This information can include keywords such as "my" and "Ms." Position information describes location, direction, and spatial relationships, including geographic coordinates and relative positions, such as "across from," "behind," and "to the left." Time information includes various data related to time, such as "now" and "on the spot." Location information refers to information related to a specific location or place, such as "work environment" and "meeting site."

[0070] In actual applications, for example, when a user asks, "What's my current lipstick shade?" the system first interprets the lipstick shade as target information, identifying it as a type of object information and confirming that the question meets the first condition. It then uses the camera to capture the user's facial image, analyzes the lip color using image recognition technology, and ultimately recommends a specific shade.

[0071] In the embodiment of the present application, by parsing the question information and extracting the target information, it is determined whether the question meets the first condition in combination with the preset matching information. In this way, the information matching-based mechanism improves the response accuracy and efficiency of the system, avoids unnecessary hardware calls, and saves resource overhead. At the same time, it also enhances the intelligence of the system, enabling it to flexibly adjust the processing strategy according to different questions, thereby providing more accurate answers combined with real-world scenarios. Automatic classification of question types is achieved, so that it can be determined whether to call sensors for real-time data collection, thereby improving the AI ​​system's perception of real-world scenarios and the quality of its answers.

[0072] In some embodiments, the step S130 of "obtaining at least one piece of environmental information associated with the problem information" may be implemented by the following steps:

[0073] Step 131: determining a sensing element that matches the target information;

[0074] Here, the target information is the specific content extracted from the question information that matches the preset matching information, such as the user's current scenario needs, intentions or behavior patterns. Sensing elements refer to hardware modules that can perceive the external environment and provide data input, such as cameras, microphones, accelerometers, gyroscopes, temperature sensors, etc. The system can intelligently select the most relevant sensing elements based on the content of the target information to ensure that the collected environmental information can effectively support the AI ​​model's answer to the question. For example, if the question is "What color lipstick is the lady opposite wearing?", the system will give priority to calling the front camera; if the question is "What is the current indoor temperature?", the temperature sensor will be called.

[0075] This step improves the pertinence and efficiency of environmental information collection through an intelligent matching mechanism, avoids unnecessary waste of resources, and improves response speed and accuracy.

[0076] Step 132: sensing the environment information of the electronic device based on the sensing element, wherein, when the electronic device is arranged with multiple sensing elements of the same type in different directions, it includes determining the sensing element in a direction that matches the target information.

[0077] When an electronic device (such as a mobile phone or tablet) has multiple built-in sensing elements of the same type (such as two front-facing cameras and multiple microphones), and these elements are distributed in different directions, the system will further determine which direction of the sensing element best meets the needs of the current target information. For example, if the user asks "What is the lipstick color of the lady opposite me", the system can choose the rear camera for detection instead of the front camera. If the user asks "Recommend a song that suits me", the system can choose the front camera for detection. This process relies on the sensor layout information inside the device and the ability to understand the user's semantics.

[0078] This step enhances the flexibility and adaptability of multi-sensor collaborative work, enabling the system to accurately obtain required information in complex environments, thereby improving overall perception capabilities and user experience.

[0079] In the embodiments of the present application, by intelligently matching target information with sensing elements and further selecting the appropriate direction when there are multiple sensing elements of the same type, the accuracy and efficiency of environmental information collection can be improved, thereby enhancing the AI ​​system's perception of real-world scenes, and thus achieving a more natural and intelligent human-computer interaction experience.

[0080] In some embodiments, the above step S220 "determining that the first condition is satisfied based on at least one piece of target information" can be implemented by the following steps:

[0081] Step 221: Determine the context information corresponding to each target information;

[0082] Here, contextual information refers to the surrounding or background information associated with the target information, which helps to more fully understand the target information's content and meaning. For example, if a user asks for a song recommendation, the target information is the song recommendation, while the contextual information may include the current time, weather conditions, the user's emotional state, recent play history, and more. By acquiring this contextual information, the system can provide more personalized and accurate responses based on the user's real-time status.

[0083] During implementation, the semantic understanding model can be used to obtain the contextual information corresponding to the target information. For example, for the sentence “I am in a good mood, please recommend some suitable songs”, the contextual information can be “I am in a good mood”.

[0084] In some embodiments, there is no matching context information for the target information. For example, for the query "Please recommend a song that suits me", if the target information is "me", there is no context information.

[0085] Step 222: When it is determined that the context information is less than a preset information amount, determine whether the target information corresponding to the context information meets the first condition.

[0086] Here, the "preset amount of information" refers to the system's pre-defined minimum valid information standard, used to determine whether the currently acquired contextual information is sufficient to support a decision. If the contextual information is insufficient, the system will assume that the target information already meets a certain basic condition, namely the first condition. For example, if no user emotion or age information is available, the system may determine that the target information corresponding to the contextual information meets the first condition.

[0087] When contextual information is limited, it can be introduced by identifying the current environment to improve the system's understanding of user intent, making responses more targeted and context-adaptive. This improves the user experience, enhances the naturalness and practicality of system interactions, and ultimately enables truly intelligent responses.

[0088] In an embodiment of the present application, by determining the contextual information of the target information and judging whether the first condition is met based on a preset amount of information, the system's understanding accuracy and response efficiency of the user's intention can be improved, thereby enhancing the intelligent perception capability of the AI ​​system and enabling better adaptation to diverse application scenarios.

[0089] In some embodiments, the present application provides a method for generating preset matching information, which can be implemented by the following steps:

[0090] Step S160: Obtain initial matching information;

[0091] Initial matching information refers to the basic feature data extracted based on the current scene or user input, which is used for subsequent model processing. For example, the initial preset matching information can be "my", "now", "work environment", etc.

[0092] Step S170: Generate the preset matching information using a generative adversarial network using the initial matching information.

[0093] Here, a Generative Adversarial Network (GAN) is a deep learning model consisting of a generator and a discriminator. The generator is responsible for generating synthetic data similar to real data, while the discriminator is responsible for determining whether the data is real data. GAN is used to enhance and optimize the initial matching information to generate more accurate and representative preset matching information.

[0094] In real-time, the system first inputs initial matching information into the generator. Based on the trained model structure, the generator generates matching data that meets the target characteristics. The discriminator then evaluates the generated data to determine its authenticity or quality level. If it meets the requirements, it outputs the final preset matching information. Otherwise, the generator parameters are further adjusted until the desired effect is achieved.

[0095] In the embodiments of the present application, initial matching information is obtained and combined with a generative adversarial network to generate preset matching information. This allows the system to automatically generate high-quality matching data without relying on large amounts of manually annotated data, thereby improving recognition accuracy and response efficiency. Furthermore, because the generative adversarial network has good generalization capabilities, the system can flexibly adapt to different scenarios, achieving more intelligent and accurate matching capabilities.

[0096] In some embodiments, the present application provides a method for training a semantic understanding model, which can be implemented by following steps A or B:

[0097] Step A: training a semantic understanding model using the preset matching information to obtain a first semantic understanding model; wherein the first semantic understanding model is used to determine the association between the question information and the current environment information;

[0098] Here, the semantic understanding model is an AI model that uses natural language processing technology to perform semantic analysis on input questions and determine whether the question is related to the current environmental information.

[0099] During implementation, the preset matching information may be used as training data to train a semantic understanding model, so as to obtain a first semantic understanding model capable of determining the association between the problem information and the current environment information.

[0100] Step B: training a semantic understanding model using the preset matching information and a weight parameter of each preset matching information to obtain the first semantic understanding model;

[0101] The weight parameter here is the importance coefficient assigned to each preset matching information, which is used to influence the degree of attention the model pays to different matching relationships during the learning process. For example, if certain matching relationships are more common or more important, their weights can be increased to enhance the model's learning effect. In actual applications, the weight parameter can be dynamically adjusted based on factors such as usage frequency in historical data, user feedback, and business needs.

[0102] During implementation, the preset matching information and the weight parameters of each preset matching information may be used as training data to train a semantic understanding model, so as to obtain a first semantic understanding model capable of determining the association between question information and current environment information.

[0103] In this way, the accuracy and adaptability of the semantic understanding model can be improved, so that it can better judge whether the problem information is related to the current environmental information, thereby improving the real-time perception and personalized service capabilities of the AI ​​system.

[0104] Correspondingly, the above step S220 of "determining whether the first condition is satisfied based on at least one piece of target information" can be implemented by the following process:

[0105] Input at least one of the target information into the first semantic understanding model to determine whether the first condition is met.

[0106] Here, the target information is the key information fragment extracted from the original question and used as input for the semantic understanding model. For example, if the original question is "Recommend a song to listen to now," the target information may include keywords such as "recommend," "suitable for listening now," and "song." After inputting this information into the first semantic understanding model, the model outputs a judgment result, determining whether the question meets the first condition, that is, whether it is relevant to the current context information.

[0107] During the implementation process, the first semantic understanding model can perform semantic analysis on the target information and output a probability value or a binary classification result to indicate whether the problem information is related to the current environmental information.

[0108] In an embodiment of the present application, by introducing a semantic understanding model and combining it with preset matching information and weight parameters for training, and using the trained first semantic understanding model to determine whether the question information is related to the current environmental information, the ability to judge the association between the question and the environmental information can be significantly improved.

[0109] In some embodiments, the present application provides a method for training a first semantic understanding model to obtain a second semantic understanding model, which can be implemented by the following process:

[0110] Training the first semantic understanding model using the preset matching information and the sensing element information associated with each of the preset matching information to obtain a second semantic understanding model;

[0111] Wherein, the second semantic understanding model is used to determine the sensing element information for collecting the environmental information;

[0112] The first semantic understanding model is a basic natural language processing model used to understand the meaning of user input questions and determine whether the question information is relevant to the environment. The second semantic understanding model is further trained on this basis to enable it to connect semantic information with actual available hardware devices. By using preset matching information and sensor element information as training data, the model can learn the mapping relationship between specific questions and specific devices.

[0113] Correspondingly, the step S130 of "obtaining at least one piece of environmental information associated with the problem information" can be implemented by the following steps:

[0114] Step 133: Input at least one target information into the second semantic understanding model to determine at least one target sensing element to be activated;

[0115] Here, target information refers to the specific content extracted from the question information that matches the preset matching information. It can refer to the key semantic elements contained in the user's question, such as keywords, intent, and context. For example, the key information in the question "Who is the stranger across the street?" is "Who is the stranger?" and "Who is the lipstick color."

[0116] The target sensing element is a hardware device selected by the system based on the output of the semantic model, such as a camera, microphone, temperature sensor, etc. For example, in the above example, the system can choose to start the rear camera to capture images.

[0117] Step 134: Start at least one of the target sensing elements to obtain at least one piece of environmental information associated with the problem information.

[0118] After determining the sensing elements that need to be activated, the system can send control instructions to the corresponding hardware modules to start their workflow and collect data. For example, starting a camera to capture the current image or starting a temperature sensor to read the ambient temperature.

[0119] In this embodiment, by introducing a second semantic understanding model, the user's semantic intent is intelligently matched with available sensing elements and the corresponding hardware devices are dynamically called. This allows for rapid acquisition of real-world scene data, providing effective support for subsequent AI reasoning.

[0120] In some embodiments, the above step S140 "generating prompt information based on the at least one environmental information and the problem information" is as follows Figure 3 As shown, this can be achieved by following the steps below:

[0121] Step S310: fusing the at least one environmental information to obtain multimodal environmental information;

[0122] Multimodal environmental information refers to the comprehensive information representation formed by fusing different types of environmental information (such as images, sounds, depth, heart rate, motion, brightness, temperature, etc.) collected by multiple sensors or devices. This information can more comprehensively reflect the current state of the real scene and provide richer context for subsequent information processing.

[0123] For example, in a classroom scenario, the system can simultaneously obtain image information (student distribution), audio information (speech content), and temperature information (indoor comfort), and integrate them through an algorithm model to form a comprehensive scene representation that includes visual, auditory, and environmental states.

[0124] Step S320: converting the multimodal environment information into text description environment information;

[0125] Textual description of environmental information is the process of converting fused multimodal environmental information into natural language. This process can be accomplished by multimodal AI models, such as using a visual language model (VLM) to convert image information into textual descriptions, or using a speech recognition model to convert audio information into text. The goal is to convert unstructured sensory data into structured, processable textual information, facilitating subsequent semantic understanding and reasoning tasks.

[0126] Step S330: Generate the prompt information based on the text description environment information and the question information.

[0127] Here, prompt information is generated by combining the contextual information described in the text with the user's question, guiding the large model to better understand and answer the question. This step can be performed by the natural language processing model. Based on the current context and user intent, it constructs contextual prompts suitable for the current situation, helping the model generate answers that are more relevant to the actual scenario.

[0128] In the embodiment of the present application, by fusing multiple environmental information and converting them into text descriptions, and then generating prompt information in combination with user questions, the AI ​​system's ability to understand real-world scenarios can be effectively improved, and the AI ​​system's contextual understanding ability and response accuracy can be significantly improved, so that it can not only answer questions, but also make more realistic judgments and feedback based on real-time scenarios.

[0129] In some embodiments, the environmental information includes an image captured by a camera; the above step S140 of "generating prompt information based on the at least one environmental information and the problem information" can be implemented by the following steps:

[0130] Step 141: Identify the image and obtain extended information of the question information;

[0131] Here, the image captured by the camera is a digital expression of the current physical environment, which can be used to identify objects, detect actions, analyze scene status, etc. Image recognition refers to the process of using artificial intelligence algorithms (such as convolutional neural networks) to parse the content in the image, extract key information and convert it into structured data. The extended information is a supplement and expansion of the original question information based on the image recognition results. It can help the system understand the user's demand background more comprehensively and generate more targeted prompt information accordingly. For example, when the user enters "Recommend a song based on my age, current mood, and work environment", the system can capture images through the camera and identify the user's age, mood and work environment information as extended information for recommended songs. It can also generate a suitable song recommendation list based on the user's historical preferences.

[0132] Step 142: Generate the prompt information based on the extended information and the question information.

[0133] During implementation, the expanded information can be combined with the question information to generate prompt information. For example, if the original question is "Recommend suitable songs for me," the user's age, gender, and mood information obtained from the image can be combined to regenerate the prompt information "Recommend suitable songs for a 30-year-old male in a happy state."

[0134] In the embodiment of the present application, by introducing camera-collected images as a source of environmental information and combining image recognition technology to generate extended information of question information, the AI ​​system's ability to understand real-world scenarios can be enhanced, thereby being able to provide more accurate and personalized responses, significantly improving the level of intelligence in human-computer interaction.

[0135] In some embodiments, the above information processing method further includes the following steps:

[0136] Step S180: When it is determined based on the problem information that environmental information will not be collected, the problem information is input into the artificial intelligence model to obtain second generated information, wherein the second generated information is used to respond to the problem information.

[0137] During implementation, when the system receives a question from a user, it first determines whether it needs to collect real-time environmental information related to the question. If analysis reveals that the question can be answered without relying on external environmental data, the system directly inputs the question information into a pre-trained AI model for processing and generates a second generated information as a response to the question.

[0138] In the embodiments of the present application, by determining whether to collect environmental information based on the problem information and directly generating a response result when collection is not required, unnecessary sensor calls and data processing overhead can be reduced, thereby improving the system's response speed and energy efficiency, thereby enhancing the user experience and extending the device's battery life.

[0139] Related AI systems generally lack the ability to perceive real-world scenarios, relying primarily on questions to answer them and failing to dynamically adjust to the user's actual context. For example, in educational settings, teachers might want to quickly count the number of students; in social settings, users might want to identify the makeup shades used by others; and in personalized recommendation scenarios, devices might need to provide appropriate music recommendations based on the user's current state. However, existing systems struggle to determine and respond based on the correlation between these questions and the real-world context.

[0140] To address these issues, this application proposes an information processing method that determines whether the problem information meets the conditions related to the current environment. If so, it uses sensing elements to obtain the corresponding environmental information, fuses the problem information with the environmental information to generate prompt information, and then inputs it into an artificial intelligence model to generate a response result that is closer to the real scene. This method not only improves the environmental perception ability of the AI ​​system, but also enhances its practicality and intelligence in different application scenarios.

[0141] Electronic devices, such as mobile phones and tablets, are equipped with hardware devices that can be used to perceive real-world scenes. For example, cameras can capture images, audio can capture sound, depth sensors can capture three-dimensional (3D) information, and event-based vision sensors (EVS) can capture motion information. In summary, these various sensors can capture information such as ambient brightness, temperature, and heart rate.

[0142] During implementation, various sensors, such as cameras, can be integrated with various AI models to create an AI system capable of perceiving real-time scenes and states. This is equivalent to equipping existing large models on the edge or in the cloud with components (sensors) such as eyes, ears, and skin that can perceive real-time scene conditions.

[0143] Figure 4 A schematic diagram of a processing flow for an AI system to achieve real-time perception is provided in an embodiment of the present application, such as Figure 4 As shown, this can be achieved by following the steps below:

[0144] Step S410: The user inputs a question (text or voice);

[0145] During the implementation process, users can input questions to be answered into the AI ​​system in the form of text or language. For example, the input questions can be "Who is absent from the classroom?", "Play me a song?", "What color is her lipstick?", etc.

[0146] Step S420: AI determines whether real-time scene information is needed;

[0147] During implementation, the AI ​​system can determine that answering an input question requires current environmental information and, in some cases, requires activating a camera or other sensory device based on whether the input question includes words like "my," "now," or "work environment" without contextual information. For example, the word "my" indicates that subsequent action requires consideration of the user's current situation (expression, age, movements, gender, etc.). The word "now" strongly indicates the need for current information.

[0148] In some embodiments, similar terms can be set by the user based on actual needs, such as "current," "real-time," and "on-site." GANs can also be used to generate expanded terms. These terms and their importance levels are then used to enhance training for the semantic understanding model, enabling it to output a judgment on whether the current environment is necessary.

[0149] Here, if the AI ​​determines that real-time scene information is needed, step S430 is executed; if the AI ​​determines that real-time scene information is not needed, step S460 is executed.

[0150] Step S430: AI determines and uses the device to collect real-time information;

[0151] During the implementation process, the AI ​​system can first determine the sensor equipment needed to obtain real-time scene information, and then use the determined sensor equipment to obtain the corresponding information.

[0152] Here, you can use audio acquisition devices to obtain sound information; cameras to collect pictures or videos; TOF sensors to collect depth information; heart rate sensors to collect heartbeat and other information; EVS to collect motion information; and brightness / temperature sensors to collect brightness / temperature.

[0153] In some embodiments, specific words can be used for enhanced training, allowing the model to output a flag array, where each element represents a sensor on or off. For example, if the words "my expression" or "my age" are involved, the flag for the front camera is set to 1; if the words "current environment", "friend across from me", or "current classroom" are involved, the flag for the rear camera is set to 1; if the words "birds chirping" or "speech spoken" are involved, the flag for the sound collector is set to 1.

[0154] Step S440: AI combines real-time information and the original question to generate text from the image;

[0155] During implementation, the AI ​​system can combine real-time information and the original question to generate text from images. For example, the AI ​​system can identify the image, obtain basic information such as the age, gender, and emotion of the person in the image, convert this basic information into text descriptions, and then generate text from the image based on the original question.

[0156] Step S450: regenerate question prompt information;

[0157] During implementation, question prompts can be regenerated based on real-time information and the image-generated text from the original question. For example, if the original question is "Recommend suitable songs for me," the user's age, gender, and mood information obtained from the image can be combined to regenerate the question prompt "Recommend suitable songs for a 30-year-old male in a happy state."

[0158] Step S460: calling the large language model;

[0159] During implementation, the existing large language model may be called to process the prompt information obtained in step S450 or step S420.

[0160] Step S470: Execute a specific application or directly output the answer.

[0161] During implementation, the large language model can perform specific applications based on prompt information or directly output answers.

[0162] In this embodiment, based on the user's question, if the user needs to use the current environment information, the system automatically determines whether to call auxiliary components (sensors) to obtain scene information and how to use them in combination. The final answer is given by combining the original question and the real-time scene information. This can effectively improve the effectiveness and accuracy of the answer.

[0163] Based on the foregoing embodiments, an embodiment of the present application provides an information processing device, which includes the modules included, each module includes sub-modules, each sub-module includes a unit, and can be implemented by a processor in an electronic device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0164] Figure 5 A schematic diagram of the structure of the information processing device provided in the embodiment of the present application is shown in FIG. Figure 5 As shown, the apparatus 500 includes:

[0165] A first acquisition module 510 is used to acquire question information;

[0166] A determination module 520 is configured to determine whether the problem information satisfies a first condition, wherein the first condition indicates that the problem information is associated with current environment information;

[0167] A second acquisition module 530 is configured to acquire at least one piece of environmental information associated with the problem information, wherein the at least one piece of environmental information is obtained based on environmental information of the electronic device being sensed by a sensing element;

[0168] A first generating module 540 is configured to generate prompt information based on the at least one piece of environmental information and the problem information;

[0169] The first information processing module 550 is used to input the prompt information into an artificial intelligence model to obtain first generated information, wherein the first generated information is used to respond to the question information.

[0170] In some embodiments, the determination module 520 includes a parsing submodule and a first determination submodule, wherein the parsing submodule is used to parse the question information to obtain at least one target information in the question information that matches the preset matching information; the first determination submodule is used to determine whether the first condition is satisfied based on at least one of the target information; the preset matching information includes at least one of the following: object information, orientation information, time information and location information.

[0171] In some embodiments, the second acquisition module 530 includes a second determination submodule and a sensing submodule, wherein the second determination submodule is used to determine the sensing element that matches the target information; and the sensing submodule is used to sense the environmental information of the electronic device based on the sensing element, wherein, when the electronic device is arranged with multiple sensing elements of the same type in different directions, it includes determining the sensing element in the direction that matches the target information.

[0172] In some embodiments, the first determination submodule includes a first determination unit and a second determination unit, wherein the first determination unit is used to determine the context information corresponding to each target information; the second determination unit is used to determine that when the context information is less than a preset amount of information, the target information corresponding to the context information meets the first condition.

[0173] In some embodiments, the device further includes a third acquisition module and a second generation module, wherein the third acquisition module is used to acquire initial matching information; and the second generation module is used to generate the preset matching information using the initial matching information through a generative adversarial network.

[0174] In some embodiments, the device also includes a first training module or a second training module, wherein the first training module is used to train the semantic understanding model using the preset matching information to obtain a first semantic understanding model; the second training module is used to train the semantic understanding model using the preset matching information and the weight parameters of each preset matching information to obtain the first semantic understanding model; correspondingly, the first determination submodule is also used to input at least one of the target information into the first semantic understanding model to determine whether the first condition is met.

[0175] In some embodiments, the device also includes a third training module, which is used to train the first semantic understanding model using the preset matching information and the sensing element information associated with each preset matching information to obtain a second semantic understanding model; wherein, the second semantic understanding model is used to determine the sensing element information for collecting the environmental information; correspondingly, the second acquisition module 530 includes a third determination submodule and an acquisition submodule, wherein the third determination submodule is used to input at least one of the target information into the second semantic understanding model to determine at least one target sensing element to be started; the acquisition submodule is used to start at least one of the target sensing elements to obtain at least one environmental information associated with the problem information.

[0176] In some embodiments, the first generation module 540 includes a fusion submodule, a conversion submodule and a first generation submodule, wherein the fusion submodule is used to fuse the at least one environmental information to obtain multimodal environmental information; the conversion submodule is used to convert the multimodal environmental information into text description environmental information; and the first generation submodule is used to generate the prompt information based on the text description environmental information and the question information.

[0177] In some embodiments, the environmental information includes an image captured by a camera; the first generation module 540 includes an identification submodule and a second generation submodule, wherein the identification submodule is used to identify the image and obtain extended information of the problem information; the second generation submodule is used to generate the prompt information based on the extended information and the problem information.

[0178] In some embodiments, the device also includes a second information processing module, which is used to input the problem information into the artificial intelligence model to obtain second generated information based on the problem information to determine not to collect environmental information, wherein the second generated information is used to respond to the problem information.

[0179] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.

[0180] It should be noted that, in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable an electronic device (which can be a mobile phone, tablet computer, laptop computer, desktop computer, etc.) to execute all or part of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0181] Correspondingly, an embodiment of the present application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the information processing method provided in the above embodiment are implemented.

[0182] Correspondingly, an embodiment of the present application provides an electronic device, Figure 6 A hardware entity diagram of an electronic device provided in an embodiment of the present application, such as Figure 6 As shown, the hardware entity of the device 600 includes: a memory 601 and a processor 602, the memory 601 stores a computer program that can be run on the processor 602, and the processor 602 implements the steps of the information processing method provided in the above embodiment when executing the program.

[0183] The memory 601 is configured to store instructions and applications executable by the processor 602, and can also cache data to be processed or processed by the processor 602 and various modules in the electronic device 600 (for example, image data, audio data, voice communication data and video communication data), which can be implemented through flash memory (FLASH) or random access memory (RAM).

[0184] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0185] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0186] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0187] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0188] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0189] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0190] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0191] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words be embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium and includes several instructions for enabling an electronic device (which can be a mobile phone, tablet computer, laptop computer, desktop computer, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0192] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0193] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0194] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0195] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. An information processing method, applied to an electronic device, comprising: Get problem information; Determining that the problem information satisfies a first condition, wherein the first condition indicates that the problem information is associated with current environment information; Acquire at least one type of environmental information associated with the problem information, where the at least one type of environmental information is obtained based on environmental information of the electronic device being sensed by a sensing element; generating prompt information based on the at least one environmental information and the problem information; The prompt information is input into an artificial intelligence model to obtain first generated information, wherein the first generated information is used to respond to the question information.

2. The method according to claim 1, wherein determining that the problem information satisfies the first condition comprises: Parsing the question information to obtain at least one target information in the question information that matches preset matching information; Determining that the first condition is satisfied based on at least one of the target information; The preset matching information includes at least one of the following: object information, orientation information, time information and location information.

3. The method according to claim 2, wherein obtaining at least one piece of environmental information associated with the problem information comprises: determining a sensing element that matches the target information; The environmental information of the electronic device is sensed based on the sensing element, wherein when the electronic device is arranged with a plurality of sensing elements of the same type in different directions, the sensing element in the direction matching the target information is determined.

4. The method according to claim 2, wherein determining that the first condition is satisfied based on at least one of the target information comprises: Determining context information corresponding to each of the target information; When it is determined that the context information is less than a preset information amount, it is determined that the target information corresponding to the context information meets the first condition.

5. The method of claim 2, further comprising: Get initial matching information; The preset matching information is generated by using the initial matching information through a generative adversarial network.

6. The method of claim 2, further comprising: Using the preset matching information to train a semantic understanding model to obtain a first semantic understanding model; or, Training a semantic understanding model using the preset matching information and a weight parameter of each of the preset matching information to obtain the first semantic understanding model; Wherein, the first semantic understanding model is used to determine whether the question information is associated with the current environment information; Correspondingly, determining that the first condition is satisfied based on at least one piece of target information includes: Input at least one of the target information into the first semantic understanding model to determine whether the first condition is met.

7. The method of claim 6, further comprising: Training the first semantic understanding model using the preset matching information and the sensing element information associated with each of the preset matching information to obtain a second semantic understanding model; Wherein, the second semantic understanding model is used to determine the sensing element information for collecting the environmental information; Correspondingly, the acquiring of at least one piece of environmental information associated with the problem information includes: Inputting at least one target information into the second semantic understanding model to determine at least one target sensing element to be activated; At least one of the target sensing elements is activated to acquire at least one type of environmental information associated with the problem information.

8. The method according to claim 1, wherein generating prompt information based on the at least one environmental information and the problem information comprises: fusing the at least one environmental information to obtain multimodal environmental information; Converting the multimodal environmental information into textual description environmental information; The prompt information is generated based on the text description environment information and the question information.

9. The method according to claim 1, wherein the environmental information comprises an image captured by a camera; The generating prompt information based on the at least one environmental information and the problem information includes: Recognizing the image and obtaining extended information of the question information; The prompt information is generated based on the extended information and the question information.

10. The method according to any one of claims 1 to 9, further comprising: When it is determined based on the problem information that environmental information will not be collected, the problem information is input into the artificial intelligence model to obtain second generated information, wherein the second generated information is used to respond to the problem information.

Citation Information

Cited By

  • Intelligent control system and method of song ordering machine

    CN120932617A