Action prediction method and device based on artificial intelligence, computer equipment and medium

By extracting key information based on visual devices, constructing a graph structure, and combining graph neural networks, large language models, and reinforcement learning policy networks for action prediction, the system solves the problems of insufficient accuracy and interpretability in traditional robot decision-making systems, achieving more efficient decision support and user trust.

CN120806115APending Publication Date: 2025-10-17PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510681960.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional robot action decision-making systems rely on simple rule matching or local environment perception, resulting in low accuracy and rationality of action decisions, a lack of deep understanding of complex scenarios and explainability of the decision-making process, affecting task execution efficiency and user trust.

Method used

The system extracts key information based on vision devices, constructs a graph structure, analyzes it using graph neural networks, combines a large language model and a reinforcement learning policy network to predict actions, and generates target explanation information, which is then output to the robot controller and the user.

Benefits of technology

It improves the accuracy and rationality of robot action decisions, makes the decision-making process explainable, and enhances user trust and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806115A_ABST
    Figure CN120806115A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to an action prediction method and device based on artificial intelligence, computer equipment and a storage medium, and the method comprises the steps: extracting key information in a target scene where a robot is located based on visual equipment; constructing a graph structure based on the key information; analyzing the graph structure based on a graph neural network to obtain an analysis result; performing action prediction on the analysis result based on a large language model to obtain a plurality of candidate actions; performing income evaluation on all the candidate actions based on a reinforcement learning strategy network so as to determine a target action from all the candidate actions; generating target interpretation information of the target action based on the demand information of the target user; and outputting the target action to a controller of the robot, and outputting the target action and the target interpretation information to the target user. In addition, the target interpretation information may be stored in the blockchain. The method can be applied to robot action prediction scenes in the financial field and the medical field, and the accuracy and rationality of action decision making are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and can be applied to the fields of financial technology and digital medical treatment, and particularly relates to an action prediction method and device based on artificial intelligence, a computer device, and a storage medium. BACKGROUND

[0002] In a traditional robot action decision system, the decision process mainly relies on simple rule matching or local environment perception, resulting in insufficient understanding of complex scene structures. Specifically, the traditional method usually makes decisions based only on the isolated attributes (such as distance, size) or local spatial relationships of obstacles, lacking depth analysis of the environment. This extensive decision-making approach often leads to a significant reduction in the accuracy and rationality of the robot's action decisions. For example, in a warehouse logistics scenario, the traditional method may only plan a path according to the physical location of the obstacles, without considering key factors such as shelf stability, passage priority, or cargo fragility. If the robot needs to detour around fragile goods in a narrow passage, the traditional method may choose a high-risk path due to the lack of global semantic understanding, resulting in damage to the goods or failure of the task. This decision bias not only affects the efficiency of task execution, but also may cause safety hazards, reducing user trust in the robot system.

[0003] In addition, the traditional robot system lacks support for the explainability of the decision-making process, and users cannot obtain the basis for action selection or the evaluation logic of candidate solutions. For example, in the field of automated trading robots in finance, the traditional method may execute trading operations based on preset quantitative indicators (such as price volatility), but cannot explain to the user why the current trading strategy is chosen instead of other alternatives. If the market environment changes suddenly, resulting in trading losses, the user lacks decision transparency and cannot evaluate the reliability of the system, and may even terminate cooperation. This "black box" decision-making mode seriously restricts the popularization and application of robot systems in key fields. Therefore, it is urgent to provide a robot action decision-making method that integrates deep understanding of scene structure and explainability of the decision-making process, to improve decision quality, enhance user trust, and promote the large-scale application of robot technology in complex scenarios. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide an action prediction method, device, computer device, and storage medium based on artificial intelligence, to solve the technical problem of low accuracy and rationality of action decisions in existing action decision-making methods that rely on simple rule matching or local environment perception.

[0005] In a first aspect, an action prediction method based on artificial intelligence is provided, comprising:

[0006] extracting key information in a target scene where the robot is located based on a preset visual device;

[0007] construct a corresponding graph structure based on the key information;

[0008] analyze the graph structure based on a preset graph neural network to obtain a corresponding analysis result;

[0009] perform action prediction on the analysis result based on a preset large language model to obtain a plurality of candidate actions;

[0010] perform benefit evaluation processing on all the candidate actions based on a preset reinforcement learning strategy network to determine a target action from all the candidate actions;

[0011] obtain demand information of a target user, and generate target explanation information corresponding to the target action based on the demand information;

[0012] output the target action to a controller corresponding to the robot, and output the target action and the target explanation information to the target user.

[0013] In a second aspect, an action prediction device based on artificial intelligence is provided, comprising:

[0014] an extraction module configured to extract key information in a target scene where a robot is located based on a preset visual device;

[0015] a construction module configured to construct a corresponding graph structure based on the key information;

[0016] an analysis module configured to analyze the graph structure based on a preset graph neural network to obtain a corresponding analysis result;

[0017] a prediction module configured to perform action prediction on the analysis result based on a preset large language model to obtain a plurality of candidate actions;

[0018] an evaluation module configured to perform benefit evaluation processing on all the candidate actions based on a preset reinforcement learning strategy network to determine a target action from all the candidate actions;

[0019] a generation module configured to obtain demand information of a target user, and generate target explanation information corresponding to the target action based on the demand information;

[0020] an output module configured to output the target action to a controller corresponding to the robot, and output the target action and the target explanation information to the target user.

[0021] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor implementing the steps of the above-mentioned action prediction method based on artificial intelligence when executing the computer program.

[0022] In a fourth aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, the computer program implementing the steps of the above-mentioned action prediction method based on artificial intelligence when executed by a processor.

[0023] In the above-mentioned action prediction method based on artificial intelligence, device, computer device, and storage medium, the key information in the target scene where the robot is located is first extracted based on a preset visual device, and a corresponding graph structure is constructed based on the key information. Then, the graph structure is analyzed based on a preset graph neural network to obtain a corresponding analysis result. Then, the analysis result is predicted for action based on a preset large language model to obtain a plurality of candidate actions. Subsequently, all the candidate actions are processed for benefit evaluation based on a preset reinforcement learning strategy network to determine a target action from all the candidate actions. Further, the demand information of a target user is obtained, and target explanation information corresponding to the target action is generated based on the demand information. Finally, the target action is output to a controller corresponding to the robot, and the target action and the target explanation information are output to the target user. In this application, the key information in the target scene where the robot is located is extracted based on the use of the visual device, and the graph structure is constructed based on the key information. Then, the graph structure is analyzed based on the use of the graph neural network to obtain the analysis result. Then, the analysis result is predicted for action based on the use of the large language model to obtain the plurality of candidate actions. Subsequently, all the candidate actions are processed for benefit evaluation based on the use of the reinforcement learning strategy network to determine the target action from all the candidate actions. Further, the target explanation information corresponding to the target action is generated based on the demand information of the target user. Finally, the target action is output to the controller corresponding to the robot, and the target action and the target explanation information are output to the target user. In this application, the behavior decision-making process of the robot is performed by combining the use of the visual device, the graph neural network, the large language model, and the reinforcement learning strategy network, which can comprehensively understand the scene structure, thereby providing more abundant information support for action decision-making, effectively improving the accuracy and rationality of action decision-making. Moreover, the target explanation information of the target action can be output according to the demand information of the target user, thereby automatically and intelligently realizing the explainability support for the action decision-making process, thereby improving the user experience and enhancing the user trust. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the solutions in the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0025] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0026] Figure 2 is a flowchart of one embodiment of the action prediction method based on artificial intelligence according to the present application;

[0027] Figure 3 is a structural schematic diagram of one embodiment of the action prediction device based on artificial intelligence according to the present application;

[0028] Figure 4 is a structural schematic diagram of one embodiment of the computer device according to the present application. DETAILED DESCRIPTION

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs; the terminology used in the specification of the present application is only for the purpose of describing specific embodiments and is not intended to limit the present application; the specification of the present application, the claims and the above description of drawings, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. The specification of the present application and the claims or the above description of drawings, the terms "first", "second" and the like are used to distinguish different objects, not to describe a specific order.

[0030] In this paper, the phrase "embodiment" means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present application. The phrase appears in the specification at various places does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment to other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0031] In order to make those skilled in the art better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings.

[0032] As Figure 1As shown, the system architecture 100 can include a terminal device 101, a network 102 and a server 103. The terminal device 101 can be a notebook computer 1011, a tablet computer 1012 or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0033] A user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0034] The terminal device 101 can be various electronic devices with display screens and supporting web browsing, in addition to the notebook computer 1011, the tablet computer 1012 or the mobile phone 1013, the terminal device 101 can also be an electronic book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer and a desktop computer, etc.

[0035] The server 103 can be a server providing various services, such as a background server supporting a page displayed on the terminal device 101.

[0036] It should be noted that the action prediction method based on artificial intelligence provided by the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the action prediction apparatus based on artificial intelligence is generally arranged in a server / terminal device.

[0037] It should be understood that Figure 1 The number of terminal devices, networks and servers in the system architecture 100 is only illustrative. According to the implementation needs, there can be any number of terminal devices, networks and servers.

[0038] Continuing to refer to the system architecture 100 Figure 2, shows a flowchart of an embodiment of the action prediction method based on artificial intelligence according to the present application. According to different needs, the order of the steps in the flowchart can be changed, and some steps can be omitted. The action prediction method based on artificial intelligence provided by the embodiment of the present application can be applied to any scenario that requires robot action prediction, and the action prediction method based on artificial intelligence can be applied to products in these scenarios, for example, robot action prediction scenarios in the financial field and the medical field. The action prediction method based on artificial intelligence includes the following steps:

[0039] Step S201: extract key information of the target scene where the robot is located based on a preset visual device.

[0040] In this embodiment, the action prediction method based on artificial intelligence is executed on the electronic device (eg Figure 1 The server / terminal device shown in the figure) can obtain key information of the target scene in which the robot is located through a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection method may include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (Ultra Wide Band) connection, and other wireless connection methods currently known or to be developed in the future. The execution subject of this application may specifically be a robot action decision system, which can be referred to as a system for short.

[0041] This application can be applied to robot action prediction scenarios in the financial and medical fields. For example, in the financial field, these robots may include at least: intelligent investment advisory dialogue robots, anti-fraud inspection robots, claims settlement service robots, intelligent securities trading robots, etc. In the medical field, these robots may include at least: surgical assistance robots, intelligent nursing robots, drug delivery robots, rehabilitation training robots, etc.

[0042] Among them, the specific implementation process of extracting key information from the target scene where the robot is located based on the preset visual equipment will be further described in detail in the subsequent specific embodiments of this application and will not be elaborated on here.

[0043] Step S202: construct a corresponding graph structure based on the key information.

[0044] In this embodiment, the specific implementation process of constructing the corresponding graph structure based on the key information will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.

[0045] Step S203: Analyze the graph structure based on a preset graph neural network to obtain corresponding analysis results.

[0046] In the embodiment, the specific implementation process of analyzing the graph structure based on the preset graph neural network to obtain the corresponding analysis result will be further described in detail in subsequent specific embodiments, and will not be described in detail here.

[0047] In step S204, the action prediction is performed on the analysis result based on the preset large language model to obtain a plurality of candidate actions.

[0048] In the embodiment, the selection of the large language model is not specifically limited, and can be selected according to actual business requirements. The large language model (LLM) can be selected. According to the analysis result of the graph neural network, possible action options (such as detour, crossing, moving obstacles) are generated, and are used as the corresponding plurality of candidate actions.

[0049] In step S205, the reward evaluation processing is performed on all the candidate actions based on the preset reinforcement learning strategy network to determine the target action from all the candidate actions.

[0050] In the embodiment, the specific implementation process of performing reward evaluation processing on all the candidate actions based on the preset reinforcement learning strategy network to determine the target action from all the candidate actions will be further described in detail in subsequent specific embodiments, and will not be described in detail here.

[0051] In step S206, the demand information of the target user is obtained, and the target explanation information corresponding to the target action is generated based on the demand information.

[0052] In the embodiment, the specific implementation process of obtaining the demand information of the target user and generating the target explanation information corresponding to the target action based on the demand information will be further described in detail in subsequent specific embodiments, and will not be described in detail here.

[0053] In step S207, the target action is output to the controller corresponding to the robot, and the target action and the target explanation information are output to the target user.

[0054] In the embodiment, the target action can be output to the controller corresponding to the robot to control the robot to complete the matching actual operation according to the target action. In addition, the target action and the target explanation information are output to the target user to output the target explanation information corresponding to the target action and meeting the demand information of the target user, effectively ensuring the pertinence and effectiveness of the target explanation information transmission, and being beneficial to improving the user experience and practicality.

[0055] The application first extracts key information in a target scene where a robot is located based on a preset visual device; and constructs a corresponding graph structure based on the key information; then analyzes the graph structure based on a preset graph neural network to obtain a corresponding analysis result; thereafter, predicts actions based on a preset large language model to obtain a plurality of candidate actions corresponding to the analysis result; subsequently, evaluates the benefits of all the candidate actions based on a preset reinforcement learning strategy network to determine a target action from all the candidate actions; further, obtains demand information of a target user, and generates target explanation information corresponding to the target action based on the demand information; finally, outputs the target action to a controller corresponding to the robot, and outputs the target action and the target explanation information to the target user. The application extracts key information in a target scene where a robot is located based on the use of a visual device, and constructs a graph structure based on the key information, then analyzes the graph structure based on the use of a graph neural network to obtain an analysis result, thereafter, predicts actions based on the use of a large language model to obtain a plurality of candidate actions, subsequently, evaluates the benefits of all the candidate actions based on the use of a reinforcement learning strategy network to determine a target action from all the candidate actions, and further generates target explanation information corresponding to the target action based on demand information of a target user, finally, outputs the target action to a controller corresponding to the robot, and outputs the target action and the target explanation information to the target user. The application combines the use of a visual device, a graph neural network, a large language model and a reinforcement learning strategy network to make behavior decision of a robot, which can comprehensively understand the scene structure, thereby providing more abundant information support for action decision, effectively improving the accuracy and rationality of action decision. Moreover, the target explanation information of the target action can be output according to the demand information of the target user, thereby automatically and intelligently realizing the explainability support of the action decision process, which can improve the user experience and enhance the user trust.

[0056] In some optional implementations, step S205 includes the following steps:

[0057] The pre-constructed reinforcement learning strategy network is called.

[0058] In this embodiment, the reinforcement learning policy network described above is a parameterized function (such as a deep neural network) that takes the current state (i.e., the output of the graph neural network) and a candidate action as input, and outputs the Q-value (expected return) of that action. The Q-value represents the long-term cumulative reward expectation of taking a certain action in a given state. The following factors are considered in the calculation: safety: such as whether to break a fragile vase (the Q-value of a high-risk action is lower). feasibility: such as whether the robot's balance ability supports crossing the box (the Q-value of a high-difficulty action is lower). efficiency: such as path length or operation time (the Q-value of an inefficient action is lower). In addition, the Q-value update process includes: experience replay: by storing historical state-action-reward-next state (SARSD) data, randomly sample to train the policy network. loss function: use mean square error loss function to minimize the difference between the predicted Q-value and the target Q-value (calculated by Bellman equation).

[0059] Based on the reinforcement learning policy network, the expected return of each candidate action is evaluated to obtain a plurality of corresponding returns.

[0060] In this embodiment, for each candidate action (such as detour, cross vase, cross box, move box), its Q-value is calculated by the reinforcement learning policy network described above to obtain the return of each candidate action. For example, the Q-value of detour = safety (high) + feasibility (medium) + efficiency (low) → Q = 0.8. The Q-value of crossing the vase = safety (low) + feasibility (high) + efficiency (high) → Q = 0.5. The Q-value of crossing the box = safety (medium) + feasibility (low) + efficiency (medium) → Q = 0.3. The Q-value of moving the box = safety (high) + feasibility (medium) + efficiency (low) → Q = 0.6.

[0061] From all the returns, the specified return with the highest value is selected, and the specified action corresponding to the specified return is obtained.

[0062] In this embodiment, all returns can be numerically compared to select the specified return with the highest value. The specified return with the highest value is obtained as the optimal solution, that is, the action with the highest Q-value is selected as the optimal solution (target action). For example, if the Q-value of detour is the highest (0.8), detour is selected.

[0063] The specified action is taken as the target action.

[0064] In this embodiment, the optimal action (target action) can be sent to the robot controller to complete the actual operation, and the corresponding reward feedback process is performed, including: immediate reward: calculate the immediate reward according to the action execution result (such as whether the vase is broken or whether the detour is successful). policy update: update the parameters of the policy network using the immediate reward to optimize future decisions.

[0065] The action decision success rate and the explanation quality can be optimized by setting a reward function to improve the performance of the system. Specifically, the reward function is set to optimize the action decision success rate and the explanation quality from two aspects: a basic task (success of task execution is rewarded, encountering failure or obstacles in the middle is punished, and making invalid actions is punished) and an explanation quality (six key aspects of the explanation are rewarded, the explanation and reasoning process is consistent, and human preferences are referenced for evaluation). In addition, the action prediction and explanation generation strategy can be updated by a policy gradient or Q-learning algorithm to maximize the cumulative reward.

[0066] The application calls a pre-constructed reinforcement learning strategy network, then evaluates the expected returns of each candidate action based on the reinforcement learning strategy network to obtain a plurality of corresponding returns, then selects a specified return with the highest value from all the returns and obtains a specified action corresponding to the specified return, and subsequently takes the specified action as the target action. The application evaluates the expected returns of each candidate action based on the reinforcement learning strategy network to obtain a plurality of corresponding returns, then selects a specified return with the highest value from all the returns, and further obtains a specified action corresponding to the specified return and takes the specified action as the final target action, so that the reinforcement learning strategy network can be used to efficiently and accurately complete the return evaluation processing of the candidate actions in combination with the analysis of various factors, and the accuracy of the target action obtained can be effectively guaranteed, thereby improving the accuracy and robustness of the action decision.

[0067] In some optional implementations of the embodiment, step S206 includes the following steps:

[0068] The demand information of the target user is obtained, wherein the demand information includes ordinary demand information, development demand information, and research demand information.

[0069] In the embodiment, the demand information of the target user can be determined by recognizing the user identity. Specifically, the user identity is determined by user login information, interaction history, or explicit input (such as a role selected by the user), for example, an ordinary user, a developer, or a researcher. If the user is an ordinary user, the corresponding demand information is ordinary demand information, because the ordinary user usually needs simple and clear explanations to avoid technical details. If the user is a developer, the corresponding demand information is development demand information, because the developer usually needs detailed technical analysis to debug or optimize the system. If the user is a researcher, the corresponding demand information is research demand information, because the researcher usually needs structured data to support further analysis or paper writing.

[0070] Alternatively, the demand information of the target user is determined according to the current scene (such as daily interaction, debugging, research). Among them, if the current scene is a daily interaction scene, the demand information of the corresponding target user is ordinary demand information; if the current scene is a debugging scene, the demand information of the corresponding target user is development demand information; and if the current scene is a research scene, the demand information of the corresponding target user is research demand information.

[0071] If the demand information is the ordinary demand information, the first preset explanation generation strategy is used to generate the target explanation information corresponding to the target action based on the ordinary demand information.

[0072] In the embodiment, the first explanation generation strategy is a processing strategy corresponding to a brief mode, and specifically includes: a brief mode: giving the reason for the current action, the target of the brief mode: providing a quick and intuitive explanation, suitable for ordinary users. Implementation process: direct cause extraction: extracting the direct cause of the optimal action from the explanation text, for example, "detour because the vase is fragile". Concise output: output in the form of a short sentence or phrase, avoiding redundant information. The explanation text refers to the explanation information generated based on the data mode, and the specific generation process of the explanation text can be referred to the process of the third explanation generation strategy described above.

[0073] If the demand information is the development demand information, the second preset explanation generation strategy is used to generate the target explanation information corresponding to the target action based on the development demand information.

[0074] In the embodiment, the second explanation generation strategy is a processing strategy corresponding to a detailed mode, and specifically includes: a detailed mode: explaining why this action is chosen and why other paths are not chosen. And give a brief analysis process. The goal of the detailed mode: provide a complete decision-making process and analysis, suitable for developers to debug or optimize the system. Implementation process: decision-making process unfolding: detailed description of scene analysis, candidate action evaluation and the reason for the final selection, for example: "the vase is fragile, crossing it will damage it, which does not meet the safety operation criteria; the box is stable, but crossing it is difficult and may affect the balance of the robot. Therefore, detour is chosen." Technical details are retained: key results of graph neural network analysis (such as "the narrow passage between the vase and the box is passable") and the logic of action selection.

[0075] If the demand information is the research demand information, the third preset explanation generation strategy is used to generate the target explanation information corresponding to the target action based on the research demand information.

[0076] In the embodiment, if the demand information is the research demand information, a third preset explanation generation strategy is used to generate the target explanation information corresponding to the target action by using the research demand information. Details will be described in subsequent embodiments.

[0077] In the embodiment, if the demand information is the research demand information, a third preset explanation generation strategy is used to generate the target explanation information corresponding to the target action by using the research demand information. Details will be described in subsequent embodiments. In the embodiment, if the demand information is the research demand information, a third preset explanation generation strategy is used to generate the target explanation information corresponding to the target action by using the research demand information. Details will be described in subsequent embodiments.

[0078] In some optional implementations, if the demand information is the research demand information, a third preset explanation generation strategy is used to generate the target explanation information corresponding to the target action by using the research demand information, including the following steps:

[0079] If the demand information is the research demand information, the analysis result is converted into a corresponding natural language description.

[0080] In this embodiment, the third explanation generation strategy described above is a processing strategy corresponding to the data mode, which specifically includes: data mode: output the current observed objects, analyze their states, list the feasible action options and their advantages and disadvantages, and finally explain why a certain action is chosen as the optimal solution and why other options are not chosen. The goal of the data mode: provide structured data suitable for researchers to analyze or further process. Implementation process: structured data output: output the following information in the form of a list or table: object attributes: such as "vase: fragile, left side, moderate size; box: stable, right side, large volume". Candidate action evaluation: such as "detour: high safety, medium feasibility, low efficiency; cross the vase: low safety, high feasibility, high efficiency". Optimal action selection: such as "choose detour because it has the highest overall safety". Original data preservation: provide the output of the graph neural network (such as the relationship representation of nodes and edges) to support in-depth analysis.

[0081] Among them, the obstacle attributes and environmental information in the analysis results generated by the graph neural network can be converted into natural language. For example: "There is a fragile vase in front, located on the left side, with moderate size; there is a stable box on the right side, with large volume. There is a narrow passage between the vase and the box."

[0082] Action analysis is performed on the candidate actions to obtain corresponding action analysis results.

[0083] In this embodiment, all candidate actions can be listed, and advantage analysis and disadvantage analysis can be performed on all candidate actions. The advantage analysis results and disadvantage analysis results of each candidate action are integrated to obtain the action analysis results described above.

[0084] The reasons for choosing the target action and the reasons for not choosing other actions are generated.

[0085] In this embodiment, the reasons for choosing the optimal action (such as "detour ensures safety") and the reasons for not choosing other actions (such as "crossing the box is difficult to operate") can be explicitly stated.

[0086] The natural language description, the action analysis results, the selection reasons, and the reasons are integrated to obtain target explanation information corresponding to the target action.

[0087] In the embodiment, the corresponding integrated content can be obtained by integrating the natural language description, the action analysis result, the selection reason and the reason, and the integrated content is taken as the target explanation information corresponding to the target action. The generated target explanation information is directly based on the obstacle attribute and the environmental relationship output by the graph neural network, ensuring the accuracy of the description and the analysis. Moreover, the generation of the candidate action and the selection of the optimal action provide the decision basis for the explanation text, so that the explanation content is completely consistent with the decision logic. In addition, through the clear scene description, the option analysis and the decision explanation contained in the target explanation information, the user can intuitively understand how the robot selects the optimal action from multiple options, that is, because the explanation content is closely combined with the decision logic, the user can clearly understand why a certain action is selected, rather than just knowing the target action itself.

[0088] If the demand information is detected as the research demand information, the analysis result is converted into corresponding natural language description; then the candidate action is subjected to action analysis to obtain corresponding action analysis result; then the selection reason corresponding to the target action is generated, and the reason for not selecting other actions is generated; and subsequently, the natural language description, the action analysis result, the selection reason and the reason are subjected to content integration to obtain the target explanation information corresponding to the target action. When the demand information is detected as the research demand information, the analysis result is converted into corresponding natural language description, the candidate action is subjected to action analysis to obtain the action analysis result, the selection reason corresponding to the target action is generated, and the reason for not selecting other actions is generated, and then the obtained natural language description, action analysis result, selection reason and reason are subjected to content integration, so that the target explanation information of the target action adapted to the research demand information can be efficiently and accurately constructed, and the accuracy and logical consistency of the generated target explanation information are effectively ensured. Moreover, because the target explanation information contains clear scene description, option analysis and decision explanation, the persuasiveness of the target explanation information is enhanced, so that the user can intuitively understand how the robot selects the optimal action from multiple options, and the user experience is improved.

[0089] In some optional implementations, step S201 includes the following steps:

[0090] Capturing a scene image of a target scene where the robot is located based on the visual device.

[0091] In the embodiment, the selection of the visual device is not specifically limited, for example, a camera or a depth sensor can be used. According to the selected visual device, a two-dimensional image or a three-dimensional image of the scene in the target scene where the robot is currently located can be captured to obtain the scene image.

[0092] The scene image is preprocessed to obtain a corresponding target scene image.

[0093] In this embodiment, the scene image can be preprocessed by denoising, enhancing contrast, and the like, to improve the accuracy of subsequent analysis, and the preprocessed image is taken as the corresponding target scene image.

[0094] An obstacle detection result is obtained by performing obstacle detection on the target scene image based on a preset target detection algorithm.

[0095] In this embodiment, the target detection algorithm can specifically be YOLO, Faster R-CNN, or the like. The obstacle in the target scene image can be identified according to the selected target detection algorithm, and the position and bounding box thereof are labeled, so as to obtain the corresponding obstacle detection result.

[0096] An attribute analysis result is obtained by performing attribute analysis on the detection result.

[0097] In this embodiment, the attribute analysis includes judging the attribute of the obstacle, such as whether it is fragile or unstable, by a classification network (such as ResNet), and outputting the corresponding attribute analysis result.

[0098] An environment analysis result is obtained by performing environment analysis on the target scene image.

[0099] In this embodiment, the environment analysis includes identifying environmental elements such as accessible paths, walls, furniture, and the like, by using a semantic segmentation technology (such as U-Net), and extracting the spatial position and relationship thereof, so as to obtain the corresponding environment analysis result.

[0100] Key information corresponding to the target scene is obtained by integrating the obstacle detection result, the attribute analysis result, and the environment analysis result.

[0101] In this embodiment, the integration information is obtained by integrating the obtained obstacle detection result, attribute analysis result, and environment analysis result, and the integration information is taken as the key information corresponding to the target scene.

[0102] The application captures a scene image in a target scene where the robot is located based on the visual device; then pre-processes the scene image to obtain a corresponding target scene image; then detects obstacles in the target scene image based on a preset target detection algorithm to obtain a corresponding obstacle detection result; analyzes the detection result to obtain a corresponding attribute analysis result; and analyzes the target scene image to obtain a corresponding environment analysis result; and subsequently integrates information of the obstacle detection result, the attribute analysis result and the environment analysis result to obtain key information corresponding to the target scene. The application captures a scene image in a target scene where the robot is located based on the use of a visual device, and pre-processes the scene image to obtain a target scene image, then detects obstacles in the target scene image based on the use of a target detection algorithm to obtain an obstacle detection result, and analyzes the detection result to obtain an attribute analysis result, and analyzes the target scene image to obtain an environment analysis result, and then integrates the obtained obstacle detection result, attribute analysis result and environment analysis result to obtain key information in the target scene where the robot is located, thereby ensuring the richness, standardization and accuracy of the obtained key information. Moreover, the extracted key information is used to provide basic data for constructing a graph structure, so as to ensure that the nodes and edges in the graph structure can accurately reflect the physical and semantic information of the scene.

[0103] In some optional implementations of the embodiment, step S202 includes the following steps:

[0104] Based on the obstacle detection result and the environment analysis result in the key information, node construction processing is performed to obtain a corresponding target node.

[0105] In the embodiment, the above-mentioned node construction processing is also referred to as node definition, which includes taking the detected obstacles and environmental elements (such as walls and furniture) as nodes in the graph, and each node is attached with its attributes (such as position, size and fragility).

[0106] Based on the attribute analysis result in the key information, edge construction processing is performed on the target node to obtain a corresponding target edge.

[0107] In the embodiment, the above-mentioned edge construction processing is also referred to as edge definition, which includes generating spatial relationship and generating similarity relationship. Specifically, generating spatial relationship includes calculating the relative position (such as distance and direction) and spatial occupation relationship (such as overlap and adjacency) between nodes to form spatial edges. Generating similarity relationship includes establishing similarity edges according to attribute similarity (such as material and function) or semantic similarity (such as both being furniture).

[0108] A preset graph database is called.

[0109] In the embodiment, the selection of the graph database is not specifically limited, and a general graph database can be selected according to actual business requirements.

[0110] The target node and the target edge are organized and processed based on the graph database to obtain a corresponding graph structure.

[0111] In the embodiment, the target node and the target edge obtained by the organization and construction can form a complete graph structure.

[0112] The application obtains the target node by performing node construction processing based on the obstacle detection result and the environment analysis result in the key information, and then obtains the target edge by performing edge construction processing on the target node based on the attribute analysis result in the key information. Then, a preset graph database is called. Subsequently, the target node and the target edge are organized and processed based on the graph database to obtain a corresponding graph structure. The application obtains the target node by performing node construction processing based on the obstacle detection result and the environment analysis result in the key information, and obtains the target edge by performing edge construction processing on the target node based on the attribute analysis result in the key information. Then, the target node and the target edge are organized and processed based on the use of the graph database, so that a graph structure matching the key information can be efficiently and accurately constructed, the quality of the generated graph structure is ensured, and the rationality of subsequent action decisions can be ensured based on the use of the graph structure.

[0113] In some optional implementation manners of the embodiment, step S203 includes the following steps:

[0114] The graph structure is feature-encoded to obtain a corresponding feature vector.

[0115] In the embodiment, the feature encoding includes node feature encoding and edge feature encoding. Specifically, the node feature encoding includes converting the attributes (such as position, size, and fragility) of each node (such as an obstacle or an environmental element) in the graph structure into a numerical vector. For example, the position is represented by two-dimensional or three-dimensional coordinates. The size is represented by volume or bounding box size. The fragility is represented by a Boolean value (True / False) or a continuous value (0-1).

[0116] The edge feature encoding includes converting the information of the edge (such as a spatial relationship or a similarity relationship) into a numerical vector. For example, the spatial relationship is represented by the Euclidean distance or the direction vector between nodes. The similarity relationship is represented by the attribute similarity (such as material and function) or the semantic similarity (such as both being furniture).

[0117] The feature vector can be obtained by integrating all the coding vectors obtained by coding the node features and the edge features of the graph structure.

[0118] The graph neural network is called.

[0119] In the embodiment, the graph neural network is not specifically limited, and can be determined according to actual business requirements. For example, a graph convolutional network (GCN) or a graph attention network (GAT) can be used.

[0120] The feature vector is input into the graph neural network, and the graph neural network is used to perform graph convolution operation on the feature vector to obtain a corresponding operation result.

[0121] In the embodiment, the graph convolution operation includes initialization, information aggregation and multi-layer propagation processing. Specifically, the initialization includes inputting the feature vectors of the nodes and edges into the graph neural network. The information aggregation includes aggregating the information of the neighbor nodes by using the GCN or the GAT. For example, the GCN performs weighted average on the features of the neighbor nodes to update the representation of the current node. The GAT dynamically calculates the importance weight of the neighbor nodes by using the attention mechanism to update the representation of the current node. The multi-layer propagation includes repeating the information aggregation operation multiple times to gradually expand the receptive field and capture the global relationship.

[0122] The operation result is subjected to relationship reasoning to obtain a corresponding relationship reasoning result.

[0123] In the embodiment, the graph neural network can be used to learn the complex relationship between the obstacle and the environment through nonlinear transformation, for example, “the vase is fragile and located at the entrance of a narrow passage, and needs to be avoided from collision.” “The box is stable and large in size, and is difficult to cross but can be moved.” A graph-level representation or a node-level representation containing global information of the scene is generated as the relationship reasoning result for subsequent action prediction.

[0124] The relationship reasoning result is used as the analysis result of the graph structure.

[0125] In the embodiment, the analysis result includes the learned graph-level or node-level representation. The analysis result can be passed to a downstream task (such as action prediction), and the attributes of the obstacle and the environmental relationship can be determined through the analysis result of the graph neural network, thereby providing a basis for generating a candidate action and selecting an optimal action.

[0126] The application encodes features of the graph structure to obtain a corresponding feature vector, then calls the graph neural network, then inputs the feature vector into the graph neural network, performs graph convolution operation on the feature vector based on the graph neural network to obtain a corresponding operation result, subsequently performs relationship reasoning on the operation result to obtain a corresponding relationship reasoning result, and takes the relationship reasoning result as the analysis result of the graph structure. The application encodes features of the graph structure to obtain a feature vector, then performs graph convolution operation on the feature vector based on the use of the graph neural network to obtain an operation result, and then performs relationship reasoning on the obtained operation result to obtain a relationship reasoning result and take the relationship reasoning result as the final analysis result of the graph structure, so that the analysis and processing of the graph structure can be efficiently and accurately completed, which is conducive to accurately capturing the complex relationship between the obstacle and the environment, and provides rich semantic and spatial information support for subsequent action decision-making.

[0127] In some optional implementations, the obtained user information seeks user consent and complies with relevant laws and relevant policies.

[0128] In addition, the non-company software tools or components appearing in the embodiments of the application are only examples and do not represent actual use.

[0129] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the application.

[0130] It should be emphasized that, in order to further ensure the privacy and security of the above target explanation information, the above target explanation information can also be stored in a node of a block chain.

[0131] The block chain referred to in the application is a new application mode of distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm and other computer technologies. The block chain (Blockch ain) is essentially a decentralized database, which is a series of data blocks associated using cryptographic methods, each data block containing information of a batch of network transactions, used to verify the validity (anti-fake) of the information and generate the next block. The block chain can include a block chain underlying platform, a platform product service layer, and an application service layer.

[0132] The embodiments of the application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence (ArtificialIntelligence, AI) is the use of digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. Theory, method, technology and application system.

[0133] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through computer readable instructions, and the computer readable instructions can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiment methods. Among them, the storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0134] It should be understood that although each step in the flowchart of the accompanying drawings is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and they can be executed in other orders. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence is not necessarily sequential, but can be alternately executed with other steps or sub-steps or stages of other steps.

[0135] Further referring to Figure 3 , as an implementation of the method shown in Figure 2 , the present application provides an embodiment of an action prediction device based on artificial intelligence, which corresponds to the method embodiment shown in Figure 2 , and the device can be specifically applied to various electronic devices.

[0136] As shown in Figure 3 , the action prediction device based on artificial intelligence 300 described in the embodiment includes an extraction module 301, a construction module 302, an analysis module 303, a prediction module 304, an evaluation module 305, a generation module 306, and an output module 307. Among them:

[0137] The extraction module 301 is configured to extract key information in a target scene where the robot is located based on a preset visual device;

[0138] The construction module 302 is configured to construct a corresponding graph structure based on the key information;

[0139] The analysis module 303 is configured to analyze the graph structure based on a preset graph neural network to obtain a corresponding analysis result;

[0140] The prediction module 304 is configured to perform action prediction on the analysis result based on a preset large language model to obtain a plurality of candidate actions.

[0141] The evaluation module 305 is configured to perform benefit evaluation processing on all the candidate actions based on a preset reinforcement learning policy network, to determine a target action from all the candidate actions.

[0142] The generation module 306 is configured to obtain demand information of a target user, and generate target explanation information corresponding to the target action based on the demand information.

[0143] The output module 307 is configured to output the target action to a controller corresponding to the robot, and output the target action and the target explanation information to the target user.

[0144] In the embodiment, the above modules or units are respectively used for performing operations corresponding to steps of the action prediction method based on artificial intelligence of the foregoing embodiments, and thus will not be described here.

[0145] In some optional implementations of the embodiment, the evaluation module 305 includes:

[0146] The first obtaining sub-module is configured to call a pre-constructed reinforcement learning policy network through the first calling sub-module.

[0147] The evaluation sub-module is configured to perform expected benefit evaluation on each of the candidate actions based on the reinforcement learning policy network, to obtain a plurality of benefits corresponding thereto.

[0148] The screening sub-module is configured to screen a specified benefit with the highest value from all the benefits, and obtain a specified action corresponding to the specified benefit.

[0149] The first determination sub-module is configured to determine the specified action as the target action.

[0150] In the embodiment, the above modules or units are respectively used for performing operations corresponding to steps of the action prediction method based on artificial intelligence of the foregoing embodiments, and thus will not be described here.

[0151] In some optional implementations of the embodiment, the generation module 306 includes:

[0152] The first obtaining sub-module is configured to obtain demand information of the target user; and the demand information includes ordinary demand information, development demand information, and research demand information.

[0153] The first generation sub-module is configured to, if the demand information is the ordinary demand information, generate target explanation information corresponding to the target action by using the ordinary demand information based on a preset first explanation generation strategy.

[0154] a second generation sub-module, configured to, if the demand information is the development demand information, generate target explanation information corresponding to the target action based on a preset second explanation generation strategy and using the development demand information;

[0155] a third generation sub-module, configured to, if the demand information is the research demand information, generate target explanation information corresponding to the target action by using the research demand information based on a preset third explanation generation strategy.

[0156] In this embodiment, the operations performed by the above modules or units correspond one by one to the steps of the action prediction method based on artificial intelligence in the foregoing embodiments, and thus are not described here again.

[0157] In some optional implementations of this embodiment, the third generation sub-module includes:

[0158] a conversion unit, configured to, if the demand information is the research demand information, convert the analysis result into corresponding natural language description;

[0159] an analysis unit, configured to perform action analysis on the candidate action to obtain corresponding action analysis result;

[0160] a generation unit, configured to generate selection reason corresponding to the target action and reason why other actions are not selected;

[0161] an integration unit, configured to integrate the natural language description, the action analysis result, the selection reason and the reason in content to obtain target explanation information corresponding to the target action.

[0162] In this embodiment, the operations performed by the above modules or units correspond one by one to the steps of the action prediction method based on artificial intelligence in the foregoing embodiments, and thus are not described here again.

[0163] In some optional implementations of this embodiment, the extraction module 301 includes:

[0164] a capture sub-module, configured to capture a scene image in a target scene where the robot is based on the visual device;

[0165] a preprocessing sub-module, configured to pre-process the scene image to obtain a corresponding target scene image;

[0166] a detection sub-module, configured to perform obstacle detection on the target scene image based on a preset target detection algorithm to obtain a corresponding obstacle detection result;

[0167] a first analysis sub-module, configured to perform attribute analysis on the detection result to obtain a corresponding attribute analysis result;

[0168] a second analysis submodule configured to perform environment analysis on the target scene image to obtain a corresponding environment analysis result;

[0169] a consolidation submodule configured to consolidate information of the obstacle detection result, the attribute analysis result, and the environment analysis result to obtain key information corresponding to the target scene.

[0170] In this embodiment, the above modules or units are respectively used for operations corresponding to the steps of the action prediction method based on artificial intelligence of the foregoing embodiments, and thus will not be described here.

[0171] In some optional implementations of this embodiment, the construction module 302 includes:

[0172] a first construction submodule configured to perform node construction processing on the obstacle detection result and the environment analysis result in the key information to obtain a target node;

[0173] a second construction submodule configured to perform edge construction processing on the target node based on the attribute analysis result in the key information to obtain a target edge;

[0174] a second calling submodule configured to call a preset graph database;

[0175] a organization submodule configured to perform organization processing on the target node and the target edge based on the graph database to obtain a graph structure.

[0176] In this embodiment, the above modules or units are respectively used for operations corresponding to the steps of the action prediction method based on artificial intelligence of the foregoing embodiments, and thus will not be described here.

[0177] In some optional implementations of this embodiment, the analysis module 303 includes:

[0178] an encoding submodule configured to perform feature encoding on the graph structure to obtain a corresponding feature vector;

[0179] a third calling submodule configured to call the graph neural network;

[0180] an operation submodule configured to input the feature vector into the graph neural network, perform graph convolution operation on the feature vector based on the graph neural network, and obtain a corresponding operation result;

[0181] an inference submodule configured to perform relationship inference on the operation result to obtain a corresponding relationship inference result;

[0182] a second determination submodule configured to take the relationship inference result as an analysis result of the graph structure.

[0183] In the embodiment, the modules or units described above are respectively used to perform operations corresponding to the steps of the action prediction method based on artificial intelligence of the foregoing embodiments, and thus no further description is provided herein.

[0184] To solve the above technical problems, the embodiment of the present application further provides a computer device. For details, please refer to Figure 4 , Figure 4 The basic structure block diagram of the computer device in the embodiment is shown in the following figure.

[0185] The computer device 4 includes a memory 41, a processor 42, and a network interface 43, which are connected to each other through a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art can understand that the computer device herein is a device capable of automatically performing numerical calculation and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0186] The computer device can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The computer device can perform human-computer interaction with a user through a keyboard, a mouse, a remote controller, a touchpad, a voice control device, and the like.

[0187] The memory 41 includes at least one type of readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or a memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 4. Of course, the memory 41 can also include both an internal storage unit and an external storage device of the computer device 4. In this embodiment, the memory 41 is generally used to store an operating system and various application software installed on the computer device 4, such as computer readable instructions of the action prediction method based on artificial intelligence, etc. In addition, the memory 41 can also be used to temporarily store various data that have been output or will be output.

[0188] The processor 42 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run computer readable instructions or process data stored in the memory 41, such as computer readable instructions of the action prediction method based on artificial intelligence.

[0189] The network interface 43 can include a wireless network interface or a wired network interface, and is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0190] The present application also provides another embodiment, i.e., to provide a computer readable storage medium storing computer readable instructions, which can be executed by at least one processor to make the at least one processor perform the steps of the action prediction method based on artificial intelligence as described above.

[0191] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the method described in each embodiment of the present application.

[0192] Obviously, the above-described embodiments are only some of the embodiments of the present application, not all the embodiments, and the drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms, and conversely, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing specific embodiments, or make equivalent replacements to some of the technical features. Any equivalent structure made by using the content of the specification and drawings, directly or indirectly applied to other related technical fields, is also within the scope of the patent protection of the present application.

Claims

1. An action prediction method based on artificial intelligence, characterized in that: The steps include: Extract key information from the target scene where the robot is located based on the preset visual equipment; Constructing a corresponding graph structure based on the key information; Analyze the graph structure based on a preset graph neural network to obtain corresponding analysis results; Perform action prediction on the analysis results based on a preset large language model to obtain corresponding multiple candidate actions; Performing benefit evaluation on all the candidate actions based on a preset reinforcement learning strategy network to determine a target action from all the candidate actions; Acquire demand information of a target user, and generate target interpretation information corresponding to the target action based on the demand information; The target action is output to a controller corresponding to the robot, and the target action and the target interpretation information are output to the target user.

2. The action prediction method based on artificial intelligence according to claim 1, characterized in that: The step of performing benefit evaluation on all the candidate actions based on the preset reinforcement learning strategy network to determine the target action from all the candidate actions specifically includes: Calling a pre-built reinforcement learning policy network; Performing an expected benefit evaluation on each of the candidate actions based on the reinforcement learning strategy network to obtain corresponding multiple benefits; Filter out the designated benefit with the highest value from all the benefits, and obtain the designated action corresponding to the designated benefit; The designated action is used as the target action.

3. The action prediction method based on artificial intelligence according to claim 1, characterized in that: The step of obtaining the target user's demand information and generating target explanation information corresponding to the target action based on the demand information specifically includes: Obtaining demand information of the target user; wherein the demand information includes general demand information, development demand information and research demand information; If the demand information is the common demand information, then based on a preset first explanation generation strategy, using the common demand information to generate target explanation information corresponding to the target action; If the requirement information is the development requirement information, then based on a preset second explanation generation strategy, using the development requirement information to generate target explanation information corresponding to the target action; If the demand information is the research demand information, target explanation information corresponding to the target action is generated using the research demand information through a preset third explanation generation strategy.

4. The action prediction method based on artificial intelligence according to claim 3, characterized in that: If the demand information is the research demand information, the step of using the research demand information to generate target explanation information corresponding to the target action through a preset third explanation generation strategy specifically includes: If the demand information is the research demand information, converting the analysis result into a corresponding natural language description; Performing action analysis on the candidate action to obtain corresponding action analysis results; Generating reasons for selecting the target action and generating reasons for not selecting other actions; The natural language description, the action analysis result, the selection reason and the justification are content-integrated to obtain target explanation information corresponding to the target action.

5. The action prediction method based on artificial intelligence according to claim 1, characterized in that: The step of extracting key information from the target scene where the robot is located based on a preset visual device specifically includes: Capturing a scene image of a target scene in which the robot is located based on the visual device; Preprocessing the scene image to obtain a corresponding target scene image; Performing obstacle detection on the target scene image based on a preset target detection algorithm to obtain a corresponding obstacle detection result; Performing attribute analysis on the detection result to obtain corresponding attribute analysis results; Performing environmental analysis on the target scene image to obtain corresponding environmental analysis results; The obstacle detection result, the attribute analysis result and the environment analysis result are integrated to obtain key information corresponding to the target scene.

6. The action prediction method based on artificial intelligence according to claim 5, characterized in that: The step of constructing a corresponding graph structure based on the key information specifically includes: Performing node construction processing based on the obstacle detection result in the key information and the environmental analysis result to obtain a corresponding target node; Performing edge construction processing on the target node based on the attribute analysis result in the key information to obtain a corresponding target edge; Call the preset graph database; The target nodes and the target edges are organized and processed based on the graph database to obtain a corresponding graph structure.

7. The action prediction method based on artificial intelligence according to claim 1, characterized in that: The step of analyzing the graph structure based on a preset graph neural network to obtain corresponding analysis results specifically includes: Performing feature encoding on the graph structure to obtain a corresponding feature vector; Calling the graph neural network; Inputting the feature vector into the graph neural network, performing a graph convolution operation on the feature vector based on the graph neural network, and obtaining a corresponding operation result; Performing relational reasoning on the operation result to obtain a corresponding relational reasoning result; The relational reasoning result is used as the analysis result of the graph structure.

8. An action prediction device based on artificial intelligence, characterized in that: include: An extraction module is used to extract key information from the target scene where the robot is located based on a preset visual device; A construction module, configured to construct a corresponding graph structure based on the key information; An analysis module, configured to analyze the graph structure based on a preset graph neural network to obtain corresponding analysis results; A prediction module, configured to perform action prediction on the analysis results based on a preset large language model to obtain a plurality of corresponding candidate actions; An evaluation module, configured to perform benefit evaluation processing on all the candidate actions based on a preset reinforcement learning strategy network, so as to determine a target action from all the candidate actions; A generating module, configured to obtain demand information of a target user and generate target explanation information corresponding to the target action based on the demand information; An output module is used to output the target action to a controller corresponding to the robot, and output the target action and the target interpretation information to the target user.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the action prediction method based on artificial intelligence as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the action prediction method based on artificial intelligence as described in any one of claims 1 to 7.