Fault identification and processing method and device based on artificial intelligence, equipment and medium
By obtaining multimodal information of intelligent robots, building a task scenario diagram and performing fault analysis, the problem of inaccurate fault positioning of robots in the existing technology is solved, and fast and accurate fault identification and processing is achieved.
Patent Information
- Application Number
- CN202510613446.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-04
AI Technical Summary
Existing intelligent robots find it difficult to quickly and accurately locate the cause of the failure through simple status indicators or sensor data feedback, making it difficult for users to effectively monitor the working status of the robot and handle the failure in a timely manner, affecting the application effect and user experience.
By obtaining multimodal information of intelligent robots performing tasks, building a task scene diagram and generating audio event tags, using a large language model to identify keyframes and text description information, performing fault analysis and error recognition, and generating fault correction tasks.
It realizes the conversion of multimodal state information of intelligent robots into easy-to-understand text descriptions, quickly and accurately locates the cause of failure and provides effective solutions, improving the efficiency and accuracy of fault handling.
Smart Images

Figure CN120245081A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and medium for fault identification and processing based on artificial intelligence. Background Art
[0002] Intelligent robots are widely used in various life and work scenarios, such as smart home, medical and health care, etc. The interaction between intelligent robots and humans is becoming increasingly frequent, and people's demand for understanding robot behavior and fault handling ability is growing day by day. However, traditional robot information display methods, such as simple status indicators or basic sensor data feedback, far from meet the needs of in-depth understanding of robot behavior and fault location.
[0003] On the one hand, the robot status information provided by simple status indicators is too little. For users, it is impossible to comprehensively understand the movement process of the robot through limited status indication information. On the other hand, it is difficult for users to obtain the internal parameter information of the robot, which undoubtedly increases the difficulty of debugging work. For example, cleaning robots in smart homes or nursing robots in medical and health usually have simple working status indicators and power indicators to display the current status and battery life of the robot. However, when the robot has an abnormality (stuck or in an infinite loop state), it is difficult to judge whether the robot is working properly through the indicator lights, and analyzing through the internal parameter information of the robot can more clearly judge the robot status. In terms of fault analysis, existing methods usually rely on preset rules in specific fields to identify faults in sensor data, lack generality and are inefficient, and it is difficult to quickly and accurately locate the cause of the fault and provide effective solutions.
[0004] The inventor realizes that the above solutions have significant deficiencies in converting the behavior and data of intelligent robots themselves into information that can be understood by humans, resulting in difficulty for users to effectively monitor the working status of the robot and timely handle faults, affecting the application effect and user experience of the robot. Summary of the Invention
[0005] The present invention provides a method, device, computer equipment and medium for fault identification and processing based on artificial intelligence to solve the technical problems that the behavior information display of intelligent robots is not comprehensive, and it is difficult to quickly and accurately locate the cause of the fault and provide effective solutions.
[0006] In a first aspect, a method for fault identification and processing based on artificial intelligence is provided, including:
[0007] Obtaining multi-modal information of an intelligent robot executing a task, where the multi-modal information includes image information and audio information;
[0008] Construct a task scenario graph based on the image information, generate audio event tags based on the audio information, and identify key frames from the task scenario graph according to the audio event tags and generate corresponding key scene text description information;
[0009] Use a large language model to identify the multiple sub-task completion information corresponding to the intelligent robot's task execution based on the key scene text description information and the corresponding audio event tags, perform fault information detection on each sub-task completion information based on a pre-set sub-task completion target, and when a fault information is detected in the sub-task completion information, perform fault analysis based on the corresponding key scene text description information to obtain a fault analysis result;
[0010] If no fault information is detected in all sub-task completion information, use a large language model to identify the execution result corresponding to the intelligent robot's task execution, perform fault information detection on the execution result, and when a fault information is detected in the execution result, obtain the task description and task plan of the intelligent robot's task execution and input them together with the execution result into the large language model for error identification to obtain an error identification result;
[0011] Generate a corresponding fault correction task for the intelligent robot's task execution according to the fault analysis result or the error identification result for the intelligent robot to execute the fault correction task.
[0012] In a second aspect, there is provided a fault identification and processing device based on artificial intelligence, including:
[0013] An information acquisition module for acquiring multi-modal information of an intelligent robot's task execution, where the multi-modal information includes image information and audio information;
[0014] An information processing module for constructing a task scenario graph based on the image information, generating audio event tags based on the audio information, identifying key frames from the task scenario graph according to the audio event tags and generating corresponding key scene text description information;
[0015] A fault analysis module for using a large language model to identify the multiple sub-task completion information corresponding to the intelligent robot's task execution based on the key scene text description information and the corresponding audio event tags, performing fault information detection on each sub-task completion information based on a pre-set sub-task completion target, and when a fault information is detected in the sub-task completion information, performing fault analysis based on the corresponding key scene text description information to obtain a fault analysis result;
[0016] An error recognition module, which is used to, if no fault information is detected in all subtask completion information, recognize the execution result corresponding to the task executed by the intelligent robot through a large language model, detect fault information in the execution result, and when fault information is detected in the execution result, obtain the task description and task plan of the task executed by the intelligent robot, and input them together with the execution result into the large language model for error recognition to obtain an error recognition result;
[0017] A fault correction module, which is used to generate a corresponding fault correction task for the task executed by the intelligent robot according to the fault analysis result or the error recognition result, so that the intelligent robot can execute the fault correction task.
[0018] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned fault recognition and processing method are implemented.
[0019] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned fault recognition and processing method are implemented.
[0020] In the solution implemented by the above-mentioned artificial intelligence-based fault recognition and processing method, device, computer device, and storage medium, multi-modal information of the task executed by the intelligent robot can be obtained, a task scenario graph can be constructed based on the image information of the multi-modal information, an audio event label can be generated based on the audio information of the multi-modal information, and key frames can be recognized from the task scenario graph according to the audio event label and corresponding key scene text description information can be generated; the key scene text description information and the corresponding audio event label are recognized through a large language model to obtain multiple subtask completion information corresponding to the task executed by the intelligent robot, and fault information detection and fault analysis are performed on each subtask completion information; fault information detection is performed on the execution result corresponding to the task executed by the intelligent robot, and when fault information is detected in the execution result, error recognition is performed on the task description and task plan of the task executed by the intelligent robot and the execution result; a corresponding fault correction task for the task executed by the intelligent robot is generated according to the fault analysis result or the error recognition result, so as to convert the multi-modal state information of the intelligent robot into text description information that is easy to understand by an artificial intelligence model, thereby quickly and accurately locating the cause of the fault and providing an effective solution. Description of the Drawings
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a schematic diagram of an application environment of a fault identification and handling method in an embodiment of the present invention;
[0023] Figure 2 It is a schematic flowchart of a fault identification and handling method in an embodiment of the present invention;
[0024] Figure 3 It is a schematic structural diagram of a fault identification and handling device in an embodiment of the present invention;
[0025] Figure 4 It is a schematic structural diagram of a computer device in an embodiment of the present invention;
[0026] Figure 5 It is another schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0028] The fault identification and handling method based on artificial intelligence provided by the embodiments of the present invention can be applied in, for example Figure 1In the application environment, the client communicates with the server through the network. The server can receive the fault identification and processing instructions sent by the client, obtain the multi-modal information of the intelligent robot performing tasks according to the instructions, construct a task scenario graph based on the image information of the multi-modal information, generate audio event tags based on the audio information of the multi-modal information, identify key frames from the task scenario graph according to the audio event tags and generate corresponding key scene text description information; identify the corresponding multiple sub-task completion information of the intelligent robot performing tasks through a large language model for the key scene text description information and the corresponding audio event tags, and perform fault information detection and fault analysis on each sub-task completion information; perform fault information detection on the execution result of the intelligent robot performing tasks, and perform error identification on the task description, task plan and execution result of the intelligent robot performing tasks when fault information is detected in the execution result; generate corresponding fault correction tasks for the intelligent robot performing tasks according to the fault analysis result or error identification result, so as to convert the multi-modal state information of the intelligent robot into an easy-to-understand text description information by means of an artificial intelligence model, thereby quickly and accurately locating the cause of the fault and providing an effective solution. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0029] Please refer to Figure 2 as shown in Figure 2 a schematic flowchart of a method for fault identification and processing based on artificial intelligence provided by an embodiment of the present invention, including the following steps:
[0030] S1: Obtain the multi-modal information of the intelligent robot performing tasks, where the multi-modal information includes image information and audio information.
[0031] The fault identification and processing method provided by the present invention can be applied to intelligent robots in various application scenarios. When performing tasks, the intelligent robot obtains multimodal information through a sensor device. For example, in the field of smart homes, when performing floor cleaning tasks, the intelligent robot obtains the surrounding image information and audio information through sensors when performing tasks; in the field of medical health, when performing nursing tasks, the intelligent robot records the nursing process through sensors to obtain a nursing video containing image information and audio information. Image information includes that the nursing robot needs to identify the position of obstacles such as beds, tables and chairs in hospital wards or nursing institutions in order to move safely; image acquisition of the display screen of medical equipment (such as infusion pumps, monitors, etc.) to monitor the operating status of the equipment and whether the parameter display is normal; image analysis is used to determine whether the robot correctly delivers drugs to patients; image analysis and diagnostic suggestions are performed by obtaining X-rays, CT images, etc. of patients in combination with multimodal large language models. Audio information includes patients being able to use voice commands to let the robot adjust the height of the bed or play music; medical staff being able to use voice commands to let the robot deliver medicines to designated wards; when the patient cries for help, the robot can quickly identify and notify medical staff; and by monitoring the sound of the robot moving, it can be determined whether there is a mechanical failure, etc.
[0032] In this embodiment, the image information is a multi-frame depth image arranged in time sequence obtained by the sensor during the intelligent robot's task execution, and the audio information is audio data with time frame marks obtained by the sensor during the intelligent robot's task execution. For example: when the intelligent robot starts to execute the task, it turns on the sensor synchronously to obtain image information and audio information during the task execution. The intelligent robot can obtain video through a multimodal sensor, extract multi-frame image information and audio information from the video, and add time frame marks to the image information and audio information according to the time sequence of the video; it can also simultaneously obtain multi-frame image information through a visual sensor (such as a camera device) and add a time frame mark to the image information according to the acquisition time, and obtain audio information with time frame marks through an auditory sensor (such as a recording device).
[0033] For example, after receiving a task to clean the floor of a room, the intelligent robot starts the sweeping function and simultaneously starts the camera to record audio and video. After the intelligent robot completes the cleaning task and finishes sweeping the floor, it stops recording to obtain a video with the recording time. Depending on the configuration of the intelligent robot's sensors, it can also take images through a camera and record sounds through a recorder to obtain multiple images with the shooting time and audio with the recording time. Through videos with recording time or multiple images with shooting time and audio with recording time, it is possible to build intelligent robot cleaning scenarios, analyze cleaning behaviors and cleaning effects, and fully understand the robot's working behaviors and status for fault analysis.
[0034] For another example, after receiving a leg care task, a nursing robot in the field of medical health activates the leg care function and simultaneously activates the camera to record audio and video, and stops recording after the nursing robot completes the leg care task, obtaining a video with the recording time. According to the configuration of the robot sensor, it can also be to take images through a camera and record sounds through a recorder, obtaining multiple images with the shooting time and audio with the recording time. Through the video with the recording time or multiple images with the shooting time and audio with the recording time, the care scenario of the nursing robot can be constructed, the care behavior and the feedback of the care recipient can be analyzed, and the working behavior and state of the robot can be comprehensively understood for fault analysis.
[0035] S2. Construct a task scenario graph based on the image information, generate audio event tags based on the audio information, and identify key frames from the task scenario graph according to the audio event tags and generate corresponding key scene text description information.
[0036] Among them, constructing a task scenario graph based on the image information includes the following steps:
[0037] Detect and semantically segment the image information to obtain multiple object objects;
[0038] Based on the depth information of the image information, convert the object objects into corresponding 3D semantic point clouds;
[0039] Aggregate multiple 3D semantic point clouds to obtain an aggregated point cloud, and construct a task scenario graph according to the aggregated point cloud.
[0040] Specifically, the image information is a depth image (RGB-D image) containing RGB three-channel color images and depth information (Depth Map). The plane detection model is used to perform graphic detection on multiple depth images in the image information respectively to obtain the object graphics of each depth image, and the convolutional neural network is used to perform semantic detection and segmentation on the object graphics to obtain multiple object objects corresponding to each depth image, and 3D semantic point cloud modeling is performed on the multiple object objects corresponding to the image in the three-dimensional space based on the depth information of the depth image.
[0041] Specifically, aggregating multiple 3D semantic point clouds to obtain an aggregated point cloud and constructing a task scenario graph according to the aggregated point cloud includes:
[0042] Identify the depth image with the earliest time point from multiple depth images in the image information as the initial image according to the time frame label, and determine the other images as the images to be aggregated;
[0043] Use the 3D semantic point cloud corresponding to the initial image as the initial point cloud, and use the 3D semantic point cloud corresponding to the image to be aggregated as the point cloud to be aggregated;
[0044] Calculate the horizontal displacement of the object according to the initial image and the image to be aggregated, and calculate the normal vector of the object according to the image to be aggregated;
[0045] Construct an alignment coordinate system based on the initial point cloud, calculate the transformed point cloud obtained by transforming the point cloud to be aggregated in the alignment coordinate system according to the horizontal displacement and the normal vector, and construct a three-dimensional network model based on the initial point cloud and the transformed point cloud to obtain a task scene graph.
[0046] For example: During the task execution of an intelligent robot, 3 task images 1, 2, and 3 are sequentially acquired. Graphic detection is performed on task image 1 to obtain 3 object graphics (such as obstacles in a smart home scenario or care objects in a medical and health scenario) and construct corresponding 3D semantic point clouds A1, B1, C1. Graphic detection is performed on task image 2 to obtain 3 object graphics and construct corresponding 3D semantic point clouds A2, B2, C2. Graphic detection is performed on task image 3 to obtain 3 object graphics and construct corresponding 3D semantic point clouds A3, B3, C3. Take the earliest acquired task image 1 as the initial image, construct an alignment coordinate system for the initial 3D semantic point clouds A1, B1, C1, transform A2 in the alignment coordinate system according to the horizontal displacement and the normal vector between A2 and A1 to obtain the transformed point cloud A2', and obtain the transformed point cloud B2' of B2, the transformed point cloud C2' of C2, the transformed point cloud A3' of A3, the transformed point cloud B3' of B3, and the transformed point cloud C3' of C3 in the same way. Construct a three-dimensional network model based on A1, B1, C1, A2', B2', C2', A3', B3', C3' in the alignment coordinate system to obtain a task scene graph.
[0047] Among them, generating an audio event label based on the audio information includes:
[0048] Perform audio segmentation on the audio information according to a preset volume threshold to obtain multiple segments of audio;
[0049] Calculate the text embedding vector corresponding to each segmented audio through a pre-trained audio-language model;
[0050] Calculate the cosine similarity between the text embedding vector corresponding to each segment of audio and the vectors of multiple preset audio event labels respectively, and identify the audio event label matched by the text embedding vector corresponding to each segment of audio based on the cosine similarity.
[0051] Specifically, performing audio segmentation on the audio information according to a preset volume threshold includes:
[0052] Continuously detect the speech intensity of the audio information;
[0053] When the speech intensity is less than the preset volume threshold, it is determined as a silent segment;
[0054] When the voice intensity is greater than or equal to the preset volume threshold, it is determined as an audible segment;
[0055] Delete the silent segments in the audio information to obtain an audio of multiple audible segments.
[0056] Among them, identifying key frames from the task scenario graph according to the audio event label and generating corresponding key scenario text description information includes:
[0057] Identifying the start time point and end time point of the corresponding audio event according to the audio event label;
[0058] Extracting the task scenario graph corresponding to the start time point and the end time point;
[0059] Traversing the task scenario graphs of each time point in chronological order, and when the task scenario graph changes, extracting the changed task scenario graph;
[0060] Extracting the task scenario graphs corresponding to the completion time points of multiple subtasks corresponding to the tasks executed by the intelligent robot;
[0061] Identifying the extracted task scenario graphs as key frames, and generating corresponding key scenario text description information for each key frame through an artificial intelligence model.
[0062] In this embodiment, the artificial intelligence model clearly describes the actions being performed by the robot in natural language, enabling the operator to better understand the robot's work process.
[0063] Specifically, identifying the start time point and end time point of the corresponding audio event according to the audio event label includes:
[0064] Obtaining the audio matching the audio event label, identifying the start time point and end time point of the audio, and taking the start time point and end time point as the start time point and end time point of the audio event corresponding to the audio event label matched by this segment of audio.
[0065] For example: In the smart home scenario, the audio event label "clean the corner" of the intelligent robot performing the sweeping task matches 4 segments of audio. According to the start time frame and end time frame of each segment of audio, the start time point and end time point corresponding to each segment of audio are identified, and 8 time points (4 start time points and 4 end time points) of the audio event corresponding to the audio event label are obtained.
[0066] For another example: In the medical and health scenario, the audio event label "leg massage" for the intelligent robot to perform nursing tasks needs to be completed 3 times, respectively matching 3 segments of audio. According to the start time frame and end time frame of each segment of audio, the start time point and end time point corresponding to each segment of audio are identified, and 6 time points (3 start time points and 3 end time points) of the audio event corresponding to the audio event label are obtained.
[0067] S30. Identify the key scenario text description information and the corresponding audio event labels through the large language model to obtain multiple sub-task completion information corresponding to the tasks performed by the intelligent robot. Based on the pre-set sub-task completion goals, perform fault information detection on each sub-task completion information. When fault information is detected in the sub-task completion information, perform fault analysis based on the corresponding key scenario text description information to obtain a fault analysis result.
[0068] Among them, identifying the key scenario text description information and the corresponding audio event labels through the large language model to obtain multiple sub-task completion information corresponding to the tasks performed by the intelligent robot includes:
[0069] Obtain the key scenario text description information and audio event labels corresponding to the multiple sub-task completion time points of the tasks performed by the intelligent robot;
[0070] Perform text summarization on the key scenario text description information and audio event labels through the large language model to obtain multiple sub-task completion information corresponding to the tasks performed by the intelligent robot.
[0071] Among them, performing fault information detection on each sub-task completion information based on the pre-set sub-task completion goals, and when fault information is detected in the sub-task completion information, performing fault analysis based on the corresponding key scenario text description information to obtain a fault analysis result includes:
[0072] Obtain the pre-set sub-task completion goals, and the pre-set sub-task completion goals include multiple completion indicators for the corresponding sub-tasks;
[0073] Compare each sub-task completion information with the corresponding completion indicators, and confirm whether the sub-task corresponding to each sub-task completion information has completed all task indicators according to the comparison result;
[0074] If so, judge that no fault information appears in the sub-task completion information;
[0075] If not, judge that fault information appears in the sub-task completion information, extract the sub-task completion information corresponding to the uncompleted indicators, obtain the corresponding key scenario text description information and perform fault analysis based on the artificial intelligence model to obtain the reason for the uncompleted indicators as the fault analysis result.
[0076] For example, in a healthcare scenario, the task performed by an intelligent robot is "leg medical care", which includes several subtasks: "leg cleaning", "leg massage", and "leg rehabilitation". Obtain the completion metrics for the "leg cleaning" subtask (such as cleanliness and time), use a large language model of artificial intelligence to perform text summarization to obtain the completion information for the "leg cleaning" subtask, and compare it with the completion metrics to determine that all metrics for the "leg cleaning" subtask have been completed. Obtain the completion metrics for the "leg massage" subtask (such as intensity, steps, and time for each step), use a large language model of artificial intelligence to perform text summarization to obtain the completion information for the "leg massage" subtask, and compare it with the completion metrics to determine that all metrics for the "leg massage" subtask have been completed. Obtain the completion metrics for the "leg rehabilitation" subtask (such as rehabilitation items and time for each item), use a large language model of artificial intelligence to perform text summarization to obtain the completion information for the "leg rehabilitation" subtask, and compare it with the completion metrics to determine that not all metrics for the "leg massage" subtask have been completed. Based on the uncompleted metrics, obtain the corresponding key scenario text description information for fault analysis, and the fault analysis result is that the duration of the rehabilitation item is less than the preset duration.
[0077] S40. If no fault information is detected in the completion information of all subtasks, use the large language model to identify the execution result corresponding to the task performed by the intelligent robot, perform fault information detection on the execution result. When fault information is detected in the execution result, obtain the task description and task plan of the task performed by the intelligent robot, and input them together with the execution result into the large language model for error identification to obtain the error identification result.
[0078] Among them, using the large language model to identify the execution result corresponding to the task performed by the intelligent robot and performing fault information detection on the execution result includes:
[0079] Obtain the completion metrics of the task performed by the intelligent robot, and compare the completion metrics with the final execution result of the task performed by the intelligent robot through the large language model;
[0080] Determine whether there are uncompleted metrics based on the comparison result. If so, determine that fault information appears in the execution result.
[0081] In this embodiment, when a fault occurs, it will automatically identify the fault and give an explanation of the fault cause for reference, helping maintenance personnel quickly locate the fault cause and reducing maintenance time and costs.
[0082] S50. Generate a corresponding fault correction task for the task performed by the intelligent robot according to the fault analysis result or the error identification result, for the intelligent robot to execute the fault correction task.
[0083] In one embodiment, the fault identification and handling method further includes:
[0084] Obtaining the task text description information of the intelligent robot executing the task;
[0085] Obtaining the parameter text description information of the intelligent robot;
[0086] Performing a logical consistency judgment on the task text description information, the parameter text description information, and the key scenario text description information, and performing fault analysis and error identification based on the judgment result.
[0087] Specifically, a fine-tuned classification model is used to classify the logical relationships of the task text description information, the parameter text description information, and the key scenario text description information. The categories of the logical relationship classification are one of mutually supportive, mutually non-supportive, and mutually contradictory. If the logical relationship classification is mutually supportive, it is determined that the task text description information, the parameter text description information, and the key scenario text description information are logically consistent; if the logical relationship classification is mutually non-supportive, the task text description information, the parameter text description information, and the key scenario text description information of the intelligent robot are regenerated for logical consistency judgment; if the logical relationship classification is mutually contradictory, a robot status description is generated according to the task text description information, the parameter text description information, and the key scenario text description information and sent to the user management terminal for fault warning and manual troubleshooting and correction.
[0088] Among them, generating the corresponding fault correction task for the intelligent robot to execute the task according to the fault analysis result or the error identification result includes:
[0089] Extracting keywords from the fault analysis result or the error identification result and converting them into prompt words in a preset format;
[0090] Generating a corresponding plurality of task execution steps by using the preset artificial intelligence language model for the prompt words;
[0091] Mapping each task execution step into a task scenario graph through a pre-trained sentence embedding model for correction, and generating a corresponding fault correction task based on the corrected plurality of task execution steps.
[0092] It can be seen that in the above solution, for intelligent robots in the fields of smart home, medical health, etc., by obtaining multi-modal information of the intelligent robot performing tasks, constructing a task scenario graph based on the image information of the multi-modal information, generating audio event tags based on the audio information of the multi-modal information, identifying key frames from the task scenario graph according to the audio event tags and generating corresponding key scenario text description information; by using a large language model to identify the key scenario text description information and the corresponding audio event tags, obtaining multiple sub-task completion information corresponding to the intelligent robot performing tasks, and performing fault information detection and fault analysis on each sub-task completion information; performing fault information detection on the execution result corresponding to the intelligent robot performing tasks, and performing error identification on the task description, task plan and execution result of the intelligent robot performing tasks when fault information is detected in the execution result; generating corresponding fault correction tasks for the intelligent robot performing tasks according to the fault analysis result or error identification result, realizing the conversion of the multi-modal state information of the intelligent robot into easily understandable text description information by an artificial intelligence model, so as to quickly and accurately locate the cause of the fault and provide an effective solution.
[0093] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0094] In one embodiment, a fault identification and processing device based on artificial intelligence is provided, and the fault identification and processing device based on artificial intelligence corresponds one-to-one with the fault identification and processing method based on artificial intelligence in the above embodiment. As Figure 3 shown, the fault identification and processing device includes an information acquisition module 101, an information processing module 102, a fault analysis module 103, an error identification module 104 and a fault correction module 105. The detailed description of each functional module is as follows:
[0095] The information acquisition module 101 is configured to acquire multi-modal information of the intelligent robot performing tasks, and the multi-modal information includes image information and audio information;
[0096] The information processing module 102 is configured to construct a task scenario graph based on the image information, generate audio event tags based on the audio information, identify key frames from the task scenario graph according to the audio event tags and generate corresponding key scenario text description information;
[0097] The fault analysis module 103 is used to identify the key scenario text description information and the corresponding audio event tags through a large language model, obtain multiple sub-task completion information corresponding to the tasks performed by the intelligent robot, detect fault information for each sub-task completion information based on a preset sub-task completion target, and when fault information is detected in the sub-task completion information, perform fault analysis based on the corresponding key scenario text description information to obtain a fault analysis result;
[0098] The error recognition module 104 is used to, if no fault information is detected in all sub-task completion information, identify the execution result corresponding to the task performed by the intelligent robot through a large language model, detect fault information for the execution result, and when fault information is detected in the execution result, obtain the task description and task plan of the task performed by the intelligent robot, and input them together with the execution result into the large language model for error recognition to obtain an error recognition result;
[0099] The fault correction module 105 is used to generate corresponding fault correction tasks for the tasks performed by the intelligent robot according to the fault analysis result or the error recognition result, for the intelligent robot to execute the fault correction tasks.
[0100] In one embodiment, the information processing module 102 is specifically used for:
[0101] Detect and semantically segment the image information to obtain multiple object objects;
[0102] Based on the depth information of the image information, convert the object objects into corresponding 3D semantic point clouds;
[0103] Aggregate multiple 3D semantic point clouds to obtain an aggregated point cloud, and construct a task scene graph according to the aggregated point cloud.
[0104] In one embodiment, the information processing module 102 is specifically used for:
[0105] Perform audio cutting on the audio information according to a preset volume threshold to obtain multiple segments of audio;
[0106] Through a pre-trained audio-language model, calculate the text embedding vectors corresponding to each segment of the cut audio;
[0107] Calculate the cosine similarity between the text embedding vector corresponding to each segment of audio and a preset multiple of audio event tags respectively, and identify the audio event tags matched by the text embedding vector corresponding to each segment of audio based on the cosine similarity.
[0108] In one embodiment, the information processing module 102 is specifically used for:
[0109] Identify the start time point and end time point of the corresponding audio event according to the audio event tag;
[0110] Extract the task scenario graph corresponding to the start time point and the end time point;
[0111] Traverse the task scenario graphs of each time point in chronological order. When the task scenario graph changes, extract the changed task scenario graph;
[0112] Extract the task scenario graphs corresponding to the completion time points of multiple subtasks executed by the intelligent robot;
[0113] Identify the extracted task scenario graphs as key frames, and generate key scene text description information corresponding to each key frame through an artificial intelligence model.
[0114] In one embodiment, the fault analysis module 103 is specifically configured to:
[0115] Obtain the key scene text description information and audio event tags corresponding to the completion time points of multiple subtasks executed by the intelligent robot;
[0116] Perform text summarization on the key scene text description information and audio event tags through a large language model to obtain the completion information of multiple subtasks executed by the intelligent robot.
[0117] In one embodiment, the fault correction module 105 is specifically configured to:
[0118] Analyze the fault analysis result or the error recognition result to obtain the uncompleted task metrics for the intelligent robot to execute tasks;
[0119] Obtain the task prompt words corresponding to the uncompleted task metrics, and input the task prompt words into a preset language model to generate corresponding task execution actions;
[0120] Map the task execution actions to the task scenario graph through a pre-trained sentence embedding model, and adjust the task execution actions based on the task scenario graph;
[0121] Generate the task steps of the fault correction task according to the adjusted task execution actions.
[0122] In one embodiment, the fault identification and processing device further includes a logical consistency module 106, specifically configured to:
[0123] Obtain the task text description information of the intelligent robot to execute tasks;
[0124] Obtain the parameter text description information of the intelligent robot;
[0125] Perform a logical consistency check on the task text description information, the parameter text description information, and the key scenario text description information, and perform fault analysis and error identification based on the judgment result.
[0126] The present invention provides a fault identification and processing device. By obtaining multi-modal information of an intelligent robot performing a task, constructing a task scenario graph based on the image information of the multi-modal information, generating an audio event label based on the audio information of the multi-modal information, identifying key frames from the task scenario graph according to the audio event label and generating corresponding key scenario text description information; identifying the key scenario text description information and the corresponding audio event label through a large language model to obtain multiple sub-task completion information corresponding to the intelligent robot performing the task, and performing fault information detection and fault analysis on each sub-task completion information; performing fault information detection on the execution result corresponding to the intelligent robot performing the task, and performing error identification on the task description, task plan and execution result of the intelligent robot performing the task when it is detected that there is fault information in the execution result; generating a corresponding fault correction task for the intelligent robot performing the task according to the fault analysis result or error identification result, so as to realize converting the multi-modal state information of the intelligent robot into easily understandable text description information by means of an artificial intelligence model, thereby quickly and accurately locating the cause of the fault and providing an effective solution.
[0127] For the specific limitations of the fault identification and processing device, reference can be made to the limitations of the intelligent question answering method in the above text, which will not be elaborated here. Each module in the above fault identification and processing device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0128] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage media. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a fault identification and processing method based on artificial intelligence.
[0129] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as shown in Figure 5 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a fault recognition and processing method based on artificial intelligence
[0130] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:
[0131] Obtain multimodal information of the intelligent robot performing a task, where the multimodal information includes image information and audio information;
[0132] Construct a task scenario graph based on the image information, generate audio event tags based on the audio information, identify key frames from the task scenario graph according to the audio event tags, and generate corresponding key scenario text description information;
[0133] Identify, through a large language model, the multimodal information corresponding to the intelligent robot performing the task, obtain multiple subtask completion information corresponding to the intelligent robot performing the task, perform fault information detection on each subtask completion information based on a preset subtask completion target. When a fault information is detected in the subtask completion information, perform fault analysis based on the corresponding key scenario text description information to obtain a fault analysis result;
[0134] If no fault information is detected in all subtask completion information, identify, through a large language model, the execution result corresponding to the intelligent robot performing the task, perform fault information detection on the execution result. When a fault information is detected in the execution result, obtain the task description and task plan of the intelligent robot performing the task, and input them together with the execution result into the large language model for error identification to obtain an error identification result;
[0135] Generate a corresponding fault correction task for the intelligent robot performing the task according to the fault analysis result or the error identification result, for the intelligent robot to execute the fault correction task.
[0136] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0137] Obtain multimodal information of the intelligent robot performing a task, where the multimodal information includes image information and audio information;
[0138] Construct a task scenario graph based on the image information, generate an audio event label based on the audio information, and identify key frames from the task scenario graph according to the audio event label and generate corresponding key scene text description information;
[0139] Identify, through a large language model, the multimodal information corresponding to the intelligent robot performing the task, obtain multiple subtask completion information corresponding to the intelligent robot performing the task, perform fault information detection on each subtask completion information based on a preset subtask completion target, and when a fault information is detected in the subtask completion information, perform fault analysis based on the corresponding key scene text description information to obtain a fault analysis result;
[0140] If no fault information is detected in all subtask completion information, identify, through a large language model, the execution result corresponding to the intelligent robot performing the task, perform fault information detection on the execution result, and when a fault information is detected in the execution result, obtain the task description and task plan of the intelligent robot performing the task and input them together with the execution result into the large language model for error identification to obtain an error identification result;
[0141] Generate a corresponding fault correction task for the intelligent robot performing the task according to the fault analysis result or the error identification result, for the intelligent robot to execute the fault correction task.
[0142] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can achieve, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0143] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0144] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0145] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use.
[0146] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A fault identification and processing method based on artificial intelligence, characterized in that, Including: Obtain multi-modal information of the intelligent robot performing a task, where the multi-modal information includes image information and audio information; Construct a task scenario graph based on the image information, generate audio event tags based on the audio information, and identify key frames from the task scenario graph according to the audio event tags and generate corresponding key scenario text description information; Use a large language model to identify the multi-subtask completion information corresponding to the intelligent robot performing the task for the key scenario text description information and the corresponding audio event tags, perform fault information detection on each subtask completion information based on a preset subtask completion target, and when a fault information is detected in the subtask completion information, perform fault analysis based on the corresponding key scenario text description information to obtain a fault analysis result; If no fault information is detected in all subtask completion information, use a large language model to identify the execution result corresponding to the intelligent robot performing the task, perform fault information detection on the execution result, and when a fault information is detected in the execution result, obtain the task description and task plan of the intelligent robot performing the task and input them together with the execution result into the large language model for error identification to obtain an error identification result; Generate a corresponding fault correction task for the intelligent robot performing the task according to the fault analysis result or the error identification result for the intelligent robot to perform the fault correction task.
2. The fault identification and handling method according to claim 1, characterized in that The constructing a task scenario graph based on the image information includes: Detect and semantically segment the image information to obtain multiple object objects; Based on the depth information of the image information, convert the object objects into corresponding 3D semantic point clouds; Aggregate multiple 3D semantic point clouds to obtain an aggregated point cloud, and construct a task scenario graph according to the aggregated point cloud.
3. The fault identification and handling method according to claim 1, characterized in that, The generating audio event tags based on the audio information includes: Perform audio cutting on the audio information according to a preset volume threshold to obtain multiple segments of audio; Through a pre-trained audio-language model, calculate the text embedding vectors corresponding to each segment of the cut audio; Calculate the cosine similarity between the text embedding vector corresponding to each segment of audio and a preset multiple of audio event tags respectively, and identify the audio event tags matched by the text embedding vector corresponding to each segment of audio based on the cosine similarity.
4. The fault identification and handling method according to claim 1, characterized in that, The identifying key frames from the task scenario graph according to the audio event tags and generating corresponding key scenario text description information includes: Identify the start time point and end time point of the corresponding audio event according to the audio event tags; Extract the task scenario graph corresponding to the start time point and the end time point; Traverse the task scenario graphs at each time point in chronological order, and when the task scenario graph changes, extract the changed task scenario graph; Extract the task scenario graphs corresponding to the multiple subtask completion time points of the intelligent robot performing the task; Identify the extracted task scenario graphs as key frames, and generate key scenario text description information corresponding to each key frame through an artificial intelligence model.
5. The fault identification and processing method according to claim 1, characterized in that, Identifying the key scenario text description information and the corresponding audio event tags through a large language model to obtain multiple subtask completion information corresponding to the tasks executed by the intelligent robot, including: Obtaining the key scenario text description information and audio event tags corresponding to multiple subtask completion time points of the tasks executed by the intelligent robot; Performing text summarization on the key scenario text description information and audio event tags through a large language model to obtain multiple subtask completion information corresponding to the tasks executed by the intelligent robot.
6. The fault identification and handling method according to claim 1, wherein Generating a corresponding fault correction task for the tasks executed by the intelligent robot according to the fault analysis result or the error identification result, including: Analyzing the fault analysis result or the error identification result to obtain the uncompleted task metrics of the tasks executed by the intelligent robot; Obtaining the task prompt words corresponding to the uncompleted task metrics, and inputting the task prompt words into a preset language model to generate corresponding task execution actions; Mapping the task execution actions into a task scenario graph through a pre-trained sentence embedding model, and adjusting the task execution actions based on the task scenario graph; Generating the task steps of the fault correction task according to the adjusted task execution actions.
7. The fault identification and handling method according to claim 1, characterized in that, The fault identification and processing method further includes: Obtaining the task text description information of the tasks executed by the intelligent robot; Obtaining the parameter text description information of the intelligent robot; Performing a logical consistency judgment on the task text description information, the parameter text description information, and the key scenario text description information, and performing fault analysis and error identification based on the judgment result.
8. An artificial intelligence-based fault identification and processing device, characterized in that, Including: An information acquisition module, configured to acquire multimodal information of the tasks executed by the intelligent robot, where the multimodal information includes image information and audio information; An information processing module, configured to construct a task scenario graph based on the image information, generate audio event tags based on the audio information, identify key frames from the task scenario graph according to the audio event tags, and generate corresponding key scenario text description information; A fault analysis module, configured to identify the key scenario text description information and the corresponding audio event tags through a large language model to obtain multiple subtask completion information corresponding to the tasks executed by the intelligent robot, perform fault information detection on each subtask completion information based on a preset subtask completion target, and when fault information is detected in the subtask completion information, perform fault analysis based on the corresponding key scenario text description information to obtain a fault analysis result; An error identification module, configured to, if no fault information is detected in all subtask completion information, identify the execution result corresponding to the tasks executed by the intelligent robot through a large language model, perform fault information detection on the execution result, and when fault information is detected in the execution result, obtain the task description and task plan of the tasks executed by the intelligent robot, and input them together with the execution result into the large language model for error identification to obtain an error identification result; A fault correction module, configured to generate corresponding fault correction tasks for the intelligent robot to execute according to the fault analysis result or the error identification result, so that the intelligent robot executes the fault correction tasks.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the fault identification and processing method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the fault identification and processing method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Fault positioning method and device
CN110750377A
Steering engine fault processing method and device thereof as well as terminal equipment
CN112536817A
Defective component diagnostic device and method for robot
JP2018008327A
Machine fault detection based on a combination of sound capture and on spot feedback
US20170178311A1