Abnormal data detection and related methods and products
Patent Information
- Application Number
- CN202610620322.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-07
- Publication Date
- 2026-09-25
AI Technical Summary
在当前VLA模型训练实践中,训练数据来源往往非常复杂,需要解决如何识别异常数据的问题
[0016]应当理解,本部分所描述的内容并非旨在标识本公开的实施例的关键或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的说明书而变得容易理解。
Smart Images

Figure CN122818138A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of large models, intelligent agents, and deep learning, specifically to a method, apparatus, intelligent agent, device, medium, and product for anomaly detection, intelligent agent training, and action generation. Background Technology
[0002] Embodied Intelligence (EAI) refers to intelligent systems with physical entities (such as robots, autonomous vehicles, or robotic arms) that autonomously adapt, learn, and complete complex physical tasks in an open world through continuous interaction with their environment.
[0003] With the development of embodied intelligence, more and more robotic systems are being trained using Vision-Language-Action (VLA) models, enabling robots to perform complex tasks based on visual information and verbal instructions. In current VLA model training practices, the sources of training data are often very complex, necessitating the solution of identifying anomalous data. Summary of the Invention
[0004] This disclosure provides a method, apparatus, agent, device, medium, and product for anomaly data detection, agent training, and action generation.
[0005] According to one aspect of this disclosure, an abnormal data detection method is provided, comprising: acquiring training data, the training data including: command data, visual data, and control data; extracting actions from the command data to obtain command action semantics; generating behavioral data and visual action semantics based on the visual data; generating control action semantics based on the control data; and determining whether the training data is abnormal data based on the command data, the command action semantics, the behavioral data, the visual action semantics, and the control action semantics.
[0006] According to another aspect of this disclosure, an agent training method is provided, comprising: deleting abnormal data in training data to obtain processed training data; each set of training data includes: command data, visual data, and control data; training an agent using the processed training data; wherein the abnormal data is detected by the method described in any of the preceding claims.
[0007] According to another aspect of this disclosure, an action generation method is provided, comprising: acquiring command data and visual data; employing an intelligent agent to process the input command data and visual data to output a current action; wherein the intelligent agent is trained using the method described in any of the preceding claims.
[0008] According to another aspect of this disclosure, an abnormal data detection device is provided, comprising: an acquisition module for acquiring training data, the training data including command data, visual data, and control data; a first processing module for extracting actions from the command data to obtain command action semantics; a second processing module for generating behavior data and visual action semantics based on the visual data; a third processing module for generating control action semantics based on the control data; and a determination module for determining whether the training data is abnormal data based on the command data, the command action semantics, the behavior data, the video action semantics, and the control action semantics.
[0009] According to another aspect of this disclosure, an intelligent agent training apparatus is provided, comprising: a processing module for deleting abnormal data in training data to obtain processed training data; each set of training data includes: command data, visual data, and control data; and a training module for training an intelligent agent using the processed training data; wherein the abnormal data is detected using the method described in any of the preceding embodiments.
[0010] According to another aspect of this disclosure, an action generation apparatus is provided, comprising: an acquisition module for acquiring command data and visual data; and a generation module for employing an intelligent agent to process the input command data and visual data to output a current action; wherein the intelligent agent is trained using the method described in any of the preceding embodiments.
[0011] According to another aspect of this disclosure, an intelligent agent is provided, comprising: an input module for inputting command data and visual data; a processing module for generating a current action based on the command data and the visual data; and an output module for outputting the current action; wherein the intelligent agent is trained using the method described in any of the preceding embodiments.
[0012] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to said at least one processor; wherein the memory stores instructions executable by said at least one processor, said instructions being executed by said at least one processor to enable said at least one processor to perform the method as described in any of the foregoing aspects.
[0013] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method according to any of the preceding aspects.
[0014] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to any of the preceding aspects.
[0015] According to embodiments of this disclosure, it is possible to efficiently and accurately determine whether training data is abnormal.
[0016] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0017] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0018] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;
[0019] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;
[0020] Figure 3 This is a schematic diagram illustrating the process of obtaining command action semantics according to embodiments of this disclosure;
[0021] Figure 4 This is a schematic diagram illustrating the process of acquiring behavioral data and visual action semantics according to embodiments of this disclosure;
[0022] Figure 5 This is a schematic diagram of the process of obtaining control action semantics according to the embodiments of this disclosure;
[0023] Figure 6 This is a schematic diagram according to the third embodiment of the present disclosure;
[0024] Figure 7 This is a schematic diagram according to the fourth embodiment of the present disclosure;
[0025] Figure 8 This is a schematic diagram according to the fifth embodiment of the present disclosure;
[0026] Figure 9 This is a schematic diagram according to the sixth embodiment of the present disclosure;
[0027] Figure 10 This is a schematic diagram according to the seventh embodiment of the present disclosure;
[0028] Figure 11 This is a schematic diagram according to the eighth embodiment of the present disclosure;
[0029] Figure 12This is a schematic diagram of an electronic device used to implement the abnormal data detection method, intelligent agent training method, or action generation method of the embodiments of this disclosure. Detailed Implementation
[0030] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0031] In related technologies, abnormal data can be detected manually. However, in the case of VLA models, the amount of training data is large, usually in the millions or even tens of millions, making it almost impossible to detect abnormal data manually.
[0032] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure. This embodiment provides an abnormal data detection method. Figure 1 As shown, the method includes:
[0033] 101. Obtain training data, which includes command data, visual data, and control data.
[0034] 102. Extract the command data into action sequences to obtain command action semantics.
[0035] 103. Generate behavioral data and visual action semantics based on the visual data.
[0036] 104. Generate control action semantics based on the control data.
[0037] 105. Based on the command data, the command action semantics, the behavior data, the visual action semantics, and the control action semantics, determine whether the training data is abnormal data.
[0038] Training data can be obtained through methods such as pre-collection or labeling.
[0039] This embodiment can be applied to the VLA model scenario, where the training data includes command data, visual data, and control data.
[0040] Command data refers to language instructions, usually natural language text, such as "pick up the cup on the table".
[0041] Visual data refers to the visual data collected by an intelligent agent (such as a robot) when executing language commands. Specifically, it can include image data or video data.
[0042] Control data refers to the motor control data when an intelligent agent (such as a robot) executes language commands, such as joint angles, joint speeds, and / or gripper states (closed or open).
[0043] The three types of data mentioned above can be pre-collected sets of mutually matched data. For example, taking visual data as image data, a set of training data can be collected as follows: first command data, multiple image data matching the first command data, and multiple control data matching the first command data. Or,
[0044] The three types of data mentioned above can also be unmatched collected data. In this case, matching training data can be obtained through classification. For example, image data can be clustered to group image data corresponding to the same command data. Similarly, control data can be clustered to group control data corresponding to the same command data, thereby obtaining a set of image data and a set of control data corresponding to each command data.
[0045] Specifically, the following training data can be obtained: picking up a water glass and its corresponding visual and control data, opening a drawer and its corresponding visual and control data, and putting down an object and its corresponding visual and control data.
[0046] Then, for each set of training data (such as picking up a water cup and its corresponding visual and control data), it is checked whether it is abnormal data.
[0047] Command action semantics refers to the key text representing actions extracted from command data.
[0048] In other words, command data generally includes not only action information, but also other information such as environment, location and objects. Command action semantics is obtained after removing and abstracting non-action information.
[0049] For example, if the command data is "pick up the cup on the table", after action extraction, the command action semantics is "take something".
[0050] Behavioral data refers to the actions of an agent obtained based on visual data, typically in the form of natural language text. For example, visual data includes video, which can be used to perform behavioral analysis to obtain behavioral description text, and then behavioral data from the agent's perspective can be obtained based on this text.
[0051] For example, after analyzing the video, the resulting behavioral description text is "the robotic arm moves to the table and picks up the cup". This behavioral description text is then converted into behavioral data such as "picking up the cup on the table".
[0052] Visual action semantics refers to the key text representing actions extracted from behavioral data.
[0053] In other words, similar to command data, behavioral data includes not only action information, but also other information such as environment, location and objects. Behavioral action semantics is obtained after removing and abstracting non-action information.
[0054] For example, if the behavioral data is "picking up the cup on the table", after action extraction, the semantic meaning of the behavioral action is "taking something".
[0055] Control action semantics refers to the key text representing actions obtained from control data.
[0056] For example, if the control data includes consecutive state data in the following order: "robotic arm moves", "gripper closes", and "robotic arm rises", then the control action semantics can be "pick up the object".
[0057] Then, based on command data, command action semantics, behavior data, visual action semantics, and control action semantics, the training data is checked for anomalies.
[0058] Specifically, the data (or semantics) at multiple levels can be compared to see if they are consistent. If the comparison results at each level are consistent, the training data is determined to be normal data. Otherwise, if the comparison results at at least one level are inconsistent, the training data is determined to be abnormal data.
[0059] In this embodiment, the abnormality of training data is detected based on command data, command action semantics, behavior data, video action semantics, and control action semantics, eliminating the need for manual detection and improving the efficiency of abnormal data detection. By using the above-mentioned multiple data for detection, information from multiple dimensions can be compared, which can improve the accuracy of abnormal data detection, thereby efficiently and accurately determining whether the training data is abnormal.
[0060] In practical implementation, semantics can be obtained based on large models or rules.
[0061] In some embodiments, the step of extracting actions from the command data to obtain command action semantics includes:
[0062] A Large Language Model (LLM) is used to extract actions from the input command data to output the semantics of the command actions.
[0063] The system can preset a prompt template, fill the command data (such as "pick up the cup on the table") into the template to obtain the current prompt information, and input the current prompt information into the LLM. The template can also contain specific instructions, such as instructing the LLM to extract the action information. Based on this, the LLM can extract the action information from the command data according to the current prompt information to obtain the command action semantics.
[0064] For example, based on the command data "pick up the cup on the table" mentioned above, the semantic meaning of the command action is "take something".
[0065] In this embodiment, LLM is used to process command data to obtain command action semantics. The advantages of LLM, such as its precision, can be utilized to obtain command action semantics efficiently and accurately.
[0066] In some embodiments, generating behavioral data and visual action semantics based on the visual data includes:
[0067] Video is obtained based on the visual data;
[0068] Obtain the behavior description text corresponding to the video;
[0069] Based on the behavior description text, obtain behavior data;
[0070] Action extraction is performed on the behavioral data to obtain the visual action semantics.
[0071] The visual data can be video, in which case the video can be obtained directly. Or,
[0072] Visual data can also include multiple images. In this case, a video can be reconstructed based on multiple images. For example, the multiple images can be reconstructed into a video using a preset period.
[0073] The preset period is not limited to the conventional 24 frames per second. It can be used to assemble the video sequentially according to the time sequence contained in the images, depending on the actual situation.
[0074] For example, for a command like "pick up the water cup", the resulting video contains a series of images in the following order: the robotic arm moves → grasps → lifts.
[0075] In this way, by processing images based on a preset period, video can be obtained easily and efficiently.
[0076] After obtaining the video, a multimodal large model can be used to obtain behavioral description text.
[0077] That is, a prompt message can be formed by combining the video and preset instructions, and the prompt message can be input into the multimodal big data model. The multimodal big data model can then extract the behavioral description text from the video based on the instruction.
[0078] For example, the behavioral description text obtained from the above video is: "There is a table in a room with multiple objects on the table. A robotic arm moves to the table, grabs a cup, and then lifts the cup up."
[0079] In this way, the processing power of multimodal large models can be utilized to extract behavioral description text based on videos, thus efficiently obtaining behavioral description text.
[0080] After obtaining the behavioral description text, LLM can be used to extract the behavioral data from it.
[0081] The behavior description text contains a large amount of environmental description information, such as room structure, lighting, and object color. The LLM can be instructed to delete environmental description information (such as deleting room, lighting, table shape, etc.) and retain the action description while converting it to the robot's first-person perspective. Based on this, the obtained behavior data is, for example, "pick up the cup on the table".
[0082] After acquiring behavioral data, visual action semantics can be extracted using LLM. For example, behavioral data and preset instructions can be combined to form a prompt. The prompt is then input into the LLM, which extracts actions from the behavioral data based on the instruction, and the output is visual action semantics.
[0083] For example, after extracting the action from the behavioral data "picking up the cup on the table", the visual action semantics are "taking something".
[0084] In this way, by performing noise reduction on the behavior description text twice using LLM, one time to obtain behavior data based on the behavior description text, and the other time to obtain visual action semantics based on the behavior data, the visual action semantics can be accurately obtained.
[0085] In this embodiment, behavioral data and visual action semantics are obtained based on visual data, which can provide visual dimension data so that subsequent consistency comparison of multi-dimensional data can be performed to identify abnormal data.
[0086] In some embodiments, generating control action semantics based on the control data includes:
[0087] Based on the control data, control events are obtained;
[0088] Based on the control event, obtain the semantics of the control action.
[0089] The training data includes control data, such as joint angles, joint velocities, and gripper states.
[0090] Control events refer to action events performed by the robot. Specifically, control data can be converted into control events according to preset rules.
[0091] For example, the preset rules include the following mapping relationships: gripper closes → grip, gripper opens → release, robotic arm rises → lifts, robotic arm descends → lowers, robotic arm moves → moves. Based on this mapping relationship, assuming the control data is gripper closes, the control event is gripping.
[0092] In this way, control data can be efficiently converted into control events based on preset rules.
[0093] Control data is usually multiple and includes time information. After converting the control data into control events, the control events can be combined according to the time information to obtain an event sequence. For example, combining them according to the order of time from first to last results in the following event sequence: move → grab → lift. Then, the control action semantics can be obtained according to the preset correspondence between the event sequence and the action. For example, if the preset event sequence (move → grab → lift) corresponds to the action of picking up something, then the obtained control action semantics is "picking up something".
[0094] In this way, control events can be efficiently converted into control action semantics based on a preset correspondence.
[0095] In this embodiment, control data can be converted into control action semantics based on preset information, and control dimension data can be provided so that subsequent consistency comparison of multi-dimensional data can be performed to identify abnormal data.
[0096] After obtaining data from multiple dimensions, their consistency can be compared to identify whether the training data is abnormal.
[0097] Specifically, this may include:
[0098] Determine the first similarity between the command data and the behavior data;
[0099] Determine a second similarity between the video action semantics and the control action semantics;
[0100] Determine the third similarity between the command action semantics and the control action semantics;
[0101] If at least one of the first similarity, the second similarity, and the third similarity is less than a preset threshold, the training data is determined to be abnormal data.
[0102] If the first similarity, the second similarity, and the third similarity are all greater than a preset threshold, the training data is determined to be normal data.
[0103] That is, three levels of consistency comparison can be performed:
[0104] First layer: Consistency between command data and behavior data;
[0105] The second layer: consistency between visual action semantics and control action semantics;
[0106] The third layer: consistency between command action semantics and control action semantics.
[0107] After passing consistency checks at all three levels, the corresponding training data is determined to be normal data; otherwise, it is considered abnormal data.
[0108] At each level, consistency detection can be performed using semantic similarity of word vectors. Taking the first level as an example, command word vectors corresponding to command data and behavior word vectors corresponding to behavior data can be obtained. The similarity between command word vectors and behavior word vectors is calculated. If the similarity is greater than a preset threshold, the consistency detection is passed. The word vectors of each data point can be converted into word vectors using existing encoders or other methods.
[0109] In this way, obtaining similarity based on word vectors makes the calculation process more stable and faster, thus enabling the batch processing of more data.
[0110] In this embodiment, the comparison of multiple dimensions can be performed to determine whether the training data is abnormal based on the above similarity, thereby improving the accuracy and reliability of abnormal data detection.
[0111] In some embodiments, the method may further include:
[0112] In response to the fact that the training data is abnormal, the training data is reviewed to obtain the review results;
[0113] Perform the corresponding operation based on the review results.
[0114] Among them, for any set of training data (including command data, visual data, and control data), if it is found to be normal data after detection, it can be used as target training data to train the VLA model.
[0115] If the data is found to be abnormal after testing, it can be reviewed. For example, manual review can be used to obtain a normal or abnormal review result.
[0116] In some embodiments, performing the corresponding operation based on the review result includes:
[0117] In response to the verification result indicating that the training data is abnormal, the training data is deleted; or,
[0118] In response to the verification result being that the training data is normal data, the training data is treated as bad examples, and preset processing for bad examples is performed.
[0119] For example, if the training data is still abnormal after review, the training data will be deleted to ensure the quality of the training data.
[0120] In this embodiment, data reliability can be improved by reviewing abnormal training data.
[0121] If the training data is found to be normal after review, it can be considered a bad case. The bad case can then be analyzed and the corresponding steps optimized.
[0122] Specifically, the causes of bad cases can be analyzed periodically, including at least one of the following: inaccurate action description extraction; video comprehension errors; text denoising and action extraction errors; and unreasonable similarity thresholds. Afterward, the system can be optimized based on the analyzed causes to continuously improve its accuracy. For example, for the first three causes, the prompt template can be adjusted; for the last cause, the similarity threshold can be adjusted.
[0123] In this embodiment, deleting or pre-processing the training data based on the verification results can improve data reliability and enhance system performance.
[0124] In conjunction with the above, the following embodiments are also provided.
[0125] Figure 2 This is a schematic diagram based on the second embodiment of the present disclosure, which provides an abnormal data detection method. For example... Figure 2 As shown, the method includes:
[0126] 201. Obtain training data, which includes command data, visual data, and control data.
[0127] 202. Obtain multiple sets of data to be matched based on the training data.
[0128] Specifically, the multiple sets of data to be matched can include three sets, namely:
[0129] The first set of data to be matched includes: command data and behavior data;
[0130] The second set of data to be matched includes: visual action semantics and control action semantics;
[0131] The third set of data to be matched includes: command action semantics and control action semantics.
[0132] Command data can be obtained from training data, while command action semantics are obtained from command data.
[0133] Behavioral data is obtained from visual data in the training data, visual action semantics is obtained from behavioral data, and control action semantics is obtained from control data in the training data.
[0134] Figure 3 This is a schematic diagram of the process of obtaining command action semantics according to the embodiments of this disclosure.
[0135] like Figure 3 As shown, command data can be input into the LLM, and prompts can be used to instruct the LLM to extract actions from the command data. The LLM then processes the command data and outputs the semantics of the command actions.
[0136] Figure 4 This is a schematic diagram illustrating the process of acquiring behavioral data and visual action semantics according to embodiments of this disclosure.
[0137] like Figure 4 As shown, taking visual data comprising multiple images as an example, multiple images can be reconstructed into a video. The video is input into a multimodal large model, and prompts are used to instruct the multimodal large model to generate behavioral description text. The multimodal large model then processes the video and outputs the behavioral description text. Next, the behavioral description text is input into an LLM (Law Management Model), and prompts are used to instruct the LLM to extract behaviors. The LLM processes the behavioral description text and outputs behavioral data. Conversely, behavioral data is input into an LLM, and prompts are used to instruct the LLM to extract actions. The LLM processes the behavioral data and outputs visual action semantics.
[0138] Figure 5 This is a schematic diagram of the process of obtaining control action semantics according to the embodiments of this disclosure.
[0139] like Figure 5 As shown, control data can be converted into control events based on preset rules, then the control events can be grouped into an event sequence based on time information, and finally the event sequence can be converted into control action semantics according to preset correspondence.
[0140] 203. Determine if the data to be matched in each group is consistent. If yes, proceed to 204; otherwise, proceed to 205.
[0141] For each group of data to be matched, the similarity can be calculated using the corresponding word vectors. If the similarity is greater than a preset threshold, it indicates that the two are consistent; otherwise, they are inconsistent.
[0142] 204. Confirm that the training data is normal data.
[0143] 205. Determine that the training data is outlier.
[0144] 206. Review the abnormal data and obtain the review results.
[0145] For example, manual verification can be used to obtain normal or abnormal verification results.
[0146] 207. Perform the corresponding operation based on the review results.
[0147] For example, if the review result is abnormal, it indicates that the final result of the training data is abnormal, and the training data is then deleted. Alternatively, if the review result is normal, it can be marked as a bad example, analyzed, and the system adjusted based on the analysis results, such as adjusting the prompt information or adjusting the preset threshold when calculating similarity.
[0148] Figure 6 This is a schematic diagram based on a third embodiment of the present disclosure, which provides a method for training an intelligent agent. For example... Figure 6 As shown, the method includes:
[0149] 601. Remove outlier data from the training data to obtain processed training data; each set of training data includes: command data, visual data, and control data.
[0150] 602. Use the processed training data to train the intelligent agent.
[0151] The abnormal data is detected using the detection method described in any of the above embodiments.
[0152] Specifically, the intelligent agent can adopt the VLA model, and each set of training data can include command data, visual data, and control data.
[0153] In this embodiment, by performing anomaly detection on the training data and deleting abnormal data, the reliability of the target training data can be improved, thereby improving the reliability of the intelligent agent.
[0154] Figure 7 This is a schematic diagram based on the fourth embodiment of the present disclosure, which provides an action generation method. For example... Figure 7 As shown, the method includes:
[0155] 701. Obtain command data and visual data.
[0156] 702. An intelligent agent is used to process the input command data and visual data to output the current action.
[0157] The intelligent agent is trained using the training method of any of the above embodiments.
[0158] In this embodiment, by employing the aforementioned intelligent agent, the accuracy of the current action can be improved.
[0159] Figure 8 The diagram is based on the fifth embodiment of the present disclosure. This embodiment provides an abnormal data detection device. The device 800 includes: an acquisition module 801, a first processing module 802, a second processing module 803, a third processing module 804, and a determination module 805.
[0160] The acquisition module 801 is used to acquire training data, which includes command data, visual data, and control data; the first processing module 802 is used to extract actions from the command data to obtain command action semantics; the second processing module 803 is used to generate behavior data and visual action semantics based on the visual data; the third processing module 804 is used to generate control action semantics based on the control data; and the determination module 805 is used to determine whether the training data is abnormal data based on the command data, the command action semantics, the behavior data, the video action semantics, and the control action semantics.
[0161] In some embodiments, the first processing module 802 is further configured to:
[0162] A large language model is used to extract actions from the input command data in order to output the semantics of the command actions.
[0163] In this embodiment, LLM is used to process command data to obtain command action semantics. The advantages of LLM, such as its precision, can be utilized to obtain command action semantics efficiently and accurately.
[0164] In some embodiments, the second processing module 803 is further configured to:
[0165] Video is obtained based on the visual data;
[0166] Obtain the behavior description text corresponding to the video;
[0167] Based on the behavior description text, obtain behavior data;
[0168] Action extraction is performed on the behavioral data to obtain the visual action semantics.
[0169] In this embodiment, behavioral data and visual action semantics are obtained based on visual data, which can provide visual dimension data so that subsequent consistency comparison of multi-dimensional data can be performed to identify abnormal data.
[0170] In some embodiments, the visual data includes: multiple images; the second processing module 803 is further configured to:
[0171] The multiple images are reconstructed into a video using a preset period.
[0172] In this way, by processing images based on a preset period, video can be obtained easily and efficiently.
[0173] In some embodiments, the second processing module 803 is further configured to:
[0174] A multimodal large model is used to process the input video to output the behavior description text.
[0175] In this way, the processing power of multimodal large models can be utilized to extract behavioral description text based on videos, thus efficiently obtaining behavioral description text.
[0176] In some embodiments, the second processing module 803 is further configured to:
[0177] A large language model is used to extract actions from the input behavioral data in order to output the visual action semantics.
[0178] In this way, by performing noise reduction on the behavior description text twice using LLM, one time to obtain behavior data based on the behavior description text, and the other time to obtain visual action semantics based on the behavior data, the visual action semantics can be accurately obtained.
[0179] In some embodiments, the third processing module 804 is further configured to:
[0180] Based on the control data, control events are obtained;
[0181] Based on the control event, obtain the semantics of the control action.
[0182] In this embodiment, control data can be converted into control action semantics based on preset information, and control dimension data can be provided so that subsequent consistency comparison of multi-dimensional data can be performed to identify abnormal data.
[0183] In some embodiments, the third processing module 804 is further configured to:
[0184] The control data is mapped to control events using preset rules.
[0185] In this way, control data can be efficiently converted into control events based on preset rules.
[0186] In some embodiments, the third processing module 804 is further configured to:
[0187] The control events are arranged into an event sequence;
[0188] Based on a preset correspondence, the event sequence is converted into control action semantics.
[0189] In this way, control events can be efficiently converted into control action semantics based on a preset correspondence.
[0190] In some embodiments, the determining module 805 is further configured to:
[0191] Determine the first similarity between the command data and the behavior data;
[0192] Determine a second similarity between the video action semantics and the control action semantics;
[0193] Determine the third similarity between the command action semantics and the control action semantics;
[0194] If at least one of the first similarity, the second similarity, and the third similarity is less than a preset threshold, the training data is determined to be abnormal data.
[0195] If the first similarity, the second similarity, and the third similarity are all greater than a preset threshold, the training data is determined to be normal data.
[0196] In this embodiment, the comparison of multiple dimensions can be performed to determine whether the training data is abnormal based on the above similarity, thereby improving the accuracy and reliability of abnormal data detection.
[0197] In some embodiments, the first similarity, the second similarity, and the third similarity are all determined based on word vectors of the corresponding data.
[0198] In this way, obtaining similarity based on word vectors makes the calculation process more stable and faster, thus enabling the batch processing of more data.
[0199] In some embodiments, the device 800 may further include:
[0200] The verification module is used to verify the training data in response to the fact that the training data is abnormal data, so as to obtain the verification result; and to perform corresponding operations based on the verification result.
[0201] In this embodiment, reliability can be improved by reviewing abnormal training data.
[0202] In some embodiments, the review module is further used for:
[0203] In response to the verification result indicating that the training data is abnormal, the training data is deleted; or,
[0204] In response to the verification result being that the training data is normal data, the training data is treated as bad examples, and preset processing for bad examples is performed.
[0205] In this embodiment, deleting or pre-processing the training data based on the verification results can improve data reliability and enhance system performance.
[0206] Figure 9 This is a schematic diagram according to the sixth embodiment of the present disclosure. This embodiment provides an intelligent agent training device 900, which includes a processing module 901 and a training module 902.
[0207] The processing module 901 is used to delete abnormal data in the initial training data to obtain target training data; the training module 902 is used to train the agent using the target training data; wherein, the abnormal data is detected by the detection method described in any of the above embodiments.
[0208] In this embodiment, by performing anomaly detection on the training data and deleting abnormal data, the reliability of the target training data can be improved, thereby improving the reliability of the intelligent agent.
[0209] Figure 10 This is a schematic diagram based on the seventh embodiment of the present disclosure. This embodiment provides an action generation device 1000, which includes: an acquisition module 1001 and a generation module 1002.
[0210] The acquisition module 1001 is used to acquire command data and visual data; the generation module 1002 is used to employ an intelligent agent to process the input command data and visual data to output the current action; wherein, the intelligent agent is trained using the training method of any of the above embodiments.
[0211] In this embodiment, by employing the aforementioned intelligent agent, the accuracy of the current action can be improved.
[0212] Figure 11 This is a schematic diagram based on the eighth embodiment of the present disclosure. This embodiment provides an intelligent agent 1100, which includes an input module 1101, a processing module 1102, and an output module 1103.
[0213] The input module 1101 is used to input command data and visual data; the processing module 1102 is used to generate the current action based on the command data and the visual data; the output module 1103 is used to output the current action; wherein, the intelligent agent is trained using the training method of any of the above embodiments.
[0214] In this embodiment, by employing the aforementioned intelligent agent, the accuracy of the current action can be improved.
[0215] It is understood that the same or similar content in different embodiments of this disclosure can be referred to each other.
[0216] It is understood that the terms "first" and "second" in the embodiments of this disclosure are only used for distinction and do not indicate the degree of importance or the order of events.
[0217] It is understandable that, unless otherwise specified, the order of steps in the process indicates that the temporal relationship between these steps is not limited.
[0218] The technical solutions disclosed herein involve the collection, storage, use, processing, transmission, provision, and disclosure of any type of information, such as user personal information, in accordance with relevant laws and regulations and without violating public order and good morals.
[0219] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0220] Figure 12 A schematic block diagram of an example electronic device 1200 that can be used to implement embodiments of the present disclosure is shown. The electronic device 1200 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0221] like Figure 12 As shown, the electronic device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. The RAM 1203 may also store various programs and data required for the operation of the electronic device 1200. The computing unit 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0222] Multiple components in electronic device 1200 are connected to I / O interface 1205, including: input unit 1206, such as keyboard, mouse, etc.; output unit 1207, such as various types of displays, speakers, etc.; storage unit 1208, such as disk, optical disk, etc.; and communication unit 1209, such as network card, modem, wireless transceiver, etc. Communication unit 1209 allows electronic device 1200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0223] The computing unit 1201 can be a variety of general-purpose and / or proprietary processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various proprietary artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as anomaly data detection methods, agent training methods, or action generation methods. For example, in some embodiments, the anomaly data detection method, agent training method, or action generation method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by the computing unit 1201, one or more steps of the anomaly data detection method, agent training method, or action generation method described above can be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured by any other suitable means (e.g., by means of firmware) to perform anomaly data detection methods, agent training methods, or action generation methods.
[0224] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), proprietary integrated circuits (ASICs), proprietary standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a proprietary or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0225] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable task processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0226] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0227] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0228] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0229] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and limited business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0230] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0231] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An abnormal data detection method, comprising: Acquire training data, which includes command data, visual data, and control data; The command data is subjected to action extraction to obtain command action semantics; Generate behavioral data and visual action semantics based on the visual data; Generate control action semantics based on the control data; Based on the command data, the command action semantics, the behavior data, the visual action semantics, and the control action semantics, determine whether the training data is abnormal data.
2. The method according to claim 1, wherein, The step of extracting actions from the command data to obtain command action semantics includes: A large language model is used to extract actions from the input command data in order to output the semantics of the command actions.
3. The method according to claim 1, wherein, The generation of behavioral data and visual action semantics based on the visual data includes: Video is obtained based on the visual data; Obtain the behavior description text corresponding to the video; Based on the behavior description text, obtain behavior data; Action extraction is performed on the behavioral data to obtain the visual action semantics.
4. The method according to claim 3, wherein, The visual data includes: multiple images; The process of acquiring video based on the visual data includes: The multiple images are reconstructed into a video using a preset period.
5. The method according to claim 3, wherein, The step of obtaining the behavior description text corresponding to the video includes: A multimodal large model is used to process the input video to output the behavior description text.
6. The method according to claim 3, wherein, The step of extracting actions from the behavioral data to obtain the visual action semantics includes: A large language model is used to extract actions from the input behavioral data in order to output the visual action semantics.
7. The method according to claim 1, wherein, The generation of control action semantics based on the control data includes: Based on the control data, control events are obtained; Based on the control event, obtain the semantics of the control action.
8. The method according to claim 7, wherein, The acquisition of control events based on the control data includes: The control data is mapped to control events using preset rules.
9. The method according to claim 7, wherein, The step of obtaining control action semantics based on the control event includes: The control events are arranged into an event sequence; Based on a preset correspondence, the event sequence is converted into control action semantics.
10. The method according to any one of claims 1-9, wherein, The step of determining whether the training data is anomalous based on the command data, the command action semantics, the behavior data, the video action semantics, and the control action semantics includes: Determine the first similarity between the command data and the behavior data; Determine a second similarity between the video action semantics and the control action semantics; Determine the third similarity between the command action semantics and the control action semantics; If at least one of the first similarity, the second similarity, and the third similarity is less than a preset threshold, the training data is determined to be abnormal data. If the first similarity, the second similarity, and the third similarity are all greater than a preset threshold, the training data is determined to be normal data.
11. The method according to claim 10, wherein, The first similarity, the second similarity, and the third similarity are all determined based on the word vectors of the corresponding data.
12. The method according to any one of claims 1-9, further comprising: In response to the fact that the training data is abnormal, the training data is reviewed to obtain the review results; Perform the corresponding operation based on the review results.
13. The method according to claim 12, wherein, The step of performing corresponding operations based on the review results includes: In response to the verification result indicating that the training data is abnormal, the training data is deleted; or, In response to the verification result being that the training data is normal data, the training data is treated as bad examples, and preset processing for bad examples is performed.
14. A method for training an agent, comprising: Remove outlier data from the training data to obtain processed training data; Each set of training data includes: command data, visual data, and control data; The processed training data is used to train the agent; The abnormal data is detected using the method described in any one of claims 1-13.
15. An action generation method, comprising: Acquire command data and visual data; An intelligent agent is used to process the input command data and visual data to output the current action; The agent is trained using the method described in claim 14.
16. An abnormal data detection device, comprising: The acquisition module is used to acquire training data, which includes command data, visual data, and control data. The first processing module is used to extract actions from the command data to obtain command action semantics; The second processing module is used to generate behavioral data and visual action semantics based on the visual data; The third processing module is used to generate control action semantics based on the control data; The determination module is used to determine whether the training data is abnormal data based on the command data, the command action semantics, the behavior data, the video action semantics, and the control action semantics.
17. An agent training device, comprising: The processing module is used to remove outlier data from the initial training data in order to obtain the target training data. The training module is used to train the agent using the target training data; The abnormal data is detected using the method described in any one of claims 1-13.
18. An action generation device, comprising: The acquisition module is used to acquire command data and visual data; A generation module is used to process the input command data and visual data using an intelligent agent to output the current action; The agent is trained using the method described in claim 14.
19. An intelligent agent, comprising: The input module is used to input command data and visual data; The processing module is used to generate the current action based on the command data and the visual data; The output module is used to output the current action; The agent is trained using the method described in claim 14.
20. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-15.
21. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-15.
22. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-15.