Human-machine collaboration dataset construction and anomaly detection method based on fusion of multi-modal large model

By acquiring multimodal data in virtual human-computer collaboration scenarios and performing timestamp alignment and model fine-tuning, the low efficiency and security issues of human-computer collaboration dataset construction in existing technologies are solved, achieving efficient multimodal dataset construction and virtual human-computer collaboration anomaly detection.

CN119538161BActive Publication Date: 2025-10-24ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411802998.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-10-24
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

Existing human-computer collaborative datasets suffer from time-consuming and labor-intensive hardware debugging and low efficiency during the collection process. They also require manual labeling, pose a high risk of potential collisions, and the collected images lack global scene information, making it difficult to construct high-quality multimodal datasets.

Method used

By acquiring collision points, multi-angle human-machine scene images, and joint posture data of virtual digital humans and virtual robotic arms in virtual human-machine collaboration scenarios, and using a multimodal large model for timestamp alignment and transformation, a visual-language instruction dataset is generated. The multimodal large model is then fine-tuned to achieve anomaly detection in virtual human-machine collaboration.

Benefits of technology

The system efficiently constructs high-quality multimodal datasets, solving the problems of low efficiency in hardware debugging and manual labeling, avoiding potential collision risks, collecting images from multiple angles, and realizing anomaly detection in virtual human-computer collaboration scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119538161B_ABST
    Figure CN119538161B_ABST
Patent Text Reader

Abstract

The application discloses a kind of man-machine cooperation data set construction and abnormality detection method of fusion multimodal big model, it is related to man-machine cooperation technical field, this method includes: based on man-machine action label control virtual digital person and virtual mechanical arm in virtual man-machine cooperation scene carries out virtual man-machine cooperation, in the process of carrying out virtual man-machine cooperation, obtain human collision part, multi-angle man-machine scene image and joint posture data, and further carry out timestamp alignment, constitute multimodal data set, convert multimodal data set into visual-linguistic instruction data set using multimodal big model, visual-linguistic instruction data set is used to fine-tune multimodal big model, and realize man-machine cooperation abnormality detection in virtual man-machine cooperation scene using fine-tuned multimodal big model.The application can efficiently construct high-quality multimodal data set, and can fine-tune training to multimodal big model, realize man-machine cooperation abnormality detection in virtual man-machine cooperation scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of human-computer collaboration, in particular to a human-computer collaboration data set construction and anomaly detection method fusing a multi-modal large model. BACKGROUND

[0002] With the continuous maturity of robot technology and the rapid development of artificial intelligence, human-computer collaboration technology, as a cutting-edge industrial production paradigm, is gradually changing the face of traditional manufacturing. In the current field of human-computer collaboration, scene anomaly detection (i.e. anomaly detection in the human-computer collaboration scene) technology is not only the key to ensuring production safety and efficiency, but also the core driving force for the continuous progress of the human-computer collaboration field. Traditional scene anomaly detection technology is based on pure visual models and often relies on single image information, making it difficult to fully capture multi-modal information in complex work environments. Multi-modal large models (also known as multi-modal large language models) can be fine-tuned for specific scene tasks through high-quality human-computer collaboration data sets, bringing a new solution to scene anomaly detection.

[0003] However, the collection of human-computer collaboration scene data in current human-computer collaboration data sets faces many challenges. First, the debugging of complex hardware devices is time-consuming and labor-intensive, resulting in long data collection cycles and low efficiency, and the need for manual data labeling is inefficient. Second, the potential collision risk in human-computer collaboration scenes not only causes data collection to be interrupted or equipment to be damaged, but also threatens the personal safety of operators, further affecting the integrity and usability of the data set. Finally, due to space limitations and perspective problems in actual scenes, single-angle images collected lack global scene information. The above problems collectively limit the quality and efficiency of human-computer collaboration data sets.

[0004] The concept of multi-modal in the human-computer collaboration scene is not limited to traditional sensory categories such as vision and hearing, but is extended to a more extensive information dimension, including but not limited to human body action posture, mechanical arm joint data, work scene images, and other types. Integrating the above different dimensional modal information generates a high-quality multi-modal data set based on the human-computer collaboration scene, which can meet the specific downstream task fine-tuning training of existing multi-modal large models, not only for anomaly detection in multiple visual tasks, but also with the significant generalization and emergence ability of multi-modal large models, achieving zero-shot understanding and reasoning for untrained complex multi-element scenes, bringing more dimensions and deeper analysis for anomaly detection in human-computer collaboration scenes, and providing more intelligent technical support for the optimization and upgrading of human-computer collaboration technology.

[0005] Therefore, there is an urgent need in the field for an efficient construction method of high-quality multi-modal data sets. SUMMARY

[0006] The application aims to provide a human-machine collaboration dataset construction and anomaly detection method fusing a multi-modal large model, which can efficiently construct a high-quality multi-modal dataset and further fine-tune a multi-modal large model to realize human-machine collaboration anomaly detection in a virtual human-machine collaboration scene.

[0007] To achieve the above-mentioned purpose, the application provides the following solutions.

[0008] In a first aspect, the application provides a human-machine collaboration dataset construction and anomaly detection method fusing a multi-modal large model, which comprises:

[0009] Controlling a virtual digital human and a virtual mechanical arm to perform virtual human-machine collaboration in a virtual human-machine collaboration scene based on human-machine action labels;

[0010] In the process of virtual human-machine collaboration, the human-machine collaboration dataset construction and anomaly detection method fuses a multi-modal large model, which comprises:

[0011] Timestamp aligning the human collision part, the multi-angle human scene image, and the joint posture data, and grouping the timestamp-aligned human collision part, multi-angle human scene image, and joint posture data to form human-machine collaboration virtual scene time series data corresponding to the human-machine action label;

[0012] Grouping each human-machine action label and the human-machine collaboration virtual scene time series data corresponding to the human-machine action label to form a multi-modal dataset;

[0013] Converting the multi-modal dataset into a visual-linguistic instruction dataset using a multi-modal large model;

[0014] Fine-tuning a multi-modal large model using the visual-linguistic instruction dataset to obtain a fine-tuned multi-modal large model; and realizing human-machine collaboration anomaly detection in a virtual human-machine collaboration scene using the fine-tuned multi-modal large model.

[0015] Optionally, the human-machine action label comprises a tester human posture label, a mechanical arm working state label, and a human-machine spatial distance label; the tester human posture label comprises standing, squatting, sitting, and lying down; the mechanical arm working state label comprises being stationary, moving but not grabbing, and grabbing; and the human-machine spatial distance label comprises being safe and normal, warning and reminding, and being dangerous and braking.

[0016] Optionally, controlling a virtual digital human and a virtual mechanical arm to perform virtual human-machine collaboration in a virtual human-machine collaboration scene based on human-machine action labels specifically comprises:

[0017] acquire human posture data of a real tester and joint motion data of a real robot arm; the human posture data and the joint motion data are data generated when the real tester and the real robot arm move based on a human-robot action label;

[0018] drive a virtual digital human to move based on the human posture data, drive a virtual robot arm to move based on the joint motion data, so that the virtual digital human and the virtual robot arm are synchronized with the real tester and the real robot arm in action, and the virtual digital human and the virtual robot arm perform virtual human-robot collaboration in a virtual human-robot collaboration scene.

[0019] Optionally, the multi-angle human scene image is an image obtained after video frame extraction and fixed frequency collection are performed on a shooting video, and the shooting video is a video obtained by shooting through a set of virtual cameras arranged in the virtual human-robot collaboration scene; the set of virtual cameras includes three virtual cameras, and the shooting angles of the three virtual cameras are orthogonal to each other.

[0020] The multi-angle human scene image includes a human scene image of the virtual human-robot collaboration scene in a front view, a human scene image of the virtual human-robot collaboration scene in a side view, and a human scene image of the virtual human-robot collaboration scene in a top view.

[0021] The method further includes naming the multi-angle human scene image by taking the human-robot action label as the name of the multi-angle human scene image.

[0022] Optionally, the human collision part, the multi-angle human scene image, and the joint posture data are time-stamped aligned, specifically including:

[0023] determining the maximum and minimum values of the time stamp of the human collision part, the time stamp of the multi-angle human scene image, and the time stamp of the joint posture data;

[0024] determining a time interval based on the maximum value and the minimum value, and time-stamping aligning the human collision part, the multi-angle human scene image, and the joint posture data by a traversal and interpolation method.

[0025] Optionally, the multi-modal data set is converted into a visual-linguistic instruction data set by using a multi-modal large model, specifically including:

[0026] automatically annotating the aligned multi-angle human scene image in the multi-modal data set by using a multi-modal large model to obtain an image title;

[0027] generate a visual-linguistic instruction dataset by using a multimodal large model, taking the multimodal dataset and the image title as input.

[0028] Optionally, the multimodal large model used to obtain the image title is CogVLM2, GLM4V, Qwen-VL-Chat or MiniCPM-V-2.5; the multimodal large model used to generate the visual-linguistic instruction dataset is ChatGPT; and the multimodal large model used to be fine-tuned for anomaly detection is CogVLM2, GLM4V, Qwen-VL-Chat or MiniCPM-V-2.5.

[0029] Optionally, the multimodal large model is used to generate a visual-linguistic instruction dataset, taking the multimodal dataset and the image title as input, specifically including:

[0030] A plurality of dialogue templates are constructed; the dialogue templates are used to guide the multimodal large model to generate instruction data with specific modes and styles, and the content of the dialogue templates includes inquiring about image content and reasoning scenarios;

[0031] A visual-linguistic instruction dataset is generated by using a multimodal large model, taking the multimodal dataset, the image title and a plurality of dialogue templates as input; the visual-linguistic instruction dataset includes a plurality of types of text data, and the types of text data include detailed description type data, long dialogue type data and complex reasoning type data.

[0032] Optionally, the multimodal large model is fine-tuned using the visual-linguistic instruction dataset to obtain a fine-tuned multimodal large model, specifically including:

[0033] The multimodal large model is fine-tuned using a fine-tuning strategy based on training loss optimization, taking the visual-linguistic instruction dataset as input, to obtain a fine-tuned multimodal large model; wherein the fine-tuning strategy based on training loss optimization includes an AdamW optimizer, a cosine learning rate scheduler and a mixed precision strategy.

[0034] Optionally, the fine-tuned multimodal large model is used to implement human-robot collaboration anomaly detection in a virtual human-robot collaboration scenario, specifically including:

[0035] The fine-tuned multimodal large model is used to perform anomaly detection on the current acquired multi-angle human scene image to obtain a tester human body posture, a mechanical arm working state, human body pixel coordinates and mechanical arm pixel coordinates;

[0036] The human body pixel coordinates are subjected to coordinate transformation to obtain human body spatial coordinates;

[0037] The mechanical arm pixel coordinates are subjected to coordinate transformation to obtain mechanical arm spatial coordinates;

[0038] calculating a Euclidean distance of the human body space coordinate and the mechanical arm space coordinate to obtain a human-machine space distance;

[0039] determining a human-machine safety degree based on the human-machine space distance and a minimum safety cooperation distance;

[0040] The method further includes:

[0041] When the human-machine space distance is less than or equal to the minimum safety cooperation distance, the human-machine safety degree is dangerous braking.

[0042] When the human-machine space distance is greater than the minimum safety cooperation distance and less than or equal to a preset distance, the human-machine safety degree is warning reminding.

[0043] When the human-machine space distance is greater than the preset distance, the human-machine safety degree is safe normal.

[0044] The method further includes:

[0045] obtaining a plurality of sampling multi-angle human field images in a sampling time period; the sampling time period is (T0, T1+T), T0 is a time point at which the human body intrudes, T1 is a time point at which the mechanical arm is triggered to retreat, and T is a time used by the mechanical arm from starting to retreat to stopping after retreating to a safe distance.

[0046] For each of the sampling multi-angle human field images, an abnormity is detected by using the sampling multi-angle human field image as an input and a fine-tuned multi-modal large model to obtain a sampling human body pixel coordinate and a sampling mechanical arm pixel coordinate.

[0047] For any two adjacent sampling multi-angle human field images, a sampling human body moving speed is obtained by calculating a ratio of the sampling human body pixel coordinate of the two adjacent sampling multi-angle human field images to a time interval of the two adjacent sampling multi-angle human field images, and a sampling mechanical arm moving speed is obtained by calculating a ratio of the sampling mechanical arm pixel coordinate of the two adjacent sampling multi-angle human field images to the time interval of the two adjacent sampling multi-angle human field images.

[0048] The minimum safety cooperation distance is calculated according to the sampling human body moving speed and the sampling mechanical arm moving speed.

[0049] The calculation formula of the minimum safety cooperation distance is:

[0050] Sp = Sh + Sr + Ss + C.

[0051] Wherein, Sp is the minimum safety cooperation distance; Sh is the related distance of the human body approaching speed, Vh(t) is the moving speed of the human body at t time, determined based on sampling with the human body moving speed; Sr is the related distance of the normal working speed of the mechanical arm, Vr(t) is the moving speed of the mechanical arm at t time along the direction of approaching the human body, determined based on sampling with the mechanical arm moving speed; Ss is the stopping path distance of the mechanical arm, Vs(t) is the moving speed of the mechanical arm at t time along the stopping path, determined based on sampling with the mechanical arm moving speed; C is the distance of the human body intrusion.

[0052] According to the specific embodiments provided in the application, the following technical effects are disclosed:

[0053] The application provides a human-machine collaboration dataset construction and anomaly detection method based on a multi-modal large model. Virtual digital people and virtual mechanical arms are controlled based on human-machine action labels to perform virtual human-machine collaboration in a virtual human-machine collaboration scene. In the process of virtual human-machine collaboration, the human body collision part when the virtual digital people collide with the physical assets in the virtual human-machine collaboration scene, multi-angle human scene images of the virtual human-machine collaboration scene, and joint posture data of the virtual mechanical arm are obtained. The human body collision part, the multi-angle human scene images, and the joint posture data are time-stamped aligned. The human body collision part, the multi-angle human scene images, and the joint posture data after time-stamping alignment form human-machine collaboration virtual scene time series data corresponding to the human-machine action labels. Each human-machine action label and the human-machine collaboration virtual scene time series data corresponding to the human-machine action label form a multi-modal dataset. The application completes virtual human-machine collaboration and data collection by constructing a digital twin system, without the need for actual human-machine collaboration and data collection processes. The application can solve the problems of time-consuming and labor-consuming hardware device debugging, the need for manual data labeling, and the existence of potential collision risks. At the same time, multi-angle human scene images are collected, solving the problem of only being able to collect single-angle images in actual scenes, thereby efficiently generating high-quality multi-modal datasets. After obtaining the multi-modal dataset, the multi-modal large model is used to convert the multi-modal dataset into a visual-linguistic instruction dataset. The multi-modal large model is fine-tuned using the visual-linguistic instruction dataset, and a fine-tuned multi-modal large model is obtained. The fine-tuned multi-modal large model is used to realize human-machine collaboration anomaly detection in a virtual human-machine collaboration scene. The application can efficiently construct high-quality multi-modal datasets and fine-tune the multi-modal large model, realizing human-machine collaboration anomaly detection in a virtual human-machine collaboration scene. BRIEF DESCRIPTION OF DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0055] Figure 1 An application environment diagram of a fusion multi-modal large model human-computer collaboration dataset construction and anomaly detection method provided for Embodiment 1 of the present application.

[0056] Figure 2 A flowchart of a fusion multi-modal large model human-computer collaboration dataset construction and anomaly detection method provided for Embodiment 1 of the present application.

[0057] Figure 3 A principle diagram of a fusion multi-modal large model human-computer collaboration dataset construction and anomaly detection method provided for Embodiment 1 of the present application.

[0058] Figure 4 An illustration of a human-computer action label category provided for Embodiment 1 of the present application.

[0059] Figure 5 An illustration of a front view perspective diagram of a virtual human-computer collaboration scene captured at a certain moment provided for Embodiment 1 of the present application.

[0060] Figure 6 An illustration of a side view perspective diagram of a virtual human-computer collaboration scene captured at a certain moment provided for Embodiment 1 of the present application.

[0061] Figure 7 An illustration of a top view perspective diagram of a virtual human-computer collaboration scene captured at a certain moment provided for Embodiment 1 of the present application.

[0062] Figure 8 A flowchart of a timestamp alignment algorithm provided for Embodiment 1 of the present application.

[0063] Figure 9 An illustration of a visual-linguistic instruction dataset subdivision type provided for Embodiment 1 of the present application.

[0064] Figure 10 A structure diagram of a computer device provided for Embodiment 2 of the present application. DETAILED DESCRIPTION

[0065] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0066] Example 1

[0067] The method for constructing a human-machine collaborative dataset and detecting anomalies based on a multimodal large-scale model provided in the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send a request for the construction of the dataset to be processed and anomaly detection (for building a multimodal dataset and performing anomaly detection) to the server. After the server receives the request for the construction of the dataset to be processed and the anomaly detection request, for the request for the construction of the dataset to be processed and the anomaly detection request, the server controls the virtual digital human and the virtual robotic arm to perform virtual human-machine collaboration in the virtual human-machine collaboration scene based on the human-machine action label; in the process of virtual human-machine collaboration, obtain the human body collision part when the virtual digital human collides with the physical assets in the virtual human-machine collaboration scene, the multi-angle human-machine scene image of the virtual human-machine collaboration scene and the joint posture data of the virtual robotic arm; for the human body collision part The method uses a large multimodal model to convert the multimodal dataset into a visual-language instruction dataset. The multimodal model is used to fine-tune the multimodal model to obtain a fine-tuned multimodal model. The fine-tuned multimodal model is used to detect anomalies in human-machine collaboration in virtual human-machine collaboration scenarios. The server can provide feedback to the terminal on the obtained multimodal dataset, the fine-tuned multimodal model, and the anomaly detection results.

[0068] In addition, in some embodiments, the human-computer collaborative dataset construction and anomaly detection method for integrating a multimodal large model can also be implemented independently by a server or a terminal. For example, the terminal can directly process the dataset construction and anomaly detection requests to be processed, or the server can obtain the dataset construction and anomaly detection requests to be processed from the data storage system and process the dataset construction and anomaly detection requests to be processed.

[0069] The terminal can be, but is not limited to, various desktop computers, notebook computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server can be implemented by a single server or a server cluster composed of multiple servers, and can also be a cloud server.

[0070] As shown in Figure 2 and Figure 3 The embodiment provides a human-machine collaboration dataset construction and anomaly detection method based on a multi-modal large model. The method is executed by a computer device, which can be executed by a terminal or a server, or by both a terminal and a server. In this embodiment, the method is applied to a server in Figure 1 The method includes the following steps:

[0071] Step S1: Based on the human-machine action label, a virtual digital person and a virtual robot arm perform virtual human-machine collaboration in a virtual human-machine collaboration scene.

[0072] Step S2: During the virtual human-machine collaboration, the human body collision part when the virtual digital person collides with a physical asset in the virtual human-machine collaboration scene, multi-angle human scene images of the virtual human-machine collaboration scene, and joint posture data of the virtual robot arm are obtained.

[0073] Step S3: The human body collision part, the multi-angle human scene images, and the joint posture data are time-stamped aligned, and the time-stamped aligned human body collision part, multi-angle human scene images, and joint posture data are combined to form human-machine collaboration virtual scene time series data corresponding to the human-machine action label.

[0074] Step S4: Each human-machine action label and the human-machine collaboration virtual scene time series data corresponding to the human-machine action label are combined to form a multi-modal dataset.

[0075] Step S5: The multi-modal dataset is converted into a visual-linguistic instruction dataset by using a multi-modal large model.

[0076] Step S6: The multi-modal large model is fine-tuned by using the visual-linguistic instruction dataset, and a fine-tuned multi-modal large model is obtained. The fine-tuned multi-modal large model is used to realize human-machine collaboration anomaly detection in a virtual human-machine collaboration scene.

[0077] By implementing the steps S1 to S6 described above, the embodiment provides a one-stop human-machine collaboration scene multi-modal data acquisition and automatic label construction scheme, solves the problems of long acquisition cycle and low efficiency in data labeling caused by the cumbersome process during real machine debugging, avoids the potential collision risk in the human-machine collaboration scene, and can acquire scene images from different angles, so as to efficiently construct a high-quality multi-modal data set and fine-tune a multi-modal large model, and further realize human-machine collaboration anomaly detection in the virtual human-machine collaboration scene by using the fine-tuned multi-modal large model.

[0078] In the embodiment, the human-machine action label-based shooting strategy is used to acquire the human collision part, multi-angle human scene image and joint posture data in the virtual human-machine collaboration process.

[0079] The core of the human-machine action label-based shooting strategy is to use the human-machine action label to classify and label the scene information such as the tester's body posture, the mechanical arm working state and the human-machine spatial distance, as the target guidance of human-machine collaboration behavior, so as to facilitate subsequent automatic acquisition of multi-modal data and automatic construction of labels in the virtual human-machine collaboration scene.

[0080] The categories of the human-machine action label are determined based on the needs of specific human-machine collaboration tasks. For example, Figure 4 As shown in the table, in the human-machine action label, the tester's body posture includes but is not limited to typical state labels such as standing, squatting (i.e. squatting and sitting), and lying down, the mechanical arm working state includes but is not limited to key action phase labels such as initial static, movement, and grabbing, wherein the movement includes translation and rotation of the mechanical arm, but the end of the mechanical arm does not perform grabbing, and grabbing refers to the end of the mechanical arm grabbing an object, and the human-machine spatial distance can be divided into multiple labels that can reflect the degree of interaction safety according to the actual human-machine collaboration, including but not limited to interaction safety labels such as safe normal, warning reminder, and danger braking, wherein when the human-machine spatial distance is less than or equal to the minimum safe collaboration distance, it corresponds to danger braking, when the human-machine spatial distance is greater than the minimum safe collaboration distance and less than or equal to a preset distance, it corresponds to warning reminder, and when the human-machine spatial distance is greater than the preset distance, it corresponds to safe normal. It should be noted that the preset distance is determined according to the safe operation distance of the specific human-machine collaboration task.

[0081] At this time, in the embodiment, the human-machine action label includes the label of the tester's body posture, the label of the mechanical arm working state and the label of the human-machine spatial distance, the label of the tester's body posture includes standing, squatting, sitting and lying down, the label of the mechanical arm working state includes static, movement but not grabbing and grabbing, and the label of the human-machine spatial distance includes safe normal, warning reminder and danger braking.

[0082] The embodiment pre-establishes a digital twin system for human-robot collaboration. Specifically, a scene modeling is performed on a real human-robot collaboration environment to obtain a virtual human-robot collaboration scene (also referred to as a human-robot collaboration digital twin scene). Then, a virtual digital human (also referred to as a digital twin human body model) and a virtual robot arm (also referred to as a digital twin robot arm model) that can perform virtual human-robot collaboration in the virtual human-robot collaboration scene are constructed in the virtual human-robot collaboration scene. The virtual digital human is driven by a real tester and is synchronized with the real tester in action. The virtual robot arm is driven by a real robot arm and is synchronized with the real robot arm in action.

[0083] The human body posture of the real tester is mapped into the virtual digital human in real time through the virtual reality device, so that the virtual digital human is synchronized with the real tester in action. The virtual reality device includes a head-mounted display device worn on the head of the real tester, a handle worn on the hand of the real tester, and a plurality of positioning base stations installed in a test area where the real tester is located. The positioning base stations are used to emit laser signals. A gyroscope and a laser receiver are arranged in the head-mounted display device and the handle. The gyroscope is used to measure the posture of the real tester. The laser receiver is used to measure the position of the real tester based on the received laser signals. The head-mounted display device is used to view the virtual human-robot collaboration between the virtual digital human and the virtual robot arm in a first-person perspective. The handle has a button for controlling the movement and hand action of the virtual digital human. The handle can control the virtual digital human to move forward, backward, left, and right, to turn, and to select and map a hand action to the virtual digital human. The hand action can physically interact with the virtual human-robot collaboration scene, such as teaching the robot arm, grasping an object, and pushing or pulling a door. After the human body posture data (i.e., the position and posture of the real tester) of the real tester is collected through the virtual reality device, the virtual digital human is further driven based on the human body posture data. Specifically, the human body posture data and a skeletal model of the virtual digital human are used as inputs to solve joint data of the skeletal model by using a skeletal inverse kinematics solver, and the virtual digital human is driven based on the joint data. When the handle has the function of selecting a hand action, the human body posture data also includes the hand action.

[0084] The joint motion data of the real robot arm is transmitted to the virtual robot arm in real time through a sensor, and the virtual robot arm is further driven to move, so that the virtual robot arm and the real robot arm move synchronously. The sensor includes a speed sensor and an angle sensor installed at each joint of the real robot arm. Each joint of the virtual robot arm is provided with a revolute pair constraint and a joint angle variable. After the joint motion data (i.e. the angle and speed of each joint) is collected by the sensor, the virtual robot arm is driven to move based on the joint motion data, specifically including: using the joint motion data as input, updating the joint angle variable of each joint of the virtual robot arm by using a state updater, and driving the virtual robot arm to move.

[0085] At this time, in the embodiment, the virtual digital human and the virtual robot arm perform virtual human-robot collaboration in the virtual human-robot collaboration scene based on the human-robot action label, specifically including:

[0086] (1) obtaining human posture data of a real tester and joint motion data of a real robot arm, the human posture data and the joint motion data being data generated when the real tester and the real robot arm move based on the human-robot action label.

[0087] It should be noted that the real tester and the real robot arm do not actually perform human-robot collaboration, but only perform their own actions based on the human-robot action label, so that the virtual digital human and the virtual robot arm perform the same virtual human-robot collaboration as the actual human-robot collaboration.

[0088] (2) driving the virtual digital human to move based on the human posture data, and driving the virtual robot arm to move based on the joint motion data, so that the virtual digital human and the real tester move synchronously, and the virtual robot arm and the real robot arm move synchronously, so that the virtual digital human and the virtual robot arm perform virtual human-robot collaboration in the virtual human-robot collaboration scene.

[0089] In the embodiment, during the virtual human-robot collaboration, the human collision part when the virtual digital human collides with the physical asset in the virtual human-robot collaboration scene, the multi-angle human scene image of the virtual human-robot collaboration scene, and the joint posture data of the virtual robot arm are obtained. The human collision part is collected when the collision occurs, the multi-angle human scene image is collected at a fixed first frequency, and the joint posture data of the virtual robot arm is collected at a fixed second frequency.

[0090] The human body collision part is obtained by determining whether a physical collision occurs between the virtual digital human and the physical assets in the virtual human-machine collaboration scene, and if so, determining and recording the human body collision part and updating the human-machine state; if not, updating the human-machine state. The determination and recording of the human body collision part specifically includes: recording the human joint posture data of the virtual digital human in the frame before the collision and the human joint posture data of the virtual digital human in the current frame of collision using a collision event distributor to obtain posture record data; determining the human body collision part using a ray collision detection function based on the posture record data; and recording the human body collision part.

[0091] The multi-angle human scene image is obtained by setting up a multi-angle image automatic acquisition system in the virtual human-machine collaboration scene, i.e., the multi-angle image automatic acquisition system is set up based on the virtual human-machine collaboration scene. Specifically, a set of virtual cameras are arranged at three orthogonal positions in the virtual human-machine collaboration scene to capture the dynamic changes of the virtual human-machine collaboration scene in real time from three different perspectives, i.e., front view, side view and overhead view. Specifically, the normal vector of the xoz plane corresponds to the front view, the normal vector of the yoz plane corresponds to the side view, and the normal vector of the xoy plane corresponds to the overhead view using the three-dimensional coordinate system in the Unreal Engine to obtain three shooting videos. For each shooting video, the scene texture capture component is used to convert the shooting video into a frame-by-frame image, and the frame-by-frame image is further collected at a fixed first frequency to obtain the multi-angle human scene image. The multi-angle human scene image is saved to a local folder, and the human-machine action label corresponding to the multi-angle human scene image is used as the name of the multi-angle human scene image through a program script. The scene texture capture component is a functional component in the software (such as the Unreal Engine) used to create a digital twin system, and the program script is a user-defined program which functions to use the human-machine action label corresponding to the multi-angle human scene image as the name of the multi-angle human scene image.

[0092] At this time, the multi-angle human scene image is obtained by video frame extraction and fixed frequency collection of the shooting video. The shooting video is obtained by setting a set of virtual cameras in the virtual human-machine collaboration scene. The set of virtual cameras includes three virtual cameras, and the shooting angles of the three virtual cameras are orthogonal to each other. Specifically, the multi-angle human scene image includes a human scene image of the virtual human-machine collaboration scene in the front view, a human scene image of the virtual human-machine collaboration scene in the side view, and a human scene image of the virtual human-machine collaboration scene in the overhead view, as shown in Figure 5 、 Figure 6 and Figure 7 .

[0093] After acquiring the multi-angle human-robot scene image, the method for constructing a human-robot collaboration dataset and performing anomaly detection by the fusion multi-modal large model of the embodiment further includes: taking the human-robot action label as the name of the multi-angle human-robot scene image to name the multi-angle human-robot scene image.

[0094] The joint posture data is acquired by collecting the joint posture of the virtual robot arm at a fixed second frequency to obtain joint posture data, which includes the angle and speed of each joint of the virtual robot arm.

[0095] The embodiment combines the human collision part, the multi-angle human-robot scene image of the virtual human-robot collaboration scene, and the joint posture data of the virtual robot arm, and constructs human-robot collaboration virtual scene time series data based on the human-robot action label through a timestamp alignment algorithm. As shown in Figure 8 the first step is to extract the timestamps of the human collision part, the timestamps of the multi-angle human-robot scene image, and the timestamps of the joint posture data. Then, the minimum and maximum values of all timestamps are determined to define the time range and determine the time interval. Next, the timestamps of different sources (i.e., the human collision part, the multi-angle human-robot scene image, and the joint posture data) are aligned to a unified time interval through methods such as traversal and interpolation. The human collision part, the multi-angle human-robot scene image, and the joint posture data are integrated according to the aligned timestamp order to achieve time series alignment and obtain human-robot collaboration virtual scene time series data. Finally, the aligned human collision part, multi-angle human-robot scene image, and joint posture data are fused with the human-robot action label, so that the human-robot collaboration virtual scene time series data integrates the classification labels of different human-robot collaboration behaviors, forming multi-modal raw data with spatiotemporal characteristics, and reflecting the dynamic changes in the human-robot collaboration scene.

[0096] At this time, in the embodiment, the human collision part, the multi-angle human-robot scene image, and the joint posture data are timestamped, and the human collision part, the multi-angle human-robot scene image, and the joint posture data after timestamping form human-robot collaboration virtual scene time series data corresponding to the human-robot action label. Specifically, the timestamp alignment of the human collision part, the multi-angle human-robot scene image, and the joint posture data includes: determining the maximum and minimum values of the timestamps of the human collision part, the timestamps of the multi-angle human-robot scene image, and the timestamps of the joint posture data; determining the time interval based on the maximum and minimum values, and aligning the timestamps of the human collision part, the multi-angle human-robot scene image, and the joint posture data through traversal and interpolation.

[0097] After obtaining the human-robot collaborative virtual scene timing data, the embodiment further groups each human-robot action label and the human-robot collaborative virtual scene timing data corresponding to the human-robot action label into a multi-modal data set, and completes the construction of the multi-modal data set based on the virtual human-robot collaborative scene.

[0098] The embodiment uses a multi-modal large model to convert the multi-modal data set of the virtual human-robot collaborative scene into a visual-linguistic instruction data set in a visual-linguistic instruction format, and further improves the adaptability of the multi-modal large model in the human-robot collaborative scene through supervised fine-tuning training.

[0099] The multi-modal data set of the virtual human-robot collaborative scene is converted into a visual-linguistic instruction data set in a visual-linguistic instruction format using a multi-modal large model, specifically including: using a multi-modal large model such as CogVLM2, GLM4V, Qwen-VL-Chat or MiniCPM-V-2.5 to automatically annotate the virtual scene images (i.e. the multi-angle human scene images after alignment) in the multi-modal data set, generating image titles, the image title is a textual description generated for the visual information contained in the multi-angle human scene image, the textual description includes specific visual information such as human body posture, human body collision part and other context-related elements, which helps to establish the association between vision and language, so that the multi-modal large model can understand the content in the multi-angle human scene image and convert it into a textual description, the image title describes the key visual elements in the multi-angle human scene image, for example, the image title is: a standing person interacting with a robot arm, the left limb collides. 100 dialogue templates are pre-constructed as seed samples, which can help guide the multi-modal large model to generate instruction data with specific patterns and styles, the content of the dialogue template can involve possible interaction modes between the user and the assistant, such as asking about the image content, reasoning about the scene, etc., for example, the dialogue template is: "Please describe the human body posture in the image.", "Is there any potential dangerous state in this scene?", "How do the actions of the robot arm and the human body position interact?". In combination with the multi-modal data set, the image title and the seed sample, the multi-modal large model such as ChatGPT is guided to generate a large-scale high-quality visual-linguistic instruction data set for instruction optimization, as shown in Figure 9 The visual-linguistic instruction data set includes three types of text data:

[0100] (1) Detailed description type data: each sample contains a round of dialogue, mainly describing the detailed information of the image and its label. The detailed description type data is a single round of description, which emphasizes the comprehensiveness of the image content and constructs a rich and comprehensive explanation. For example, the detailed description type data can be:

[0101]

[0102]

[0103] (2) Long dialogue type data: each sample involves multiple rounds of dialogue, simulating the process of user interaction with the assistant on image content. Long dialogue type data is used to simulate multiple rounds of dialogue, gradually deepening the understanding of abnormal detection in virtual human-machine collaboration, such as asking for the safe distance of human-machine interaction, abnormal state, etc. For example, long dialogue type data can be:

[0104]

[0105]

[0106] (3) Complex reasoning type data: each sample focuses on scenes and problems that require visual reasoning. Complex reasoning type data is used to promote the model to generate visual-based reasoning analysis, such as inferring abnormal conditions from image content, detecting possible safety risks, etc. For example, complex reasoning type data can be:

[0107]

[0108] At this time, in the embodiment, the multi-modal large model is used to convert the multi-modal data set into a visual-linguistic instruction data set, specifically including:

[0109] (1) Using a multi-modal large model to automatically annotate the aligned multi-angle human scene images in the multi-modal data set to obtain image titles.

[0110] Among them, the multi-modal large model used to obtain the image title is CogVLM2, GLM4V, Qwen-VL-Chat or MiniCPM-V-2.5.

[0111] (2) Using a multi-modal large model to generate a visual-linguistic instruction data set with the multi-modal data set and the image title as input.

[0112] Among them, using a multi-modal large model to generate a visual-linguistic instruction data set with the multi-modal data set and the image title as input, specifically including: constructing multiple dialogue templates, the dialogue templates are used to guide the multi-modal large model to generate instruction data with specific patterns and styles, the content of the dialogue template includes asking image content and reasoning scene; using a multi-modal large model to generate a visual-linguistic instruction data set with the multi-modal data set, the image title and the multiple dialogue templates as input, the visual-linguistic instruction data set includes multiple types of text data, and the type of text data includes detailed description type data, long dialogue type data and complex reasoning type data.

[0113] Among them, the multi-modal large model used to generate the visual-linguistic instruction data set is ChatGPT.

[0114] It should be noted that each sample in the visual-linguistic instruction dataset still includes the aligned human collision part, multi-angle human-robot scene image and joint pose data, and human-robot action label, but the data information is described in the form of natural language.

[0115] The adaptability of the multi-modal large model in the human-robot collaboration scene is improved through supervised fine-tuning training, specifically including: fine-tuning the multi-modal large model through the constructed visual-linguistic instruction dataset, in the fine-tuning process, adopting a fine-tuning strategy based on training loss optimization, the fine-tuning strategy based on training loss optimization adopts an AdamW optimizer, a cosine learning rate scheduler and a mixed precision strategy combining float16 and float32, and subsequently using the fine-tuned multi-modal large model to identify the abnormal state in the human-robot collaboration process, to realize abnormal detection, and further respond to the abnormal state, to generate a text response.

[0116] At this time, in the embodiment, the multi-modal large model is fine-tuned using the visual-linguistic instruction dataset to obtain a fine-tuned multi-modal large model, specifically including: taking the visual-linguistic instruction dataset as input, fine-tuning the multi-modal large model using a fine-tuning strategy based on training loss optimization to obtain a fine-tuned multi-modal large model, wherein the fine-tuning strategy based on training loss optimization includes an AdamW optimizer, a cosine learning rate scheduler and a mixed precision strategy.

[0117] Among them, the multi-modal large model fine-tuned to realize abnormal detection is CogVLM2, GLM4V, Qwen-VL-Chat or MiniCPM-V-2.5.

[0118] The fine-tuned multi-modal large model is used to realize human-robot collaboration abnormal detection in a virtual human-robot collaboration scene, and the fine-tuned multi-modal large model is used to realize human-robot collaboration abnormal detection in a virtual human-robot collaboration scene, specifically including: taking the currently acquired multi-angle human-robot scene image to be detected as input, using the fine-tuned multi-modal large model to perform abnormal detection, to obtain the tester's human body posture, the mechanical arm working state, the human body pixel coordinates and the mechanical arm pixel coordinates; performing coordinate transformation on the human body pixel coordinates to obtain human body space coordinates; performing coordinate transformation on the mechanical arm pixel coordinates to obtain mechanical arm space coordinates; calculating the Euclidean distance between the human body space coordinates and the mechanical arm space coordinates to obtain the human-robot space distance; determining the human-robot safety degree based on the human-robot space distance and the minimum safe collaboration distance, specifically when the human-robot space distance is less than or equal to the minimum safe collaboration distance, the human-robot safety degree is dangerous braking, when the human-robot space distance is greater than the minimum safe collaboration distance and less than or equal to a preset distance, the human-robot safety degree is warning, and when the human-robot space distance is greater than the preset distance, the human-robot safety degree is safe and normal, specifically including the following steps:

[0119] (1) Coordinate extraction and analysis: The fine-tuned multi-modal large model first receives the input multi-angle human-robot scene image to be detected. Based on its multi-modal capability, the fine-tuned multi-modal large model further combines the visual pre-training model (ViT-biG) of the target task to analyze the content of the multi-angle human-robot scene image to be detected, and generates the coordinate frame and key points of the human body and the robot arm (i.e. human body pixel coordinates and robot arm pixel coordinates).

[0120] (2) Human-robot coordinate and spatial relationship calculation: Based on the coordinate frame and key points of the human body and the robot arm detected, the tester's human body posture, the robot arm working state, and the human-robot spatial distance are calculated to determine the relative spatial position of the human body and the robot arm. The calculation method of the human-robot spatial distance is as follows: based on the obtained human body pixel coordinates and robot arm pixel coordinates, a spatial distance calculation module is used to calculate the human-robot spatial distance. The spatial distance calculation module is responsible for extracting human body pixel coordinates and robot arm pixel coordinates from the image, converting the detected pixel coordinates into actual spatial coordinates (x, y, z), and specifically using a multi-angle camera to perform triangulation to convert the extracted 2D pixel coordinates into 3D spatial coordinates to obtain human body spatial coordinates and robot arm spatial coordinates. The Euclidean distance between the two sets of coordinate points (i.e. human body spatial coordinates and robot arm spatial coordinates) is calculated to obtain the human-robot spatial distance: where D is the human-robot spatial distance, (x1, y1, z1) represents the human body spatial coordinates, and (x2, y2, z2) represents the robot arm spatial coordinates.

[0121] (3) Application of abnormal rules: After the human-robot spatial distance is extracted, based on the rule base of collision and safety, which can be automatically learned and optimized by the pre-defined training data set during the fine-tuning process, the spatial distance calculation module determines the human-robot safety level according to the calculated human-robot spatial distance D: if D≤ minimum safe collaboration distance, it is judged as dangerous braking; if minimum safe collaboration distance < D < D1, D1 is a preset distance, it is judged as warning reminder; if D > D1, it is judged as safe normal, and the final human-robot safety level is fed back to the multi-modal large model.

[0122] (4) Abnormal type discrimination and response generation: When an abnormal situation (such as too close distance or collision) is detected, the model can generate a corresponding text response.

[0123] wherein the calculation method of the minimum safe collaboration distance includes:

[0124] (1) Obtain a plurality of sampling multi-angle human-robot scene images in a sampling time period, the sampling time period is (T0, T1+T), T0 is the time point when the human body intrudes, that is, the specific time point when the human body is detected to enter the virtual human-robot collaborative scene, T1 is the time point when the human body is less than a certain distance from the robot arm, triggering the robot arm to retreat, T is the time used for the robot arm to start retreating to the safe distance and stop, which is a preset value determined by the manufacturer.

[0125] (2) For each sampling multi-angle human-robot scene image, use the fine-tuned multi-modal large model to perform anomaly detection on the sampling multi-angle human-robot scene image as input, and obtain the sampling human body pixel coordinates and the sampling robot arm pixel coordinates.

[0126] (3) For any two adjacent sampling multi-angle human-robot scene images, calculate the ratio of the sampling human body pixel coordinates of the two adjacent sampling multi-angle human-robot scene images to the time interval of the two adjacent sampling multi-angle human-robot scene images, to obtain the sampling human body moving speed, and calculate the ratio of the sampling robot arm pixel coordinates of the two adjacent sampling multi-angle human-robot scene images to the time interval of the two adjacent sampling multi-angle human-robot scene images, to obtain the sampling robot arm moving speed.

[0127] The two adjacent images inputted are sampled, and the moving speed of the human body and the robot arm in the time interval dt is calculated, wherein the human body moving speed is obtained by dividing the change of the human body pixel coordinates in the two adjacent images by dt, and the robot arm moving speed is obtained by dividing the change of the robot arm pixel coordinates in the two adjacent images by dt.

[0128] (4) Calculate the minimum safe collaborative distance according to the sampling human body moving speed and the sampling robot arm moving speed.

[0129] The calculation formula of the minimum safe collaborative distance is:

[0130] Sp=Sh+Sr+Ss+C;

[0131] Wherein, Sp is the minimum safe collaborative distance; Sh is the related distance of the human body approaching speed, Vh(t) is the moving speed of the human body at time t, which is determined based on the sampling human body moving speed; Sr is the related distance of the normal working speed of the robot arm, Vr(t) is the moving speed of the robot arm in the direction of approaching the human body at time t, which is determined based on the sampling robot arm moving speed; Ss is the stopping path distance of the robot arm, Vs(t) is the moving speed of the robot arm along the stop path at time t, determined based on the sampling robot arm moving speed, generally determined by the manufacturer; C is the distance of the human body intrusion, that is, the distance of the human body from the specific boundary (through which the human body enters the virtual human-robot collaborative scene) when the human body is detected to have just entered the virtual human-robot collaborative scene.

[0132] wherein generating a text response refers to: when an abnormal situation is detected, the model generates a suitable text response according to the learned visual-linguistic instruction data set, prompting the user to take measures or make adjustments. That is, the text response refers to the text automatically generated by the fine-tuned multi-modal large model for reminding the user or providing solutions when the model is performing human-robot collaboration anomaly detection, and the content can include the following aspects:

[0133] (1) Description of abnormal state: the model can provide a brief description of the abnormal situation, for example, the robot arm is too close, there is a risk of collision.

[0134] (2) Safety suggestions or warnings: if potential danger is detected, the model may give a warning prompt, for example, please adjust the position of the robot arm to maintain a safe distance, for example, the robot arm is currently in the wrong working mode.

[0135] (3) Solution or operation suggestion: if the abnormal state is related to a specific operation, the model may suggest how the user adjusts the working parameters or recalibrates the device, for example, please place the robot arm in the initial position and restart.

[0136] (4) Task status update: in some cases, the model can also provide the status of the current collaboration task, for example, the task is suspended, waiting for safety confirmation, for example, the task continues, the current state is safe.

[0137] As an example, assuming that the model detects that the robot arm is close to the human body pose and the distance is too close, and is in a high-risk grabbing stage, the text response of the model may be: Warning: the current robot arm is too close to the human body and may collide, please adjust the position of the robot arm to maintain a safe distance; Prompt: please ensure that the end of the robot arm does not perform grabbing operation to prevent accidents.

[0138] The embodiment provides a human-machine collaboration data set construction and anomaly detection method fusing a multi-modal large model, and relates to the technical field of human-machine collaboration, and the method comprises the following steps: building a multi-angle image automatic acquisition system in a virtual human-machine collaboration scene, acquiring multi-angle human scene images in a human-machine collaboration process by using a shooting strategy based on human-machine action labels, combining human collision parts, joint posture data of a collaboration mechanical arm and the multi-angle human scene images, constructing human-machine collaboration virtual scene time sequence data based on human-machine action labels by using a timestamp alignment algorithm, further constructing a multi-modal data set, converting the multi-modal data set into a visual-language instruction data set by using a multi-modal large model, improving the adaptability of the multi-modal large model in a human-machine collaboration scene by using supervised fine-tuning training, and finally realizing anomaly detection by using the fine-tuned multi-modal large model. The embodiment provides an all-in-one human-machine collaboration scene multi-modal data acquisition and automatic label construction scheme, meets the training requirements of the current multi-modal large model, and improves the scene generalization understanding and intelligent detection level of the human-machine collaboration task.

[0139] The method of the embodiment has the following advantages:

[0140] (1) The all-in-one human-machine collaboration scene multi-modal data acquisition and automatic label construction scheme solves the problems of long acquisition cycle and low efficiency in data labeling caused by the cumbersome process during real machine debugging, avoids potential collision risks in the human-machine collaboration scene, and can acquire scene images at different angles.

[0141] (2) The human-machine collaboration virtual scene time sequence data integrates different modal data such as human collision parts, joint posture data of a collaboration mechanical arm and multi-angle human scene images, and provides more dimensional and more detailed scene information data for human-machine collaboration.

[0142] (3) The large-scale high-quality visual-language instruction data set meets the training requirements of the current multi-modal large model task, and improves the scene generalization understanding and intelligent detection level of the human-machine collaboration task.

[0143] The application further provides an application scenario of the method for constructing a human-machine collaboration data set and performing anomaly detection by using a multi-modal large model. Specifically, the method for constructing a human-machine collaboration data set and performing anomaly detection by using a multi-modal large model can be applied in a human-machine collaboration anomaly detection scenario. The human-machine collaboration anomaly detection scenario includes a data set construction link, a model fine-tuning link, and an anomaly detection link. The data set construction link is configured to construct a multi-modal data set based on digital twinning. The model fine-tuning link is configured to fine-tune a multi-modal large model based on the multi-modal data set to obtain a fine-tuned multi-modal large model. The anomaly detection link is configured to perform anomaly detection in a human-machine collaboration process by using the fine-tuned multi-modal large model. The method for constructing a human-machine collaboration data set and performing anomaly detection by using a multi-modal large model provided in the application belongs to the data set construction link, the model fine-tuning link, and the anomaly detection link.

[0144] Embodiment 2

[0145] In an exemplary embodiment, a computer device, which can be a server or a terminal, is provided. An internal structure diagram of the computer device can be as shown in FIG. 1. Figure 10 The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store data. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement a method for constructing a human-machine collaboration data set and performing anomaly detection by using a multi-modal large model.

[0146] Those skilled in the art can understand that Figure 10 The structure shown in FIG. 1 is only a block diagram of part of the structure related to the scheme of the application, and does not constitute a limitation on the computer device to which the scheme of the application is applied. Specifically, the computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0147] In an example embodiment, a computer device is also provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the computer program to implement the method for constructing a human-machine collaboration dataset and detecting an anomaly of a fusion multi-modal large model according to embodiment 1.

[0148] Embodiment 3

[0149] The embodiments of the present application provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method for constructing a human-machine collaboration dataset and detecting an anomaly of a fusion multi-modal large model according to embodiment 1.

[0150] Embodiment 4

[0151] The embodiments of the present application provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the method for constructing a human-machine collaboration dataset and detecting an anomaly of a fusion multi-modal large model according to embodiment 1.

[0152] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0153] The technical features of the above embodiments can be combined arbitrarily, and to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present application.

[0154] The principles and implementation modes of the present application are described by applying specific examples in the present application, and the above embodiment descriptions are only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In conclusion, the content of the present application should not be understood as a limitation.

Claims

1. A method for constructing a human-machine collaboration dataset and detecting anomalies by fusing multi-modal large models, characterized in that, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: In the process of virtual human-robot collaboration, the human body collision part, multi-angle human scene image and joint posture data of the virtual robot are obtained when the virtual digital human collides with the physical assets in the virtual human-robot collaboration scene. The human body collision part, multi-angle human scene image and joint posture data are timestamped, and the timestamped human body collision part, multi-angle human scene image and joint posture data are combined to form human-robot collaboration virtual scene time sequence data corresponding to the human-robot action label. Each human-robot action label and the human-robot collaboration virtual scene time sequence data corresponding to the human-robot action label form a multi-modal data set. The multi-modal data set is converted into a visual-linguistic instruction data set by using a multi-modal large model. The multi-modal large model is fine-tuned using the visual-linguistic instruction data set to obtain a fine-tuned multi-modal large model. The fine-tuned multi-modal large model is used to realize human-robot collaboration anomaly detection in a virtual human-robot collaboration scene. The fine-tuned multi-modal large model is used to realize human-robot collaboration anomaly detection in a virtual human-robot collaboration scene. The tester's human body posture, robot working state, human body pixel coordinates and robot pixel coordinates are obtained by using the fine-tuned multi-modal large model to perform anomaly detection on the current obtained multi-angle human scene image. The human body pixel coordinates are transformed to obtain human body spatial coordinates. The robot pixel coordinates are transformed to obtain robot spatial coordinates. The Euclidean distance between the human body spatial coordinates and the robot spatial coordinates is calculated to obtain the human-robot spatial distance. The human-robot safety degree is determined based on the human-robot spatial distance and the minimum safe collaboration distance. When the human-robot spatial distance is less than or equal to the minimum safe collaboration distance, the human-robot safety degree is dangerous braking. When the human-robot spatial distance is greater than the minimum safe collaboration distance and less than or equal to the preset distance, the human-robot safety degree is warning. When the human-robot spatial distance is greater than the preset distance, the human-robot safety degree is safe and normal. The method for calculating the minimum safe collaboration distance comprises the following steps: A plurality of sampling multi-angle human scene images are obtained within a sampling time period. For each of the sampling multi-angle human scene images, the human body pixel coordinates and the robot pixel coordinates are obtained by using the fine-tuned multi-modal large model to perform anomaly detection on the sampling multi-angle human scene image. For any two adjacent sampling multi-angle human scene images, the ratio of the sampling human pixel coordinates of the two adjacent sampling multi-angle human scene images to the time interval of the two adjacent sampling multi-angle human scene images is calculated to obtain the sampling human moving speed, and the ratio of the sampling mechanical arm pixel coordinates of the two adjacent sampling multi-angle human scene images to the time interval of the two adjacent sampling multi-angle human scene images is calculated to obtain the sampling mechanical arm moving speed; The minimum safe cooperation distance is calculated according to the sampling human moving speed and the sampling mechanical arm moving speed; The calculation formula of the minimum safe cooperation distance is: Sp=Sh+Sr+Ss+C; Wherein, Sp is the minimum safety cooperation distance; Sh is the related distance of the human body approaching speed, Vh(t) is the moving speed of the human body at t time, determined based on sampling with the human body moving speed; Sr is the related distance of the normal working speed of the mechanical arm, Vr(t) is the moving speed of the mechanical arm along the direction of approaching the human body at t time, determined based on sampling with the mechanical arm moving speed; Ss is the stopping path distance of the mechanical arm, Vs(t) is the moving speed of the mechanical arm along the stopping path at t time, determined based on sampling with the mechanical arm moving speed; C is the distance of the human body intrusion.

2. The method of claim 1, wherein the method further comprises: The human-machine action label includes: a label of a tester's human posture, a label of a mechanical arm working state, and a label of a human-machine spatial distance; the label of the tester's human posture includes standing, squatting, sitting, and lying down; the label of the mechanical arm working state includes being stationary, moving but not grabbing, and grabbing; and the label of the human-machine spatial distance includes being safe and normal, being warning and reminding, and being dangerous and braking.

3. The method of claim 1, wherein the method further comprises: The virtual digital human and the virtual mechanical arm are controlled to perform virtual human-machine cooperation in a virtual human-machine cooperation scene based on the human-machine action label, and specifically include: Obtaining human posture data of a real tester and joint motion data of a real mechanical arm; the human posture data and the joint motion data are data generated when the real tester and the real mechanical arm move based on the human-machine action label; The virtual digital human is driven to move based on the human posture data, and the virtual mechanical arm is driven to move based on the joint motion data, so that the virtual digital human and the real tester move synchronously, and the virtual mechanical arm and the real mechanical arm move synchronously, so that the virtual digital human and the virtual mechanical arm perform virtual human-machine cooperation in the virtual human-machine cooperation scene.

4. The method of claim 1, wherein the method further comprises: The multi-angle human scene image is an image obtained after video frame extraction and fixed frequency collection are performed on a shooting video, and the shooting video is a video obtained by shooting through a set of virtual cameras arranged in the virtual human-machine cooperation scene; the set of virtual cameras includes three virtual cameras, and the shooting angles of the three virtual cameras are orthogonal to each other; The multi-angle human scene image includes a human scene image of the virtual human-machine cooperation scene in a front view, a human scene image of the virtual human-machine cooperation scene in a side view, and a human scene image of the virtual human-machine cooperation scene in a top view; The method further includes naming the multi-angle human scene image by taking the human-machine action label as the name of the multi-angle human scene image.

5. The method of claim 1, wherein the method further comprises: The human body collision part, the multi-angle human scene image, and the joint posture data are time stamped and aligned, specifically including: Determining the maximum and minimum values of the time stamp of the human body collision part, the time stamp of the multi-angle human scene image, and the time stamp of the joint posture data; Determining a time interval based on the maximum value and the minimum value, and time stamping and aligning the human body collision part, the multi-angle human scene image, and the joint posture data by traversing and interpolating.

6. The method of claim 1, wherein the method further comprises: Converting the multi-modal data set into a visual-linguistic instruction data set by using a multi-modal large model, specifically comprising: Automatically annotating the aligned multi-angle human scene images in the multi-modal data set by using a multi-modal large model to obtain image titles. Taking the multi-modal data set and the image titles as inputs, generating a visual-linguistic instruction data set by using a multi-modal large model.

7. The method of claim 6, wherein the method further comprises: The multi-modal large model used to obtain the image titles is CogVLM2, GLM4V, Qwen-VL-Chat or MiniCPM-V-2.5; the multi-modal large model used to generate the visual-linguistic instruction data set is ChatGPT; the multi-modal large model used to be fine-tuned for anomaly detection is CogVLM2, GLM4V, Qwen-VL-Chat or MiniCPM-V-2.

5.

8. The method of claim 7, wherein the method further comprises: Taking the multi-modal data set and the image titles as inputs, generating a visual-linguistic instruction data set by using a multi-modal large model, specifically comprising: Constructing a plurality of dialogue templates; the dialogue templates are used to guide the multi-modal large model to generate instruction data with specific modes and styles, and the content of the dialogue templates includes asking image content and reasoning scenarios; Taking the multi-modal data set, the image titles and a plurality of dialogue templates as inputs, generating a visual-linguistic instruction data set by using a multi-modal large model; the visual-linguistic instruction data set includes a plurality of types of text data, and the types of text data include detailed description type data, long dialogue type data and complex reasoning type data.

9. The method of claim 1, wherein the method further comprises: Fine-tuning the multi-modal large model by using the visual-linguistic instruction data set to obtain a fine-tuned multi-modal large model, specifically comprising: Taking the visual-linguistic instruction data set as input, fine-tuning the multi-modal large model by using a fine-tuning strategy based on training loss optimization to obtain a fine-tuned multi-modal large model; wherein the fine-tuning strategy based on training loss optimization includes an AdamW optimizer, a cosine learning rate scheduler and a mixed precision strategy.

Citation Information

Patent Citations

  • Personalized man-machine cooperation assembly safety detection and early warning method based on digital twinning

    CN116460857A

  • Multi-modal data exception identification method and device, electronic equipment and storage medium

    CN118379755A