Station action decomposition identification and data management method and system

By using multimodal data acquisition and feature fusion technology, workstation actions are identified and associated with the operation objects, solving the problems of inaccurate action boundary definition and data fragmentation in complex industrial sites. This enables structured data management and improves the traceability and control efficiency of the production process.

CN122067320AInactive Publication Date: 2026-05-19ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG SCI-TECH UNIV
Filing Date
2026-04-21
Publication Date
2026-05-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In complex industrial settings, single-vision-based workstation motion recognition methods struggle to accurately define motion boundaries and operational objects. This results in motion data failing to establish a queryable and analyzable structured association with material information, personnel information, and process document libraries, creating 'motion data silos' that cannot be used for in-depth production efficiency analysis and real-time compliance verification.

Method used

A multimodal data acquisition method is adopted, which combines visual, acoustic and inertial sensors to acquire multimodal raw data. The action type is identified by time alignment and feature fusion technology, the operation object identifier is associated, and composite actions are combined based on the preset workstation standard operation sequence to establish a data association index with the material, personnel and process document library.

Benefits of technology

It enables accurate decomposition and identification of workstation actions, solving the problems of inaccurate action identification and missing association of operation objects, and improving the traceability and control efficiency of the production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122067320A_ABST
    Figure CN122067320A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent manufacturing, and particularly discloses a station action decomposition identification and data management method and system, and the method comprises the steps: obtaining the multi-mode original data of a worker in a target station operation region, and obtaining a plurality of candidate action information of the continuous operation of the worker according to the multi-mode original data; and for each piece of candidate action information, obtaining a corresponding multi-modal action feature. According to the invention, by acquiring the multi-modal original data, identifying the action type, associating the operation object identifier, combining the composite action and establishing the data association index, the visual and light interference is overcome, the station action is accurately decomposed and identified, the operation object is reliably associated, and the structured data management is realized; the method has the advantages that the problems of inaccurate station action recognition, operation object association missing and data splitting can be effectively solved, the integration of action decomposition recognition and structured data management is realized, and the traceability and management and control efficiency of the production process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent manufacturing technology, and in particular to a method and system for decomposing, identifying, and managing workstation actions. Background Technology

[0002] The workstation motion decomposition, recognition, and data management system is a core digital equipment for operations that adapts to the needs of human-machine collaboration, flexible production, and refined control in modern intelligent manufacturing environments. It integrates multimodal perception, temporal motion analysis, and production context association technologies, aiming to achieve accurate perception, structured analysis, and traceable management of personnel operation behaviors.

[0003] In complex industrial settings, the presence of obstructions, varying lighting conditions, and interference from multiple objects in the work area, coupled with the numerous intricate, continuous, and complex sub-action sequences involved in worker operations, makes it difficult for single-vision-based recognition methods to accurately define action boundaries, differentiate between similar actions (such as "tightening" and "loosening"), or effectively identify the specific objects being manipulated (such as parts or tools of different specifications). Furthermore, the recognition results are typically isolated action labels or time segments, failing to establish a searchable and analyzable structured association with specific objects in the material information database, operator identities in the personnel information database, and standard operating procedures in the process document database. This creates "action data silos," preventing the use of massive amounts of operational behavior data for in-depth production efficiency analysis, real-time compliance verification, personnel skill assessment, and digital traceability of the production process. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for decomposing, identifying, and managing workstation actions, in order to solve the technical problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A method for decomposing, identifying, and managing workstation actions, comprising:

[0007] Acquire multimodal raw data of workers in the target workstation area, and obtain multiple candidate action information of continuous work based on the multimodal raw data;

[0008] For each candidate action information, obtain its corresponding multimodal action features, and obtain the action type based on the multimodal action features;

[0009] Based on the action type, obtain the operation object identifier associated with the occurrence of the candidate action information;

[0010] Obtain the preset workstation standard operation sequence, and combine the action type with the operation object identifier based on the preset workstation standard operation sequence to obtain the composite action;

[0011] The action record is obtained based on the composite action, and the action record is stored in the time-series database. Based on the operation object identifier and the operator identifier, a data association index is established with the material information database, the personnel information database, and the process document database.

[0012] Preferably, the step of obtaining multiple candidate action information of continuous operation of personnel based on the multimodal raw data includes:

[0013] Based on the timestamps of each modal data source in the original multimodal data, a time-aligned cross-modal synchronous data stream is generated;

[0014] Based on the cross-modal synchronization data stream, the visual pose skeleton sequence, acoustic event spectrum sequence, and inertial motion vector sequence are extracted respectively.

[0015] The motion trajectories of key joints are obtained based on the visual posture skeleton sequence, significant acoustic event markers are obtained based on the acoustic event spectrum sequence, and peak motion signal points are obtained based on the inertial motion vector sequence.

[0016] Based on the synchronous changes of the key joint motion trajectories, significant acoustic event markers, and peak motion signal points on the time axis, multiple significant motion intervals are obtained through analysis.

[0017] Based on the start and end points of each action salient interval, the corresponding data segments are extracted from the cross-modal synchronization data stream to obtain the information of the multiple candidate actions.

[0018] Preferably, the step of obtaining the corresponding multimodal action features for each candidate action information, and obtaining the action type based on the multimodal action features, includes:

[0019] Based on the candidate action information, the corresponding visual data segment, acoustic data segment, and inertial measurement data segment are separated.

[0020] Based on the visual data segment, extract the spatial coordinate set of joint points and the limb movement direction vector;

[0021] Based on the acoustic data segment, extract the spectral energy distribution characteristics and the zero-crossing rate change profile;

[0022] Based on the inertial measurement data segment, extract the triaxial acceleration amplitude envelope and angular velocity variation spectrum;

[0023] Based on the joint point spatial coordinate set, limb movement direction vector, spectral energy distribution characteristics, zero-crossing rate change profile, triaxial acceleration amplitude envelope and angular velocity change spectrum, a unified perception feature vector of the candidate action information is constructed.

[0024] Based on the real-time quality assessment results of each component sub-vector in the unified perception feature vector, the dynamic contribution weight of each component sub-vector in the classification decision is obtained.

[0025] The unified perception feature vector is weighted and integrated according to the dynamic contribution weight to generate an enhanced discriminative feature representation.

[0026] The action semantic category to which the candidate action information belongs is obtained by performing multi-dimensional matching between the enhanced discriminative feature representation and the preset action feature template library.

[0027] Preferably, the step of obtaining the operation object identifier associated with the occurrence of the candidate action information according to the action type includes:

[0028] Based on the action type, obtain the target spatial region and time window for object recognition;

[0029] Based on the target spatial region, extract region image blocks from the visual data segments corresponding to the candidate action information;

[0030] Image visual features are obtained from the image blocks in the region, and the category of the object to be operated is identified based on the image visual features;

[0031] Based on the time window, acoustic event segments are extracted from the acoustic data segments corresponding to the candidate action information;

[0032] Acoustic features are obtained from the acoustic event fragments, and sound patterns associated with the operation object category are matched based on the acoustic features;

[0033] If both the image visual features and the acoustic features can identify the category of the operation object, then the operation object identifier is determined according to the preset object category and identifier mapping table;

[0034] If the category of the operation object can only be identified by one of the image visual features or acoustic features, then the confidence level is evaluated based on the identification result and the historical operation object behavior database, and the operation object identifier is determined based on the evaluation result.

[0035] If the category of the operation object cannot be identified through image visual features or acoustic features, then the electronic tag of the material or tool within the time window is read from the radio frequency identification data segment corresponding to the candidate action information, and the code corresponding to the electronic tag is used as the identifier of the operation object.

[0036] Preferably, the steps of obtaining a preset workstation standard operation sequence and combining the action type with the operation object identifier based on the preset workstation standard operation sequence to obtain a composite action include:

[0037] Obtain the corresponding preset job sequence template based on the current workstation identifier. The preset job sequence template includes predefined action steps arranged in chronological order and their corresponding operation object type constraints.

[0038] Based on the time sequence of the predefined action steps and the constraints of the type of the operation object, generate multi-level action matching rules based on time windows;

[0039] Based on the identified action type and operation object identifier, construct the current action object association sequence in chronological order;

[0040] The current action object associated sequence is compared step by step according to the multi-level action matching rules to identify action object combination fragments that match the predefined action steps;

[0041] The type of composite action is determined based on the position of the action object combination fragment in the preset task sequence template;

[0042] If there is a combination of consecutive action objects that perfectly matches the predefined action steps in terms of type and time sequence, then the combination of consecutive action objects is merged to generate a standard composite action.

[0043] If there are partially matched but missing or additional action object combination fragments, then the action object combination fragments are corrected according to the preset fault tolerance rules to generate a corrected composite action.

[0044] The generated standard compound action or modified compound action, as well as the matching state between the action object combination fragment and the predefined action steps, are taken as the compound action.

[0045] Preferably, the step of storing the action record in a time-series database and establishing a data association index with the material information database, personnel information database, and process document database based on the operation object identifier and operator identifier includes:

[0046] Based on the action record, extract the timestamp, workstation identifier, operator identifier, composite action type, operation object identifier, actual action time, and sequence compliance status contained therein;

[0047] Based on the timestamp, the action records are written into the time-series database in chronological order, and a corresponding time-series data point is generated for each action record;

[0048] Based on the operation object identifier, obtain the corresponding material attribute information from the material information database;

[0049] Based on the material attribute information and the composite action type in the action record, a material action association mapping is generated;

[0050] Based on the operator's identifier, retrieve the corresponding personnel attribute information from the personnel information database;

[0051] Based on the personnel attribute information and the composite action type and actual action time in the action record, a personnel performance association mapping is generated.

[0052] Based on the workstation identifier and the type of compound action, obtain the corresponding standard operating procedure information from the process document library;

[0053] Based on the standard operating procedure information and the actual time consumption and sequence compliance status of the actions in the action records, a process compliance association mapping is generated.

[0054] The material action association mapping, personnel efficiency association mapping, and process compliance association mapping are established using the operation object identifier and operator identifier as association keys to create a multi-dimensional data association index.

[0055] This invention also provides a workstation motion decomposition, identification, and data management system, comprising:

[0056] The candidate action information acquisition module is used to acquire multimodal raw data of workers in the target workstation work area, and acquire multiple candidate action information of continuous work by workers based on the multimodal raw data.

[0057] The action type acquisition module is used to acquire the corresponding multimodal action features for each candidate action information, and to acquire the action type based on the multimodal action features;

[0058] The operation object identifier acquisition module is used to acquire the operation object identifier associated with the occurrence of the candidate action information according to the action type.

[0059] The composite action acquisition module is used to acquire the preset workstation standard operation sequence, and combine the action type with the operation object identifier based on the preset workstation standard operation sequence to obtain the composite action;

[0060] The data association index acquisition module is used to acquire action records based on the composite action, store the action records in the time-series database, and establish a data association index with the material information database, personnel information database, and process document database based on the operation object identifier and operator identifier.

[0061] Preferably, the candidate action information acquisition module includes:

[0062] The synchronous data stream acquisition unit is used to generate a time-aligned cross-modal synchronous data stream based on the timestamps of each modal data source in the multimodal raw data.

[0063] The motion vector sequence acquisition unit is used to extract the visual pose skeleton sequence, acoustic event spectrum sequence, and inertial motion vector sequence respectively based on the cross-modal synchronization data stream.

[0064] The motion signal point acquisition unit is used to acquire the motion trajectory of key joints according to the visual posture skeleton sequence, acquire significant acoustic event markers according to the acoustic event spectrum sequence, and acquire peak motion signal points according to the inertial motion vector sequence.

[0065] The action saliency interval acquisition unit is used to analyze and obtain multiple action saliency intervals based on the synchronous change of the key joint motion trajectory, significant acoustic event markers and peak motion signal points on the time axis.

[0066] The candidate action acquisition unit is used to extract corresponding data segments from the cross-modal synchronization data stream based on the start and end points of the salient intervals of each action, and obtain the information of the multiple candidate actions.

[0067] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of a workstation action decomposition, identification and data management method.

[0068] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of a workstation action decomposition, recognition, and data management method.

[0069] The beneficial effects of this application are as follows: This invention overcomes visual and light interference by acquiring multimodal raw data, identifying action types, associating operation object identifiers, combining compound actions, and establishing a data association index. It accurately decomposes and identifies workstation actions, reliably associates operation objects, and realizes structured data management. It can effectively solve the problems of inaccurate workstation action identification, missing operation object associations, and data fragmentation, and realize the integration of action decomposition and identification with structured data management, thereby improving the traceability and control efficiency of the production process. Attached Figure Description

[0070] Figure 1 This is a schematic diagram of a method flow according to an embodiment of this application.

[0071] Figure 2 This is a schematic diagram of the system structure according to an embodiment of this application.

[0072] Figure 3 This is a schematic diagram of the internal structure of a computer device according to an embodiment of this application.

[0073] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0074] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0075] like Figure 1 As shown, this application provides a method for workstation motion decomposition, identification, and data management, including:

[0076] S1. Obtain multimodal raw data of workers in the target work area, and obtain multiple candidate action information of continuous work based on the multimodal raw data;

[0077] S2. Obtain the corresponding multimodal action features for each candidate action information, and obtain the action type based on the multimodal action features;

[0078] S3. Obtain the operation object identifier associated with the occurrence of the candidate action information according to the action type;

[0079] S4. Obtain the preset workstation standard operation sequence, and combine the action type and the operation object identifier based on the preset workstation standard operation sequence to obtain the composite action;

[0080] S5. Obtain action records based on the composite action, store the action records in the time-series database, and establish a data association index with the material information database, personnel information database, and process document database based on the operation object identifier and operator identifier.

[0081] As described in steps S1-S5 above, this invention deploys independent visual sensors, acoustic sensors, and inertial sensors in the target workstation area to collect video streams, audio streams, and motion data, respectively. Then, a threshold-based method can be used, such as setting a video frame difference threshold, an audio energy threshold, or an inertial sensor acceleration change threshold. When any modal data exceeds a preset threshold, it is considered that an action may have occurred. Thus, the raw data of all modalities within that time period are extracted as candidate action information. For each candidate action information, keyframe image features can be extracted from the visual data, loudness or pitch of the sound can be extracted from the acoustic data, and the average acceleration value can be extracted from the inertial data. These single-modal features are then simply concatenated into a feature vector. Based on this feature vector, a pre-trained simple classifier can be compared with a template library containing a limited number of action categories to obtain the action type to which the candidate action information belongs. An image of the corresponding region is then extracted from the visual data corresponding to the candidate action information. Subsequently, a simple image recognition algorithm identifies the objects in the image and searches for the corresponding operation object identifier from a preset object category and identifier mapping table. For example, after recognizing the "pick up" action, a red square is identified on the screen and labeled "Part A". Alternatively, operators can be required to manually input or scan a barcode to input the identifier of the object being operated on into the system when performing a specific action. A standard work sequence for a specific workstation is pre-stored, describing the action name and possible object names for each step in text form. After recognizing the action type and object identifier, the system can simply concatenate them chronologically and perform string matching against the text description in the standard work sequence. If a match is successful, the matched action-object combination is marked as a composite action. After obtaining the composite action, the system can generate a text record containing a timestamp, workstation identifier, operator identifier, composite action type, and object identifier. This text record is then directly written into a general relational database, using the time of the composite action as the primary index. To establish data relationships, multiple independent tables can be created in a relational database, such as "Action Table," "Material Table," "Personnel Table," and "Process Table." By storing simple foreign key references for material IDs, personnel IDs, and process IDs in the "Action Table," relationships with other information databases can be achieved.

[0082] In one embodiment, the step of obtaining multiple candidate action information for continuous human operation based on the multimodal raw data includes:

[0083] S101. Generate a time-aligned cross-modal synchronous data stream based on the timestamps of each modal data source in the multimodal raw data.

[0084] S102. Based on the cross-modal synchronization data stream, extract the visual pose skeleton sequence, acoustic event spectrum sequence, and inertial motion vector sequence respectively;

[0085] S103. Obtain the motion trajectory of key joints according to the visual posture skeleton sequence, obtain significant acoustic event markers according to the acoustic event spectrum sequence, and obtain peak motion signal points according to the inertial motion vector sequence.

[0086] S104. Based on the synchronous changes of the key joint motion trajectory, significant acoustic event markers, and peak motion signal points on the time axis, multiple significant motion intervals are obtained through analysis.

[0087] S105. Based on the start and end points of each action salient interval, extract the corresponding data segments from the cross-modal synchronization data stream to obtain the multiple candidate action information.

[0088] As described in steps S101-S105 above, multimodal raw data refers to raw data collected from different sensors, such as video data from cameras, audio data from microphones, and accelerometer and gyroscope data from inertial measurement units (IMUs). These data describe the operational behavior of workers at their workstations from different dimensions. The timestamp is a marker recording the moment of data acquisition, typically based on a system clock or hardware clock. The time-aligned cross-modal synchronized data stream refers to precisely calibrating and synchronizing raw data from different modalities using their respective timestamps to form a continuous and consistent data sequence in the time dimension. The visual posture skeleton sequence refers to the sequence of spatial coordinates and their connections between key human body points (such as head, shoulders, elbows, wrists, etc.) continuously extracted from video data, changing over time. This sequence can intuitively reflect changes in a person's posture and limb movement trajectories. The acoustic event spectrum sequence refers to the sequence extracted from audio data, reflecting the energy distribution of sound events at different frequencies over time. This sequence can capture specific sounds generated during operation (such as tool knocking sounds, friction sounds) and their spectral characteristics. The inertial motion vector sequence refers to the sequence extracted from inertial measurement unit data, reflecting the changes in kinematic parameters such as acceleration and angular velocity of personnel or tools over time. It can accurately describe the dynamic characteristics of motion, such as intensity, speed, and direction. The key joint motion trajectory refers to the continuous motion path in three-dimensional space of a representative joint (e.g., wrist, finger, elbow) selected from the visual posture skeleton sequence for workstation operation. Its changes can indicate the start, end, or critical stage of the motion. The degree of synchronous change refers to the degree to which the three types of multimodal features—key joint motion trajectories, significant acoustic event markers, and peak motion signal points—show coordinated changes or mutual corroboration on the time axis. For example, when a hand joint begins to move rapidly, it may be accompanied by the impact sound of the tool and a peak acceleration on the inertial sensor. The start and end points refer to the precise boundaries on the time axis of each significant motion interval obtained through the above analysis. The cross-modal synchronous data stream refers to the original multimodal data after time alignment processing, containing the original information of all modalities, including visual, acoustic, and inertial. The candidate action information refers to segments containing complete multimodal raw data that are precisely extracted from the cross-modal synchronous data stream based on the start and end points of the action salient interval. These segments are the basis for subsequent action recognition and analysis.

[0089] This application effectively addresses the problems of inaccurate action boundary definition and unreliable action range extraction caused by time inconsistencies in multimodal data sources and environmental interference in complex industrial environments. Specifically, by generating time-aligned cross-modal synchronous data streams, it ensures high consistency of different modal data across time dimensions, laying a solid foundation for subsequent accurate analysis. Furthermore, it extracts visual posture skeleton sequences, acoustic event spectrum sequences, and inertial action vector sequences, comprehensively capturing action features from three complementary dimensions: visual, auditory, and kinematic, significantly enhancing the representation capability of complex and intricate actions.

[0090] In one embodiment, the step of obtaining the corresponding multimodal action features for each candidate action information and obtaining the action type based on the multimodal action features includes:

[0091] S201. Based on the candidate action information, separate the corresponding visual data segment, acoustic data segment, and inertial measurement data segment;

[0092] S202. Based on the visual data segment, extract the spatial coordinate set of joint points and the limb movement direction vector;

[0093] S203. Based on the acoustic data segment, extract the spectral energy distribution characteristics and the zero-crossing rate change profile;

[0094] S204. Based on the inertial measurement data segment, extract the triaxial acceleration amplitude envelope and angular velocity variation spectrum;

[0095] S205. Based on the joint point spatial coordinate set, limb movement direction vector, spectral energy distribution characteristics, zero-crossing rate change profile, triaxial acceleration amplitude envelope and angular velocity change spectrum, construct a unified perception feature vector for the candidate action information.

[0096] S206. Based on the real-time quality assessment results of each component sub-vector in the unified perception feature vector, obtain the dynamic contribution weight of each component sub-vector in the classification decision.

[0097] S207. The unified perception feature vector is weighted and integrated according to the dynamic contribution weight to generate an enhanced discriminative feature representation;

[0098] S208. Perform multi-dimensional matching between the enhanced discriminative feature representation and the preset action feature template library to obtain the action semantic category to which the candidate action information belongs.

[0099] As described in steps S201-S208 above, the present invention separates the corresponding visual data segments, acoustic data segments, and inertial measurement data segments based on the candidate action information. This aims to accurately segment and extract the visual, acoustic, and inertial measurement data associated with specific candidate action information from the original multimodal data. Its function is to provide clean, focused data input for subsequent modal feature extraction, avoiding interference from different actions or background noise. The joint point spatial coordinate set describes the position of key parts of the human body in three-dimensional space, while the limb movement direction vector reflects the direction and trend of movement of these joint points or limb segments. The spectral energy distribution characteristics describe the energy intensity of sound at different frequencies, reflecting attributes such as timbre and loudness, such as the sound of tools striking, rubbing, or machines operating. The zero-crossing rate change profile represents the number of times the sound wave signal crosses the zero axis per unit time, often used to distinguish between audible and silent segments, or to identify the roughness and frequency changes of sound. The step of extracting the triaxial acceleration amplitude envelope and angular velocity change spectrum based on the inertial measurement data segments aims to capture the dynamic characteristics of actions from the inertial measurement data. The triaxial acceleration amplitude envelope reflects the overall intensity and trend of linear acceleration during movement, effectively characterizing actions such as impact, vibration, or rapid movement. The angular velocity variation spectrum describes the velocity and frequency components of an object's rotational motion, playing a crucial role in identifying rotational, torsional, or swaying movements. The step of constructing a unified perceptual feature vector for candidate action information based on the joint point spatial coordinate set, limb movement direction vector, spectral energy distribution characteristics, zero-crossing rate variation profile, triaxial acceleration amplitude envelope, and angular velocity variation spectrum aims to integrate heterogeneous features extracted from different modalities to form a comprehensive, multi-dimensional feature representation. The unified perceptual feature vector comprehensively reflects the visual, auditory, and dynamic physical characteristics of the action, thus providing richer and more discriminative information for subsequent action classification. By evaluating the data quality of each modality in real time and dynamically adjusting its weight in the final classification decision, the negative impact of low-quality data on the recognition results can be effectively reduced. By comparing the enhanced discriminative feature representation with pre-stored standard action feature templates, the system can identify the specific action semantic category corresponding to the current candidate action information.

[0100] In one embodiment, the step of obtaining the operation object identifier associated with the occurrence of the candidate action information according to the action type includes:

[0101] S301. Obtain the target spatial region and time window for object recognition based on the action type;

[0102] S302. Extract region image blocks from the visual data segment corresponding to the candidate action information according to the target spatial region;

[0103] S303. Obtain image visual features based on the region image block, and identify the category of the operation object based on the image visual features;

[0104] S304. Extract acoustic event segments from the acoustic data segments corresponding to the candidate action information according to the time window;

[0105] S305. Obtain acoustic features based on the acoustic event fragments, and match sound patterns associated with the category of the operation object based on the acoustic features;

[0106] If both the image visual features and the acoustic features can identify the category of the operation object, then the operation object identifier is determined according to the preset object category and identifier mapping table;

[0107] If the category of the operation object can only be identified by one of the image visual features or acoustic features, then the confidence level is evaluated based on the identification result and the historical operation object behavior database, and the operation object identifier is determined based on the evaluation result.

[0108] If the category of the operation object cannot be identified through image visual features or acoustic features, then the electronic tag of the material or tool within the time window is read from the radio frequency identification data segment corresponding to the candidate action information, and the code corresponding to the electronic tag is used as the identifier of the operation object.

[0109] As described in steps S301-S305 above, this invention obtains the target spatial region and time window for object recognition based on the action type, aiming to narrow the scope of object recognition and improve recognition efficiency and accuracy. For example, for the "tightening" action, the system can preset that it usually occurs in a specific area of ​​the workbench, thereby limiting the focus of visual recognition to this area and accurately extracting the image portion that may be related to the target object from the original visual data. This can be achieved by applying a two-dimensional bounding box or mask corresponding to the target spatial region on the visual frame, thereby retaining only the pixel data within the bounding box and removing background interference. The image visual features can be traditional feature descriptors, such as scale-invariant feature transforms. These features can capture the local texture, shape, and edge information of the image and separate the sound events related to the interaction with the operation object from the continuous acoustic data. This can be achieved by accurately matching timestamps and extracting the audio waveform data within the corresponding time window from the original acoustic data stream. Acoustic features can include Mel-frequency cepstral coefficients (MFCC), short-time energy, zero-crossing rate, spectral centroid, etc., which can characterize the timbre, loudness, and frequency distribution of the sound. By matching these acoustic features with pre-established sound patterns associated with different categories of manipulated objects (e.g., the sounds of different tools striking different materials), the category of the manipulated object can be identified. If both the visual features and the acoustic features can identify the category of the manipulated object, then the object identifier is determined according to a preset object category and identifier mapping table. This means that when both visual and acoustic modalities provide clear and consistent object category identification results, the system will confirm the manipulated object with high confidence. The pre-defined object category and identifier mapping table is a database or lookup table that stores the correspondence between each identified object category and its unique operation object identifier. At this point, the system utilizes a historical operation object behavior database, which records the sequence, frequency, and contextual information of commonly used operation objects under specific action types. By comparing the current single-modal recognition result with historical data, its reasonableness and confidence level are evaluated. When neither visual nor acoustic recognition can provide effective information, the system activates Radio Frequency Identification (RFID) as a final fallback mechanism. This typically involves deploying RFID readers in the workstation area. When an electronic tag (such as an RFID tag) attached to a material or tool enters the reading range and is detected within a specified time window, its internally stored unique code is directly read. This code, as the most direct and reliable physical identifier, is determined as the operation object identifier.

[0110] In one embodiment, the step of obtaining a preset workstation standard operation sequence and combining the action type with the operation object identifier based on the preset workstation standard operation sequence to obtain a composite action includes:

[0111] S401. Obtain the corresponding preset job sequence template according to the current workstation identifier. The preset job sequence template includes predefined action steps arranged in chronological order and their corresponding operation object type constraints.

[0112] S402. Generate multi-level action matching rules based on time windows according to the time sequence of the predefined action steps and the constraints of the type of operation object.

[0113] S403. Based on the identified action type and operation object identifier, construct the current action object association sequence in chronological order;

[0114] S404. The current action object association sequence is compared step by step according to the multi-level action matching rule to identify the action object combination fragment that matches the predefined action step.

[0115] S405. Determine the corresponding composite action type based on the position of the action object combination fragment in the preset operation sequence template;

[0116] If there is a combination of consecutive action objects that perfectly matches the predefined action steps in terms of type and time sequence, then the combination of consecutive action objects is merged to generate a standard composite action.

[0117] If there are partially matched but missing or additional action object combination fragments, then the action object combination fragments are corrected according to the preset fault tolerance rules to generate a corrected composite action.

[0118] S406. The generated standard composite action or modified composite action, and the matching state between the action object combination fragment and the predefined action steps, are taken as composite actions.

[0119] As described in steps S401-S406 above, this invention retrieves the corresponding template file from a pre-stored standard work template library based on the input workstation identifier (e.g., workstation number, workstation name) through database query. It defines a state for each step according to the predefined time sequence of action steps and defines the transition conditions between states based on the operation object type constraint. For example, it can be defined that "tightening the screw" must follow "placing the washer," and the operation object must be a "screwdriver" and a "screw." The time window can be defined as the expected duration of each step or the maximum allowed interval between steps. A graph-based or sequence pattern mining method is employed. The predefined action steps and their constraints are represented as a directed graph or sequence pattern, where nodes represent action steps, and edges represent time sequence and operation object constraints. Matching rules can include exact matching, fuzzy matching, and sequential matching within a time window. Multi-level matching can refer to matching the main action first, then matching the secondary action or operation object, organizing the discrete action types and operation object identifiers identified in real time according to their chronological order to form a continuous data stream that can be compared with a standard sequence. A queue or list data structure is used. Whenever the system identifies an action type and its associated object identifier, it adds it as a tuple (action type, object identifier, timestamp) to the end of the sequence. Elements in the sequence are strictly ordered in ascending order of timestamps. The real-time constructed action sequence is compared with preset multi-level matching rules to find action-object combinations that match the standard operating procedure. Dynamic programming algorithms or sequence alignment algorithms (e.g., Levenshtein distance, variants of the Smith-Waterman algorithm) are used. The current action-object associated sequence is compared one-by-one or segmentally with predefined action steps in the preset operating sequence template. Matching scores are calculated, and high-scoring consecutive segments are identified. Based on the position of the identified action-object combination segment in the standard operating sequence, it is assigned a composite action label with business meaning. This is done through table lookup or mapping. Each predefined action step in the preset operating sequence template is associated with a unique composite action type identifier. When a combination of action objects is identified as matching a predefined action step in the template, the system directly retrieves the composite action type corresponding to that predefined action step. For action fragments that are completely consistent with the standard operation sequence, a clear and unbiased composite action record is generated. When the comparison results show that the current combination of action objects is completely consistent with the predefined action steps in terms of type and time sequence, the system marks the fragment as "standard" and encapsulates it with the corresponding composite action type, start time, end time, and other information into a standard composite action object. For action fragments with deviations (missing or additional actions), adjustments are made through a fault-tolerance mechanism to generate composite action records that conform to the standard as much as possible, reflecting subtle differences in actual operation. Edit distance is used.When a partial match is found, the edit distance between the current action object combination fragment and the predefined action steps is calculated, and a predetermined fault tolerance threshold is used to determine whether it can be corrected. Correction rules may include: ignoring a small number of irrelevant actions, inserting missing expected actions, replacing similar actions, etc. After correction, a composite action marked "corrected" is generated, and the standard composite action, the corrected composite action, and their matching status are output uniformly, providing a structured and traceable job record for subsequent data storage and analysis.

[0120] By introducing preset standard workstation sequence templates and multi-level action matching rules based on time windows, the system achieves accurate comparison between the real-time identified action types and operation object identification sequences and the standard work process, thus ensuring that the generation of composite actions conforms to the expected work specifications. Simultaneously, by handling partial matching cases through preset error-tolerant rules, it can identify and correct missing or additional actions that may occur in actual operation, significantly improving the adaptability of composite action recognition to complex industrial scenarios. The generated standard composite actions and corrected composite actions, along with their matching status, can more accurately and comprehensively reflect the completeness and compliance of the work, providing high-quality and highly reliable foundational data for subsequent data association indexing with material information databases, personnel information databases, and process document databases.

[0121] In one embodiment, the step of storing the action record in a time-series database and establishing a data association index with the material information database, personnel information database, and process document database based on the operation object identifier and operator identifier includes:

[0122] S501. Based on the action record, extract the timestamp, workstation identifier, operator identifier, composite action type, operation object identifier, actual action time, and sequence compliance status contained therein;

[0123] S502. Based on the timestamp, write the action records into the time-series database in chronological order, and generate corresponding time-series data points for each action record;

[0124] S503. Based on the operation object identifier, obtain the corresponding material attribute information from the material information database;

[0125] S504. Generate a material action association mapping based on the material attribute information and the composite action type in the action record;

[0126] S505. Based on the operator identifier, obtain the corresponding personnel attribute information from the personnel information database;

[0127] S506. Generate a personnel performance association mapping based on the personnel attribute information and the composite action type and actual action time in the action record;

[0128] S507. Based on the workstation identifier and the composite action type, obtain the corresponding standard operating procedure information from the process document library;

[0129] Based on the standard operating procedure information and the actual time consumption and sequence compliance status of the actions in the action records, a process compliance association mapping is generated.

[0130] S508. The material action association mapping, personnel efficiency association mapping, and process compliance association mapping are established using the operation object identifier and operator identifier as association keys to create a multi-dimensional data association index.

[0131] As described in steps S501-S508 above, this invention accurately extracts the attributes constituting the core elements from the identified action records. These attributes include the timestamp of the action, the workstation identifier, the operator identifier, the identified composite action type, the identifier of the involved operation object, the actual time consumed by the action, and the compliance status of the action sequence. The extracted action records are stored in an orderly manner in a time-series database based on their timestamps, and a unique time-series data point is generated for each record. The time-series database is optimized for processing data with time tags and can efficiently store, query, and analyze time-series data. Based on the operation object identifier contained in the action record, detailed material attribute information corresponding to the operation object is retrieved from the material information database. Subsequently, this material attribute information is combined with the composite action type in the action record to generate a material-action association mapping, linking the action record with the operator performing the action for personnel performance evaluation and management. Specifically, based on the operator identifier in the action record, the attribute information of the operator is obtained from the personnel information database. Then, combining this personnel attribute information, the composite action types in the action records, and the actual action time, a personnel efficiency correlation mapping is generated. The action records are then compared with preset process standards to assess the compliance of the operation. Specifically, based on the workstation identifier and composite action type in the action records, standard operating procedure (SOP) information related to the current operation is retrieved from the process document library. Subsequently, this SOP information is compared with the actual action time and sequence compliance status in the action records to generate a process compliance correlation mapping. Finally, this is crucial for achieving deep data correlation and breaking down data silos. It constructs a multi-dimensional data correlation index using the previously generated material action correlation mapping, personnel efficiency correlation mapping, and process compliance correlation mapping, with the operation object identifier and operator identifier as correlation keys.

[0132] like Figure 2 As shown, the present invention also provides a workstation action decomposition, identification, and data management system, including:

[0133] The candidate action information acquisition module is used to acquire multimodal raw data of workers in the target workstation work area, and acquire multiple candidate action information of continuous work by workers based on the multimodal raw data.

[0134] The action type acquisition module is used to acquire the corresponding multimodal action features for each candidate action information, and to acquire the action type based on the multimodal action features;

[0135] The operation object identifier acquisition module is used to acquire the operation object identifier associated with the occurrence of the candidate action information according to the action type.

[0136] The composite action acquisition module is used to acquire the preset workstation standard operation sequence, and combine the action type with the operation object identifier based on the preset workstation standard operation sequence to obtain the composite action;

[0137] The data association index acquisition module is used to acquire action records based on the composite action, store the action records in the time-series database, and establish a data association index with the material information database, personnel information database, and process document database based on the operation object identifier and operator identifier.

[0138] In one embodiment, the candidate action information acquisition module includes:

[0139] The synchronous data stream acquisition unit is used to generate a time-aligned cross-modal synchronous data stream based on the timestamps of each modal data source in the multimodal raw data.

[0140] The motion vector sequence acquisition unit is used to extract the visual pose skeleton sequence, acoustic event spectrum sequence, and inertial motion vector sequence respectively based on the cross-modal synchronization data stream.

[0141] The motion signal point acquisition unit is used to acquire the motion trajectory of key joints according to the visual posture skeleton sequence, acquire significant acoustic event markers according to the acoustic event spectrum sequence, and acquire peak motion signal points according to the inertial motion vector sequence.

[0142] The action saliency interval acquisition unit is used to analyze and obtain multiple action saliency intervals based on the synchronous change of the key joint motion trajectory, significant acoustic event markers and peak motion signal points on the time axis.

[0143] The candidate action acquisition unit is used to extract corresponding data segments from the cross-modal synchronization data stream based on the start and end points of the salient intervals of each action, and obtain the information of the multiple candidate actions.

[0144] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0145] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0146] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent results or equivalent process transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for decomposing, identifying, and managing workstation actions, characterized in that, include: Acquire multimodal raw data of workers in the target workstation area, and obtain multiple candidate action information of continuous work based on the multimodal raw data; For each candidate action information, obtain its corresponding multimodal action features, and obtain the action type based on the multimodal action features; Based on the action type, obtain the operation object identifier associated with the occurrence of the candidate action information; Obtain the preset workstation standard operation sequence, and combine the action type with the operation object identifier based on the preset workstation standard operation sequence to obtain the composite action; The action record is obtained based on the composite action, and the action record is stored in the time-series database. Based on the operation object identifier and the operator identifier, a data association index is established with the material information database, the personnel information database, and the process document database.

2. The method for decomposing, identifying, and managing workstation actions according to claim 1, characterized in that, The steps for obtaining multiple candidate action information for continuous operation of personnel based on the multimodal raw data include: Based on the timestamps of each modal data source in the original multimodal data, a time-aligned cross-modal synchronous data stream is generated; Based on the cross-modal synchronization data stream, the visual pose skeleton sequence, acoustic event spectrum sequence, and inertial motion vector sequence are extracted respectively. The motion trajectories of key joints are obtained based on the visual posture skeleton sequence, significant acoustic event markers are obtained based on the acoustic event spectrum sequence, and peak motion signal points are obtained based on the inertial motion vector sequence. Based on the synchronous changes of the key joint motion trajectories, significant acoustic event markers, and peak motion signal points on the time axis, multiple significant motion intervals are obtained through analysis. Based on the start and end points of each action salient interval, the corresponding data segments are extracted from the cross-modal synchronization data stream to obtain the information of the multiple candidate actions.

3. The method for decomposing, identifying, and managing workstation actions according to claim 1, characterized in that, The steps of obtaining the corresponding multimodal action features for each candidate action information and obtaining the action type based on the multimodal action features include: Based on the candidate action information, the corresponding visual data segment, acoustic data segment, and inertial measurement data segment are separated. Based on the visual data segment, extract the spatial coordinate set of joint points and the limb movement direction vector; Based on the acoustic data segment, extract the spectral energy distribution characteristics and the zero-crossing rate change profile; Based on the inertial measurement data segment, extract the triaxial acceleration amplitude envelope and angular velocity variation spectrum; Based on the joint point spatial coordinate set, limb movement direction vector, spectral energy distribution characteristics, zero-crossing rate change profile, triaxial acceleration amplitude envelope and angular velocity change spectrum, a unified perception feature vector of the candidate action information is constructed. Based on the real-time quality assessment results of each component sub-vector in the unified perception feature vector, the dynamic contribution weight of each component sub-vector in the classification decision is obtained. The unified perception feature vector is weighted and integrated according to the dynamic contribution weight to generate an enhanced discriminative feature representation. The action semantic category to which the candidate action information belongs is obtained by performing multi-dimensional matching between the enhanced discriminative feature representation and the preset action feature template library.

4. The method for decomposing, identifying, and managing workstation actions according to claim 1, characterized in that, The step of obtaining the operation object identifier associated with the occurrence of the candidate action information according to the action type includes: Based on the action type, obtain the target spatial region and time window for object recognition; Based on the target spatial region, extract region image blocks from the visual data segments corresponding to the candidate action information; Image visual features are obtained from the image blocks in the region, and the category of the object to be operated is identified based on the image visual features; Based on the time window, acoustic event segments are extracted from the acoustic data segments corresponding to the candidate action information; Acoustic features are obtained from the acoustic event fragments, and sound patterns associated with the operation object category are matched based on the acoustic features; If both the image visual features and the acoustic features can identify the category of the operation object, then the operation object identifier is determined according to the preset object category and identifier mapping table; If the category of the operation object can only be identified by one of the image visual features or acoustic features, then the confidence level is evaluated based on the identification result and the historical operation object behavior database, and the operation object identifier is determined based on the evaluation result. If the category of the operation object cannot be identified through image visual features or acoustic features, then the electronic tag of the material or tool within the time window is read from the radio frequency identification data segment corresponding to the candidate action information, and the code corresponding to the electronic tag is used as the identifier of the operation object.

5. The method for decomposing, identifying, and managing workstation actions according to claim 1, characterized in that, The steps of obtaining a preset workstation standard operation sequence and combining the action type with the operation object identifier based on the preset workstation standard operation sequence to obtain a composite action include: Obtain the corresponding preset job sequence template based on the current workstation identifier. The preset job sequence template includes predefined action steps arranged in chronological order and their corresponding operation object type constraints. Based on the time sequence of the predefined action steps and the constraints of the type of the operation object, generate multi-level action matching rules based on time windows; Based on the identified action type and operation object identifier, construct the current action object association sequence in chronological order; The current action object associated sequence is compared step by step according to the multi-level action matching rules to identify action object combination fragments that match the predefined action steps; The type of composite action is determined based on the position of the action object combination fragment in the preset task sequence template; If there is a combination of consecutive action objects that perfectly matches the predefined action steps in terms of type and time sequence, then the combination of consecutive action objects is merged to generate a standard composite action. If there are partially matched but missing or additional action object combination fragments, then the action object combination fragments are corrected according to the preset fault tolerance rules to generate a corrected composite action. The generated standard compound action or modified compound action, as well as the matching state between the action object combination fragment and the predefined action steps, are taken as the compound action.

6. The method for decomposing, identifying, and managing workstation actions according to claim 1, characterized in that, The steps of storing the action records in a time-series database and establishing a data association index with the material information database, personnel information database, and process document database based on the operation object identifier and operator identifier include: Based on the action record, extract the timestamp, workstation identifier, operator identifier, composite action type, operation object identifier, actual action time, and sequence compliance status contained therein; Based on the timestamp, the action records are written into the time-series database in chronological order, and a corresponding time-series data point is generated for each action record; Based on the operation object identifier, obtain the corresponding material attribute information from the material information database; Based on the material attribute information and the composite action type in the action record, a material action association mapping is generated; Based on the operator's identifier, retrieve the corresponding personnel attribute information from the personnel information database; Based on the personnel attribute information and the composite action type and actual action time in the action record, a personnel performance association mapping is generated. Based on the workstation identifier and the type of compound action, obtain the corresponding standard operating procedure information from the process document library; Based on the standard operating procedure information and the actual time consumption and sequence compliance status of the actions in the action records, a process compliance association mapping is generated. The material action association mapping, personnel efficiency association mapping, and process compliance association mapping are established using the operation object identifier and operator identifier as association keys to create a multi-dimensional data association index.

7. A workstation motion decomposition, recognition, and data management system, characterized in that, include: The candidate action information acquisition module is used to acquire multimodal raw data of workers in the target workstation work area, and acquire multiple candidate action information of continuous work by workers based on the multimodal raw data. The action type acquisition module is used to acquire the corresponding multimodal action features for each candidate action information, and to acquire the action type based on the multimodal action features; The operation object identifier acquisition module is used to acquire the operation object identifier associated with the occurrence of the candidate action information according to the action type. The composite action acquisition module is used to acquire the preset workstation standard operation sequence, and combine the action type with the operation object identifier based on the preset workstation standard operation sequence to obtain the composite action; The data association index acquisition module is used to acquire action records based on the composite action, store the action records in the time-series database, and establish a data association index with the material information database, personnel information database, and process document database based on the operation object identifier and operator identifier.

8. The workstation motion decomposition, identification, and data management system according to claim 7, characterized in that, The candidate action information acquisition module includes: The synchronous data stream acquisition unit is used to generate a time-aligned cross-modal synchronous data stream based on the timestamps of each modal data source in the multimodal raw data. The motion vector sequence acquisition unit is used to extract the visual pose skeleton sequence, acoustic event spectrum sequence, and inertial motion vector sequence respectively based on the cross-modal synchronization data stream. The motion signal point acquisition unit is used to acquire the motion trajectory of key joints according to the visual posture skeleton sequence, acquire significant acoustic event markers according to the acoustic event spectrum sequence, and acquire peak motion signal points according to the inertial motion vector sequence. The action saliency interval acquisition unit is used to analyze and obtain multiple action saliency intervals based on the synchronous change of the key joint motion trajectory, significant acoustic event markers and peak motion signal points on the time axis. The candidate action acquisition unit is used to extract corresponding data segments from the cross-modal synchronization data stream based on the start and end points of the salient intervals of each action, and obtain the information of the multiple candidate actions.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.