Fine-grained action classification and regression

The method addresses the lack of granularity in existing action classification by integrating pose and object datasets for precise action assessment, enhancing quality assurance in assembly processes and reducing reliance on skilled personnel.

HK40135071APending Publication Date: 2026-07-17HONG KONG APPLIED SCI & TECH RES INST

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Applications
Current Assignee / Owner
HONG KONG APPLIED SCI & TECH RES INST
Filing Date
2025-03-28
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing methods for human action classification and regression lack the necessary granularity to accurately assess compliance with specific procedural standards, particularly in scenarios requiring nuanced evaluations such as worker assembly actions or elderly motor skills, leading to inefficiencies and reliance on skilled personnel for quality assurance.

Method used

A method involving spatial-temporal video analysis that integrates pose and object datasets to generate a compound data structure, which is input into a machine learning model for fine-grained action classification and regression, enabling precise adherence to procedural standards.

Benefits of technology

Enables efficient, accurate classification and regression of human actions, reducing the need for skilled personnel and improving quality assurance by identifying deviations early in the assembly process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Methods and non-transitory computer-readable storage media for fine motion classification and / or regression are disclosed. The method includes: receiving a video stream capturing a series of human actions; identifying a reference object having a space-time relationship with the action sequence; extracting a gesture data set representing the action sequence; extracting an object data set representing the spatial position of the reference object; generating a composite data structure integrating the gesture data set and the object data set; and inputting the composite data structure into a trained machine learning model for classification and / or regression.
Need to check novelty before this filing date? Find Prior Art

Description

W O 2 02 6 / 08 59 00 A 1 I Hi ll Ill i II llll ll H ill H ill H ill III II I I II H ill H ill Ill i H ill H ill Il li l llll ll I lli Ill i II I (12) INTERNATIONAL APPLICATION PUBLISHED UNDER THE PATENT COOPERATION TREATY (PCT) (19) World Intellectual Property Organization International Bureau (43) International Publication Date 30 April 2026 (30.04.2026) (10) International Publication Number WO 2026 / 085900 Al WIPO!PCT (51) International Patent Classification: G06V 40 / 20 (2022.01) (21) International Application Number: PCT / CN2024 / 127978 (22) International Filing Date: 29 October 2024 (29.10.2024) (25) Filing Language: English (26) Publication Language: English (30) Priority Data: 18 / 922,855 22 October 2024 (22.10.2024) US (71) Applicant: HONG KONG APPLIED SCIENCE AND TECHNOLOGY RESEARCH INSTITUTE CO” LTD [CN / CN]; 5 / F, Photonics Centre, 2 Science Park East Av­ enue, Hong Kong Science Park, Shatin, Hong Kong (CN). (72) Inventors: CHOW, King Wai; Flat G, 21 / F, Tower 7, YOHO Town, YuenLong, New Territories, Hong Kong (CN). WONG, Chung Wai; Flat B 26 / F, Heng Tien Man­ sion, Taikoo Shing, Hong Kong (CN). (74) Agent: SHANGHAI SAVVY INTELLECTUAL PROP­ ERTY AGENCY; Unit 606, Shenergy International Build­ ing 1 Middle Fuxing Road, Huangpu District, Shanghai 200021 (CN). (81) Designated States (unless otherwise indicated, for every kind of national protection available) '. AE, AG, AL, AM, AO, AT, AU, AZ, BA, BB, BG, BH, BN, BR, BW, BY; BZ, CA, CH, CL, CN, CO, CR, CU, CV CZ, DE, DJ, DK, DM, DO, DZ, EC, EE, EG, ES, FI, GB, GD, GE, GH, GM, GT, HN, HR, HU, ID, IL, IN, IQ, IR, IS, IT, JM, JO, JP, KE, KG, KH, KN, KP, KR, KW, KZ, LA, LC, LK, LR, LS, LU, LY, MA, MD, MG, MK, MN, MU, MW, MX, MY, MZ, NA, NG, NI, NO, NZ, OM, PA, PE, PG, PH, PL, PT, QA, RO, RS, RU, RW, SA, SC, SD, SE, SG, SK, SL, ST, SV SY, TH, TJ, TM, TN, TR, IT, TZ, UA, UG, US, UZ, VC, VN, WS, ZA, ZM, ZW. (54) Title: FINE-GRAINED ACTION CLASSIFICATION AND REGRESSION 100 (57) Abstract: Methods and a non-transitorycomputer-readable storage medium for fine-grained action classification and / or regression are disclosed. The method includes: receiving a video stream capturing a sequence of human subject actions; identifying reference objects with spatial-temporal relationships to the action sequence; extracting a pose dataset representing the action sequence; extracting object datasets representing spatial positions of the reference objects; generating a compound data structure integrating the pose dataset and object datasets; and inputting the compound data structure into a trained machine learning model for classification and / or regression. [Continued on next page] WO 2026 / 085900 Al IIIIIIIIIM (84) Designated States (unless otherwise indicated, for every kind of regional protection available)'. ARIPO (BW, CV, GH, GM, KE, LR, LS, MW, MZ, NA, RW, SC, SD, SL, ST, SZ, TZ, UG, ZM, ZW), Eurasian (AM, AZ, BY, KG, KZ, RU, TJ, TM), European (AL, AT, BE, BG, CH, CY, CZ, DE, DK, EE, ES, FI, FR, GB, GR, HR,HU, IE, IS, IT, LT, LU, LV, MC, ME, MK, MT, NL, NO, PL, PT, RO, RS, SE, SI, SK, SM, TR), OAPI (BF, BJ, OF, CG, CI, CM, GA, GN, GQ, GW, KM, ML, MR, NE, SN, TD, TG). Published: — with international search report (Art. 21(3)) WO 2026 / 085900 PCT / CN2024 / 127978 FINE-GRAINED ACTION CLASSIFICATION AND REGRESSION FIELD OF THE INVENTION

[0001] .The present invention relates to video action recognition. Specifically, the present invention relates to fine-grained action classification and regression. BACKGROUND OF THE INVENTION

[0002] .Machine learning has revolutionized human action classification and assessment in recent years. Traditional approaches relied heavily on hand-crafted features and rule-based systems, which were often limited in their ability to generalize across diverse scenarios. With the advent of deep learning techniques, particularly Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs), researchers have developed more robust and accurate models. These networkscan automatically learn hierarchical features from raw input data, such as video frames or motion capture data, enabling them to classify complex human actions.

[0003] . Patent application No. US20210275107A1 discloses a computer-implemented method for human gait analysis extracts three-dimensional gait information from a video stream of an individual's walk. The three-dimensional gait information includes estimates of joint locations, including foot locations, on each frame. The method determines gait parameters based on foot locations in local extrema frames, providing a comprehensive understanding of the individual's gait.

[0004] ,Patent application No. US20220079472A1 discloses a fall-detection system detects personal falls while maintaining privacy by receiving a sequence of video images of a monitored person. The system processes each image, identifying the person and extracting a skeletal figure. The system then labels each figure with an action among predetermined actions,generating a fall / non-fall decision for the detected person. 1 WO 2026 / 085900 PCT / CN2024 / 127978

[0005] .Patent application No. US20240037977A1 discloses an apparatus consists of a joint-determination module, a pose estimation module, and an action-identification module. It analyzes an image containing one or more people using a computational neural network, derives pose estimates from these candidates, and analyzes a region of interest to identify an action.

[0006] .However, current methods face limitations when more nuanced evaluation is required. In scenarios where the degree of compliance or quality of specific actions needs assessment (such as worker assembly actions or elderly motor skills), a finer level of granularity is necessary. SUMMARY OF THE DESCRIPTION

[0007] .The invention addresses this need by introducing a spatial-temporal video dataset for fine-grained action classification and regression.

[0008] .One aspect of the embodiment of the present invention discloses a method ofimplementing fine-grained action classification and / or regression by a system comprising a processor and a non-transitory computer-readable storage medium storing instructions that, when executed, cause the system to perform the method comprising: receiving (SI 10) at least one video stream capturing a sequence of human subject actions; identifying (SI 40) at least one reference object with spatial-temporal relationships to the sequence of human subject actions; extracting (SI30) a pose dataset representing the sequence of human subject actions; extracting (S160) at least one object dataset representing the spatial positions of the at least one reference object; generating (S170) a compound data structure that integrates the pose dataset and the at least one object dataset; and inputting (SI80) the compound data structure as into a trained machine learning model for classification and / or regression.

[0009] . Another aspect of the embodiment of the present invention discloses a method ofimplementing quality prediction or compliance prediction by a system comprising a processor and a non-transitory computer-readable storage medium storing instructions that, when executed, cause the system to perform the method, comprising: receiving (S710) N 2 WO 2026 / 085900 PCT / CN2024 / 127978 compound data structures each generated according to the method of claim 1, wherein compound data structures 1 to N are each associated with a sequence of human actions which should comply with a set of specified procedural standards, and each of the sequences of human actions 1 to N is associated with an assembly portion of a final product; adding (S720) a timestamp from a global clock for each compound data structure, wherein each of the compound data structures includes timestamp information for each extracted frame; concatenating (S730) the N compound data structures according to their timestamp information to form a temporal sequence of data structures.

[0010] .Another aspect of the presentinvention provides a non-transitory computer- readable storage medium storing instructions that, when executed by one or more processors, cause a processing system to perform the method for fine-grained action classification and / or regression as disclosed herein. This computer-readable medium embodies the method in a form that can be directly utilized by computing devices to implement the invention’s functionalities. BRIEF DESCRIPTION OF THE DRAWINGS [001 l].The present invention is illustrated by way of example and not limitation in the FIG.s of the accompanying drawings in which like references indicate similar elements.

[0012] . FIG. 1 is a flow chart illustrating an exemplary method of an inference phase for fine-grained action classification and regression, according to an embodiment of the disclosure.

[0013] .FIGS. 2A and 2B are an exemplary schematic diagram depicting the generation of a compound data structure for fine-grained action classification and regression, according to anembodiment of the disclosure.

[0014] . FIG. 3 is a flow chart illustrating an exemplary method of a training phase for fine-grained action classification and regression, according to an embodiment of the disclosure. 3 WO 2026 / 085900 PCT / CN2024 / 127978

[0015] .FIG. 4 is a schematic diagram depicting an example scene for gait assessment, according to an embodiment of the disclosure.

[0016] .FIG. 5 is a flow chart illustrating an exemplary method of a training phase for gait assessment, according to an embodiment of the disclosure.

[0017] .FIG. 6 is a schematic diagram depicting an exemplary assembly line for production of ink cartridges, according to an embodiment of the disclosure.

[0018] .FIG. 7 is a flow chart illustrating an exemplary method of an inference phase for quality prediction of a final product, according to another embodiment of the disclosure.

[0019] .FIG.8 is a flow chart illustrating an exemplary method of an inference phase for compliance prediction of a final product,according to another embodiment of the disclosure. [0020J.FIG.9 is a flow chart illustrating an exemplary method of a training phase for compliance prediction of a final product, according to another embodiment of the disclosure. 4 WO 2026 / 085900 PCT / CN2024 / 127978 DETAILED DESCRIPTION

[0021] . Various embodiments and aspects of the inventions will be described with reference to details discussed below, and the accompanying drawings will illustrate the various embodiments. The following description and drawings are illustrative of the invention and are not to be construed as limiting the invention. Numerous specific details are described to provide a thorough understanding of various embodiments of the present invention. However, in certain instances, well-known or conventional details are not described in order to provide a concise discussion of embodiments of the present inventions.

[0022] . Reference in the specification to “one embodiment” or “an embodiment” or “another embodiment”means that a particular feature, structure, or characteristic described in conjunction with the embodiment can be included in at least one embodiment of the invention. The appearances of the phrase “in one embodiment” in various places in the specification do not necessarily all refer to the same embodiment.

[0023] .The first embodiment of this disclosure pertains to fine-grained action classification or assessment in the context of the assembly of printer ink cartridges on a factory production line.

[0024] .Fig. 6 is a schematic diagram depicting an exemplary assembly line for production of ink cartridges 630, according to an embodiment of the disclosure. The assembly line comprises Assembly Stations 610-1 through 610-N, where each assembly station is responsible for completing a portion of the printer cartridge assembly work in accordance with a predetermined workflow sequence. Workers at each assembly station are required to follow their respective standard operating procedures tocomplete the work at that particular station.

[0025] . Figs. 2A and 2B is an exemplary schematic diagram depicting the generation of a compound data structure for fine-grained action classification and regression, according to an embodiment of the disclosure. 5 WO 2026 / 085900 PCT / CN2024 / 127978

[0026] .Picture 220 in Figs. 2A and 2B illustrates one of Assembly Stations along the production line, which can be Assembly Station 610-1 shown in Fig.6.

[0027] . According to the standard operating procedure (SOP) for Assembly Station 610-1, workers are required to perform a series of intricate actions. One such critical task involves the worker using a handheld nozzle 202 to apply adhesive precisely to designated areas on each ink cartridge component 210-1 to 210-N.

[0028] . In conventional processes, ensuring adherence to the SOP across all assembly stations typically relies on downstream quality assurance (QA) procedures. These QA checks involve inspecting the fully assembled ink cartridges atthe end of the production line. For instance, QA personnel manually examine whether all cartridge components 210-1 to 210-N have been properly glued.

[0029] .This traditional approach, however, presents significant challenges. It requires QA staff to possess an in-depth understanding of how improperly glued components appear, which can be subtle and difficult to detect. This level of expertise is crucial for effectively identifying assembly errors, such as missed adhesive applications.

[0030] .The reliance on post-assembly QA checks not only demands highly skilled personnel but also introduces potential inefficiencies.

[0031] .To address the limitations of traditional quality control methods in assembly line operations, there is a need for analysis of worker actions through video footage. This approach aims to evaluate whether assembly procedures at each station adhere to the standard operating procedure, as any deviation could result in defective components in the final product.

[0032] .Existing technology, such as Vision Transformers, has been used for human action classification and assessment. This approach divides images into small pixel patches, which are then processed through a tokenization phase. After training on labeled video data, the model excels in two key areas: predicting action classes (like drinking water or brushing teeth) and assessing action quality (such as evaluating the correct form 6 WO 2026 / 085900 PCT / CN2024 / 127978 in physical exercises). However, the approach lacks the granularity needed to accurately classify or score actions based on specific procedural standards.

[0033] . According to the embodiment disclosed in this disclosure, a novel approach is proposed for fine-grained action classification and regression. This method leverages the spatial-temporal relationships of specific keypoints derived from human pose estimation, along with their interactions with selected reference objects in the surrounding environment.

[0034] .Specifically, according to the embodiments of this disclosure, the processes described herein with reference to the flowcharts can be implemented as computer programs. For example, the embodiments of this disclosure provide a computer program product that includes a computer program carried on a computer-readable medium, where the computer program contains program code for executing at least one step in the method embodiments of this disclosure.

[0035] .FIG. 1 is a flow chart 100 illustrating an exemplary method of an inference phase for fine-grained action classification and regression, according to an embodiment of the disclosure.

[0036] .In the embodiment of the disclosure, a method of implementing quality prediction is provided. This method is executed by a system comprising a processor and a non- transitory computer-readable storage medium storing instructions that, when executed, cause the system to perform the method.

[0037] . At Step SI 10, the system receives at least one videostream capturing a sequence of human subject actions. In this embodiment, the video streams are captured from one or more cameras located around Assembly Station 610-1 at predetermined intervals. This ensures comprehensive coverage of the area where the actions of human subject are taking place. The camera could be of type RGB, Infrared, or depth or any combination of the aforementioned three sensing modalities.

[0038] . At Step S120: human pose estimation is applied to a plurality of frames of the received video stream(s) to generate a human pose data stream. The human pose 7 WO 2026 / 085900 PCT / CN2024 / 127978 estimation task aims to first form a skeleton-based representation and then process it according to the needs of the final application. 2D and 3D pose estimation techniques are widely employed in the field of human pose analysis. 2D pose estimation involves detecting and localizing key body joints in image or video frames, typically representing them as a set of 2D coordinates (x,y) in the image plane. This approach is computationally efficient and works well for many applications, but lacks depth information.3D pose estimation, on the other hand, aims to recover the full 3D configuration of the human body, representing joint positions in a 3D coordinate system (x, y, z). The 2D / 3D coordinates of key body joints from multiple frames in a video sequence form a human pose data stream.

[0039] .In one embodiment of the disclosure, as shown in Fig.2B, the key body joints of a worker are identified in picture 220. These joints include, but are not limited to, the Right hand joint 201-1, Right hand joint 201-2, and Left hand joint 201-M. The human pose estimation process is applied to each frame of the received video stream(s), resulting in a human pose data stream that contains the 2D or 3D coordinates of these key body joints for each processed frame.

[0040] .For example, consider a worker assembling a batch of 5 cartridge prototypes at Assembly station 601. Thecomplete assembly cycle for this batch takes approximately 90 seconds. If the video of this process is captured at a standard rate of 30 frames per second, it would result in a total of 2700 image frames for the entire cycle (90 seconds * 30 frames / second = 2700 frames). Consequently, the human pose stream generated from this video would comprise 2700 human pose sets. Each of these sets contains the 2D or 3D coordinates for each of the identified key body joints, extracted from its corresponding frame.

[0041] . At Step S130: Extracting pose dataset from the human pose data stream. This pose dataset comprises a metadata segment and a plurality of data segments. The metadata segment may include, but is not limited to, the names of keypoints and the names of coordinate systems used. The data segments contain the actual 2D / 3D 8 WO 2026 / 085900 PCT / CN2024 / 127978 coordinates of keypoints, with the structure and meaning of these coordinates defined by the metadata segment.

[0042] . Pose dataset201 shown in Fig.2 listed pose dataset of a frame from the human pose data stream. The metadata segment from the pose dataset includes the names of keypoints: Right hand joint 1, Right hand joint 2 Left hand joint M, which correspond to the keypoints of the worker shown in the picture 220. The coordinate systems include: X- coordinate, Y-coordinate, and Z-coordinate, and confidence level. The confidence level for each keypoint of the human body, e.g., between 0.0 and 1.0, where 0.0 means no confidence or the key point is typically suppressed whereas 1.0 means almost certain that the key point is present. In pose dataset 201, the actual 2D / 3D coordinates of each key body joints are listed with the structure and meaning of these coordinates defined by the metadata segment.

[0043] . At Step S140: Identifying at least one reference object with spatial-temporal relationships to the sequence of human subject actions. These reference objects provide context for the human actions and arecrucial for accurate action classification and regression. In this embodiment, handheld nozzle 202 and ink cartridge component 210-1 to 210-N shown in picture 220 are identified as reference objects with bounding boxes respectively.

[0044] . At Step SI50: Applying domain-specific object detection and segmentation algorithm to the plurality of frames to generate at least one object dataset stream. Faster R-CNN, YOLO (You Only Look Once), SSD (Single Shot Detector) are commonly used for object detection, and U-Net, Mask R-CNN, DeepLab are commonly used for segmentation.

[0045] .Like the pose dataset, each object dataset comprises a metadata segment and a plurality of data segments. The metadata segment from the object dataset includes the names of keypoints of the object. As shown in picture 220 of Figs. 2A and 2B, for example, the keypoints of handheld nozzle 202, and ink cartridge components 210-1 to 210-N are Comer 1, Comer 2, Comer 3, Comer 4 of each of their bounding boxes. 9 WO2026 / 085900 PCT / CN2024 / 127978

[0046] . At Step S160: Extracting object dataset for each of reference objects. Object datasets 202, 210-1 to 210-N shown in Fig.2A are object datasets from the same frame as pose dataset 201. In the example described above regarding the worker assembling a batch of 5 cartridge prototypes at Assembly station 601 for a 90-second video stream, the object datasets comprise 2700 sets for the reference objects.

[0047] , At Step S170: Generating a compound data structure that integrates the pose dataset and the at least one object dataset. For each frame of the video stream, the compound data comprise 2D / 3D coordinates of the selected human body's keypoints and 2D / 3D coordinates of the keypoints for each of the identified one or more reference object, aligned according to coordinate system. Picture 230 in Fig. 2B schematically illustrates the compound dataset for one frame of the video stream. Datasets 201-M, 201- 1, and 201-2 contain the 2D or 3D coordinates ofthe Left hand joint 201-M, Right hand joint 201-1, and Right hand joint 201-2, respectively. Datasets 202 and 210 contain 2D or 3D coordinates of the comers of the bounding boxes of the handheld nozzle 202, and ink cartridge components 210-1 to 210-N.

[0048] .In the example described above regarding the worker assembling a batch of 5 cartridge prototypes at Assembly station 601 for a 90-second video stream, the compound data structure includes 2700 sets of human pose data and 2700 sets of object data. Alternatively, one could form a compound set constructed / mapped from the human and target object(s) for each image frame, resulting in a temporal sequence of 2700 compound sets

[0049] ..In one embodiment, a trainable mapping is applied to the compound data structure to generate a fused data structure. The weights of this mapping are learned during a training phase. Notably, the length of the fused data structure is smaller than the sum of the lengths of the pose dataset and the objectdatasets, allowing for more efficient processing.

[0050] . At Step SI80: The compound data structure (or the fused data structure) is input into a trained machine learning model for fine-grained action classification and regression. 10 WO 2026 / 085900 PCT / CN2024 / 127978

[0051] . At Step S190:The trained machine learning model outputs a classification result indicating whether the captured series of human actions comply with a set of specified procedural standards.

[0052] . Fig.3 is a flow chart 300 illustrating an exemplary method of a training phase for fine-grained action classification and regression, according to an embodiment of the disclosure.

[0053] . During the training phase of the ML model, domain expert(s) are required to perform Quality Assurance (QA) on product components that have passed through the Assembly station 610-1. If the QA process determines that there are issues with the products, they will label the corresponding products accordingly. The training set for the MLmodel is then created using the compound data structures (or fused data structures) associated with the labeled products, as identified through the Quality Assurance results. At Step S3 80, this labeled compound data structure by domain expert(s) or via QA’s Result is used to train the ML model, resulting in a trained ML model capable of classifying whether a product component is good or not good.

[0054] . Steps S310 to S370 in the ML model training phase method 300 of Fig. 3 are implemented similarly to steps SI 10 to S170 in the inference phase of method 100 in Fig. 1. For specific implementation details, please refer to the previous description of method 100 in Fig. 1. [(X)55]. Optionally, the first embodiment of this disclosure described above can be implemented in the context of the tokenization of the Al Transformer or similar framework. Specifically, the compound data structure or fused data structure can be in form of token for a transformer-based machine learning frame work.

[0056] . According to a second embodiment of this disclosure, the fine-grained action classification and regression method of this disclosure can be used for gait assessment. For example, Tinetti-POMA, which stands for Tinetti Performance Oriented Mobility Assessment, is a widely used tool to assess balance and gait in older adults. It's designed 11 WO 2026 / 085900 PCT / CN2024 / 127978 to evaluate a person's risk of falling by observing their performance in various mobility tasks. The test includes various activities such as: Balance tests: sitting balance, rising from a chair, standing balance (with eyes open and closed), and turning 360 degrees; Gait tests: initiation of gait, step length and height, step symmetry, step continuity, path deviation, trunk stability, and walking stance.

[0057] . There are multiple items to be tested throughout the Tinetti-POMA, each with its own scoring criteria. Each scoring item can have a score of 0 or 1. If the total score is too low, the individual wouldbe assessed as having a relatively high risk of falling in this test. For example, in a walking test for the gait assessment, there are the following four scoring items labeled A, B, C, and D. 12 WO 2026 / 085900 PCT / CN2024 / 127978

[0058] . The inference phase and training phase of the method of fine-grained action classification and regression illustrated in Figs. 1 and 3 can apply to gait assessment. Each test is considered an Assembly station in the first embodiment described above. A Step length right heel swings past left big toe = 1 left heel swings past right big toe = 1 B Foot clearance right foot completely clears floor = 1 left foot complete clears floor = 1 C Step symmetry right and left step length equal = 1 D Step continuity steps appear continuous = 1

[0059] . Fig. 4 is a schematic diagram 400 depicting an example scene for gait assessment, according to an embodiment of the disclosure.

[0060] . Fig. 4 illustrates an example of gait assessment by analyzing a video sequence of ahuman subject 401 for the walking test, according to an embodiment. In this example, the camera 403 is set up in front of the human subject's 401 path, recording a video sequence of the human subject walking towards the camera 403 and then turning to walk away from the camera 403

[0061] . During the training phase, a domain expert (physiotherapist / medical doctor) 402 produces a score after observing the “performance” of the human subject.

[0062] . Fig. 5 is a flow chart 500 illustrating an exemplary method of a training phase for gait assessment, according to an embodiment of the disclosure

[0063] . At Step S510, after receiving video sequences of the entire walking test from camera 403, human pose estimation is applied to multiple frames of the received video stream(s) at Step S520 to generate a human pose data stream. 2D and 3D pose estimation are used to obtain 2D / 3D coordinates of key body joints from multiple frames in a video sequence, forming a human pose data stream. In gaitassessments, joint points of the human subject, such as the patient's hands or feet, are key points that require special attention. 13 WO 2026 / 085900 PCT / CN2024 / 127978

[0064] .Next, at Step S530, the pose dataset is extracted from the human pose data stream. This includes a metadata segment containing the names of keypoints and coordinate systems used, and a data segment containing the actual 2D / 3D coordinates of key body joints. The structure and meaning of these coordinates are defined by the metadata segment.

[0065] . At Step S540, at least one reference object with spatial-temporal relationships to the sequence of human subject actions is identified. In this implementation, the reference object can be the ground plane of the floor with bounding box 410 in Fig. 4, or the bounding boxes of the armrests of the chair (not shown).

[0066] .Then at Step S55O, domain-specific object detection and segmentation algorithms are applied to multiple frames to generate at least one object dataset.For example, the generated object dataset may include a metadata segment naming the four comers (Corner 1, Comer 2, Comer 3, Comer 4) of the bounding box indicating the reference object in each frame, and a data segment indicating the 2D / 3D coordinates of these four corners respectively.

[0067] . At Step S560, the object dataset for the reference object is extracted.

[0068] . At Step S570, a compound data structure is generated that integrates the pose dataset and the object dataset. For each frame of the video stream, the compound data comprise 2D / 3D coordinates of the selected human body's keypoints and 2D / 3D coordinates of the keypoints for the reference object, aligned according to the coordinate system.

[0069] . At Step S580, scores obtained from the domain expert serve as the ground truth. Then at Step S590, the scores obtained at Step S580 and the compound data structure obtained at Step S570 are used as the training set to begin the training process of the machine learning model.

[0070] .Optionally, during the training phase, Step S591 can be used to fine-tune the model: optimizing and adjusting the model based on preliminary training results. Then at Step S592, it's determined whether the model has reached the expected performance 14 WO 2026 / 085900 PCT / CN2024 / 127978 level. If the model's performance is unsatisfactory, it returns to the training step for further training and tuning. If the model's performance is satisfactory, the training process is completed.

[0071] . Optionally, the Second embodiment of this disclosure described above can be implemented in the context of the tokenization of the Al Transformer or similar framework. Specifically, the compound data structure or fused data structure can be in form of token for a transformer-based machine learning frame work.

[0072] . The third embodiment of this disclosure pertains to quality prediction and compliance prediction for the final product in the context of the assembly of printer ink cartridges on afactory production line, as illustrated in Fig. 6.

[0073] .The quality of the final product, the printer ink cartridge, depends on whether the workers at each assembly station adhere to the corresponding standard operating procedures when working on specific parts of the printer cartridge. The final printer cartridge product, assembled through the process from Assembly Station 610-1 to Assembly Station N, will subsequently undergo a quality assurance (QA) process to inspect the assembled printer cartridges on the production line. Each cartridge after QA is classified as “Good” or “No Good”, which can be served as a label for training the ML model for quality prediction.

[0074] . In the event of a “No Good” quality assurance (QA) result, a Production or Industrial Engineer conducts a post-assembly analysis. This analysis serves two primary purposes. First, the engineer endeavors to identify the underlying cause of the product failure. Second, they trace the assembly process backwards todetermine at which specific assembly station or stations the error occurred. It's important to note the possibility that multiple assembly stations may have contributed to the product failure. The assembly station error measure could be in the form of probability from 0.0 to 1.0. The aforementioned analysis outcome can be served as the label for training the ML model for compliance prediction. 15 WO 2026 / 085900 PCT / CN2024 / 127978

[0075] . As illustrated in Fig. 1 and described in the related paragraphs, the videos of worker actions captured at each Assembly Station in the production of ink cartridges can generate a corresponding compound data structure that integrates the pose dataset and at least one object dataset.

[0076] . According to an embodiment of this disclosure, all the compound data structures related to the production of a final product, obtained from each assembly station on the production line, can be concatenated to generate a temporal sequence of data structure. Thetemporal sequence of data structure is unique for the final product and can be used to predict the quality of the final product or predict which portion(s) of the assembly process was non-compliant when the quality of the final product is not satisfied.

[0077] .Fig. 7 is a flow chart 700 illustrating an exemplary method of an inference phase for quality prediction of a final product, according to another embodiment of the disclosure.

[0078] . As shown in Fig.6, after the final printer cartridge 630 is assembled through the process from Assembly Station 610-1 to Assembly Station 610-N, all the video streams capturing a worker’s actions at each Assembly Station are processed as described in relation to steps from S110 to S170 of Fig. 1 to generate compound data structure 1 to compound data structure N. It can be understood that compound data structures 1 to N are each associated with workers’ actions at each Assembly Station 610-1 to Assembly Station 610-N respectively, as shown in Fig.6.

[0079] . In one embodiment of the disclosure, if conventional QA is NOT performed, a method of implementing quality prediction is provided. This method is executed by a system comprising a processor and a non-transitory computer-readable storage medium storing instructions that, when executed, cause the system to perform the method.

[0080] . At Step S710, the system for implementing quality prediction or compliance prediction receives the compound data structure 1 to compound data structure N. It's important to note that compound data structures 1 to N are each associated with the 16 WO 2026 / 085900 PCT / CN2024 / 127978 worker’s actions when working on specific parts of the printer cartridge at each of Assembly Station 610-1 to Assembly Station 610-N respectively.

[0081] .Following the reception of the compound data structures, at Step S720, the system adds a timestamp from a global clock for each compound data structure. It should be noted that each of the compound data structures alreadyincludes timestamp information for each extracted frame. This additional timestamp from the global clock provides a unified time reference across all compound data structures.

[0082] . At Step S730, the system concatenates the N compound data structures according to their timestamp information. This concatenation results in the formation of a temporal sequence of data structures. This temporal sequence provides a chronological representation of the assembly process for the final product.

[0083] . At Step S740, the system provides the temporal sequence of data structures as input to a trained machine learning model for quality prediction. The model has been previously trained with labeled QA classification results or final product of cartridge.

[0084] . Subsequently, at Step S750, the trained machine learning model for quality prediction outputs a classification result “Good” or “No Good” for the final product. This classification result predicts the quality of the final product based onthe analysis of the temporal sequence of data structures. The model can serve as pre-screen tool to predict product quality

[0085] . Fig. 8 is a flow chart 800 illustrating an exemplary method of an inference phase for compliance prediction of a final product, according to another embodiment of the disclosure.

[0086] . In this embodiment of the disclosure, if conventional QA is NOT performed, a method of implementing compliance prediction is provided.

[0087] ,The initial steps of this embodiment (S710, S720, S730) are identical to those described in the previous embodiment. These steps involve receiving N compound data 17 WO 2026 / 085900 PCT / CN2024 / 127978 structures, adding global timestamps, and concatenating the structures to form a temporal sequence. [(X)88].Following the formation of the temporal sequence of data structures, at Step S820, the system receives a result of Quality Assurance (QA) for the final product, i.e., “Good” or “No Good” for the final product.

[0089] .If the resultof the conventional QA indicates that the quality of the final product is unsatisfactory (i.e., a failure), the system proceeds with the following steps.

[0090] . At Step S830, the system provides two key inputs to a trained machine learning model for compliance prediction: a) The result of QA for the quality of the final product, and b) The temporal sequence of data structures (generated at Step S730).

[0091] . At Step S840, the trained machine learning model for compliance prediction processes the inputs and outputs an identification of which assembly portion(s) of the final product deviated from the specified procedural standards.

[0092] .The machine learning model for compliance prediction can predict where non- compliant steps / take place and the relevant corrective action can be administered.

[0093] . Fig.9 illustrates a flow chart 900 for a training phase of compliance prediction for a final product, according to the other embodiment of the present disclosure.

[0094] ,The trainingprocess begins with a human-driven step.

[0095] . At Step S910, an engineer performs a post-assembly analysis when a quality assurance (QA) result for a final product indicates a failure. This analysis involves a thorough examination of the failed product and its assembly process to identify potential causes of the failure.

[0096] .Following the post-assembly analysis, at Step S920, the engineer identifies one or more assembly portions potentially contributing to the failure. For each identified assembly portion, at Step S930, the engineer determines a probability of error contribution. This probability represents the likelihood that the particular assembly 18 WO 2026 / 085900 PCT / CN2024 / 127978 portion contributed to the product failure. The probability of error contribution is represented as a value ranging from 0.0 to 1.0. For example, a value of 0.0 would indicate that the assembly portion definitely did not contribute to the failure, and a value of 1.0 would indicate that the assemblyportion was certainly responsible for the failure.

[0097] .Finally, at Step S950, the machine learning model for compliance prediction is trained using the generated set of labeled training data. The model learns to predict potential assembly errors and their probabilities based on the input data.

[0098] . Optionally, the Third embodiment of this disclosure described above can be implemented in the context of the tokenization of the Al Transformer or similar framework. Specifically, the temporal sequence of data structures can be in form of token for a transformer-based machine learning frame work.

[0099] . It should be clear to those skilled in the art that, for the sake of convenience and brevity, the specific working processes of the systems, apparatus, devices, and modules described above can be referred to in the corresponding processes in the aforementioned method embodiments, and will not be repeated here.

[0100] .By studying the drawings, disclosure content, and the attachedclaims, those skilled in the art, when practicing the subject matter to be protected, can understand and implement variations of the disclosed embodiments. In the claims, the phrase "A and / or B" refers to A, B, or A and B; the word "includes" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude multiples. The words "first," "second," "third," "fourth" are merely used to distinguish elements or steps and do not indicate the order of elements or steps. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. 19 WO 2026 / 085900 PCT / CN2024 / 127978 CLAIMS 1. A method of implementing fine-grained action classification and / or regression by a system comprising a processor and a non-transitory computer-readable storage medium storing instructions that, when executed, cause the system to perform the method comprising: receiving (SI 10) atleast one video stream capturing a sequence of human subject actions; identifying (SI40) at least one reference object with spatial-temporal relationships to the sequence of human subject actions; extracting (SI 30) a pose dataset representing the sequence of human subject actions; extracting (SI60) at least one object dataset representing the spatial positions of the at least one reference object; generating (S170) a compound data structure that integrates the pose dataset and the at least one object dataset; and inputting (SI80) the compound data structure as into a trained machine learning model for classification and / or regression. 2. The method according to claim 1, further comprising: applying (S120) human pose estimation to a plurality of frames of at least one video stream to generate human pose data stream, and extracting the pose dataset from the human pose data stream, wherein the pose dataset comprises a metadata segment and a plurality of data segments; applying (SI50)domain-specific object detection to the plurality of frames to generate at least one object dataset, wherein the pose dataset comprises a metadata segment and a plurality of data segments; wherein the method further comprising: outputting, by the trained machine learning model for classification and / or regression, a classification result indicating whether the captured series of human actions comply with a set of specified procedural standards. 20 WO 2026 / 085900 PCT / CN2024 / 127978 3. The method according to claim 1, wherein the pose dataset comprises: for each extracted frame, 2D / 3D coordinates of the selected human body's keypoints, and a confidence level for each keypoint of the human body, wherein the confidence level is between 0.0 and 1.0. 4. The method according to claim 3, further comprises, adding a bounding box for each of the at least one reference object, and wherein for each extracted frame, the object dataset comprises : 2D / 3D coordinates of the keypoints of the boundingbox, and a confidence level for each keypoint of the reference objects. 5. The method according to claim 4, wherein generating the compound data structure comprises: for each extracted frame, combining the determined 2D / 3D coordinates of the selected human body's keypoints and the determined 2D / 3D coordinates of the keypoints for each of the identified one or more reference objects, according to the metadata segment. 6. The method according to claim 4, further comprises: applying a trainable mapping to the compound data structure to generate a fused data structure, wherein weights of the mapping are learned during a training phase; wherein the length of the fused data structure is smaller than the sum of the lengths of the pose dataset and the object datasets. 7. The method according to claim 1, wherein the at least one video streams are captured from one or more cameras located around a scene at predetermined intervals. 8. The method according to claim 1 or claim 6, furthercomprising, labeling, by domain expert(s) or through Quality Assurance result, for the compound data structure or the fused data structure, 21 WO 2026 / 085900 PCT / CN2024 / 127978 training the machine learning model to obtain the trained machine learning model for classification and / or regression. 9. A method of implementing quality prediction or compliance prediction by a system comprising a processor and a non-transitory computer-readable storage medium storing instructions that, when executed, cause the system to perform the method, comprising: receiving (S710) N compound data structures each generated according to the method of claim 1, wherein compound data structures 1 to N are each associated with a sequence of human actions which should comply with a set of specified procedural standards, and each of the sequences of human actions 1 to N is associated with an assembly portion of a final product; adding (S720) a timestamp from a global clock for each compound data structure, whereineach of the compound data structures includes timestamp information for each extracted frame; concatenating (S730) the N compound data structures according to their timestamp information to form a temporal sequence of data structures. 10. The method according to claim 9, further comprises: providing (S740) the temporal sequence of data structures as input to a trained machine learning model for quality prediction; and outputting (S750), by the trained machine learning model for quality prediction, a classification result that predicts the quality of the final product. 11. The method according to claim 9, further comprises: receiving (S820) a result of QA for the quality of the final product, if the result of conventional QA indicates that the quality of the final product is failure, providing (S830) the result of QA for the quality of the final product and the temporal sequence of data structures as input to a trained machine learning model for compliance prediction; 22 WO 2026 / 085900PCT / CN2024 / 127978 outputting (S840), by the trained machine learning model for compliance prediction, identification which assembly portion(s) of the final product that deviated from the specified procedural standards. 12. The method according to claim 11, further comprising training the machine learning model for compliance prediction, comprising: performing (S910), by an engineer, a post-assembly analysis when a quality assurance (QA) result for a final product result is failure; identifying (S920) one or more assembly portions potentially contributing to the failure; determining (S930), for each identified assembly portions, a probability of error contribution; generating (S940) a set of labeled training data based on the determined probabilities; and training (S950) the machine learning model for compliance prediction using the generated set of labeled training data to predict potential assembly errors and their probabilities. 13. The method of claim 12, wherein the probability oferror contribution is represented as a value ranging from 0.0 to 1.0. 14. A non-transitory computer-readable storage medium having stored therein instructions which, when executed by one or more processors of a processing system, causes the processing system to perform the method according to any one of claims 1-13. 23 WO 2026 / 085900 PCT / CN2024 / 127978 FIG. 1 1 / 10 WO 2026 / 085900 PCT / CN2024 / 127978 Pose dataset 201 200 \ :::' ■■" _ _ ,, yji $n ?WlK ;$ ? J J. ... II <? ji <? n ,x, Object 1 dataset 202 201-1 201-2 Object 2 dataset 210-1 KiYPs>i!>t Cssrsw i. C^r^r] Ccsnw 5 | 1 ■Cssfn&f 4 X-s.04sr<Sh^$x x^PTQ <?K4 J yJ^、'„PK3 case of e^TCl 1 Object N dataset 210-N ► 3 FIG. 2A 2 / 10 WO 2026 / 085900 PCT / CN2024 / 127978 230 201-M 201-1 201-2 202 210 FIG. 2B 3 / 10 WO 2026 / 085900 PCT / CN2024 / 127978 ^S390 ! Training the ML model for classification and regression FIG. 3 4 / 10 WO 2026 / 085900 PCT / CN2024 / 127978 400 FIG. 4 5 / 10 WO 2026 / 085900 PCT / CN2024 / 127978 S591 S593 FIG. 5 6 / 10 Ti m e WO 2026 / 085900PCT / CN2024 / 127978 620_N Compound data structure N 620 2 El 620 1 Compound data structure 2 Compound data structure 1 630 FIG. 6 7 / 10 WO 2026 / 085900 PCT / CN2024 / 127978 700 \ FIG. 7 8 / 10 WO 2026 / 085900 PCT / CN2024 / 127978 800 FIG. 8 9 / 10 WO 2026 / 085900 PCT / CN2024 / 127978 900 \ FIG. 9 10 / 10 INTERNATIONAL SEARCH REPORT International application No. PCT / CN2024 / 127978 A. CLASSIFICATION OF SUBJECT MATTER G06V40 / 20(2022.01)i According to International Patent Classification (IPC) or to both national classification and IPC B. FIELDS SEARCHED Minimum documentation searched (classification system followed by classification symbols) IPC: G06V, G06N Documentation searched other than minimum documentation to the extent that such documents are included in the fields searched Electronic data base consulted during the international search (name of data base and, where practicable, search terms used) VEN, CNABS, CNTXT, WOTXT, EPTXT, USTXT, CNKI, IEEE: body, motion, behavior, posture, pose, gait, reference,surroundings, association, relation, correlation, video, key point, joint, position, 2D, 3D, fusion, classification, regression, quality, machine learning, ML, training C. DOCUMENTS CONSIDERED TO BE RELEVANT Category* Citation of document, with indication, where appropriate, of the relevant passages Relevant to claim No. X CN 117765609 A (TONGJI UNIVERSITY) 26 March 2024 (2024-03-26) Description, paragraphs 0034-0042 1-14 A CN 114581613 A ( HANGZHOU YILAN TECHNOLOGY CO., LTD.) 03 June 2022 (2022-06-03) The whole document 1-14 A CN 115690874 A (SOUTH CHINA UNIVERSITY OF TECHNOLOGY, et al.) 03 February 2023 (2023-02-03) The whole document 1-14 A US 2022383639 Al (SPORTLOGIQ INC.) 01 December 2022 (2022-12-01) The whole document 1-14 A US 2024252263 Al (DIGITAL SURGERY LIMITED) 01 August 2024 (2024-08-01) The whole document 1-14 I | Further documents are listed in the continuation of Box C. See patent family annex. * Special categories of cited documents: “T” later document publishedafter the international filing date or priority “A" document defining the general state of the art which is not considered date and not in conflict with the application but cited to understand the to be of particular relevance principle or theory underlying the invention "D” document cited by the applicant in the interaational application “X” document of particular relevance; the claimed invention cannot be “E” earlier application or patent but pubUshed on or after the international ctmsidered novel or cannot be considered to involve an inventive step filing dote when the document is token Hlone •>L” document which may throw doubts on priority claim(s) or which is “Y” document of particular relevance; the claimed invention cannot be cited to establish the publication date of another citation or other considered to involve an inventive step when the document is special reason (as specified) combined with one or more other such documents, such combination “0” document referring to anoral disclosure, use, exhibition or other being obvious to a person skilled in the art means document member of the same patent family “P” document published prior to the international filing date but later than the priority date claimed Date of the actual completion of the international search 13 July 2025 Date of mailing of the international search report 17 July 2025 Name and mailing address of the ISA / CN CHINA NATIONAL INTELLECTUAL PROPERTY ADMINISTRATION 6, Xitucheng Rd., Jimen Bridge, Haidian District, Beijing 100088, China Authorized officer WEIJiaLi Telephone No. (+86) 010-53961398 Form PCT / ISA / 210 (second sheet) (July 2022) INTERNATIONAL SEARCH REPORT Information on patent family members International application No. PCT / CN2024 / 127978 Patent document cited in search report Publication date (day / month>year) Patent family member(s) Publication date (day / month / year) CN 117765609 A 26 March 2024 None CN 114581613 A 03 June 2022 None CN 115690874 A 03 February 2023 None US2022383639 Al Dec 01, 2022 EP 4085374 Al Nov 09, 2022 EP 4085374 A4 Jan 17, 2024 WO 2021189145 Al Sep 30, 2021 CA 3167079 Al Sep 30, 2021 US 2024252263 Al Aug 01, 2024 WO 2022263430 Al Dec 22, 2022 EP 4355247 Al Apr 24, 2024 Form PCT / ISA / 210 (patent family annex) (July 2022) (19) State Intellectual Property Office of the People's Republic of China (12) Invention Patent Application (10) Publication Number of Application CN 119604907 A (43) Publication Date Mar 11, 2025 (21) Application Number 202480002727.5 (22) Application Date Oct 29, 2024 (30) Priority Data 18 / 922,855 Oct 22, 2024 US (85) Date of Entry of PCT International Application into the National Phase Nov 25, 2024 (51) Int.CI. G06V 40 / 20(2022.01) G06V 70 / 764(2022.01) G06V 70 / 72(2022.01) G06V 70 / 766(2022.01) G06V 70 / 774(2022.01) G06N 20 / 00 (2019.01) (86) Application Data of PCT International Application PCT / CN2024 / 127978 Oct 29, 2024 (71) Applicant Hong Kong Applied Science and Technology Research Institute Co., Ltd. Address 5 / F, Optoelectronics Centre, 2 Science Park East, Technology Drive, Shatin, Hong Kong, China (72) Inventors Zhou Jingwei; Wang Zhongwei (74) Patent Agency Shanghai Siwei Intellectual Property Agency (General Partnership) 31237 Patent Attorney Gu Danli Number of Claims 2 Number of Pages of Specification 9 Number of Pages of Drawings 10 (54) Invention Title Fine Motion Classification and Regression (57) Abstract Methods and non-transitory computer-readable storage media for fine motion classification and / or regression are disclosed. The method includes: receiving a video stream capturing a series of human actions; identifying a reference object having a spatio-temporal relationship with the action sequence; extracting a pose data set representing the action sequence; extracting an object data set representing the spatial position of the reference object; generating a composite data structure integrating the pose data set and the object data set; and inputting the composite data structure into a trained machine learning model for classification and / or regression.Receiving one or more video streams ~~ Identifying reference objects, Human pose estimation, Specific domain object detection { Extracting a pose data set, Extracting an object data set. Generating a composite data structure, Generating an ML model for classification and regression sim 1 | Ejecting who V Z06S96I T—I 6 CN 119604907 A Claims 1 / 2 pages 1. A method for implementing fine-grained action classification and / or regression using a system comprising a processor and a non-transitory computer-readable storage medium storing instructions, wherein the instructions, when executed, cause the system to perform a method comprising the following steps: Receiving (S110) at least one video stream that captures a sequence of human subject actions; Identifying (S140) at least one reference object having a spatio-temporal relationship with the sequence of human subject actions; Extracting (S130) a pose data set representing the sequence of human subject actions; Extracting (S160) at least one object data set representing the spatial positions of the at least one reference object; Generating (S170) a composite data structure integrating the pose data set and the at least one object data set; and Inputting (S180) the composite data structure into a trained machine learning model for classification and / or regression. 2. The method according to claim 1, further comprising: Applying human pose estimation (S120) to multiple frames of the at least one video stream to generate a human pose data stream, and extracting the pose data set from the human pose data stream, wherein the pose data set includes a metadata segment and multiple data segments; Applying (S150) specific domain object detection to the multiple frames to generate at least one object data set, wherein the object data set includes a metadata segment and multiple data segments; wherein the method further comprises: Outputting a classification result by the trained machine learning model for classification and / or regression, the classification result indicating whether the captured sequence of human actions conforms to a set of specified procedural criteria. 3. The method according to claim 1, wherein the pose data set includes: For each extracted frame, the 2D / 3D coordinates of selected human key points, and the confidence of each of the human key points, wherein the confidence is between 0.0 and 1.0. 4. The method according to claim 3, further comprising adding a bounding box for each of the at least one reference object, wherein, for each extracted frame, the object data set includes: The 2D / 3D coordinates of the key points of the bounding box, and the confidence of each key point of the reference object. 5. The method according to claim 4, wherein generating the composite data structure comprises: Combining the determined 2D / 3D coordinates of the selected human key points with the determined 2D / 3D coordinates of the key points of the identified one or more reference objects according to the metadata segment. 6.The method of claim 4, further comprising: applying a trainable mapping to the composite data structure to generate a fused data structure, wherein the weights of the mapping are learned during a training phase; wherein the length of the fused data structure is less than the sum of the lengths of the pose dataset and the object dataset. 7. The method of claim 1, wherein the at least one video stream is captured from one or more cameras located around the scene at predetermined time intervals. 8. The method of claim 1 or claim 6, further comprising: labeling the composite data structure or the fused data structure by a domain expert or through quality assurance results; training the machine learning model to obtain the trained machine learning model for classification and / or regression. 9. A method for implementing quality prediction or compliance prediction using a system comprising a processor and a non-transitory computer-readable storage medium; the storage medium storing instructions, when executed, causing the system to perform the method, comprising: receiving (S710) N composite data structures generated by the method of claim 1, wherein composite data structures 1 to N are respectively associated with sequences of human actions to conform to a set of specified procedural standards, and each of the human action sequences 1 to N is associated with an assembly portion of a final product; adding (S720) a timestamp from a global clock to each composite data structure, wherein each composite data structure includes timestamp information for each extracted frame; concatenating (S730) the N composite data structures according to the timestamp information of the N composite data structures to form a time series of data structures. 10. The method of claim 9, further comprising: providing (S740) the time series of the data structures as input for quality prediction of a trained machine learning model; and outputting (S750) a classification result predicting the quality of the final product from the trained machine learning model. 11. The method of claim 9, further comprising: receiving (S820) a quality assurance result of the quality of the final product; if the result of conventional quality assurance indicates that the quality of the final product is non-conforming, providing (S830) the quality assurance result of the quality of the final product and the time series of the data structure as input to a trained machine learning model for compliance prediction; and outputting (S840) from the trained machine learning model for compliance prediction indicating which assembly parts in the final product deviate from the specified procedure standard. 12. The method of claim 11, further comprising training the machine learning model for compliance prediction, including: when the quality assurance result of the final product is non-conforming, having an engineer perform (S910) a post-assembly analysis; identifying (S920)...One or more assembly parts that may cause the defect; determining (S930) the error contribution probability of each identified assembly part; generating (S940) a set of labeled training data based on the determined probability; and training (S950) the machine learning model using the generated set of labeled training data to predict potential assembly errors and their probabilities. 13. The method of claim 12, wherein the error contribution probability is expressed as a value between 0.0 and 1.0. 14. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors of a processing system, cause the processing system to perform the method of any one of claims 1-13. 3 CN 119604907 A Specification 1 / 9 pages Fine-grained motion classification and regression technical field

[0001] The present invention relates to video motion recognition. In particular, the present invention relates to fine-grained motion classification and regression. Background Art

[0002] In recent years, machine learning has revolutionized the classification and evaluation of human motion. Traditional methods rely heavily on handcrafted features and rule-based systems, which often have limited ability to generalize across scenarios. With the advent of deep learning technologies, particularly convolutional neural networks (CNNs) and recurrent neural networks (RNNs), researchers have developed more powerful and accurate models. These networks can automatically learn hierarchical features from raw input data (such as video frames or motion capture data) to classify complex human movements.

[0003] Patent application number US20210275107A1 discloses a computer-implemented method for human gait analysis that extracts three-dimensional gait information from a video stream of an individual walking. The three-dimensional gait information includes estimates of joint positions (including foot positions) on each frame. The method determines gait parameters based on foot positions in local extreme frames, thereby gaining a comprehensive understanding of an individual's gait.

[0004] Patent application number US20220079472A1 discloses a fall detection system that detects individual falls by receiving a sequence of video images of a monitored person while protecting privacy. The system processes each image, identifies the monitored person, and extracts a skeletal map. Then, the system adds a label to each action in the predetermined actions, thereby generating a fall / non-fall judgment for the monitored person.

[10005] Patent application number US20240037977A1 ​​discloses an apparatus consisting of a joint judgment module, a posture estimation module, and an action recognition module. It uses a computational neural network to analyze images containing one or more people, derives posture estimates from these candidates, and analyzes regions of interest to identify actions.

[0006] However, existing methods face limitations when more refined evaluation is required. When it is necessary to evaluate specific actions...In cases of varying degrees of accuracy or quality (e.g., worker assembly movements or elderly motor skills), finer granularity is required.

[0007] Therefore, the present invention addresses this need by introducing a spatial-temporal video dataset for fine-grained movement classification and regression.

[0008] One aspect of the present invention discloses a method for fine-grained movement classification and / or regression implemented through a system comprising a processor and a non-volatile computer-readable storage medium. The storage medium stores instructions that, when executed, cause the system to perform a method comprising the steps of: receiving (S110) at least one video stream that captures a sequence of human subject movements; identifying (S140) at least one reference object having a spatiotemporal relationship with the sequence of human subject movements; extracting (S130) a pose dataset representing the sequence of human subject movements; extracting (S160) at least one object dataset representing the spatial location of the at least one reference object; generating (S170) a composite data structure integrating the pose dataset and the at least one object dataset; and inputting (S180) the composite data structure into a trained machine learning model for classification and / or regression.

[0009] Another aspect of the present invention discloses a method for quality prediction or compliance prediction using a system comprising a processor and a non-volatile computer-readable storage medium. The storage medium stores instructions that, when executed, cause the system to perform the method, comprising: receiving (S710) N composite data structures generated by the method according to claim 1, wherein composite data structures 1 to N are respectively associated with sequences of human actions that should conform to a set of specified procedural standards, and each of the human action sequences 1 to N is associated with an assembly portion of a final product; adding (S720) a timestamp from a global clock to each composite data structure, wherein each composite data structure includes timestamp information for each extracted frame; concatenating (S730) the N composite data structures according to the timestamp information of the composite data structures to form a time series of data structures.

[0010] Another aspect of the present invention provides a non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause a processing system to perform the methods disclosed herein for fine-grained action classification and / or regression. The computer-readable medium embodies the method in a form that can be directly utilized by a computing device to achieve the functions of the invention. Brief Description of the Drawings

[0011] The invention will now be described by way of example with reference to the accompanying drawings, which are not intended to be limiting. Similar reference numerals in the drawings denote similar parts.

[0012] FIG1 is a flowchart of an exemplary method for the inference phase of fine-grained action classification and regression according to an embodiment of the present disclosure.

[0013] Figures 2A and 2B are exemplary schematic diagrams illustrating the generation of a composite data structure for fine-grained action classification and regression according to an embodiment of the present disclosure.

[0014] Figure 3 is a flowchart illustrating an exemplary method for a training phase for fine-grained action classification and regression according to an embodiment of the present disclosure.

[0015] Figure 4 is a schematic diagram illustrating an example scenario of gait evaluation according to an embodiment of the present disclosure.

[0016] Figure 5 is a flowchart illustrating an example method of gait evaluation according to an embodiment of the present disclosure.

[0017] Figure 6 is a schematic diagram illustrating an exemplary assembly line for producing ink cartridges according to an embodiment of the present disclosure.

[0018] Figure 7 is a flowchart illustrating an exemplary method for a final product quality prediction inference phase according to an embodiment of the present disclosure.

[0019] Figure 8 is a flowchart illustrating an exemplary method for a final product compliance prediction inference phase according to an embodiment of the present disclosure.

[0020] Figure 9 is a flowchart illustrating an exemplary method for a final product compliance prediction training phase according to an embodiment of the present disclosure. Detailed Description [(X)21] Various embodiments and aspects of the present invention will be discussed in detail below, while the accompanying drawings will illustrate various embodiments. The following description and figures are used to illustrate the invention, but should not be construed as limiting the invention. Many specific details are described in order to provide a thorough understanding of various embodiments of the invention. However, in some cases, well-known or conventional details are not described in order to briefly discuss embodiments of the invention.

[0022] In the specification, "one embodiment" or "an embodiment" or "another embodiment" means that a particular feature, structure or characteristic described together with the embodiment may be included in at least one embodiment of the invention. The phrase "in one embodiment" appearing in multiple places in the specification does not necessarily refer to the same embodiment.

[0023] A first embodiment of the present disclosure relates to the fine movements performed when assembling printer cartridges on a factory production line, classified or evaluated.

[0024] FIG6 is a schematic diagram of an exemplary assembly line for producing cartridge 630 according to an embodiment of the present disclosure. The assembly line includes assembly stations 610-1 to 610-N, each of which is responsible for completing a portion of the printer cartridge assembly work in a predetermined workflow sequence. Workers at each assembly station must complete the work of that particular station in accordance with their respective standard operating procedures.

[0025] Figures 2A and 2B are exemplary schematic diagrams illustrating the generation of a composite data structure for fine-grained action classification and regression according to an embodiment of the present disclosure.

[0026] Sub-figure 220 in Figures 2A and 2B shows an assembly station on a production line, which may be assembly station 610-1 shown in Figure 6.

[0027] According to the Standard Operating Procedure (SOP) of assembly station 610-1, workers need to perform a series of complex operations. One key task involves workers using a handheld nozzle 202 to precisely apply adhesive to designated areas of each cartridge assembly 210-1 to 210-N.

[0028] In traditional processes, ensuring that all assembly stations adhere to the SOP typically relies on downstream quality assurance (QA) procedures. These QA checks include inspecting fully assembled cartridges at the end of the production line. For example, QA personnel manually check that all cartridge assemblies 210-1 to 210-N are properly bonded.

[0029] However, this traditional approach presents significant challenges. It requires QA personnel to have a deep understanding of the appearance of improperly bonded components, which can be subtle and difficult to detect. This expertise is crucial for effectively identifying assembly errors, such as missed adhesive application.

[0030] Reliance on post-assembly QA checks not only requires highly skilled personnel but can also lead to inefficiencies.

[0031] To address the limitations of traditional quality control methods in assembly line operations, it is necessary to analyze worker behavior through video clips. This method aims to assess whether the assembly process at each workstation adheres to standard operating procedures, as any deviation can lead to defects in the final product.

[0032] Existing technologies, such as visual transformers, have been used for human behavior classification and evaluation. This method divides images into small pixel blocks and then processes them through a labeling stage. After training on labeled video data, the model performs well in two key areas: predicting action categories (such as drinking water or brushing teeth) and evaluating action quality (such as assessing the correct form of physical exercise). However, this method lacks the granularity required for precise classification or scoring of actions according to specific procedural criteria.

[0033] According to embodiments disclosed in this disclosure, a novel method for fine-grained action classification and regression is proposed. This method utilizes the spatiotemporal relationships of specific key points derived from human pose estimation, as well as their interactions with selected reference objects in the surrounding environment.

[0034] Specifically, according to embodiments of this disclosure, the processes described herein with reference to the flowchart can be implemented as a computer program. For example, embodiments of this disclosure provide a computer program product comprising a computer program carried on a computer-readable medium, wherein the computer program includes program code for performing at least one step in embodiments of this disclosure.

[0035] FIG1 is a flowchart 100. According to embodiments of this disclosure, flowchart 100 illustrates an exemplary method for the inference phase of fine-grained action classification and regression.

[0036] In embodiments of this disclosure, a method for implementing quality prediction is provided. The method is performed by a system including a processor and a non-transitory computer-readable storage medium storing instructions that, when executed, perform the method described above.

[0037] In step S110, the system receives at least one video stream that captures a series of movements of a human subject. In this embodiment, the video stream is captured at predetermined time intervals from one or more cameras located around assembly station 610-1. This ensures comprehensive coverage of the area where the human subject's movements occur. The cameras can be RGB, infrared, depth, or any combination of the above three sensing modes.

[0038] In step S120, human pose estimation is applied to multiple frames of the received video stream to generate a human pose data stream. The task of human pose estimation is to first form a skeleton-based representation, and then process it according to the needs of the final application. 2D and 3D pose estimation techniques are widely used in the field of human pose analysis. 2D pose estimation involves detecting and locating key body joints in an image or video frame, typically represented as a set of 2D coordinates (x, y) on the image plane. This method is computationally efficient and suitable for many applications, but lacks depth information. 3D pose estimation aims to recover the complete 3D configuration of the human body, representing joint positions in a 3D coordinate system (x, y, z). The 2D / 3D coordinates of key body joints in multiple frames of a video sequence form a human pose data stream.

[0039] In one embodiment of this disclosure shown in FIG2B, key body joints of the worker are identified in sub-Figure 220. These joints include, but are not limited to, right hand joint 201-1, right hand joint 201-2, and left hand joint 201-Mo. The human pose estimation process is applied to each frame of the received video stream, thereby generating a human pose data stream containing the 2D or 3D coordinates of these key body joints for each processed frame.

[0040] For example, suppose a worker assembles a batch of 5 ink cartridge prototypes at assembly station 601. The complete assembly cycle for this batch of products takes approximately 90 seconds. If this process is videotaped at a standard rate of 30 frames per second, a total of 2700 image frames will be generated for the entire cycle (90 seconds * 30 frames / second = 2700 frames). Therefore, the human pose stream generated from the video will contain 2700 human pose sets. Each pose set contains the 2D or 3D coordinates of each identified key body joint extracted from its corresponding frame.

[0041] In step S130, a pose dataset is extracted from the human pose data stream. The pose dataset includes a metadata segment and multiple data segments. The metadata segment may include, but is not limited to, the names of key points and the names of the coordinate systems used. The data segments contain the actual 2D / 3D coordinates of the key points, the structure and meaning of which are defined by the metadata segment.

[0042] The pose dataset 201 shown in Figure 2 lists the pose dataset of a frame in the human pose data stream. The metadata part of the pose dataset includes the names of key points: right hand joint 1, right hand joint 2... left hand joint M, respectively corresponding toThe keypoint coordinate system of the worker shown in subfigure 220 includes: X coordinate, Y coordinate, and Z coordinate, as well as confidence. The confidence of each keypoint of the human body is, for example, between 0.0 and 1.0, where 0.0 indicates no confidence or the keypoint is generally suppressed, while 1.0 indicates that the keypoint is almost certain to exist. In the pose dataset 201, the actual 2D / 3D coordinates of each key body joint are listed, and the structure and meaning of these coordinates are defined by the metadata segment.

[0043] In step S140, at least one reference object that has a spatiotemporal relationship with the sequence of human subject movements is identified. These reference objects provide context for human movements and are essential for accurate movement classification and regression. In this embodiment, the handheld nozzle 202 and the cartridge assemblies 210-1 to 210-N shown in FIG. 220 are identified as reference objects with bounding boxes.

[0044] In step S150, domain-specific object detection and segmentation algorithms are applied to multiple frames to generate at least one object dataset stream. Faster R-CNN, YOLO (You Only Look Once), and SSD (Single Shot Detector) are typically used for object detection, while U-Net, Mask R-CNN, and DeepLab are typically used for segmentation.

[0045] Similar to the pose dataset, each object dataset includes a metadata segment and multiple data segments. The metadata segment of the object dataset includes the keypoint names of the objects. As shown in subfigure 220 in Figures 2A and 2B, for example, the keypoints of the handheld nozzle 202 and cartridge assemblies 210-1 to 210-N are the corners 1, 2, 3, and 4 of their bounding boxes.

[0046] In step S160, an object dataset is extracted for each reference object. The object dataset 202, CN 119604907 A & $ + 5 / 9 pages 210-1 to 210-N shown in Figure 2A comes from the same frame as the pose dataset 201. In the example above about workers assembling a batch of 5 cartridge prototypes at assembly station 601 for a 90-second video stream, the object dataset includes a set of 2700 reference objects.

[0047] In step S170, a composite data structure is generated that integrates the pose dataset and at least one object dataset. For each frame of the video stream, the composite data includes the 2D / 3D coordinates of selected human keypoints and the 2D / 3D coordinates of keypoints of each identified one or more reference objects, which are aligned according to a coordinate system. Figure 230 in Figure 2B illustrates the composite dataset of a frame in the video stream in schematic form. Datasets 201-M, 201-1, and 201-2 contain the 2D or 3D coordinates of the left hand joint 201-M, the right hand joint 201-1, and the right hand joint 201-2, respectively. Data sets 202 and 210 containThe 2D or 3D coordinates of the bounding box corners of the handheld nozzle 202 and cartridge assemblies 210-1 to 210-N.

[0048] In the example above where workers are assembling a batch of 5 cartridge prototypes at assembly station 601 for a 90-second video stream, the composite data structure includes 2700 sets of human pose data and 2700 sets of object data. Alternatively, a composite set consisting of people and target objects can be constructed / mapped for each image frame, thereby forming a time series of 2700 composite sets.

[0049] In one embodiment, a trainable mapping is applied to the composite data structure to generate a fused data structure. The weights of this mapping are learned during the training phase. It is worth noting that the length of the fused data structure is less than the sum of the lengths of the pose dataset and the object dataset, thereby enabling more efficient processing.

[0050] In step S180, the composite data structure (or fused data structure) is input into a trained machine learning model for fine-grained action classification and regression.

[0051] In step S190, the trained machine learning model outputs a classification result indicating whether a series of captured human behaviors conform to a set of specified procedural criteria.

[0052] Figure 3 is a flowchart 300. According to an embodiment of the present invention, this flowchart 300 illustrates an exemplary method for fine-grained action classification and regression training phases.

[0053] During the training phase of the machine learning model, a domain expert is required to perform quality assurance (QA) on the product components passing through assembly station 610-1. If the QA process determines that a product has a problem, the corresponding product is marked. Then, a training set for the ML (machine learning) model is created using composite data structures (or fused data structures) associated with the marked products, which are determined by the quality assurance results. In step S380, the composite data structures marked by the domain expert or by the quality assurance results are used to train the ML model, thereby obtaining a trained ML model capable of classifying whether product components are good or bad.

[0054] Steps S310 to S370 in the ML model training phase of method 300 in Figure 3 are similar to steps S110 to S170 in the inference phase of method 100 in Figure 1. For specific implementation details, please refer to the previous description of method 100 in Figure 1.

[0055] Optionally, the first embodiment of this disclosure described above can be implemented in the context of a tokenized framework such as A1 Transformer. Specifically, composite data structures or fused data structures can be based on the tokenized form of the Transformer machine learning framework.

[0056] According to a second embodiment of this disclosure, the fine-grained motion classification and regression methods of this disclosure can be used for gait assessment. For example, Tinetti-POMA (i.e., Tinetti Performance Oriented Mobility Assessment) is a...A widely used tool for assessing balance and gait in older adults. It assesses fall risk by observing a person’s performance on various mobility tasks. The test includes a variety of activities, such as:

[0057] Balance test: seated balance, getting up from a chair, standing balance (with and without eyes), 360-degree turn;

[0058] Gait test: gait initiation, stride length and height, gait symmetry, gait continuity, path deviation, trunk stability, and walking posture.

[0059] The Tinetti-POMA test includes multiple test items, each with its own scoring criteria. Each scoring item can be scored as 0 or 1. If the total score is too low, it indicates that the person has a relatively high risk of falling in the test. For example, in the walking test for gait assessment, there are four scoring items, labeled A, B, C, and D. A. Stride length: Right heel swings past left big toe = 1; Left heel swings past right big toe = 1. B. Foot gap: Right foot completely leaves the ground = 1; Left foot completely leaves the ground = 1. C. Stride length symmetry: Left and right stride lengths are equal = 1. D. Stride length continuity: Stride length is continuous = 1.

[0061] The inference and training phases of the fine motor classification and regression methods shown in Figures 1 and 3 can be applied to gait assessment. In the first embodiment described above, each test is considered an assembly station.

[0062] Figure 4 is a schematic diagram 400 of an example gait assessment scenario described according to an embodiment of this disclosure.

[0063] Figure 4 illustrates gait assessment by analyzing a video sequence of a human subject 401 used for a walking test, according to an embodiment. In this embodiment, a camera 403 is positioned in front of the path of a human subject 401, recording a video sequence of the human subject walking towards the camera 403 and then turning away from the camera 403

[0064] . During the training phase, a domain expert (physiotherapist / doctor) 402 gives a score after observing the human subject's "performance".

[0065] Figure 5 is a flowchart 500. According to an embodiment of the invention, this flowchart 500 illustrates an exemplary method for the gait assessment training phase

[0066] . In step S510, after receiving the video sequence of the entire walking test from the camera 403, in step S520, human pose estimation is applied to multiple frames of the received video stream to generate a human pose data stream. 2D / 3D coordinates of key body joints are obtained from multiple frames in the video sequence using 2D and 3D pose estimation, thereby forming the human pose data stream - in gait assessment, the joints of the human subject, such as the patient's hands or feet, are key points that require special attention.

[0067] Next, in step S530, a pose dataset is extracted from the human pose data stream. This includes a metadata segment containing the names of the keypoints and the coordinate system used, and a data segment containing the key body joints.Actual 2D / 3D coordinates. The structure and meaning of these coordinates are defined by the metadata segment.

[0068] In step S540, at least one reference object with a spatiotemporal relationship to the human motion sequence is identified. In this embodiment, the reference object can be the ground plane of bounding box 410 in Figure 4, or the bounding box of the armrest of a chair (not shown).

[0069] Then in step S550, a domain-specific object detection and segmentation algorithm is applied to multiple frames to generate at least I object datasets. For example, the generated object datasets may include a metadata segment that names the four corners (corner 1, corner 2, corner 3, corner 4) of the bounding box indicating the reference object in each frame, and a data segment that indicates the 2D / 3D coordinates of these four corners respectively.

[0070] In step S560, the object dataset of the reference object is extracted.

[0071] In step S570, a composite data structure is generated that integrates the pose dataset and the object dataset. For each frame of the video stream, the composite data includes the 2D / 3D coordinates of the selected human keypoints and the 2D / 3D coordinates of the keypoints of the reference object, which are aligned according to the coordinate system.

[0072] In step S580, the score given by the domain expert is used as the true value. Then in step S590, the score obtained in step S580 and the composite data structure obtained in step S570 are used as the training set to start the training process of the machine learning model.

[0073] In the training phase, optionally, step S591 can be used to fine-tune the model: optimize and adjust the model based on the preliminary training results. Then in step S592, it is determined whether the model has reached the expected performance level. If the performance of the model is not satisfactory, the training step is returned for further training and adjustment. If the performance of the model is satisfactory, the training process is completed.

[0074] Optionally, the above second embodiment can be implemented in the context of a tokenized framework of A1 Transformer or similar framework. Specifically, the composite data structure or fused data structure can be based on the tokenized form of the Transformer machine learning framework.

[0075] A third embodiment of this disclosure relates to quality and compliance prediction of the final product during the assembly of printer cartridges on a factory production line, as shown in Figure 6.

[0076] The quality of the final product, i.e., the printer cartridge, depends on whether the workers at each assembly station follow the corresponding standard operating procedures when handling specific components of the printer cartridge. The final product of the printer cartridge, assembled from assembly station 610-1 to assembly station N, will then undergo a quality assurance (QA) process to inspect the assembled printer cartridges on the production line. Each cartridge will be classified as "qualified" or "unqualified" after QA, which can serve as a label for training the ML model for quality prediction.

[0077] If the Quality Assurance (QA) result is “non-compliant,” production or industrial engineers will perform a post-assembly analysis. This analysis has two main purposes. First, engineers strive to identify the root cause of the product failure. Second, they trace the assembly process to determine which or which specific assembly stations the error occurred at. It is worth noting that multiple assembly stations may have contributed to the product failure. The assembly station error metric can be expressed as a probability from 0.0 to 1.0. The results of the above analysis can be used as labels to train an ML model for compliance prediction.

[0078] As described in Figure 1 and related paragraphs, during cartridge production, worker operation videos captured at each assembly station can generate a corresponding composite data structure that integrates a pose dataset and at least one object dataset.

[0079] According to one embodiment of this disclosure, all composite data structures related to the production of the final product obtained from each assembly station on the production line can be concatenated to generate a time series of data structures. The time series of the data structures is unique to the final product and can be used to predict the quality of the final product, or to predict non-compliant parts of the assembly process when the final product quality is substandard.

[0080] FIG7 is a flowchart 700, which illustrates an exemplary method for the final product quality prediction inference stage according to another embodiment of the present invention.

[0081] As shown in FIG6, after the final printer cartridge 630 is assembled through the process from assembly station 610-1 to assembly station 610-N, all video streams capturing the actions of workers at each assembly station are generated according to steps S110 to S170 in FIG1 to generate composite data structures 1 to N. It can be understood that composite data structures 1 to N are respectively associated with the behavior of workers at assembly stations 610-1 to assembly stations 610-N, as shown in FIG6.

[0082] In one embodiment of the present disclosure, a method for implementing quality prediction is provided if conventional QA is not performed. The method is performed by a system including a processor and a non-transitory computer-readable storage medium storing instructions that, when executed, the system performs the method.

[0083] In step S710, the system for performing quality prediction or compliance prediction receives composite data structures 1 to N. It is worth noting that composite data structures 1 to N are respectively associated with the operations performed by workers when handling specific parts of printer cartridges at assembly stations 610-1 to 610-N.

[0084] In step S720, after receiving the composite data structures, the system adds a timestamp from the global clock to each composite data structure. It should be noted that each composite data structure already contains the timestamp information for each extracted frame.Additional timestamps from the global clock provide a unified time reference for all composite data structures.

[0085] In step S730, the system concatenates N composite data structures according to the timestamp information. This concatenation forms a time series of the data structures. The time series describes the assembly process of the final product in chronological order.

[0086] In step S740, the system provides the time series of the data structures as input to a trained machine learning model for quality prediction. The model has previously been trained with labeled QA classification results or cartridge final products.

[0087] Subsequently, in step S750, the trained quality prediction machine learning model outputs a classification result of "good" or "bad" for the final product. The classification result predicts the quality of the final product based on the time series analysis of the data structures. The model can serve as a pre-screening tool for predicting product quality.

[0088] Figure 8 is a flowchart 800, which illustrates an exemplary method for the inference phase of predicting final product compliance according to another embodiment of this disclosure.

[0089] In this embodiment of this disclosure, a method for implementing compliance prediction is provided if conventional QA is not performed.

[0090] The initial steps (S710, S720, S730) of this embodiment are the same as those described in the previous embodiments. These steps include receiving N composite data structures, adding global timestamps, and concatenating the structures to form a time series.

[0091] After forming the time series of the data structures, in step S820, the system receives the quality assurance (QA) result of the final product, i.e., whether the final product is "qualified" or "unqualified".

[0092] If the result of conventional QA is that the quality of the final product is unsatisfactory (i.e., unqualified), the system will perform the following steps.

[0093] In step S830, the system provides two key inputs to the trained machine learning model for compliance prediction: a) the QA result of the final product quality; b) the time series of the data structures (generated in step S730).

[0094] In step S840, the trained compliance prediction machine learning model processes the inputs and outputs which assembly parts in the final product deviate from the prescribed procedural standards.

[0095] The compliance prediction machine learning model can predict the steps / locations of non-compliance and can take relevant corrective actions.

[0096] Figure 9 shows a flowchart 900 of the final product compliance prediction training phase according to other embodiments of this disclosure.

[0097] The training process begins with manually driven steps.

[0098] In step S910, when the quality assurance (QA) results of the final product indicate a failure, the engineer performs a post-assembly analysis. This analysis includes a thorough inspection of the failed product and its assembly process to determine the potential causes of the failure.[00"] After the post-assembly analysis, in step S920, the engineer identifies one or more assembly parts that may cause failure. For each identified assembly part, in step S930, the engineer determines the probability of causing failure. This probability represents the likelihood that a particular assembly part will cause product failure. The error contribution probability is represented by a value between 0.0 and 1.0. For example, 0.0 indicates that the assembly part definitely did not cause failure, while 1.0 indicates that the assembly part definitely caused failure.

[0100] Finally, in step S950, a machine learning model for compliance prediction is trained using the generated labeled training dataset. The model learns to predict potential assembly errors and their probabilities based on the input data.

[0101] Optionally, the third embodiment of the present invention described above can be implemented in the context of a labeled A1 Transformer or similar framework. Specifically, the time series of the data structure can exist in the labeled form of page 9 / 9 of the specification based on the Transformer machine learning framework.

[0102] Those skilled in the art should understand that, for convenience and brevity, the specific workflow of the above-mentioned systems, devices, equipment and modules can be referred to the corresponding workflow in the above-mentioned method embodiments, and will not be repeated here.

[0103] By reading the accompanying drawings, the disclosure and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments when implementing the subject matter to be protected. In the claims, "A and / or B" means A, B or A and B; the word "comprising" does not exclude other elements or steps, and the indefinite article V or "an" does not exclude a plural. The words "first", "second", "third", "fourth" etc. are only used to distinguish elements or steps and do not indicate the order of elements or steps. The repetition of certain measures in mutually independent independent claims does not mean that the combination of these measures cannot exert an advantage. 12 CN 119604907 A Specification Drawings 1 / 10 pages Figure 1 13 CN 119604907 A Specification Drawings 2 / 10 pages 200 Posture Dataset 201 Key Points Right Hand Joint 1 Right Hand Joint 2 Left Hand Joint M xjl x J2 x JM V J1 yJ2 VJM z coordinate <optional - if 3D pose / object result exists) Zjl ZJM cji______ c J2 C JM Object 1 Dataset 202 201-1 201-2 Object 2 Dataset 210-1 -> Keypoint Angle 1 Angle 2 Angle 3 Angle 4 X coordinate x.PTC1 x PTC2 X PTC3 X PTC4 coordinate Y.PTC1 V.PTC2 V.PTC3 V.PTC4 Z coordinate (optional - if 3D pose / object result exists) Z_PTC12.PTC2 2_PTC3 Z.PTC4 Confidence c RTC1 c PTC2 C PTC3 C PTC4 Object N Dataset 210-N Keypoint Angle 1 Angle 2 Angle 3 Angle 4 X Coordinate x.PRC1 x PRC2 x PRC3 x PRC4 Y Coordinate V.PRC1 V PRC3 V PRC4 Z Coordinate (Optional - if 3D pose / object results exist) z_PRC1 z_PRC2 Z_PRC3 Z.PRC4 Confidence c PRC1 c PRC 2 c PRC3 C PRC4 Figure 2A 14 CN 119604907 A Instruction Manual Appendix 3 / 10 Page 230 201-M 201-1 201-2 202 210 Figure 2B 15 CN 119604907 A Instruction Manual Appendix 4 / 10 Page 300 Figure 3 16 CN 119604907 A Instruction Manual, Figure 5 / 10, Page 17 CN 119604907 A Instruction Manual, Figure 6 / 10, Figure 5 18 CN 119604907 A Instruction Manual, Figure 7 / 10, 620_N Composite Data Structure N 620_2 I 620_l Composite Data Structure 2 Composite Data Structure 1 Back-to-Back Figure 6 19 CN 119604907 A Instruction Manual, Figure 8 / 10, Page 700, Figure 7 20 CN 119604907 A Instruction Manual, Figure 9 / 10, Page 800, Figure 8 21 CN 119604907 A Instruction Manual, Figure 10 / 10, Page 900, Figure 9 22 FINE-GRAINED ACTION CLASSIFICATION AND REGRESSION Fine-grained Action Classification and Regression China application no.: 202480002727.5 Abstract: Methods and a non-transitory computer-readable storage medium for fine-grained action classification and / or regression are disclosed. The method includes: receiving a video stream capturing a sequence of human subject actions; identifying reference objects with spatial-temporalrelationships to the action sequence; extracting a pose dataset representing the action sequence; extracting object datasets representing spatial positions of the reference objects; generating a compound data structure integrating the pose dataset and object datasets; and inputting the compound data structure into a trained machine learning model for classification and / or regression. Abstract

Claims

1. A method for implementing fine motion classification and / or regression using a system comprising a processor and a non-transitory computer-readable storage medium storing instructions, wherein the instructions, when executed, cause the system to perform a method comprising the following steps: receiving ( S110 ) at least one video stream capturing a sequence of human subject actions; identifying ( S140 ) at least one reference object having a spatiotemporal relationship with the human subject action sequence; extracting (S130) a posture dataset representing a sequence of actions of the human subject; extracting (S160) at least one object data set representing a spatial position of the at least one reference object; generating ( S170 ) a composite data structure integrating the posture dataset and the at least one object dataset; as well as The composite data structure is input (S180) into a trained machine learning model for classification and / or regression.

2. The method according to claim 1, further comprising: Applying human pose estimation (S120) to a plurality of frames of the at least one video stream to generate a human pose data stream, and extracting the pose dataset from the human pose data stream, wherein the pose dataset comprises a metadata segment and a plurality of data segments; applying (S150) domain-specific object detection to the plurality of frames to generate at least one object dataset, wherein the object dataset comprises a metadata segment and a plurality of data segments; Wherein, the method further comprises: A classification result is output by the trained machine learning model for classification and / or regression, which indicates whether the captured human action sequence meets a set of specified program standards.

3. The method according to claim 1, wherein the posture data set comprises: For each extracted frame, the 2D / 3D coordinates of the selected human key points, and the confidence of each of the human key points, where the confidence is between 0.0 and 1.

0.

4. The method of claim 3, further comprising adding a bounding box to each of the at least one reference object, wherein For each extracted frame, the object dataset includes: The 2D / 3D coordinates of the key points of the bounding box, and the confidence of each key point of the reference object.

5. The method according to claim 4, wherein: Generating the composite data structure comprises: The determined 2D / 3D coordinates of the selected human body key points are combined with the determined 2D / 3D coordinates of the identified key points of the one or more reference objects based on the metadata segments.

6. The method according to claim 4, further comprising: applying a trainable mapping to the composite data structure to generate a fused data structure, wherein weights of the mapping are learned in a training phase; The length of the fused data structure is smaller than the sum of the lengths of the posture dataset and the object dataset.

7. The method of claim 1, wherein the at least one video stream is captured from one or more cameras located around the scene at predetermined time intervals.

8. The method according to claim 1 or claim 6, further comprising: The composite data structure or the fused data structure is labeled by a domain expert or through quality assurance results, The machine learning model is trained to obtain the trained machine learning model for classification and / or regression.

9. A method of implementing quality prediction or compliance prediction by a system comprising a processor and a non-transitory computer-readable storage medium; The storage medium stores instructions which, when executed, cause the system to perform the method, including: receiving (S710) N composite data structures generated by the method according to claim 1, wherein the composite data structures 1 to N are respectively associated with human action sequences that should comply with a set of specified program standards, and each of the human action sequences 1 to N is associated with an assembly part of a final product; adding (S720) a timestamp from a global clock to each composite data structure, wherein each of the composite data structures includes timestamp information for each extracted frame; According to the timestamp information of the N composite data structures, the N composite data structures are connected in series (S730) to form a time series of data structures.

10. The method according to claim 9, further comprising: providing (S740) the time series of the data structure as a quality prediction input of a trained machine learning model; as well as The trained machine learning model outputs (S750) a classification result predicting the quality of the final product.

11. The method according to claim 9, further comprising: receiving (S820) a quality assurance result of the quality of the final product, If the result of the traditional quality assurance indicates that the quality of the final product is unqualified, providing (S830) the quality assurance result of the quality of the final product and the time series of the data structure as input to a trained machine learning model for compliance prediction; The output (S840) of the trained machine learning model for compliance prediction indicates which assembly parts in the final product deviate from the specified process standard.

12. The method according to claim 11, further comprising training a machine learning model for compliance prediction, comprising: When the quality assurance result of the final product is unqualified, the engineer performs (S910) post-assembly analysis; identifying ( S920 ) one or more assembly parts that may cause the nonconformity; determining ( S930 ) an error contribution probability of each identified assembly part; generating (S940) a set of labeled training data according to the determined probability; as well as The machine learning model is trained ( S950 ) using the generated set of labeled training data to predict potential assembly errors and their probabilities.

13. The method according to claim 12, wherein: The error contribution probability is expressed as a value between 0.0 and 1.

0.

14. A non-transitory computer-readable storage medium having stored therein instructions which, when executed by one or more processors of a processing system, cause the processing system to perform the method according to any one of claims 1-13.