Vision-based sewing action recognition and positioning method and system, medium and equipment
By constructing a sewing action recognition model and using a secure dilated convolutional network and a reverse temporal convolutional network to extract sewing action features, the problem of labor shortage and low production efficiency in the garment processing industry was solved. This enabled the accurate positioning and recognition of sewing actions, thereby improving production efficiency.
Patent Information
- Application Number
- CN202511076548.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-14
AI Technical Summary
The garment processing industry suffers from labor shortages and low production efficiency, particularly in new employee training and production line analysis and identification, leading to delivery delays and economic losses.
A sewing action recognition model is constructed, including a convolution module, a temporal modeling module, and a temporal behavior recognition module. Sewing action features are extracted through a secure dilated convolutional network and a reverse temporal convolutional network. Combined with a gating fusion mechanism and loss function optimization, the model achieves accurate localization and recognition of sewing actions.
It enables precise positioning and recognition of sewing actions in complex environments, improves production efficiency, is applicable to various industrial inspection scenarios, and features modular independence and easy deployment.
Smart Images

Figure CN120954092A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video data processing technology, and in particular to a vision-based method, system, medium, and device for sewing action recognition and positioning. Background Technology
[0002] The domestic and international apparel industry currently boasts a market size of trillions of US dollars, employing approximately 120 million people. However, with rising labor costs and a declining number of young workers, the industry is experiencing a widespread labor shortage. Currently, after recruiting new employees, apparel factories rely on industrial engineers (IE) to conduct line inspections, record operational videos, analyze the processes, and provide guidance to improve sewing efficiency. This manual method is inefficient, requires extensive experience, and has a limited pool of qualified IE professionals. Furthermore, the overall efficiency of a garment production line can be affected by specific processes, requiring timely identification and location of bottlenecks. Current methods of manual line inspection and capacity analysis are relatively outdated, severely hindering delivery times and causing economic losses for garment factories during rush orders.
[0003] Therefore, how to quickly obtain employee time sequence information based on employee production operation videos and provide data for production line analysis has become an important issue that needs to be addressed by existing technologies. Summary of the Invention
[0004] The purpose of this application is to provide a vision-based method, system, medium, and equipment for recognizing and locating sewing actions, in order to solve the technical problem of low efficiency in garment processing and production caused by the low efficiency of manually analyzing data to determine process time.
[0005] To achieve the above and other related objectives, the first aspect of this application provides a vision-based method for recognizing and locating sewing actions. The vision-based method for recognizing and locating sewing actions includes:
[0006] Construct a sewing action recognition model, including a convolution module, a temporal modeling module, and a temporal behavior recognition module;
[0007] Acquire a sewing video and perform stage processing on the sewing video to obtain multi-stage continuous frame images. Input the continuous frame images of each stage into the sewing action recognition model for processing.
[0008] The sewing action feature sequence of the consecutive frame images is extracted based on the convolution module;
[0009] Based on the temporal modeling module, the secure dilated feature sequence of the sewing action feature sequence in the secure dilated convolutional network and the inverse temporal feature sequence in the inverse temporal convolutional network are calculated respectively. The secure dilated feature sequence and the inverse temporal feature sequence are fused based on the gating fusion mechanism, and the initial predicted feature sequence of the sewing action feature sequence is determined according to the fusion result and the sewing action feature sequence.
[0010] Based on the temporal behavior recognition module, the initial predicted feature sequence is subjected to safety dilation residual processing to obtain the sewing action label sequence corresponding to the continuous frame images;
[0011] The sewing action label sequence corresponding to the consecutive frame images of the multi-stage process is stacked to output the sewing action recognition result.
[0012] In some embodiments of the first aspect of this application, a sewing action recognition model is constructed, including:
[0013] Obtain sewing video samples and corresponding sewing action recognition result samples;
[0014] The sewing video test result sample is determined based on the initial sewing action recognition model;
[0015] The comprehensive loss value of the sewing video test results and the sewing action recognition results samples is calculated based on the comprehensive loss function, and the model parameters of the initial sewing action recognition model are adjusted based on the comprehensive loss value, and the training is carried out iteratively.
[0016] The initial sewing action recognition model that minimizes the overall loss value is determined as the sewing action recognition model.
[0017] In some embodiments of the first aspect of this application, the comprehensive loss function includes: ordinary cross-entropy and FocalLoss loss terms, short-time action loss term, and temporal smoothing loss term; calculating the comprehensive loss value of the sewing video test result and the sewing video sample based on the comprehensive loss function includes:
[0018] The classification loss value of the sewing video test result and the sewing action recognition result sample is calculated based on the ordinary cross-entropy and Focal Loss loss term.
[0019] Calculate the short-time motion loss value of the sewing video test result and the sewing motion recognition result sample based on the short-time motion loss term;
[0020] The temporal smoothing loss value of the sewing video test result and the sewing action recognition result sample is calculated based on the temporal smoothing loss term.
[0021] The weighted sum of the classification loss value, the short-term action loss value, and the temporal smoothing loss value is determined as the comprehensive loss value.
[0022] In some embodiments of the first aspect of this application, the dilation strategy of the secure dilated convolutional network is executed in segments along the network depth direction: the first segment of the convolutional layer adopts an exponentially increasing dilation rate strategy, and the second segment of the convolutional layer adopts a linearly increasing dilation rate strategy; wherein, the first segment of the convolutional layer accounts for the first one-third of the total number of layers of the secure dilated convolutional network, and the second segment of the convolutional layer accounts for the last two-thirds of the total number of layers of the secure dilated convolutional network.
[0023] In some embodiments of the first aspect of this application, calculating the secure dilated feature sequence of the sewing action feature sequence in a secure dilated convolutional network includes:
[0024] The secure dilated convolutional network is divided into a first convolutional layer and a second convolutional layer according to the dilation strategy;
[0025] The dilation rate of the first and second convolutional layers is calculated layer by layer according to the dilation strategy, and a safety constraint verification is performed.
[0026] The sewing action feature sequence is processed based on the verified expansion rate to obtain a safe expansion feature sequence.
[0027] In some embodiments of the first aspect of this application, the initial predicted feature sequence is subjected to secure dilation residual processing based on the temporal action recognition module to obtain the sewing action label sequence corresponding to the consecutive frame images, including:
[0028] Receive the initial predicted feature sequence to obtain the sequence length of the initial predicted feature sequence;
[0029] Perform dilated convolution processing on the initial predicted feature sequence to obtain the main feature sequence, and perform dilated convolution processing with a dilation rate of 1 on the initial predicted features to obtain the compensation feature sequence;
[0030] The main feature sequence and the compensation feature sequence are added element by element to obtain the residual output feature sequence;
[0031] Based on the residual output feature sequence, obtain the sewing action label sequence corresponding to the continuous frame images of the first stage.
[0032] In some embodiments of the first aspect of this application, the timing modeling module includes:
[0033] A one-dimensional convolutional layer employs multi-branch processing, including a first branch of a secure dilated convolutional network and a second branch of a reverse temporal convolutional network; wherein, the dilation strategy of the secure dilated convolutional network is executed segmentally along the network depth direction;
[0034] The gating activation layer employs an attention mechanism.
[0035] The residual connection layer includes a ReLU activation function built into the convolution operation.
[0036] To achieve the above and other related objectives, a second aspect of this application provides a vision-based sewing action recognition and localization system. The vision-based sewing action recognition and localization system includes: a model construction and acquisition module for constructing a sewing action recognition model, including a convolution module, a temporal modeling module, and a temporal behavior recognition module; acquiring a sewing video and performing staged processing on the sewing video to obtain multi-stage continuous frame images;
[0037] The acquisition module is used to acquire sewing video and perform stage processing on the sewing video to obtain multi-stage continuous frame images, and input the continuous frame images of each stage into the sewing action recognition model for processing.
[0038] The feature extraction module is used to extract the sewing action feature sequence of the consecutive frame images based on the convolution module;
[0039] The feature processing module is used to calculate the feature extraction results of the sewing action feature sequence in the secure dilated convolutional network and the inverse temporal convolutional network based on the temporal modeling module, respectively, fuse the feature extraction results based on the gating fusion mechanism, and determine the initial predicted feature sequence of the sewing action feature sequence based on the fusion result and the sewing action feature sequence.
[0040] The action recognition module is used to perform safety dilation residual processing on the initial predicted feature sequence based on the temporal behavior recognition module to obtain the sewing action label sequence corresponding to the continuous frame images;
[0041] An output module is used to stack the sewing action label sequence corresponding to the multi-stage continuous frame images to output the sewing action recognition result. To achieve the above and other related objectives, a third aspect of this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the vision-based sewing action recognition and localization method according to any one of the first aspects of this application.
[0042] To achieve the above and other related objectives, a fourth aspect of this application provides an electronic device, the electronic device comprising: a processor and a memory, the memory being used to store a computer program; the processor being communicatively connected to the memory, and executing the vision-based sewing action recognition and positioning method described in any of the first aspects of this application when the computer program is invoked.
[0043] As described above, the vision-based sewing action recognition and positioning method, system, medium, and device of this application have the following beneficial effects:
[0044] This application adopts an architecture of "safe dilated residual layer" and "safe prediction generation + multi-stage safe optimization" to impose an upper limit constraint on the standard dilation rate and introduce compensating convolution to avoid information sparsity caused by excessive receptive field; at the same time, it achieves robust capture of long-term dependencies and tail actions through forward / backward bi-branch dilated convolution and gating dynamic fusion.
[0045] With limited computing power, an improved version of the PP-TSM V2 network is used for spatial feature extraction, providing a detailed and context-rich representation for subsequent time series analysis.
[0046] To balance boundary accuracy and overall coherence in action segmentation, this application integrates the following into the loss function: hybrid cross-entropy + Focal Loss (balancing common frames and difficult samples), temporal weights (highlighting switching boundaries), short-term action loss (enhancing the identification of instantaneous trigger segments), and dynamically controlled temporal smoothing loss (suppressing jitter in non-switching segments and preserving true switching). This multi-dimensional and multi-objective training signal can accurately locate short-lived actions while ensuring stable output for long-term operations.
[0047] This application can accurately locate short-lived movements and ensure stable output during long-term operation. It can identify and locate sewing movements in real time and accurately under complex fabric textures, lighting changes, and occlusion interference.
[0048] Each functional module of the vision-based sewing action recognition and positioning system of this application has good module independence and interface definition, which facilitates replacement, optimization or incremental upgrade in actual deployment and is suitable for a variety of industrial inspection scenarios. Attached Figure Description
[0049] Figure 1 The diagram shown illustrates an application scenario of an electronic device according to an embodiment of this application.
[0050] Figure 2a The diagram shown is a flowchart illustrating a vision-based sewing action recognition and positioning method according to an embodiment of this application.
[0051] Figure 2bThe diagram shown is a flowchart illustrating the vision-based sewing action recognition and positioning method provided in an embodiment of this application.
[0052] Figure 3 The diagram shows the location of the video acquisition device in an embodiment of this application.
[0053] Figure 4a The diagram shown is a flowchart illustrating the process of obtaining the initial predicted feature sequence according to an embodiment of this application.
[0054] Figure 4b The diagram shown is a structural schematic of a secure dilated convolutional network according to an embodiment of this application.
[0055] Figure 5 The diagram shows a flowchart of obtaining the sewing action label sequence according to an embodiment of this application.
[0056] Figure 6 The diagram shown is a flowchart illustrating a vision-based sewing action recognition and positioning method according to an embodiment of this application.
[0057] Figure 7 The diagram shown is a structural schematic of a vision-based sewing action recognition and positioning system according to an embodiment of this application.
[0058] Figure 8 The diagram shown is a structural schematic of an electronic device according to an embodiment of this application.
[0059] Component designation
[0060] 11 mobile phones
[0061] 12 tablet computers
[0062] 13 Laptops
[0063] 3. Video capture equipment
[0064] 7. Vision-based sewing motion recognition and positioning system
[0065] 71 Model Building Module
[0066] 72 Acquisition Module
[0067] 73 Feature Extraction Module
[0068] 74 Feature Processing Module
[0069] 75 Action Recognition Module
[0070] 76 Output Module
[0071] 8 Electronic devices
[0072] 801 processor
[0073] 802 memory
[0074] 8021 operating system
[0075] 8022 Application
[0076] 803 Network Interface
[0077] 804 bus system
[0078] 805 User Interface
[0079] Steps S21 to S26
[0080] Steps S241~S243 Detailed Implementation
[0081] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0082] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0083] The PP-TSM V2 (Paddle Temporal Shift Module Version 2) model is a high-efficiency and practical video classification model developed by the PaddleVideo team. It is an improvement upon the TSM (Temporal Shift Module), inheriting the core ideas of the time-shift module and effectively improving model accuracy while maintaining the original number of parameters. Its main components include:
[0084] Time Shift Module (TSM): Its core lies in time shift operation, which captures the temporal information in the video by shifting and sharing some feature channels in the time dimension. No additional parameters or computation are required, enabling the model to utilize the temporal information of the video.
[0085] Backbone: Lightweight models such as MobileNetV3 and LitePose can be used as the backbone network.
[0086] Classifier: After extracting features from the backbone network, the frame-level features are first averaged to obtain video-level features, and then a fully connected layer is used for classification to reduce interference and improve accuracy.
[0087] The vision-based sewing action recognition and localization method of this application can be applied to, for example... Figure 1 The electronic devices shown in this application may include mobile phones 11 with wireless charging capabilities, tablet computers 12, laptop computers 13, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. The specific types of electronic devices are not limited in this application embodiment.
[0088] For example, the electronic device may be a station (STAION, ST) in a WLAN with wireless charging capability, a cellular phone, cordless phone, Session Initiation Protocol (SIP) phone, Wireless Local Loop (WLL) station, Personal Digital Assistant (PDA) device, handheld device with wireless charging capability, computing device or other processing device, computer, laptop computer, handheld communication device, handheld computing device, and / or other devices for communicating over a wireless system, as well as next-generation communication systems, etc.
[0089] For example, the electronic device can communicate with networks and other devices wirelessly. The aforementioned wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), BT, GNSS, WLAN, NFC, FM, and / or IR technologies. The technical solutions in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0090] The principles and implementation methods of the vision-based sewing action recognition and positioning method, system, medium, and device described in this application will be explained in detail below with reference to the accompanying drawings, so that those skilled in the art can understand the vision-based sewing action recognition and positioning method, system, medium, and device of this embodiment without creative effort.
[0091] To facilitate understanding of the embodiments of this application, firstly, in conjunction with Figures 2a-2b Detailed explanation. Figure 2a The flowchart shown is a vision-based sewing action recognition and positioning method provided in an embodiment of this application. Figure 2b The diagram shown illustrates a flowchart of a vision-based sewing action recognition and positioning method provided in an embodiment of this application. Figure 2a As shown, the process includes steps S21 to S26.
[0092] Step S21: Construct a sewing action recognition model, including a convolution module, a temporal modeling module, and a temporal behavior recognition module.
[0093] In this embodiment, constructing a sewing action recognition model includes: acquiring sewing video samples and corresponding sewing action recognition result samples; determining sewing video test result samples of the sewing video samples based on an initial sewing action recognition model; calculating the comprehensive loss value of the sewing video test results and the sewing action recognition result samples according to a comprehensive loss function, and adjusting the model parameters of the initial sewing action recognition model based on the comprehensive loss value, and iteratively training; determining the initial sewing action recognition model with the smallest comprehensive loss value as the sewing action recognition model.
[0094] The comprehensive loss function includes: ordinary cross-entropy and Focal Loss loss terms, short-time action loss term, and temporal smoothing loss term.
[0095] The comprehensive loss value of the sewing video test result and the sewing video sample is calculated based on the comprehensive loss function, including: calculating the classification loss value of the sewing video test result and the sewing action recognition result sample based on the ordinary cross-entropy and Focal Loss loss terms; calculating the short-time action loss value of the sewing video test result and the sewing action recognition result sample based on the short-time action loss term; calculating the temporal smoothing loss value of the sewing video test result and the sewing action recognition result sample based on the temporal smoothing loss term; and determining the weighted sum of the classification loss value, the short-time action loss value, and the temporal smoothing loss value as the comprehensive loss value.
[0096] Step S22: Acquire the sewing video and perform stage processing on the sewing video to obtain multi-stage continuous frame images. Input the continuous frame images of each stage into the sewing action recognition model for processing.
[0097] In this embodiment, the device for acquiring the sewing video is a camera, which is mounted on the waist of the sewing machine, such as... Figure 3 As shown. This camera is a Galaxycore CV2003 model camera, equipped with a Rockchip RV1126 chip.
[0098] Specifically, a video decoder is used to read the sewn video frame by frame in order to decode the sewn video into a series of image frames.
[0099] According to a preset strategy, continuous frames are divided into multiple stages. Each frame in each stage is preprocessed, including size normalization and pixel normalization, to obtain continuous frame images in multiple stages.
[0100] To avoid losing information at stage boundaries, overlapping division can be used, where some frames from adjacent stages are overlapped.
[0101] Optionally, for longer sewing videos, in order to reduce computational load, time downsampling can be performed within each segment to compress the segment length.
[0102] Step S23: Extract the sewing action feature sequence of the continuous frame images based on the convolution module.
[0103] In this embodiment, the feature extraction model on which the convolution module relies is the PP-TSM V2 model, and the scene video data of the sewing workers are used as the training dataset for the feature extraction model.
[0104] For the acquired multi-stage continuous frame images, they need to be input into the trained feature extraction model for feature extraction to obtain the sewing action feature sequence corresponding to the continuous frame images.
[0105] Step S24: Based on the temporal modeling module, calculate the secure dilated feature sequence of the sewing action feature sequence in the secure dilated convolutional network and the inverse temporal feature sequence in the inverse temporal convolutional network, respectively. Based on the gating fusion mechanism, fuse the secure dilated feature sequence and the inverse temporal feature sequence, and determine the initial predicted feature sequence of the sewing action feature sequence according to the fusion result and the sewing action feature sequence.
[0106] In this embodiment, as Figure 4a As shown, step S24 includes:
[0107] Step S241: Process the sewing action feature sequence using a secure dilated convolutional network and a reverse temporal convolutional network to obtain the secure dilated feature sequence and the reverse temporal feature sequence, respectively.
[0108] Step S242: Perform feature fusion on the safety expansion feature sequence, the reverse feature sequence, and the sewing action feature sequence to obtain a fused feature sequence;
[0109] Step S243: Use residual connection to fuse the sewing action feature sequence with the fused feature sequence to obtain initial prediction features.
[0110] In this embodiment, the temporal modeling module includes: a one-dimensional convolutional layer employing multi-branch processing, comprising a first branch of a securely dilated convolutional network and a second branch of a reverse temporal convolutional network; a gated activation layer employing an attention mechanism; and a residual connection layer comprising a ReLU activation function built into the convolution operation.
[0111] The dilation strategy of the secure dilated convolutional network is executed in segments along the network depth direction: the first segment of the convolutional layer adopts an exponential dilation rate strategy, and the second segment of the convolutional layer adopts a linear dilation rate strategy; the first segment of the convolutional layer accounts for the first third of the total number of layers in the secure dilated convolutional network, and the second segment of the convolutional layer accounts for the last two-thirds of the total number of layers in the secure dilated convolutional network.
[0112] Specifically, the secure dilated convolutional network employs an innovative two-stage dilation strategy in the depth dimension, strictly dividing the entire convolutional network into two complementary processing stages:
[0113] The first convolutional layer, also known as the exponential growth layer, accounts for the first third of the total number of layers in the secure dilated convolutional network. The dilation rate growth strategy employs exponential expansion so that the shallow network features obtained in the shallow network of the secure dilated convolutional network can quickly expand the receptive field and capture long-range dependencies.
[0114] The second convolutional layer, or linear growth layer, accounts for the last two-thirds of the total number of layers in the secure dilated convolutional network. The dilation rate adopts a linear and stable growth strategy to avoid excessive sparsity of the deep network features obtained in the deep network of the secure dilated convolutional network, and to maintain the ability to perceive local details.
[0115] like Figure 4b As shown, the sewing action feature sequence is processed according to the safety dilated convolutional network, including: dividing the safety dilated convolutional network into a first convolutional layer and a second convolutional layer according to the dilation strategy; calculating the dilation rate of the first convolutional layer and the second convolutional layer layer by layer according to the dilation strategy and performing safety constraint verification; and processing the sewing action feature sequence based on the verified dilation rate to obtain the safety dilated feature sequence.
[0116] The sewing action feature sequence is processed using a reverse temporal convolutional network, including: reversing the sewing action feature sequence to obtain a reverse sequence; applying a standard convolution operation to the obtained reverse sequence to generate an intermediate feature sequence; and reversing the intermediate feature sequence step-by-step to generate a reverse feature sequence. The reverse feature sequence obtained through the reverse temporal convolutional network can capture the dependencies between future and current time steps and avoids the gradient conflict problem of traditional bidirectional architectures.
[0117] Furthermore, the safety expansion feature sequence, the reverse feature sequence, and the sewing action feature sequence are fused to obtain a fused feature sequence, including: fusing the safety expansion feature sequence, the reverse feature sequence, and the sewing action feature sequence through a gating fusion mechanism.
[0118] Calculate the similarity matrix of the safety expansion feature sequence, the reverse feature sequence, and the sewing action feature sequence. Generate dynamic gating weights based on the obtained similarity matrix. Fuse the safety expansion feature sequence, the reverse feature sequence, and the sewing action feature sequence according to the dynamic gating weights to obtain the fused feature sequence.
[0119] A residual connection is used to fuse the sewing action feature sequence with the fused feature sequence to obtain initial prediction features. This includes: element-wise addition of the sewing action feature sequence and the fused feature sequence, followed by activation function processing of the addition result to obtain the initial prediction features. By fusing shallow detail features (sewing action feature sequence) and deep semantic features (fused feature sequence) through residual connections, the initial prediction features retain the spatiotemporal integrity of the original data while possessing advanced action intent understanding capabilities, laying a solid foundation for subsequent accurate predictions.
[0120] Step S25: Based on the temporal behavior recognition module, perform secure dilation residual processing on the initial predicted feature sequence to obtain the sewing action label sequence corresponding to the continuous frame images.
[0121] In this embodiment, the safety expansion residual layer consists of the following structure:
[0122] Main convolution branch: a stack of three dilated convolutions with increasing dilation rate, where the dilation rate sequence of the three dilated convolutions is [1,2,5].
[0123] Compensated convolution branch: The dilation rate of its three-level convolution is 1;
[0124] In the main convolutional branch and the compensated convolutional branch: inter-level insertion of batch normalized layers with the Swish activation function;
[0125] The outputs of the main convolution branch and the outputs of the compensated convolution branch are added together through residual connections.
[0126] In this embodiment, as Figure 5 As shown, step S25 includes:
[0127] Step S251: Receive the initial predicted feature sequence to obtain the sequence length of the initial predicted feature sequence;
[0128] Step S252: Perform dilated convolution processing on the initial predicted feature sequence to obtain the main feature sequence, and perform dilated convolution processing with a dilation rate of 1 on the initial predicted features to obtain the compensation feature sequence.
[0129] Step S253: Add the main feature sequence and the compensation feature sequence element by element to obtain the residual output feature sequence;
[0130] Step S254: Obtain the sewing action label sequence corresponding to the continuous frame images of the first stage based on the residual output feature sequence.
[0131] Obtaining the sewing action label sequence corresponding to the continuous frame images of a first stage based on the residual output feature sequence includes: performing temporal pooling on the residual output feature sequence to obtain a global feature vector; and inputting the global feature vector into a classifier to generate a sewing action label sequence.
[0132] Step S26: Stack the sewing action label sequences corresponding to the continuous frame images of multiple stages to output the sewing action recognition result.
[0133] In this embodiment, step S26 includes: splicing the acquired multi-stage sewing action label sequence according to the channel dimension, and reducing the channel dimension to the category dimension through a convolutional compression layer to generate sewing action recognition results.
[0134] Please see Figure 6, Figure 6 The diagram shown is a flowchart illustrating a vision-based sewing action recognition and positioning method according to an embodiment of this application.
[0135] Multi-stage (1 to n) continuous frame images: The acquired sewing video is decoded into continuous image frames, and the continuous image frames are divided into stages according to a preset strategy to obtain 1 to n stage continuous frame images.
[0136] The following example, using the processing of consecutive frames in stage 1, briefly describes the vision-based sewing action recognition and localization method described in this embodiment.
[0137] The feature sequence of the first-stage sewing action was obtained by feature extraction using the PP-TSM V2 model.
[0138] Sequence: Obtained by processing the feature sequence of the 1-stage sewing action according to the preset safe dilated convolutional network; wherein the preset safe dilated convolutional network includes: the first convolutional layer, which accounts for the first third of the total number of layers in the safe dilated convolutional network, and adopts an exponential expansion strategy for the dilation rate growth; the second convolutional layer, which accounts for the last two-thirds of the total number of layers in the safe dilated convolutional network, and adopts a linear and stable growth strategy for the dilation rate growth.
[0139] Reverse time feature sequence: The sewing action feature sequence is reversed and reconstructed, and a standard convolution operation is performed to generate an intermediate feature sequence. The intermediate feature sequence is then reversed and reconstructed according to time steps to generate a reverse time feature sequence.
[0140] Fusion Feature Sequence: Calculate the similarity matrix of the safety expansion feature sequence, the reverse feature sequence, and the sewing action feature sequence. Generate dynamic gating weights based on the obtained similarity matrix. Fusion feature sequence, reverse feature sequence, and sewing action feature sequence are fused according to the dynamic gating weights to obtain the fused feature sequence.
[0141] Initial prediction features in stage 1: The sewing action feature sequence and the fused feature sequence are added element by element, and the addition result is processed by an activation function to obtain the initial prediction features in stage 1.
[0142] Stage 1 sewing action label sequence: The initial predicted features of Stage 1 are processed according to the preset safe dilation residual layer to obtain the Stage 1 sewing action label sequence; wherein, the preset safe dilation residual layer includes: main convolution branch: a stack of three-level dilated convolutions with increasing dilation rate, wherein the dilation rate sequence of the three-level dilated convolutions is [1, 2, 5]; compensation convolution branch: the dilation rate of its three-level convolutions is 1; in the main convolution branch and the compensation convolution branch: batch normalization layers and Swish activation functions are inserted between levels; the output of the main convolution branch and the output of the compensation convolution branch are added through residual connections.
[0143] Based on the same processing procedure as the continuous frame images in stage 1 above, the sewing action label sequence of order 2 to n is obtained.
[0144] Sewing action recognition results: The acquired multi-stage sewing action label sequence is spliced according to the channel dimension, and the channel dimension is reduced to the category dimension through a convolutional compression layer to generate sewing action recognition results.
[0145] It should be noted that, in the embodiments of this application, the words "exemplary" or "for example" indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0146] The scope of protection of the vision-based sewing action recognition and positioning method in this application is not limited to the execution order of the steps listed in this embodiment. Any solution implemented by adding, subtracting, or replacing steps in the prior art based on the principles of this application is included within the scope of protection of this application.
[0147] This application also provides a vision-based sewing action recognition and positioning system. The vision-based sewing action recognition and positioning system can implement the vision-based sewing action recognition and positioning method of this application. However, the implementation device of the vision-based sewing action recognition and positioning method of this application includes, but is not limited to, the structure of the vision-based sewing action recognition and positioning system listed in this embodiment. All structural modifications and substitutions of the prior art made in accordance with the principles of this application are included within the protection scope of this application.
[0148] Please see Figure 7 The diagram shows a structural schematic of a vision-based sewing action recognition and positioning system according to an embodiment of this application.
[0149] like Figure 7 As shown, the vision-based sewing action recognition and localization system 7 includes: a model building module 71, an acquisition module 72, a feature extraction module 73, a feature processing module 74, an action recognition module 75, and an output module 76. Among these,
[0150] Model building module 71 is used to build a sewing action recognition model, including a convolution module, a temporal modeling module, and a temporal behavior recognition module;
[0151] The acquisition module 72 is used to acquire the sewing video and perform stage processing on the sewing video to obtain multi-stage continuous frame images, and input the continuous frame images of each stage into the sewing action recognition model for processing.
[0152] Feature extraction module 73 is used to extract the sewing action feature sequence of the consecutive frame images based on the convolution module;
[0153] The feature processing module 74 is used to calculate the feature extraction results of the sewing action feature sequence in the secure dilated convolutional network and the inverse temporal convolutional network based on the temporal modeling module, fuse the feature extraction results based on the gating fusion mechanism, and determine the initial predicted feature sequence of the sewing action feature sequence based on the fusion result and the sewing action feature sequence.
[0154] Action recognition module 75 is used to perform safety dilation residual processing on the initial predicted feature sequence based on the temporal behavior recognition module to obtain the sewing action label sequence corresponding to the continuous frame images;
[0155] Output module 76 is used to stack the sewing action label sequence corresponding to the continuous frame images of multiple stages to output the sewing action recognition result.
[0156] The working content of each module in the vision-based sewing action recognition and positioning system of this application embodiment is the same as that of the vision-based sewing action recognition and positioning method of this application embodiment, so they will not be described in detail here.
[0157] It should be noted that the above division of modules is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, module x can be a separate processing element, or it can be integrated into a chip in the aforementioned device. Alternatively, it can be stored as program code in the memory of the aforementioned device, and its function can be called and executed by a processing element of the device. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. In the implementation process, the steps of the above method or the various modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.
[0158] According to the vision-based sewing action recognition and positioning method provided in the embodiments of this application, this application also provides a computer program product, which includes one or more computer instructions. When the one or more computer instructions are executed on a computer, the computer causes the computer to execute the instructions shown in Figures 2 to 3. Figure 6 The embodiments shown illustrate vision-based sewing action recognition and localization methods.
[0159] When the computer program code is executed on a computer, it produces, in whole or in part, the processes or functions according to the embodiments of this application. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0160] According to the vision-based sewing action recognition and positioning method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code. When the program code is run on a computer, the computer executes the commands shown in Figures 2 to 3. Figure 6 The embodiments shown illustrate vision-based sewing action recognition and localization methods.
[0161] Figure 8 This is a schematic block diagram of the electronic device provided in an embodiment of this application. Figure 8 As shown, electronic device 8 includes at least one processor 801, a memory 802, at least one network interface 803, and a user interface 805. The various components in the device are coupled together via a bus system 804. It is understood that the bus system 804 is used to implement communication between these components. In addition to a data bus, the bus system 804 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 8 The general will label all buses as bus systems.
[0162] The user interface 805 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.
[0163] It is understood that memory 802 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable categories of memory.
[0164] In an exemplary embodiment, the electronic device 8 may be used by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs) to perform the aforementioned method.
[0165] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0166] In summary, this application provides a vision-based method, system, medium, and device for sewing action recognition and localization. This application employs a "safe dilation residual layer" and a "safe prediction generation + multi-stage safe optimization" architecture to impose an upper limit constraint on the standard dilation rate and introduce compensating convolutions, avoiding information sparsity caused by excessive receptive fields. Simultaneously, through forward / backward bi-branch dilation convolutions and gated dynamic fusion, robust capture of long-term dependencies and tail actions is achieved. Under limited computing power, an improved version of the PP-TSM V2 network is used for spatial feature extraction, providing detailed and context-rich representations for subsequent temporal analysis. To balance boundary accuracy and overall coherence in action segmentation, the loss function incorporates: hybrid cross-entropy + FocalLoss (balancing common frames and difficult samples), temporal weights (highlighting switching boundaries), short-term action loss (enhancing instantaneous trigger segment recognition), and dynamically controlled temporal smoothing loss (suppressing jitter in non-switching segments and preserving true switching). This multi-dimensional, multi-objective training signal enables precise localization of short-lived movements while ensuring stable output during long-term operations. This application can accurately locate short-lived movements and ensure stable output during long-term operations, enabling real-time and accurate identification and localization of sewing actions even under complex fabric textures, lighting variations, and occlusion interference. Each functional module of the vision-based sewing action recognition and localization system in this application has good module independence and interface definition, facilitating replacement, optimization, or incremental upgrades during actual deployment, and is applicable to various industrial inspection scenarios. Therefore, this application effectively overcomes the various shortcomings of existing technologies and has high industrial application value.
[0167] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A vision-based method for recognizing and locating sewing actions, characterized in that, The method includes: Construct a sewing action recognition model, including a convolution module, a temporal modeling module, and a temporal behavior recognition module; Acquire a sewing video and perform stage processing on the sewing video to obtain multi-stage continuous frame images. Input the continuous frame images of each stage into the sewing action recognition model for processing. The sewing action feature sequence of the consecutive frame images is extracted based on the convolution module; Based on the temporal modeling module, the secure dilated feature sequence of the sewing action feature sequence in the secure dilated convolutional network and the inverse temporal feature sequence in the inverse temporal convolutional network are calculated respectively. The secure dilated feature sequence and the inverse temporal feature sequence are fused based on the gating fusion mechanism, and the initial predicted feature sequence of the sewing action feature sequence is determined according to the fusion result and the sewing action feature sequence. Based on the temporal behavior recognition module, the initial predicted feature sequence is subjected to safety dilation residual processing to obtain the sewing action label sequence corresponding to the continuous frame images; The sewing action label sequence corresponding to the consecutive frame images of the multi-stage process is stacked to output the sewing action recognition result.
2. The vision-based sewing action recognition and localization method according to claim 1, characterized in that, Construct a sewing action recognition model, including: Obtain sewing video samples and corresponding sewing action recognition result samples; The sewing video test result sample is determined based on the initial sewing action recognition model; The comprehensive loss value of the sewing video test results and the sewing action recognition results samples is calculated based on the comprehensive loss function, and the model parameters of the initial sewing action recognition model are adjusted based on the comprehensive loss value, and the training is carried out iteratively. The initial sewing action recognition model that minimizes the overall loss value is determined as the sewing action recognition model.
3. The vision-based sewing action recognition and localization method according to claim 2, characterized in that, The comprehensive loss function includes: ordinary cross-entropy and FocalLoss loss terms, short-time action loss term, and temporal smoothing loss term; the comprehensive loss value of the sewing video test result and the sewing video sample is calculated based on the comprehensive loss function, including: The classification loss value of the sewing video test result and the sewing action recognition result sample is calculated based on the ordinary cross-entropy and Focal Loss loss term. Calculate the short-time motion loss value of the sewing video test result and the sewing motion recognition result sample based on the short-time motion loss term; The temporal smoothing loss value of the sewing video test result and the sewing action recognition result sample is calculated based on the temporal smoothing loss term. The weighted sum of the classification loss value, the short-term action loss value, and the temporal smoothing loss value is determined as the comprehensive loss value.
4. The vision-based sewing action recognition and localization method according to claim 1, characterized in that, The dilation strategy of the secure dilated convolutional network is executed in segments along the network depth direction: the first segment of the convolutional layer adopts an exponentially increasing dilation rate strategy, and the second segment of the convolutional layer adopts a linearly increasing dilation rate strategy; wherein, the first segment of the convolutional layer accounts for the first one-third of the total number of layers of the secure dilated convolutional network, and the second segment of the convolutional layer accounts for the last two-thirds of the total number of layers of the secure dilated convolutional network.
5. The vision-based sewing action recognition and localization method according to claim 4, characterized in that, Calculating the secure dilated feature sequence of the sewing action feature sequence in a secure dilated convolutional network includes: The secure dilated convolutional network is divided into a first convolutional layer and a second convolutional layer according to the dilation strategy; The dilation rate of the first and second convolutional layers is calculated layer by layer according to the dilation strategy, and a safety constraint verification is performed. The sewing action feature sequence is processed based on the verified expansion rate to obtain a safe expansion feature sequence.
6. The vision-based sewing action recognition and localization method according to claim 1, characterized in that, Based on the temporal behavior recognition module, the initial predicted feature sequence is subjected to safety dilation residual processing to obtain the sewing action label sequence corresponding to the consecutive frame images, including: Receive the initial predicted feature sequence to obtain the sequence length of the initial predicted feature sequence; Perform dilated convolution processing on the initial predicted feature sequence to obtain the main feature sequence, and perform dilated convolution processing with a dilation rate of 1 on the initial predicted features to obtain the compensation feature sequence; The main feature sequence and the compensation feature sequence are added element by element to obtain the residual output feature sequence; Based on the residual output feature sequence, obtain the sewing action label sequence corresponding to the continuous frame images of the first stage.
7. The vision-based sewing action recognition and localization method according to claim 1, characterized in that, The time series modeling module includes: A one-dimensional convolutional layer employs multi-branch processing, including a first branch of a secure dilated convolutional network and a second branch of a reverse temporal convolutional network; wherein, the dilation strategy of the secure dilated convolutional network is executed segmentally along the network depth direction; The gating activation layer employs an attention mechanism. The residual connection layer includes a ReLU activation function built into the convolution operation.
8. A vision-based sewing action recognition and positioning system, characterized in that, The system includes: The model building module is used to build a sewing action recognition model, including a convolution module, a temporal modeling module, and a temporal behavior recognition module; The acquisition module is used to acquire sewing video and perform stage processing on the sewing video to obtain multi-stage continuous frame images, and input the continuous frame images of each stage into the sewing action recognition model for processing. The feature extraction module is used to extract the sewing action feature sequence of the consecutive frame images based on the convolution module; The feature processing module is used to calculate the feature extraction results of the sewing action feature sequence in the secure dilated convolutional network and the inverse temporal convolutional network based on the temporal modeling module, respectively, fuse the feature extraction results based on the gating fusion mechanism, and determine the initial predicted feature sequence of the sewing action feature sequence based on the fusion result and the sewing action feature sequence. The action recognition module is used to perform safety dilation residual processing on the initial predicted feature sequence based on the temporal behavior recognition module to obtain the sewing action label sequence corresponding to the continuous frame images; The output module is used to stack the sewing action label sequence corresponding to the continuous frame images of multiple stages to output the sewing action recognition result.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the vision-based sewing action recognition and positioning method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, The electronic device includes: A memory that stores a computer program; The processor, which is communicatively connected to the memory, executes the vision-based sewing action recognition and positioning method according to any one of claims 1 to 7 when calling the computer program.
Citation Information
Patent Citations
Video sentiment classification method combining multi-level branch convolution and expansion interactive sampling
CN115965898A
UWB coal mine underground positioning method based on bidirectional time convolution network
CN118283524A
Space-time self-supervised runoff data completion method and device and storage medium
CN119088793A
Temporal bottleneck attention architecture for video action recognition
US11270124B1