Multi-modal gait recognition method and device for breeding personnel identity recognition

By extracting gait sequence frames from aquaculture scenarios to generate human silhouette and analytical images, and using a multimodal gait recognition model to fuse feature information, the problem of insufficient gait recognition accuracy in aquaculture scenarios is solved, achieving high-accuracy identity recognition in complex environments.

CN120977002APending Publication Date: 2025-11-18BEIJING RES CENT FOR INFORMATION TECH & AGRI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510986228.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In aquaculture scenarios, existing gait recognition models have insufficient accuracy in situations such as occlusion, blurred vision at long distances, and residual noise, making it difficult to accurately identify the identity of aquaculture workers in complex environments.

Method used

By detecting workers in monitoring videos of aquaculture scenes, gait sequence frames are extracted to generate human silhouette images and human analytical images. These are then input into a multimodal gait recognition model, which integrates global shape and fine structure feature information. The P3D-ECANet and 3D Swin Transformer structures are used to collaboratively process local details and global context, thereby improving recognition accuracy.

Benefits of technology

It significantly improves the accuracy of personnel identification in complex aquaculture environments, can accurately identify the identity of workers in multi-person collaborative operation scenarios, and enhances the discriminative power of gait representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977002A_ABST
    Figure CN120977002A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal gait recognition method and device for breeding personnel identity recognition, and relates to the technical field of gait recognition, and the method comprises the steps: detecting an operator in a breeding scene monitoring video, carrying out the target tracking, and extracting a gait sequence frame based on a target frame, and extracting a human body connected domain from the gait sequence frame to generate a human body silhouette image, performing human body analysis on the gait sequence frame to segment a human body analysis image of seven types of semantic parts, and inputting the human body silhouette image and the human body analysis image into the multi-modal gait recognition model to obtain a personnel identity recognition result output by the multi-modal gait recognition model. Wherein the multi-modal gait recognition model is obtained based on multi-modal data training of a plurality of gait samples with personnel identity tags, the multi-modal data comprises a human body silhouette and a human body analysis graph, and the multi-modal gait recognition model is used for fusing feature information of the human body silhouette and the human body analysis graph frame by frame. And personnel identity recognition is carried out based on the fused feature information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of gait recognition, and in particular to a multi-modal gait recognition method and device for identifying the identity of a breeding personnel. BACKGROUND

[0002] In a breeding scenario, real-time identification and collection of the identity information of breeding personnel can provide useful information input for the development of an intelligent production management system. Gait recognition is also applicable to a traceability system in a breeding scenario. However, in an actual breeding scenario, there are generally situations such as occlusion, blurred long-distance view, and noise residue caused by carrying tools. In a gait recognition model of related technology, a convolutional neural network (CNN) focuses on the collection of time sequence feature extraction and ignores the modeling of inter-frame dependency. A residual network (ResNet) has a too concentrated local receptive field, which suppresses its global expression ability. The recognition accuracy of the model under a real monitoring system is still insufficient. SUMMARY

[0003] The present application provides a multi-modal gait recognition method and device for identifying the identity of a breeding personnel, to improve the accuracy of personnel identity recognition in a breeding environment.

[0004] In a first aspect, the present application provides a multi-modal gait recognition method for identifying the identity of a breeding personnel, which comprises: detecting a working personnel in a breeding scenario monitoring video and performing target tracking, and extracting a gait sequence frame based on a target frame; extracting a human connected domain from the gait sequence frame to generate a human silhouette graph, and performing human parsing on the gait sequence frame to segment a human parsing graph of 7 types of semantic parts; inputting the human silhouette graph and the human parsing graph into a multi-modal gait recognition model to obtain a personnel identity recognition result output by the multi-modal gait recognition model; The multi-modal gait recognition model is trained based on multi-modal data of a plurality of gait samples with personnel identity labels, and the multi-modal data includes the human silhouette graph and the human parsing graph. The multi-modal gait recognition model is used to fuse the feature information of the human silhouette graph and the human parsing graph frame by frame, and to perform personnel identity recognition based on the fused feature information.

[0005] In a second aspect, the present application further provides a multi-modal gait recognition device for identifying the identity of a breeding personnel, which comprises: an extraction unit configured to detect a working personnel in a breeding scenario monitoring video and perform target tracking, and extract a gait sequence frame based on a target frame; The generating unit is configured to generate a human silhouette graph by extracting a human connected domain from the gait sequence frame, and generate a human parsing graph by performing human parsing segmentation on the gait sequence frame to segment 7 types of semantic parts; The identifying unit is configured to input the human silhouette graph and the human parsing graph into a multi-modal gait recognition model to obtain a personnel identity recognition result output by the multi-modal gait recognition model. The multi-modal gait recognition model is trained based on multi-modal data of a plurality of gait samples with personnel identity labels, the multi-modal data includes the human silhouette graph and the human parsing graph, and the multi-modal gait recognition model is configured to fuse feature information of the human silhouette graph and the human parsing graph frame by frame, and perform personnel identity recognition based on the fused feature information.

[0006] In a third aspect, the present application further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the multi-modal gait recognition method for personnel identity recognition in aquaculture according to the first aspect when executing the computer program.

[0007] In a fourth aspect, the present application further provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the multi-modal gait recognition method for personnel identity recognition in aquaculture according to the first aspect.

[0008] In a fifth aspect, the present application further provides a computer program product, which includes a computer program, and the computer program is executable on a processor to implement the multi-modal gait recognition method for personnel identity recognition in aquaculture according to the first aspect.

[0009] The multi-modal gait recognition method and device for personnel identity recognition in aquaculture provided by the present application can detect and target track workers in an aquaculture scene monitoring video, extract gait sequence frames to generate a human silhouette graph and a human parsing graph, and output the human silhouette graph and the human parsing graph to a multi-modal gait recognition model to finally obtain a personnel identity recognition result. The silhouette with global shape and the human parsing with fine structure are fused, the discriminability of gait representation is enhanced, the multi-modal gait recognition model can significantly improve the attention coverage of human key regions by coordinating local details and global context, and can extract more discriminative feature representations. In a complex aquaculture environment and a multi-person collaborative work scene, the personnel identity can still be accurately recognized, and the accuracy of personnel identity recognition in an aquaculture environment is improved. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to make the technical solutions in the present application or the prior art clearer, the accompanying drawings needed in the embodiments or the prior art description will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present application, and all other embodiments obtained by a person of ordinary skill in the art without creative work on the basis of these drawings also belong to the protection scope of the present application.

[0011] Figure 1 is a flowchart of a multi-modal gait recognition method for breeding personnel identity recognition provided by the present application.

[0012] Figure 2 is a video data acquisition schematic diagram of breeding workers provided by an embodiment of the present application.

[0013] Figure 3 is a flowchart of a multi-modal gait recognition method provided by an embodiment of the present application.

[0014] Figure 4 is a P3D-ECANet structure diagram provided by an embodiment of the present application.

[0015] Figure 5 is a 3D Swin Transformer structure diagram provided by an embodiment of the present application.

[0016] Figure 6 is a gait data set quality enhancement processing schematic diagram provided by an embodiment of the present application.

[0017] Figure 7 is a multi-modal gait recognition method model structure schematic diagram provided by an embodiment of the present application.

[0018] Figure 8 is a structure schematic diagram of a multi-modal gait recognition device for breeding personnel identity recognition provided by the present application.

[0019] Figure 9 is a structure schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION

[0020] In order to make the technical solutions in the present application or the prior art clearer, the accompanying drawings needed in the embodiments or the prior art description will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present application, and all other embodiments obtained by a person of ordinary skill in the art without creative work on the basis of these drawings also belong to the protection scope of the present application.

[0021] The term "and / or", used in the present application, describes the association relationship of the associated objects, and indicates that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.

[0022] The term "multiple" in the present application refers to two or more, and other quantifiers are similar.

[0023] The terms "first", "second", and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second" are usually of the same type and do not limit the number of objects, for example, the first object can be one or more.

[0024] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.

[0025] In order to more clearly understand the technical solutions of the embodiments of the present application, first, some technical contents related to the embodiments of the present application are introduced.

[0026] In the breeding scene, real-time identification and collection of the identity information of the breeding personnel can provide useful information input for the development of an intelligent production management system. It not only can monitor the identity of the personnel in the process of picking up dead fish, applying medicine, and inspecting and other farm work links in real time, to ensure the compliance of the production process, but also can provide the basis for tracing after the quality and safety event of agricultural products, and provide technical support for the quality and safety of aquatic products. At present, in the supply chain of aquaculture, the identity recognition system of the operating personnel is not perfect, and it is difficult to provide real-time and accurate information records. Early recording methods such as manual signature are prone to omissions and errors, and have weak traceability; existing fingerprint, iris, and face recognition methods require personnel cooperation, which not only increases the workload of personnel, but also easily produces untrustworthy results. Therefore, solving the automatic identification system of the breeding personnel identity in the uncontrolled scene is crucial to improving the production efficiency and the level of traceability of agricultural product quality and safety.

[0027] In the monitoring video of the open farming scene, a large amount of color and texture information irrelevant to gait is covered. Single modal input cannot fully represent the global information of gait. Silhouette retains the overall shape by background difference, but it is difficult to capture the details of limb movement and lacks structured representation. Skeleton can reflect the kinematic characteristics of joints, but it loses the most discriminative appearance information. In addition, in actual aquaculture, there are generally occlusions, blurred long-distance views, and noise residues caused by carrying tools. The recognition accuracy of existing models in real monitoring systems is still insufficient. CNN focuses on set timing feature extraction, ignoring the modeling of inter-frame dependency. The local receptive field of ResNet is too concentrated, which suppresses its global expression ability.

[0028] Based on this, the embodiments of the present application provide a solution, by detecting and target tracking the workers in the aquaculture scene monitoring video, extracting gait sequence frames, generating human silhouette and human analysis diagram according to the gait sequence frames, and outputting the human silhouette and human analysis diagram to the multi-modal gait recognition model, finally obtaining the result of personnel identity recognition. By frame-by-frame fusion of the silhouette of the global shape and the human analysis of the fine structure, the discriminability of gait representation can be effectively enhanced. Through data quality enhancement, the quality enhanced data tends to be classified as the center, improving the intra-class compactness. Using P3D-ECANet structure and 3D Swin Transformer structure to cooperate local details and global context can significantly improve the attention coverage of human key areas, and can perform better recognition ability in the actual aquaculture scene, and improve the accuracy of personnel identity recognition in the aquaculture environment.

[0029] Figure 1 is the flowchart of the multi-modal gait recognition method for aquaculture personnel identity recognition provided by the present application, as shown in Figure 1 The method comprises the following steps 101, 102 and 103.

[0030] Step 101, detecting workers in the aquaculture scene monitoring video and target tracking, extracting gait sequence frames based on the target frame.

[0031] Specifically, the detection process can identify the workers in the breeding scene and obtain the target box of the human body position of the workers. The target tracking is a process of associating the position of the same worker in each subsequent frame of the breeding scene monitoring video based on the worker target detected in the initial frame of the breeding scene monitoring video, thereby forming a continuous motion trajectory of the worker. The target box is a minimum rectangular box containing the worker, and the position and range of the worker in the image can be determined through the coordinate parameters of the target box. The gait sequence is a set of continuous image frames describing the walking process of the worker, and each frame records the posture and position of the worker at this moment, and the playback of the entire gait sequence can restore the dynamic process of the worker walking.

[0032] In some embodiments, the breeding scene monitoring video is obtained by a camera, which can be adjusted to be at an angle of 45 degrees with the horizontal ground, or other suitable angles, which are not limited herein. The video acquisition resolution can be set to 1920x1080, and the frame rate can be set to 30fps, or other suitable resolutions and frame rates, which are not limited herein. Figure 2 is a breeding worker video data acquisition schematic diagram provided by the embodiment of the present application, as shown in the figure, the first behavior is a picture obtained by cutting the worker video obtained by the indoor camera, the second behavior is a picture obtained by cutting the worker video obtained by the outdoor camera, the first column is a picture obtained by cutting the worker video of the picking gait, the second column is a picture obtained by cutting the worker video of the pesticide application gait, and the third column is a picture obtained by cutting the worker video of the inspection gait.

[0033] In some embodiments, a pre-trained YoLox model can be used to detect the workers in the breeding scene monitoring video, and other target detection models can also be used for detection, which are not limited herein. After detecting the workers in the breeding scene monitoring video, the workers in the breeding scene can be identified, and the target box of the human body position of the workers in the breeding scene monitoring video can be obtained. Based on the detection result, a pre-trained Bytetrack model can be used to track the workers in the breeding scene monitoring video, and other target tracking models can also be used for tracking, which are not limited herein. The target tracking process can associate the same worker in different frames of the breeding scene monitoring video, form a continuous walking process of the worker, and obtain the complete trajectory of the motion process of the worker in the breeding scene.

[0034] In some embodiments, after the target frame is obtained, the target frame is filled as a gait image frame. The filling method can be direct scaling, edge filling after equal scaling, or cutting after equal scaling. Other image filling methods can also be used, which are not limited herein. The filled gait image frame can be an image of 192x192 pixels, or other pixel sizes, which are not limited herein.

[0035] In some embodiments, the obtained gait image frame can be preset as a sequence of 30-80 frames to obtain a gait sequence. The method of presetting the gait image frame as a gait sequence is not limited herein, and other frame numbers of gait sequences can also be preset, which are not limited herein. Each gait sequence can be saved as a pkl format file to increase the operation speed of the model, or saved as other format files, which are not limited herein.

[0036] Step 102, extracting a human connected domain from the gait sequence frame to generate a human silhouette image, and performing human parsing on the gait sequence frame to segment a human parsing image of 7 semantic parts.

[0037] Specifically, the human connected domain is a region in which the pixel values belonging to the human body are connected in the image. The human silhouette image is a binary image containing only the contour of the human body, which is generated by separating the human body from the background through image segmentation. Human parsing refers to the process of segmenting pixel points in a human image according to different semantic attributes to obtain semantic part regions corresponding to different semantic attributes. The 7 semantic parts can be head, torso, upper arm, lower arm, upper limb, lower limb, and background. The human parsing image is an image obtained by dividing the human image according to semantic attributes, which is used to visually present the structure of the human body.

[0038] In some embodiments, after obtaining the gait sequence of the worker in step 101, the human connected domain can be extracted from the gait sequence frame based on the SCL semantic segmentation algorithm to generate the human silhouette image. Other semantic segmentation algorithms can also be used for extraction, which are not limited herein. The human body can also be parsed using the CE2P human parsing model to segment the human parsing image of 7 semantic parts. Other human parsing models can also be used for parsing, which are not limited herein.

[0039] Step 103, inputting the human silhouette image and the human parsing image into the multi-modal gait recognition model to obtain a personnel identity recognition result output by the multi-modal gait recognition model.

[0040] The multi-modal gait recognition model is trained based on multi-modal data of a plurality of gait samples with personnel identity labels, and the multi-modal data includes human silhouette graphs and human analysis graphs.

[0041] Specifically, the personnel identity label is an identifier identifying a specific personnel individual to which the gait sample belongs, and each gait sample corresponds to one personnel identity label. The feature information refers to information obtained by performing feature extraction on the human silhouette graph and the human analysis graph of each frame of the breeding scene monitoring video. Frame-by-frame fusion refers to a process of frame-by-frame fusion of the feature information extracted for each frame.

[0042] In some embodiments, after obtaining the human silhouette graph and the human analysis graph in step 102, a data set of the multi-modal gait recognition model can be made using the human silhouette graph and the human analysis graph. The data set can be divided into a training data set and a test data set according to a ratio of 6:4, or can be divided according to other ratios, which are not limited herein. The test data set can be divided into query samples and control samples, 9 control samples can be set for each personnel identity recognition label, and the rest can be set as query samples.

[0043] In some embodiments, in order to simulate the diversity of real breeding scene data, dynamic data augmentation can be performed on the data. The data augmentation method can use random horizontal rotation, random angle rotation, random perspective transformation, or other data augmentation methods, which are not limited herein.

[0044] In some embodiments, according to the real personnel information corresponding to each data, the data set can be labeled with the corresponding personnel identity label, and each data sample corresponds to one personnel identity label. The multi-modal gait recognition model is trained using the training data set. The human silhouette graph and the human analysis graph are input into the trained multi-modal gait recognition model, the multi-modal gait recognition model extracts features from the human silhouette graph and the human analysis graph to obtain feature information, and frame-by-frame fusion is performed on the feature information obtained for each frame, and finally the personnel identity recognition result is output.

[0045] Figure 3 is a multi-modal gait recognition method flowchart provided by the embodiment of the present application, as shown in Figure 3As shown, step S1: the video of the breeding scene operation personnel can be collected by a Hikvision camera. Step S2: the video is preprocessed, which can include using a pre-trained YoLox model to detect the operation personnel in the breeding scene monitoring video, and then using a pre-trained Bytetrack model to track the operation personnel in the breeding scene monitoring video, filling the obtained operation personnel target frame into a gait image frame, extracting a human connected domain from the gait sequence frame to generate a human silhouette image based on an SCL semantic segmentation algorithm, and also segmenting the human body into a human analysis image of 7 semantic parts based on a CE2P human analysis model. Step S3: the human silhouette image and the human analysis image can be used to make a data set of a multi-modal gait recognition model, and the data set is labeled with corresponding personnel identity labels, and each data sample corresponds to a personnel identity label. Step S4: a multi-modal gait recognition model is constructed. Step S5: the network parameters of the multi-modal gait recognition model are initialized, and the multi-modal gait recognition model can be trained using the data set. Step S6: the gait features (human silhouette image, human analysis image) obtained from real-time gait samples are input into the multi-modal gait recognition model, and the identity of the operation personnel is output, and the identity of the operation personnel is automatically recorded.

[0046] The multi-modal gait recognition method for breeding personnel identity recognition provided by the present application detects and tracks the operation personnel in the breeding scene monitoring video, extracts gait sequence frames to generate a human silhouette image and a human analysis image, and outputs the human silhouette image and the human analysis image to a multi-modal gait recognition model, and finally obtains the result of personnel identity recognition. The silhouette of the global shape and the human analysis of the fine structure are fused, the discriminability of the gait representation is enhanced, the multi-modal gait recognition model can realize the cooperation of local details and global context, significantly improve the attention coverage of human key areas, and can extract more discriminative feature representations. In a complex breeding environment and a multi-person collaborative operation scene, the identity of the operation personnel can still be accurately identified, and the accuracy of personnel identity recognition in a breeding environment is improved.

[0047] In some embodiments, the multi-modal gait recognition model includes a first feature extraction module, a second feature extraction module, and a classification prediction module. The first feature extraction module is used to perform preliminary feature extraction on the human silhouette image and the human analysis image based on a 2D-ECANet network, and the extracted silhouette gait features and analysis gait features are combined by using an early attention fusion mechanism to obtain fusion features; The second feature extraction module is used to input the fusion features into a P3D-ECANet network to aggregate spatial-temporal information and strengthen local dynamic features, and then input the features output by the P3D-ECANet network into a 3D Swin Transformer network to capture long-distance gait motion information; The classification prediction module is configured to process the features output by the 3D Swin Transformer network into local feature vectors through a temporal pooling layer and a horizontal pooling layer, map the local feature vectors to an identity metric space, and output a personnel identity recognition result.

[0048] Specifically, the 2D-ECANet network is an efficient channel attention (ECA) network for preliminary feature extraction of two-dimensional data. The early attention fusion mechanism refers to a mechanism for assigning weights to features of different sources through an attention mechanism at an early stage after feature extraction, and then obtaining fused features.

[0049] The spatio-temporal information aggregation refers to the fusion of information in the time dimension and the space dimension. The local dynamic feature enhancement refers to the enhancement of dynamic change features in a local region. The fused features can be input to the P3D-ECANet for spatio-temporal information aggregation and local dynamic feature enhancement. The long-distance gait motion information refers to the motion information of the gait under long-distance conditions. Long distance can refer to a long time interval between gait sequence frames in the time dimension, or a large distance between different motion part regions in the same frame in the space dimension. For the case where the time interval between gait sequence frames is long or the distance between different motion part regions in the same frame is large, the Swin Transformer network can be used to capture long-distance gait motion information.

[0050] The temporal pooling layer is a processing module for down-sampling features in the time dimension, and the horizontal pooling layer is a processing module for down-sampling features in the space dimension. The local feature vector refers to a vector extracted from a local region of an image and used to describe the features of the local region. The identity metric space refers to a space in which features of the same identity sample are close to each other and features of different identity samples are far away from each other.

[0051] In some embodiments, a human silhouette image frame and a human parsing image frame corresponding to a gait sequence frame can be obtained, the human silhouette image frame can be represented as , and the human parsing image frame can be represented as where T is the length of the gait sequence frame, C is the number of channels of the human silhouette image and the human parsing image frame, H and W are the height and width of the human silhouette image and the human parsing image frame, respectively. The first feature extraction module can input the human silhouette image and the human parsing image to the 2D-ECANet network for feature extraction to obtain corresponding silhouette gait features and parsing gait features. After preliminary feature extraction, an early attention fusion mechanism can be used to assign different weights to the silhouette gait features and the parsing gait features, and the weights can be combined according to the weights to obtain a fused feature , This represents a four-dimensional tensor with the shape of .

[0052] In some embodiments, after obtaining the fused features, the second feature extraction module can input the fused features into the P3D-ECANet network, and output the features after spatiotemporal information aggregation and local dynamic feature enhancement. Due to the limitation of the small receptive field of convolution, it is difficult to directly model long-distance spatiotemporal dependencies. However, the Swin Transformer network can directly model global or long-distance spatiotemporal dependencies and is good at capturing spatiotemporal interactions of different parts of the body. Therefore, the Swin Transformer network can be used to capture long-distance gait motion information and output sequential motion information.

[0053] In some embodiments, the classification prediction module can first aggregate the sequential motion information output by the Swing Transformer network into a temporal pooling layer. The overall aggregated features can be represented as follows: Where t represents the number of time frames, This represents the features aggregated from the corresponding first time frame. This represents the aggregated features corresponding to the t-th time frame.

[0054] In some embodiments, after time pooling, the aggregated features can be... After being divided into P sub-sub ... , Indicates the number of output channels.

[0055] In some embodiments, after obtaining the local feature vectors, a fully connected layer can be used to further process the local feature vectors. Mapping to the identity metric space brings positive sample pairs closer together and negative sample pairs further apart. For each batch of gait sequence samples, this can be obtained by mapping using a fully connected layer. Each identity metric space can be obtained based on the identity label. and A number of positive and negative sample pairs. During the training of the multimodal gait recognition model, the triplet loss of the fully connected layer can be calculated. The triplet lost It can be represented as: Where j represents The index of a sample pair in a sample pair, k represents The index of a sample pair in a set of P sample pairs, where i represents the index of the gait local feature vector formed in the P sub-blocks. and Let represent the distances between the i-th gait local features in the j-th positive sample pair and the k-th negative sample pair, respectively. This indicates taking a positive function. Indicates the loss tolerance interval.

[0056] In some embodiments, a batch normalization layer can be used to further eliminate identity metric bias, feature vector Identity prediction tags The probability can be expressed as Given the real label During the training of a multimodal gait recognition model, the cross-entropy loss of the batch normalization layer can be calculated. The cross-entropy loss It can be represented as: Where r represents The index of a sample in an identity metric space, when When, predict probability ,when When, predict probability .

[0057] In some embodiments, the joint loss used to train the multimodal gait recognition model It can be represented as: in, Representing the triplet loss of the fully connected layer, respectively. and the cross-entropy loss of the batch normalization layer The corresponding weight hyperparameters, A multimodal gait recognition model can output personnel identification results.

[0058] The multimodal gait recognition method for identifying livestock workers provided by this invention first uses a 2D-ECANet network to perform preliminary feature extraction from human silhouette and human analytical images. Then, the fused features are input into a P3D-ECANet network for spatiotemporal information aggregation and local dynamic feature enhancement. Afterward, the data is input into a 3D Swin Transformer network to capture long-distance gait motion information. The output is processed into local feature vectors and mapped to an identity metric space to obtain the personnel identification result. By employing a fusion strategy to co-optimize silhouette and human analytical modalities, and through spatiotemporal information aggregation, local dynamic feature enhancement, and the capture of long-distance gait motion information, the fusion of multimodal features effectively enhances the discriminative power of gait representation, leading to more accurate identification of livestock workers' identities.

[0059] In some embodiments, the P3D-ECANet network is a P3D network with an ECA module introduced. The ECA module is connected after the second spatial convolution of the P3D network and is used for local dynamic feature enhancement.

[0060] Specifically, the P3D network is a network structure that uses spatial and temporal convolutions. The P3D network can reduce memory consumption without weakening the ability to model spatiotemporal features. The second spatial convolution in the P3D network is a spatial convolution, and the ECA module is connected after this spatial convolution, which can be used to enhance the representation of local dynamic features of gait.

[0061] In some embodiments, Figure 4 The P3D-ECANet structure diagram provided in the embodiments of the present invention is as follows: Figure 4 As shown, each convolutional (Conv) layer uses a kernel size of T2×H2×W2, where T2 is the length of the input fused feature image sequence, representing temporal feature extraction. H2 and W2 are the height and width of the fused feature image, respectively, representing spatial feature extraction. After the fused features are input into the P3D-ECANet structure, they first pass through a 1×3×3 convolutional layer. This convolutional layer is a spatial convolution, which extracts spatial features of a single frame image using a 3×3 kernel in both the height and width directions. The extracted spatial features then pass through a 3×1×1 convolutional layer, which is a temporal convolution. T2=3 indicates that temporal features related to each frame are extracted every three frames. Next, the residual unit adds the spatial features obtained from the 1×3×3 spatial convolution and the temporal features obtained from the 3×1×1 temporal convolution, performing residual fusion of the spatial and temporal features. A Rectified Linear Unit (ReLU) activation function is then used to add non-linearity. The residual unit output then passes through a 1×3×3 convolutional layer, which is also a spatial convolution. This layer extracts spatial features from the single-frame image using 3×3 kernels in both the height and width directions, further refining the spatial features from the residual-fused features.

[0062] In some embodiments, an ECA module can be connected after the second spatial convolution of the P3D network, such as... Figure 4 As shown, Figure 4The specific architecture of the ECA module is shown on the right side. First, different scales of spatiotemporal features can be extracted through an Inception structure. The Inception structure is a network structure assembled by multiple convolution or pooling operations. The Inception V1 structure in the existing Inception structure can be used, or the Inception V2 structure in the existing Inception structure can be used, or other suitable Inception structures can be used, which are not limited here. After the Inception structure, a global average pooling operation can be used to obtain the input feature image information Aggregate each channel descriptor The calculation formula can be represented as: Among them, the input feature image information The output channel descriptor B can represent the batch of input feature images, T3 can represent the length of the input feature image sequence, H3 and W3 are the height and width of the input feature image respectively, C3 can represent the channel number of the input feature image, and i and j are the coordinate indexes of the input feature image row and column respectively. After the global average pooling operation, for a given channel number, the cross-channel dependency relationship can be obtained through one-dimensional channel adaptive convolution. The convolution kernel size k represents the approximate range of channel interaction information. The calculation formula can be represented as: Among them, represents the nearest odd number, and can be set to 2 and 1, or other suitable values, which are not limited here. After one-dimensional channel adaptive convolution, the Sigmoid function can be used to depict all channel feature weights The calculation formula can be represented as: Among them is one-dimensional convolution, is a Sigmoid function, and the channel feature weight . Then the normalized weight can be weighted to the feature of each channel, and the calculation formula can be represented as: Among them, the ECA module output .

[0063] The application provides a multi-modal gait recognition method for breeding personnel identity recognition.

[0064] In some embodiments, in the 3D SW-MSA module of the 3D Swin Transformer network, a scaled cosine function is used to replace the original dot product to calculate the position similarity.

[0065] Specifically, based on the limitation of the convolution small receptive field, it is difficult to directly model long-distance spatiotemporal dependencies, and the use of the 3DSwin Transformer network can directly model global or long-distance spatiotemporal dependencies, and is good at capturing the spatiotemporal interaction of different parts of the body. Figure 5 is a 3D Swin Transformer structure diagram provided by the embodiment of the application, as Figure 5 shown, first, the 3D Window Multi-Head Self-Attention (3DW-MSA) module and the 3D Shifted Window Multi-Head Self-Attention (3DSW-MSA) module and two layers of Feed-Forward Neural (FFN) network can be divided by a three-dimensional window, the middle can be connected by a Gaussian Error Linear Unit (GELU) nonlinear, and the layer normalization (Layer Normalization, LN) operation is performed before each self-attention module and the feed-forward neural network, then the residual connection is performed to output the features, and the calculation formula can be represented as: wherein, l can represent the index of the 3D Swin Transformer structure layer, z l-1 may represent the output of the previous layer of the current layer, which is also the input of the current layer. z l-1 After the layer normalization processing and the 3DW-MSA structure, the output is added to the original input z l-1 of the current layer to obtain , After the layer normalization processing and the FFN network, is added to again to obtain , The output of the current layer, which is also the input of the next layer of the current layer. l After layer normalization processing and 3DSW-MSA structure, the output is added to the original input z l of the current layer to obtain residual fusion After layer normalization processing and FFN network, the output is added to again to obtain residual fusion z l+1 , z l+1 The output of the next layer of the current layer.

[0066] As shown in Figure 5 , Figure 5 the right side of the middle is the structure of 3DSW-MSA in the 3D Swin Transformer. Assuming that the input of the 3DSW-MSA structure is Z1, the input Z1 can first obtain the query, key, and value matrices through three linear layers. Among them, W Q can represent the query weight matrix, and the query matrix Q can be represented as , W K can represent the key weight matrix, and the key matrix K can be represented as , W V can represent the value weight matrix, and the value matrix V can be represented as .

[0067] In some embodiments, after obtaining the query matrix Q, the key matrix K, and the value matrix V corresponding to the input Z1, since the cosine similarity range is 0-1, the similarity value distribution is smoother, and the scaled cosine attention method can be used instead of the original dot product to calculate the position and the position Similarity of the position, avoiding extreme differences in numerical scales of dot product similarity, thereby reducing the problem of excessive concentration of attention weights. The original dot product can calculate the position and the position Similarity of the position can be represented as , and the scaled cosine attention method can calculate the position and the position Similarity of the position can be represented as , where is the query vector of the nth position in the query matrix Q, is the query vector of the mth position in the key matrix K, is a learnable parameter.

[0068] The application provides a multi-modal gait recognition method for breeding personnel identity recognition.

[0069] In some embodiments, in the 3D SW-MSA module of the 3D Swin Transformer network, a logarithmic interval continuous position bias is used to replace the original linear interval bias.

[0070] Specifically, in the 3D SW-MSA module of the 3D Swin Transformer network, the original linear interval bias is used, for example, assuming that the coordinates of position are ( , ), the coordinates of position are ( , ), and the linear interval bias of position , , can represent the linear interval bias of the pixel pair in the horizontal direction, that is, the difference between the horizontal coordinates of the two pixel points, can represent the linear interval bias of the pixel pair in the vertical direction, that is, the difference between the vertical coordinates of the two pixel points.

[0071] In order to obtain a larger attention window, a logarithmic interval continuous position bias can be introduced to replace the original linear interval bias, and the long-distance position representation range can be compressed. The bias value of an arbitrary coordinate range can be generated by a small fully connected layer neural network according to the shielding condition of the input image, and the position bias can be dynamically adjusted. The logarithmic interval continuous position bias can be represented as: wherein, is the logarithmic interval continuous position bias of the pixel pair in the horizontal direction, is the logarithmic interval continuous position bias of the pixel pair in the vertical direction.

[0072] After obtaining the logarithmic interval continuous position bias, the and can be input into a small fully connected layer neural network. The small fully connected layer neural network can be an FFN network or other suitable small fully connected layer neural network, which is not limited herein. The relative position bias of position and position ​ It can be represented as: Here, G can represent a small neural network. (The location is then determined.) and location relative position deviation and location and location After determining the positional similarity, the relative positional deviation and positional similarity can be summed, and the summed output is... It can be represented as: Then you can Attention weights are generated by normalization using the softmax function, and the output of the 3D SW-MSA module is obtained by multiplying the attention weights with the corresponding value matrix.

[0073] The multimodal gait recognition method for identifying livestock farmers provided by this invention uses log-interval continuous position offsets instead of the original linear interval offsets in the 3D SW-MSA module, which can compress the long-distance position representation range. Based on the occlusion of the input image, a small fully connected neural network generates offset values ​​for any coordinate range and dynamically adjusts the position offsets, which can more accurately identify the identity information of livestock farmers.

[0074] In some embodiments, the first feature extraction module is further configured to perform quality enhancement processing on the human silhouette image and the human analytical image before performing feature extraction on the human silhouette image and the human analytical image based on the 2D-ECANet network.

[0075] Specifically, quality enhancement processing is a process performed to improve image quality. Before inputting human silhouette images and human detailed images into the 2D-ECANet network, quality enhancement processing can be performed on the human silhouette images and human detailed images to remove unqualified samples and reduce noise interference.

[0076] In some embodiments, in complex and crowded outdoor scenes, human silhouette images may contain segmentation errors due to background noise, fishing nets, and pesticide residue, affecting training performance. Therefore, the maximum connected region of the silhouette can be calculated and a threshold can be set. When the largest connected region When, the entire silhouette can be discarded. At this time, noise residue in non-connected small areas can be eliminated, while the main area is preserved. Figure 6 This is a schematic diagram of gait dataset quality enhancement processing provided in an embodiment of the present invention, as shown below. Figure 6 As shown, Figure 6(a) The upper picture in the middle does not use quality enhancement, and the background has noise residues. After quality enhancement, the lower picture can eliminate the noise residues in the non-connected small area, only the area of the worker subject is retained, and the interference of the background noise on detection can be avoided.

[0077] In some embodiments, due to long-distance visual blur, the environmental background is confused with the blurred human body, attention is distracted in extracting the human body analysis, and the uncertainty of human body analysis extraction is increased. Therefore, the enhanced silhouette can be used to generate a binary mask, and the human body silhouette graph and the human body analysis graph can be background masked to suppress the interference of the complex background and generate a more complete human body analysis. As shown in Figure 6 The, Figure 6 (b) The upper left picture in the middle does not use quality enhancement, the environmental background is similar to the blurred human body, and the extraction and analysis of the human body can be difficult. After using the enhanced silhouette to generate a binary mask for background masking, the lower left picture can be obtained, the upper right picture is the human body analysis graph generated by the upper left picture without quality enhancement, and the lower right picture is the human body analysis graph generated by the lower left picture with quality enhancement. After quality enhancement, the generated human body analysis graph is more accurate and clear.

[0078] In some embodiments, when severe occlusion and non-human body shape caused by super long distance occur, a pre-trained prior template can be used to screen samples. The template can include 7 standard view human body shapes at intervals of 30° in the range of 0°-180°, and can also include other view human body shapes, which are not limited herein. The similarity of the input human body silhouette graph and the human body analysis graph to the template can be calculated, and the human body silhouette graph and the human body analysis graph with low similarity can be removed to balance the quality distribution of the data set. As shown in Figure 6 The, Figure 6 (c) The left side in the middle is a pre-trained prior template for screening samples. By calculating the similarity of the input human body silhouette graph and the human body analysis graph to the template, the right side picture with low similarity to the left side template can be screened out. The right side picture is a non-human body shape caused by severe occlusion or long distance.

[0079] The multi-modal gait recognition method for breeding personnel identity recognition provided by the present application can filter out noise irrelevant to the workers by quality enhancement processing of the human body silhouette graph and the human body analysis graph, and can improve the stability of the model. The data after quality enhancement tends to be the classification center, which can allow a larger margin to further compress the feature space and improve the intra-class compactness.

[0080] Figure 7 The multi-modal gait recognition method model structure diagram provided by the embodiment of the present application is shown in Figure 7As shown, first, the breeding scene monitoring video (Video) is collected, a pre-trained YoLox model can be used to detect the workers in the breeding scene monitoring video, identify the workers in the breeding scene, obtain the worker body position target frame, and then use a pre-trained Bytetrack model to track the workers in the breeding scene monitoring video, associate the positions of the same workers in each subsequent frame of the video, and form continuous worker motion trajectories. Fill in the target frame as a gait image frame, divide it into multiple gait sequences, and then perform image segmentation (Segmentation) on the gait sequence frame to generate a human silhouette graph (silhouette). The white part of the human silhouette graph represents the worker, and the black part represents the background. A human parsing model can also be used to perform human parsing (Human Paring) on the gait sequence frame to obtain a human parsing graph. Different colors in the human parsing graph represent different semantic parts of the worker. The human silhouette graph and the human parsing graph are augmented (Augmentation) to obtain the quality-enhanced human silhouette graph and the human parsing graph. Then the picture can be input into the 2D-ECANet network, the extracted silhouette gait features and parsing gait features are weighted and combined using a fusion mechanism to obtain fusion features, and then the fusion features are input into the P3D network with an added ECA module, i.e. P3D-ECANet network, for spatiotemporal aggregation and local dynamic feature enhancement. The output features are input into the 3D Swin Transformer network, and the 3D SW-MSA module in the 3D Swin Transformer network is improved. The scaling cosine function can be used instead of the original dot product to calculate the position similarity, and the logarithmic interval continuous position bias can be used instead of the original linear interval bias. The features output by the 3D Swin Transformer network can be processed through a fully connected layer (Fully Connected Layer, FC) and a batch normalization (Batch Normalization, BN) layer to output the personnel identity recognition result person01~ person05.

[0081] The multi-modal gait recognition device for breeding personnel identity recognition provided by the present application is described below. The multi-modal gait recognition device for breeding personnel identity recognition described below can be referred to in conjunction with the multi-modal gait recognition method for breeding personnel identity recognition described above.

[0082] Figure 8 The multi-modal gait recognition device for breeding personnel identity recognition provided by the present application is described below. The multi-modal gait recognition device for breeding personnel identity recognition described below can be referred to in conjunction with the multi-modal gait recognition method for breeding personnel identity recognition described above. Figure 8 As shown, the device comprises: The extraction unit 810 is configured to detect a worker in the breeding scene monitoring video and perform target tracking, and extract a gait sequence frame based on a target frame; The generation unit 820 is configured to generate a human silhouette graph by extracting a human connected domain from the gait sequence frame, and perform human parsing on the gait sequence frame to segment a human parsing graph of seven types of semantic parts; The recognition unit 830 is configured to input the human silhouette graph and the human parsing graph into a multi-modal gait recognition model to obtain a personnel identity recognition result output by the multi-modal gait recognition model. The multi-modal gait recognition model is trained based on multi-modal data of a plurality of gait samples with personnel identity labels, the multi-modal data includes the human silhouette graph and the human parsing graph, and the multi-modal gait recognition model is used for frame-by-frame fusion of feature information of the human silhouette graph and the human parsing graph, and personnel identity recognition based on the fused feature information.

[0083] In some embodiments, the multi-modal gait recognition model includes a first feature extraction module, a second feature extraction module, and a classification prediction module. The first feature extraction module is configured to perform preliminary feature extraction on the human silhouette graph and the human parsing graph based on a 2D-ECANet network, and combine the extracted silhouette gait features and parsing gait features using an early attention fusion mechanism to obtain fusion features. The second feature extraction module is configured to input the fusion features into a P3D-ECANet network to aggregate spatial and temporal information and strengthen local dynamic features, and then input the features output by the P3D-ECANet network into a 3D Swin Transformer network to capture long-distance gait motion information. The classification prediction module is configured to process the features output by the 3D Swin Transformer network into a local feature vector through a time pooling layer and a horizontal pooling layer, map the local feature vector to an identity metric space, and output a personnel identity recognition result.

[0084] In some embodiments, the P3D-ECANet network is a P3D network with an ECA module, the ECA module is connected after the second spatial convolution of the P3D network, and is used for local dynamic feature enhancement.

[0085] In some embodiments, in the 3D SW-MSA module of the 3D Swin Transformer network, a scaled cosine function is used to replace the original dot product to calculate the position similarity.

[0086] In some embodiments, in the 3D SW-MSA module of the 3D Swin Transformer network, a logarithmic interval continuous position bias is used to replace the original linear interval bias.

[0087] In some embodiments, the first feature extraction module is further configured to perform quality enhancement processing on the human silhouette image and the human parsing image before performing feature extraction on the human silhouette image and the human parsing image based on the 2D-ECANet network.

[0088] It should be noted that the multi-modal gait recognition device for identifying the identity of the breeding personnel provided by the present application can realize all the method steps realized by the method embodiments and achieve the same technical effects, and the same parts and beneficial effects in the present embodiment as the method embodiments will not be described in detail.

[0089] Figure 9 is a structural schematic diagram of an electronic device provided by the present application, as Figure 9 shown, the electronic device can include a processor 910, a communications interface 920, a memory 930 and a communications bus 940, wherein the processor 910, the communications interface 920 and the memory 930 communicate with each other through the communications bus 940. The processor 910 can invoke the logical instructions in the memory 930 to execute the multi-modal gait recognition method for identifying the identity of the breeding personnel.

[0090] In addition, the logical instructions in the memory 930 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0091] It should be noted that the electronic device provided by the present application can realize all the method steps realized by the method embodiments and achieve the same technical effects, and the same parts and beneficial effects in the present embodiment as the method embodiments will not be described in detail.

[0092] In another aspect, the present application also provides a non-transitory computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the above-mentioned multi-modal gait recognition method for breeding personnel identity recognition.

[0093] It should be noted that the above-mentioned non-transitory computer readable storage medium provided by the present application can realize all method steps realized by the method embodiment and achieve the same technical effects, and the same parts and beneficial effects in the present embodiment as the method embodiment will not be described in detail.

[0094] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program being stored on a non-transitory computer readable storage medium, and when the computer program is executed by a processor, the computer can execute the above-mentioned multi-modal gait recognition method for breeding personnel identity recognition.

[0095] It should be noted that the above-mentioned computer program product provided by the present application can realize all method steps realized by the method embodiment and achieve the same technical effects, and the same parts and beneficial effects in the present embodiment as the method embodiment will not be described in detail.

[0096] The device embodiments described above are only schematic, and the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e. they may be located in one place, or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement it without creative labor.

[0097] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary general hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0098] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A multimodal gait recognition method for identifying livestock farmers, characterized in that, include: Detect and track workers in monitoring videos of aquaculture scenes, and extract gait sequence frames based on the target bounding boxes; Human connected components are extracted from the gait sequence frames to generate human silhouette images, and human parsing is performed on the gait sequence frames to segment human images into 7 semantic parts; The human silhouette image and the human analytical image are input into the multimodal gait recognition model to obtain the personnel identification result output by the multimodal gait recognition model; The multimodal gait recognition model is trained based on multimodal data of multiple gait samples with personnel identification labels. The multimodal data includes human silhouette images and human analytical images. The multimodal gait recognition model is used to fuse the feature information of the human silhouette images and the human analytical images frame by frame, and to perform personnel identification based on the fused feature information.

2. The multimodal gait recognition method for identifying livestock farmers according to claim 1, characterized in that, The multimodal gait recognition model includes a first feature extraction module, a second feature extraction module, and a classification prediction module; The first feature extraction module is used to perform preliminary feature extraction on the human silhouette image and the human analytical image based on the 2D-ECANet network, and to use an early attention fusion mechanism to weight and combine the extracted silhouette gait features and analytical gait features to obtain fused features; The second feature extraction module is used to input the fused features into the P3D-ECANet network for spatiotemporal information aggregation and local dynamic feature enhancement, and then input the features output by the P3D-ECANet network into the 3D Swin Transformer network to capture long-distance gait motion information; The classification prediction module is used to process the features output by the 3D Swin Transformer network into local feature vectors through temporal pooling layers and horizontal pooling layers, and then map the local feature vectors to the identity metric space to output the personnel identity recognition result.

3. The multimodal gait recognition method for identifying livestock farmers according to claim 2, characterized in that, The P3D-ECANet network is a P3D network that incorporates an ECA module. The ECA module is connected after the second spatial convolution of the P3D network and is used for local dynamic feature enhancement.

4. The multimodal gait recognition method for identifying livestock farmers according to claim 2, characterized in that, In the 3D SW-MSA module of the 3D Swin Transformer network, the scaled cosine function is used instead of the original dot product to calculate positional similarity.

5. The multimodal gait recognition method for identifying livestock workers according to claim 2 or 4, characterized in that, In the 3D SW-MSA module of the 3D Swin Transformer network, logarithmically spaced continuous position biases are used instead of the original linear biases.

6. The multimodal gait recognition method for identifying livestock farmers according to claim 2, characterized in that, The first feature extraction module is further configured to perform quality enhancement processing on the human silhouette image and the human analytical image before performing feature extraction on the human silhouette image and the human analytical image based on the 2D-ECANet network.

7. A multimodal gait recognition device for identifying livestock farmers, characterized in that, include: The extraction unit is used to detect workers in the monitoring video of the aquaculture scene and perform target tracking, and extract gait sequence frames based on the target bounding box; The generation unit is used to extract human connected components from the gait sequence frame to generate a human silhouette image, and to perform human body parsing and segmentation on the gait sequence frame to obtain a human body parsing image with 7 semantic parts. The recognition unit is used to input the human silhouette image and the human analytical image into the multimodal gait recognition model to obtain the personnel identification result output by the multimodal gait recognition model; The multimodal gait recognition model is trained based on multimodal data of multiple gait samples with personnel identification labels. The multimodal data includes human silhouette images and human analytical images. The multimodal gait recognition model is used to fuse the feature information of the human silhouette images and the human analytical images frame by frame, and to perform personnel identification based on the fused feature information.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the multimodal gait recognition method for identifying livestock personnel as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multimodal gait recognition method for identifying livestock personnel as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the multimodal gait recognition method for identifying livestock personnel as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Adaptive gait recognition method based on multi-mode dynamic complementation

    CN121527852A

  • An adaptive gait recognition method based on multi-modal dynamic complement

    CN121527852B