A method, device, medium and product for processing human body trajectory sequence

By extracting global and local features from human trajectory sequences to remove noise and splitting problems and merging trajectory sequences of the same human body, the accuracy and stability issues of human weight recognition in complex and changing scenarios are solved, and efficient model updates and accuracy improvements are achieved.

CN120183001BActive Publication Date: 2025-09-12ZHE JIANG SHEN XIANG ZHI NENG KE JI YOU XIAN GONG SI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510641611.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-12
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

Existing technologies perform poorly in human weight recognition in complex and changing scenarios. This is limited by the storage space and inference speed of edge devices, and there are noise and fragmentation problems, resulting in low accuracy, poor generalization ability, and unstable updates of human weight recognition models.

Method used

By obtaining multiple human body trajectory sequences, extracting global and local features respectively, removing non-target human images, and merging trajectory sequences belonging to the same human body, the human weight recognition model is updated using incremental learning, and the clustering model of global and local features is used to remove noise and splitting problems, thereby improving the accuracy and precision of the model.

Benefits of technology

It effectively removes noise and splitting problems in human trajectory sequences, improves the accuracy and precision of the human weight recognition model, adapts to complex and changing scenarios, reduces error accumulation, and enhances the model's generalization ability and update stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183001B_ABST
    Figure CN120183001B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a method, device, medium, and product for processing a human body trajectory sequence. The method includes: obtaining multiple human body trajectory sequences, wherein the human body trajectory sequences include: human body images of at least one human body collected at multiple different times; extracting the global features and local features of each human body image in the human body trajectory sequence; based on the global features and the local features, removing human body images of non-target human bodies in the human body trajectory sequence, and determining multiple filtered human body trajectory sequences; wherein, in the human body trajectory sequence, the number of human body images of non-target human bodies is less than the number of human body images of target human bodies; based on the global features of each human body image in the multiple filtered human body trajectory sequences, determining human body trajectory sequences belonging to the same human body and merging them to determine a merged human body trajectory sequence. The method can reduce the noise in the trajectory sequence and improve the accuracy of the trajectory sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a method, device, medium, and product for processing a human body trajectory sequence. Background Art

[0002] With the rapid development of artificial intelligence, traditional offline retail outlets such as shopping malls and supermarkets are also undergoing a digital transformation. The ubiquity of various smart devices within shopping malls not only provides convenience for consumers but also improves management efficiency and reduces labor costs. The ubiquitous cameras within shopping malls and supermarkets also provide computer vision-based solutions for applications such as loss prevention, pedestrian detection, customer flow counting, and consumer behavior statistics, providing a theoretical basis for businesses to understand their operational status.

[0003] Currently, most technologies for loss prevention, pedestrian detection, and customer counting in shopping malls and supermarkets rely on human weight recognition. Human weight recognition typically uses algorithms based on the Transformer (an attention-based algorithm) and deep convolutional neural networks. However, its application is significantly limited by the storage space constraints and inference speed requirements of edge devices.

[0004] In addition, in some special scenarios, such as complex and changeable factors, such as location changes, seasonal changes, lighting changes, shooting angle changes, occlusion, etc., the human weight recognition effect is unsatisfactory and the accuracy is low. Summary of the Invention

[0005] The embodiments of the present application illustrate a method, device, medium, and product for processing a human body trajectory sequence.

[0006] In a first aspect, an embodiment of the present application illustrates a method for processing a human body trajectory sequence, the method comprising:

[0007] Acquire multiple human body trajectory sequences, where each human body trajectory sequence includes at least: human body images of at least one human body acquired at multiple different moments;

[0008] For any human body trajectory sequence among the multiple human body trajectory sequences, respectively extracting global features of each human body image in the human body trajectory sequence, and respectively extracting local features of each human body image in the human body trajectory sequence;

[0009] According to the global features and the local features, human body images of non-target humans are removed from the human body trajectory sequence to determine a plurality of screened human body trajectory sequences; wherein, in the human body trajectory sequence, the number of human body images of non-target humans is less than the number of human body images of target humans;

[0010] According to the global features of the human body images in the multiple screened human body trajectory sequences, human body trajectory sequences belonging to the same human body are determined and merged to determine a merged human body trajectory sequence.

[0011] In an optional implementation, removing human body images of non-target human bodies in the human body trajectory sequence based on the global features and the local features includes:

[0012] Removing human body images of non-target human bodies from the human body trajectory sequence according to global features of each human body image in the human body trajectory sequence;

[0013] According to the local features of the remaining human body images in the human body trajectory sequence, human body images of non-target human bodies are removed from the remaining human body images in the human body trajectory sequence.

[0014] In an optional implementation, the local features of the human body image are a plurality of different types of local features of the human body image;

[0015] The removing of non-target human body images from the remaining human body images in the human body trajectory sequence according to the local features of the remaining human body images in the human body trajectory sequence comprises:

[0016] According to a preset order between different types of local features, the remaining human body trajectory sequences in the human body trajectory sequence are screened according to the local features to remove human body images of non-target human bodies;

[0017] The preset order of priority between the different types of local features is obtained by sorting the different types of local features in descending order according to their respective abilities to represent the characteristics of the human body.

[0018] In an optional implementation, removing human body images of non-target human bodies in the human body trajectory sequence based on the global features of each human body image in the human body trajectory sequence includes:

[0019] Using a first clustering model corresponding to the global feature, clustering the human body images in the human body trajectory sequence according to the global feature, and identifying human body images of non-target humans;

[0020] In the human body trajectory sequence, removing human body images of identified non-target human bodies;

[0021] The first radius of the first clustering model is used to define the range of the neighborhood, and the first number of the second clustering model is used to define the data points required to form a core point in the neighborhood.

[0022] In an optional implementation, removing human body images of non-target human bodies from the remaining human body images in the human body trajectory sequence based on local features of the remaining human body images in the human body trajectory sequence includes:

[0023] Using a second clustering model corresponding to the local features, clustering the remaining human body images in the human body trajectory sequence according to the local features, and identifying human body images of non-target humans;

[0024] Removing the identified non-target human body images from the remaining human body images in the human body trajectory sequence;

[0025] The second radius of the second clustering model is used to define the range of the neighborhood, and the second number of the second clustering model is used to define the data points required to form a core point in the neighborhood; the second radius is smaller than the first radius, and / or the second number is greater than the first number.

[0026] In an optional implementation, the second radius of the clustering model whose local features are ranked higher in type is greater than the second radius of the clustering model whose local features are ranked lower in type;

[0027] And / or, the second number of clustering models whose types of local features are ranked higher is less than or equal to the second number of clustering models whose types of local features are ranked lower.

[0028] In an optional implementation, determining human body trajectory sequences belonging to the same human body and merging them according to the global features of the human body images in the multiple screened human body trajectory sequences to determine the merged human body trajectory sequence includes:

[0029] Selecting a first human body trajectory sequence from the screened human body trajectory sequences, and obtaining first global features of M human body images from the first human body trajectory sequence;

[0030] Acquire a second global feature of each human body image from a second human body trajectory sequence, where the second human body trajectory sequence is a human body trajectory sequence other than the first human body trajectory sequence in the screened human body trajectory sequence;

[0031] determining a feature similarity between the first global feature and the second global feature;

[0032] According to the feature similarity, a human body trajectory sequence that belongs to the same human body as the target human body trajectory is selected from the second human body trajectory sequence as the human body trajectory sequence to be merged;

[0033] The first human body trajectory sequence is merged with the human body trajectory sequence to be merged to determine a merged human body trajectory sequence.

[0034] In an optional implementation, selecting, in the second human trajectory sequence, a human trajectory sequence that belongs to the same human body as the target human trajectory according to the feature similarity as the human trajectory sequence to be merged includes:

[0035] Select the top N feature similarities from the acquired feature similarities;

[0036] According to the first N feature similarities, the second human body trajectory sequence to which the corresponding human body image belongs is screened, and a human body trajectory sequence to be merged that belongs to the same human body as the target human body trajectory sequence is determined.

[0037] In an optional implementation, selecting the first N feature similarities from the acquired feature similarities includes: for any human body image in the M human body images in the first human body trajectory sequence, respectively screening the first K feature similarities.

[0038] In an optional implementation, screening the second human body trajectory sequence to which the corresponding human body image belongs based on the first N feature similarities, and determining the human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence, includes:

[0039] Determine, based on the first N feature similarities, a second human body trajectory sequence to which the corresponding human body image belongs, and the number of feature similarities corresponding to the second human body trajectory sequence;

[0040] According to the number of feature similarities corresponding to the second human body trajectory sequence, a human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence is selected.

[0041] In an optional implementation, screening the human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence according to the number of feature similarities contained in the second human body trajectory sequence includes:

[0042] Calculate the ratio of the number of feature similarities contained in the second human trajectory sequence to N;

[0043] When the ratio is greater than a preset threshold, the corresponding second human body trajectory sequence is determined as a human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence.

[0044] In an optional implementation, the method further includes: performing recognition processing on the combined human body trajectory sequence to determine trajectory information of the target user.

[0045] In a second aspect, an embodiment of the present application shows an electronic device, which includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method shown in any of the aforementioned aspects.

[0046] In a third aspect, an embodiment of the present application shows a non-temporary computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method shown in any of the aforementioned aspects.

[0047] In a fourth aspect, an embodiment of the present application shows a computer program product, including a computer program / computer executable instructions, wherein the computer program / computer executable instructions, when executed by a processor in an electronic device, implement the method shown in any of the aforementioned aspects.

[0048] Compared with the prior art, the embodiments of the present application have the following advantages:

[0049] In an embodiment of the present application, multiple human trajectory sequences are acquired, and one human trajectory sequence includes at least one human body image captured at multiple different times. For any human body trajectory sequence, global features of each human body image in the human body trajectory sequence are extracted, and local features of each human body image in the human body trajectory sequence are extracted. Based on the global features and the local features, human body images of non-target human bodies are removed from the human body trajectory sequence, and multiple filtered human body trajectory sequences are determined. In the human body trajectory sequence, the number of human body images of non-target human bodies is less than the number of human body images of target human bodies, thereby removing human body images that are noise in the human body trajectory sequence and improving the accuracy of the human body trajectory sequence. Based on the global features of each human body image in the multiple filtered human body trajectory sequences, human body trajectory sequences belonging to the same human body are determined and merged to determine the merged human body trajectory sequence, further improving the accuracy of the human body trajectory sequence. This also facilitates subsequent processing based on the human body trajectory sequence. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a flowchart of a method for processing a human body trajectory sequence according to an embodiment of the present application;

[0051] Figure 2 This is a flowchart of the steps of a method for removing non-target human bodies from a human body image according to an embodiment of the present application;

[0052] Figure 3 This is a flowchart of the steps of a method for merging human body trajectory sequences according to an embodiment of the present application;

[0053] Figure 4 This is a flowchart of the steps of a method for obtaining a human body trajectory sequence to be merged according to an embodiment of the present application;

[0054] Figure 5 This is a structural block diagram of a device for processing a human body trajectory sequence according to an embodiment of the present application;

[0055] Figure 6 This is a structural block diagram of a device according to an embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the embodiments of the present application are further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0057] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0058] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0059] In order to cope with the challenges brought by complex and changing scenarios, limited by the requirements for model complexity and inference speed, an automatic iteration solution of the model that can perform incremental learning is needed to cope with complex and changing scenarios, ensure the human weight recognition effect in different scenarios, and realize true Thousand Mirrors and Thousand Models.

[0060] For the automatic iteration scheme of the incremental learning model, it is necessary to obtain incremental training data sets regularly or continuously and update the human weight recognition model based on the incremental training data sets.

[0061] With respect to obtaining an incremental training dataset, there is at least one training data, and one training data may include a human body trajectory sequence and the annotation data of the human body trajectory sequence. A human body trajectory sequence may include human body images of a human body at multiple different moments, and the annotation data of the human body trajectory sequence may be the unique identification information (label or session) of the human body corresponding to the human body trajectory sequence.

[0062] However, the human body trajectory sequence of a certain human body may contain noise problems (for example, the human body trajectory sequence of a certain human body contains human body images of other human bodies) and / or splitting problems (for example, human body images belonging to the same human body are mistakenly divided into multiple different human body trajectory sequences, that is, among multiple human body trajectory sequences, there may be more than two human body trajectory sequences belonging to the human body trajectory sequence of the same human body).

[0063] If there are noise and / or splitting problems in the human trajectory sequence, the accuracy of the subsequent human weight recognition model updated based on the human trajectory sequence will be low or the error will be large.

[0064] For example, training a human weight recognition model is a typical classification problem, and the loss function used in training (such as cross-entropy loss) is designed for classification problems. In classification problems, human trajectory sequences labeled as different categories (different people) must actually belong to different categories (different people). However, when there is a split problem, at least two or more human trajectory sequences contain images of the same person, indicating that the training data labels are incorrect, which will interfere with the update of the human weight recognition model and lead to the following defects:

[0065] The classification accuracy of the human weight recognition model is low. The human weight recognition model has difficulty in accurately distinguishing the true categories of human trajectory sequences, resulting in a decrease in the overall classification accuracy.

[0066] The generalization ability of the human weight recognition model is low and the performance on unseen data may deteriorate because the human trajectory sequence model does not learn the true category features.

[0067] The update process of the human weight recognition model is unstable, the loss value of the loss function may fluctuate greatly, the convergence speed becomes slower, and it may even fail to converge to the ideal solution.

[0068] In particular, when there are more human trajectory sequences with splitting problems and / or noise problems, the cumulative error is worse, and the impact on the overall classification accuracy of the human weight recognition model is greater.

[0069] To this end, the solution of the embodiment of the present application is proposed.

[0070] Before introducing the solutions of the embodiments of the present application, the professional terms that may be involved in the embodiments of the present application are first explained.

[0071] DiKR, Disentangled Key point Regression, decoupled key point regression.

[0072] DBSCAN, Density-Based Spatial Clustering of Applications with Noise, density-based spatial clustering algorithm with noise.

[0073] ViT-large, Vision Transformer-large, visual self-attention.

[0074] ReID, Re-identification, re-identification.

[0075] YOLO x, You Only Look Once x version, the Xth version of an efficient target detection algorithm.

[0076] DETR, DEtection Transformer, an object detection model based on the Transformer architecture.

[0077] Human body weight recognition is an image retrieval task that detects whether human body images in a video are of the same person.

[0078] Incremental learning is a learning method that can continuously learn new knowledge from new samples while preserving most of the previously learned knowledge.

[0079] Session is a unique identifier used to identify an individual person.

[0080] See also Figure 1 , shows a method for processing a human body trajectory sequence according to an embodiment of the present application, which is applied to an electronic device, which may include a terminal or a server. The method may include:

[0081] In step S101 , a plurality of human body trajectory sequences are acquired, wherein one human body trajectory sequence at least includes: human body images of at least one human body acquired at a plurality of different moments.

[0082] In one embodiment of the present application, human bodies can be tracked and detected in a captured video to obtain multiple human trajectory sequences. The video may include video of a certain area captured by a camera, in which a person passes through the area, so that the captured video contains human images. The present application does not limit the specific tracking and detection method.

[0083] In one example, for videos captured in complex and ever-changing scenes at the edge, human body detection technology (e.g., DiKR, etc.) and tracking algorithms (e.g., Kalman filtering, etc.) can be used to obtain multiple uninterrupted human body images (forming human body trajectories). A clustering algorithm (e.g., DBSCAN, etc.) is used to preliminarily combine human body images belonging to the same human body into a human body trajectory sequence. If there are multiple human bodies in the video, multiple human body trajectory sequences can be obtained.

[0084] However, the combined trajectory sequence of a certain human body may contain noise and splitting problems.

[0085] Therefore, it is necessary to solve the noise problem and the splitting problem through the process of step S102 to step S104 in the embodiment of the present application.

[0086] Alternatively, in another embodiment of the present application, the multiple human body trajectory sequences may be the multiple human body trajectory sequences finally obtained after executing step S104, that is, at least some of the multiple human body trajectory sequences obtained in step S101 are human body trajectory sequences merged through step S104.

[0087] There is a sorting order between the multiple human body images in a human body trajectory sequence, and the sorting order is the order of the acquisition time corresponding to the human body images from early to late.

[0088] In step S102 , for any human body trajectory sequence among the multiple human body trajectory sequences, global features of each human body image in the human body trajectory sequence are extracted, and local features of each human body image in the human body trajectory sequence are extracted.

[0089] In step S103, based on the global features and the local features, human body images of non-target humans are removed from the human body trajectory sequence to determine a plurality of screened human body trajectory sequences, wherein the number of human body images of non-target humans in the human body trajectory sequence is less than the number of human body images of target humans.

[0090] For each other human body trajectory sequence in the multiple human body trajectory sequences, the operations of step S102 to step S103 are similarly performed, thereby achieving the removal of human body images of non-target human bodies in each human body trajectory sequence in the multiple human body trajectory sequences.

[0091] In the embodiments of the present application, the target human body is not an absolute concept, and the non-target human body is not an absolute concept. The target human body and the non-target human body are relative concepts. The target human body does not specifically refer to a specific / fixed human body, but is based on the concept of clustering algorithm. In any human body trajectory sequence, if the number of human body images of a certain human body is greater than the number of human body images of other human bodies, then the certain human body is the target human body, and the other human bodies are non-target human bodies.

[0092] As for “respectively extracting local features of each human body image in the human body trajectory sequence”, half-body features of the human body in each human body image in the human body trajectory sequence may be extracted respectively, and shoe features of the shoes worn by the human body in each human body image in the human body trajectory sequence may be extracted respectively.

[0093] The specific process of "removing non-target human body images in the human body trajectory sequence based on the global feature and the local feature" can be found in the following Figure 2 The embodiment shown will not be described in detail here.

[0094] In one embodiment of the present application, a global human weight recognition model can be trained in advance. For example, based on a large number of human images obtained in complex and changing scenes, a ReID model based on ViT-large is trained. The loss functions used in the training process may include: triplet loss, cluster loss or cross entropy loss, etc. The model is continuously iterated until the model converges, thereby obtaining a global human weight recognition model.

[0095] In this way, the global features of each human body image in the human body trajectory sequence can be extracted based on the global human body weight recognition model.

[0096] In another embodiment of the present application, a local human weight recognition model can be trained in advance. For example, based on currently existing detection technologies (such as Yolox and / or DETR, etc.), local images (for example, shoe images of shoes worn by a person, images of the upper body of a person, images of the lower body of a person, or images of the face of a person, etc.) in a large number of human images obtained in complex and changeable scenes can be detected. Then, based on a large number of local images of the same type (for example, either all images of shoes worn by a person, or all images of the upper body of a person, or all images of the lower body of a person, or all images of the face of a person, etc.), a ReID model based on ViT-large is trained. The loss function used in the training process may include: triplet loss, cluster loss, or cross-entropy loss, etc. The model is continuously iterated until the model converges, thereby obtaining a local human weight recognition model.

[0097] In this way, the local images in each human body image in the human body trajectory sequence can be detected based on the currently existing detection technology (such as Yolox and / or DETR, etc.), and then the local features of the local images in each human body image in the human body trajectory sequence can be extracted based on the local human body weight recognition model.

[0098] In an optional implementation, different types of local features can be extracted from each human body image in the human body trajectory sequence. The different types of local features include shoe features of a shoe image of a human body, upper body features of an upper body image of a human body, lower body features of a lower body image of a human body, or facial features of a facial image of a human body.

[0099] At least two of the local human weight recognition models corresponding to shoe images, upper body images, lower body images, and face images can be trained in advance.

[0100] In this way, for shoe images, the shoe images of the shoes worn by the human body in each human body image in the human body trajectory sequence can be detected based on currently existing detection technologies (such as Yolox and / or DETR, etc.), and then the shoe features of the shoe images of the shoes worn by the human body in each human body image in the human body trajectory sequence can be extracted respectively based on the local human body weight recognition model corresponding to the shoe image.

[0101] For another example, for the lower body image, the lower body image of the human body in each human body image in the human body trajectory sequence can be detected based on currently existing detection technologies (such as Yolox and / or DETR, etc.), and then the upper body features of the lower body image in each human body image in the human body trajectory sequence can be extracted respectively based on the local human body weight recognition model corresponding to the lower body image.

[0102] For another example, for facial images, the facial images of the human body in each human body image in the human body trajectory sequence can be detected based on currently existing detection technologies (such as Yolox and / or DETR, etc.), and then the facial features of the facial images in each human body image in the human body trajectory sequence can be extracted based on the local human body weight recognition model corresponding to the facial image.

[0103] For another example, for upper body images, the upper body images of the human body in each human body image in the human body trajectory sequence can be detected based on currently existing detection technologies (such as Yolox, DETR and / or Yolo-pose, etc.), and then the upper body features of the upper body images in each human body image in the human body trajectory sequence can be extracted respectively based on the local human body weight recognition model corresponding to the upper body image.

[0104] Among them, for any human body image in the human body trajectory sequence, the human body key points in the human body image can be detected based on the currently existing detection technology. The human body key points include: the top of the head, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip and right hip, etc.

[0105] The vertical coordinate of the head vertex in the human body image can be obtained and used as the maximum vertical coordinate required for intercepting the partial image.

[0106] The vertical coordinate of the left hip in the human body image, the vertical coordinate of the right hip in the human body image, the vertical coordinate of the left wrist in the human body image, and the vertical coordinate of the right wrist in the human body image can be obtained, and then the largest vertical coordinate among the vertical coordinates of the left hip in the human body image, the vertical coordinate of the right hip in the human body image, the vertical coordinate of the left wrist in the human body image, and the vertical coordinate of the right wrist in the human body image is selected as the minimum vertical coordinate required to capture the partial image.

[0107] The horizontal coordinates of the left shoulder in the human body image, the horizontal coordinates of the right shoulder in the human body image, the horizontal coordinates of the left elbow in the human body image, the horizontal coordinates of the right elbow in the human body image, the horizontal coordinates of the left wrist in the human body image, the horizontal coordinates of the right wrist in the human body image, the horizontal coordinates of the left hip in the human body image, and the horizontal coordinates of the right hip in the human body image can be obtained. Among the horizontal coordinates of the left shoulder in the human body image, the horizontal coordinates of the right shoulder in the human body image, the horizontal coordinates of the left elbow in the human body image, the horizontal coordinates of the right elbow in the human body image, the horizontal coordinates of the left wrist in the human body image, the horizontal coordinates of the right wrist in the human body image, the horizontal coordinates of the left hip in the human body image, and the horizontal coordinates of the right hip in the human body image, the smallest horizontal coordinate is used as the minimum horizontal coordinate required for intercepting the partial image, and the largest horizontal coordinate is used as the maximum horizontal coordinate required for intercepting the partial image.

[0108] Then, the upper body image of the human body is intercepted in the human body image according to the minimum vertical coordinate required for intercepting the partial image, the maximum vertical coordinate required for intercepting the partial image, the minimum horizontal coordinate required for intercepting the partial image, and the maximum horizontal coordinate required for intercepting the partial image.

[0109] In step S104 , based on the global features of the human body images in the plurality of screened human body trajectory sequences, human body trajectory sequences belonging to the same human body are determined and merged to determine a merged human body trajectory sequence.

[0110] According to the global features of each human body image in a plurality of human body trajectory sequences in which no non-target human body images are present, human body trajectory sequences belonging to the same human body are merged in the plurality of human body trajectory sequences in which no non-target human body images are present.

[0111] In one embodiment of the present application, assuming that the multiple human body trajectory sequences obtained in step S101 are n human body trajectory sequences, if there are human body images of non-target humans in all of the n human body trajectory sequences, then in steps S102-S103, the human body images of non-target humans will be removed from the n human body trajectory sequences respectively. In this way, the "multiple human body trajectory sequences without human body images of non-target humans" in step S104 can be: n human body trajectory sequences from which human body images of non-target humans have been removed respectively in steps S102-S103.

[0112] Alternatively, in another embodiment of the present application, assuming that the multiple human trajectory sequences obtained in step S101 are n human trajectory sequences, if there are human images of non-target human bodies in x (x is less than n) human trajectory sequences, If there is no human body image of non-target human body in the human body trajectory sequence, then in steps S102-S103, human body images of non-target human body will be removed from the x human body trajectory sequences respectively. Thus, the “multiple human body trajectory sequences without human body images of non-target human body” in step S104 can be: the x human body trajectory sequences with human body images of non-target human body removed in steps S102-S103 and the human body trajectory sequences without human body images of non-target human body obtained in step S101 Individual body trajectory sequences.

[0113] The details of step S104 can be found in the following Figure 3 The embodiment shown will not be described in detail here.

[0114] For example, there are n human body trajectory sequences in human body images without non-target human bodies. After merging human body trajectory sequences belonging to the same human body in the n human body trajectory sequences in human body images without non-target human bodies, q human body trajectory sequences are obtained, where q is less than n.

[0115] In an embodiment of the present application, multiple human body trajectory sequences are acquired, and one human body trajectory sequence includes at least: human body images of at least one human body acquired at multiple different times. For any human body trajectory sequence, global features of each human body image in the human body trajectory sequence are extracted, and local features of each human body image in the human body trajectory sequence are extracted. Based on the global features and the local features, human body images of non-target human bodies are removed from the human body trajectory sequence, wherein the number of human body images of non-target human bodies in the human body trajectory sequence is less than the number of human body images of target human bodies. Based on the global features of each human body image in the multiple human body trajectory sequences that do not contain human body images of non-target human bodies, human body trajectory sequences belonging to the same human body in the multiple human body trajectory sequences that do not contain human body images of non-target human bodies are merged.

[0116] The embodiments of the present application are based on the global features and local features of each human body image in the human body trajectory sequence. In the human body trajectory sequence, human body images of non-target human bodies are progressively removed, thereby removing human body images that are noise in the human body trajectory sequence, solving the noise problem. The human body trajectory sequences belonging to the same human body in multiple human body trajectory sequences after noise removal are merged, solving the splitting problem, ensuring the purity or quality of the final human body trajectory sequence, reducing errors, and reducing the degree of error accumulation. Furthermore, the final human body trajectory sequence can be used as training data to update the human body weight recognition model, which can improve the accuracy of the human body weight recognition model in complex and changing scenarios.

[0117] Furthermore, in order to remove the noise caused by merging multiple human trajectory sequences, the operations of steps S102-S103 are performed on each human trajectory sequence obtained after the merger in step S104 to remove the noise from each human trajectory sequence obtained after the merger, thereby obtaining a merged human trajectory sequence with the noise removed. This reduces the noise (non-target human body) in the merged human trajectory sequence, improves the quality or purity of the merged human trajectory sequence, and avoids error accumulation.

[0118] Furthermore, historical human trajectory sequences have been previously obtained through the methods of the embodiments of this application and have been combined into a training dataset. Thus, the noise-removed, merged human trajectory sequences can be added to the training dataset as training data to enrich the training dataset. Subsequently, when needed (or periodically), the training data in the training dataset (including the noise-removed, merged human trajectory sequences, i.e., incremental data) can be used to update the human weight recognition model (including the global human weight recognition model and multiple different types of local human weight recognition models), further reducing cumulative error and improving the accuracy of the global and local human weight recognition models.

[0119] Among them, after the training data set is enriched, the amount of data in the training data set and the covered data space distribution range are larger. Each automatic update of the human weight recognition model is to incrementally add new training data on the basis of the existing training data. In this way, it can adapt to complex and changing scenarios and improve the generalization ability or accuracy of the human weight recognition model (including the global human weight recognition model and multiple different types of local human weight recognition models).

[0120] In one embodiment, updating multiple different types of local human weight recognition models based on the enriched training data set includes: updating the local human weight recognition model corresponding to the upper body image based on the enriched training data set, updating the local human weight recognition model corresponding to the shoe image, updating the local human weight recognition model corresponding to the lower body image, updating the local human weight recognition model corresponding to the face image, etc.

[0121] The embodiments of the present application can realize the training, testing, iterative update and deployment of the human weight recognition model at the edge, ensuring the effect of human weight recognition at the edge. The automatic labeling based on human weight recognition (labeling the entire human trajectory sequence once, rather than labeling each human image in a human trajectory sequence separately) reduces the labeling cost, shortens the iteration cycle of the human weight recognition model, makes full use of idle machine resources at the edge, adapts to complex and changing scenarios, provides an adaptive human weight recognition model for complex and diverse scenarios, and realizes thousands of mirrors and thousands of models.

[0122] In another embodiment of the present application, see Figure 2 The process of "removing human body images of non-target human bodies in the human body trajectory sequence based on the global feature and the local feature" in step S102 includes:

[0123] In step S201 , human body images of non-target human bodies are removed from the human body trajectory sequence according to global features of the human body images in the human body trajectory sequence.

[0124] In this embodiment of the present application, a clustering model corresponding to the global features (or a first clustering model) can be used to identify non-target human images in the human trajectory sequence based on the global features of each human image in the human trajectory sequence. The first clustering model corresponding to the global features can be used to cluster the human images in the human trajectory sequence according to the global features to identify non-target human images.

[0125] Then, in the human body trajectory sequence, the human body images of the identified non-target human bodies are removed.

[0126] For example, the global features of each human body image in the human body trajectory sequence are input into the clustering model corresponding to the global features, so that the first clustering model corresponding to the global features performs clustering processing on the global features of each human body image in the human body trajectory sequence, obtains the human body image of the non-target human body in the human body trajectory sequence, and outputs the human body image of the non-target human body in the human body trajectory sequence. The electronic device can obtain the human body image of the non-target human body in the human body trajectory sequence output by the clustering model corresponding to the global features.

[0127] In an embodiment of the present application, the clustering model includes a radius and the number of core points, the radius is used to define the range of the neighborhood, and the number of core points includes: defining the minimum number of data points required to form a core point in the neighborhood.

[0128] Among them, the radius of the range used to define the neighborhood in the first clustering model corresponding to the global feature is the first radius, and the points within this radius are regarded as neighbors. The minimum number of data points required to define the core point in the neighborhood in the first clustering model corresponding to the global feature is the first number.

[0129] In one embodiment of the present application, the clustering model corresponding to the global feature may include DBSCAN, etc.

[0130] The first clustering model corresponding to the global feature is based on the global feature of each human body image in the human body trajectory sequence, and can find core points, outliers and noise points in each human body image in the human body trajectory sequence.

[0131] The first radius may be set to be loose, for example, it may include 1.1, 1.05 or 1.0, etc.

[0132] The first number may be set to be smaller, for example, the first number may be 2.

[0133] According to the definition of the DBSCAN clustering model, eps (the radius used to define the range of the neighborhood) corresponds to the distance between data. In the human weight recognition problem, the feature vectors are usually L2 normalized. All normalized feature vectors can be regarded as distributed on a high-dimensional sphere with a radius of 1. Therefore, the distance between any two feature vectors ranges from 0 to 2. The value of eps is determined by the distribution of samples in the dataset, that is, the feature vectors. If the distance between two samples is less than eps, they will be regarded as adjacent points by the DBSCAN model and may be divided into the same cluster.

[0134] minPts (which defines the minimum number of data points required to form a core point in the neighborhood) is a lower bound on the number of neighbors a core point must have to form a cluster. Its value is also determined by the distribution of the samples in the dataset, i.e., the feature vectors. Alternatively, minPts refers to the minimum number of neighbors to be considered a core point, i.e., a core point must have at least minPts points within the eps radius.

[0135] The distribution of feature vectors in the dataset is related to the acquisition method of the dataset itself. For example, if a large number of images are saved at a small time interval in each human trajectory sequence in the video, then these images are distributed very closely as sample points, otherwise they are distributed relatively sparsely. On the other hand, it is also related to the representation ability of the human body recognition model to represent the characteristics of the human body. For example, a model with strong representation ability can make the feature vectors distributed in a smaller spatial range when extracting features from human images, as long as these human images are from the same person, regardless of the lighting, distance, perspective, clarity and other conditions of the human images saved for this person. The distance from human images that are not of this person is larger. On the contrary, if the representation ability of the model is weak, the distribution of sample features is relatively sparse, and the distance to human images of other people may be closer.

[0136] DBSCAN is a classic density clustering model that is particularly suitable for discovering clusters of arbitrary shapes and can automatically identify noise points. The basic principle of DBSCAN is to identify clusters through the concept of density.

[0137] Key parameters in DBSCAN include eps (epsilon, which defines the radius of the neighborhood) and minPts (which defines the minimum number of data points required to form a core point in the neighborhood).

[0138] EPS defines the maximum distance between two points so that they can be considered as points in the neighborhood. That is, if the distance between two points is less than or equal to EPS, then the two points are considered to be directly reachable by density, which is the basis for forming clusters.

[0139] Choosing an appropriate eps value is crucial to the success of clustering: if eps is too small, many points may be regarded as noise; if eps is too large, it may lead to an overly simple model and multiple clusters may be merged into one.

[0140] MinPts is typically set to at least the dimension of the dataset plus one (for example, for two-dimensional data, the minimum value should be 3). MinPts affects cluster formation and noise detection. If set too low, many points may be mistaken for core points; if set too high, cluster formation may be difficult.

[0141] Core Point: A point is considered a core point if its eps neighborhood contains at least MinPts points. Core points are the starting point for cluster formation and growth. Core points are the foundation of clustering; they represent high-density areas, and clusters are discovered and expanded through these points.

[0142] Border Points: Border points are located within the EPS neighborhood of a core point, but their own EPS neighborhood does not have enough points to qualify as a core point. Border points are attached to the cluster but do not have the ability to extend the cluster. Therefore, they belong to the edge of the cluster determined by the neighboring core points.

[0143] Noise Points: Noise points are neither core points nor boundary points. They do not belong to any cluster. These points may represent abnormal data, irregular values, or isolated points that are not dense enough to form a cluster.

[0144] Introduction to the DBSCAN process:

[0145] Initialization: Select any unvisited point as the starting point.

[0146] Expand the cluster: Check the eps neighborhood of the point. If the number of points is not less than MinPts, then the point is a core point and a new cluster can be formed. Add all points in the neighborhood of the core point to the cluster and check each point in the neighborhood to determine whether they are new core points.

[0147] Cluster Growth: Repeat the neighborhood search and cluster expansion for each newly added core point. Through this series of checks and neighborhood connections, the cluster continues to grow. If a cluster boundary encounters border points, these border points are also marked as part of the cluster, but the cluster is not further expanded.

[0148] Noise Handling: If the number of points in a point's neighborhood is less than MinPts and does not belong to any other cluster, the point is labeled as noise. By identifying clusters in a dataset through density reachability (i.e., paths directly or indirectly connected through core points), DBSCAN can naturally handle irregularly shaped clusters without requiring a priori specification of the number or type of clusters. This makes DBSCAN highly effective in analyzing real-world datasets with complex structures.

[0149] In the embodiment of the present application, for any human body trajectory sequence, only the largest cluster set is retained, and the rest are regarded as outliers.

[0150] If the human body images in the human body trajectory sequence are all human body images of the same human body, then after DBSCAN clustering processing, all human body images in the human body trajectory sequence will generally be clustered into one cluster.

[0151] Alternatively, among the human body images in the human body trajectory sequence, if a small number of human body images do not belong to the same human body as the majority of other human body images, then two situations will occur with these small number of human body images (outliers): one is that they will be clustered into a smaller cluster (when their number is large and the clustering parameter conditions are met), and the other is that they cannot be clustered into any cluster and become noise points.

[0152] In the embodiment of the present application, the human body images and noise point images in the smaller clusters can be regarded as outliers, that is, human body images of non-target humans, which actually do not belong to the human body trajectory sequence.

[0153] In step S202 , human body images of non-target human bodies are removed from the remaining human body images in the human body trajectory sequence according to the local features of the remaining human body images in the human body trajectory sequence.

[0154] In an embodiment of the present application, a clustering model corresponding to local features (or a second clustering model) can be used to identify non-target human images from the remaining human images in the human trajectory sequence based on the local features of each remaining human image in the human trajectory sequence. The second clustering model corresponding to local features is used to cluster the remaining human images in the human trajectory sequence according to the local features to identify non-target human images; and the identified non-target human images are removed from the remaining human images in the human trajectory sequence. The second radius of the second clustering model is used to define the range of the neighborhood, and the second number of the second clustering model is used to define the minimum number of data points required to form a core point within the neighborhood. The second radius is smaller than the first radius, and / or the second number is greater than the first number.

[0155] Then, the human body images of the identified non-target human bodies are removed from the remaining human body images in the human body trajectory sequence.

[0156] For example, the local features of each remaining human body image in the human body trajectory sequence are input into the second clustering model corresponding to the local features, so that the second clustering model corresponding to the local features performs clustering processing on the local features of each remaining human body image in the human body trajectory sequence, obtains the human body image of the non-target human body in the human body trajectory sequence, and outputs the human body image of the non-target human body in the human body trajectory sequence. The electronic device can obtain the human body image of the non-target human body in the human body trajectory sequence output by the clustering model corresponding to the local features.

[0157] The radius of the range used to define the neighborhood in the second clustering model corresponding to the local feature is a second radius, and the minimum number of data points required to form a core point in the neighborhood in the clustering model corresponding to the local feature is a second number. The second preset threshold is less than the first preset threshold and / or the second number is greater than the first number.

[0158] In one embodiment of the present application, the second clustering model corresponding to the local features may include DBSCAN, etc.

[0159] The second clustering model corresponding to the local features can find core points, outliers and noise points in each human body image in the human body trajectory sequence based on the local features of the human body images in the human body trajectory sequence.

[0160] The second radius may be set more strictly, for example, it may include 0.8, 0.85 or 0.9, etc.

[0161] The first number can be set to be larger, for example, the first number is 3.

[0162] In one embodiment of the present application, the local features of the human body image are multiple different types of local features of the human body image. The different types of local features include: shoe features of a shoe image of a human body, upper body features of an upper body image of a human body, lower body features of a lower body image of a human body, or facial features of a facial image of a human body.

[0163] In this way, when removing human body images of non-target human bodies from the remaining human body images in the human body trajectory sequence based on the local features of each remaining human body image in the human body trajectory sequence, the remaining human body trajectory sequence in the human body trajectory sequence can be screened according to the local features according to the preset order between different types of local features to remove human body images of non-target human bodies.

[0164] For example, the local features of a human body image are P different types of local features of the human body image, where P is a positive integer.

[0165] First, according to the local features of the first rank of each human body image in the human body trajectory sequence, the human body images of non-target human bodies are removed from the remaining human body trajectory sequences in the human body trajectory sequence obtained in step S201 to obtain human body trajectory sequence 1. Then, according to the local features of the second rank of each human body image in the human body trajectory sequence, the human body images of non-target human bodies are removed from the human body trajectory sequence 1 to obtain human body trajectory sequence 2. ... Then, according to the local features of the P rank of each human body image in the human body trajectory sequence, the human body images of non-target human bodies are removed from the human body trajectory sequence 2. In the human body image, non-target human bodies are removed to obtain the final human body trajectory sequence.

[0166] The preset order of priority between the different types of local features is obtained by sorting the different types of local features in descending order according to their respective abilities to represent the characteristics of the human body.

[0167] For example, the shoe features of a shoe image of a person wearing shoes can represent the characteristics of the person. The ability of the upper body features of the upper body image to represent the characteristics of the human body The upper body features of the lower body image of the human body can represent the characteristics of the human body.

[0168] The second radius of the neighborhood range defined in the clustering model corresponding to the local feature with the highest ranking is larger than the second radius of the neighborhood range defined in the clustering model corresponding to the local feature with the lowest ranking. That is, the second radius of the clustering model corresponding to the local feature with the highest ranking is larger than the second radius of the clustering model corresponding to the local feature with the lowest ranking.

[0169] And / or, the second number of data points required to define a core point within a neighborhood in a clustering model corresponding to a local feature of a higher-ranked type is less than or equal to the second number of data points required to define a core point within a neighborhood in a clustering model corresponding to a local feature of a lower-ranked type. That is, the second number of clustering models for the local feature of a higher-ranked type is less than or equal to the second number of clustering models for the local feature of a lower-ranked type.

[0170] The radius of the range used to define the neighborhood in the first clustering model corresponding to the global feature The radius of the range used to define the neighborhood in the second clustering model corresponding to the top-ranked local features The radius used to define the neighborhood in the second clustering model for local features ranked lower in the ranking is gradually tightened. This allows for smoother outlier removal. This means that by first removing "significant" outliers, and then further clustering with a stricter radius, these already-removed outliers won't interfere with the clustering model. In other words, different radii are used because global features and different types of local features have different abilities to represent the human body.

[0171] The purpose of this embodiment is to remove outliers in the human body trajectory sequence, so it is necessary to pay attention to the distance between the features of different human body images.

[0172] In the first clustering model corresponding to all features and the second clustering model corresponding to local features, the global features used in the first clustering model corresponding to global features have a stronger ability to characterize human features. The distance between outliers (other human images) in a given human trajectory sequence and the human images of that human in that trajectory sequence is relatively far, and different outliers are also relatively scattered. Therefore, a looser threshold can be used to remove outliers in the human trajectory sequence. The local features used in the second clustering model corresponding to local features have a weaker ability to characterize human features. Generally speaking, even if local features belong to different people, the distance between them is relatively close. Therefore, if the distance between outliers (other human partial images) in a given human trajectory sequence and the partial images of that human in that trajectory sequence is relatively close, a strict radius is required to remove the outliers.

[0173] In another embodiment of the present application, see Figure 3 , step S104 includes:

[0174] In step S301 , a first human body trajectory sequence is selected from the screened human body trajectory sequences, and first global features of M human body images are obtained from the first human body trajectory sequence.

[0175] In step S302 , a second global feature of each human body image is obtained from a second human body trajectory sequence, where the second human body trajectory sequence is a human body trajectory sequence other than the first human body trajectory sequence in the screened human body trajectory sequence.

[0176] In step S303 , the feature similarity between the first global feature and the second global feature is determined.

[0177] For a first human trajectory sequence among multiple human trajectory sequences that do not contain images of non-target humans, feature similarities are obtained between the global features of the M human images in the first human trajectory sequence and the global features of each human image in a second human trajectory sequence. The first human trajectory sequence is any human trajectory sequence among the multiple human trajectory sequences that do not contain images of non-target humans, and the second human trajectory sequence is any human trajectory sequence among the multiple human trajectory sequences that do not contain images of non-target humans, excluding the first human trajectory sequence.

[0178] M is a positive integer.

[0179] For any human image in the M human images in the first human trajectory sequence, feature similarity between the global features of the human image and the global features of each human image in the second human trajectory sequence can be obtained. Feature similarity can include cosine similarity, etc. The above operation is performed similarly for each other human image in the M human images in the first human trajectory sequence.

[0180] In step S304 , based on the feature similarity, a human body trajectory sequence that belongs to the same human body as the target human body trajectory is selected from the second human body trajectory sequence as a human body trajectory sequence to be merged.

[0181] In one embodiment of the present application, among the feature similarities between the global features of the M human images in the first human trajectory sequence and the global features of each human image in the second human trajectory sequence, the largest multiple feature similarities (the first multiple feature similarities sorted in descending order) are selected.

[0182] Since a feature similarity represents the feature similarity between the global features of a human body image in the first human body trajectory sequence and the global features of a human body image in the second human body trajectory sequence, a feature similarity involves the global features of a human body image in the first human body trajectory sequence and the global features of a human body image in the second human body trajectory sequence.

[0183] In this way, the second human body trajectory sequence to which the human body image corresponding to the global feature to which each feature similarity among the selected maximum feature similarities belongs may be statistically calculated.

[0184] Then, based on the statistical second human body trajectory sequence, a human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence is obtained. For example, the statistical second human body trajectory sequence is determined as the human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence.

[0185] In addition, this step can also be referred to as Figure 4The embodiment shown will not be described in detail here.

[0186] In step S305 , the first human body trajectory sequence and the human body trajectory sequence to be merged are merged to determine a merged human body trajectory sequence.

[0187] For example, the human body images in the first human body trajectory sequence and the human body images in the human body trajectory sequence to be merged are merged into a new human body trajectory sequence.

[0188] In this embodiment of the present application, each human trajectory sequence in a plurality of human trajectory sequences that do not contain human images of non-target human bodies may be treated as a first human trajectory sequence to execute the step of "obtaining feature similarity between the global features of the M human images in the first human trajectory sequence and the global features of each human image in the second human trajectory sequence." If a human trajectory sequence has not yet been considered as the first human trajectory sequence to execute steps S301 to S305, but has already been merged with the second human trajectory sequence as a human trajectory sequence to be merged, then the human trajectory sequence is no longer considered as the first human trajectory sequence to execute steps S301 to S305.

[0189] In another embodiment of the present application, see Figure 4 , step S304 includes:

[0190] In step S401, the first N feature similarities are selected from the acquired feature similarities. In one example, , you can select the front feature similarity, K is a preset positive integer.

[0191] In an optional implementation, for any human image in the M human images in the first human trajectory sequence, top K feature similarities are selected from the feature similarities between the global features of the human image and the global features of each human image in the second human trajectory sequence. For any human image in the M human images in the first human trajectory sequence, the top K feature similarities are selected, that is, the above operation is performed on each of the other human images in the M human images in the first human trajectory sequence, thereby obtaining the top K feature similarities. feature similarity.

[0192] In step S402, based on the first N feature similarities, the second human body trajectory sequence to which the corresponding human body image belongs is screened to determine a human body trajectory sequence to be merged that belongs to the same human body as the target human body trajectory sequence.

[0193] Before statistics can be Each feature similarity in the feature similarities is respectively related to a second human body trajectory sequence to which the human body image corresponding to the global feature belongs.

[0194] Since a feature similarity represents the feature similarity between the global features of a human body image in the first human body trajectory sequence and the global features of a human body image in the second human body trajectory sequence, a feature similarity involves the global features of a human body image in the first human body trajectory sequence and the global features of a human body image in the second human body trajectory sequence.

[0195] In this way, we can count the Each feature similarity in the feature similarities is respectively related to a second human body trajectory sequence to which the human body image corresponding to the global feature belongs.

[0196] K may include 20, 25 or 30, etc., which is not limited in the embodiment of the present application.

[0197] Furthermore, based on the statistical second human body trajectory sequence, a human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence is obtained. In one embodiment of the present application, the statistical second human body trajectory sequence can be determined as a human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence.

[0198] Alternatively, in another embodiment of the present application, the second human body trajectory sequence to which the corresponding human body image belongs and the number of feature similarities corresponding to the second human body trajectory sequence are determined based on the first N feature similarities. Among the feature similarities, the number of feature similarities of the human body image corresponding to the global feature involved belonging to the second human body trajectory sequence is counted. For each other second human body trajectory sequence in the counted second human body trajectory sequence, the above operation is also performed to obtain the number of feature similarities of the human body image corresponding to the global feature involved belonging to each counted second human body trajectory sequence.

[0199] Then, based on the counted number, the human body trajectory sequences to be merged that belong to the same person as the first human body trajectory sequence are selected from the counted second human body trajectory sequences. That is, based on the number of feature similarities corresponding to the second human body trajectory sequences, the human body trajectory sequences to be merged that belong to the same person as the first human body trajectory sequence are selected.

[0200] Calculate the ratio of the number of feature similarities contained in the corresponding second human trajectory sequence to N; if the ratio is greater than a preset threshold, the corresponding second human trajectory sequence is determined as a human trajectory sequence to be merged that belongs to the same human body as the first human trajectory sequence. For any second human trajectory sequence in the statistical second human trajectory sequence, the number of feature similarities of the human image corresponding to the global feature involved and the number of human images belonging to the second human trajectory sequence are calculated. When the ratio between the two is greater than a preset threshold, the second human trajectory sequence is determined as a human trajectory sequence to be merged that belongs to the same human body as the first human trajectory sequence, or the number of feature similarities between the human image corresponding to the global feature involved and the number of feature similarities between the human image corresponding to the global feature involved and the second human trajectory sequence. When the ratio between the second and first human body trajectory sequences is less than or equal to a preset threshold, the second human body trajectory sequence is not determined as a human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence, but it can be determined that the second human body trajectory sequence and the first human body trajectory sequence belong to different human bodies.

[0201] The above operation is performed similarly for each other second human body trajectory sequence in the counted second human body trajectory sequence.

[0202] Among them, the specific value of the preset threshold can be set in advance by technicians according to actual conditions. For example, it can include 30%, 35% or 40%, etc., which can be determined according to actual conditions. The embodiments of this application do not limit this.

[0203] Based on the above embodiment, the embodiment of the present application can perform recognition processing on the merged human trajectory sequence to determine the trajectory information of the target user. For example, the merged human trajectory sequence can be recognized to determine the user information corresponding to the human trajectory sequence, and the trajectory information corresponding to the human trajectory sequence can be determined to determine the trajectory information of the target user.

[0204] Based on the above user information and trajectory information, corresponding processing can be performed, such as counting the flow of people in the shopping mall, counting the flow of people corresponding to the goods in the store, etc.

[0205] In some other embodiments of the present application, the combined human body trajectory sequence can also be used as training data to train a recognition model, such as a human weight recognition model.

[0206] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the order of the actions described, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by the embodiments of the present application.

[0207] Reference Figure 5 , shows a structural block diagram of a device for processing a human body trajectory sequence according to an embodiment of the present application, which is applied to an electronic device, and the device includes:

[0208] The acquisition module 11 is configured to acquire a plurality of human body trajectory sequences, wherein a human body trajectory sequence at least includes: human body images of at least one human body acquired at a plurality of different moments.

[0209] The first extraction module 12 is configured to extract global features of each human body image in any human body trajectory sequence among the plurality of human body trajectory sequences.

[0210] The second extraction module 13 is configured to extract local features of each human body image in any human body trajectory sequence among the plurality of human body trajectory sequences.

[0211] The removal module 14 is configured to remove human body images of non-target humans from the human body trajectory sequence based on the global features and the local features, and determine a plurality of screened human body trajectory sequences; wherein, in the human body trajectory sequence, the number of human body images of non-target humans is less than the number of human body images of target humans.

[0212] The merging module 15 is configured to determine the human body trajectory sequences belonging to the same human body according to the global features of the human body images in the plurality of screened human body trajectory sequences, merge the human body trajectory sequences, and determine a merged human body trajectory sequence.

[0213] In an optional implementation, the removal module includes:

[0214] A first removal unit is configured to remove human body images of non-target human bodies in the human body trajectory sequence according to global features of each human body image in the human body trajectory sequence;

[0215] The second removing unit is configured to remove human body images of non-target human bodies from the remaining human body images in the human body trajectory sequence according to local features of the remaining human body images in the human body trajectory sequence.

[0216] In an optional implementation, the local features of the human body image are a plurality of different types of local features of the human body image;

[0217] The second removal unit includes:

[0218] a second removal subunit, configured to screen the remaining human body trajectory sequences in the human body trajectory sequence according to the local features in accordance with a preset order between different types of local features, and remove human body images of non-target human bodies;

[0219] The preset order of priority between the different types of local features is obtained by sorting the different types of local features in descending order according to their respective abilities to represent the characteristics of the human body.

[0220] In an optional implementation, the first removing unit includes:

[0221] a first recognition subunit, configured to cluster the human body images in the human body trajectory sequence according to the global features using a first clustering model corresponding to the global features, and identify human body images of non-target humans;

[0222] A second removal subunit is configured to remove human body images of identified non-target human bodies from the human body trajectory sequence;

[0223] The first radius of the first clustering model is used to define the range of the neighborhood, and the first number of the first clustering model is used to define the minimum number of data points required to form a core point in the neighborhood.

[0224] In an optional implementation, the second removing unit includes:

[0225] a second recognition subunit, configured to cluster the remaining human body images in the human body trajectory sequence according to the local features using a second clustering model corresponding to the local features, and identify human body images of non-target humans;

[0226] a third removing subunit, configured to remove human body images of identified non-target human bodies from the remaining human body images in the human body trajectory sequence;

[0227] The second radius of the second clustering model is used to define the range of the neighborhood, and the second number of the second clustering model is used to define the minimum number of data points required to form a core point in the neighborhood; the second radius is smaller than the first radius, and / or the second number is greater than the first number.

[0228] In an optional implementation, the second radius of the clustering model whose local features are ranked higher in type is greater than the second radius of the clustering model whose local features are ranked lower in type;

[0229] And / or, the second number of clustering models whose types of local features are ranked higher is less than or equal to the second number of clustering models whose types of local features are ranked lower.

[0230] In an optional implementation, the merging module includes:

[0231] A first acquisition unit is configured to acquire a second global feature of each human body image from a second human body trajectory sequence, where the second human body trajectory sequence is a human body trajectory sequence other than the first human body trajectory sequence in the screened human body trajectory sequence; and determine a feature similarity between the first global feature and the second global feature; the first human body trajectory sequence is any human body trajectory sequence from a plurality of human body trajectory sequences in which no non-target human body is present, and the second human body trajectory sequence is a human body trajectory sequence other than the first human body trajectory sequence from a plurality of human body trajectory sequences in which no non-target human body is present;

[0232] a second acquiring unit, configured to select, from the second human body trajectory sequence, a human body trajectory sequence that belongs to the same human body as the target human body trajectory, as a human body trajectory sequence to be merged;

[0233] The merging unit is configured to merge the first human body trajectory sequence with the human body trajectory sequence to be merged, and determine a merged human body trajectory sequence.

[0234] In an optional implementation, the second acquiring unit includes:

[0235] A selection subunit is used to select the top N feature similarities from the acquired feature similarities;

[0236] The statistical subunit is used to screen the second human body trajectory sequence to which the corresponding human body image belongs based on the first N feature similarities, and determine the human body trajectory sequence to be merged that belongs to the same human body as the target human body trajectory sequence.

[0237] In an optional implementation, the selection subunit is specifically configured to: for any human image in the M human images in the first human trajectory sequence, screen the top K feature similarities. For any human image in the M human images in the first human trajectory sequence, select TOPK feature similarities from the feature similarities between the global features of the human image and the global features of each human image in the second human trajectory sequence.

[0238] In an optional implementation, the statistical subunit is specifically configured to determine, based on the first N feature similarities, a second human body trajectory sequence to which the corresponding human body image belongs, and the number of feature similarities corresponding to the second human body trajectory sequence; and, based on the number of feature similarities corresponding to the second human body trajectory sequence, screen a human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence.

[0239] In an optional implementation, the statistical subunit is specifically configured to calculate a ratio between the number of feature similarities corresponding to the second human trajectory sequence and N; and when the ratio is greater than a preset threshold, determine the corresponding second human trajectory sequence as a human trajectory sequence to be merged that belongs to the same human body as the first human trajectory sequence.

[0240] In an optional implementation, the system further includes: an identification module configured to perform identification processing on the combined human trajectory sequence to determine trajectory information of the target user.

[0241] In an optional implementation, the system further includes a training module for using the combined human trajectory sequence as training data to train a recognition model.

[0242] In an embodiment of the present application, multiple human body trajectory sequences are acquired, and one human body trajectory sequence includes at least: human body images of at least one human body acquired at multiple different times. For any human body trajectory sequence, global features of each human body image in the human body trajectory sequence are extracted, and local features of each human body image in the human body trajectory sequence are extracted. Based on the global features and the local features, human body images of non-target human bodies are removed from the human body trajectory sequence, wherein the number of human body images of non-target human bodies in the human body trajectory sequence is less than the number of human body images of target human bodies. Based on the global features of each human body image in the multiple human body trajectory sequences that do not contain human body images of non-target human bodies, human body trajectory sequences belonging to the same human body in the multiple human body trajectory sequences that do not contain human body images of non-target human bodies are merged.

[0243] The embodiments of the present application are based on the global features and local features of each human body image in the human body trajectory sequence. In the human body trajectory sequence, human body images of non-target human bodies are progressively removed, thereby removing human body images that are noise in the human body trajectory sequence, solving the noise problem. The human body trajectory sequences belonging to the same human body in multiple human body trajectory sequences after noise removal are merged, solving the splitting problem, ensuring the purity or quality of the final human body trajectory sequence, reducing errors, and reducing the degree of error accumulation. Furthermore, the final human body trajectory sequence can be used as training data to update the human body weight recognition model, which can improve the accuracy of the human body weight recognition model in complex and changing scenarios.

[0244] An embodiment of the present application further provides a non-volatile readable storage medium, which stores one or more modules (programs). When the one or more modules are applied to a device, the device can execute instructions (instructions) of each method step in the embodiment of the present application.

[0245] The present application provides one or more machine-readable media having instructions stored thereon, which, when executed by one or more processors, cause an electronic device to perform one or more of the methods described in the above embodiments. In the present application, the electronic device includes a server, a gateway, an electronic device, etc., and the sub-device is an Internet of Things device or other device.

[0246] The embodiments of the present disclosure may be implemented as an apparatus configured as desired using any appropriate hardware, firmware, software, or any combination thereof, which may include a server (cluster), a terminal device such as an IoT device, and other electronic devices.

[0247] Figure 6 An exemplary apparatus 600 that can be used to implement various embodiments of the present application is schematically shown.

[0248] For one embodiment, Figure 6 An exemplary apparatus 600 is shown having one or more processors 602, a control module (chip set) 604 coupled to at least one of the processor(s) 602, a memory 606 coupled to the control module 604, a non-volatile memory (NVM) / storage device 608 coupled to the control module 604, one or more input / output devices 610 coupled to the control module 604, and a network interface 612 coupled to the control module 604.

[0249] The processor 602 may include one or more single-core or multi-core processors, and the processor 602 may include any combination of general-purpose processors or dedicated processors (such as graphics processors, application processors, baseband processors, etc.). In some embodiments, the apparatus 600 can serve as a server device such as a gateway in the embodiments of the present application.

[0250] In some embodiments, the apparatus 600 may include one or more computer-readable media (e.g., memory 606 or NVM / storage 608) having instructions 614 and one or more processors 602 configured in combination with the one or more computer-readable media to execute the instructions 614 to implement a module to perform actions in the present disclosure.

[0251] For one embodiment, the control module 604 may include any suitable interface controller to provide any suitable interface to at least one of the processor(s) 602 and / or any suitable device or component in communication with the control module 604 .

[0252] The control module 604 may include a memory controller module to provide an interface to the memory 606. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0253] The memory 606 can be used, for example, to load and store data and / or instructions 614 for the device 600. For one embodiment, the memory 606 can include any suitable volatile memory, such as a suitable DRAM. In some embodiments, the memory 606 can include double data rate quad synchronous dynamic random access memory (DDR4 SDRAM).

[0254] For one embodiment, the control module 604 may include one or more input / output controllers to provide interfaces to the NVM / storage device 608 and the input / output device(s) 610 .

[0255] For example, NVM / storage 608 may be used to store data and / or instructions 614. NVM / storage 608 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).

[0256] NVM / storage device 608 may include storage resources that are physically part of the device on which apparatus 600 is installed, or it may be accessible to the device without being part of the device. For example, NVM / storage device 608 may be accessible over a network via input / output device(s) 610.

[0257] (One or more) input / output devices 610 may provide an interface for apparatus 600 to communicate with any other appropriate device. Input / output devices 610 may include a communication component, a phonetic component, a sensor component, etc. Network interface 612 may provide an interface for apparatus 600 to communicate via one or more networks. Apparatus 600 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, for example, accessing a wireless network based on a communication standard, such as Wi-Fi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.

[0258] For one embodiment, at least one of the processor(s) 602 may be packaged together with the logic of one or more controllers (e.g., a memory controller module) of the control module 604. For one embodiment, at least one of the processor(s) 602 may be packaged together with the logic of one or more controllers of the control module 604 to form a system-in-package (SiP). For one embodiment, at least one of the processor(s) 602 may be integrated on the same die with the logic of one or more controllers of the control module 604. For one embodiment, at least one of the processor(s) 602 may be integrated on the same die with the logic of one or more controllers of the control module 604 to form a system-on-chip (SoC).

[0259] In various embodiments, apparatus 600 may be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, apparatus 600 may have more or fewer components and / or a different architecture. For example, in some embodiments, apparatus 600 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0260] An embodiment of the present application provides an electronic device, comprising: one or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, enable the electronic device to execute one or more methods in the embodiments of the present application.

[0261] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0262] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0263] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, and the combination of the processes and / or boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable information processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable information processing terminal device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0264] These computer program instructions may also be stored in a computer readable memory that can guide a computer or other programmable information processing terminal device to work in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0265] These computer program instructions can also be loaded onto a computer or other programmable information processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable terminal device. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0266] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0267] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0268] The above is a detailed introduction to the method, device, medium and product for processing human body trajectory sequences provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the embodiments of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the embodiments of the present application. At the same time, for those skilled in the art, according to the ideas of the embodiments of the present application, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the embodiments of the present application.

Claims

1. A method for processing a human body trajectory sequence, characterized in that: The method comprises: Acquire a plurality of human body trajectory sequences, wherein the human body trajectory sequences include: human body images of at least one human body acquired at a plurality of different moments; For any human body trajectory sequence among the multiple human body trajectory sequences, respectively extracting global features of each human body image in the human body trajectory sequence, and respectively extracting local features of each human body image in the human body trajectory sequence; Using a first clustering model corresponding to the global feature, clustering the human body images in the human body trajectory sequence according to the global feature, and identifying human body images of non-target humans; In the human body trajectory sequence, human body images of identified non-target human bodies are removed; wherein the first radius of the first clustering model is used to define the range of the neighborhood, and the first number of the first clustering model is used to define the minimum number of data points required to form a core point in the neighborhood; Using a second clustering model corresponding to the local features, clustering the remaining human body images in the human body trajectory sequence according to the local features, and identifying human body images of non-target humans; Removing the identified non-target human images from the remaining human images in the human trajectory sequence; wherein the second radius of the second clustering model is used to define the range of the neighborhood, and the second number of the second clustering model is used to define the minimum number of data points required to form a core point in the neighborhood; the second radius is smaller than the first radius, and / or the second number is larger than the first number, and in the human trajectory sequence, the number of non-target human images is smaller than the number of target human images; According to the global features of the human body images in the plurality of screened human body trajectory sequences, human body trajectory sequences belonging to the same human body are determined and merged to determine a merged human body trajectory sequence.

2. The method according to claim 1, characterized in that The local features of the human body image are a plurality of different types of local features of the human body image; the method further includes: According to a preset order between the different types of local features, performing the step of removing the human body image of the non-target human body identified by using the second clustering model corresponding to the local features; The preset order of priority between the different types of local features is obtained by sorting the different types of local features in descending order according to their respective abilities to represent the characteristics of the human body.

3. The method according to claim 2, characterized in that The second radius of the clustering model whose local feature types are ranked higher is greater than the second radius of the clustering model whose local feature types are ranked lower; And / or, the second number of clustering models whose types of local features are ranked higher is less than or equal to the second number of clustering models whose types of local features are ranked lower.

4. The method according to claim 1, wherein The method of determining human body trajectory sequences belonging to the same human body and merging the human body trajectory sequences based on the global features of the human body images in the plurality of screened human body trajectory sequences to determine the merged human body trajectory sequence comprises: Selecting a first human body trajectory sequence from the screened human body trajectory sequences, and obtaining first global features of M human body images from the first human body trajectory sequence; Acquire a second global feature of each human body image from a second human body trajectory sequence, where the second human body trajectory sequence is a human body trajectory sequence other than the first human body trajectory sequence in the screened human body trajectory sequence; determining a feature similarity between the first global feature and the second global feature; According to the feature similarity, a human body trajectory sequence that belongs to the same human body as the target human body trajectory is selected from the second human body trajectory sequence as the human body trajectory sequence to be merged; The first human body trajectory sequence is merged with the human body trajectory sequence to be merged to determine a merged human body trajectory sequence.

5. The method according to claim 4, characterized in that The step of selecting a human body trajectory sequence belonging to the same human body as the target human body trajectory in the second human body trajectory sequence according to the feature similarity as the human body trajectory sequence to be merged includes: Select the top N feature similarities among the acquired feature similarities; According to the first N feature similarities, the second human body trajectory sequence to which the corresponding human body image belongs is screened, and a human body trajectory sequence to be merged that belongs to the same human body as the target human body trajectory sequence is determined.

6. The method according to claim 5, characterized in that The selecting the top N feature similarities from the acquired feature similarities includes: For any human body image in the M human body images in the first human body trajectory sequence, the top K feature similarities are respectively screened.

7. The method according to claim 5, characterized in that The step of screening the second human body trajectory sequence to which the corresponding human body image belongs based on the first N feature similarities, and determining the human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence, includes: Determine, based on the first N feature similarities, a second human body trajectory sequence to which the corresponding human body image belongs, and the number of feature similarities corresponding to the second human body trajectory sequence; According to the number of feature similarities corresponding to the second human body trajectory sequence, a human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence is selected.

8. The method according to claim 7, characterized in that The step of screening the human body trajectory sequence to be merged, which belongs to the same human body as the first human body trajectory sequence, according to the number of feature similarities corresponding to the second human body trajectory sequence, includes: Calculate the ratio of the number of feature similarities contained in the second human trajectory sequence to N; When the ratio is greater than a preset threshold, the corresponding second human body trajectory sequence is determined as a human body trajectory sequence to be merged that belongs to the same human body as the first human body trajectory sequence.

9. The method according to claim 1, characterized in that Also includes: The combined human body trajectory sequence is identified and processed to determine the trajectory information of the target user.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When a processor executes the program, the method according to any one of claims 1 to 9 is implemented.

11. A computer-readable storage medium, characterized in that A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

12. A computer program product comprising a computer program / computer executable instructions, wherein: When the computer program / computer executable instructions are executed by a processor in an electronic device, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Online multi-target tracking method and system based on pedestrian attribute guidance

    CN116977365A

  • Pedestrian counting and detection at a traffic intersection based on object movement within a field of view

    US9460613B1