Recognition, people flow statistics, tracking, detection and alarm methods, apparatuses and devices
By constructing a feature extractor with a tree-like branch structure, the problem of decreased feature extraction speed in multi-granularity networks is solved, achieving more efficient feature extraction and recognition.
Patent Information
- Application Number
- CN202010746548.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-29
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2040-07-29
AI Technical Summary
In existing technologies, adding branches to multi-granularity networks during human feature extraction leads to an increase in the number of network parameters, a decrease in feature extraction speed, and a reduction in pedestrian recognition efficiency.
A feature extractor with a tree-like branching structure is adopted. By constructing convolutional layers and pooling computation modules with tree-like branches, global features, horizontal local features, and vertical local features of the image are extracted, reducing the number of nodes to improve the speed and accuracy of feature extraction.
While improving recognition accuracy, the same number of branches are achieved with fewer nodes, thus improving feature extraction speed and recognition efficiency.
Smart Images

Figure CN114092957B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an identification method and apparatus, a people flow statistics method and apparatus, a trajectory tracking method and apparatus, a detection method and apparatus, an alarm method and apparatus, an electronic device, and a storage medium. Background Technology
[0002] Pedestrian re-identification, also known as pedestrian re-identification, is a technique that uses computer vision to determine whether a specific pedestrian exists in an image or video sequence. It aims to overcome the visual limitations of fixed cameras and can be combined with pedestrian detection / tracking technologies, finding applications in fields such as intelligent video surveillance and intelligent security.
[0003] Among these, the extraction of human features is one of the key steps in pedestrian re-identification. In existing technologies, multi-granularity networks are commonly used for human feature extraction. A multi-granularity network is a multi-branch deep network that obtains multi-granularity local feature representations by dividing an image into multiple local stripes and changing the number of stripes in different local branches.
[0004] To improve the accuracy of extracted human features, specific local feature branches can be added; however, adding branches will increase the number of network parameters, which will lead to a decrease in feature extraction speed and an increase in time consumption, thereby reducing the efficiency of pedestrian recognition. Summary of the Invention
[0005] This application provides an identification method that improves identification efficiency while ensuring improved identification accuracy.
[0006] Accordingly, embodiments of this application also provide an identification device, an electronic device, and a storage medium to ensure the implementation and application of the above methods.
[0007] This application provides a method for counting people, which improves both the accuracy and efficiency of counting people.
[0008] Accordingly, embodiments of this application also provide a people flow counting device, an electronic device, and a storage medium to ensure the implementation and application of the above method.
[0009] This application provides a trajectory tracking method to improve the efficiency of trajectory tracking and recognition while ensuring improved trajectory tracking accuracy.
[0010] Accordingly, embodiments of this application also provide a trajectory tracking device, an electronic device, and a storage medium to ensure the implementation and application of the above method.
[0011] This application provides a detection method that improves detection efficiency while ensuring improved detection accuracy.
[0012] Accordingly, embodiments of this application also provide a detection device, an electronic device, and a storage medium to ensure the implementation and application of the above method.
[0013] This application provides an alarm method to reduce the false alarm rate.
[0014] Accordingly, embodiments of this application also provide an alarm device, an electronic device, and a storage medium to ensure the implementation and application of the above methods.
[0015] To address the aforementioned issues, this application discloses an identification method, comprising: acquiring a target image; extracting feature information corresponding to an object in the target image using a feature extractor, wherein the feature extractor has a tree-like branching structure; and identifying whether the target image contains a target object based on the feature information corresponding to the object in the target image.
[0016] Optionally, the step of identifying whether the target image has a target object based on the feature information corresponding to the object in the target image includes: determining the distance between the feature information corresponding to the object in the target image and the feature information corresponding to the target object, and determining whether the target image has a target object based on the distance.
[0017] Optionally, the step of extracting feature information corresponding to the object in the target image using a feature extractor includes: inputting the target image into the feature extractor to obtain global feature information, horizontal local feature information, and vertical local feature information corresponding to the target image; and using the global feature information, horizontal local feature information, and vertical local feature information corresponding to the target image to generate feature information corresponding to the object in the target image.
[0018] Optionally, the method further includes the step of constructing the feature extractor: constructing a convolutional layer with a tree-like branch structure; and constructing a convolutional computation module and a pooling computation module sequentially after the convolutional layer to obtain the feature extractor.
[0019] Optionally, the convolutional layer that constructs the tree-like branch structure includes: constructing a convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner.
[0020] Optionally, the step of constructing a convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner includes: constructing a convolutional layer shared by M layers; and using the Mth convolutional layer shared by M layers as the root node, constructing an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner.
[0021] Optionally, the step of constructing an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, using the shared convolutional layer of the Mth layer as the root node, includes: determining the first number of global feature branches, the second number of horizontal local feature branches, and the third number of vertical local feature branches; determining the horizontal local information corresponding to the second number of horizontal local feature branches, and determining the vertical local information corresponding to the third number of vertical local feature branches; and constructing an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, based on the first number, the second number, the third number, the horizontal local information, and the vertical local information.
[0022] Optionally, in the last N convolutional layers, the number of convolutional layers in the next layer of two adjacent layers is greater than or equal to the number of convolutional layers in the previous layer.
[0023] This application also discloses a method for pedestrian flow statistics. The method includes: acquiring target video data within a set time period, the target video data including multiple frames of images; using a feature extractor to extract human body feature information corresponding to pedestrians in the multiple frames of images, wherein the feature extractor has a tree-like branching structure; identifying pedestrians based on the human body feature information corresponding to pedestrians in the multiple frames of images, determining images with the same pedestrians and obtaining multiple image groups, each image group including multiple images, with different pedestrians corresponding to different image groups; and counting pedestrian flow within the set time period based on the number of image groups and the number of images outside the image groups.
[0024] This application also discloses a trajectory tracking method, the method comprising: acquiring target video data and target pedestrian images, the target video data including multiple frames of images to be detected; using a feature extractor to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected, and extracting human body feature information of the target pedestrian in the target pedestrian image, wherein the feature extractor has a tree-like branching structure; comparing the human body feature information corresponding to pedestrians in the multiple frames of images to be detected with the human body feature information of the target pedestrian in the target pedestrian image, and selecting a target image containing the target pedestrian from the multiple frames of images to be detected; and generating the motion trajectory of the target pedestrian based on the target image containing the target pedestrian.
[0025] This application also discloses a detection method, the method comprising: determining the target shelf where the unpaid goods are located when it is determined that an unpaid item is lost; acquiring target video data, the target video data including multiple frames of images to be detected; extracting human body feature information corresponding to pedestrians in the multiple frames of images to be detected using a feature extractor, wherein the feature extractor has a tree-like branching structure; determining images with the same pedestrians based on the human body feature information corresponding to pedestrians in the multiple frames of images to be detected and obtaining multiple image groups, the image groups including multiple images, and different pedestrians corresponding to different image groups; generating motion trajectories of multiple pedestrians based on the multiple image groups respectively, and determining the target pedestrian passing through the target shelf based on the motion trajectories of the multiple pedestrians; and detecting whether the target pedestrian has taken the unpaid goods.
[0026] This application also discloses an alarm method, the method comprising: acquiring a target image; extracting human body feature information corresponding to pedestrians in the target image using a feature extractor, wherein the feature extractor has a tree-like branch structure; comparing the human body feature information corresponding to pedestrians in the target image with the human body feature information corresponding to target pedestrians in a preset blacklist; and performing an alarm processing when it is determined that a target pedestrian exists in the target image.
[0027] This application also discloses a detection method, the method comprising: determining when a preset product is lost; acquiring target video data, the target video data including multiple frames of images to be detected; using a feature extractor to extract feature information corresponding to the product in the multiple frames of images to be detected, wherein the feature extractor has a tree-like branching structure; determining a target image containing the preset product in the multiple frames of images to be detected based on the feature information corresponding to the product in the multiple frames of images to be detected; and detecting a target pedestrian taking the preset product from the target image.
[0028] This application also discloses an identification device, the device comprising: a first acquisition module for acquiring a target image; a first feature extraction module for extracting feature information corresponding to an object in the target image using a feature extractor, wherein the feature extractor has a tree-like branch structure; and a first identification module for identifying whether the target image has a target object based on the feature information corresponding to the object in the target image.
[0029] Optionally, the first recognition module is used to determine the distance between the feature information corresponding to the object in the target image and the feature information corresponding to the target object, and to determine whether the target image has a target object based on the distance.
[0030] Optionally, the first feature extraction module is used to input the target image into the feature extractor to obtain global feature information, horizontal local feature information, and vertical local feature information corresponding to the target image; and to use the global feature information, horizontal local feature information, and vertical local feature information corresponding to the target image to generate feature information corresponding to the object in the target image.
[0031] Optionally, the apparatus further includes: a construction module for constructing the feature extractor; the construction module includes: a convolutional layer construction submodule for constructing a tree-branched convolutional layer; and a computation module construction submodule for sequentially constructing a convolution computation module and a pooling computation module after the convolutional layer to obtain the feature extractor.
[0032] Optionally, the convolutional layer construction submodule is used to construct a convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner.
[0033] Optionally, the convolutional layer construction submodule is used to construct a shared convolutional layer of M layers; taking the shared convolutional layer of the Mth layer as the root node, and constructing an N-layer convolutional layer containing global feature branches, horizontal local feature branches and vertical local feature branches in a tree-like branching manner.
[0034] Optionally, the convolutional layer construction submodule is used to determine the first number of global feature branches, the second number of horizontal local feature branches, and the third number of vertical local feature branches; determine the horizontal local information corresponding to the second number of horizontal local feature branches, and determine the vertical local information corresponding to the third number of vertical local feature branches; and construct an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, based on the first number, the second number, the third number, the horizontal local information, and the vertical local information.
[0035] Optionally, in the last N convolutional layers, the number of convolutional layers in the next layer of two adjacent layers is greater than or equal to the number of convolutional layers in the previous layer.
[0036] This application also discloses a pedestrian flow counting device, the device comprising: a second acquisition module, configured to acquire target video data within a set time period, the target video data including multiple frames of images; a second feature extraction module, configured to extract human body feature information corresponding to pedestrians in the multiple frames of images using a feature extractor, wherein the feature extractor has a tree-like branch structure; a second recognition module, configured to recognize pedestrians based on the human body feature information corresponding to pedestrians in the multiple frames of images, determine images with the same pedestrians and obtain multiple image groups, the image group including multiple images, different image groups corresponding to different pedestrians; and a pedestrian flow counting module, configured to count the pedestrian flow within the set time period based on the number of image groups and the number of images outside the image groups.
[0037] This application also discloses a trajectory tracking device, comprising: a third acquisition module for acquiring target video data and target pedestrian images, the target video data including multiple frames of images to be detected; a third feature extraction module for extracting human body feature information corresponding to pedestrians in the multiple frames of images to be detected, and extracting human body feature information of the target pedestrian in the target pedestrian image, wherein the feature extractor has a tree-like branch structure; a selection module for comparing the human body feature information corresponding to pedestrians in the multiple frames of images to be detected with the human body feature information of the target pedestrian in the target pedestrian image, and selecting a target image containing the target pedestrian from the multiple frames of images to be detected; and a trajectory generation module for generating the motion trajectory of the target pedestrian based on the target image containing the target pedestrian.
[0038] This application also discloses a detection device, comprising: a shelf determination module, used to determine the target shelf where the unpaid goods are located when the unpaid goods are lost; a fourth acquisition module, used to acquire target video data, the target video data including multiple frames of images to be detected; a fourth feature extraction module, used to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected using a feature extractor, wherein the feature extractor has a tree-like branch structure; an identification grouping module, used to determine images with the same pedestrians and obtain multiple image groups based on the human body feature information corresponding to pedestrians in the multiple frames of images to be detected, the image groups including multiple images, and different pedestrians corresponding to different image groups; a first target pedestrian determination module, used to generate the motion trajectories of multiple pedestrians corresponding to the multiple image groups, and determine the target pedestrian passing through the target shelf based on the motion trajectories of the multiple pedestrians; and a detection module, used to detect whether the target pedestrian has taken the unpaid goods.
[0039] This application also discloses an alarm device, which includes: a fifth acquisition module for acquiring a target image; a fifth feature extraction module for extracting human body feature information corresponding to pedestrians in the target image using a feature extractor, wherein the feature extractor has a tree-like branch structure; a comparison module for comparing the human body feature information corresponding to pedestrians in the target image with the human body feature information corresponding to target pedestrians in a preset blacklist; and an alarm module for performing an alarm when it is determined that a target pedestrian exists in the target image.
[0040] This application also discloses a detection device, comprising: a sixth acquisition module, configured to acquire target video data when a preset product is lost, the target video data including multiple frames of images to be detected; a sixth feature extraction module, configured to extract feature information corresponding to the product in the multiple frames of images to be detected using a feature extractor, wherein the feature extractor has a tree-like branch structure; an image determination module, configured to determine a target image containing the preset product in the multiple frames of images to be detected based on the feature information corresponding to the product in the multiple frames of images to be detected; and a second target pedestrian determination module, configured to detect a target pedestrian taking the preset product from the target image.
[0041] This application also discloses an electronic device, including: a processor; and a memory storing executable code thereon, which, when executed, causes the processor to perform one or more methods as described in this application.
[0042] This application also discloses one or more machine-readable media storing executable code thereon, which, when executed, causes a processor to perform one or more methods as described in this application.
[0043] Compared with the prior art, the embodiments of this application have the following advantages:
[0044] In this embodiment, after acquiring the target image, a feature extractor can be used to extract feature information corresponding to the object in the target image; then, based on the feature information corresponding to the object in the target image, identification is performed to determine whether the target image contains a target object. The feature extractor has a tree-like branching structure, which can increase the number of branches, thereby increasing the accuracy of feature extraction and improving recognition accuracy. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of the branch, the feature extractor with the tree-like branching structure in this embodiment only needs to reach the target number of branches at the end of the branch; therefore, compared to the prior art, this application can achieve the same number of branches using fewer nodes, thus enabling faster feature extraction and higher recognition efficiency. Attached Figure Description
[0045] Figure 1A This is a schematic diagram of the data processing procedure of an identification method according to an embodiment of this application;
[0046] Figure 1B This is a schematic diagram of the structural framework of a feature extractor according to an embodiment of this application;
[0047] Figure 1C This is a flowchart illustrating the steps of one embodiment of the identification method of this application;
[0048] Figure 2 This is a flowchart of the steps of an optional embodiment of the identification method of this application;
[0049] Figure 3 This is a flowchart illustrating the steps of an embodiment of a people flow statistics method according to this application;
[0050] Figure 4 This is a flowchart illustrating the steps of an embodiment of the trajectory tracking method of this application;
[0051] Figure 5 This is a flowchart illustrating the steps of one embodiment of the detection method of this application;
[0052] Figure 6 This is a flowchart illustrating the steps of another embodiment of the detection method of this application;
[0053] Figure 7 This is a flowchart illustrating the steps of an embodiment of an alarm method according to this application;
[0054] Figure 8 This is a structural block diagram of one embodiment of the identification device of this application;
[0055] Figure 9 This is a structural block diagram of an optional embodiment of the identification device of this application;
[0056] Figure 10This is a structural block diagram of an embodiment of a people flow counting device according to this application;
[0057] Figure 11 This is a structural block diagram of an embodiment of a trajectory tracking device according to this application;
[0058] Figure 12 This is a structural block diagram of one embodiment of the detection device of this application;
[0059] Figure 13 This is a structural block diagram of another embodiment of the detection device of this application;
[0060] Figure 14 This is a structural block diagram of an embodiment of an alarm device according to this application;
[0061] Figure 15 This is a schematic diagram of the structure of a device provided in an embodiment of this application. Detailed Implementation
[0062] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0063] This application provides a recognition method capable of identifying whether a target object exists in an image. This recognition method can perform recognition based on feature information extracted by a feature extractor with a tree-like branching structure. The tree-like branching structure allows for the addition of branches, thereby increasing the accuracy of feature extraction and improving recognition accuracy. Furthermore, while existing parallel branching structures require a target number of branches at the beginning of each branch, the tree-like branching structure feature extractor in this application only needs to reach the target number of branches at the end. Therefore, compared to existing technologies, this application can achieve the same number of branches using fewer nodes, thus improving the speed of feature extraction and increasing recognition efficiency.
[0064] Reference Figure 1A The diagram illustrates a data processing procedure for a recognition method according to an embodiment of this application. An image can be input into a feature extractor, which extracts the feature information of the image; then, based on the feature information, recognition is performed to determine whether the image contains a target object.
[0065] In one embodiment of this application, a feature extractor can be pre-built and then used to extract features.
[0066] The step of constructing the feature extractor may include: constructing a convolutional layer with a tree-like branch structure; and sequentially constructing a convolutional computation module and a pooling computation module after the convolutional layer to obtain the feature extractor. The convolutional layer is connected to the convolutional computation module, and the convolutional computation module is connected to the pooling module; (Refer to...) Figure 1B The diagram illustrates a structural framework of a feature extractor according to an embodiment of this application. The feature extraction process, after an image is input to the feature extractor, can proceed as follows: the image is input to a convolutional layer for calculation, and the corresponding result is output to the convolutional calculation module; the convolutional calculation module performs convolutional calculations on the input data and outputs the corresponding result to the pooling calculation module; the pooling calculation module pools the input data and outputs the feature information corresponding to the input data.
[0067] In this embodiment, global and local feature information (including horizontal and vertical local feature information) of the image can be extracted; based on the global and local feature information, recognition accuracy can be improved. Correspondingly, a convolutional layer can also be constructed based on the global and local feature information; the construction of the tree-branched convolutional layer includes: constructing a convolutional layer containing a global feature branch, a horizontal local feature branch, and a vertical local feature branch in a tree-branching manner. The global feature branch can be used to extract global feature information, the horizontal local feature branch is used to extract horizontal local feature information, and the vertical local feature branch is used to extract vertical local feature information. Thus, a convolutional layer capable of extracting global, horizontal, and vertical local feature information can be constructed.
[0068] In this embodiment, constructing a convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches using a tree-like branching method includes: constructing an M-layer shared convolutional layer; using the M-layer shared convolutional layer as the root node, constructing an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches using a tree-like branching method. M and N are positive integers, and this embodiment does not limit the size of M and N (e.g., ...). Figure 1B In the case where N=3), the sum of M and N represents the total number of convolutional layers. We can first construct M convolutional layers, which have no branches and are shared by the subsequent N convolutional layers that have branches. For example... Figure 1B As shown in X1. Then, after the Mth layer, N convolutional layers with a branching structure are constructed; wherein, the Mth convolutional layer can be used as the root node, and N convolutional layers with multiple branches can be constructed in a tree-like branching manner. The branches contained in these N convolutional layers are: global feature branches, horizontal local feature branches, and vertical local feature branches.
[0069] In one optional embodiment of this application, the step of constructing an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, using the shared convolutional layer of the Mth layer as the root node, includes: determining a first number corresponding to the global feature branches, a second number corresponding to the horizontal local feature branches, and a third number corresponding to the vertical local feature branches; determining the horizontal local information corresponding to the second number of horizontal local feature branches, and determining the vertical local information corresponding to the third number of vertical local feature branches; and constructing an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, based on the first number, second number, third number, horizontal local information, and vertical local information.
[0070] Specifically, the number of global feature branches, the number of horizontal local feature branches, and the number of vertical local feature branches can be determined in advance based on requirements such as recognition accuracy and efficiency. In other words, the total number of branches to be constructed needs to be determined. In this embodiment, there are no restrictions on the number of branches used by the convolutional layer to extract global feature information, the number of branches used by the convolutional layer to extract horizontal local feature information, and the number of branches used by the convolutional layer to extract vertical local feature information. Figure 1B The convolutional layer has 8 branches. The first number of global feature branches is 1, corresponding to branch 1; the second number of horizontal local feature branches is 6, corresponding to branches 2 to 6; and the third number of vertical local branches is 2, corresponding to branches 7 and 8. Of course, the first, second, and third numbers can be other values, and this embodiment does not limit them.
[0071] Then, based on the requirements and the second number of horizontal local feature branches, determine which horizontal local feature information to extract for each horizontal local feature branch; and based on the requirements and the third number of vertical local feature branches, determine which vertical local feature information to extract for each vertical local feature branch.
[0072] In one example of this application, the convolutional layer of the feature extractor can divide the image into multiple sub-blocks when extracting local feature information; then, it extracts the feature information of at least one of these sub-blocks to obtain local feature information. Specifically, when extracting horizontal local feature information, the image can be horizontally divided into multiple sub-blocks; when extracting vertical local feature information, the image can be vertically divided into multiple sub-blocks. For example, assuming the image is horizontally divided into 6 sub-blocks, the number of horizontal local feature branches is 5; one branch, such as branch 2, can be used to extract the feature information of sub-blocks 1-5 and sub-blocks 2-6, and the horizontal local information corresponding to this branch can be: horizontal sub-blocks 1-5 and horizontal sub-blocks 2-6. Another branch, such as branch 3, extracts the feature information of sub-blocks 1-4, 2-5, and 3-6 respectively; the horizontal local information corresponding to this branch can be: horizontal sub-blocks 1-4, horizontal sub-blocks 2-5, and horizontal sub-blocks 3-6. Another branch, such as branch 4, extracts feature information from sub-blocks 1-3, 2-4, 3-5, and 4-6 respectively. The corresponding horizontal local information for this branch could be: horizontal sub-blocks 1-3, 2-4, 3-5, and 4-6; and so on. For example, suppose the image is vertically divided into 8 sub-blocks, with 2 horizontal local feature branches. One branch, such as branch 7, extracts feature information from the middle 4 sub-blocks; the corresponding vertical local information for this branch could be: horizontal sub-blocks 3-6. The other branch, such as branch 8, extracts feature information from the middle 6 sub-blocks; the corresponding vertical local information for this branch could be: horizontal sub-blocks 2-7.
[0073] Then, using the Mth convolutional layer as the root node, a first number of global feature branches that can extract global feature information can be constructed in a tree-like manner, a second number of horizontal local feature branches that can extract corresponding horizontal local information and corresponding local feature information can be constructed, and a third number of vertical local feature branches that can extract corresponding vertical local information and corresponding local feature information can be constructed; thus, an N-layer convolutional layer with the number of branches being the sum of the first number, the second number, and the third number can be obtained.
[0074] In this embodiment of the application, the number of branches in each branch is not limited, and the number of branches in each branch can be greater than or equal to two; for example Figure 1B The diagram shows that the number of branches in each branch is 2. Furthermore, in the subsequent N convolutional layers, each convolutional layer can have at least one branch.
[0075] In an optional embodiment of this application, the number of convolutional layers in the subsequent N convolutional layers is greater than or equal to the number of convolutional layers in the preceding layer. That is, in the subsequent N layers, the branching interval is not limited; branching can occur in every layer, or with a gap of one or more layers. For example, Figure 1B This shows the case where each of the last three layers branches.
[0076] In one example of this application, the feature extractor may be a neural network structure.
[0077] In this embodiment, assuming the convolutional layer has 8 branches and 10 layers, with branching starting from the 8th layer, the existing parallel branching structure would require 8 branches at the 8th layer. Consequently, each of the last 3 layers would contain 8 convolutional layers, resulting in a total of (8+8+8) = 24 convolutional layers. The tree-like branching structure of this application, as described in the embodiment... Figure 1B The branching method results in a total of (2+4+8) = 14 convolutional layers for the last three layers; that is, the 8th layer includes 2 convolutional layers, the 9th layer includes 4 convolutional layers, and the 10th layer includes 8 convolutional layers. Alternatively, starting from the 8th convolutional layer, there are 4 branches; the 9th layer does not branch (meaning the second-to-last layer has 4 branches); and each branch in the 10th layer is further divided into 2 branches (meaning the last layer has 8 branches). Therefore, the total number of the last three convolutional layers is (4+4+8) = 16 layers. It is evident that regardless of the form of tree-like branching used in this embodiment, the number of nodes (i.e., the number of convolutional layers) in the feature extractor of this embodiment is less than that of existing feature extractors. Consequently, the feature extractor of this embodiment extracts features faster, thereby improving recognition efficiency. Furthermore, the branching method in this embodiment can add specific local feature branches, which can improve the accuracy of feature extraction, thereby improving recognition accuracy.
[0078] After the image is input to the feature extractor, the feature extractor performs the feature extraction process as follows: the image is input into the convolutional layer for calculation; then the output of the convolutional layer of each branch is convolved, and the convolutional calculation result of each branch is pooled to output the feature information corresponding to the image.
[0079] Reference Figure 1C The diagram shows a flowchart of one embodiment of the identification method of this application.
[0080] Step 102: Obtain the target image.
[0081] When it is necessary to detect whether an image contains a target object, the image can be acquired and identified as the target image. The target object can be a person or an object; this embodiment of the application does not impose any limitations on this.
[0082] Step 104: Use a feature extractor to extract the feature information corresponding to the object in the target image, wherein the feature extractor has a tree-like branching structure.
[0083] In this embodiment, the feature extractor with a tree-like branching structure described above can be used to extract feature information corresponding to objects in the target image; thereby improving the speed and accuracy of feature extraction.
[0084] Step 104 may include the following sub-steps S1042-S1044:
[0085] S1042. Input the target image into the feature extractor to obtain the global feature information, horizontal local feature information and vertical local feature information corresponding to the target image.
[0086] The target image can be input into a feature extractor, whose convolutional layers can extract features from the target image. Each branch of the convolutional layer can extract different features; some branches extract global features of the object in the target image, while others extract local features. These local features can include horizontal and vertical features. The horizontal features can be obtained by dividing the target image horizontally into multiple sub-blocks and then extracting features from each sub-block. Similarly, the vertical features can be obtained by dividing the target image vertically into multiple sub-blocks and then extracting features from each sub-block.
[0087] In one example of this application, in the multiple branches of the convolutional layer used to extract horizontal local feature information, each branch can extract multiple horizontal local feature information, and the number of horizontal local feature information extracted by each branch can be different. For example Figure 1BIn this model, the convolutional layer has five branches for extracting horizontal local features: branch 2, branch 3, branch 4, branch 5, and branch 6. Branch 2 extracts 2 horizontal local features, branch 3 extracts 3, branch 4 extracts 4, branch 5 extracts 5, and branch 6 extracts 6. Assuming the target image is horizontally divided into 6 sub-blocks, branch 2 extracts features from sub-blocks 1-5 and sub-blocks 2-6, resulting in 2 horizontal local features. Correspondingly, branch 3 extracts features from sub-blocks 1-4, 2-5, and 3-6, resulting in 3 horizontal local features. Branch 4 extracts features from sub-blocks 1-3, 2-4, 3-5, and 4-6, resulting in 4 horizontal local features; and so on.
[0088] In one example of this application, in the multiple branches of the convolutional layer used to extract vertical local feature information, each branch can extract one vertical local feature information, and the local areas corresponding to the vertical local feature information extracted by each branch are different. For example, Figure 1B In the diagram, the convolutional layer has two branches for extracting vertical local features: branch 7 and branch 8. Both branch 7 and branch 8 extract one vertical local feature. Assuming the target image is vertically divided into eight sub-blocks, branch 7 extracts the feature information of the middle four sub-blocks, obtaining the corresponding vertical local feature information; branch 8 extracts the feature information of the middle six sub-blocks, obtaining the corresponding vertical local feature information.
[0089] Then, convolution and pooling calculations are performed sequentially on the features output from each branch of the convolutional layer, and the results are output. This yields the global feature information, horizontal local feature information, and vertical local feature information corresponding to the target image.
[0090] S1044. Using the global feature information, horizontal local feature information and vertical local feature information corresponding to the target image, generate feature information corresponding to the object in the target image.
[0091] Specifically, the global feature information, horizontal local feature information, and vertical local feature information corresponding to the target image can be used to generate the feature information corresponding to the object in the target image; thus, the feature information corresponding to the object in the target image can be obtained through the above method.
[0092] Step 106: Identify the target image based on the feature information corresponding to the object in the target image, and determine whether the target image has a target object.
[0093] In an optional embodiment of this application, step 106 may include: determining the distance between the feature information corresponding to the object in the target image and the feature information corresponding to the target object, and determining whether the target image has a target object based on the distance. The methods for determining the distance between the feature information corresponding to the object in the target image and the feature information corresponding to the target object can be varied, and the calculated distance can also include various methods such as Euclidean distance, cosine distance, etc., and this embodiment of the application does not limit this.
[0094] In one embodiment of this application, determining whether the target image has a target object based on the distance includes: determining whether the distance is less than a preset distance threshold; if the distance is less than or equal to the preset distance threshold, then determining that the target image has a target object; if the distance is greater than the preset distance threshold, then determining that the target image does not have a target object. The preset distance threshold can be set as needed, and this embodiment does not impose any restrictions on it.
[0095] In summary, in this embodiment, after acquiring the target image, a feature extractor can be used to extract the feature information corresponding to the object in the target image; then, based on the feature information corresponding to the object in the target image, identification is performed to determine whether the target image contains a target object. The feature extractor has a tree-like branching structure, which can increase the number of branches, thereby increasing the accuracy of feature extraction and improving recognition accuracy. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of the branch, while the feature extractor with the tree-like branching structure in this embodiment only needs to reach the target number of branches at the end of the branch, this application, compared to existing technologies, can achieve the same number of branches using fewer nodes, thus enabling faster feature extraction and higher recognition efficiency.
[0096] Based on the above embodiments, the embodiments of this application can also be applied to determine whether at least two images contain the same object.
[0097] Reference Figure 2 The diagram shows a flowchart of an optional embodiment of the identification method of this application.
[0098] Step 202: Obtain at least two target images.
[0099] When it is necessary to detect whether there is the same object in two or more images, these two or more images can be acquired and identified as target images; that is, the acquired target images can include at least two images.
[0100] The target image can be obtained from data collected by one image acquisition device or from data collected by multiple image acquisition devices, depending on the specific requirements. This application does not impose any restrictions on this.
[0101] Step 204: Use a feature extractor to extract feature information corresponding to objects in the at least two target images respectively, wherein the feature extractor has a tree-like branching structure.
[0102] Step 204 may include the following sub-steps 2042 to 2044:
[0103] Sub-step 2042: Input the at least two target images into the feature extractor respectively to obtain the global feature information, horizontal local feature information and vertical local feature information corresponding to the at least two target images respectively;
[0104] Sub-step 2044: Using the global feature information, horizontal local feature information and vertical local feature information corresponding to the two target images respectively, generate feature information corresponding to the objects in the at least two target images.
[0105] Sub-steps 2042 to 2044 are similar to sub-steps 1042 to 1044 described above, and will not be repeated here.
[0106] Step 206: Identify the objects in the at least two target images based on their corresponding feature information, and determine whether the at least two images have the same object.
[0107] When there are two target images, the feature information corresponding to the object in the two target images can be compared to determine whether the two target images contain the target object.
[0108] When there are more than two target images, the feature information corresponding to the objects in each pair of target images can be compared sequentially to determine whether the two pairs of target images have the same object. Then, based on the results of the determination of whether the two pairs of target images have the same object, it can be determined whether the multiple target images have the same object. When any two target images do not have the same object, it can be determined that the multiple target images do not have the same object.
[0109] In an optional embodiment of this application, step 206 may include: determining the distance between feature information corresponding to objects in the at least two target images, and determining whether the at least two target images have the same object based on the distance.
[0110] In one embodiment of this application, determining whether the at least two target images have the same object based on the distance includes: determining whether the distance is less than a preset distance threshold; if the distance is less than or equal to the preset distance threshold, then determining that the at least two target images have the same object; if the distance is greater than the preset distance threshold, then determining that the at least two target images do not have the same object.
[0111] Based on the identification method described in the above embodiments, this application also discloses a people flow statistics method, which can be applied to shopping malls, supermarkets, office areas, roads, shops, etc., and can quickly and accurately count people flow.
[0112] Reference Figure 3 The diagram shows a flowchart of an embodiment of a people flow statistics method according to this application.
[0113] Step 302: Obtain target video data within a set time period, wherein the target video data includes multiple frames of images.
[0114] In this embodiment, when it is necessary to count pedestrian traffic within a set time period, video data within that time period can be acquired and identified as target video data. Then, based on the identification method described above and the target video data within the set time period, pedestrian traffic within that time period is counted. The set time period can be set according to requirements, such as 12:00–2:00, 21:00–23:00, etc., and this embodiment does not impose any restrictions on this.
[0115] For example, if you need to count the foot traffic in a shopping mall / supermarket between 12:00 and 2:00 noon, you can obtain video data collected by image acquisition devices installed at various entrances and exits of the mall / supermarket during the same period and use that video data as the target video data.
[0116] For example, if you need to count the pedestrian traffic of a pedestrian street between 8:00 and 10:00 in the morning, you can obtain the surveillance video data of the pedestrian street during the same period and use that surveillance video data as the target video data.
[0117] The target video data may include multiple frames of images.
[0118] Step 304: Use a feature extractor to extract human body feature information corresponding to pedestrians in the multiple frames of images, wherein the feature extractor has a tree-like branching structure.
[0119] Then, a feature extractor with a tree-like branching structure can be used to extract human body feature information corresponding to pedestrians in each frame of the target video data; this is similar to step 104 above, and will not be repeated here.
[0120] Step 306: Identify the human body feature information corresponding to the pedestrians in the multiple frames of images, determine the images with the same pedestrians and obtain multiple image groups. Each image group includes multiple images, and different image groups correspond to different pedestrians.
[0121] In real-world scenarios, the same pedestrian may appear multiple times at the same location, such as repeatedly entering and exiting a shopping mall, or may remain at the same location for a certain period of time, such as waiting for a friend at the mall entrance. This can result in multiple images in the target video data corresponding to the same pedestrian. Furthermore, some pedestrians may appear multiple times at the same location or remain at the same location for a certain period of time, leading to multiple images in the target video data corresponding to the same pedestrian, but each time the multiple images correspond to a different pedestrian. Therefore, this embodiment of the application can identify pedestrians based on the human feature information corresponding to pedestrians in multiple frames of the target video data, determine images with the same pedestrian, and obtain multiple image groups. Each image group includes multiple images, and different image groups correspond to different pedestrians.
[0122] Specifically, the human feature information of any two frames in the multi-frame image data of the target video can be compared to determine whether the two frames contain images of the same pedestrian; and images containing the same pedestrian are grouped together as an image group. The method described in step 206 above can be used to determine whether any two frames contain images of the same pedestrian, which will not be elaborated further here.
[0123] Step 308: Based on the number of images in the image group and the number of images outside the image group, count the pedestrian flow within a set time period.
[0124] Specifically, the number of image groups can be determined, and this number is used as a portion of the pedestrian traffic within a set time period (referred to as the first pedestrian traffic for ease of explanation). The number of images contained in each image group is also determined, along with the total number of images contained in each image group. Then, the total number of images contained in the target video data is subtracted from the total number of images contained in the image groups to obtain the number of images outside the image groups; this number is then used as another portion of the pedestrian traffic within the set time period (referred to as the second pedestrian traffic for ease of explanation). Finally, the first and second pedestrian traffic are added together to obtain the total pedestrian traffic within the set time period.
[0125] For example, target video data for a shopping mall during the period from 12:00 PM to 2:00 PM includes 1000 frames of images; among them, 80 frames correspond to pedestrian 1, 180 frames correspond to pedestrian 2, 300 frames correspond to pedestrian 3, and the remaining 440 images are images outside of the aforementioned image group. It can be determined that the pedestrian flow in the shopping mall during the period from 12:00 PM to 2:00 PM is 443.
[0126] In summary, in this embodiment, target video data within a set time period can be acquired. Then, a feature extractor is used to extract human body feature information corresponding to pedestrians from multiple frames of the target video data. Based on the human body feature information of the pedestrians in the multiple frames, images with the same pedestrians are identified, resulting in multiple image groups. Each image group includes multiple images, with different pedestrians corresponding to different image groups. Based on the number of image groups and the number of images outside the image groups, the pedestrian flow within the set time period is statistically analyzed. The feature extractor has a tree-like branching structure, which can increase the number of branches, thereby increasing the accuracy of feature extraction and improving the accuracy of pedestrian flow statistics. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of the branch, the feature extractor with the tree-like branching structure in this embodiment only needs to reach the target number of branches at the end of the branch. Therefore, compared to existing technologies, this application can achieve the same number of branches using fewer nodes, resulting in faster feature extraction and improved efficiency in pedestrian flow statistics.
[0127] Based on the identification method described in the above embodiments, this application also discloses a trajectory tracking method that can quickly and accurately track specific pedestrians.
[0128] Reference Figure 4 The diagram shows a flowchart of the steps of an embodiment of the trajectory tracking method of this application.
[0129] Step 402: Obtain target video data and target pedestrian images, wherein the target video data includes multiple frames of images to be detected.
[0130] In this embodiment of the application, when it is necessary to track the trajectory of a pedestrian, the image of the pedestrian can be acquired; wherein, for the convenience of subsequent explanation, the pedestrian for whom trajectory tracking is required can be referred to as the target pedestrian, and the acquired image of the target pedestrian can be referred to as the target pedestrian image.
[0131] This involves acquiring surveillance video data from various locations and roads, and determining target video data based on this data. The target video data can be determined by filtering the surveillance video data based on other information, such as filtering surveillance video data from certain roads / locations. Alternatively, all surveillance video data can be used as the target video data; this embodiment does not impose such limitations. The target video data includes multiple frames of images to be detected.
[0132] Step 404: Use a feature extractor to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected, and extract human body feature information of the target pedestrian in the target pedestrian image, wherein the feature extractor has a tree-like branching structure.
[0133] Then, a feature extractor with a tree-like branching structure can be used to extract the human body feature information corresponding to the pedestrian in each frame of the image to be detected; and a feature extractor with a tree-like branching structure can be used to extract the human body feature information of the target pedestrian in the target pedestrian image. This is similar to step 104 above, and will not be repeated here.
[0134] Step 406: Compare the human body feature information corresponding to pedestrians in the multi-frame images to be detected with the human body feature information of the target pedestrian in the target pedestrian image, and extract the target image containing the target pedestrian from the multi-frame images to be detected.
[0135] Then, the human body feature information corresponding to the pedestrian in each image to be detected can be compared with the human body feature information of the target pedestrian in the target pedestrian image to determine whether there is a target pedestrian in the image to be detected; if there is a target pedestrian in the image to be detected, the image to be detected can be identified as the target image.
[0136] The process of determining whether a target pedestrian exists in the image to be detected is the process of determining whether the image to be detected and the target pedestrian image have the same object; this is similar to step 206 above, and will not be repeated here.
[0137] Step 408: Generate the motion trajectory of the target pedestrian based on the target image containing the target pedestrian.
[0138] After filtering out the target image from multiple images to be detected in the target video data, the identifier of the target image, such as its ID, can be obtained. Then, based on the identifier of each target image, the shooting time and location of each image are determined; and based on the shooting time and location of each image, the motion trajectory of the target pedestrian is generated.
[0139] In summary, in this embodiment, target video data and target pedestrian images can be acquired, and a feature extractor is used to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected, as well as human body feature information of the target pedestrian in the target pedestrian image. Then, the human body feature information corresponding to pedestrians in the multiple frames of images to be detected is compared with the human body feature information of the target pedestrian in the target pedestrian image. A target image containing the target pedestrian is determined from the multiple frames of images to be detected, and the motion trajectory of the target pedestrian is generated based on the target image containing the target pedestrian. The feature extractor has a tree-like branching structure, which can increase the number of branches, thereby increasing the accuracy of feature extraction and improving the accuracy of pedestrian tracking. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of the branch, the feature extractor with the tree-like branching structure in this embodiment only needs to reach the target number of branches at the end of the branch. Therefore, compared with the prior art, this application can achieve the same number of branches using fewer nodes, thus enabling faster feature extraction and improving the efficiency of pedestrian tracking.
[0140] Based on the identification method described in the above embodiments, this application also discloses a detection method that can be applied to scenarios such as shops and supermarkets, and can quickly and accurately detect target pedestrians taking goods based on images of the same pedestrians.
[0141] Reference Figure 5 The diagram shows a flowchart of one embodiment of the detection method of this application.
[0142] Step 502: If it is determined that the unpaid goods are lost, determine the target shelf where the paid goods are located.
[0143] Step 504: Obtain target video data, which includes multiple frames of images to be detected.
[0144] When it is determined that an unpaid item is lost, the target shelf where the unpaid item was located can be identified. Then, surveillance video data of the area corresponding to the target shelf can be obtained and used as the target video data. Based on the target video data and a feature extractor, pedestrian trajectories in that area are generated. Then, based on the trajectories of each pedestrian, target pedestrians who passed the target shelf are identified, and from these target pedestrians, those who took the unpaid item are excluded.
[0145] Step 506: Use a feature extractor to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected, wherein the feature extractor has a tree-like branch structure.
[0146] Step 506 is similar to step 104 above, and will not be described again here.
[0147] Step 508: Based on the human body feature information corresponding to pedestrians in the multiple frames of images to be detected, determine the images with the same pedestrians and obtain multiple image groups. Each image group includes multiple images, and different image groups correspond to different pedestrians.
[0148] Step 508 is similar to step 206 above, and will not be described again here.
[0149] Step 510: Generate the motion trajectories of multiple pedestrians based on the multiple image groups respectively, and determine the target pedestrians passing through the target shelf based on the motion trajectories of the multiple pedestrians.
[0150] Then, based on each image group, the motion trajectory of the pedestrian corresponding to that image group can be generated, which is similar to step 408 above and will not be repeated here. The motion trajectory of each pedestrian is then analyzed to determine the target pedestrians who have passed the target shelf.
[0151] Step 512: Detect whether the target pedestrian has taken the unpaid goods.
[0152] In one embodiment of this application, image recognition can be performed on the image corresponding to the target pedestrian; when the unpaid goods are identified from the image corresponding to the target pedestrian, it can be determined that the target pedestrian has taken the unpaid goods.
[0153] In one example of this application, surveillance video data from other locations and roads can be acquired. Based on the image of the pedestrian who took the unpaid goods and the surveillance video data, the trajectory of the pedestrian can be tracked using the method described in the above embodiment. In another example of this application, the image of the pedestrian who took the unpaid goods can also be reported to relevant departments for alarm purposes.
[0154] In summary, in this embodiment, when it is determined that an unpaid item is lost, the target shelf where the lost item is located can be identified, and target video data can be obtained. Then, a feature extractor is used to extract human body feature information corresponding to pedestrians in multiple frames of images to be detected in the target video data. Based on the human body feature information corresponding to pedestrians in the multiple frames of images to be detected, images with the same pedestrian are identified, resulting in multiple image groups. Each image group includes multiple images, and different pedestrians correspond to different image groups. Then, motion trajectories of multiple pedestrians are generated based on the multiple image groups, and target pedestrians passing through the target shelf are determined based on the motion trajectories of the multiple pedestrians. Whether the target pedestrian has taken the unpaid item is detected. The feature extractor has a tree-like branching structure, which can increase the number of branches, thereby increasing the accuracy of feature extraction and improving the accuracy of detection. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of a branch, while the feature extractor of this application has a tree-like branching structure, it only needs to reach the target number of branches at the end of the branch; thus, compared with the prior art, this application can achieve the same number of branches with fewer nodes; thereby enabling faster detection speed and improving the efficiency of detecting pedestrians taking unpaid goods.
[0155] Based on the identification method described in the above embodiments, this application also discloses a detection method that can be applied to scenarios such as shops and supermarkets, and can quickly and accurately detect pedestrians taking preset goods based on images with preset goods.
[0156] Reference Figure 6 The diagram shows a flowchart of another embodiment of the detection method of this application.
[0157] Step 602: When a preset product is lost, acquire target video data, which includes multiple frames of images to be detected.
[0158] In this embodiment of the application, the preset goods can be set according to needs, such as goods specified by the user, easily lost goods, valuable goods, etc.
[0159] When it is determined that a preset product is lost, surveillance video data of the location where the preset product is located can be obtained and used as target video data.
[0160] Step 604: Use a feature extractor to extract the feature information corresponding to the goods in the multiple frames of images to be detected, wherein the feature extractor has a tree-like branch structure.
[0161] Step 606: Based on the feature information corresponding to the goods in the multi-frame images to be detected, determine the target image with the preset goods in the multi-frame images to be detected.
[0162] In this embodiment of the application, a feature extractor can be used to extract the feature information corresponding to the goods in multiple frames of detection images in the target video data; then, based on the feature information corresponding to the goods in each frame of the detection images, a target image with a preset goods can be found from the multiple frames of the detection images.
[0163] Specifically, the feature information corresponding to each frame of the image to be detected can be compared with the feature information corresponding to a preset product. Images whose distance to the feature information corresponding to the preset product is less than a preset distance are identified and determined as target images. Multiple target images may be included.
[0164] Steps 604-606 can be referred to as steps 204-206 in the above embodiment, and will not be repeated here.
[0165] Step 608: Detect the target pedestrians who take the preset goods from the target image.
[0166] In this embodiment, for a target image, the relationship between the position of a pedestrian's hand and the position of a preset product in the target image can be identified to determine whether the pedestrian has taken the preset product. When it is determined that the pedestrian has taken the preset product, an image of the pedestrian paying for the preset product can be searched in the target image. If no image of the pedestrian paying for the preset product is found, it means that the pedestrian has not paid for the taken preset product, and at this time, the pedestrian can be identified as the target pedestrian.
[0167] In one optional embodiment of this application, after identifying the target pedestrian, surveillance video data from other locations and roads can be acquired. Based on the target image containing the target pedestrian and the surveillance video data, the trajectory of the target pedestrian can be tracked using the method described in the above embodiment. In another example of this application, the image of the target pedestrian can also be reported to relevant departments for alarm purposes, etc.
[0168] In summary, in this embodiment, when a preset product is determined to be lost, target video data can be acquired, and then a feature extractor is used to extract the feature information corresponding to the product in the multiple frames of images to be detected. Based on the feature information corresponding to the product in the multiple frames of images to be detected, a target image containing the preset product is determined, and then the target pedestrian taking the preset product is detected from the target image. The feature extractor has a tree-like branching structure, which can increase the number of branches, thereby increasing the accuracy of feature extraction and improving detection accuracy. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of the branch, while the feature extractor with the tree-like branching structure in this embodiment only needs to reach the target number of branches at the end of the branch, this application, compared to the prior art, can achieve the same number of branches using fewer nodes, thus enabling faster detection and improving the efficiency of detecting pedestrians taking the preset product.
[0169] Based on the identification method described in the above embodiments, this application also discloses an alarm method that can be applied to scenarios such as shops and supermarkets, and can quickly and accurately detect and trigger an alarm.
[0170] Reference Figure 7 The diagram shows a flowchart of an embodiment of an alarm method according to this application.
[0171] Step 702: Obtain the target image.
[0172] Step 704: Use a feature extractor to extract human body feature information corresponding to pedestrians in the target image, wherein the feature extractor has a tree-like branching structure.
[0173] Step 706: Compare the human body feature information corresponding to the pedestrian in the target image with the human body feature information corresponding to the target pedestrian in the target pedestrian image in the preset blacklist.
[0174] Step 708: When it is determined that there is a target pedestrian in the target image, an alarm is triggered.
[0175] In this embodiment, a blacklist can be generated in advance using images of target pedestrians; the target pedestrians can be designated pedestrians or criminals, and this embodiment does not impose any restrictions on this. Alternatively, a feature extractor can be used to extract the human body feature information of the target pedestrians from the images, and then this human body feature information can be used to generate the blacklist; this embodiment does not impose any restrictions on this.
[0176] During the operation of the monitoring equipment, each video frame captured in real time by the monitoring equipment can be used as the target image. Then, a feature extractor with a tree-like branching structure is used to extract the human body feature information corresponding to the pedestrian in the target image. The human body feature information corresponding to the pedestrian in the target image is then compared with the human body feature information corresponding to the target pedestrian in the target pedestrian image to determine whether the target pedestrian exists in the target image. This is similar to the above embodiment and will not be described again here.
[0177] In this embodiment, after acquiring the target image, a feature extractor with a tree-like branching structure can be used to extract the human body feature information corresponding to the target pedestrian in the target pedestrian image; alternatively, after generating a blacklist using the target pedestrian image, a feature extractor with a tree-like branching structure can be used to extract the human body feature information corresponding to the target pedestrian in each target pedestrian image in the blacklist, so as to improve the efficiency of determining the presence of the target pedestrian in the target image and to promptly perform alarm processing; this application embodiment does not limit this.
[0178] When it is determined that a target pedestrian exists in the target image, an alarm can be triggered, such as issuing a voice alarm, playing an alarm tone, or reporting to relevant departments. When it is determined that no target pedestrian exists in the target image, the next target image can be acquired to determine whether a target pedestrian exists in the next target image.
[0179] In one example of this application, the alarm method can be applied to a supermarket. A blacklist can be generated by using historically identified images of pedestrians who have taken unpaid goods. Then, real-time images of the target pedestrians captured by the supermarket's entrance and exit monitoring equipment can be obtained, and steps 604-606 can be executed. Upon confirming the presence of a target pedestrian in the image, an alarm can be sent to the supermarket's security department, providing information such as the pedestrian's location and image, facilitating appropriate action by the supermarket's security personnel.
[0180] In summary, in this embodiment, after acquiring the target image in real time, a feature extractor can be used to extract the human feature information corresponding to the target image and the human feature information corresponding to the target pedestrian image in a preset blacklist. Then, the human feature information corresponding to the target image is compared with the human feature information corresponding to the target pedestrian image, and an alarm is triggered when a target pedestrian is found in the target image. The feature extractor has a tree-like branching structure, which increases the number of branches and thus increases the accuracy of feature extraction, thereby reducing the false alarm rate. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of the branch, the feature extractor with the tree-like branching structure in this embodiment only needs to reach the target number of branches at the end of the branch. Therefore, compared to existing technologies, this application can achieve the same number of branches using fewer nodes, thus enabling timely alarm triggering.
[0181] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of this application.
[0182] Based on the above-described identification method embodiments, this embodiment also provides an identification device that can be applied to electronic devices such as terminal devices and servers.
[0183] Reference Figure 8 The diagram shows a structural block diagram of an embodiment of the identification device of this application, which may specifically include the following modules:
[0184] The first acquisition module 802 is used to acquire the target image;
[0185] The first feature extraction module 804 is used to extract feature information corresponding to objects in the target image using a feature extractor, wherein the feature extractor has a tree-like branching structure;
[0186] The first recognition module 806 is used to identify whether the target image has a target object based on the feature information corresponding to the object in the target image.
[0187] Reference Figure 9 The diagram shows a structural block diagram of an optional embodiment of the identification device of this application.
[0188] In one optional embodiment of this application, the first identification module 802 is used to determine the distance between the feature information corresponding to the object in the target image and the feature information corresponding to the target object, and to determine whether the target image has a target object based on the distance.
[0189] In one optional embodiment of this application, the first feature extraction module 804 is used to input the target image into the feature extractor to obtain global feature information, horizontal local feature information and vertical local feature information corresponding to the target image; and to generate feature information corresponding to the object in the target image using the global feature information, horizontal local feature information and vertical local feature information corresponding to the target image.
[0190] In one optional embodiment of this application, the apparatus further includes:
[0191] Module 808 is used to construct the feature extractor:
[0192] The construction module 808 includes:
[0193] Convolutional layer construction submodule 8082 is used to construct convolutional layers with a tree-like branching structure;
[0194] The computation module construction submodule 8084 is used to sequentially construct a convolution computation module and a pooling computation module after the convolutional layer to obtain the feature extractor.
[0195] In one optional embodiment of this application, the convolutional layer construction submodule 8082 is used to construct a convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner.
[0196] In one optional embodiment of this application, the convolutional layer construction submodule 8082 is used to construct a convolutional layer shared by M layers; taking the convolutional layer shared by the Mth layer as the root node, and constructing an N-layer convolutional layer containing global feature branches, horizontal local feature branches and vertical local feature branches in a tree-like branching manner.
[0197] In one optional embodiment of this application, the convolutional layer construction submodule 8082 is used to determine the first number of global feature branches, the second number of horizontal local feature branches, and the third number of vertical local feature branches; determine the horizontal local information corresponding to the second number of horizontal local feature branches, and determine the vertical local information corresponding to the third number of vertical local feature branches; and construct an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, based on the first number, the second number, the third number, the horizontal local information, and the vertical local information.
[0198] In one optional embodiment of this application, the number of convolutional layers in the next N convolutional layers is greater than or equal to the number of convolutional layers in the previous layer.
[0199] In summary, in this embodiment, after acquiring the target image, a feature extractor can be used to extract the feature information corresponding to the object in the target image; then, based on the feature information corresponding to the object in the target image, identification is performed to determine whether the target image contains a target object. The feature extractor has a tree-like branching structure, which can increase the number of branches, thereby increasing the accuracy of feature extraction and improving recognition accuracy. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of the branch, while the feature extractor with the tree-like branching structure in this embodiment only needs to reach the target number of branches at the end of the branch, this application, compared to the prior art, can achieve the same number of branches using fewer nodes, thus enabling faster feature extraction and higher recognition efficiency.
[0200] Based on the above-described method for counting people, this embodiment also provides a device for counting people, which can be used in electronic devices such as terminal devices and servers.
[0201] Reference Figure 10 The diagram shows a structural block diagram of an embodiment of a people flow counting device according to this application.
[0202] The second acquisition module 1002 is used to acquire target video data within a set time period, wherein the target video data includes multiple frames of images.
[0203] The second feature extraction module 1004 is used to extract human body feature information corresponding to pedestrians in the multi-frame images using a feature extractor, wherein the feature extractor has a tree-like branch structure;
[0204] The second recognition module 1006 is used to recognize the human body feature information corresponding to the pedestrians in the multiple frames of images, determine the images with the same pedestrians and obtain multiple image groups, wherein the image group includes multiple images and the pedestrians corresponding to different image groups are different.
[0205] The pedestrian flow statistics module 1008 is used to count the pedestrian flow within a set time period based on the number of images in the image group and the number of images outside the image group.
[0206] In summary, in this embodiment, target video data within a set time period can be acquired. Then, a feature extractor is used to extract human body feature information corresponding to pedestrians from multiple frames of the target video data. Based on the human body feature information of the pedestrians in the multiple frames, images with the same pedestrians are identified, resulting in multiple image groups. Each image group includes multiple images, with different pedestrians corresponding to different image groups. Based on the number of image groups and the number of images outside the image groups, the pedestrian flow within the set time period is statistically analyzed. The feature extractor has a tree-like branching structure, which can increase the number of branches, thereby increasing the accuracy of feature extraction and improving the accuracy of pedestrian flow statistics. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of the branch, the feature extractor with the tree-like branching structure in this embodiment only needs to reach the target number of branches at the end of the branch. Therefore, compared to existing technologies, this application can achieve the same number of branches using fewer nodes, resulting in faster feature extraction and improved efficiency in pedestrian flow statistics.
[0207] Based on the above-described trajectory tracking method embodiments, this embodiment also provides a trajectory tracking device, which can be applied to electronic devices such as terminal devices and servers.
[0208] Reference Figure 11 The diagram shows a structural block diagram of an embodiment of the trajectory tracking device of this application.
[0209] The third acquisition module 1102 is used to acquire target video data and target pedestrian images, wherein the target video data includes multiple frames of images to be detected;
[0210] The third feature extraction module 1104 is used to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected, and to extract human body feature information of the target pedestrian in the target pedestrian image, wherein the feature extractor has a tree-like branch structure;
[0211] The selection module 1106 is used to compare the human body feature information corresponding to pedestrians in the multi-frame images to be detected with the human body feature information of the target pedestrian in the target pedestrian image, and select the target image containing the target pedestrian from the multi-frame images to be detected.
[0212] The trajectory generation module 1108 is used to generate the motion trajectory of the target pedestrian based on the target image containing the target pedestrian.
[0213] In summary, in this embodiment, target video data and target pedestrian images can be acquired, and a feature extractor is used to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected, as well as human body feature information of the target pedestrian in the target pedestrian image. Then, the human body feature information corresponding to pedestrians in the multiple frames of images to be detected is compared with the human body feature information of the target pedestrian in the target pedestrian image. A target image containing the target pedestrian is determined from the multiple frames of images to be detected, and the motion trajectory of the target pedestrian is generated based on the target image containing the target pedestrian. The feature extractor has a tree-like branching structure, which can increase the number of branches, thereby increasing the accuracy of feature extraction and improving the accuracy of pedestrian tracking. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of the branch, the feature extractor with the tree-like branching structure in this embodiment only needs to reach the target number of branches at the end of the branch. Therefore, compared with the prior art, this application can achieve the same number of branches using fewer nodes, thus enabling faster feature extraction and improving the efficiency of pedestrian tracking.
[0214] Based on the above-described detection method embodiments, this embodiment also provides a detection device that can be applied to electronic devices such as terminal devices and servers.
[0215] Reference Figure 12 The diagram shows a structural block diagram of an embodiment of the detection device of this application.
[0216] The shelf determination module 1202 is used to determine the target shelf where the unpaid goods are located when the unpaid goods are lost.
[0217] The fourth acquisition module 1204 is used to acquire target video data, which includes multiple frames of images to be detected.
[0218] The fourth feature extraction module 1206 is used to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected using a feature extractor, wherein the feature extractor has a tree-like branch structure;
[0219] The identification grouping module 1208 is used to determine images with the same pedestrian and obtain multiple image groups based on the human body feature information corresponding to pedestrians in the multiple frames of images to be detected. The image group includes multiple images, and different image groups correspond to different pedestrians.
[0220] The first target pedestrian determination module 1210 is used to generate the motion trajectories of multiple pedestrians based on the multiple image groups respectively, and to determine the target pedestrians passing through the target shelf based on the motion trajectories of the multiple pedestrians.
[0221] The detection module 1212 is used to detect whether the target pedestrian has taken the unpaid goods.
[0222] In summary, in this embodiment, when it is determined that an unpaid item is lost, the target shelf where the lost item is located can be identified, and target video data can be obtained. Then, a feature extractor is used to extract human body feature information corresponding to pedestrians in multiple frames of images to be detected in the target video data. Based on the human body feature information corresponding to pedestrians in the multiple frames of images to be detected, images with the same pedestrian are identified, resulting in multiple image groups. Each image group includes multiple images, and different pedestrians correspond to different image groups. Then, motion trajectories of multiple pedestrians are generated based on the multiple image groups, and target pedestrians passing through the target shelf are determined based on the motion trajectories of the multiple pedestrians. Whether the target pedestrian has taken the unpaid item is detected. The feature extractor has a tree-like branching structure, which can increase the number of branches, thereby increasing the accuracy of feature extraction and improving the accuracy of detection. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of a branch, while the feature extractor of this application has a tree-like branching structure, it only needs to reach the target number of branches at the end of the branch; thus, compared with the prior art, this application can achieve the same number of branches with fewer nodes; thereby enabling faster detection speed and improving the efficiency of detecting pedestrians taking unpaid goods.
[0223] Based on the above-described detection method embodiments, this embodiment also provides a detection device that can be applied to electronic devices such as terminal devices and servers.
[0224] Reference Figure 13 The diagram shows a structural block diagram of an optional embodiment of the detection device of this application.
[0225] The sixth acquisition module 1302 is used to acquire target video data when a preset product is lost, wherein the target video data includes multiple frames of images to be detected;
[0226] The sixth feature extraction module 1304 is used to extract feature information corresponding to the goods in the multiple frames of images to be detected using a feature extractor, wherein the feature extractor has a tree-like branch structure;
[0227] The image determination module 1306 is used to determine the target image with a preset product in the multi-frame images to be detected based on the feature information corresponding to the product in the multi-frame images to be detected;
[0228] The second target pedestrian determination module 1308 is used to detect the target pedestrians taking the preset goods from the target image.
[0229] In summary, in this embodiment, when a preset product is determined to be lost, target video data can be acquired, and then a feature extractor is used to extract the feature information corresponding to the product in the multiple frames of images to be detected. Based on the feature information corresponding to the product in the multiple frames of images to be detected, a target image containing the preset product is determined, and then the target pedestrian taking the preset product is detected from the target image. The feature extractor has a tree-like branching structure, which can increase the number of branches, thereby increasing the accuracy of feature extraction and improving detection accuracy. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of the branch, while the feature extractor with the tree-like branching structure in this embodiment only needs to reach the target number of branches at the end of the branch, this application, compared to the prior art, can achieve the same number of branches using fewer nodes, thus enabling faster detection and improving the efficiency of detecting pedestrians taking the preset product.
[0230] Based on the above-described alarm method embodiments, this embodiment also provides an alarm device that can be applied to electronic devices such as terminal devices and servers.
[0231] Reference Figure 14 The diagram shows a structural block diagram of an optional embodiment of the detection device of this application.
[0232] The fifth acquisition module 1402 is used to acquire the target image;
[0233] The fifth feature extraction module 1404 is used to extract human body feature information corresponding to pedestrians in the target image using a feature extractor, wherein the feature extractor has a tree-like branch structure;
[0234] The comparison module 1406 is used to compare the human body feature information corresponding to the pedestrian in the target image with the human body feature information corresponding to the target pedestrian in the target pedestrian image in the preset blacklist.
[0235] The alarm module 1408 is used to trigger an alarm when it is determined that a target pedestrian exists in the target image.
[0236] In summary, in this embodiment, after acquiring the target image in real time, a feature extractor can be used to extract the human feature information corresponding to the target image and the human feature information corresponding to the target pedestrian image in a preset blacklist. Then, the human feature information corresponding to the target image is compared with the human feature information corresponding to the target pedestrian image, and an alarm is triggered when a target pedestrian is found in the target image. The feature extractor has a tree-like branching structure, which increases the number of branches and thus increases the accuracy of feature extraction, thereby reducing the false alarm rate. Furthermore, since existing parallel branching structures require a target number of branches at the beginning of the branch, the feature extractor with the tree-like branching structure in this embodiment only needs to reach the target number of branches at the end of the branch. Therefore, compared to existing technologies, this application can achieve the same number of branches using fewer nodes, thus enabling timely alarm triggering.
[0237] This application also provides a non-volatile readable storage medium storing one or more modules (programs). When these modules are applied to a device, they enable the device to execute the instructions for the method steps in this application.
[0238] This application provides one or more machine-readable media storing instructions that, when executed by one or more processors, cause an electronic device to perform one or more of the methods described in the above embodiments. In this application, the electronic device includes various types of devices such as terminal devices and servers (clusters).
[0239] The embodiments of this disclosure can be implemented as an apparatus configured as desired using any suitable hardware, firmware, software, or any combination thereof, including electronic devices such as terminal devices, servers (clusters), etc. Figure 15 An exemplary apparatus 1500 is schematically shown that can be used to implement the various embodiments described in this application.
[0240] In one embodiment, Figure 15 An exemplary device 1500 is shown, which includes one or more processors 1502, a control module (chipset) 1504 coupled to at least one of the processors 1502, a memory 1506 coupled to the control module 1504, a non-volatile memory (NVM) / storage device 1508 coupled to the control module 1504, one or more input / output devices 1510 coupled to the control module 1504, and a network interface 1512 coupled to the control module 1504.
[0241] Processor 1502 may include one or more single-core or multi-core processors, and processor 1502 may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, device 1500 can serve as a terminal device, server (cluster), or other device as described in the embodiments of this application.
[0242] In some embodiments, apparatus 1500 may include one or more computer-readable media (e.g., memory 1506 or NVM / storage device 1508) having instructions 1514 and one or more processors 1502 that are combined with the one or more computer-readable media and configured to execute instructions 1514 to implement a module thereby performing the actions described in this disclosure.
[0243] In one embodiment, the control module 1504 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1502 and / or any suitable device or component communicating with the control module 1504.
[0244] The control module 1504 may include a memory controller module to provide an interface to the memory 1506. The memory controller module may be a hardware module, a software module, and / or a firmware module.
[0245] Memory 1506 may be used, for example, to load and store data and / or instructions 1514 for device 1500. In one embodiment, memory 1506 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, memory 1506 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).
[0246] In one embodiment, the control module 1504 may include one or more input / output controllers to provide interfaces to the NVM / storage device 1508 and (one or more) input / output devices 1510.
[0247] For example, NVM / storage device 1508 may be used to store data and / or instructions 1514. NVM / storage device 1508 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more optical disc drives (CDs), and / or one or more digital universal optical disc (DVD) drives).
[0248] NVM / storage device 1508 may include storage resources that are physically part of a device on which device 1500 is mounted, or that can be accessed by the device without needing to be part of the device. For example, NVM / storage device 1508 may be accessed via a network via one or more input / output devices 1510.
[0249] One or more input / output devices 1510 may provide an interface for device 1500 to communicate with any other suitable device. Input / output devices 1510 may include communication components, audio components, sensor components, etc. Network interface 1512 may provide an interface for device 1500 to communicate via one or more networks. Device 1500 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G, 5G, etc., or combinations thereof.
[0250] In one embodiment, at least one of the processors 1502 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 1504. In one embodiment, at least one of the processors 1502 may be logically packaged with one or more controllers of the control module 1504 to form a system-in-package (SiP). In one embodiment, at least one of the processors 1502 may be integrated with the logic of one or more controllers of the control module 1504 on the same die. In one embodiment, at least one of the processors 1502 may be integrated with the logic of one or more controllers of the control module 1504 on the same die to form a system-on-a-chip (SoC).
[0251] In various embodiments, device 1500 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop, handheld computing device, tablet, netbook, etc.). In various embodiments, device 1500 may have more or fewer components and / or a different architecture. For example, in some embodiments, device 1500 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0252] The detection device may use a main control chip as a processor or control module, and sensor data, position information, etc. may be stored in a memory or NVM / storage device. The sensor group may be used as an input / output device, and the communication interface may include a network interface.
[0253] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0254] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0255] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0256] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0257] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable terminal equipment, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0258] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0259] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0260] The foregoing has provided a detailed description of a method and apparatus for identification, people flow statistics, tracking, detection and alarm, an electronic device and a storage medium provided by this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A recognition method, characterized in that, The method includes: Acquire the target image; A feature extractor is used to extract feature information corresponding to objects in the target image, wherein the feature extractor has a tree-like branching structure; Based on the feature information corresponding to the object in the target image, it is determined whether the target image contains a target object; The method further includes the step of constructing the feature extractor: Construct a convolutional layer shared by M layers; Determine the first number corresponding to the global feature branch, the second number corresponding to the horizontal local feature branch, and the third number corresponding to the vertical local feature branch, respectively; Determine the horizontal local information corresponding to the second number of horizontal local feature branches, and determine the vertical local information corresponding to the third number of vertical local feature branches; Using the Mth convolutional layer as the root node, and in a tree-like branching manner, based on the first quantity, the second quantity, the third quantity, horizontal local information, and vertical local information, construct an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches. After the N convolutional layers, a convolutional computation module and a pooling computation module are constructed sequentially to obtain the feature extractor.
2. The method according to claim 1, characterized in that, The step of identifying whether the target image contains a target object based on the feature information corresponding to the object in the target image includes: Determine the distance between the feature information corresponding to the object in the target image and the feature information corresponding to the target object, and determine whether the target image has a target object based on the distance.
3. The method according to claim 1, characterized in that, The step of extracting feature information corresponding to objects in the target image using a feature extractor includes: The target image is input into the feature extractor to obtain global feature information, horizontal local feature information and vertical local feature information corresponding to the target image; The feature information corresponding to the object in the target image is generated by using the global feature information, horizontal local feature information and vertical local feature information corresponding to the target image.
4. The method according to claim 1, characterized in that, In the N convolutional layers, the number of convolutional layers in the next layer is greater than or equal to the number of convolutional layers in the previous layer.
5. A method for counting pedestrian flow, characterized in that, The method includes: Acquire target video data within a set time period, wherein the target video data includes multiple frames of images; A feature extractor is used to extract human body feature information corresponding to pedestrians in the multiple frames of images, wherein the feature extractor has a tree-like branching structure; Based on the human body feature information of pedestrians in the multiple frames of images, images with the same pedestrian are identified and multiple image groups are obtained. Each image group includes multiple images, and different image groups correspond to different pedestrians. Based on the number of images in the image group and the number of images outside the image group, the pedestrian flow within a set time period is counted. The method further includes the step of constructing the feature extractor: Construct a convolutional layer shared by M layers; Determine the first number corresponding to the global feature branch, the second number corresponding to the horizontal local feature branch, and the third number corresponding to the vertical local feature branch, respectively; Determine the horizontal local information corresponding to the second number of horizontal local feature branches, and determine the vertical local information corresponding to the third number of vertical local feature branches; Using the Mth convolutional layer as the root node, and in a tree-like branching manner, based on the first quantity, the second quantity, the third quantity, horizontal local information, and vertical local information, construct an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches. After the N convolutional layers, a convolutional computation module and a pooling computation module are constructed sequentially to obtain the feature extractor.
6. A trajectory tracking method, characterized in that, The method includes: Acquire target video data and target pedestrian images, wherein the target video data includes multiple frames of images to be detected; A feature extractor is used to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected, and to extract human body feature information of the target pedestrian in the target pedestrian image. The feature extractor has a tree-like branching structure. The human body feature information corresponding to pedestrians in the multi-frame images to be detected is compared with the human body feature information of the target pedestrian in the target pedestrian image, and the target image containing the target pedestrian is selected from the multi-frame images to be detected. Based on the target image containing the target pedestrian, generate the movement trajectory of the target pedestrian; The method further includes the step of constructing the feature extractor: Construct a convolutional layer shared by M layers; Determine the first number corresponding to the global feature branch, the second number corresponding to the horizontal local feature branch, and the third number corresponding to the vertical local feature branch, respectively; Determine the horizontal local information corresponding to the second number of horizontal local feature branches, and determine the vertical local information corresponding to the third number of vertical local feature branches; Using the Mth convolutional layer as the root node, and in a tree-like branching manner, based on the first quantity, the second quantity, the third quantity, horizontal local information, and vertical local information, construct an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches. After the N convolutional layers, a convolutional computation module and a pooling computation module are constructed sequentially to obtain the feature extractor.
7. A detection method, characterized in that, The method includes: When it is determined that an unpaid item is lost, the target shelf where the unpaid item is located is determined. Acquire target video data, which includes multiple frames of images to be detected; A feature extractor is used to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected, wherein the feature extractor has a tree-like branching structure; Based on the human body feature information of pedestrians in the multiple frames of images to be detected, images with the same pedestrians are identified and multiple image groups are obtained. Each image group includes multiple images, and different image groups correspond to different pedestrians. Based on the multiple image groups, corresponding motion trajectories of multiple pedestrians are generated, and the target pedestrians passing through the target shelf are determined based on the motion trajectories of the multiple pedestrians. Detect whether the target pedestrian has taken the unpaid goods; The method further includes the step of constructing the feature extractor: Construct a convolutional layer shared by M layers; Determine the first number corresponding to the global feature branch, the second number corresponding to the horizontal local feature branch, and the third number corresponding to the vertical local feature branch, respectively; Determine the horizontal local information corresponding to the second number of horizontal local feature branches, and determine the vertical local information corresponding to the third number of vertical local feature branches; Using the Mth convolutional layer as the root node, and in a tree-like branching manner, based on the first quantity, the second quantity, the third quantity, horizontal local information, and vertical local information, construct an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches. After the N convolutional layers, a convolutional computation module and a pooling computation module are constructed sequentially to obtain the feature extractor.
8. An alarm method, characterized in that, The method includes: Acquire the target image; A feature extractor is used to extract human body feature information corresponding to pedestrians in the target image, wherein the feature extractor has a tree-like branching structure; The human body feature information corresponding to the pedestrian in the target image is compared with the human body feature information corresponding to the target pedestrian in the target pedestrian image in the preset blacklist; An alarm is triggered when a pedestrian is detected in the target image. The method further includes the step of constructing the feature extractor: Construct a convolutional layer shared by M layers; Determine the first number corresponding to the global feature branch, the second number corresponding to the horizontal local feature branch, and the third number corresponding to the vertical local feature branch, respectively; Determine the horizontal local information corresponding to the second number of horizontal local feature branches, and determine the vertical local information corresponding to the third number of vertical local feature branches; Using the Mth convolutional layer as the root node, and in a tree-like branching manner, based on the first quantity, the second quantity, the third quantity, horizontal local information, and vertical local information, construct an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches. After the N convolutional layers, a convolutional computation module and a pooling computation module are constructed sequentially to obtain the feature extractor.
9. A detection method, characterized in that, The method includes: When a preset product is determined to be lost, target video data is acquired, the target video data including multiple frames of images to be detected; A feature extractor is used to extract feature information corresponding to the goods in the multiple frames of images to be detected, wherein the feature extractor has a tree-like branching structure; Based on the feature information corresponding to the goods in the multi-frame images to be detected, a target image containing the preset goods is determined in the multi-frame images to be detected; Detect the target pedestrian taking the preset product from the target image; The method further includes the step of constructing the feature extractor: Construct a convolutional layer shared by M layers; Determine the first number corresponding to the global feature branch, the second number corresponding to the horizontal local feature branch, and the third number corresponding to the vertical local feature branch, respectively; Determine the horizontal local information corresponding to the second number of horizontal local feature branches, and determine the vertical local information corresponding to the third number of vertical local feature branches; Using the Mth convolutional layer as the root node, and in a tree-like branching manner, based on the first quantity, the second quantity, the third quantity, horizontal local information, and vertical local information, construct an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches. After the N convolutional layers, a convolutional computation module and a pooling computation module are constructed sequentially to obtain the feature extractor.
10. An identification device, characterized in that, include: The first acquisition module is used to acquire the target image; The first feature extraction module is used to extract feature information corresponding to objects in the target image using a feature extractor, wherein the feature extractor has a tree-like branching structure; The first recognition module is used to identify whether the target image has a target object based on the feature information corresponding to the object in the target image. The apparatus further includes: a construction module for constructing the feature extractor; the construction module includes: A convolutional layer construction submodule is used to construct a convolutional layer shared by M layers; it determines the first number of global feature branches, the second number of horizontal local feature branches, and the third number of vertical local feature branches; it determines the horizontal local information corresponding to the second number of horizontal local feature branches, and the vertical local information corresponding to the third number of vertical local feature branches; using the convolutional layer shared by the M layers as the root node, it constructs an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, based on the first number, the second number, the third number, the horizontal local information, and the vertical local information. The computation module construction submodule is used to sequentially construct a convolution computation module and a pooling computation module after the N convolutional layers to obtain the feature extractor.
11. The apparatus according to claim 10, characterized in that, The first recognition module is used to determine the distance between the feature information corresponding to the object in the target image and the feature information corresponding to the target object, and to determine whether the target image has a target object based on the distance.
12. The apparatus according to claim 10, characterized in that, The first feature extraction module is used to input the target image into the feature extractor to obtain global feature information, horizontal local feature information and vertical local feature information corresponding to the target image; and to generate feature information corresponding to the object in the target image using the global feature information, horizontal local feature information and vertical local feature information corresponding to the target image.
13. The apparatus according to claim 10, characterized in that, In the N convolutional layers, the number of convolutional layers in the next layer is greater than or equal to the number of convolutional layers in the previous layer.
14. A people flow counting device, characterized in that, The device includes: The second acquisition module is used to acquire target video data within a set time period, wherein the target video data includes multiple frames of images. The second feature extraction module is used to extract human body feature information corresponding to pedestrians in the multiple frames of images using a feature extractor, wherein the feature extractor has a tree-like branch structure; The second recognition module is used to identify pedestrians based on the human body feature information corresponding to the pedestrians in the multiple frames of images, determine images with the same pedestrians and obtain multiple image groups, wherein the image group includes multiple images and different pedestrians correspond to different image groups; The pedestrian flow statistics module is used to count the pedestrian flow within a set time period based on the number of images in the image group and the number of images outside the image group; The apparatus further includes: a construction module for constructing the feature extractor; the construction module includes: A convolutional layer construction submodule is used to construct a convolutional layer shared by M layers; it determines the first number of global feature branches, the second number of horizontal local feature branches, and the third number of vertical local feature branches; it determines the horizontal local information corresponding to the second number of horizontal local feature branches, and the vertical local information corresponding to the third number of vertical local feature branches; using the convolutional layer shared by the M layers as the root node, it constructs an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, based on the first number, the second number, the third number, the horizontal local information, and the vertical local information. The computation module construction submodule is used to sequentially construct a convolution computation module and a pooling computation module after the N convolutional layers to obtain the feature extractor.
15. A trajectory tracking device, characterized in that, The device includes: The third acquisition module is used to acquire target video data and target pedestrian images, wherein the target video data includes multiple frames of images to be detected. The third feature extraction module is used to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected, and to extract human body feature information of the target pedestrian in the target pedestrian image, wherein the feature extractor has a tree-like branch structure; The selection module is used to compare the human body feature information corresponding to pedestrians in the multi-frame images to be detected with the human body feature information of the target pedestrian in the target pedestrian image, and select the target image containing the target pedestrian from the multi-frame images to be detected; The trajectory generation module is used to generate the movement trajectory of the target pedestrian based on the target image containing the target pedestrian; The apparatus further includes: a construction module for constructing the feature extractor; the construction module includes: A convolutional layer construction submodule is used to construct a convolutional layer shared by M layers; it determines the first number of global feature branches, the second number of horizontal local feature branches, and the third number of vertical local feature branches; it determines the horizontal local information corresponding to the second number of horizontal local feature branches, and the vertical local information corresponding to the third number of vertical local feature branches; using the convolutional layer shared by the M layers as the root node, it constructs an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, based on the first number, the second number, the third number, the horizontal local information, and the vertical local information. The computation module construction submodule is used to sequentially construct a convolution computation module and a pooling computation module after the N convolutional layers to obtain the feature extractor.
16. A detection device, characterized in that, The device includes: The shelf determination module is used to determine the target shelf where the unpaid goods are located when they are lost. The fourth acquisition module is used to acquire target video data, which includes multiple frames of images to be detected; The fourth feature extraction module is used to extract human body feature information corresponding to pedestrians in the multiple frames of images to be detected using a feature extractor, wherein the feature extractor has a tree-like branch structure; The identification grouping module is used to determine images with the same pedestrian and obtain multiple image groups based on the human body feature information corresponding to pedestrians in the multiple frames of images to be detected. Each image group includes multiple images, and different image groups correspond to different pedestrians. The first target pedestrian determination module is used to generate motion trajectories of multiple pedestrians based on the multiple image groups, and to determine the target pedestrians passing through the target shelf based on the motion trajectories of the multiple pedestrians. The detection module is used to detect whether the target pedestrian has taken the unpaid goods; The apparatus further includes: a construction module for constructing the feature extractor; the construction module includes: A convolutional layer construction submodule is used to construct a convolutional layer shared by M layers; it determines the first number of global feature branches, the second number of horizontal local feature branches, and the third number of vertical local feature branches; it determines the horizontal local information corresponding to the second number of horizontal local feature branches, and the vertical local information corresponding to the third number of vertical local feature branches; using the convolutional layer shared by the M layers as the root node, it constructs an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, based on the first number, the second number, the third number, the horizontal local information, and the vertical local information. The computation module construction submodule is used to sequentially construct a convolution computation module and a pooling computation module after the N convolutional layers to obtain the feature extractor.
17. An alarm device, characterized in that, The device includes: The fifth acquisition module is used to acquire the target image; The fifth feature extraction module is used to extract human body feature information corresponding to pedestrians in the target image using a feature extractor, wherein the feature extractor has a tree-like branching structure; The comparison module is used to compare the human body feature information corresponding to the pedestrian in the target image with the human body feature information corresponding to the target pedestrian in the target pedestrian image in the preset blacklist. The alarm module is used to trigger an alarm when it determines that a target pedestrian exists in the target image. The apparatus further includes: a construction module for constructing the feature extractor; the construction module includes: A convolutional layer construction submodule is used to construct a convolutional layer shared by M layers; it determines the first number of global feature branches, the second number of horizontal local feature branches, and the third number of vertical local feature branches; it determines the horizontal local information corresponding to the second number of horizontal local feature branches, and the vertical local information corresponding to the third number of vertical local feature branches; using the convolutional layer shared by the M layers as the root node, it constructs an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, based on the first number, the second number, the third number, the horizontal local information, and the vertical local information. The computation module construction submodule is used to sequentially construct a convolution computation module and a pooling computation module after the N convolutional layers to obtain the feature extractor.
18. A detection device, characterized in that, The device includes: The sixth acquisition module is used to acquire target video data when a preset product is lost, the target video data including multiple frames of images to be detected; The sixth feature extraction module is used to extract feature information corresponding to the goods in the multiple frames of images to be detected using a feature extractor, wherein the feature extractor has a tree-like branch structure; The image determination module is used to determine the target image containing the preset product in the multi-frame images to be detected based on the feature information corresponding to the product in the multi-frame images to be detected; The second target pedestrian determination module is used to detect the target pedestrian taking the preset product from the target image; The apparatus further includes: a construction module for constructing the feature extractor; the construction module includes: A convolutional layer construction submodule is used to construct a convolutional layer shared by M layers; it determines the first number of global feature branches, the second number of horizontal local feature branches, and the third number of vertical local feature branches; it determines the horizontal local information corresponding to the second number of horizontal local feature branches, and the vertical local information corresponding to the third number of vertical local feature branches; using the convolutional layer shared by the M layers as the root node, it constructs an N-layer convolutional layer containing global feature branches, horizontal local feature branches, and vertical local feature branches in a tree-like branching manner, based on the first number, the second number, the third number, the horizontal local information, and the vertical local information. The computation module construction submodule is used to sequentially construct a convolution computation module and a pooling computation module after the N convolutional layers to obtain the feature extractor.
19. An electronic device, characterized in that, include: processor; and A memory having executable code stored thereon, which, when executed, causes the processor to perform the method as described in one or more of claims 1-9.
20. One or more machine-readable media having executable code stored thereon, which, when executed, causes a processor to perform the method as described in one or more of claims 1-9.
Citation Information
Patent Citations
Monitoring method and device, server and computer readable storage medium
CN111405249A