Cross-shot target object extraction method, device, electronic device and storage medium

By establishing a positional relationship mapping model in multi-monitoring node scenarios and using deep neural network training, efficient cross-lens target object positioning and collection are achieved, solving the problems of large amount of computing and low efficiency in the existing technology, and reducing the computing power requirements of equipment.

CN114267018BActive Publication Date: 2025-07-08CHONGQING UNISINSIGHT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111633193.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-08
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

In the multi-monitoring node scenario, when finding the same target object through human feature comparison method, the calculation amount is large and the efficiency is low, making it difficult to efficiently and quickly locate and collect suspicious target data across the lens.

Method used

Establish a positional relationship mapping model, use the position information of the first monitoring node to map the location of the target object in other monitoring nodes, reduce the calculation and comparison of human body or face features, and use deep neural network to train the relationship mapping model to improve the accuracy of position information.

Benefits of technology

It reduces the computing power demand of equipment, improves the positioning and collection efficiency of target objects in multi-monitoring node scenarios, and reduces the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114267018B_ABST
    Figure CN114267018B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, electronic device, and storage medium for extracting target objects across lenses, which is used to improve the efficiency of extracting target object images across lenses. In an embodiment of the present application, first, first position information of a target object in a first image is obtained; then the first position information is input into a position relationship mapping model, and second position information of the target object in second images collected by each second monitoring node output by the position relationship mapping model is obtained; finally, for any second monitoring node, an image of the target object is extracted based on the second position information of the target object in the second image collected by the second monitoring node. In the present application, according to the position information and relationship mapping model in the first monitoring node, the positions of the target object in each second monitoring node are obtained, without the need for human body or face feature calculation, extraction, and comparison, greatly reducing the computing power and thus improving the efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] At the entrances of subways, airports, large-scale exhibitions, etc., multiple monitoring nodes are often used to monitor the same scene. Once something abnormal occurs, the monitoring personnel need to collect and sort out the information of the target object by browsing multiple monitoring cameras. Obviously, manual sorting has low processing efficiency.

[0003] In related technologies, a method for automatically finding the same target object in different monitoring nodes of the same scene is also provided. This method is mainly implemented by comparing human body features. That is: through the feature detection and comparison algorithm, all pedestrian targets that appear in a certain period are detected, and then the human body features of each pedestrian are extracted through the human body feature comparison algorithm. The same target object is found by the similarity of the human body features, and then the same target object in different monitoring nodes is associated. However, this method requires parsing multiple videos and constantly calculating, extracting, and comparing the features of pedestrians, which requires high computing power. Summary of the Invention

[0004] The purpose of this application is to provide a method, device, electronic device, and storage medium for extracting target objects across lenses, which is used to reduce the computing power for extracting target object images across lenses.

[0005] In a first aspect, an embodiment of this application provides a method for extracting target objects across lenses. The first monitoring node and at least one second monitoring node have a common shooting area. The method includes:

[0006] Obtain the first position information of the target object in the first image, where the first image is an image collected by the first monitoring node;

[0007] Input the first position information into the position relationship mapping model to obtain the second position information of the target object in the second images collected by each of the second monitoring nodes output by the position relationship mapping model;

[0008] For any of the second monitoring nodes, based on the second position information of the target object in the second image collected by the second monitoring node, extract the image of the target object.

[0009] In this application, according to the position information and relationship mapping model in the first monitoring node, the positions of the target object in each second monitoring node are obtained, without the need to calculate, extract, and compare human body or face features, greatly reducing the computing power.

[0010] In some possible embodiments, for any second monitoring node among the at least one second monitoring node, execute respectively:

[0011] Construct a first training sample, where the first training sample includes a first sample position of a sample object in a first sample frame collected by the first monitoring node, and second sample positions of the sample object in second sample frames collected by each of the second monitoring nodes, and the collection times of the first sample frame and each of the second sample frames are the same;

[0012] Input the first sample position into the relationship mapping model, and train the relationship mapping model with the second sample positions of each of the second sample frames as the expected outputs.

[0013] In this application, by training the relationship mapping model, the position information of the same target object in the second monitoring node is obtained according to the position information in the first monitoring node, which greatly saves the calculation cost.

[0014] In some possible embodiments, the inputting the first sample position into the relationship mapping model and training the relationship mapping model with the second sample positions of each of the second sample frames as the expected outputs includes:

[0015] Input the first sample position into the relationship mapping model, and use the second sample positions of each of the second sample frames as the expected outputs to obtain the generated position information respectively output by the relationship mapping model for each second sample position;

[0016] Determine a loss value according to the generated position information and the second sample position;

[0017] Adjust the parameters of the relationship mapping model based on the loss value.

[0018] In some possible embodiments, the method further includes:

[0019] Obtain the generated position information respectively output by the relationship mapping model for each second sample position based on the first sample position, and construct a second training sample, where the second training sample includes the first sample position and each of the generated position information;

[0020] Determine the intersection over union of each of the generated position information and the corresponding second sample position;

[0021] If the intersection over union of each of the generated position information and the corresponding second sample position is greater than or equal to a first preset value, then construct a third training sample using the generated position information and the first sample position;

[0022] Train the relationship mapping model using the third training sample.

[0023] In this application, by constructing a third training sample set to train the relationship mapping model, the training of the relationship mapping model is made more comprehensive, improving the accuracy of the relationship mapping model.

[0024] In some possible embodiments, after determining the intersection over union (IoU) of each of the generated position information and the corresponding second sample position, the method further includes:

[0025] If the IoU corresponding to any generated position information is less than a second preset value, then construct a fourth training sample using each of the second sample positions and the first sample position;

[0026] Merge the third training sample and the fourth training sample into a fifth training sample;

[0027] Train the relationship mapping model using the fifth training sample.

[0028] In this application, by constructing a fifth training sample set to train the relationship mapping model, the accuracy of the relationship mapping model is further improved.

[0029] In some possible embodiments, obtaining the first position information of the target object in the first image includes:

[0030] Perform target detection on the target object in the first image to obtain the position information of the target object in the first image, where the position information includes the vertex coordinates of the rectangular box and the length information and width information of the rectangular box.

[0031] In this application, the method of obtaining the rectangular box of the target object can accurately lock the target object.

[0032] In some possible embodiments, the performing target detection on the target object in the first image includes:

[0033] Perform target detection on the target object in the first image using a face detection method or a human body detection method.

[0034] In a second aspect, this application also provides a cross-shot target object extraction device. The first monitoring node and at least one second monitoring node have a common shooting area. The device includes:

[0035] A first position information acquisition module, configured to acquire the first position information of the target object in the first image, where the first image is an image collected by the first monitoring node;

[0036] A second position information acquisition module, configured to input the first position information into a position relationship mapping model, and obtain second position information of the target object in second images collected by each of the second monitoring nodes output by the position relationship mapping model;

[0037] An extraction module, configured to, for any of the second monitoring nodes, extract an image of the target object based on the second position information of the target object in the second image collected by the second monitoring node.

[0038] In some possible embodiments, the second position information acquisition module separately executes, for any second monitoring node among the at least one second monitoring node:

[0039] Construct a first training sample, where the first training sample includes a first sample position of a sample object in a first sample frame collected by the first monitoring node, and second sample positions of the sample object in second sample frames collected by each of the second monitoring nodes, and the collection times of the first sample frame and each of the second sample frames are the same;

[0040] Input the first sample position into the relationship mapping model, and use the second sample positions of each of the second sample frames as expected outputs to train the relationship mapping model.

[0041] In some possible embodiments, when the second position information acquisition module executes inputting the first sample position into the relationship mapping model and using the second sample positions of each of the second sample frames as expected outputs to train the relationship mapping model, it is configured to:

[0042] Input the first sample position into the relationship mapping model, use the second sample positions of each of the second sample frames as expected outputs, and obtain generated position information respectively corresponding to each second sample position output by the relationship mapping model;

[0043] Determine a loss value according to the generated position information and the second sample position;

[0044] Adjust parameters of the relationship mapping model based on the loss value.

[0045] In some possible embodiments, the apparatus further includes:

[0046] A second training sample construction module, configured to obtain the generated position information respectively corresponding to each second sample position output by the relationship mapping model based on the first sample position, and construct a second training sample, where the second training sample includes the first sample position and each of the generated position information;

[0047] An intersection over union determination module, configured to determine the intersection over union of each of the generated position information and the corresponding second sample position;

[0048] A third training sample construction module, configured to construct a third training sample using the generated position information and the first sample position if the intersection over union of each of the generated position information and the corresponding second sample position is greater than or equal to a first preset value;

[0049] A training module, configured to train the relationship mapping model using the third training sample.

[0050] In some possible embodiments, after the intersection over union determination module determines the intersection over union of each of the generated position information and the corresponding second sample position, it is further configured to:

[0051] If the intersection over union corresponding to any generated position information is less than a second preset value, construct a fourth training sample using each of the second sample positions and the first sample position;

[0052] Combine the third training sample and the fourth training sample into a fifth training sample;

[0053] Train the relationship mapping model using the fifth training sample.

[0054] In some possible embodiments, when the first position information acquisition module acquires the first position information of the target object in the first image, it is configured to:

[0055] Perform target detection on the target object in the first image to obtain the position information of the target object in the first image, where the position information includes the vertex coordinates of the rectangular frame and the length information and width information of the rectangular frame.

[0056] In some possible embodiments, when the first position information acquisition module performs target detection on the target object in the first image, it is configured to: perform target detection on the target object in the first image using a face detection method or a human body detection method.

[0057] In a third aspect, another embodiment of the present application further provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any method provided in the first aspect embodiment of the present application.

[0058] Fourthly, another embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is used to cause a computer to execute any method provided in the embodiment of the first aspect of the present application.

[0059] Other features and advantages of the present application will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. Obviously, the following introduced drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0061] Figure 1A It is an application scenario diagram of a method for extracting a target object across cameras provided in an embodiment of the present application;

[0062] Figure 1B It is a schematic diagram of a common area of a method for extracting a target object across cameras provided in an embodiment of the present application;

[0063] Figure 2A It is an overall flowchart of a method for extracting a target object across cameras provided in an embodiment of the present application;

[0064] Figure 2B It is a schematic diagram of obtaining a target object in a first image of a method for extracting a target object across cameras provided in an embodiment of the present application;

[0065] Figure 3 It is a flowchart of training a relationship mapping model of a method for extracting a target object across cameras provided in an embodiment of the present application;

[0066] Figure 4 It is a flowchart of constructing a first training sample set of a method for extracting a target object across cameras provided in an embodiment of the present application;

[0067] Figure 5A It is a schematic diagram of a common area of a method for extracting a target object across cameras provided in an embodiment of the present application;

[0068] Figure 5B It is a schematic diagram of determining the position of a sample object in a first sample frame of a monitoring node of a method for extracting a target object across cameras provided in an embodiment of the present application;

[0069] Figure 5C Schematic diagram of a relationship mapping model for a method for extracting target objects across camera shots provided by an embodiment of the present application;

[0070] Figure 6A Schematic diagram of input and output of a relationship mapping model for a method for extracting target objects across camera shots provided by an embodiment of the present application;

[0071] Figure 6B Internal schematic diagram of a relationship mapping model for a method for extracting target objects across camera shots provided by an embodiment of the present application;

[0072] Figure 7 Schematic diagram of a training relationship mapping model for a method for extracting target objects across camera shots provided by an embodiment of the present application;

[0073] Figure 8 Schematic diagram of a training relationship mapping model for a method for extracting target objects across camera shots provided by an embodiment of the present application;

[0074] Figure 9 Schematic diagram of a device for a method for extracting target objects across camera shots provided by an embodiment of the present application;

[0075] Figure 10 Schematic diagram of an electronic device for a method for extracting target objects across camera shots provided by an embodiment of the present application. Detailed implementation manners

[0076] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0077] It should be noted that the terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0078] The entrances of subways, airports, large exhibitions, etc. often use multiple monitoring nodes to monitor the same scene. Once an abnormal situation occurs, the monitoring personnel need to collect and organize the information of the target object by browsing multiple monitoring cameras. Manual sorting is obviously inefficient. The relevant technology also provides a method for automatically finding the same target object in different monitoring nodes in the same scene. This method is mainly based on human feature comparison and association, that is, through the feature detection and comparison algorithm, all pedestrian targets that appear in a certain period of time are detected, and then the human features of each pedestrian are extracted through the human feature comparison algorithm. The same target object is found through the similarity of human features, and then the same target object in different monitoring nodes is associated. However, this method requires parsing multiple videos, and the human or facial features of pedestrians are calculated, extracted, and compared all the time, which requires high computing power. In addition, in some scenarios, there will be multiple cameras for collaborative monitoring, so it is particularly important to associate the same target in the same time and space, solve the problem of efficient and rapid positioning, collect suspicious target data across lenses, and greatly shorten the search for the original video.

[0079] In view of this, the present application proposes a cross-lens target object extraction method, device, electronic device and storage medium to solve the above problems. The inventive concept of the present application can be summarized as follows: for a common shooting area with multiple monitoring nodes, a position relationship mapping model can be established, which is used to obtain the position information of the target object in the images of other monitoring nodes in the common shooting area based on the position information of the target object in the image of the monitoring node after the target object is collected by any monitoring node.

[0080] In this way, the position of the same target object in the images of different monitoring nodes in the common shooting area can be clearly identified, and the target object can be extracted based on the position information without searching and comparing the same target object in the entire area of ​​each monitoring node image. Therefore, compared with the low complexity of the method of performing target detection and human feature comparison from the overall image of each monitoring node in the related art, the solution provided by the embodiment of the present application can reduce the requirements for device computing power.

[0081] For ease of understanding, a cross-lens target object extraction method provided by an embodiment of the present application is described in detail below with reference to the accompanying drawings:

[0082] Take the example of four monitoring nodes with a common shooting area, such as Figure 1AAs shown in the figure, it is an application scenario diagram of a method for extracting target objects across cameras in an embodiment of the present application. The figure includes: a network 10, a server 20, a memory 30, a first monitoring node 40, and a second monitoring node set 50. The second monitoring node set 50 includes a second monitoring node 501, a second monitoring node 502, and a second monitoring node 503. The first monitoring node 40 and the second monitoring nodes 501, 502, and 503 have a common shooting area. Among them, the common shooting area is as Figure 1B shown, where:

[0083] The server 20 first obtains the first position information of the target object in the first image, where the first image is the image collected by the first monitoring node 40. Then, the first position information is input into the position relationship mapping model to obtain the second position information of the target object in the second images collected by the second monitoring nodes 501, 502, and 503 respectively output by the position relationship mapping model. Among them, the relationship mapping model is installed in the memory 30. For any second monitoring node, based on the second position information of the target object in the second image collected by the second monitoring node, the image of the target object is extracted.

[0084] In the description of the present application, only a single server or monitoring node is described in detail. However, those skilled in the art should understand that the shown first monitoring node 40, second monitoring nodes 501, 502, 503, server 20, and memory 30 are intended to represent the operations of the terminal devices, servers, and memories involved in the technical solution of the present application. The detailed description of a single server and memory is for convenience of illustration only, and does not imply any limitation on the number, type, or location of the monitoring nodes and servers. It should be noted that if additional modules are added to or individual modules are removed from the illustrated environment, the underlying concept of the exemplary embodiments of the present application will not be changed. Additionally, although a two-way arrow from the memory 30 to the server 20 is shown in Figure 1A for convenience of illustration, those skilled in the art can understand that the above data transmission and reception also need to be implemented through the network 10.

[0085] It should be noted that the memory in the embodiments of the present application can be, for example, a cache system, a hard disk storage, a memory storage, etc. In addition, the method for extracting target objects across cameras proposed in the present application is not only applicable to Figure 1A the application scenario shown, but also applicable to any device with a need to extract target objects across cameras.

[0086] As Figure 2A shown, it is a flow diagram of a method for extracting target objects across cameras provided by an embodiment of the present application, where:

[0087] In step 201: Obtain the first position information of the target object in the first image, where the first image is the image collected by the first monitoring node;

[0088] In some embodiments, when obtaining the first position information of the target object in the first image, it can be specifically implemented as: performing target detection on the target object in the first image to obtain the position information of the target object in the first image, where the position information includes the vertex coordinates of the rectangular box, the length information, and the width information of the rectangular box. For example: as Figure 2B shown, perform target detection on the target object in the first image, and obtain that the length information of the rectangular box of the target object is h, the width information is w, and the vertex coordinates are (x, y). Therefore, the position information of the target object in the first image can be obtained as (x, y, w, h).

[0089] Among them, the target detection method can be implemented by comparing with known features of the target object such as face detection, human body detection, etc. to detect the first position information of the target object in the first monitoring node.

[0090] In step 202: Input the first position information into the position relationship mapping model to obtain the second position information of the target object in each second image collected by the second monitoring nodes output by the position relationship mapping model;

[0091] In step 203: For any second monitoring node, extract the image of the target object based on the second position information of the target object in the second image collected by the second monitoring node.

[0092] For ease of understanding, the steps in Figure 2A will be described in detail below:

[0093] In order to realize cross-shot extraction of the image information of the target object, a relationship mapping model is trained in this application, and the training steps of the relationship mapping model are as Figure 3 shown:

[0094] In step 301: Construct the first training sample set;

[0095] In some embodiments, when constructing the first training sample set, it can be specifically implemented as the steps as Figure 4 shown:

[0096] In step 401: Obtain the first images containing the sample object collected by the first monitoring node to form the first sample frame; and, obtain the first sample position of the sample object in the first sample frame;

[0097] In the embodiments of this application, for example: as Figure 5AAs shown, monitoring nodes A, B, C, and D cover the same common area. Taking monitoring node A as the first monitoring node and monitoring nodes B, C, and D as the second monitoring nodes, images collected by monitoring nodes A, B, C, and D at the same moment during the process of the sample object walking in the common area are collected. Assuming that monitoring nodes A, B, C, and D have all collected images of the sample object at moments 1, 2, 3, 4, and 5, the images collected by monitoring node A at moments 1, 2, 3, 4, and 5 respectively constitute the first sample frames. As Figure 5B shown, then the position of the sample object in each first sample frame of monitoring node A is determined. The position of the sample object in each sample frame is the first sample position corresponding to that sample frame.

[0098] In step 402: For any second monitoring node, obtain the second image corresponding to the first image collected by the second monitoring node and containing the sample object to form a second sample frame; and obtain the second sample position of the sample object in the second sample frame;

[0099] Continuing with Figure 5A as an example, the images collected by monitoring nodes B, C, and D respectively constitute their own second sample frames. Taking monitoring node B as an example for illustration, assuming Figure 5B are the images of the sample object collected by monitoring node B at moments 1, 2, 3, 4, and 5, the image corresponding to moment 1 constitutes a second sample frame, the image corresponding to moment 2 constitutes a second sample frame, the image corresponding to moment 3 constitutes a second sample frame, the image corresponding to moment 4 constitutes a second sample frame, and the image corresponding to moment 5 constitutes a second sample frame; then the position of the sample object in each second sample frame of monitoring node B is determined, and the position of the sample object in each second sample frame is the second sample position corresponding to that sample frame.

[0100] In step 403: Construct a first training sample set according to the first sample position and the second sample positions of each second sample frame.

[0101] In summary, the first training sample includes the first sample position of the sample object in the first sample frame collected by the first monitoring node and the second sample positions of the sample object in the second sample frames collected by each second monitoring node, and the acquisition times of the first sample frame and each second sample frame are the same. For example: Taking Figure 5A and Figure 5B as an example, the first training sample includes the position of the sample object in the image collected by monitoring node A at moment 1, the position of the sample object in the image collected by monitoring node B at moment 1, the position of the sample object in the image collected by monitoring node C at moment 1, and the position of the sample object in the image collected by monitoring node D at moment 1.

[0102] In step 302: Input the first sample position into the relationship mapping model, and train the relationship mapping model with the second sample positions of each second sample frame as the expected output.

[0103] Taking the example that four monitoring nodes have a common shooting area, the relationship mapping model is as Figure 5C shown. This relationship mapping model has one input and three outputs, that is, input the position of the target object in one monitoring node, and the positions of the target object in the other three monitoring nodes can be output. In implementation, this relationship mapping model can be a Deep Neural Networks (DNN) model.

[0104] As Figure 6A shown, use the first sample position of the sample object collected by monitoring node A and the second sample positions of the sample objects collected by monitoring nodes B, C, and D to train the relationship mapping model, and the expected output is the second sample positions of the sample objects collected by monitoring nodes B, C, and D. Further, in order to enable the relationship mapping model to map the position of the target object in any monitoring node to the positions of the target object in other monitoring nodes when in use, therefore, after training the relationship mapping model with the first sample position of the sample object collected by monitoring node A and the second sample positions of the sample objects collected by monitoring nodes B, C, and D, monitoring nodes B, C, and D can be successively used as the first monitoring node, and the other monitoring nodes can be used as the second monitoring node to continue training this relationship mapping model. For example: First, use monitoring node A as the first monitoring node, and monitoring nodes B, C, and D as the second monitoring node, and use the first sample position of the sample object collected by monitoring node A and the second sample positions of the sample objects collected by monitoring nodes B, C, and D to train the relationship mapping model; then use monitoring node B as the first monitoring node, and monitoring nodes A, C, and D as the second monitoring node, and use the first sample position of the sample object collected by monitoring node B and the second sample positions of the sample objects collected by monitoring nodes A, C, and D to train the relationship mapping model.

[0105] Finally, the internal schematic diagram of the relationship mapping model after training convergence is as Figure 6B shown. There are 4 monitoring nodes in this scenario, and each monitoring node can be used as the first monitoring node. Training convergence is, for example, that the number of iterative training reaches a specified number, or for another example, the position accuracy output by the position relationship mapping model reaches a specified accuracy.

[0106] In the embodiments of the present application, there will inevitably be some errors in the data generated by the relationship mapping model. In order to further improve the accuracy of the relationship mapping model, the relationship mapping model is further trained in the present application, and the specific implementation can be the steps as Figure 7 shown:

[0107] In step 701: Obtain the generated position information for each second sample position output by the relationship mapping model based on the first sample position, and construct a second training sample, where the second training sample includes the first sample position and each generated position information;

[0108] In step 702: Determine the intersection over union (IoU) between each generated position information and the corresponding second sample position;

[0109] In some embodiments, the IoU can be determined using formula 1:

[0110]

[0111] where P b ' is the generated position information, P b is the second sample position, and IOU is the intersection over union.

[0112] In step 703: If the IoU between each generated position information and the corresponding second sample position is greater than or equal to a first preset value, then construct a third training sample using the generated position information and the first sample position;

[0113] If the IoU is greater than or equal to the first preset value, it indicates that the generated position information generated this time is relatively accurate, so it can be regarded as a high-quality sample. Therefore, in this application, this type of generated position information is used as the third training sample.

[0114] In step 704: Train the relationship mapping model using the third training sample.

[0115] For example: Monitoring nodes A, B, C, and D cover the same common area. Taking monitoring node A as the first monitoring node and monitoring nodes B, C, and D as the second monitoring nodes, the first training sample includes: the second sample positions of the sample object at times 1, 2, 3, 4, and 5 in the second sample frames of monitoring nodes B, C, and D, and the first sample position of the sample object at times 1, 2, 3, 4, and 5 in the first sample frame of monitoring node A. If the IoU between the second sample position at time 2 in the second sample frames of monitoring nodes B, C, and D and the corresponding generated position information is greater than the first preset value, then determine the third training sample as: the position of the sample object at time 2 in the second sample frames of monitoring nodes B, C, and D, and the position of the sample object at time 2 in the first sample frame of monitoring node A.

[0116] In some embodiments, when collecting the coordinate position data of a sample object at the same moment in different monitoring nodes, there are differences in the heights of pedestrians in the monitoring scenario. Therefore, when collecting sample data, it is often impossible to cover the entire height of the sample object, resulting in insufficient richness of the collected samples, and thus errors in the trained model. Therefore, in this application, in order to further improve the accuracy of the relationship mapping model, after determining the intersection over union (IoU) between each generated position information and the corresponding second sample position, the steps as shown in Figure 8 can be implemented:

[0117] In step 801: If the IoU corresponding to any generated position information is less than a second preset value, then use each of the second sample positions and the first sample position to construct a fourth training sample;

[0118] In specific implementation, if the IoU is less than the second preset value, it means that the current mapping is an incorrect mapping. Therefore, it is necessary to collect the second sample position and the first sample position corresponding to this generated position information, and use the second sample position and the first sample position to continue training the relationship mapping model. However, since the samples corresponding to the incorrect mapping are few, it is necessary to continue to execute 802 to expand the samples.

[0119] In step 802: Merge the third training sample and the fourth training sample into a fifth training sample;

[0120] In step 803: Use the fifth training sample to train the relationship mapping model.

[0121] For example: Monitoring nodes A, B, C, and D cover the same common area. Taking monitoring node A as the first monitoring node and monitoring nodes B, C, and D as the second monitoring nodes, use the target detection algorithm to determine the third position information of the sample object at moments 1, 2, 3, 4, and 5 in the second sample frames of monitoring nodes B, C, and D. Then determine the generated position information of the sample object in the second training sample set at moments 1, 2, 3, 4, and 5 in the second sample frames of monitoring nodes B, C, and D. Determine the IoU between the third position information of monitoring node B at moment 1 and the generated position information corresponding to moment 1 of monitoring node B, and compare this IoU with the preset value. If it is less than the second preset value, it means that the generated position information generated by the relationship mapping model is inaccurate. Therefore, collect the second sample position and the first sample position corresponding to this generated position information.

[0122] In this application, the process of using each sample to train the relationship mapping model is the same. Therefore, for the sake of easy understanding, the following takes the process of inputting the first sample position into the relationship mapping model and training the relationship mapping model with the second sample positions of each second sample frame as the expected output as an example for detailed description:

[0123] First, input the first sample position into the relationship mapping model, use the second sample positions of each second sample frame as the expected outputs, and obtain the generated position information for each second sample position output by the relationship mapping model; then determine the loss value based on the generated position information and the second sample positions; finally, adjust the parameters of the relationship mapping model based on the loss value.

[0124] In summary, according to the position information in the first monitoring node and the relationship mapping model in this application, the positions of the target object in each second monitoring node are obtained without performing human body or face feature calculation, extraction, and comparison, greatly reducing the computing power and thus improving the efficiency.

[0125] As Figure 9 shown, based on the same inventive concept, a cross-shot target object extraction device 900 is proposed, including:

[0126] A first position information acquisition module 9001, configured to acquire the first position information of the target object in the first image, where the first image is an image collected by the first monitoring node;

[0127] A second position information acquisition module 9002, configured to input the first position information into the position relationship mapping model to obtain the second position information of the target object in each second image collected by each second monitoring node output by the position relationship mapping model;

[0128] An extraction module 9003, configured to extract the image of the target object for any second monitoring node based on the second position information of the target object in the second image collected by the second monitoring node.

[0129] In some possible embodiments, for any second monitoring node among the at least one second monitoring node, respectively execute:

[0130] Construct a first training sample, where the first training sample includes the first sample position of the sample object in the first sample frame collected by the first monitoring node and the second sample positions of the sample object in the second sample frames collected by each second monitoring node, and the acquisition times of the first sample frame and each second sample frame are the same;

[0131] Input the first sample position into the relationship mapping model, and train the relationship mapping model with the second sample positions of each second sample frame as the expected outputs.

[0132] In some possible embodiments, when the second position information acquisition module executes inputting the first sample position into the relationship mapping model and training the relationship mapping model with the second sample positions of each second sample frame as the expected outputs, it is configured to:

[0133] Input the first sample position into the relationship mapping model, use the second sample positions of the second sample frames as the expected outputs, and obtain the generated position information output by the relationship mapping model for each second sample position respectively;

[0134] Determine a loss value according to the generated position information and the second sample position;

[0135] Adjust the parameters of the relationship mapping model based on the loss value.

[0136] In some possible embodiments, the apparatus further includes:

[0137] A second training sample construction module, configured to obtain the generated position information output by the relationship mapping model for each second sample position respectively based on the first sample position, and construct a second training sample, where the second training sample includes the first sample position and each generated position information;

[0138] An intersection over union (IoU) determination module, configured to determine the IoU of each generated position information and the corresponding second sample position;

[0139] A third training sample construction module, configured to, if the IoU of each generated position information and the corresponding second sample position is greater than or equal to a first preset value, construct a third training sample using the generated position information and the first sample position;

[0140] A training module, configured to train the relationship mapping model using the third training sample.

[0141] In some possible embodiments, after the IoU determination module executes to determine the IoU of each generated position information and the corresponding second sample position, it is further configured to:

[0142] If the IoU corresponding to any generated position information is less than a second preset value, construct a fourth training sample using each second sample position and the first sample position;

[0143] Combine the third training sample and the fourth training sample into a fifth training sample;

[0144] Train the relationship mapping model using the fifth training sample.

[0145] In some possible embodiments, when the first position information acquisition module executes to acquire the first position information of the target object in the first image, it is configured to:

[0146] Perform object detection on the target object in the first image to obtain the position information of the target object in the first image, where the position information includes the vertex coordinates of the rectangular box and the length information and width information of the rectangular box.

[0147] In some possible embodiments, when the first position information acquisition module performs object detection on the target object in the first image, it is configured to: perform object detection on the target object in the first image by using a face detection method or a human body detection method.

[0148] After introducing the cross-shot target object extraction method and device of the exemplary embodiments of the present application, next, an electronic device according to another exemplary embodiment of the present application is introduced.

[0149] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.

[0150] In some possible embodiments, the electronic device according to the present application may at least include at least one processor and at least one memory. Among them, the memory stores program code, and when the program code is executed by the processor, the processor executes the steps in the cross-shot target object extraction method according to various exemplary embodiments of the present application described above in this specification.

[0151] Next, refer to Figure 10 to describe the electronic device 130 according to this embodiment of the present application. Figure 10 The displayed electronic device 130 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present application.

[0152] As Figure 10 shown, the electronic device 130 is presented in the form of a general electronic device. The components of the electronic device 130 may include but are not limited to: the above-mentioned at least one processor 131, the above-mentioned at least one memory 132, and a bus 133 connecting different system components (including the memory 132 and the processor 131).

[0153] The bus 133 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a processor, or a local bus using any bus structure in a variety of bus structures.

[0154] The memory 132 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 1321 and / or cache memory 1322, and may further include read-only memory (ROM) 1323.

[0155] The memory 132 may further include a program / utilities 1325 having a set (at least one) of program modules 1324. Such program modules 1324 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0156] The electronic device 130 may also communicate with one or more external devices 134 (such as a keyboard, a pointing device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 130, and / or may communicate with any device that enables the electronic device 130 to communicate with one or more other electronic devices (such as a router, a modem, etc.). Such communication may be through an input / output (I / O) interface 135. Moreover, the electronic device 130 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 136. As shown in the figure, the network adapter 136 communicates with other modules for the electronic device 130 through a bus 133. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 130, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0157] In some possible implementation manners, various aspects of a method for extracting a target object across lenses provided in this application may also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in a method for extracting a target object across lenses according to various exemplary implementation manners of this application described above in this specification.

[0158] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0159] The program product for cross-shot target object extraction according to the embodiments of the present application may adopt a portable compact disk read-only memory (CD-ROM) and include program code, and may run on an electronic device. However, the program product of the present application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0160] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including - but not limited to - an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0161] The program code contained on the readable medium may be transmitted by any appropriate medium, including - but not limited to - wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0162] The program code for performing the operations of the present application can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's electronic device, partially on the user's device, executed as an independent software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In the case of a remote electronic device, the remote electronic device can be connected to the user's electronic device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external electronic device (e.g., by connecting through the Internet using an Internet service provider).

[0163] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described units can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0164] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that the operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0165] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0166] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in a flow or multiple flows in the flowchart and / or a block or multiple blocks in the block diagram.

[0167] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in a flow or multiple flows in the flowchart and / or a block or multiple blocks in the block diagram.

[0168] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in a flow or multiple flows in the flowchart and / or a block or multiple blocks in the block diagram.

[0169] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.

Claims

1. A method for extracting target objects across shots, characterized in that, The first monitoring node and at least one second monitoring node have a common monitoring area, and the method includes: Obtain first position information of a target object in a first image, where the first image is an image collected by the first monitoring node; Input the first position information into a position relationship mapping model to obtain second position information of the target object in second images collected by each of the second monitoring nodes output by the position relationship mapping model; wherein, the position relationship mapping model is trained based on a first training sample of a sample object in the first training sample collected by the first monitoring node and second training samples of the sample object in the second training samples collected by each of the second monitoring nodes; obtain generated position information of the position relationship mapping model for each second sample position based on a first sample position, and construct a second training sample including the first sample position and each generated position information; determine the intersection-over-union ratio of each generated position information and the corresponding second sample position, if each intersection-over-union ratio is greater than or equal to a first preset value, then construct a third training sample using the generated position information and the first sample position, and further train the position relationship mapping model based on the third training sample; For any of the second monitoring nodes, extract an image of the target object based on the second position information of the target object in the second image collected by the second monitoring node.

2. The method according to claim 1, wherein For each of the at least one second monitoring node, perform the following respectively: Construct a first training sample, which includes a first sample position of a sample object in a first sample frame collected by the first monitoring node and second sample positions of the sample object in second sample frames collected by each of the second monitoring nodes, and the acquisition times of the first sample frame and each second sample frame are the same; Input the first sample position into the position relationship mapping model, and train the position relationship mapping model with the second sample positions of each second sample frame as the expected output.

3. The method according to claim 2, wherein The step of inputting the first sample position into the position relationship mapping model and training the position relationship mapping model with the second sample positions of each second sample frame as the expected output includes: Input the first sample position into the position relationship mapping model, and with the second sample positions of each second sample frame as the expected output, obtain generated position information of the position relationship mapping model for each second sample position; Determine a loss value according to the generated position information and the second sample position; Adjust the parameters of the position relationship mapping model based on the loss value.

4. The method according to claim 1, characterized in that, After determining the intersection-over-union ratio of each generated position information and the corresponding second sample position, the method further includes: If the intersection-over-union ratio corresponding to any generated position information is less than a second preset value, then construct a fourth training sample using each second sample position and the first sample position, Combine the third training sample and the fourth training sample into a fifth training sample; Train the position relationship mapping model using the fifth training sample.

5. The method according to any one of claims 1-4, characterized in that Obtain the first position information of the target object in the first image, including: Perform target detection on the target object in the first image to obtain the position information of the target object in the first image, where the position information includes the vertex coordinates of the rectangular box and the length information and width information of the rectangular box.

6. The method according to claim 5, characterized in that, The performing target detection on the target object in the first image includes: Adopt a face detection method or a human body detection method to perform target detection on the target object in the first image.

7. An apparatus for extracting a target object across shots, characterized in that, The first monitoring node and at least one second monitoring node have a common monitoring area, and the device includes: A first position information acquisition module, configured to obtain the first position information of the target object in the first image, where the first image is an image collected by the first monitoring node; A second position information acquisition module, configured to input the first position information into a position relationship mapping model to obtain the second position information of the target object in the second images collected by each of the second monitoring nodes output by the position relationship mapping model; wherein, the position relationship mapping model is trained based on the first training samples collected by the first monitoring node of the sample object and the second training samples collected by the sample object at each of the second monitoring nodes; A second training sample construction module, configured to obtain the generated position information respectively for each second sample position output by the position relationship mapping model based on the first sample position, and construct a second training sample including the first sample position and each of the generated position information; An intersection over union determination module, configured to determine the intersection over union of each of the generated position information and the corresponding second sample position; A third training sample construction module, configured to, if each of the intersection over unions is greater than or equal to a first preset value, construct a third training sample by using the generated position information and the first sample position; A training module, configured to further train the position relationship mapping model based on the third training sample; An extraction module, configured to, for any of the second monitoring nodes, extract the image of the target object based on the second position information of the target object in the second image collected by the second monitoring node.

8. An electronic device, characterized in that, Include at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-6.

9. A computer storage medium, characterized in that, The computer storage medium stores a computer program, and the computer program is used to cause a computer to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Target identification and detection association method and device, medium and computer program product

    CN113642404A