A method, apparatus and device for obtaining location information
By using deep learning models in the target tracking algorithm for feature extraction, and combining target detection and related filtering algorithms, the accuracy problem of target tracking in the existing technology in complex scenarios is solved, and more efficient target position determination is achieved, supporting accurate inspections of epidemiological investigation personnel.
Patent Information
- Application Number
- CN202110475736.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-29
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-04-29
AI Technical Summary
In the prior art, the target tracking algorithm has low accuracy when facing scenes such as appearance deformation, lighting changes, rapid motion, motion blur, background similar interference, out-of-plane rotation, in-plane rotation, scale changes, occlusion and out-of-field field of view, resulting in difficulties in the investigation of contact personnel when the epidemiological investigation personnel are in contact.
By obtaining the designated monitoring video, the position information of the target object in the first frame image is determined, its characteristic information is extracted, and subsequent frame images are featured. Combined with the object detection algorithm and related filtering algorithm, the position information of the target object in each frame image is determined.
It improves the accuracy of target tracking, can more effectively determine the location information of target objects in complex scenarios, and supports epidemiological investigation personnel to more accurately detect contact personnel.
Smart Images

Figure CN113255460B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and particularly to a method, device, and equipment for obtaining location information. Background Art
[0002] Entering the cold winter season, the survival period of the virus is long. Currently, the effective way to prevent infection is to avoid crowd gathering, but there are still situations where crowd gathering cannot be avoided. Especially in bank branches, the flow of people is complex. Once a positive case is detected, it is necessary to investigate the people who have come into contact.
[0003] In the prior art, given surveillance videos within 14 days, target tracking based on deep learning can assist epidemiological investigators in identifying contact persons. However, in the field of computer vision, the target tracking algorithms currently face the following difficulties in scenarios: appearance deformation, illumination change, fast movement, motion blur, similar background interference, out-of-plane rotation, in-plane rotation, scale variation, occlusion, and out-of-view, etc., resulting in relatively low accuracy of the results determined by the target tracking algorithms.
[0004] Therefore, there is an urgent need in the industry for a technical solution that can solve the above technical problems. Summary of the Invention
[0005] Embodiments of this specification provide a method, device, and equipment for obtaining location information, which can make the results of target tracking more accurate.
[0006] The method, device, and equipment for obtaining location information provided in this specification are implemented in the following manner.
[0007] A method for obtaining location information includes: obtaining a specified surveillance video, determining the location information of a target object in the first frame image of the specified surveillance video; extracting the feature information of the target object from the first frame image according to the location information of the target object in the first frame image; performing feature extraction on the second frame image of the specified surveillance video to obtain the feature information of the second frame image, where the second frame image is the frame image after the first frame image; determining the location information of the target object in the second frame image based on the feature information of the target object and the feature information of the second frame image; and when it is determined that the second frame image is the last frame image of the specified surveillance video, obtaining the location information of the target object in each frame image of the specified surveillance video.
[0008] A device for obtaining location information, comprising: a first obtaining module, configured to obtain a specified surveillance video and determine the location information of a target object in the first frame image of the specified surveillance video; an extraction module, configured to extract the feature information of the target object from the first frame image according to the location information of the target object in the first frame image; an obtaining module, configured to perform feature extraction on the second frame image of the specified surveillance video to obtain the feature information of the second frame image; wherein the second frame image is the frame image after the first frame image; a determination module, configured to determine the location information of the target object in the second frame image based on the feature information of the target object and the feature information of the second frame image; a second obtaining module, configured to obtain the location information of the target object in each frame image of the specified surveillance video when the second frame image is the last frame image of the specified surveillance video.
[0009] A device for obtaining location information, comprising at least one processor and a memory storing computer-executable instructions, wherein when the processor executes the instructions, the steps of any one of the method embodiments in the embodiments of the present specification are implemented.
[0010] A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed, the steps of any one of the method embodiments in the embodiments of the present specification are implemented.
[0011] A method, device and equipment for obtaining location information provided in this specification. In some embodiments, a specified surveillance video can be obtained, the location information of a target object in the first frame image of the specified surveillance video can be determined, and according to the location information of the target object in the first frame image, the feature information of the target object can be extracted from the first frame image. Further, feature extraction can be performed on the second frame image of the specified surveillance video to obtain the feature information of the second frame image; wherein the second frame image is the next frame image after the first frame image, and based on the feature information of the target object and the feature information of the second frame image, the location information of the target object in the second frame image can be determined. It is also possible to obtain the location information of the target object in each frame image of the specified surveillance video when it is determined that the second frame image is the last frame image of the specified surveillance video. Since the present application uses a deep learning model for feature extraction, uses a target detection algorithm to determine the location of the target object, and uses a correlation filtering algorithm to determine the object associated with the target object, the result of target tracking using the solution of the present application is more accurate. By adopting the implementation scheme provided in this specification, the result of target tracking can be made more accurate. Description of the Drawings
[0012] The drawings described herein are used to provide a further understanding of this specification, form a part of this specification, and do not limit this specification. In the drawings:
[0013] Figure 1 It is a schematic flowchart of an embodiment of a method for obtaining location information provided in this specification;
[0014] Figure 2 It is a schematic diagram of the Mish activation function image and the ReLU activation function image provided in this specification;
[0015] Figure 3 It is a schematic diagram of the module structure of an embodiment of a device for obtaining location information provided in this specification;
[0016] Figure 4 It is a block diagram of the hardware structure of an embodiment of a server for obtaining location information provided in this specification. Detailed implementation manners
[0017] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only some of the embodiments in this specification, rather than all the embodiments. Based on one or more embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope protected by the embodiments of this specification.
[0018] The implementation scheme of this specification will be described below by taking a specific application scenario as an example. Specifically, Figure 1 It is a schematic flowchart of an embodiment of a method for obtaining location information provided in this specification. Although this specification provides method operation steps or device structures as shown in the following embodiments or drawings, based on routine or non-creative labor, more or some combined and fewer operation steps or module units may be included in the method or device.
[0019] An implementation scheme provided in this specification can be applied to a client, a server, etc. The client may include terminal devices, such as smart phones, tablet computers, etc. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed system, etc.
[0020] It should be noted that the following embodiments do not limit the technical solutions in other extensible application scenarios based on this specification. A specific example is Figure 1 As shown, in an embodiment of a method for obtaining location information provided in this specification, the method may include the following steps.
[0021] S0: Obtain a specified surveillance video, and determine the location information of the target object in the first frame image of the specified surveillance video.
[0022] Among them, the specified surveillance video can be obtained from the surveillance system. The specified surveillance video can be a video corresponding to the target object. The target object can be an object for which location information needs to be tracked. The target object can include one or more.
[0023] In some embodiments, when there are multiple target objects in the specified surveillance video, the KM (Kuhn - Munkres) algorithm can be used to perform association matching on the target objects in different frame images. For example, in some implementation scenarios, if there are several targets to be tracked in the current frame, each target can be tagged. Due to the situation of target disappearance and the increase in the number of targets, the KM algorithm can be used for multi - target association matching to continuously tag the corresponding targets, ensuring that the target labeled as target 1 still has the label of target 1 in the next frame image.
[0024] In some implementation scenarios, after determining the target object to be tracked, the surveillance video including the target object can be retrieved from the surveillance system according to the information of the target object. Among them, the information of the target object can include name, identity identifier, residential address, etc.
[0025] In some implementation scenarios, the video including the target object retrieved from the surveillance system may be very long and occupy a relatively large space. To improve the subsequent processing efficiency and accuracy, the retrieved video can be pre - processed. Among them, the pre - processing methods include but are not limited to cropping, integration, etc. For example, in some implementation scenarios, after retrieving the video including the target object from the surveillance system, a 14 - day video can be intercepted as the specified surveillance video. In some implementation scenarios, the specified surveillance video can include at least two frame images.
[0026] In some embodiments, after obtaining the specified surveillance video, the position information of the target object in the first frame image of the specified surveillance video can be determined by using a target detection algorithm. Among them, the position information can include position coordinates, position region boxes, etc. Among them, the target detection algorithm can be used to detect the positions of all human bodies in the current frame image and obtain the position information. The target detection algorithm can include yolov5, SSD (Single Shot MultiBox Detector), etc.
[0027] In some implementation scenarios, the first frame image can be the initial frame image in the specified surveillance video, or the current frame image, or other frame images. This specification does not make a limitation on this.
[0028] S2: Extract the feature information of the target object from the first frame image according to the position information of the target object in the first frame image.
[0029] In the embodiments of this specification, after determining the position information of the target object in the first frame image of the specified surveillance video, the feature information of the target object can be extracted from the first frame image according to the position information of the target object in the first frame image. Among them, the feature information can include spatial information, semantic information, etc. Among them, the spatial information can include the contour information corresponding to each object when recognizing the image. The semantic information can include, at the pixel level when recognizing the image, that is, similar to semantic segmentation of the image, assigning a category to each pixel in the image.
[0030] In some embodiments, before extracting the feature information of the target object from the first frame image, it may further include: determining whether there is a first object in the first frame image; when it is determined that there is, using an object detection algorithm to determine the position of the first object; based on the position of the first object, obtaining the position information of the first object; according to the position information of the target object in the first frame image and the position information of the first object, determining whether the first object is an associated object of the target object. Among them, the first object may be other objects in the first frame image except the target object. The first object may include one or more. In some implementation scenarios, the target object may be a detected positive case person, and the associated object of the target object may be a person in close contact with the positive case person. Of course, the above is only for illustrative purposes, and the target object is not limited to the above examples. Those skilled in the art may make other changes under the inspiration of the technical essence of this application, but as long as the functions and effects achieved are the same or similar to those of this application, they should all be covered within the protection scope of this application.
[0031] In some implementation scenarios, after obtaining the position information of the target object in the first frame image and the position information of the first object, the Euclidean distance between the target object and the first object can be calculated according to the position information, and further it can be determined whether the Euclidean distance is within a preset threshold. If so, it can be determined that the first object is an associated object of the target object. Among them, the preset threshold can be set according to the actual scenario, and this specification does not limit this.
[0032] In some implementation scenarios, when determining that the first object is an associated object of the target object, the first object is taken as the first target object. Correspondingly, a specified surveillance video can be obtained, and the position information of the first target object in the first frame image of the specified surveillance video can be determined; according to the position information of the first target object in the first frame image, the feature information of the first target object can be extracted from the first frame image; feature extraction is performed on the second frame image of the specified surveillance video to obtain the feature information of the second frame image; wherein, the second frame image is a frame image after the first frame image; based on the feature information of the first target object and the feature information of the second frame image, the position information of the first target object in the second frame image is determined; in the case where the second frame image is the last frame image of the specified surveillance video, the position information of the first target object in each frame image of the specified surveillance video is obtained.
[0033] For example, in some implementation scenarios, after obtaining the integrated surveillance video, the target to be tracked can be selected in the first frame first, and at the same time, the position of the human body around the tracked target is detected using a target detection algorithm, and the distance between each human body and the target is obtained. If this distance is within a preset threshold, it can be stated that the human body is a close contact of the tracked target. At this time, this close contact can be taken as another target for tracking.
[0034] In some embodiments, extracting the feature information of the target object from the first frame image may include: using a deep learning model to extract the feature information of the target object from the first frame image; wherein, the structure of the deep learning model may include a convolutional layer, a residual layer, a pooling layer, and a fully connected layer; the convolutional layer may be 44 layers, and the activation function used by the deep learning model may be the Mish activation function. Among them, in a deep neural network, as the network depth deepens, the speed will become slower and slower, but the accuracy will be improved more and more. Therefore, in order to ensure the speed while improving the accuracy, the number of convolutional layers is initially selected as 44 layers in this application. Of course, the above is only for illustrative purposes, and the number of convolutional layers is not limited to the above examples. Those skilled in the art may make other changes under the inspiration of the technical essence of this application, but as long as the functions and effects achieved are the same or similar to those of this application, they should all be covered within the protection scope of this application.
[0035] As Figure 2 shown, Figure 2It is a schematic diagram of the Mish activation function image and the ReLU activation function image provided in this specification. Among them, the first figure is the Mish activation function image, and the second figure is the ReLU activation function image. The abscissa represents the value of the independent variable, and the ordinate represents the function value. As can be seen from the figure, the Mish activation function has no upper limit, which can ensure that there is no saturation area and there will be no gradient disappearance problem during the training process; the Mish activation function has a lower limit, which can ensure a certain strong regularization effect (appropriately fitting the model); the Mish activation function has non-monotonicity, which helps to maintain small negative values, thereby stabilizing the network gradient flow; the Mish activation function has infinite-order continuity and smoothness, which will have better generalization ability and effective optimization ability of the results, thereby improving the quality of the results. The differential of the ReLU activation function is 0, which cannot guarantee negative values, resulting in most neurons not being updated.
[0036] Since a smooth activation function can allow better information to penetrate into the neural network, resulting in better accuracy and generalization. Therefore, in the embodiments of this specification, the Mish activation function is used to replace the commonly used activation function, which can effectively avoid saturation caused by capping. Theoretically, the slight allowance for negative values allows for better gradient flow, rather than the hard zero boundary in the ReLU activation function. Of course, the above is only for illustrative purposes, and the activation function is not limited to the above examples. Those skilled in the art may make other changes under the inspiration of the technical essence of this application, but as long as the functions and effects achieved are the same or similar to those of this application, they should all be covered within the protection scope of this application.
[0037] S4: Extract features from the second frame image of the specified surveillance video to obtain the feature information of the second frame image; wherein, the second frame image is the frame image after the first frame image.
[0038] In the embodiments of this specification, after extracting the feature information of the target object from the first frame image, the feature extraction can be performed on the second frame image of the specified surveillance video to obtain the feature information of the second frame image. Among them, the second frame image is the frame image after the first frame image. For example, if the first frame image is the current frame image in the surveillance video, then the second frame image is the next frame image of the surveillance video.
[0039] In some embodiments, the extracting features from the second frame image of the specified surveillance video to obtain the feature information of the second frame image may include: dividing the second frame image into a preset number of image blocks; using a deep learning model to extract features from each image block to obtain the feature information of the second frame image; wherein, the feature information of the second frame image includes the feature information corresponding to each image block.
[0040] In some implementation scenarios, when extracting feature information from the second-frame image of a specified surveillance video using a deep learning model, the frame image can first be divided into multiple small image blocks, and then the deep learning model can be used to extract feature information from the divided small image blocks respectively to obtain the feature information of the second-frame image. Among them, the structure of the deep learning model can include a convolutional layer, a residual layer, a pooling layer, and a fully connected layer. The convolutional layer can be 44 layers, and the activation function used by the deep learning model can be the Mish activation function. The feature information can include spatial information, semantic information, etc. The process of the deep learning model extracting feature information can be understood as a process in which an image matrix undergoes a convolution operation with a convolution kernel to obtain another matrix. Among them, the obtained matrix is also called a feature map.
[0041] In the embodiments of this specification, by selecting a deep learning model with an appropriate depth for feature extraction, the efficiency and accuracy of feature extraction can be higher.
[0042] In the embodiments of this specification, by using the Mish activation function to replace the commonly used activation function, saturation caused by capping can be avoided, and theoretically, a slight allowance for negative values allows for better gradient flow, so that the deep learning model can obtain better accuracy and generalization ability.
[0043] S6: Based on the feature information of the target object and the feature information of the second-frame image, determine the position information of the target object in the second-frame image.
[0044] In the embodiments of this specification, after obtaining the feature information of the second-frame image, the position information of the target object in the second-frame image can be determined based on the feature information of the target object and the feature information of the second-frame image.
[0045] In some embodiments, the determining the position information of the target object in the second-frame image based on the feature information of the target object and the feature information of the second-frame image may include: calculating the correlation between the feature information of the target object and the feature information corresponding to each image block in the feature information of the second-frame image using a correlation filtering algorithm; obtaining the position information corresponding to the image block with the highest correlation; and using the position information corresponding to the image block with the highest correlation as the position information of the target object in the second-frame image.
[0046] In some implementation scenarios, after extracting the feature information of the target object from the first-frame image and obtaining the feature information of the second-frame image, the correlation filtering method can be used to find the image block in the second-frame image with the highest correlation with the feature information of the target object extracted from the previous frame image, and then use the position information of the image block in the second-frame image as the position information of the target object in the second-frame image.
[0047] In some implementation scenarios, a filtering template can be designed, and the correlation operation is performed between the template and the target candidate region, and the position of the maximum output response is taken as the position of the target object in the second-frame image. Among them, the target candidate region can be each image block after the division of the above-mentioned second-frame image. Specifically, it can be expressed by the following formula:
[0048]
[0049] Among them, y represents the response output, x represents the target candidate region, and w represents the filtering template. In some implementation scenarios, the above-mentioned correlation operation can be converted into a dot product with less computational complexity by using relevant theorems:
[0050]
[0051] Among them, are the Fourier transforms of y, x, and w respectively.
[0052] In some implementation scenarios, after determining the position of the target object in the first-frame image of the specified surveillance video, sampling can be performed near this position to train a ridge regression function, and then the response of a small window sampling is calculated by using the trained ridge regression function, and then the position of the target object in the next-frame image is determined according to the response. Among them, the response of the small window sampling can be understood as the response of each image block in the next-frame image.
[0053] In some implementation scenarios, sampling can be performed near a specified position (the specified position can be the position of the target object in the first-frame image) through a cyclic shift operation. Among them, the cyclic shift operation can be used to approximate the displacement of the sampling window. In some implementation scenarios, when training the ridge regression function, a series of displacement samplings around the specified position can be represented by a two-dimensional block circulant matrix X, and the (i, j) block can represent the result of shifting the original image down by i rows and right by j columns. Among them, X can complete the linear operation by using the fast Fourier transform.
[0054] Since ridge regression presents a simple closed-form solution and can achieve performance close to that of more complex methods, in the training stage, a function f(x i ) = w T x i needs to be found so that the squared error reaches the minimum value, that is, Among them, x i is the sample, y i is the target to be regressed, λ is the regularization parameter for controlling overfitting, and ω is a column vector representing the weight coefficient.
[0055] In some implementation scenarios, after obtaining the ridge regression function, the next frame of the image can be entered, the next frame of the image can be divided into image blocks, and then the deep learning model can be used to extract the feature information of each image block (sample). Finally, the feature information is input into the trained ridge regression function to determine the response of each sample, and the position corresponding to the image block with the maximum response is used as the position of the target object in the frame of the image.
[0056] Since the correlation filtering method has a very high response speed, the position of the target object determined by the correlation filtering method in each frame of the image in the embodiments of the present specification is more accurate.
[0057] S8: In the case where it is determined that the second frame of the image is the last frame of the specified surveillance video, obtain the position information of the target object in each frame of the specified surveillance video.
[0058] In the embodiments of the present specification, after determining the position information of the target object in the second frame of the image, it can be determined whether the second frame of the image is the last frame of the specified surveillance video. If so, the position information of the target object in each frame of the specified surveillance video can be obtained.
[0059] In some implementation scenarios, the specified surveillance video may include multiple frames of images. Thus, when it is determined that the second frame of the image is not the last frame of the specified surveillance video, the feature information of the target object can be extracted from the second frame of the image according to the position information of the target object. Further, the feature extraction of the third frame of the image of the specified surveillance video can be performed to obtain the feature information of the third frame of the image; wherein, the third frame of the image is the next frame after the second frame of the image; based on the feature information of the target object and the feature information of the third frame of the image, determine the position information of the target object in the third frame of the image; determine whether the third frame of the image is the last frame of the specified surveillance video. When it is determined to be so, obtain the position information of the target object in each frame of the specified surveillance video.
[0060] In some implementation scenarios, when it is determined that the first object is the associated object of the target object, that is, when determining the close contact corresponding to the target object, the close contact can be used as another target object, and then the method for obtaining the position information provided in the present specification can be used to obtain the position information of the target object in the specified surveillance video, so as to achieve target tracking.
[0061] In some implementation scenarios, after determining the associated object of the target object, the information of the associated object can be output to a preset file so as to perform target tracking on the associated object according to this information. Among them, the information of the associated object can include location information, basic information (such as name, identity identifier, etc.), and so on. The preset file can be a txt file, an Excel file, etc., and this specification does not limit this.
[0062] In some implementation scenarios, since the determined close contacts can include multiple ones, in this way, multi-target tracking needs to be carried out. In some implementation scenarios, when there are multiple targets to be tracked, an association algorithm (such as the KM algorithm) can be used for target association matching. Among them, the association algorithm can mainly perform the matching of multiple targets between frames, including the appearance of new targets and the disappearance of old targets, as well as the matching of targets between the previous frame and the current frame. The KM algorithm is an optimal matching solution for a weighted bipartite graph.
[0063] In some implementation scenarios, after obtaining the position of the target object in each frame image of the specified surveillance video, disinfection and sterilization, etc. can be performed on the area corresponding to this position.
[0064] In some implementation scenarios, after determining the associated object of the target object, the associated object can be isolated.
[0065] In the embodiments of this specification, by combining the speed of the correlation filtering method and the accuracy of deep learning, the real-time performance, accuracy, and precision of target tracking can be enhanced.
[0066] The above test method will be described below with specific implementation scenarios. However, it should be noted that this specific embodiment is only for better explaining this application and does not constitute an improper limitation to this application.
[0067] For example, in some implementation scenarios, when performing single-target tracking, the position of the target object in the initial frame image can be determined first, denoted as x p , and then a deep learning model is used to extract the feature information of the area where this position is located to obtain the feature information of the target object in the initial frame image, denoted as T. Further, sampling can be performed near x p to train a ridge regression function. Then enter the next frame image, sample near x p , use the deep learning model to extract the feature information of the current image sample, denoted as T i , input the feature information T i into the ridge regression function trained in the previous step to judge the response of each image sample, and use the image sample with the strongest response as the position of the target object in the current frame, denoted as x o。Repeat the above steps until all frames of the video are processed to obtain the position information of the target object in each frame of the image. It should be noted that the above method for tracking a single target can be simply referred to as the single-target tracking method hereinafter.
[0068] In some implementation scenarios, when performing multi-target tracking, the position of the target object in the current frame of the image can be determined first, denoted as x1, and then the above single-target tracking method can be used to determine the position of the target object in the next frame of the image. At the same time, the target detection algorithm yolov5 is used to detect all humans in the current frame of the image, and the Euclidean distance between other humans and the target object is calculated. If the Euclidean distance is less than the set threshold T0, it means that this human belongs to a close contact, and at this time, this human can be marked as the target object x2 to be tracked, and so on, marked as x i 。Furthermore, for each target object to be tracked in the current frame of the image, the above single-target tracking method is used for target tracking. Among them, if there are several target objects to be tracked in the current frame of the image, each target object can be tagged. Due to the situation of target disappearance and the increase in the number of targets, the KM algorithm is used for multi-target association matching to continuously mark the corresponding target objects, ensuring that the target object with the label of target 1 still has the label of target 1 in the next frame of the image. Repeat the above steps until all frames of the video are processed to obtain the position information of each target object in each frame of the image.
[0069] Of course, the above is only an exemplary description. The embodiments of this specification are not limited to the above examples. Those skilled in the art may make other changes under the inspiration of the technical essence of this application. However, as long as the functions and effects achieved are the same or similar to those of this application, they should all be covered within the protection scope of this application.
[0070] From the above description, it can be seen that the embodiments of this application can obtain a specified surveillance video, determine the position information of the target object in the first frame of the image of the specified surveillance video, and extract the feature information of the target object from the first frame of the image according to the position information of the target object in the first frame of the image. Further, the feature extraction of the second frame of the image of the specified surveillance video can be performed to obtain the feature information of the second frame of the image; where the second frame of the image is the next frame after the first frame of the image. Based on the feature information of the target object and the feature information of the second frame of the image, the position information of the target object in the second frame of the image is determined. It is also possible to obtain the position information of the target object in each frame of the specified surveillance video when it is determined that the second frame of the image is the last frame of the specified surveillance video. Since this application uses a deep learning model for feature extraction, uses a target detection algorithm to determine the position of the target object, and uses a correlation filtering algorithm to determine the object associated with the target object, the result of target tracking using the solution of this application is more accurate.
[0071] In the embodiments of the above method in this specification, each embodiment is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. For the relevant parts, refer to the partial description of the method embodiments.
[0072] Based on the above-mentioned method for obtaining location information, one or more embodiments of this specification also provide a device for obtaining location information. The described device may include a system (including a distributed system), software (application), module, component, server, client, etc. that uses the method described in the embodiments of this specification and combines the necessary implementation hardware. Based on the same inventive concept, the devices in one or more embodiments provided in the embodiments of this specification are as described in the following embodiments. Since the implementation solutions for the device to solve problems are similar to the method, the implementation of the specific device in the embodiments of this specification can refer to the implementation of the foregoing method, and the repeated parts will not be elaborated. As used hereinafter, the term "unit" or "module" may be a combination of software and / or hardware that can implement a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0073] Specifically, Figure 3 is a schematic diagram of the module structure of an embodiment of a device for obtaining location information provided in this specification. As Figure 3 shown, a device for obtaining location information provided in this specification may include: a first acquisition module 120, an extraction module 122, an acquisition module 124, a determination module 126, and a second acquisition module 128.
[0074] The first acquisition module 120 may be used to acquire a specified surveillance video and determine the location information of the target object in the first frame image of the specified surveillance video;
[0075] The extraction module 122 may be used to extract the feature information of the target object from the first frame image according to the location information of the target object in the first frame image;
[0076] The acquisition module 124 may be used to perform feature extraction on the second frame image of the specified surveillance video to obtain the feature information of the second frame image; wherein, the second frame image is the frame image after the first frame image;
[0077] The determination module 126 may be used to determine the location information of the target object in the second frame image based on the feature information of the target object and the feature information of the second frame image;
[0078] The second acquisition module 128 can be used to acquire the position information of the target object in each frame image of the specified surveillance video when the second frame image is the last frame image of the specified surveillance video.
[0079] It should be noted that the above-described device may also include other implementation manners according to the description of the method embodiments. The specific implementation manners may refer to the description of the relevant method embodiments and will not be elaborated herein one by one.
[0080] This specification also provides an embodiment of a device for acquiring position information, including a processor and a memory for storing processor-executable instructions. When the instructions are executed by the processor, the following steps are implemented: acquiring a specified surveillance video, determining the position information of a target object in the first frame image of the specified surveillance video; extracting the feature information of the target object from the first frame image according to the position information of the target object in the first frame image; performing feature extraction on the second frame image of the specified surveillance video to obtain the feature information of the second frame image; where the second frame image is the frame image after the first frame image; determining the position information of the target object in the second frame image based on the feature information of the target object and the feature information of the second frame image; and when it is determined that the second frame image is the last frame image of the specified surveillance video, acquiring the position information of the target object in each frame image of the specified surveillance video.
[0081] It should be noted that the above-described device may also include other implementation manners according to the description of the method or device embodiments. The specific implementation manners may refer to the description of the relevant method embodiments and will not be elaborated herein one by one.
[0082] The method embodiments provided in this specification can be executed on a mobile terminal, a computer terminal, a server, or a similar computing device. Taking running on a server as an example, Figure 4 is a hardware structure block diagram of an embodiment of a server for acquiring position information provided in this specification. The server can be the device for acquiring position information or the device for acquiring position information in the above embodiments. As Figure 4 shown, the server 10 may include one or more (only one is shown in the figure) processors 100 (the processor 100 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 200 for storing data, and a transmission module 300 for communication functions. Those of ordinary skill in the art can understand that Figure 4 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the server 10 may further include more Figure 4more or fewer components shown, for example, may further include other processing hardware, such as a database or a multi-level cache, a GPU, or have a configuration different from that Figure 4 shown.
[0083] The memory 200 can be used to store software programs and modules of application software, such as program instructions / modules corresponding to the method of obtaining location information in the embodiments of this specification. The processor 100 executes various functional applications and data processing by running the software programs and modules stored in the memory 200. The memory 200 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 200 may further include a memory remotely located relative to the processor 100, and these remote memories can be connected to the computer terminal through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0084] The transmission module 300 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of a computer terminal. In one instance, the transmission module 300 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission module 300 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0085] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order from that in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0086] The methods or devices described in the above embodiments provided in this specification can implement business logic through computer programs and record them on a storage medium. The storage medium can be read and executed by a computer to achieve the effects of the solutions described in the embodiments of this specification. The storage medium can include physical devices for storing information, usually storing information after digitization and then using media such as electricity, magnetism, or optics. The storage medium can include: devices that store information using electrical energy, such as various memories, such as RAM, ROM, etc.; devices that store information using magnetic energy, such as hard disks, floppy disks, magnetic tapes, magnetic core memories, bubble memories, USB flash drives; devices that store information using optical methods, such as CDs or DVDs. Of course, there are also other ways of readable storage media, such as quantum memories, graphene memories, and so on.
[0087] The method or device embodiments for obtaining location information provided in this specification can be implemented in a computer by a processor executing corresponding program instructions, such as being implemented in a PC using the C++ language with the Windows operating system, being implemented in a Linux system, or being implemented in a smart terminal using programming languages such as Android or iOS, and being implemented based on the processing logic of a quantum computer, etc.
[0088] It should be noted that the devices, equipment, and systems described in the above specification may also include other implementation manners according to the descriptions of the relevant method embodiments. The specific implementation manners can refer to the descriptions of the corresponding method embodiments and will not be elaborated here one by one.
[0089] Each embodiment in this application is described in a progressive manner. For the parts that are the same or similar among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the hardware + program type embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial descriptions of the method embodiments.
[0090] For the convenience of description, when describing the above devices, they are described separately as various modules according to functions. Of course, when implementing one or more of this specification, the functions of some modules can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc.
[0091] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and systems according to embodiments of the present invention. It should be understood that it can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate a device for realizing the specified functions. These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one or more processes and / or blocks Figure 1 in one process or multiple processes and / or blocks Figure 1 or in one block or multiple blocks.
[0092] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.
[0093] The above are only embodiments of one or more embodiments of this specification and are not used to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims.
Claims
1. A method for obtaining location information, characterized in that, Including: Obtain a specified surveillance video, and determine the position information of the target object in the first frame image of the specified surveillance video; Extract the feature information of the target object from the first frame image according to the position information of the target object in the first frame image; Perform feature extraction on the second frame image of the specified surveillance video to obtain the feature information of the second frame image; wherein, the second frame image is the frame image after the first frame image; Based on the feature information of the target object and the feature information of the second frame image, determine the position information of the target object in the second frame image; When it is determined that the second frame image is the last frame image of the specified surveillance video, obtain the position information of the target object in each frame image of the specified surveillance video; Before extracting the feature information of the target object from the first frame image, it further includes: Judge whether there is a first object in the first frame image; When it is determined that there is one, use an object detection algorithm to determine the position of the first object; Based on the position of the first object, obtain the position information of the first object; According to the position information of the target object in the first frame image and the position information of the first object, determine whether the first object is an associated object of the target object; When it is determined that the first object is an associated object of the target object, use the first object as the first target object; Correspondingly, obtain a specified surveillance video, and determine the position information of the first target object in the first frame image of the specified surveillance video; Extract the feature information of the first target object from the first frame image according to the position information of the first target object in the first frame image; Perform feature extraction on the second frame image of the specified surveillance video to obtain the feature information of the second frame image; wherein, the second frame image is the frame image after the first frame image; Based on the feature information of the first target object and the feature information of the second frame image, determine the position information of the first target object in the second frame image; When the second frame image is the last frame image of the specified surveillance video, obtain the position information of the first target object in each frame image of the specified surveillance video.
2. The method according to claim 1, characterized in that, Extracting the feature information of the target object from the first frame image includes: Use a deep learning model to extract the feature information of the target object from the first frame image; wherein, the structure of the deep learning model includes a convolutional layer, a residual layer, a pooling layer, and a fully connected layer; the convolutional layer has 44 layers, and the activation function used by the deep learning model is the Mish activation function.
3. The method according to claim 2, characterized in that, Performing feature extraction on the second frame image of the specified surveillance video to obtain the feature information of the second frame image includes: Divide the second frame image into a preset number of image blocks; Use a deep learning model to perform feature extraction on each image block to obtain the feature information of the second frame image; wherein, the feature information of the second frame image includes the feature information corresponding to each image block.
4. The method according to claim 3, characterized in that, Based on the feature information of the target object and the feature information of the second frame image, determining the position information of the target object in the second frame image includes: Calculate the correlation degree of the feature information of the target object and the feature information corresponding to each image block in the feature information of the second frame image by using the correlation filtering algorithm; Obtain the position information corresponding to the image block with the highest correlation degree; Use the position information corresponding to the image block with the highest correlation degree as the position information of the target object in the second frame image.
5. The method according to claim 1, characterized in that, It further includes: In the case where it is determined that the second frame image is not the last frame image of the specified surveillance video, extract the feature information of the target object from the second frame image according to the position information of the target object in the second frame image; Correspondingly, perform feature extraction on the third frame image of the specified surveillance video to obtain the feature information of the third frame image; where the third frame image is the next frame image after the second frame image; based on the feature information of the target object and the feature information of the third frame image, determine the position information of the target object in the third frame image; determine whether the third frame image is the last frame image of the specified surveillance video, and when it is determined to be so, obtain the position information of the target object in each frame image of the specified surveillance video.
6. The method according to claim 1, characterized in that, When there are multiple target objects in the specified surveillance video, it further includes: Use the KM algorithm to perform association matching on the target objects in different frame images.
7. A device for obtaining location information, characterized in that, It includes: The first acquisition module is used to acquire the specified surveillance video and determine the position information of the target object in the first frame image of the specified surveillance video; The extraction module is used to extract the feature information of the target object from the first frame image according to the position information of the target object in the first frame image; The acquisition module is used to perform feature extraction on the second frame image of the specified surveillance video to obtain the feature information of the second frame image; where the second frame image is the frame image after the first frame image; The determination module is used to determine the position information of the target object in the second frame image based on the feature information of the target object and the feature information of the second frame image; The second acquisition module is used to acquire the position information of the target object in each frame image of the specified surveillance video when the second frame image is the last frame image of the specified surveillance video; Before extracting the feature information of the target object from the first frame image, it further includes: Determine whether there is a first object in the first frame image; When it is determined that there is one, use the target detection algorithm to determine the position of the first object; Based on the position of the first object, obtain the position information of the first object; According to the position information of the target object in the first frame image and the position information of the first object, determine whether the first object is an associated object of the target object; When it is determined that the first object is an associated object of the target object, use the first object as the first target object; Correspondingly, acquire the specified surveillance video and determine the position information of the first target object in the first frame image of the specified surveillance video; Extract the feature information of the first target object from the first frame image according to the position information of the first target object in the first frame image; Extract features from the second frame image of the specified surveillance video to obtain the feature information of the second frame image; wherein, the second frame image is the frame image after the first frame image; Based on the feature information of the first target object and the feature information of the second frame image, determine the position information of the first target object in the second frame image; In the case where the second frame image is the last frame image of the specified surveillance video, obtain the position information of the first target object in each frame image of the specified surveillance video.
8. An apparatus for obtaining location information, characterized in that, Comprising at least one processor and a memory storing computer-executable instructions, the processor, when executing the instructions, implements the steps of the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Stored thereon are computer instructions which, when executed, implement the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Target tracking method and device suitable for accurate monitoring and computer equipment
CN110516559A
Target tracking method and device, electronic equipment and readable storage medium
CN110910422A