Cross-lens personnel retrieval method and system based on cloud edge collaboration, electronic equipment and storage medium
Through the cloud-edge collaborative architecture, the object detection and personnel retrieval algorithms are delegated to edge devices, and cross-lens personnel retrieval is achieved using feature vector matching, which solves the problem of missing targets in the blind spot of the field of vision, improves the accuracy and real-time tracking, and reduces the computing resource requirements.
Patent Information
- Application Number
- CN202510544839.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-19
AI Technical Summary
The prior art cannot continue to track the target after the blind spot of the field of vision is lost, and the computing performance requirements are high, resulting in insufficient personnel tracking accuracy and real-time.
The cloud-edge collaborative architecture is adopted to delegate the object detection and personnel search algorithms to edge devices for calculations. By obtaining the search tasks issued by the cloud, cross-lens personnel search is achieved using feature vector matching, reducing the cloud computing performance requirements.
After losing targets in the blind spot in the field of vision, the target can be repositioned, improve the accuracy and real-timeness of personnel tracking, and reduce the demand for computing resources.
Smart Images

Figure CN120510562A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a cross-lens personnel retrieval method, system, electronic device and storage medium based on cloud-edge collaboration. Background Art
[0002] With the development of deep learning, people tracking technology is increasingly being used in video security. For example, in smart city security management, suspicious targets (such as terrorists and criminal suspects) may flee the scene after committing a crime in a specific area. Managers need to use on-site surveillance equipment to track them across cameras, quickly identifying the target's latest location and notifying security personnel for interception and capture. Children are often lost in scenic spots and public places. Using cross-camera tracking technology, managers can locate missing children in a timely manner, greatly improving the efficiency of personnel searches.
[0003] Related technologies usually use target detection algorithms combined with multi-target tracking algorithms to achieve cross-lens tracking. This tracking method is limited by the blind spots of the camera's field of view. Since tracking relies more on field of view coherence and temporal continuity, if the target is lost in the blind spot, subsequent tracking cannot be completed.
[0004] In addition, as the number of targets and cameras that need to be matched increases, the tracking algorithm requires more computing performance to meet the real-time performance of the system, and has higher requirements for the deployed hardware. Summary of the Invention
[0005] The main purpose of the embodiments of this application is to propose a cross-lens personnel retrieval method, system, electronic device and storage medium based on cloud-edge collaboration, aiming to relocate the target after losing the target in the blind spot of the field of view, reduce computing performance requirements, and improve the accuracy and real-time performance of personnel tracking.
[0006] To achieve the above objectives, one aspect of an embodiment of the present application proposes a cross-camera personnel retrieval method based on cloud-edge collaboration. The method is applied to an edge device, and the edge device is bound to several cameras. The method includes the following steps:
[0007] Obtaining a search task issued by the cloud, wherein the search task includes a target image and a search area;
[0008] Acquire the video stream of the camera according to the search area to obtain a monitoring image;
[0009] Performing personnel detection processing on the surveillance image to obtain an image to be retrieved;
[0010] Performing a person search process based on the target image and the image to be searched to obtain a search result;
[0011] The search results are converted into warning information and sent to the cloud.
[0012] In some embodiments, performing personnel detection processing on the surveillance image to obtain the image to be retrieved includes the following steps:
[0013] Performing target detection processing on the surveillance image to obtain category information and a predicted bounding box of the detected target;
[0014] Screening the detection targets according to the category information to obtain targets to be retrieved;
[0015] According to the predicted bounding box corresponding to the target to be retrieved, image interception processing is performed on the monitoring image to obtain the image to be retrieved.
[0016] In some embodiments, performing target detection processing on the surveillance image to obtain category information and a predicted bounding box of the detected target includes the following steps:
[0017] Scaling the monitoring image to obtain a proportional image;
[0018] Filling the proportional image with a fixed value to obtain a filled image;
[0019] performing normalization processing on the filled image to obtain a normalized image;
[0020] The normalized image is input into the target detection model for target detection processing to obtain category information and predicted bounding box of the detected target.
[0021] In some embodiments, performing a person search process based on the target image and the image to be searched to obtain a search result includes the following steps:
[0022] Performing feature extraction processing on the target image to obtain a target feature vector;
[0023] Performing feature extraction processing on the image to be retrieved to obtain a query feature vector;
[0024] Determining a vector distance based on the target feature vector and the query feature vector, and determining whether the vector distance is less than a preset threshold, to obtain a determination result;
[0025] If the judgment result is that the vector distance is less than the preset threshold, the image to be retrieved corresponding to the query feature vector and the camera device number of the image source camera are used as the retrieval result.
[0026] In some embodiments, performing feature extraction processing on the target image to obtain a target feature vector includes the following steps:
[0027] The target image is input into a lightweight neural network for encoding processing to obtain a target feature vector of fixed dimension.
[0028] In some embodiments, the cross-shot personnel retrieval method based on cloud-edge collaboration further includes the following steps:
[0029] Performing an associated query based on the camera device number of the image source camera to obtain an associated area;
[0030] Acquire the video stream of the camera according to the associated area to obtain a tracking image;
[0031] The retrieval result is updated according to the tracking image.
[0032] In some embodiments, performing an association query based on the camera device number of the image source camera to obtain the associated area includes the following steps:
[0033] Determining a first area according to a monitoring overlap range based on a camera device number of a camera from which the image originates;
[0034] Determining a second area based on the camera device number of the image source camera and path continuity;
[0035] An associated area is obtained according to the first area and the second area.
[0036] To achieve the above objectives, another aspect of the present application provides a cross-camera personnel retrieval system based on cloud-edge collaboration. The system is applied to an edge device, which is bound to several cameras. The system includes:
[0037] The first module is used to obtain a search task issued by the cloud, wherein the search task includes a target image and a search area;
[0038] The second module is used to obtain the video stream of the camera according to the search area to obtain a monitoring image;
[0039] The third module is used to perform personnel detection processing on the monitoring image to obtain an image to be retrieved;
[0040] The fourth module is used to perform a person search process based on the target image and the image to be searched to obtain a search result;
[0041] The fifth module is used to convert the search results into alarm information and send it to the cloud.
[0042] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned method when executing the computer program.
[0043] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above-mentioned method when executed by a processor.
[0044] The embodiments of the present application include at least the following beneficial effects: the present application provides a cross-lens personnel retrieval method, system, electronic device and storage medium based on cloud-edge collaboration. The solution obtains a retrieval task issued by the cloud, and the retrieval task includes a target image and a retrieval area; obtains the video stream of the camera according to the retrieval area to obtain a monitoring image; performs personnel detection processing on the monitoring image to obtain an image to be retrieved; performs personnel retrieval processing based on the target image and the image to be retrieved to obtain a retrieval result; converts the retrieval result into an alarm information and sends it to the cloud. The present application can relocate the target after the target is lost in the blind spot of the field of view, reduce computing performance requirements, and improve the accuracy and real-time performance of personnel tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a flowchart of a cross-lens personnel retrieval method based on cloud-edge collaboration provided in an embodiment of the present application;
[0046] Figure 2 This is a schematic diagram of the cloud-edge collaborative access form provided by an embodiment of the present application;
[0047] Figure 3 This is a structural diagram of the YOLOv8 network provided in an embodiment of the present application;
[0048] Figure 4 This is a flowchart of the image processing of the personnel detection algorithm provided in the embodiment of the present application;
[0049] Figure 5 This is a flowchart of a personnel search algorithm provided by an embodiment of the present application;
[0050] Figure 6 2 is a schematic diagram of a framework of a cross-lens personnel retrieval method based on cloud-edge collaboration provided by another embodiment of the present application;
[0051] Figure 7 This is a structural diagram of a cross-lens personnel retrieval system based on cloud-edge collaboration provided by an embodiment of the present application;
[0052] Figure 8 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0054] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0055] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0057] Before describing the embodiments of the present application in detail, the nouns involved in the embodiments of the present application and some related technologies are described as follows.
[0058] (1) Network Video Recorder (NVR) is the storage and forwarding part of the network video surveillance system. Its core function is to store and forward video streams. It usually works in conjunction with a video encoder or network camera to complete the video recording, storage and forwarding functions.
[0059] (2) The Real Time Streaming Protocol (RTSP) is a text-based multimedia playback control protocol. It is an application layer protocol that operates in a client / server manner and is often used to control real-time multimedia on-demand services. However, it is not used to transmit streaming media data itself and must rely on the services provided by the Real-time Transport Protocol (RTP) and the Real-time Transport Control Protocol (RTCP) to complete the transmission of streaming media data. RTSP is responsible for defining specific control information, operation methods, status codes, and describing the interaction between it and RTP.
[0060] Cross-camera tracking relies heavily on the camera's field of view consistency and temporal continuity. In actual deployments, surveillance cameras often have blind spots—areas not covered by cameras, or areas where the coverage of different cameras is not fully connected. When a target enters these blind spots, the system loses access to its continuous position, resulting in tracking interruption. Furthermore, even with comprehensive camera coverage, targets may temporarily disappear during movement due to factors such as occlusion and changing lighting.
[0061] At the same time, most common personnel tracking methods use multi-target tracking. The server accesses the video stream intranet or NVR and performs personnel detection and target tracking simultaneously. Since multi-target tracking is commonly completed using filtering algorithms and matching algorithms such as Kalman Filter and Hungarian matching, its computing resource requirements increase with the increase in tracking targets. Common cross-lens tracking may require the association of hundreds or even thousands of cameras, which places high requirements on the server configuration on the server side.
[0062] In view of this, an embodiment of the present application provides a cross-lens personnel retrieval method, system, electronic device and storage medium based on cloud-edge collaboration. The solution obtains a retrieval task issued by the cloud, and the retrieval task includes a target image and a retrieval area; obtains the video stream of the camera according to the retrieval area to obtain a monitoring image; performs personnel detection processing on the monitoring image to obtain an image to be retrieved; performs personnel retrieval processing based on the target image and the image to be retrieved to obtain a retrieval result; converts the retrieval result into an alarm information and sends it to the cloud. The present application can relocate the target after the target is lost in the blind spot of the field of view, reduce computing performance requirements, and improve the accuracy and real-time performance of personnel tracking.
[0063] The cross-lens personnel retrieval method based on cloud-edge collaboration provided in the embodiment of the present application relates to the field of image processing technology. The cross-lens personnel retrieval method based on cloud-edge collaboration provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements a cross-lens personnel retrieval method based on cloud-edge collaboration, etc., but is not limited to the above forms.
[0064] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0065] Figure 1 This is an optional flowchart of the cross-lens personnel retrieval method based on cloud-edge collaboration provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S105.
[0066] Step S101: Obtain a search task sent from the cloud, where the search task includes a target image and a search area.
[0067] Step S102: acquiring the video stream of the camera according to the search area to obtain a monitoring image.
[0068] Step S103: Perform personnel detection processing on the surveillance image to obtain an image to be retrieved.
[0069] Step S104: performing person retrieval processing based on the target image and the image to be retrieved to obtain a retrieval result.
[0070] Step S105: convert the search results into alarm information and send it to the cloud.
[0071] In this embodiment, a cloud-edge collaborative architecture is adopted to fully utilize the computing power of cloud devices and edge devices, and the target detection algorithm and personnel retrieval algorithm are decentralized to the edge side for calculation. The cloud only needs to complete the dispatch of personnel retrieval tasks, and the edge device completes the personnel detection algorithm and personnel retrieval algorithm of the bound camera image. Figure 2 The cloud service platform can be deployed in the central computer room on the user side and adopt a B / S network architecture. Users can bind edge devices and cameras on the platform side, thereby issuing retrieval tasks and receiving structured alarm information reported by edge devices.
[0072] Understandably, the cloud can be deployed as a central server or a local server to support private deployment. This design allows users to flexibly adjust the cloud architecture design, optimize functional configuration, and improve performance based on specific business needs, thereby providing greater flexibility and adaptability. Through private deployment, users can better control data and information security. Furthermore, the cloud can be integrated with edge devices via the intranet, further improving data transmission speeds and meeting application scenarios with high real-time requirements.
[0073] Exemplarily, the user side determines the target image by uploading a picture of the target to be retrieved and customizing the frame, or selects the detection target in the historical alarm information as the target image to issue the retrieval task. The cloud records the task query ID of the task and waits for the retrieval result of the edge device. Once the retrieval is successful, the user is informed in the form of an alarm.
[0074] Edge devices are deployed on surveillance cameras or NVRs, within the same local area network as the connected cameras, and are responsible for reading surveillance images in real time. In this embodiment, the surveillance images are the camera's Real-Time Streaming Protocol (RTSP) data, which can be decoded to obtain image frames.
[0075] It should be noted that as the NPU performance of edge devices gradually increases, many algorithms can be quantitatively deployed to the edge side for computing, thereby reducing the computing pressure on cloud servers. In addition, edge-side deployment is easier to expand, and the number of edge devices can be dynamically adjusted and deployed according to the number of connected cameras, greatly improving the flexibility of deployment. Through this cloud-edge collaborative architecture, each edge device only needs to process a limited number of camera images, which can reduce the demand for computing resources. In addition, the personnel detection and personnel retrieval algorithms that require higher computing power are all run on the edge side, which not only reduces the network bandwidth demand and latency, but also reduces the dependence on cloud server performance, making it easier for users to deploy according to the number of cameras in different areas.
[0076] Furthermore, when the cloud server issues a cross-camera search task, a search region can be selected. This region refers to the search range specified by the user when performing a cross-camera personnel search, which determines which cameras are searched. For example, the user can specify a global search (all cameras) or a specific region (a subset of cameras).
[0077] It is understandable that users can dynamically adjust the search area according to actual needs. For example, in a global search, if the target is found in a specific area, the user can narrow the search scope to that area to improve the search efficiency.
[0078] After receiving the retrieval task sent from the cloud, the edge device reads the images of the cameras bound to the area in real time according to the retrieval area, runs the person detection algorithm to detect whether there is a person in the current image, and after detecting a person, uses the area image information of the person as the image to be retrieved.
[0079] After obtaining the target image and the image to be retrieved, the target image and the image to be detected are input into the person detection algorithm and encoded into feature vectors, and then matched to see if they are the same person. If the retrieval is successful, the alarm information is reported to the cloud service platform.
[0080] This embodiment uses image retrieval to locate people. Image retrieval is a matrix multiplication method with high efficiency and low computational requirements. It also fully utilizes the computing power of cloud and edge devices, offloading the target detection and person retrieval algorithms to edge devices for calculation, which can reduce the computing performance requirements of the cloud. Person retrieval is achieved through feature vector matching, which does not rely on the target's continuous motion trajectory. Even if the target is lost in the blind spot, the target can still be relocated through feature vector matching, solving the problem of difficult tracking of targets lost in the blind spot and improving the accuracy and real-time performance of person tracking.
[0081] In some embodiments, step S103 may include but is not limited to steps S201 to S203.
[0082] Step S201 : performing target detection processing on the surveillance image to obtain category information and predicted bounding boxes of the detected targets.
[0083] Step S202: screening the detection targets according to the category information to obtain the targets to be retrieved.
[0084] Step S203 : performing image capture processing on the monitoring image according to the predicted bounding box corresponding to the target to be retrieved, to obtain the image to be retrieved.
[0085] In this embodiment, the personnel detection algorithm serves as the input prerequisite for the personnel retrieval algorithm and adopts the YOLOv8 target detection algorithm based on a deep learning neural network. This algorithm is a one-stage target detection algorithm that can detect the predicted bounding boxes and category information of multiple targets in the surveillance image at one time. It not only has high accuracy but also good real-time performance.
[0086] Reference Figure 3 The YOLOv8 network structure consists of three parts: the backbone network (Backbone), the feature enhancement network (Neck), and the detection head (Head). By collecting a specific personnel data set and training, a suitable target detection model is obtained. The monitoring image is input into the trained target detection model, and the predicted bounding box (Bboxes) and category information (Cls) of the target identified in the monitoring image are output. Among them, the predicted bounding box is used to represent the image position information of the target in the image, including the horizontal coordinate x of the target center point, the vertical coordinate y of the target center point, the width w of the bounding box, and the height h of the bounding box; the category information is used to represent the category of the target identified in the detection, such as personnel, vehicles, or animals.
[0087] It should be noted that the YOLOv8 network does not use the manually set anchor box (Anchor Base) training method consistent with the YOLOv3 and YOLOv5 series. Instead, it uses anchor points without anchor boxes (Anchor Free) to predict the target box and the distribution of positive and negative samples, reducing the recognition accuracy difference caused by anchor box setting and improving the robustness of the person detection algorithm.
[0088] In actual scenarios, surveillance images may contain multiple targets at the same time. Non-human targets (such as cars or animals) are filtered out through category information, and detection targets classified as people are selected as targets to be retrieved to ensure that only people are processed subsequently.
[0089] After obtaining the target to be retrieved, the coordinates of the four corners of the bounding box are calculated based on the predicted bounding box of each target to be retrieved. Based on the determined bounding box, an image processing tool is used to take a screenshot of the corresponding surveillance image to obtain the image to be retrieved. Each image to be retrieved represents a person identified by the person detection algorithm.
[0090] Optionally, for a specific edge device, quantization technology is used to convert the parameter accuracy of the trained model from float32 to int8, and then the model is deployed on the specific edge device for personnel detection.
[0091] In some embodiments, step S201 may include but is not limited to steps S301 to S304.
[0092] Step S301: scaling the monitoring image to obtain a proportional image.
[0093] Step S302 , performing filling processing on the proportional image using a fixed value to obtain a filled image.
[0094] Step S303: normalize the filled image to obtain a normalized image.
[0095] Step S304: Input the normalized image into the target detection model for target detection processing to obtain category information and predicted bounding box of the detected target.
[0096] In this embodiment, the monitoring image is scaled to the network input size without changing the aspect ratio of the image to obtain a proportional image, thereby avoiding image distortion.
[0097] For example, referring to Figure 4 The size of the input surveillance image is 1920 (width) x 1080 (height), and the size of the network input size is 640x640. The surveillance image is scaled according to the aspect ratio of 9:16 until the height and width of the surveillance image are less than or equal to the height and width of the network input size.
[0098] When the size of the scaled image is smaller than the network input size, a fixed value can be used to fill the missing pixels above and below or left and right of the image to make the size of the scaled image equal to the size of the network input size to obtain a filled image. For example, Figure 4 After the input image is scaled, its width is 640 but its height is only 360. Then the difference pixels can be filled with gray with an RGB value of 114.
[0099] Normalize the pixel values in the padded image to adjust the pixel values of the image to a uniform range, for example, normalize the pixel values from [0, 225] to [0, 1] or [-1, 1] so that the model can process the data more efficiently.
[0100] The normalized image is input into the trained target detection model to obtain the network output, and the network output is encoded to obtain the category information and predicted bounding box of the detected target.
[0101] In some embodiments, step S104 may include but is not limited to steps S401 to S404.
[0102] Step S401: Perform feature extraction on the target image to obtain a target feature vector.
[0103] Step S402: performing feature extraction processing on the image to be retrieved to obtain a query feature vector.
[0104] Step S403 , determining a vector distance based on the target feature vector and the query feature vector, and determining whether the vector distance is less than a preset threshold to obtain a determination result.
[0105] In step S404, if the judgment result is that the vector distance is less than the preset threshold, the image to be retrieved corresponding to the query feature vector and the camera device number of the image source camera are used as the retrieval result.
[0106] In this embodiment, the personnel search algorithm defines the personnel search task as a 1-N matching problem. Figure 5 ,The personnel retrieval algorithm includes feature extraction module, feature vector library and ,vector distance calculation module.
[0107] The feature extraction module is responsible for encoding image information into feature vectors of specific dimensions. The image sources include the target image sent from the cloud and the image to be retrieved from the output of the person detection algorithm.
[0108] After the edge device obtains the target image of the retrieval task sent by the cloud, it performs feature extraction on the target image through the feature extraction module, encodes it into a feature vector to obtain the target feature vector.
[0109] The feature vector library is used to store the target feature vectors of the target image after being encoded by the feature extraction module, and serves the subsequent vector distance calculation module.
[0110] Whenever the person detection algorithm outputs an image to be retrieved, the image to be retrieved is input into the feature extraction module for feature extraction processing, and encoded into a feature vector to obtain a query feature vector.
[0111] The vector distance calculation module is responsible for calculating the Euclidean distance between the query feature vector and the target feature vector in the feature vector library, and judging whether they are the same person based on the preset threshold.
[0112] For example, assuming that the preset threshold is 1.0, if the Euclidean distance between the two vectors is less than 1.0, the person currently being retrieved is considered to be the target person in the retrieval task, and the image to be retrieved and the camera device number of the camera from which the image is source are used as the retrieval results, so that the user can further verify based on the image to be retrieved, and determine the current geographic location of the person based on the camera device number.
[0113] In some embodiments, step S401 may include but is not limited to step S501.
[0114] Step S501: Input the target image into a lightweight neural network for encoding processing to obtain a target feature vector of fixed dimension.
[0115] In this embodiment, since the image information stores RGB information and cannot be used as input for calculating the distance, the lightweight neural network MobileNetV3 is used as the feature extraction module of the person retrieval algorithm. The MobileNetV3 network is optimized for mobile devices and embedded systems. It combines NAS (neural architecture search) and NetAdapt methods on the basis of MobileNetV2, and integrates the Squeeze-and-Excitation (SE) module and h-swish (hard-swish) activation function, which can improve the accuracy of the model while maintaining a low computational complexity.
[0116] After selecting the feature vector dimension as 2048, the target image is input into the MobileNetV3 network, and the neural network is used to extract the salient feature information of the people in the image. There is no need to manually define attributes. The neural network learns the relevant attribute information of interest by itself, encodes the target image through multi-layer convolution and pooling operations, and extracts low-level features (such as edges, textures, etc.) and high-level features (such as clothing patterns, gait, gender, etc.) in the image at one time to obtain a target feature vector of fixed dimension.
[0117] In step S402 of some embodiments, the principle of performing feature extraction processing on the image to be retrieved by the feature processing module is consistent with step S501 and will not be repeated here.
[0118] In some embodiments, the cross-shot personnel retrieval method based on cloud-edge collaboration may also include but is not limited to steps S601 to S603.
[0119] Step S601: perform an association query based on the camera device number of the image source camera to obtain an associated area.
[0120] Step S602: acquiring a video stream of a camera according to the associated area to obtain a tracking image.
[0121] Step S603: Update the search result according to the tracking image.
[0122] In this embodiment, the cloud server can also define camera association areas when binding cameras. The association area refers to an area where the monitoring ranges of a group of cameras have spatial overlap or path continuity. It is divided based on the physical position relationship between the cameras and the logic of the monitoring range to optimize the scope and efficiency of cross-lens personnel retrieval.
[0123] By defining associated regions, after a global search, when a camera belonging to a specific associated region reports having successfully retrieved a specific target person, a related query is performed based on the camera ID of the image source camera to obtain the associated region. Based on the associated region, cross-lens person search tasks for cameras not in the associated region are canceled, and the video streams of all cameras in the associated region are obtained to obtain tracking images. The tracking images are then used to update the search results, making the entire person search process more flexible.
[0124] Alternatively, if a camera within a specific association area reports that a target person has been lost in a blind spot, the association area prioritizes searching for other cameras with a coherent path to that camera. Association areas don't rely on the target's continuous motion trajectory, so there's no need for the camera's field of view to be coherent and continuous in time. The target person can be re-identified in surveillance images from other cameras with coherent paths, updating the search results.
[0125] In some embodiments, step S601 may include but is not limited to steps S701 to S703.
[0126] Step S701: determining a first area according to a monitoring overlap range based on the camera device number of the image source camera.
[0127] Step S702 : determining a second area based on the camera device number of the image source camera and path continuity.
[0128] Step S703: Obtain an associated area according to the first area and the second area.
[0129] In this embodiment, if the surveillance images of one camera Camera A overlap with those of another camera Camera B, it can be defined that Camera A and Camera B are associated. The surveillance images of other surrounding cameras can be queried based on the camera device number of the image source camera, and all cameras that overlap with the surveillance images of the image source camera can be determined as the first area according to the monitoring overlap range.
[0130] Alternatively, if the walkable roads in the monitoring areas of cameras Camera A and Camera B are connected, it is also possible to define an association between Camera A and Camera B. Based on the device number of the image source camera, the walkable roads monitored by other surrounding cameras are queried, and the camera connected to the walkable road monitored by the image source camera is determined as the second area based on the path continuity.
[0131] The first area and the second area are integrated to obtain all associated areas with associated relationships, so that when a cross-lens search task is issued to the cloud, a certain associated area can be selected for personnel search.
[0132] In some embodiments, in scenarios where the number of cameras is small and centrally distributed, or the camera network architecture has computing capabilities, no additional edge devices are required. Therefore, the cross-lens personnel retrieval method provided in the embodiments of the present application can be deployed on the network camera architecture. The camera network architecture can implement the cross-lens personnel retrieval method with reference to the above embodiments and will not be repeated here.
[0133] In this embodiment, a suitable deployment mode can be selected according to actual needs to meet the performance and security requirements in different scenarios.
[0134] The following describes and explains the solution of the embodiment of the present invention in detail with reference to specific application examples.
[0135] Reference Figure 6 , Figure 6 This is a framework diagram of a cross-lens personnel retrieval method based on cloud-edge collaboration provided by another embodiment of the present application.
[0136] First, define the neural network framework for the person detection algorithm. The YOLOv8 backbone network uses CSPDarknet as the feature extraction network, primarily consisting of 3x3 convolutional layers and batch normalization layers, with SiLu as the activation function. The feature fusion enhancement module (Neck) uses FPN as the upsampling module and PAN as the downsampling module. The detection head is divided into a classification branch and a bounding box regression branch.
[0137] We manually collect footage from specific surveillance cameras, selecting multiple cameras in different areas and collecting footage containing people during specific time periods: morning, noon, afternoon, and evening. After collecting a certain amount of footage, we hand it over to annotators to annotate the people in the footage with bounding boxes. The resulting data format is [person classification number id, target center point x-coordinate, target center point y-coordinate, bounding box width w, bounding box height h].
[0138] After obtaining the labeled data, image enhancement is performed on the image, including proportional scaling, pixel padding, random color gamut conversion, random flipping, etc. The image and annotation information are input into the neural network, and the neural network parameters are updated by backpropagation by calculating the classification loss, bounding box CIOU loss, and DFL distribution loss between the neural network prediction value and the annotation value.
[0139] After a certain number of training iterations, if the loss value reaches the expected target, it means that the training of the personnel detection algorithm is completed, and the corresponding parameter model is derived as the model of the personnel detection algorithm.
[0140] The personnel retrieval algorithm includes a feature extraction module, a feature vector library, and a vector distance calculation module. The lightweight neural network MobileNetV3 is used as the feature extraction network of the personnel retrieval algorithm, and the feature vector dimension is selected as 2048 to facilitate calculation by the vector distance calculation module.
[0141] Add a camera to the edge device in the corresponding area. The cross-camera personnel search method provided by the embodiments of this application is deployed on this edge device. Enter the corresponding camera information, including but not limited to the camera device IP, camera manufacturer, camera login username, camera login password, and other information. After adding the camera, you can check whether the camera is online on the edge device. If it is not online, you need to check the problem.
[0142] On the cloud side, add the edge device that has been bound to the camera and enter its device ID, device serial number, and network information. After successfully adding it, check whether it is online. If it is not online, you need to investigate the problem.
[0143] The user opens a browser on any device that can access the cloud server, enters the platform URL to log in, and after logging in to the remote access platform, clicks on the corresponding camera display page to start the personnel detection algorithm. When the user needs to start a personnel search task, he can choose to upload the target image himself, or select the target image that has been alarmed in the historical alarm record of the platform to issue the search task, such as Figure 6 As shown, select the target to be searched using a custom frame. If you need to enable a specific associated area for cross-lens personnel search, you can select the search area when creating the task. This can be a specific area or global. Multiple selections are supported for specific areas.
[0144] It's understandable that if the user selects a global search area, after a camera in a specific area finds a target, the user can choose to change the search area to a specific area in the browser. If the user selects a specific area for person search, the cloud server will send a request to cancel the cross-camera person search to the edge device outside the search area, improving search efficiency.
[0145] After a search task is issued, the cloud server stores the image and the corresponding query ID in a backend database and sends the task information to the edge device associated with the corresponding camera, awaiting the edge device's search results. It should be understood that if a search task includes a search area, the task will only be sent to edge devices within that specific search area.
[0146] After receiving a search task, the edge device parses the target image and query ID. The target image is then fed into the feature extraction module of the person search algorithm. After encoding, the module generates a target feature vector and stores it in a feature vector library.
[0147] At the same time, the edge device reads the bound camera image in real time, scales, pads, and normalizes the image, and then feeds it into the person detection algorithm to obtain the coordinate information of the person in the frame (i.e., the predicted bounding box). Using the person coordinate information and the image, the pixel area containing the person is intercepted and fed into the feature extraction module of the person retrieval algorithm to encode it into a feature vector, which then generates the corresponding query feature vector.
[0148] After L2 regularization, the target feature vector and the query feature vector are calculated. The Euclidean distance between the vectors is then determined to be less than a set threshold to determine whether the person is the target. If so, the target has been retrieved; otherwise, the search continues.
[0149] If the target is retrieved, the corresponding image and camera information will be structured into alarm information and sent to the cloud server. Users can view the search results on the browser side and perform operations such as viewing detailed content, confirming alarms, and historical queries.
[0150] In summary, the cross-lens personnel retrieval method based on cloud-edge collaboration provided in the embodiment of the present application is mainly composed of a cloud-edge collaboration architecture, a personnel detection algorithm, and a personnel retrieval algorithm, which can achieve cross-lens personnel retrieval requirements under limited computing power and device resources.
[0151] On the premise that the user provides the target image to be retrieved or selects a specific target from the warning image output by the personnel detection algorithm, the personnel detection task is performed from the video images of multiple cameras in a limited area, and the output detection results and the target to be retrieved are input into the image retrieval algorithm, and the latest location information of the target that matches the target to be retrieved is output.
[0152] By employing person retrieval technology and leveraging feature vector matching to achieve cross-lens retrieval, the system doesn't rely on field of view continuity. Even if a target is lost in a blind spot, it can still be relocated through person retrieval. Furthermore, image retrieval relies on matrix multiplication, which is highly efficient and computationally demanding. This significantly reduces device performance requirements compared to multi-target tracking.
[0153] At the same time, the embodiment of the present application makes full use of the computing power of cloud devices and edge devices, and delegates the target detection algorithm and personnel retrieval algorithm to the edge side for calculation. The cloud side only needs to complete the distribution of personnel retrieval tasks. To subsequently expand the number of camera routes, only edge devices need to be added on the end side, and there is no need to modify the cloud side.
[0154] Reference Figure 7 The embodiment of the present application further provides a cross-lens personnel retrieval system based on cloud-edge collaboration, which can implement the above-mentioned cross-lens personnel retrieval method based on cloud-edge collaboration. The system is applied to an edge device, and the edge device is bound to several cameras. The system includes:
[0155] The first module is used to obtain the retrieval task issued by the cloud. The retrieval task includes the target image and the retrieval area.
[0156] The second module is used to obtain the video stream of the camera according to the search area to obtain the monitoring image.
[0157] The third module is used to perform personnel detection processing on the monitoring image to obtain the image to be retrieved.
[0158] The fourth module is used to perform person retrieval processing based on the target image and the image to be retrieved to obtain the retrieval result.
[0159] The fifth module is used to convert the search results into alarm information and send it to the cloud.
[0160] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0161] The present application also provides an electronic device comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned cross-camera personnel search method based on cloud-edge collaboration. The electronic device can be any smart terminal, including a tablet computer and an in-vehicle computer.
[0162] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0163] Reference Figure 8 , Figure 8 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0164] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0165] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the cross-lens personnel retrieval method based on cloud-edge collaboration in the embodiments of this application.
[0166] The input / output interface 903 is used to implement information input and output.
[0167] The communication interface 904 is used to realize communication interaction between this device and other devices. Communication can be realized through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0168] The bus 905 transmits information between various components of the device (eg, the processor 901 , the memory 902 , the input / output interface 903 , and the communication interface 904 ).
[0169] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0170] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned cross-shot personnel retrieval method based on cloud-edge collaboration.
[0171] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0172] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0173] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0174] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0175] The system embodiment described above is merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0176] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0177] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0178] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A cross-shot personnel retrieval method based on cloud-edge collaboration, characterized in that: The method is applied to an edge device, wherein the edge device is bound to several cameras, and the method comprises the following steps: Obtaining a search task issued by the cloud, wherein the search task includes a target image and a search area; Acquire the video stream of the camera according to the search area to obtain a monitoring image; Performing personnel detection processing on the surveillance image to obtain an image to be retrieved; Performing a person search process based on the target image and the image to be searched to obtain a search result; The search results are converted into warning information and sent to the cloud.
2. The method according to claim 1, characterized in that The step of performing personnel detection processing on the surveillance image to obtain the image to be retrieved comprises the following steps: Performing target detection processing on the surveillance image to obtain category information and a predicted bounding box of the detected target; Screening the detection targets according to the category information to obtain targets to be retrieved; According to the predicted bounding box corresponding to the target to be retrieved, image interception processing is performed on the monitoring image to obtain the image to be retrieved.
3. The method according to claim 2, characterized in that The target detection processing is performed on the monitoring image to obtain category information and a predicted bounding box of the detected target, including the following steps: Scaling the monitoring image to obtain a proportional image; Filling the proportional image with a fixed value to obtain a filled image; performing normalization processing on the filled image to obtain a normalized image; The normalized image is input into the target detection model for target detection processing to obtain category information and predicted bounding box of the detected target.
4. The method according to claim 1, wherein The process of performing a person search process based on the target image and the image to be searched to obtain a search result includes the following steps: Performing feature extraction processing on the target image to obtain a target feature vector; Performing feature extraction processing on the image to be retrieved to obtain a query feature vector; Determining a vector distance based on the target feature vector and the query feature vector, and determining whether the vector distance is less than a preset threshold to obtain a determination result; If the judgment result is that the vector distance is less than the preset threshold, the image to be retrieved corresponding to the query feature vector and the camera device number of the image source camera are used as the retrieval result.
5. The method according to claim 4, characterized in that The step of performing feature extraction on the target image to obtain a target feature vector comprises the following steps: The target image is input into a lightweight neural network for encoding processing to obtain a target feature vector of fixed dimension.
6. The method according to claim 4, characterized in that The cross-shot personnel retrieval method based on cloud-edge collaboration also includes the following steps: Performing an associated query based on the camera device number of the image source camera to obtain an associated area; Acquire the video stream of the camera according to the associated area to obtain a tracking image; The retrieval result is updated according to the tracking image.
7. The method according to claim 6, characterized in that The method of performing an associated query based on the camera device number of the image source camera to obtain an associated area includes the following steps: Determining a first area according to a monitoring overlap range based on a camera device number of a camera from which the image originates; Determining a second area based on the camera device number of the image source camera and path continuity; An associated area is obtained according to the first area and the second area.
8. A cross-lens personnel retrieval system based on cloud-edge collaboration, characterized by: The system is applied to an edge device, which is bound to several cameras. The system includes: The first module is used to obtain a search task issued by the cloud, wherein the search task includes a target image and a search area; The second module is used to obtain the video stream of the camera according to the search area to obtain a monitoring image; The third module is used to perform personnel detection processing on the monitoring image to obtain an image to be retrieved; The fourth module is used to perform a person search process based on the target image and the image to be searched to obtain a search result; The fifth module is used to convert the search results into alarm information and send it to the cloud.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Cross-border tracking method for personnel in park
CN112365522A
Data processing system and method based on cloud side end
CN114067235A
Target tracking method, system and device based on cloud edge collaboration and medium
CN114241002A
Cross-camera target object tracking method and device, equipment and storage medium
CN114821430A
Suspicious person tracking method, device and equipment based on cloud edge collaboration and medium
CN117649428A
Cited By
End-side cloud collaborative big data analysis method and system
CN121935410A
An edge-cloud collaborative big data analysis method and system
CN121935410B