A method and apparatus for supervision, device, program product, storage medium

By combining multiple fixed cameras and optoelectronic systems in low-altitude surveillance, using the YOLO network to identify targets and adaptively tracking them through the optoelectronic system, the problem of supervising non-cooperative drones has been solved, achieving efficient and accurate monitoring in complex environments.

CN122120407APending Publication Date: 2026-05-29CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD
Filing Date
2024-11-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing low-altitude surveillance technologies are mainly designed for cooperative drones, and cannot effectively monitor non-cooperative drones or other targets that lack communication capabilities. This results in insufficient surveillance security and tracking accuracy, increasing safety risks in low-altitude airspace.

Method used

The solution employs a combination of multiple fixed cameras and photoelectric systems. The fixed cameras are used for large-area monitoring, and the YOLO network is used to identify targets and select points of interest. The photoelectric system performs adaptive tracking to achieve data association and precise supervision.

Benefits of technology

In situations where base stations cannot be installed or conditions are unsuitable, efficient and accurate monitoring of non-cooperative targets has been achieved, improving the safety and regulatory efficiency of low-altitude airspace.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120407A_ABST
    Figure CN122120407A_ABST
Patent Text Reader

Abstract

The application discloses a kind of supervision method and device, equipment, program product, storage medium, the method comprises: obtaining the first video set of first camera set;Wherein, the video in the first video set corresponds with the camera in the first camera set;According to the first video set determines target interest content;Obtain the first information corresponding to the target interest content;The first information includes the first sub-information of the target interest content and the second sub-information of target camera;The target camera is the camera corresponding to the video where the target interest content is located;According to the first information determines data association result, the data association result is the result of the data association of the target camera and first system;The association result is used to the supervision of the target interest content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of vertical industries, and in particular to a regulatory method and apparatus, equipment, program product, and storage medium. Background Technology

[0002] Sensor-based systems play a crucial role in low-altitude airspace surveillance. First, the sensing system can detect and identify various targets within low-altitude airspace, including drones, aircraft, and birds, providing regulators with real-time target information. Second, the sensing system provides situational awareness of the low-altitude airspace, monitoring parameters such as target position, speed, and altitude, helping regulators understand dynamic changes within the airspace. Furthermore, the sensing system can monitor and identify violations, such as overflight and exceeding altitude limits, providing regulators with a basis for timely action. Through data collection, analysis, and prediction, the sensing system provides regulators with more comprehensive information support, helping them effectively manage and control low-altitude airspace and ensure aviation safety. Therefore, improving the efficiency and safety of low-altitude airspace surveillance is an urgent issue to be addressed. Summary of the Invention

[0003] To address the aforementioned technical problems, this application provides a regulatory method, apparatus, equipment, program product, and storage medium.

[0004] The regulatory approaches provided in this application include:

[0005] Obtain the first video set of the first camera set; wherein, the videos in the first video set correspond to the cameras in the first camera set;

[0006] Determine target interest content based on the first video set;

[0007] Obtain the first information corresponding to the target interest content; the first information includes the first sub-information of the target interest content and the second sub-information of the target camera; the target camera is the camera corresponding to the video where the target interest content is located.

[0008] The data association result is determined based on the first information. The data association result is the result of the data association between the target camera and the first system. The data association result is used for the supervision of the target's content of interest.

[0009] The monitoring device provided in this application includes:

[0010] The acquisition unit is used to acquire a first video set of the first camera set; wherein the videos in the first video set correspond to the cameras in the first camera set;

[0011] The determining unit is used to determine the target interest content based on the first video set;

[0012] The acquisition unit is further configured to acquire first information corresponding to the target interest content; the first information includes first sub-information of the target interest content and second sub-information of the target camera; the target camera is the camera corresponding to the video where the target interest content is located;

[0013] The determining unit is used to determine the data association result based on the first information. The data association result is the result of data association between the target camera and the first system. The data association result is used for monitoring the target's content of interest.

[0014] The monitoring device provided in this application includes a processor and a memory, the memory being used to store computer programs, and the processor being used to call and run the computer programs stored in the memory to perform the aforementioned monitoring method.

[0015] This application provides a computer program product, comprising: a computer program that implements the above-described method when executed by a processor.

[0016] The computer-readable storage medium provided in this application is used to store a computer program that causes a computer to perform the above-described method.

[0017] In the technical solution of this application, a first video set of a first camera set is obtained; wherein, the videos in the first video set correspond to the cameras in the first camera set; a target interest content is determined based on the first video set; first information corresponding to the target interest content is obtained; the first information includes first sub-information of the target interest content and second sub-information of the target camera; the target camera is the camera corresponding to the video where the target interest content is located; a data association result is determined based on the first information, the data association result being the result of data association between the target camera and the first system; the data association result is used for monitoring the target interest content. Thus, by determining the target through the shooting results of the first camera set and the recognition results of the first network, and by realizing real-time monitoring of the monitored area based on the data association result between the second camera and the target camera, the efficiency and security of monitoring are improved. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the regulatory method provided in the embodiments of this application. Figure 1 ;

[0019] Figure 2 This is a flowchart illustrating the regulatory method provided in the embodiments of this application. Figure 2 ;

[0020] Figure 3 This is a schematic diagram of the optical imaging principle provided in the embodiments of this application;

[0021] Figure 4This is a schematic diagram of two fixed cameras provided in an embodiment of this application;

[0022] Figure 5 This is a schematic diagram of a video screenshot of a photoelectric system tracking provided in an embodiment of this application;

[0023] Figure 6 This is a schematic diagram of the structural composition of the monitoring device provided in the embodiments of this application;

[0024] Figure 7 This is a schematic structural diagram of a monitoring device provided in an embodiment of this application;

[0025] Figure 8 This is a schematic structural diagram of the chip according to an embodiment of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0027] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0028] It should also be noted that the terms "first, second, and third" used in the embodiments of this application are only used to distinguish similar objects and do not represent a specific order of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0029] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship. It should also be understood that "instruction" mentioned in the embodiments of this application can be a direct instruction, an indirect instruction, or an indication of a related relationship. For example, A instructing B can mean that A directly instructs B, for example, B can be obtained through A; it can also mean that A indirectly instructs B, for example, A instructs C, B can be obtained through C; or it can mean that there is a related relationship between A and B. It should also be understood that "correspondence" mentioned in the embodiments of this application can indicate a direct or indirect correspondence between two objects, or an related relationship between two objects, or a relationship of instruction and being instructed, configuration and being configured, etc.

[0030] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and they all fall within the protection scope of the embodiments of this application.

[0031] Sensor-based systems play a crucial role in low-altitude airspace surveillance. First, the sensing system can detect and identify various targets within low-altitude airspace, including drones, aircraft, and birds, providing regulators with real-time target information. Second, the sensing system provides situational awareness of low-altitude airspace, monitoring parameters such as target position, speed, and altitude, helping regulators understand dynamic changes within the airspace. Furthermore, the sensing system can monitor and identify violations, such as overflight and exceeding altitude limits, providing regulators with a basis for timely action. Through data collection, analysis, and prediction, the sensing system provides regulators with more comprehensive information support, helping them effectively manage and control low-altitude airspace and ensure aviation safety.

[0032] Current methods for effectively regulating low-altitude airspace primarily focus on cooperative drones. These methods typically rely on drones actively reporting their position and attitude information for tracking. However, this approach has significant limitations in regulating non-cooperative drones or other target objects. Non-cooperative drones do not proactively provide attitude information, making them difficult to monitor and manage effectively using traditional methods. To address this issue, new regulatory methods need to be developed that can accurately track non-cooperative drones or other target objects without requiring attitude information from the regulated entity, thereby ensuring the safety and order of low-altitude airspace.

[0033] Existing low-altitude surveillance technologies primarily target cooperative drones, tracking them by receiving position and attitude information proactively reported by the drones. However, this method has significant limitations when dealing with non-cooperative drones or other targets lacking communication capabilities.

[0034] One implementation proposes a solution to the monitoring problem of non-cooperative targets. This patent relies on the collaborative operation of three types of sensors: the first is a 5G-A device for sensing target position and velocity information; the second is multiple fixed cameras; and the third is an optoelectronic system (which can be considered a freely rotating camera device). However, in some scenarios, the application of this solution is limited due to the inability to install base stations or unsuitable installation conditions.

[0035] Because real-time pose information cannot be obtained from the target, the security of low-altitude target surveillance is low, and tracking accuracy cannot be guaranteed. Existing technologies are inadequate in detecting and tracking non-cooperative targets. This results in an inability to effectively deal with unauthorized or non-cooperative aircraft in low-altitude airspace surveillance, increasing safety risks. Therefore, improving the efficiency and security of surveillance has become a problem that needs to be considered. To this end, the following technical solutions based on embodiments of this application are proposed.

[0036] To facilitate understanding of the technical solutions of the embodiments of this application, the technical solutions of this application are described in detail below through specific embodiments. The above-mentioned related technologies are optional solutions and can be arbitrarily combined with the technical solutions of the embodiments of this application, all of which fall within the protection scope of the embodiments of this application. The embodiments of this application include at least some of the following contents.

[0037] Figure 1 This is a flowchart illustrating the regulatory method provided in the embodiments of this application. Figure 1 ,like Figure 1 As shown, the regulatory method includes the following steps:

[0038] Step 101: Obtain the first video set of the first camera set; wherein, the videos in the first video set correspond to the cameras in the first camera set.

[0039] In some implementations, the cameras in the first camera set are fixed cameras, and the videos in the first video set correspond to the cameras in the first camera set. Multiple fixed cameras are installed within the monitored area, and their positions are adjusted so that their total field of view covers the entire monitored area. These cameras all belong to the first camera set.

[0040] In some implementations, obtaining a first video set from a first set of cameras includes: obtaining an initial video set from the first set of cameras; the videos in the initial video set are captured by cameras in the first set of cameras; performing target detection on the videos in the initial video set through a first network to obtain second information; wherein the second information includes target-related information; and determining the first video set based on the second information and the initial video set.

[0041] In some implementations, the cameras in the first set of cameras capture videos of the areas monitored by the respective cameras, and the videos captured by these cameras all belong to the initial video set.

[0042] In some implementations, the first network is a YOLO network, which divides the image into a fixed-size grid and performs object detection in each grid cell. Therefore, before performing object detection on videos in the initial video set based on the first network, a YOLO network corresponding to the initial video set needs to be trained. Specifically, the YOLO network is based on a deep neural network and can identify different categories of targets by training on a large number of labeled images. The training set of the YOLO network is related to the content to be monitored. For example, when drones are to be monitored within an area, the network can identify different types of drones. Drones with significantly different sizes are also labeled as different categories in the training set. Specifically, the training set of the first network is related to the monitored area and the content of the monitored area; this application does not impose specific limitations on this.

[0043] In some implementations, each camera in the first camera set runs a YOLO target detection system in the background. After the cameras in the first camera set capture the initial video, each initial video needs to pass through the YOLO network. The network divides the image into a grid of fixed size and performs target detection in each grid cell to obtain second information, wherein the second information includes target-related information.

[0044] In some implementations, the second information includes the bounding box coordinates of the target in the video and the class probability of the target. For example, after the YOLO network identifies a target, for each target, the YOLO algorithm outputs a bounding box and its coordinates, as well as the class probability of the target.

[0045] In some implementations, the first video set includes an initial video set and second information. That is, the initial video set after superimposing the output information of the first network is the first video set. It can be understood that the first video set includes the initial video set, as well as the target bounding boxes, the coordinates of the target bounding boxes, and the category probabilities after the first network identifies the targets in the initial video set.

[0046] In some implementations, the initial video information, after being overlaid with the output information of the first network, is streamed to the front-end interface for display. It is understood that the display interface includes videos from the initial video set, as well as the bounding boxes of the targets identified by the first network in those videos, the coordinates of the bounding boxes, and the probability of the identified target's category. For example, the identified target boxes are red rectangles, and the category probability of the target is displayed in the upper left corner of the target box. This display area for the category probability is transparent for easy viewing of the video content. Each target box can also be labeled with a corresponding number. The specific settings of the target boxes and the position of the identified target's category probability displayed in the video can be determined according to the actual situation; this application does not impose specific limitations on this.

[0047] Step 102: Determine the target interest content based on the first video set.

[0048] In some implementations, the content of interest can be the content that needs to be monitored, and the target content of interest is the selected content of interest. The target content of interest includes targets. For example, when it is necessary to monitor the situation of drones and other flying objects in the monitored area, the selected target content of interest can be content of interest that includes drones, and the drones are the targets in the target content of interest.

[0049] In some implementations, determining the target interest content based on the first video set includes: determining the target interest content based on the second information and the initial video set.

[0050] In some implementations, the first network performs target recognition on the initial video set to obtain the bounding box coordinates and category probabilities of the targets. Based on the target bounding box coordinates and category probabilities, it selects the target content of interest. For example, when the category probabilities of the target recognition include targets that need to be monitored, the bounding box of the target can be selected, and the video content information of the selected bounding box can also be obtained.

[0051] In some implementations, during the front-end monitoring process, the supervisor observes the situation within the monitored area in real time through the monitoring interface. When a target of interest is detected, the supervisor can select it by clicking the target's bounding box with the mouse, or by clicking the area inside the bounding box, or by selecting the target based on the corresponding number on the bounding box. When the screen is a touchscreen, the supervisor can also directly click the target box on the screen. The specific method of selecting the target can be determined according to the actual situation, and this application does not impose specific limitations on it.

[0052] In some implementations, specific targets can be set, and when a specific target is detected, it can be automatically selected as the target point of interest. For example, when a camera in the first camera set detects a drone, that target can be automatically selected as a point of interest.

[0053] Step 103: Obtain the first information corresponding to the target interest content; the first information includes the first sub-information of the target interest content and the second sub-information of the target camera; the target camera is the camera corresponding to the video where the target interest content is located.

[0054] In some implementations, first information corresponding to the target content of interest is obtained. This first information includes first sub-information and second sub-information. The first sub-information is information related to the target content of interest, and the second sub-information is information related to the target camera. The target camera is the camera corresponding to the video containing the selected target content of interest. For example, if the target content of interest in the video captured by camera 1 is selected, then camera 1 is the target camera.

[0055] In some implementations, obtaining the first information corresponding to the target interest content includes: determining the first sub-information of the target interest content based on the second information and the initial video set; determining the ID information of the target camera based on the target interest content; and determining the second sub-information of the target camera based on the ID information.

[0056] In some implementations, first sub-information is determined based on the coordinates of the bounding box of the target recognition, the category probability, and the initial video set; the ID information of the target camera is determined based on the target's interest content; and second sub-information is determined based on the ID information. It is understandable that, after determining the target's interest content, the ID information of the target camera can also be directly obtained, and then the second sub-information of the target camera can be determined based on the ID information.

[0057] In some implementations, the relevant information of the first camera set is stored in a corresponding database. Each camera in the first camera set has a corresponding ID and other relevant information. Therefore, after determining the ID information of the target camera based on the target interest content, the second sub-information of the corresponding target camera is queried in the database based on the ID information.

[0058] In some implementations, the first sub-information includes one or more of the following: bounding box information of the target interest content, target category probability information of the target interest content, and image information of the target interest content.

[0059] In some implementations, the first sub-information is information related to the target interest content, including the bounding box information of the target interest content, such as bounding box coordinates and bounding box display mode; the image information of the target interest content is the video information corresponding to the target interest content.

[0060] In some implementations, the second sub-information includes one or more of the following: the target camera's identity ID information, the target camera's intrinsic parameter information, the target camera's spatial location information, and the target camera's installation posture information.

[0061] In some implementations, the target camera's intrinsic parameters, spatial location, and installation attitude are retrieved from the database based on the target camera's ID information.

[0062] In some implementations, the intrinsic parameter information of all cameras is represented in the form of a matrix. For example, the intrinsic parameter information matrix of the cameras is called a Matrix. id As shown in formula (1):

[0063]

[0064] Among them, f x and f y c represents the focal length of the camera along the x-axis and y-axis in the image plane, respectively. x and c y These represent the coordinates of the image center point along the x-axis and y-axis, respectively.

[0065] In some implementations, the spatial location information P of all cameras id As shown in formula (2):

[0066] P id =[lat id lng id h id (2);

[0067] Among them, lat id , lng id h id These represent the latitude, longitude, and altitude information of the camera ID, respectively.

[0068] In some implementations, the mounting orientation information of all cameras is... Wherein θ id It's the camera mounting angle. This refers to the camera mounting tilt angle. The camera mounting tilt angle is the angle between the camera coordinate system's Z-axis and the horizontal plane xoy of the reference coordinate system, with the top angle being positive. The camera mounting offset angle is the angle between the projection of the camera coordinate system's Z-axis onto the horizontal plane xoy of the reference coordinate system and the ox-axis, with the clockwise direction being positive.

[0069] Step 104: Determine the data association result based on the first information. The data association result is the result of data association between the target camera and the first system. The data association result is used for monitoring the target's content of interest.

[0070] In some implementations, the first system is an optoelectronic system, which includes a second set of cameras and a gimbal. The cameras in the second set of cameras are rotatable, unlike the fixed cameras in the first set of cameras. The gimbal can control the cameras in the second set to rotate at a desired angle. After acquiring the first information of the target interest content, the information is sent to the gimbal. After acquiring the first information, the gimbal associates the data with the target camera and initiates adaptive tracking of the target in the target interest content based on the association result.

[0071] In some implementations, the optoelectronic system can be replaced with other devices that include a servo system and a camera.

[0072] In some implementations, determining the data association result based on the first information includes: determining the first location information of the target interest content based on the first information; and determining the data association result based on the first location information through a second network.

[0073] In some implementations, the first location information is the coordinate range of the target interest content in the reference coordinate system of the first system, that is, the coordinate range in the reference coordinate system of the photoelectric system. The second network determines the data association result based on the first location information.

[0074] In some implementations, determining the first location information of the target interest content based on the first information includes: determining the first height of the target based on the target category probability information of the target interest content; determining the camera's elevation and horizontal frame angles based on the bounding box information of the target interest content and the intrinsic parameter information of the target camera; determining the camera's elevation and horizontal line-of-sight angles based on the camera's elevation and horizontal frame angles and the target camera's mounting posture information; determining the second location information based on the camera's elevation and horizontal line-of-sight angles and the first height; and determining the first location information based on the second location information and the spatial location information of the target camera.

[0075] In some implementations, based on the target category probability of the target interest content, the corresponding real-world physical height of the target, i.e., the first height, is queried from the database and represented by l. t It indicates that the unit is meters.

[0076] In some implementations, the bounding box information of the target interest content includes the coordinates of the bounding box, which can be specifically represented as: [x t ,y t [,w,h], where x t y t Here are the coordinates of the center point of the bounding box, and w and h are the width and height of the bounding box. Camera intrinsic parameters include the camera's focal length f along the x-axis on the image plane. x and the focal length f in the y-axis direction yThe camera elevation frame angle refers to the angle between the optical axis (connected to the origin of the camera coordinate system and the target point) and the Z-axis of the camera coordinate system in the vertical plane, with the upward direction of the optical axis being positive. The camera horizontal frame angle refers to the angle between the optical axis (connected to the origin of the camera coordinate system and the target point) and the Z-axis of the camera coordinate system in the reference coordinate system xoy plane, with the clockwise direction being positive. Based on the bounding box information of the target interest content and the intrinsic parameters of the target camera, the camera elevation frame angle q of the target interest content from the viewpoint of the target camera is calculated. α and horizontal frame angle q β The specific calculation methods are shown in formulas (3) and (4):

[0077] q a =atan2(-y t +h / 2,f y (3);

[0078] q β =atan2(x t -w / 2,f x (4);

[0079] Where 'a' is a coefficient, which can be set according to the actual situation, and this application does not impose any specific restrictions on it.

[0080] In some implementations, the elevation and slenderness angle q of the target in the target interest content as viewed from the target camera is used. α and horizontal frame angle q β In addition to the installation pose information of the target camera, the elevation angle q of the selected point of interest target under the view of the target camera is calculated. f and the camera's horizontal line of sight q h The specific calculation methods are shown in formulas (5) and (6):

[0081] q f =q a +θ id (5);

[0082]

[0083] Where, θ id It's the camera mounting angle. It's the camera mounting angle.

[0084] In some implementations, the camera's elevation angle q, based on the aforementioned target interest content, is determined from the target camera's perspective. f and the camera's horizontal line of sight q h And the physical height of the target interest content in the real world. tThe coordinate range of the target within the selected content of interest is calculated in the reference coordinate system corresponding to the target camera. The specific calculation method is as follows:

[0085] First, assume there is a percentage error p between the actual size of the target and the value recorded in the database. Therefore, the actual size of the target, l r The range is l t •(1±p). The three-dimensional position (x, y) of the target in the target content of interest in the camera reference coordinate system. p ,y p ,z p It should be noted that, here, the commonly used northeast-northeast coordinate system is used as the camera reference coordinate system, that is, the due north direction is set as the x-axis, the due east direction is set as the y-axis, and the z-axis is determined according to the right-hand rule. The origin of the coordinate system is defined as the installation position of the camera. The camera coordinate system refers to the camera optical center as the origin. The x-axis is parallel to the x-axis of the pixel coordinate system, and is positive to the right; the y-axis is parallel to the y-axis of the pixel coordinate system, and is positive downwards; the z-axis is determined by the right-hand rule. The distance from the camera lens to the target is calculated as shown in formula (7):

[0086]

[0087] Where z is the distance from the camera lens to the target, h is the height of the bounding box, and l t f represents the physical height of the target in real time. y This represents the focal length of the camera along the y-axis in the image plane.

[0088] The three-dimensional position (x, y) of the point of interest target in the camera reference coordinate system. p ,y p ,z p The calculation methods are as follows: (8), (9) and (10):

[0089]

[0090]

[0091] Because of l r ∈[l t ·(1-p),l t Substituting (1+p)] into formulas (8), (9), and (10), we can obtain the range of the point of interest in the reference coordinate system as: point A = [x pmin ,y pmin ,z pmin ] T B = [x pmax ,y pmax ,z pmax ] T .

[0092] In some implementations, the coordinate range of the target point of interest in the reference coordinate system of the photoelectric system is calculated based on the range of the target in the target interest content within the reference coordinate system corresponding to the target camera and the spatial position information of the target camera. For example, firstly, based on the coordinates of points A and B and the spatial position information of the target camera (i.e., the latitude, longitude, and altitude information of the target camera), the longitude, latitude, and altitude information of A and B are obtained respectively. Secondly, using the geodetic projection method and the longitude, latitude, and altitude information of the known installation location of the photoelectric system, the coordinates A′ and B′ of points A and B in the reference coordinate system of the photoelectric system are obtained.

[0093] In some implementations, the second network is a SiamRPN network.

[0094] In some implementations, the data association result is determined by a second network based on the first location information, including: using the SiamRPN network to perform data association between the target camera and the optoelectronic system based on the coordinate range of the target in the target interest content in the reference coordinate system of the optoelectronic system.

[0095] In some implementations, the target within the target interest content must lie between lines A′ and B′. In this case, the photoelectric system scans between A′ and B′, and simultaneously activates the SiamRPN network to correlate the data. Specifically, the SiamRPN network application method is as follows:

[0096] 1) The target bounding box initialized by SiamRPN is selected as the image content of the target interest.

[0097] 2) Extract the middle regions of all scanned images between lines A′ and B′ and perform feature extraction using a Siamese network. The Siamese network encodes the features of the target bounding box and the middle regions of all scanned images, enabling comparison in the feature space. Through correlation operations with the target template, the similarity between the candidate region and the target template is calculated. This similarity value represents the possible location of the target; the final target location is determined by maximizing the similarity. This completes the data association process.

[0098] In some implementations, after obtaining the data association results, the target in the target interest content is tracked.

[0099] In some implementations, after the optoelectronic system completes data association, image tracking is performed using a traditional SiamRPN network. Control signals are determined based on the target's bounding box information to ensure the target's bounding box is centered within the viewfinder of the camera used to monitor the target in the second camera set, and that the bounding box's size remains constant within the viewfinder. For example, the SiamRPN network tracks the target's bounding box information [x...].b ,y b ,w b ,h b The data is transmitted to the photoelectric servo system, which uses the feedback information to generate control information to keep the target's bounding box centered on the screen and maintain its size within the screen. The control signal calculation is shown in formulas (11), (12), and (13).

[0100] pan = pid * (w / 2 - x) b (11);

[0101] tilt = pid * (h / 2 - y) b (12);

[0102] zoom = pid * (k * wh - w) b *h b (13);

[0103] Where w and h are the width and height of the image, pid is the PID controller, and k is the percentage of the desired target size in the entire image.

[0104] In some implementations, the images from the cameras in the second set of cameras are displayed at the front end.

[0105] The technical solution of this application embodiment involves obtaining a first video set from a first set of cameras; wherein the videos in the first video set correspond to the cameras in the first set of cameras; determining target content of interest based on the first video set; obtaining first information corresponding to the target content of interest; the first information includes first sub-information of the target content of interest and second sub-information of the target camera; the target camera is the camera corresponding to the video containing the target content of interest; determining a data association result based on the first information, the data association result being the result of data association between the target camera and the first system; the data association result is used for monitoring the target content of interest. Thus, by determining the target through the shooting results of the first set of cameras and the recognition results of the first network, and by realizing real-time monitoring of the monitored area based on the data association result between the first system and the target camera, the efficiency and security of monitoring are improved.

[0106] The technical solutions of the embodiments of this application are illustrated below with specific application examples.

[0107] With the development of aviation technology and the rapid growth of the drone market, low-altitude airspace regulation has become an increasingly important area. Specifically, the advantages of low-altitude airspace regulation include:

[0108] Safety Management: Low-altitude surveillance can improve flight safety and prevent drones from colliding with other aircraft, buildings, or people. By monitoring low-altitude areas, accidental collisions and aerial incidents can be reduced, ensuring safety both in the air and on the ground.

[0109] Air traffic management: Effective low-altitude regulation can promote the development of air traffic management, including managing the routes, altitudes and speeds of drones and other low-altitude vehicles in various environments such as cities, rural areas and industrial zones, as well as coordinating air traffic flow between different aircraft.

[0110] Drone service providers and manufacturers: For drone manufacturers and service providers, low-altitude regulation means increased demand for their products and services. Effective regulation can facilitate more commercial applications, including aerial photography, security patrols, agricultural monitoring, and logistics delivery.

[0111] Existing low-altitude airspace surveillance technologies primarily target cooperative drones, tracking them by receiving position and attitude information proactively reported by the drones. However, this method has significant limitations for non-cooperative drones or other targets lacking communication capabilities. Due to the inability to obtain real-time pose information from the target, the security of low-altitude target surveillance is low, and tracking accuracy cannot be guaranteed. Existing technologies are inadequate in detecting and tracking non-cooperative targets. This results in an inability to effectively address unauthorized or non-cooperative aircraft in low-altitude airspace surveillance, increasing safety risks. Therefore, there is an urgent need to develop new surveillance methods to solve the problem of real-time tracking and monitoring of non-cooperative targets, thereby improving the security and efficiency of low-altitude airspace surveillance.

[0112] Some technical solutions address the monitoring of non-cooperative targets. These solutions rely on the coordinated operation of three types of sensors: first, 5G-A devices for sensing target position and velocity; second, multiple fixed cameras; and third, an optoelectronic system (which can be considered a freely rotating camera device). However, in certain scenarios, the application of this solution is limited due to the inability to install base stations or unsuitable installation conditions. In such cases, a combination of multiple fixed cameras and an optoelectronic system is necessary to track non-cooperative targets, ensuring effective target monitoring even in areas lacking base station support.

[0113] This combined monitoring method uses multiple fixed cameras to cover the entire monitored area. Monitors select points of interest from the camera feeds and have them tracked by an optoelectronic system. While this method introduces the challenge of data correlation between the cameras and the optoelectronic system, this application proposes a novel solution to ensure that the systems can still work effectively together under these conditions to achieve accurate target monitoring.

[0114] This application embodiment designs a low-altitude surveillance solution based on multiple fixed cameras and an optoelectronic system. The method includes: first, placing multiple fixed cameras in the surveillance area, ensuring that the total field of view of the cameras covers the entire surveillance area, and pushing the images from each camera to the front end for display in real time; when an illegal intrusion object appears in the image displayed by the fixed camera, the supervisor uses a mouse to circle the intrusion object; after the optoelectronic system obtains the intrusion object information sent by the supervisor, it tracks the object and adaptively adjusts the focal length of the optoelectronic system.

[0115] Low-altitude airspace regulation primarily involves the management of airspace several hundred meters above the ground. Its core purpose is to ensure aviation safety, maintain public safety, and support urban development. With the increasing prevalence of drones and low-altitude aircraft, effective low-altitude airspace regulation can prevent collisions between drones and manned aircraft, reduce aviation accidents, prevent illegal flight activities, and enhance public safety.

[0116] This application addresses the technical need for monitoring non-cooperative targets in complex environments, particularly when base stations cannot be installed or the installation conditions are unsuitable. Traditional low-altitude surveillance methods primarily rely on pose information reported by cooperative drones for tracking; however, pose information for non-cooperative targets is unavailable, rendering traditional methods inapplicable. Therefore, this application proposes a monitoring scheme based on multiple fixed cameras and an optoelectronic system.

[0117] In practical applications, when the environment is unsuitable for installing base stations, the monitored area cannot utilize 5G-A to perceive target information. To address this issue, this application proposes a monitoring scheme combining multiple fixed cameras and an optoelectronic system. Specifically, multiple fixed cameras are deployed within the monitored area to provide comprehensive coverage and capture potential target objects. The optoelectronic system is installed in other fixed locations and features pan-tilt-zoom functionality, allowing for flexible adjustment of the viewing angle to continuously track the target. Furthermore, since the equipment cost of the optoelectronic system is significantly higher than that of fixed cameras, relying solely on the optoelectronic system for full coverage of the target area would significantly increase costs. Therefore, this application combines fixed cameras with an optoelectronic system, achieving an optimal balance between cost-effectiveness and monitoring performance.

[0118] In the overall system architecture, fixed cameras are primarily used to monitor a large area and capture all possible targets. The photoelectric system, on the other hand, performs precise tracking based on points of interest selected by the supervisor within the fixed camera's view. The photoelectric system can adjust its viewing angle in real time to keep the target centered in the frame, thus achieving high-precision target tracking.

[0119] Through this combination, the solution in this application embodiment can still effectively monitor non-cooperative targets and provide real-time tracking data even when base stations are missing or installation conditions are unsuitable. This not only improves the applicability and flexibility of the monitoring system but also provides reliable technical support for low-altitude surveillance in complex environments.

[0120] Figure 2 This is a flowchart illustrating the regulatory method provided in the embodiments of this application. Figure 2 ,like Figure 2 As shown in the figure, the low-altitude surveillance solution based on multiple fixed cameras and an optoelectronic system proposed in this application uses fixed cameras to cover the entire surveillance area. Each fixed camera runs a target detection algorithm at its backend and streams the synthesized detection result video to the front end. Surveillance personnel select points of interest at the front end, and the corresponding camera sends the relevant information of the points of interest to the optoelectronic system. The optoelectronic system performs data association on the points of interest and then performs tracking and adaptive zoom. A detailed description follows:

[0121] Fixed camera:

[0122] Multiple fixed cameras are installed within the monitored area, and their placement is adjusted to ensure that the total field of view of all fixed cameras covers the entire monitored area. Each camera runs a YOLO target detection system in the background. The YOLO algorithm is based on a deep neural network and, through training on a large number of labeled images, has learned to identify different categories of targets, including different types of drones (drones with significant size differences are labeled as different categories in the training set). The network divides the image into a fixed-size grid and performs target detection in each grid cell. For each target, the YOLO algorithm outputs the coordinates of a bounding box and the corresponding class probability. The video, after overlaying the YOLO output information, is then streamed to the front-end interface for display. During front-end monitoring, the supervisor observes the situation within the monitored area in real time through the monitoring interface. When a target of interest is detected, the supervisor can select it by clicking on the target's bounding box. After selection, the camera that captured the target of interest (e.g., the target camera) packages the bounding box information of the target of interest, the target class probability, the target camera's own ID, and the corresponding image information within the bounding box and sends it to the photoelectric system.

[0123] Furthermore, this system supports automated point-of-interest (POI) selection. When the target camera detects a specific target (such as a drone), the system can automatically designate that target as a POI and send the relevant information to the photoelectric system for tracking. This feature significantly reduces the burden of manual operation and improves the system's response speed and overall monitoring efficiency.

[0124] Data association:

[0125] The optoelectronic system includes a camera and a gimbal, which allows the camera to rotate at the desired angle. After receiving the packet information from the camera, the optoelectronic system correlates the data with the target, thereby initiating adaptive tracking of the target.

[0126] First, we define the relevant coordinate system and some angles:

[0127] Camera reference coordinate system: Here, the commonly used northeast-northeast coordinate system is adopted, with true north as the x-axis, true east as the y-axis, and the z-axis determined according to the right-hand rule. The origin of the coordinate system is defined as the camera's installation position.

[0128] Pixel coordinate system: The top left corner of the image output by each camera is the origin. The x-axis coordinate u of a pixel represents its column number in the image array, and the y-axis coordinate v represents its row number in the image array.

[0129] Camera coordinate system: The origin is the camera's optical center. The x-axis is parallel to the x-axis of the pixel coordinate system, with positive to the right; the y-axis is parallel to the y-axis of the pixel coordinate system, with positive downwards; the z-axis follows the right-hand rule.

[0130] Camera mounting tilt angle: The angle between the camera coordinate system Z-axis and the horizontal plane xoy of the reference coordinate system, with the upward tilt being positive.

[0131] Camera mounting angle: The angle between the projection of the camera coordinate system's Z-axis onto the horizontal plane of the reference coordinate system (xoy) and the ox-axis, with clockwise direction being positive.

[0132] Camera elevation and elevation viewing angles: The direction of the line connecting the origin of the camera coordinate system and the target point is the optical axis, and the angle between this optical axis and the xoy plane of the reference coordinate system is positive upwards;

[0133] Camera horizontal line of sight angle: The direction of the line connecting the origin of the camera coordinate system and the target point is the optical axis. The angle between the projection of this optical axis onto the xoy plane of the reference coordinate system and the ox axis is positive in the clockwise direction.

[0134] Camera height frame angle: The direction of the line connecting the origin of the camera coordinate system and the target point is the optical axis. The angle between this optical axis and the Z-axis of the camera coordinate system in the vertical plane is positive upwards.

[0135] Camera horizontal frame angle: The direction of the line connecting the origin of the camera coordinate system and the target point is the optical axis. The angle between this optical axis and the Z-axis of the camera coordinate system on the xoy plane of the reference coordinate system is positive in the clockwise direction.

[0136] The camera reference coordinate system, pixel coordinate system, camera coordinate system, camera mounting tilt angle, camera mounting angle, camera elevation and slant angle, camera horizontal viewing angle, camera elevation and slant frame angle, and camera horizontal frame angle for each fixed camera are defined as above.

[0137] The specific steps for data association are as follows:

[0138] Step 1: Based on the camera ID information in the packaging information, query the corresponding intrinsic parameter information (Matrix) of that camera in the database. id Spatial location information P id And the installation posture information of the fixed camera. Where θ id It's the camera mounting angle. It is the camera mounting angle. The camera intrinsic parameter information matrix is ​​shown in formula (1), where f x and f y c represents the focal length of the camera along the x-axis and y-axis in the image plane, respectively. x and c y These represent the coordinates of the image center point along the x-axis and y-axis, respectively.

[0139] The spatial location information of the camera is represented as shown in formula (2), where lat id , lng id h id These represent the latitude, longitude, and altitude information of the camera ID, respectively.

[0140] Step 2: Based on the target category information (i.e., the aforementioned category probability) in the packaging information, query the database for the corresponding physical height l of the target in the real world. t rice.

[0141] Step 3: Based on the bounding box information and camera intrinsic parameters of the point of interest target in the packaging information, calculate the camera elevation and framing angle q of the selected point of interest target from the camera's viewpoint. α and horizontal frame angle q β .

[0142] The bounding box coordinates of the point of interest target are [x t ,y t [,w,h], where x t ,y t Here are the coordinates of the center point of the bounding box, and w and h are the width and height of the bounding box; the camera intrinsic parameters include the focal length f of the camera on the x-axis of the image plane. x and the focal length f in the y-axis direction y .

[0143] Therefore, the height and low frame angle q under the camera's perspective α and horizontal frame angle q β The calculation methods are shown in formulas (3) and (4).

[0144] Step 4: Determine the elevation and framing angles q of the target point of interest from the perspective of the target camera.α and horizontal frame angle q β In addition to the installation pose information of the target camera, the elevation angle q of the selected point of interest target under the view of the target camera is calculated. f and the camera's horizontal line of sight q h The specific calculation methods for the camera's vertical and horizontal viewing angles are shown in the aforementioned formulas (5) and (6).

[0145] Step 5: Based on the point of interest obtained in Step 4, determine the camera's elevation angle q from the target camera's viewpoint. f and the camera's horizontal line of sight q h And the physical height of the point of interest in the real world. t Calculate the coordinate range of the selected point of interest target in the reference coordinate system corresponding to the target camera.

[0146] First, assume there is a percentage error p between the actual size of the target and the value recorded in the database. Therefore, the actual size of the target, l r The range is l t ·(1±p).

[0147] The three-dimensional position (x, y) of the point of interest target in the camera reference coordinate system. p ,y p ,z p ). Figure 3 This is a schematic diagram of the optical imaging principle provided in the embodiments of this application, such as... Figure 3 As shown, the distance L from the camera lens to the target is obtained as follows:

[0148]

[0149] Where H is the height of the target in the reference coordinate system, A is the height of the target in the pixel coordinate system, and L is the distance from the camera lens to the target.

[0150] After substituting the target information of the point of interest into formula (14), the distance z from the target point of interest to the camera is obtained as shown in formula (7).

[0151] Therefore, the three-dimensional position (x, y) of the point of interest target in the camera reference coordinate system. p ,y p ,z p The calculation methods are shown in formulas (8), (9) and (10).

[0152] Because of l r ∈[l t ·(1-p),l tSubstituting (1+p)] into formulas (8), (9), and (10), we can obtain the range of the point of interest in the reference coordinate system as: point A = [x pmin ,y pmin ,z pmin ] T B = [x pmax ,y pmax ,z pmax ] T

[0153] Step Six: Based on the range of the point of interest target in the reference coordinate system corresponding to the target camera obtained in Step Five and the spatial position information of the target camera, calculate the coordinate range of the point of interest target in the reference coordinate system of the photoelectric system.

[0154] First, based on the coordinates of points A and B and the spatial location information of the target camera (i.e., the latitude, longitude, and altitude information of the target camera), the longitude, latitude, and altitude information of A and B are obtained respectively.

[0155] Secondly, using the geodetic projection method and the longitude, latitude, and altitude information of the known installation location of the photoelectric system, the coordinates A′ and B′ of points A and B in the reference coordinate system of the photoelectric system are obtained.

[0156] Step 7: Using the SiamRPN network, perform data association between the target camera and the optoelectronic system based on the coordinate range of the target point in the reference coordinate system of the optoelectronic system.

[0157] The target must lie between lines A′ and B′. The photoelectric system then scans between A′ and B′, while simultaneously activating the SiamRPN network to correlate the data. The SiamRPN network application method is as follows:

[0158] 1) The target bounding box in SiamRPN initialization is selected as the image content in the camera packet information;

[0159] 2) Extract the middle regions of all scanned images between lines A′ and B′ and perform feature extraction using a Siamese network. The Siamese network encodes the features of the target bounding box and the middle regions of all scanned images, enabling comparison in the feature space. Through correlation operations with the target template, the similarity between the candidate region and the target template is calculated. This similarity value represents the possible location of the target; the final target location is determined by maximizing the similarity. This completes the data association process.

[0160] By employing the data association method in this application embodiment, front-end personnel can directly select targets of interest from the real-time video stream of a fixed camera (corresponding to the target camera mentioned above), and then a mobile camera (i.e., the optoelectronic system mentioned above) automatically tracks the target of interest. The entire process relies solely on image information, eliminating the need for active sensors such as radar to perceive distance, thereby improving the system's stealth and preventing easy detection or destruction by the enemy. Furthermore, by fully utilizing parameters such as the fixed camera's intrinsic parameters, spatial position, and installation attitude information during the data association process, the accuracy of data association can be improved, further enhancing the precision of the optoelectronic system in identifying targets of interest. In addition, since the cost of a fixed camera is significantly lower than that of a mobile camera equipped with a pan-tilt system, this solution can significantly reduce the number of mobile cameras deployed while effectively monitoring the area, thus lowering the overall cost.

[0161] Step 8: Track points of interest.

[0162] Once the optoelectronic system completes data association, image tracking is performed using a traditional SiamRPN network. The bounding box information of the tracked target [x] is then transferred to the target. b ,y b ,w b ,h b The signal is transmitted to the photoelectric servo system, which uses the feedback information to generate control information to center the target's bounding box in the image and maintain the bounding box's size within the image. The control signal calculation is shown in formulas (11), (12), and (13). Figure 4 This is a schematic diagram of two fixed cameras provided in an embodiment of this application; Figure 5 This is a schematic diagram of a video screenshot of a photoelectric system tracking provided in an embodiment of this application, thus achieving target tracking.

[0163] This application provides a low-altitude surveillance solution based on multiple fixed cameras and an optoelectronic system. It aims to ensure accurate tracking of non-cooperative targets through passive monitoring of image information and system linkage, thereby improving the effectiveness and applicability of low-altitude surveillance. Specifically, the fixed cameras package and send the bounding box information, target category information, image information within the bounding box, and the camera ID of the target camera corresponding to the identified point of interest (POI) to the optoelectronic system. The optoelectronic system performs data association between the optoelectronic system and the fixed cameras based on the packaged data, and continues to track the POI based on the data association results. The system calculates the camera's vertical and horizontal frame angles (q) from the target camera's viewpoint based on the bounding box information and camera intrinsic parameters of the POI in the packaged data. βBased on the camera's high and low frame angles, horizontal frame angles, the target camera's mounting pose information, and the physical height of the point of interest (POI), the coordinate range of the POI in the reference coordinate system of the optoelectronic system is calculated. Then, the SiamRPN network is used to perform data association between the target camera and the optoelectronic system based on the coordinate range of the POI in the reference coordinate system of the optoelectronic system.

[0164] The technical effects of the embodiments of this application are as follows:

[0165] 1. Integrated Surveillance Combining Multiple Fixed Cameras and Optoelectronic Systems: This application embodiment achieves effective surveillance of low-altitude targets through the wide-area coverage of multiple fixed cameras and the precise tracking of the optoelectronic system. The cameras handle large-scale monitoring, while the optoelectronic system selects points of interest within the camera feed for detailed tracking, providing a comprehensive solution for monitoring non-cooperative targets.

[0166] 2. Enhanced Passive Information Monitoring and Stealth: This application's embodiments rely entirely on image information for monitoring, avoiding the use of active information. This passive monitoring method not only broadens the scope of application for monitoring, making it particularly suitable for non-cooperative targets where active information cannot be obtained, but also improves the system's stealth, giving it strong confidentiality in military applications.

[0167] 3. Innovative Algorithm Enables Linkage Between Camera and Optoelectronic System: An innovative algorithm correlates image data from different locations, resolving the linkage issue between the fixed camera and the optoelectronic system. This step ensures that the optoelectronic system can adjust its viewing angle in real time, keeping the target always centered in the field of view, thereby improving the accuracy and stability of monitoring.

[0168] 4. Cost-efficient design: The solution combining multiple fixed cameras and an optoelectronic system achieves an optimal balance between cost and performance. Fixed cameras provide wide coverage, while the optoelectronic system provides precise tracking. This combination effectively reduces overall cost and improves the system's economic efficiency.

[0169] Compared to traditional methods, this embodiment adopts the following approach: First, multiple fixed cameras are installed within the monitored area, ensuring the total field of view covers the entire area. The images from each camera are then displayed in real-time at the front end. When an unauthorized intrusion appears in the image displayed by a fixed camera, the monitor uses a mouse to circle the intrusion. The photoelectric system, upon receiving the intrusion information from the monitor, tracks the object and adaptively adjusts its focal length, displaying the tracked image in real-time at the front end. The advantage of this solution is that it combines multiple fixed cameras and a photoelectric system for low-altitude monitoring, leveraging the strengths of each sensor. The linkage between the photoelectric system and multiple cameras broadens the applicability of current low-altitude monitoring, enabling the monitoring of non-cooperative targets. Since the entire monitoring process uses only passive sensor information, it is less likely to be detected by third parties, improving the security of the monitoring system and making it suitable for military applications. Furthermore, a novel data association algorithm is introduced to address the linkage issue between the photoelectric system and multiple cameras. Finally, the photoelectric system adaptively tracks the associated objects, displaying the tracking image at the front end, allowing monitors to take further action against the monitored objects.

[0170] This application addresses the limitations of traditional low-altitude surveillance methods in monitoring non-cooperative targets in complex environments by proposing a base station-free monitoring scheme. By combining multiple fixed cameras and an optoelectronic system, this application achieves wide-area coverage and high-precision tracking of targets within the monitored area. Fixed cameras capture all potential targets over a large area, while the optoelectronic system precisely tracks targets based on points of interest selected by the monitor. While the optoelectronic system is relatively expensive, it offers flexible viewing angle adjustments, achieving an optimal balance between monitoring effectiveness and cost-effectiveness. Especially in situations where base stations cannot be installed or the installation conditions are unsuitable, this scheme still provides stable and reliable monitoring capabilities, achieving efficient surveillance of non-cooperative targets solely based on image information.

[0171] Through this combination, the embodiments of this application significantly improve the applicability and flexibility of the low-altitude airspace monitoring system, especially in complex environments, effectively ensuring the safety and efficient management of low-altitude airspace.

[0172] Figure 6 This is a schematic diagram of the structural composition of the monitoring device provided in the embodiments of this application, as shown below. Figure 6 As shown, the monitoring device includes:

[0173] The acquisition unit 601 is used to acquire a first video set of the first camera set; wherein the videos in the first video set correspond to the cameras in the first camera set;

[0174] The determining unit 602 is used to determine target interest content based on the first video set;

[0175] The acquisition unit 601 is further configured to acquire first information corresponding to the target interest content; the first information includes first sub-information of the target interest content and second sub-information of the target camera; the target camera is the camera corresponding to the video where the target interest content is located;

[0176] The determining unit 602 is used to determine the data association result based on the first information. The data association result is the result of data association between the target camera and the first system. The data association result is used for monitoring the target's content of interest.

[0177] In some embodiments, the acquisition unit 601 is used to acquire an initial video set of a first camera set; the videos in the initial video set are captured by cameras in the first camera set; target detection is performed on the videos in the initial video set through a first network to obtain second information; wherein the second information includes target-related information; and the first video set is determined based on the second information and the initial video set.

[0178] In some implementations, the determining unit 602 is used to determine target interest content based on the second information and the initial video set.

[0179] In some embodiments, the acquisition unit 601 is configured to determine first sub-information of the target interest content based on the second information and the initial video set; determine the ID information of the target camera based on the target interest content; and determine the second sub-information of the target camera based on the ID information.

[0180] In some implementations, the first sub-information includes one or more of the following: bounding box information of the target interest content, target category probability information of the target interest content, and image information of the target interest content; the second sub-information includes one or more of the following: identity ID information of the target camera, intrinsic parameter information of the target camera, spatial location information of the target camera, and installation attitude information of the target camera.

[0181] In some implementations, the determining unit 602 is used to determine the first location information of the target interest content based on the first information; and to determine the data association result based on the first location information through a second network.

[0182] In some embodiments, the determining unit 602 is configured to: determine a first height of the target based on the target category probability information of the target interest content; determine the camera's elevation and horizontal frame angles based on the bounding box information of the target interest content and the intrinsic parameter information of the target camera; determine the camera's elevation and horizontal line-of-sight angles based on the camera's elevation and horizontal frame angles and the target camera's mounting posture information; determine second position information based on the camera's elevation and horizontal line-of-sight angles and the first height; and determine first position information based on the second position information and the target camera's spatial position information.

[0183] Those skilled in the art should understand that Figure 6 The functions of each unit in the monitoring device shown can be understood by referring to the relevant description of the aforementioned method. Figure 6 The functions of each unit in the monitoring device shown can be implemented through a program running on a processor or through specific logic circuits.

[0184] Figure 7 This is a schematic structural diagram of a monitoring device 700 provided in an embodiment of this application. Figure 7 The monitoring device 700 shown includes a processor 710, which can call and run computer programs from memory to implement the methods in the embodiments of this application.

[0185] Optionally, such as Figure 7 As shown, the monitoring device 700 may further include a memory 720. The processor 710 can retrieve and run computer programs from the memory 720 to implement the methods described in this embodiment.

[0186] The memory 720 can be a separate device independent of the processor 710, or it can be integrated into the processor 710.

[0187] Optionally, such as Figure 7 As shown, the monitoring device 700 may also include a transceiver 730, which the processor 710 can control to communicate with other devices. Specifically, it can send information or data to other devices or receive information or data sent by other devices.

[0188] The transceiver 730 may include a transmitter and a receiver. The transceiver 730 may further include antennas, and the number of antennas may be one or more.

[0189] The monitoring device 700 can implement the corresponding processes implemented by the monitoring device in the various methods of the embodiments of this application, which will not be described in detail here for the sake of brevity.

[0190] Figure 8 This is a schematic structural diagram of the chip according to an embodiment of this application. Figure 8 The chip 800 shown includes a processor 810, which can call and run computer programs from memory to implement the methods in the embodiments of this application.

[0191] Optionally, such as Figure 8 As shown, chip 800 may further include memory 820. Processor 810 can retrieve and run computer programs from memory 820 to implement the methods described in this embodiment.

[0192] The memory 820 can be a separate device independent of the processor 810, or it can be integrated into the processor 810.

[0193] Optionally, the chip 800 may also include an input interface 830. The processor 810 can control the input interface 830 to communicate with other devices or chips; specifically, it can acquire information or data sent by other devices or chips.

[0194] Optionally, the chip 800 may also include an output interface 840. The processor 810 can control the output interface 840 to communicate with other devices or chips, specifically, to output information or data to other devices or chips.

[0195] This chip can implement the corresponding processes implemented by the monitoring device in the various methods of the embodiments of this application, which will not be described in detail here for the sake of brevity.

[0196] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0197] It should be understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0198] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0199] It should be understood that the above-described memory is exemplary and not a limiting description. For example, the memory in the embodiments of this application may also be static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DR RAM), etc. That is to say, the memory in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.

[0200] This application also provides a computer program product, including a computer program.

[0201] When executed by the processor, the computer program implements the corresponding processes implemented by the monitoring device in the various methods of the embodiments of this application, which will not be described in detail here for the sake of brevity.

[0202] This application also provides a computer-readable storage medium for storing computer programs.

[0203] The computer program causes the computer to execute the corresponding processes implemented by the monitoring device in the various methods of the embodiments of this application, which will not be described in detail here for the sake of brevity.

[0204] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0205] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0206] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0207] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0208] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0209] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0210] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A regulatory method, characterized in that, The method includes: Obtain a first video set from the first camera set; wherein the videos in the first video set correspond to the cameras in the first camera set; Determine the target interest content based on the first video set; Obtain first information corresponding to the target interest content; the first information includes first sub-information of the target interest content and second sub-information of the target camera; the target camera is the camera corresponding to the video where the target interest content is located. The data association result is determined based on the first information. The data association result is the result of data association between the target camera and the first system. The data association result is used for monitoring the target's content of interest.

2. The method according to claim 1, characterized in that, The acquisition of the first video set from the first camera set includes: Obtain an initial video set from the first camera set; the videos in the initial video set are captured by the cameras in the first camera set. The first network is used to perform target detection on the videos in the initial video set to obtain second information; wherein, the second information includes target-related information; The first video set is determined based on the second information and the initial video set.

3. The method according to claim 2, characterized in that, The step of determining the target interest content based on the first video set includes: The target interest content is determined based on the second information and the initial video set.

4. The method according to claim 2, characterized in that, The step of obtaining the first information corresponding to the target interest content includes: Based on the second information and the initial video set, the first sub-information of the target interest content is determined; The ID information of the target camera is determined based on the target interest content, and the second sub-information of the target camera is determined based on the ID information.

5. The method according to any one of claims 1 to 4, characterized in that, The first sub-information includes one or more of the following: bounding box information of the target interest content, target category probability information of the target interest content, and image information of the target interest content; the second sub-information includes one or more of the following: identity ID information of the target camera, intrinsic parameter information of the target camera, spatial location information of the target camera, and installation posture information of the target camera.

6. The method according to claim 5, characterized in that, The step of determining the data association result based on the first information includes: Based on the first information, determine the first location information of the target interest content; The data association result is determined by the second network based on the first location information.

7. The method according to claim 6, characterized in that, The step of determining the first location information of the target interest content based on the first information includes: The first height of the target is determined based on the target category probability information of the target interest content; The camera's high and low frame angles and horizontal frame angles are determined based on the bounding box information of the target interest content and the intrinsic parameter information of the target camera. The camera's elevation and horizontal viewing angles are determined based on the camera's elevation and horizontal frame angles, the horizontal frame angles, and the target camera's mounting posture information. The second position information is determined based on the camera's elevation and lateral viewing angles, the camera's horizontal viewing angle, and the first height. The first location information is determined based on the second location information and the spatial location information of the target camera.

8. A monitoring device, characterized in that, The device includes: The acquisition unit is used to acquire a first video set of the first camera set; wherein the videos in the first video set correspond to the cameras in the first camera set; A determining unit is configured to determine target interest content based on the first video set; The acquisition unit is further configured to acquire first information corresponding to the target interest content; the first information includes first sub-information of the target interest content and second sub-information of the target camera; the target camera is the camera corresponding to the video where the target interest content is located; The determining unit is used to determine a data association result based on the first information, wherein the data association result is the result of data association between the target camera and the first system; the data association result is used for monitoring the target's content of interest.

9. A monitoring device, characterized in that, include: A processor and a memory for storing a computer program, the processor for calling and running the computer program stored in the memory to perform the method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, include: A computer program that, when executed by a processor, implements the method according to any one of claims 1 to 7.

11. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the method as described in any one of claims 1 to 7.