Multi-mode photoelectric information fusion farmland monitoring method, device, equipment and medium

Through multimodal photoelectric information fusion technology, combined with radar and cameras, the targets in the farmland are determined and tracked, and the improved long-term and short-term memory network and attention mechanism are used to predict trajectory and feature fusion, solving the accuracy of target recognition and tracking in the existing technology, and achieving high-precision farmland safety monitoring.

CN120143130AActive Publication Date: 2025-06-13HAI NAN ZHI YUAN KE JI YOU XIAN GONG SI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510615028.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

Existing farmland safety monitoring technology is difficult to achieve accurate target identification and tracking in complex target tracking and changing farmland environments.

Method used

The multimodal photoelectric information fusion method is adopted, combined with radar and multi-eye cameras to determine the tracking target. By building a continuous trajectory and using improved long and short-term memory networks for trajectory prediction, the drone is controlled for tracking and shooting, and the additive attention mechanism is used to fuse images and trajectory features for target tracking and classification.

Benefits of technology

It improves the accuracy and real-time nature of target detection, tracking and classification, and achieves high-accuracy identification and alarm of targets in agricultural insurance scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120143130A_ABST
    Figure CN120143130A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-mode photoelectric information fusion farmland monitoring method, device, equipment and medium, and belongs to the technical field of farmland monitoring under radar and video information fusion. According to the method, a fixed point radar and a multi-view camera are matched to screen a tracking target and construct a continuous trajectory of the tracking target, a lost trajectory is supplemented by a long-short term memory network improved by an additional door during construction, and after an unmanned aerial vehicle approaches the tracking target to shoot a tracking image, the tracking image and the continuous trajectory at the same moment are fused by an additive attention mechanism. And gradually tracking and identifying the target based on similarity comparison of fusion features of two adjacent moments. According to the method, the improved long and short term memory network and the attention mechanism enhancement method are combined, the multi-modal photoelectric information is fused and utilized, the tracking target is determined through the multi-source data, the target position is gradually determined through adjacent fusion feature comparison, the accuracy and real-time performance of target detection, tracking and classification can be effectively improved, and the accuracy and real-time performance of target detection, tracking and classification are improved. And high-accuracy identification alarm of the target in the agricultural insurance scene is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of farmland monitoring technology under the fusion of radar and video information, and in particular to a farmland monitoring method, device, equipment and medium under the fusion of multi-modal photoelectric information. Background Art

[0002] With the development of agricultural science and technology, smart agriculture has gradually become an important development direction of modern agriculture. The use of the Internet of Things, artificial intelligence, drones, remote sensing technology, etc. to achieve refined management of farmland environment has become one of the important means of agricultural production. In the subdivision of smart agricultural management, farmland safety monitoring has received more and more attention, especially in application scenarios such as scientific research test fields, high-value crop planting areas, and unmanned farms, which have put forward higher requirements for intrusion monitoring, theft prevention, and wild animal expulsion.

[0003] For farmland safety monitoring, accurate tracking and positioning of intrusion targets is a prerequisite for realizing their type identification and thus completing targeted processing. Although there are relevant technical solutions in the current farmland safety monitoring segment that integrate multiple sensing methods to identify intrusion targets, the way such solutions make comprehensive use of various sensor data is usually weighted fusion with fixed weights.

[0004] For the agricultural protection scenario of farmland safety monitoring, most of the targets to be identified (such as animals and birds) have the characteristics of variable and irregular movement paths and large fluctuations in movement speed; at the same time, most of the protected areas (i.e. farmlands) in the agricultural protection scenario are in outdoor environments, which leads to a large degree of environmental impact on the farmland safety monitoring process. Combining the above two reasons, the existing fixed-weight multi-mode data fusion target tracking method is difficult to cope with the complex state of the targets to be identified and the changeable farmland environment in the agricultural protection scenario due to the lack of adaptive ability and reliable complex target tracking ability, resulting in it being unable to meet the requirements of accurate tracking and positioning of targets in the agricultural protection scenario, and unable to provide accurate and reliable fusion data for target type identification, and ultimately unable to effectively guarantee farmland safety.

[0005] Therefore, the current farmland monitoring methods in agricultural protection scenarios have the technical problem of insufficient target recognition reliability. Summary of the invention

[0006] In view of this, the embodiments of the present invention provide a farmland monitoring method, device, equipment and medium with multi-modal optoelectronic information fusion to solve the technical problem that the farmland monitoring method in the current agricultural protection scenario has insufficient reliability in target recognition.

[0007] In a first aspect, a method for farmland monitoring using multimodal optoelectronic information fusion is provided, the method comprising: Determine a first effective target with a monitoring point radar outside the preset distance of the monitoring point, determine a second effective target with a monitoring point multi-camera within the preset distance of the monitoring point, and use the first effective target and the second effective target as tracking targets; Take the speed and position of the tracking target at the same moment as the trajectory at that moment, construct a continuous trajectory based on the trajectory, and when the trajectory is lost at the current moment, input the trajectory of the previous moment into the trained long short-term memory network with an additional gate to output a predicted trajectory as the trajectory at the current moment; After the drone reaches the tracking target according to the continuous trajectory, track and continuously photograph the tracking target to obtain a tracking image, extract a first type of feature from the tracking image, extract a second type of feature from the continuous trajectory, fuse the first type of feature and the second type of feature at the same moment into a fused feature with an additive attention mechanism, determine the position of the tracking target at the latter moment in the adjacent two moments based on the similarity of the fused features at the adjacent two moments to achieve the tracking, and determine the category of the tracking target from the fused feature for processing.

[0008] In a second aspect, a farmland monitoring device for multimodal optoelectronic information fusion is provided, and the device includes: A tracking target selection module, configured to determine a first effective target with a monitoring point radar outside the preset distance of the monitoring point, determine a second effective target with a monitoring point multi-camera within the preset distance of the monitoring point, and use the first effective target and the second effective target as tracking targets; A target trajectory construction module, configured to take the speed and position of the tracking target at the same moment as the trajectory at that moment, construct a continuous trajectory based on the trajectory, and when the trajectory is lost at the current moment, input the trajectory of the previous moment into the trained long short-term memory network with an additional gate to output a predicted trajectory as the trajectory at the current moment; A drone tracking and warning module, configured to, after the drone reaches the tracking target according to the continuous trajectory, track and continuously photograph the tracking target to obtain a tracking image, extract a first type of feature from the tracking image, extract a second type of feature from the continuous trajectory, fuse the first type of feature and the second type of feature at the same moment into a fused feature with an additive attention mechanism, determine the position of the tracking target at the latter moment in the adjacent two moments based on the similarity of the fused features at the adjacent two moments to achieve the tracking, and give an alarm after determining the category of the tracking target.

[0009] In a third aspect, an embodiment of the present invention provides a computer device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the farmland monitoring method for multimodal optoelectronic information fusion as described in the first aspect is implemented.

[0010] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the multi-modal optoelectronic information fusion-based farmland monitoring method as described in the first aspect.

[0011] The beneficial effects of the present invention compared with the prior art are as follows: In this method of the present invention, first, at a fixed monitoring point, a radar and a multi-camera are used in cooperation to screen and track a target and determine the position and speed of the tracked target. Then, a continuous trajectory of the tracked target is constructed, and during the construction process, a long short-term memory network improved by an additional gate is used to predict and supplement the lost trajectory. After that, a drone is controlled to approach the tracked target to capture its tracking image, and the tracking image of the tracked target and the continuous trajectory at the same moment are fused by an additive attention mechanism. Then, based on the similarity comparison of the fusion features at two adjacent moments, step-by-step iterative tracking of the target based on multi-modal optoelectronic information fusion is realized, and the target category is determined by the fusion features to complete the alarm. The present invention combines an improved long short-term memory network and an attention mechanism strengthening method to utilize multi-modal optoelectronic information fusion of fixed video, radar, and mobile video. At the stage of screening the target to be recognized, multi-modal data are supplemented and coordinated. At the stage of target tracking, the accurate step-by-step iteration of the target position is realized through the comparison of the fusion features at two adjacent moments, and the target category is determined based on the fusion features. Finally, the accuracy and real-time performance of target detection, tracking, and classification can be effectively improved, and high-accuracy recognition and alarm of the target in the agricultural protection scenario can be realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts.

[0013] Figure 1 FIG. 1 is a schematic diagram of an application environment of a multi-modal optoelectronic information fusion-based farmland monitoring method provided in Embodiment 1 of the present invention; Figure 2 FIG. 2 is a schematic flowchart of a multi-modal optoelectronic information fusion-based farmland monitoring method provided in Embodiment 1 of the present invention; Figure 3 FIG. 3 is a schematic structural diagram of a multi-modal optoelectronic information fusion-based farmland monitoring device provided in Embodiment 3 of the present invention; Figure 4 FIG. 4 is a schematic structural diagram of a computer device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] In the following description, specific details such as specific system architectures, technologies, etc. are presented for purposes of illustration and not limitation, so as to provide a thorough understanding of the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present invention.

[0015] It should be understood that when used in the specification and claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0016] It should also be understood that the term "and / or" as used in the specification and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0017] As used in the specification and claims of the present invention, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.

[0018] In addition, in the description of the specification and claims of the present invention, the terms "first", "second", "third", etc. are only used for differentiating descriptions and cannot be understood as indicating or implying relative importance.

[0019] Reference to "an embodiment" or "some embodiments" or the like described in the specification of the present invention means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0020] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, techniques, and application systems.

[0021] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0022] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0023] To illustrate the technical solution of the present invention, it will be described below through specific embodiments.

[0024] A farmland monitoring method for multimodal optoelectronic information fusion provided in Embodiment 1 of the present invention can be applied in an application environment such as Figure 1 where the client communicates with the server. Among them, the client includes but is not limited to terminal devices such as palm computers, desktop computers, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, cloud terminal devices, and personal digital assistants (PDAs). The server can be implemented by an independent server or a server cluster composed of multiple servers.

[0025] The above-mentioned farmland monitoring method for multimodal optoelectronic information fusion can be applied to the server in FIG. 1. The computer device corresponding to the server is connected to a corresponding database, rule base, etc. to obtain corresponding data in the database. The above computer device is connected to the client to collect data and control instructions sent by the client user, and send relevant monitoring parameter data and video data to the client for retention. The above computer device is also connected to a sensing device to obtain relevant monitoring data collected by the sensing device. The sensing device includes but is not limited to video sensors on unmanned aerial vehicles, spherical cameras at fixed monitoring points, multi-eye cameras at fixed monitoring points, and radars at fixed monitoring points.

[0026] See Figure 2, is a schematic flowchart of a multi-modal optoelectronic information fusion-based farmland monitoring method provided in the first embodiment of the present invention. The method may include the following steps: Step S101, determine a first effective target at a preset distance outside the monitoring point by the monitoring point radar, determine a second effective target within the preset distance of the monitoring point by the monitoring point multi-camera, and use the first effective target and the second effective target as tracking targets.

[0027] This farmland monitoring method in this embodiment is executed in an unmanned manner. Therefore, to achieve unmanned monitoring, necessary hardware devices need to be deployed in advance. In this embodiment, a vertical pole is set as a bearing platform and denoted as the monitoring point. A radar, a multi-camera, and a drone airport required for monitoring are arranged on the vertical pole, serving as the hardware basis for the implementation of this farmland monitoring method in this embodiment.

[0028] For unified coordinate reference, in this embodiment, the vertical pole is taken as the center origin, the west-east direction is the x-axis direction, and the north-south direction is the y-axis direction to form a local coordinate system of the monitoring area. All the following observation coordinates are based on this coordinate system.

[0029] In this embodiment, 4 radars are used to form a 360° observation range, and the aperture of each radar is 90°, which is used to detect the monitoring area, that is, the farmland. Although a 360° observation range can be formed by deploying 4 radars, due to the limitations of the working principle of the radar itself, its recognition accuracy for objects closer to itself is poor or it cannot recognize them, that is, there is a recognition blind area. To avoid the tracking targets in the blind area in the tracking targets obtained by the radar being misdetected targets, in this embodiment, it is set to use the multi-camera on the monitoring point, that is, the vertical pole, to fuse the radar information to supplement the acquisition of tracking targets. For this reason, in this embodiment, a preset distance is determined according to the distance between the blind area boundary of the radar and the monitoring point. The area outside the preset distance from the monitoring point is used as the radar monitoring area, and the area within the preset distance from the monitoring point is used as the multi-camera monitoring area. The radar identifies and filters the moving objects in the radar monitoring area, and the multi-camera identifies and filters the moving objects in the multi-camera monitoring area.

[0030] Specifically, the radar uses the Doppler effect and beam scanning to detect the speed of the moving object in the radar monitoring area 、the echo time t and the azimuth relative to the center of the radar module , and further the distance = ct / 2, where c represents the speed of light, and the speed is the radial speed according to the Doppler radar principle. Also, the echo intensity p can be described according to the power of the radar received signal. Therefore, the output parameters are the radial speed 、position and echo intensity p, where , .

[0031] For moving objects in the monitoring area, some belong to categories where threats are clearly determinable as non-existent, such as operators performing planting and maintenance operations in farmland, smaller objects, etc. Therefore, for the moving objects detected by the radar in the radar monitoring area, it is first necessary to preliminarily screen the moving objects based on the moving speed and echo intensity of the moving objects obtained by the radar to obtain the first effective targets in the radar monitoring area, so as to exclude such targets as operators.

[0032] Then, for the moving objects in the radar monitoring area, the speed threshold and echo intensity threshold method are used to screen the first effective target Aobj from them. Here, the speed range is set as the minimum radial speed , the maximum radial speed , the minimum echo intensity pmin = 120, and the maximum echo intensity pmax = 200. Assume that the moving object detected by the radar scan is obj, then the first effective target is as follows: Aobj = obj (pmin < p < pmax and < < ), and one or more of these targets can be selected for tracking subsequently.

[0033] The multi-camera uses its own target recognition algorithm to recognize the objects in the monitoring area of the multi-camera, obtaining the vertical pixel number Hpix and horizontal pixel number Wpix of the target. The target recognition algorithm can adopt any existing feasible solution, such as target recognition based on YOLOv8, and determines the moving speed of the object through the moving object tracking algorithm on the multi-camera , and the moving object tracking algorithm on the multi-camera can also adopt any existing feasible solution, and automatically adjusts the viewing angle of the multi-camera through the PTZ (Pan-Tilt-Zoom) function to ensure that the object in the monitoring area of the multi-camera is always at the center of the field of view of the multi-camera. Then, the azimuth of the object relative to the multi-camera can be determined according to the rotation angle of the multi-camera , and based on the parameters of the multi-camera and the position of the object in the field of view of the multi-camera, the distance between the object and the monitoring point can be estimated , so as to determine the position of the object detected by the multi-camera according to the distance and the azimuth , where , , .

[0034] Similarly, since some of the moving objects in the monitoring area clearly belong to categories that pose no threat, it is also necessary to screen the objects in the monitoring area of the multi-camera. Since the size S and focal length F of the sensors of the multi-camera are both known, the actual physical size of the objects in the monitoring area of the multi-camera can be calculated first: Hreal = Hpix × ri × S / F, Wreal = Wpix × ri × S / F. According to the actual physical size, screen the objects in the monitoring area of the multi-camera to determine the second effective target: Bobj = obj (Hreal > 50 cm and Wreal > 50 cm).

[0035] Thus, the tracking targets that need to be processed and judged for the entire monitoring area can be determined based on Bobj and Aobj. Subsequently, it is necessary to drive and control the takeoff of the UAV according to the tracking targets determined by Bobj and Aobj, and provide the UAV with the tracking target positions to assist the UAV in planning the route, approaching for observation, and then carrying out subsequent tracking and warning operations.

[0036] Step S102: Take the speed and position of the tracking target at the same time as the trajectory at that moment, construct a continuous trajectory based on the trajectory, and when the trajectory at the current moment is lost, input the trajectory at the previous moment into the trained long short-term memory network with an additional gate to output the predicted trajectory as the trajectory at the current moment.

[0037] Before controlling the UAV to start tracking and warning the tracking target, in this embodiment, the motion trajectory of the tracking target needs to be constructed first as the data for subsequent recognition of the tracking target category. The construction of the motion trajectory adopts the existing method. The following is an example of the construction process of the motion trajectory in this embodiment.

[0038] Specifically, based on the position and speed of the tracking target monitored by the radar and the multi-camera at the same time as the trajectory at that moment:

[0039] Among them, represents the trajectory of the tracking target at time k, represents the abscissa of the tracking target at time k, represents the ordinate of the tracking target at time k, represents the horizontal speed of the tracking target at time k, , represents the vertical speed of the tracking target at time k, , represents the vector speed of the tracking target at time k, represents the azimuth angle of the tracking target relative to the monitoring point.

[0040] Based on the trajectory of the tracking target at time k, the predicted trajectory at the next time, i.e., time k+1, can be obtained:

[0041]

[0042]

[0043] Among them, represents the predicted trajectory at time k+1, is the noise at time k, F is the mapping matrix, is the sampling interval, 、 、 and are the noise corresponding to the abscissa, the noise corresponding to the ordinate, the noise corresponding to the horizontal velocity, and the noise corresponding to the vertical velocity at time k, respectively.

[0044] Therefore, there is:

[0045] Among them, represents the predicted value of the abscissa at time k+1, represents the predicted value of the ordinate at time k+1, represents the horizontal velocity at time k+1, represents the vertical velocity at time k+1.

[0046] If there are actual sampling values at time k+1, it is easy to understand that there should be:

[0047] However, there will be an error between the predicted value and the actual sampling value at the actual time k+1, so it can be determined that:

[0048] And the noise obtained at time k+1 is used for the trajectory prediction at time k+2, thereby completing the trajectory iteration and finally forming a continuous trajectory.

[0049] Meanwhile, during the construction of the continuous trajectory, it is inevitable that the radar or multi-camera will miss the recognition of moving objects in their respective monitoring areas at a certain or certain moments, that is, the trajectory information at a certain or certain moments is lost. Then, if the trajectory is lost at the current moment, the long short-term memory network (LSTM) improved by the additional gate is used to intelligently predict the trajectory of the moving object at the current moment. The long short-term memory network (LSTM) improved by the additional gate consists of an input gate, a forget gate, an additional gate, and an output gate, which is used to capture the dependencies in the continuous trajectory information, so as to reconstruct the lost trajectory information and solve the problem of target loss.

[0050] Specifically, first, based on the above trajectory construction method, the known trajectory is used as the training set and input into the long short-term memory network improved by the additional gate for training, that is, through the continuous trajectory known or constructed so far up to the current M moment , where , to train the long short-term memory network model improved by the additional gate and obtain the trained long short-term memory network improved by the additional gate. Then, when the trajectory is lost at the current moment, the trajectory of the previous moment is input into the trained long short-term memory network improved by the additional gate, so that the network outputs the predicted value of the trajectory at the current moment, and the predicted value of the trajectory at the current moment is used as the trajectory at the current moment and the continuous trajectory construction continues.

[0051] Step S103, after the UAV reaches the tracking target according to the continuous trajectory, it tracks and continuously shoots the tracking target to obtain a tracking image, extracts a first type of feature from the tracking image, extracts a second type of feature from the continuous trajectory, and fuses the first type of feature and the second type of feature at the same moment into a fusion feature by an additive attention mechanism. The position of the tracking target at the latter moment in the adjacent two moments is determined by the similarity of the fusion features at the adjacent two moments to realize the tracking, and the category of the tracking target is determined by the fusion feature and then an alarm is given.

[0052] During the construction of the continuous trajectory, the UAV can determine the latest position of the tracking target to be recognized and alarmed based on the latest trajectory in the continuous trajectory, so as to control the UAV in the UAV airport to reach the tracking target to be recognized and alarmed, and in the subsequent process, the tracking image information of the tracking target captured by the video sensor on the UAV and the continuous trajectory information corresponding to the tracking target are used to fuse the features of the two aspects to construct a fusion feature, so as to complete the tracking and recognition and alarm of the tracking target.

[0053] Specifically, after the UAV reaches the position where the tracking target is located, it first takes a tracking image of the tracking target through the video sensor on it at the arrival moment, and it is easy to understand that there is also a continuous trajectory corresponding to the tracking target at the arrival moment.

[0054] Then, in order to achieve the integrated utilization of the two aspects of information, namely the tracking image information and the continuous trajectory information, in this embodiment, first, the existing CNN network (such as ResNet, but not limited to ResNet) is used to extract features from the tracking image as a type of feature :

[0055] Among them, is a type of feature extracted by the CNN network from the tracking image, represents the mapping function of the CNN network model, is the weight pre-learned by the CNN network, is the tracking image.

[0056] Then, the continuous trajectory obtained by using the improved LSTM above is subjected to local window mapping to obtain a short window mapping trajectory image, and a second type of feature is extracted from this short window mapping trajectory image by a CNN + attention mechanism (such as Yolov8) network model :

[0057] Among them, represents the mapping function of the CNN + attention mechanism network model, is the weight pre-learned by the CNN network, is the short window mapping trajectory image.

[0058] Then, cross-modal feature fusion is performed on the extracted first-type and second-type features. Since the second-type features are mainly obtained by radar and the first-type features are obtained by video sensors, and the radar and video feature channels are different, so we use the sigmoid function to first perform weighted adjustment on the two features:

[0059]

[0060] Among them, and are the weighting of the first-type feature and the second-type feature respectively, and are the weights pre-learned by the supervised learning method, is the sigmoid function.

[0061] Then, the enhancement of the spatial attention mechanism is carried out. Specifically, the additive attention mechanism is used to fuse the first-type feature and the second-type feature to obtain the fused feature :

[0062]

[0063] Among them, is the spatial attention weight pre-learned by the supervised learning method, is the fused attention weight matrix, is the normalization exponential function.

[0064] Thus, the fused feature corresponding to the arrival time when the UAV reaches the tracking target position can be obtained.

[0065] There are two purposes for obtaining the fused feature in this embodiment. On the one hand, the target type is recognized based on the fused feature obtained by fusing the multi-modal optoelectronic information acquired under radar, fixed video, and mobile video to improve the recognition accuracy, and warning and driving away are performed based on the target category. On the other hand, since the UAV warning processing needs to be continuously carried out, the fused feature containing multi-modal optoelectronic information is also used to complete the continuous and accurate tracking of the UAV for the tracking target.

[0066] Then, to achieve the tracking of the UAV for the tracking target, after obtaining the fused feature corresponding to the arrival time in this embodiment, at the next moment, the UAV continues to capture the tracking image of the tracking target, and the short-window mapping trajectory image corresponding to the next moment is obtained by mapping the continuous trajectory of the tracking target corresponding to the next moment through a local window. And as the method for obtaining the fused feature corresponding to the arrival time described above, based on the tracking image and the short-window mapping trajectory image at the next moment, the fused feature corresponding to the next moment is obtained.

[0067] Thus, the fused features of two adjacent moments, namely the arrival time and the next moment after the arrival time, are obtained. Here, we regard the implementation process of the fused feature as a function, and calculate the similarity (cross-correlation) between the multi-modal optoelectronic information (X1, Y1) of the previous moment and the multi-modal optoelectronic information (X2, Y2) of the next moment among two adjacent moments in the Siamese network, that is, calculate the similarity of the fused features of these two adjacent moments:

[0068] Among them, is the similarity score, represents the cross-correlation calculation, represents the tracking image of the previous moment among two adjacent moments, represents the tracking image of the next moment among two adjacent moments, represents the short-window mapping trajectory image of the previous moment among two adjacent moments, represents the short-window mapping trajectory image of the next moment among two adjacent moments, represents the fused feature of the previous moment among two adjacent moments, represents the fused feature at the latter moment among two adjacent moments.

[0069] However, it should be noted that since the position of the tracking target at the latter moment is random, there are actually multiple possibilities for the content of the tracking image taken by the UAV at the latter moment. That is, before the position of the tracking target at the latter moment is determined, the tracking image at the latter moment can be understood as having multiple candidate images, and each candidate image can obtain a corresponding fused feature at the latter moment, so each candidate image corresponds to a similarity score.

[0070] That is, since the position of the tracking target at the latter moment among two adjacent moments is not yet determined for the UAV, there are multiple possible values for the above similarity calculation result. It is easy to understand that when the similarity value is the largest, it means that the similarity between the position at the latter moment and the known position of the tracking target at the previous moment among the two adjacent moments is the largest, that is, the similarity between the distribution of the tracking target in the images at the two moments is the largest. At this time, the position at the latter moment is most likely the actual position of the tracking target at the latter moment, and the above candidate image at the latter moment can be recorded as the tracking image at the latter moment and used as the tracking image at the previous moment in the next tracking iteration. Thus, based on the position of the tracking target at the arrival moment, combined with the similarity value of the fused features between the arrival moment and the next moment after the arrival moment, the position of the tracking target at the next moment after the arrival moment can be determined, that is, the position of the tracking target at the latter moment among two adjacent moments can be determined:

[0071] Among them, represents the position at the latter moment among two adjacent moments corresponding to the maximum similarity value of the tracking target, represents determination The position coordinates corresponding to when taking the maximum value.

[0072] Repeat the above process of determining the position of the tracking target by the UAV at the latter moment among two adjacent moments. Starting from the arrival moment when the UAV reaches the position of the tracking target, the UAV can realize the tracking of the tracking target, continuously obtain the fused features at each moment during the tracking process, and based on the obtained fused features, use existing classification models such as YOLOv8 to classify the target: a. Harmful targets: people, large animals. b. Harmless targets: medium-sized and smaller animals. Then, according to the preset processing strategy, alarm processing such as light warning, shouting, or pushing information is performed on the current tracking target.

[0073] Preferably, in this embodiment, the UAV arrives at the tracking target, that is, the fusion feature corresponding to the above arrival time, to complete the classification of the target. In other embodiments, the operator can use the UAV to complete the target classification with the fusion feature corresponding to any moment during the tracking of the target. At the same time, during the implementation of tracking and warning the target, key data videos and data are retained as evidence and uploaded to the cloud server.

[0074] The embodiment of the present invention combines an improved long short-term memory network with an attention mechanism reinforcement method to utilize the multi-modal optoelectronic information fusion of fixed videos, radars, and mobile videos. At the stage of screening the target to be recognized, multi-modal data is supplemented and coordinated. At the target tracking stage, the accurate step-by-step iteration of the target position is realized by comparing the fusion features of two adjacent moments, and the target category is determined based on the fusion features. The collaborative supplementation and optimization of multiple deep learning frameworks are realized in multiple stages, and finally, the accuracy and real-time performance of target detection, tracking, and classification can be effectively improved, and high-accuracy recognition and warning of targets in the agricultural insurance scenario can be realized.

[0075] In the second embodiment, the trained additional gate improved long short-term memory network in step S102 specifically includes: Forget gate: , where , and are the parameters of the additional gate improved LSTM network model trained, is the current trajectory state, is the hidden state of the previous time step, and σ is the sigmoid activation function. The overall output range of the forget gate is (0, 1), which determines whether the information before the lost trajectory is useful when the trajectory of the moving object is lost.

[0076] Additional gate: , and the role of this additional gate is to increase 's weight and enhance the recall ability. Where , and are the parameters of the additional gate improved LSTM network model trained. The content of this additional gate in this embodiment actually introduces an additional gate mechanism for enhancing the weight of the historical memory unit . By adding a stridable historical memory control gate outside the traditional LSTM gating structure, this gate dynamically adjusts 's proportion in the current unit , and realizes the enhanced memory of long-distance dependent information. Different from the existing mainstream LSTM variant Peephole LSTM that only focuses on , this solution supports explicit modeling of the memory impact over a longer time interval, enhances the sensitivity and fidelity to long-term dependence features in sequence modeling, ultimately improves the prediction accuracy of the trajectories of moving objects, and constructs a more accurate continuous trajectory.

[0077] Input gate: , , , here determines whether the new information at the current moment is written into the cell state, is the candidate cell state (value range from -1 to 1), is the updated cell state, is the cell state at the previous time step, is the historical cell state, and ⊙ represents element-wise multiplication. The input gate mainly combines the current information and the past information to determine which new information needs to be stored. Among them , , , and , are the parameters of the LSTM network model improved by the additional gates trained.

[0078] Output gate: , , controls the influence of the cell state at the current moment on the output. is the hidden state at the current time step. Among them , and are the parameters of the LSTM network model improved by the additional gates trained.

[0079] Through the in the foregoing embodiments, where is used to train the model to obtain , , , , , , , and , , , , the trained LSTM network model improved by the additional gates is obtained. Further, according to predict to supplement the lost trajectories that cannot be linearly predicted during the construction of the continuous trajectory.

[0080] In this embodiment of the present invention, the cell state update mechanism is improved through an additional gate, enhancing the sensitivity and fidelity of the LSTM network to long-term dependence features during trajectory prediction, and without significantly increasing the computational complexity of the LSTM model during the trajectory prediction process. Ultimately, a more efficient and accurate prediction supplement for the lost trajectory of a moving object can be achieved.

[0081] Corresponding to the method in the above embodiment, Figure 3 The structural block diagram of the farmland monitoring device for multimodal optoelectronic information fusion provided in Embodiment 3 of the present invention is shown. This farmland monitoring device is applied to a computer device, and the computer device is connected to a target database through a preset application programming interface. When the target database is driven to run to execute corresponding tasks, corresponding task logs will be generated, and the above task logs can be collected through the API. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown.

[0082] See Figure 3 , this farmland monitoring device includes: A tracking target selection module 31, configured to determine a first effective target outside a preset distance from the monitoring point by a monitoring point radar, determine a second effective target within the preset distance from the monitoring point by a monitoring point multi-camera, and use the first effective target and the second effective target as tracking targets; A target trajectory construction module 32, configured to use the speed and position of the tracking target at the same moment as the trajectory at that moment, construct a continuous trajectory according to the trajectory, and when the trajectory is lost at the current moment, input the trajectory of the previous moment into a trained long short-term memory network improved by an additional gate to output a predicted trajectory as the trajectory at the current moment; A UAV tracking and warning module 33, configured to after the UAV reaches the tracking target according to the continuous trajectory, track and continuously photograph the tracking target to obtain a tracking image, extract a first type of feature from the tracking image, extract a second type of feature from the continuous trajectory, fuse the first type of feature and the second type of feature at the same moment into a fusion feature by an additive attention mechanism, determine the position of the tracking target at the latter moment in the adjacent two moments according to the similarity of the fusion features at the adjacent two moments to achieve the tracking, and issue a warning after determining the category of the tracking target by the fusion feature.

[0083] Optionally, the tracking target selection module 31 includes: A preset distance determination unit, configured to limit the preset distance to be determined according to the distance between the blind area boundary of the radar and the monitoring point.

[0084] Optionally, the target trajectory construction module 32 includes: An input gate structure limitation unit, configured to limit the update of the cell state in the input gate of the trained long short-term memory network improved by an additional gate is:

[0085] Among them, is the output value of the forget gate, is the cell state at the previous time step, is the activation value of the input gate, is the candidate cell state, is the output value of the additional gate, is the historical cell state, represents element-wise multiplication.

[0086] Optionally, the drone tracking and warning module 33 includes: A feature extraction method limiting unit for limiting the CNN network to extract the first type of feature from the tracking image; Perform local window mapping on the continuous trajectory to obtain a short window mapped trajectory image, and extract the second type of feature from the short window mapped trajectory image using a CNN + attention mechanism network model.

[0087] Optionally, the drone tracking and warning module 33 further includes: A fused feature acquisition unit for limiting the fusion of the first type of feature and the second type of feature at the same time using an additive attention mechanism, including: Use the sigmoid function to determine the weighting of the first type of feature and the weighting of the second type of feature respectively; Based on the weighting of the first type of feature and the weighting of the second type of feature, use the additive attention mechanism to obtain the fused feature.

[0088] Optionally, the drone tracking and warning module 33 further includes: A feature acquisition refinement unit for limiting the weighting of the first type of feature and the weighting of the second type of feature are respectively:

[0089]

[0090] Among them, is the first type of feature, is the second type of feature, and are weights pre-learned through a supervised learning method, is the sigmoid function; The fused feature is:

[0091]

[0092] Among them, is the fused attention weight matrix, is the spatial attention weight pre-learned by the supervised learning method, is the normalized exponential function.

[0093] Optionally, the UAV tracking and warning module 33 further includes: A tracking implementation unit, which is used to determine the position of the tracking target at the latter moment in the two adjacent moments by the similarity of the fused features at two adjacent moments to implement the tracking, including: Taking the position at the latter moment corresponding to the maximum similarity value as the position of the tracking target that the UAV needs to track at the latter moment to implement the tracking.

[0094] It should be noted that the information interaction, execution process, etc. between the above modules, due to being based on the same concept as the method embodiment of the present invention, for its specific functions and the technical effects brought, please specifically refer to the method embodiment part, and will not be elaborated here.

[0095] Figure 4 This is a schematic structural diagram of a computer device provided in Embodiment 4 of the present invention. As Figure 4 shown, the computer device of this embodiment includes: at least one processor ( Figure 4 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, it implements the steps in any of the above-mentioned embodiments of the multi-modal optoelectronic information fusion-based farmland monitoring method.

[0096] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 4 merely an example of a computer device, which does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or combine some components, or different components.

[0097] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0098] The memory includes a readable storage medium, internal memory, etc. Among them, the internal memory may be the memory of the computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be the hard disk of the computer device, and in some other embodiments, it may also be an external storage device of the computer device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory may also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders, data, and other programs, etc. The other programs such as the program code of the computer program, etc. The memory may also be used to temporarily store the data that has been output or will be output.

[0099] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working process of the units and modules in the above device can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiment of the present invention, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiment can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0100] To implement all or part of the processes in the above method embodiment of the present invention, it can also be completed by a computer program product. When the computer program product runs on a computer device, it enables the computer device to execute and implement the steps in the above method embodiment.

[0101] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0102] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0103] In the embodiments provided by the present invention, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0104] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0105] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A farmland monitoring method based on multi-modal optoelectronic information fusion, characterized in that: The method comprises: A first valid target is determined outside a preset distance of the monitoring point by a radar at the monitoring point, a second valid target is determined within the preset distance of the monitoring point by a multi-view camera at the monitoring point, and the first valid target and the second valid target are tracked targets; The speed and position of the target being tracked are recorded at the same time as the trajectory at that moment, a continuous trajectory is constructed based on the trajectory, and when the trajectory at the current moment is lost, the trajectory at the previous moment is input into the trained additional gate improved long short-term memory network output predicted trajectory as the trajectory at the current moment; After the UAV reaches the tracking target according to the continuous trajectory, it tracks and continuously shoots the tracking target to obtain a tracking image, extracts a first type of feature from the tracking image, extracts a second type of feature from the continuous trajectory, fuses the first type of feature and the second type of feature at the same moment into a fused feature by an additive attention mechanism, determines the position of the tracking target at the latter of the two adjacent moments by the similarity of the fused features at two adjacent moments to achieve the tracking, and issues an alarm after determining the category of the tracking target by the fused features.

2. The farmland monitoring method based on multi-modal optoelectronic information fusion according to claim 1 is characterized in that: The preset distance is determined according to the distance between the blind area boundary of the radar and the monitoring point.

3. The farmland monitoring method of multimodal optoelectronic information fusion according to claim 1, characterized in that: The trained additional gate improves the cell state in the input gate of the long short-term memory network for: ,in, is the output value of the forget gate, is the cell state at the previous time step, is the input gate activation value, is the candidate cell state, is the additional gate output value, is the historical cell state, Represents element-wise multiplication.

4. The farmland monitoring method of multimodal optoelectronic information fusion according to claim 1, characterized in that: Extracting the first type of features from the tracking image by a CNN network; The continuous trajectory is subjected to local window mapping to obtain a short window mapping trajectory image, and the two types of features are extracted from the short window mapping trajectory image using a CNN+attention mechanism network model.

5. The farmland monitoring method of multimodal optoelectronic information fusion according to claim 1 or 4, characterized in that: The method of fusing the first type of features and the second type of features at the same time into a fused feature by using an additive attention mechanism includes: Using a sigmoid function to determine the weight of the first type of features and the weight of the second type of features respectively; Based on the weighting of the first type of features and the weighting of the second type of features, an additive attention mechanism is used to obtain the fused features.

6. The farmland monitoring method of multimodal optoelectronic information fusion according to claim 5 is characterized in that: The weight of the feature The weighting of the two features They are: , ,in, is a class of features, is the second-class feature, and are weights pre-learned through supervised learning methods, is the sigmoid function; The fusion feature for: , ,in, is the fused attention weight matrix, is the spatial attention weight pre-learned by supervised learning method, is the normalized exponential function.

7. The farmland monitoring method of multimodal optoelectronic information fusion according to claim 1 or 6, characterized in that: The step of determining the tracking target position at the latter moment of the two adjacent moments by using the similarity of the fused features at the two adjacent moments to achieve the tracking comprises: The position at the next moment corresponding to the maximum similarity value is used as the position of the tracking target to be tracked by the drone at the next moment to achieve the tracking.

8. A farmland monitoring device with multi-modal optoelectronic information fusion, characterized in that: The device comprises: A tracking target selection module, used to determine a first valid target outside a preset distance of the monitoring point by using a radar at the monitoring point, determine a second valid target within the preset distance of the monitoring point by using a multi-view camera at the monitoring point, and take the first valid target and the second valid target as tracking targets; A target trajectory construction module is used to take the speed and position of the tracking target recorded at the same time as the trajectory at that moment, construct a continuous trajectory based on the trajectory, and when the trajectory at the current moment is lost, input the trajectory at the previous moment into the trained additional gate improved long short-term memory network output predicted trajectory as the trajectory at the current moment; The UAV tracking and warning module is used for tracking and continuously photographing the tracking target to obtain a tracking image after the UAV reaches the tracking target according to the continuous trajectory, extracting a first type of feature from the tracking image, extracting a second type of feature from the continuous trajectory, fusing the first type of feature and the second type of feature at the same moment into a fused feature by an additive attention mechanism, determining the position of the tracking target at the latter of the two adjacent moments by the similarity of the fused features at the two adjacent moments to achieve the tracking, and issuing an alarm after determining the category of the tracking target by the fused features.

9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the farmland monitoring method of multimodal optoelectronic information fusion as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the farmland monitoring method of multimodal optoelectronic information fusion as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Target object tracking method and device, computer equipment and storage medium

    CN116681730A

  • Radar parameter adaptive method for multi-target tracking

    CN118707510A

  • Method to detect and manage situations where a large vehicle may hit a vehicle when turning

    US20240416899A1