Farmland Monitoring Method, Device, Equipment and Medium for Multimodal Optoelectronic Information Fusion

Through multimodal photoelectric information fusion, radar and cameras are used to screen targets, build and predict trajectories, and combined with drone image feature fusion, the problem of insufficient target recognition reliability in farmland monitoring is solved, and high-precision target tracking and classification are achieved.

CN120143130BActive Publication Date: 2025-07-11HAI NAN ZHI YUAN KE JI YOU XIAN GONG SI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510615028.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-11
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing farmland monitoring methods are insufficient in the agricultural insurance scenarios to identify targets, making it difficult to cope with complex states and changing environments, resulting in inaccurate target tracking and positioning.

Method used

The multimodal photoelectric information fusion method is adopted, and the tracking target is screened and tracked by radar and multi-eye cameras, continuous trajectories are built and long-term memory networks are improved through additional gates to predict the trajectory. Combined with the image taken by the drone, feature fusion is adopted for feature fusion, and the gradual iterative tracking and classification of the target is achieved.

Benefits of technology

It improves the accuracy and real-time nature of target detection, tracking and classification, and realizes high-accuracy identification and alarms in agricultural insurance scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120143130B_ABST
    Figure CN120143130B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-modal optoelectronic information fusion-based farmland monitoring method, device, equipment and medium, belonging to the technical field of farmland monitoring under radar and video information fusion. The method uses a fixed-point radar and a multi-camera to cooperate to screen and track targets and construct their continuous trajectories. When constructing, a long short-term memory network improved with an additional gate is used to supplement the lost trajectories. Then, a drone approaches the tracked target to take tracking images, and the tracking images and continuous trajectories at the same moment are fused with an additive attention mechanism. Based on the similarity comparison of the fusion features at two adjacent moments, the target is gradually tracked and identified. The present invention combines an improved long short-term memory network and an attention mechanism enhancement method to utilize multi-modal optoelectronic information fusion, determine the tracking target with multi-source data, and gradually determine the target position by comparing the fusion features at adjacent moments, which can effectively improve the accuracy and real-time performance of target detection, tracking and classification, and achieve high-accuracy identification and warning of targets in the agricultural insurance scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of farmland monitoring technology under the fusion of radar and video information, and in particular to a farmland monitoring method, device, equipment and medium under the fusion of multi-modal photoelectric information. Background Art

[0002] With the development of agricultural science and technology, smart agriculture has gradually become an important development direction of modern agriculture. The use of the Internet of Things, artificial intelligence, drones, remote sensing technology, etc. to achieve refined management of farmland environment has become one of the important means of agricultural production. In the subdivision of smart agricultural management, farmland safety monitoring has received more and more attention, especially in application scenarios such as scientific research test fields, high-value crop planting areas, and unmanned farms, which have put forward higher requirements for intrusion monitoring, theft prevention, and wild animal expulsion.

[0003] For farmland safety monitoring, accurate tracking and positioning of intrusion targets is a prerequisite for realizing their type identification and thus completing targeted processing. Although there are relevant technical solutions in the current farmland safety monitoring segment that integrate multiple sensing methods to identify intrusion targets, the way such solutions make comprehensive use of various sensor data is usually weighted fusion with fixed weights.

[0004] For the agricultural protection scenario of farmland safety monitoring, most of the targets to be identified (such as animals and birds) have the characteristics of variable and irregular movement paths and large fluctuations in movement speed; at the same time, most of the protected areas (i.e. farmlands) in the agricultural protection scenario are in outdoor environments, which leads to a large degree of environmental impact on the farmland safety monitoring process. Combining the above two reasons, the existing fixed-weight multi-mode data fusion target tracking method is difficult to cope with the complex state of the targets to be identified and the changeable farmland environment in the agricultural protection scenario due to the lack of adaptive ability and reliable complex target tracking ability, resulting in it being unable to meet the requirements of accurate tracking and positioning of targets in the agricultural protection scenario, and unable to provide accurate and reliable fusion data for target type identification, and ultimately unable to effectively guarantee farmland safety.

[0005] Therefore, the current farmland monitoring methods in agricultural protection scenarios have the technical problem of insufficient target recognition reliability. Summary of the invention

[0006] In view of this, the embodiments of the present invention provide a farmland monitoring method, device, equipment and medium with multi-modal optoelectronic information fusion to solve the technical problem that the farmland monitoring method in the current agricultural protection scenario has insufficient reliability in target recognition.

[0007] In a first aspect, a method for farmland monitoring using multimodal optoelectronic information fusion is provided, the method comprising:

[0008] Determine a first effective target with a monitoring point radar outside the preset distance of the monitoring point, determine a second effective target with a monitoring point multi-camera within the preset distance of the monitoring point, and use the first effective target and the second effective target as tracking targets;

[0009] Take the speed and position of the tracking target at the same time as the trajectory at that moment, construct a continuous trajectory according to the trajectory, and when the trajectory is lost at the current moment, input the trajectory of the previous moment into the trained long short-term memory network improved with an additional gate to output a predicted trajectory as the trajectory at the current moment;

[0010] After the UAV reaches the tracking target according to the continuous trajectory, track and continuously photograph the tracking target to obtain a tracking image, extract a first type of feature from the tracking image, extract a second type of feature from the continuous trajectory, and use an additive attention mechanism to fuse the first type of feature and the second type of feature at the same time into a fusion feature, determine the position of the tracking target at the latter moment in the adjacent two moments with the similarity of the fusion features at the adjacent two moments to achieve the tracking, and determine the category of the tracking target from the fusion feature for processing.

[0011] In a second aspect, a farmland monitoring device for multimodal optoelectronic information fusion is provided, and the device includes:

[0012] A tracking target selection module, configured to determine a first effective target with a monitoring point radar outside the preset distance of the monitoring point, determine a second effective target with a monitoring point multi-camera within the preset distance of the monitoring point, and use the first effective target and the second effective target as tracking targets;

[0013] A target trajectory construction module, configured to take the speed and position of the tracking target at the same time as the trajectory at that moment, construct a continuous trajectory according to the trajectory, and when the trajectory is lost at the current moment, input the trajectory of the previous moment into the trained long short-term memory network improved with an additional gate to output a predicted trajectory as the trajectory at the current moment;

[0014] A UAV tracking and warning module, configured to, after the UAV reaches the tracking target according to the continuous trajectory, track and continuously photograph the tracking target to obtain a tracking image, extract a first type of feature from the tracking image, extract a second type of feature from the continuous trajectory, use an additive attention mechanism to fuse the first type of feature and the second type of feature at the same time into a fusion feature, determine the position of the tracking target at the latter moment in the adjacent two moments with the similarity of the fusion features at the adjacent two moments to achieve the tracking, and give an alarm after determining the category of the tracking target.

[0015] In a third aspect, an embodiment of the present invention provides a computer device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the multi-modal optoelectronic information fusion-based farmland monitoring method as described in the first aspect.

[0016] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the multi-modal optoelectronic information fusion-based farmland monitoring method as described in the first aspect.

[0017] The beneficial effects of the present invention compared with the prior art are as follows:

[0018] In this method of the present invention, first, at a fixed monitoring point, a radar and a multi-camera are used in cooperation to screen and track a target and determine the position and speed of the tracked target. Then, a continuous trajectory of the tracked target is constructed, and during the construction process, a long short-term memory network improved by an additional gate is used to predict and supplement the lost trajectory. After that, a drone is controlled to approach the tracked target to capture its tracking image, and the tracking image of the tracked target and the continuous trajectory at the same moment are fused by an additive attention mechanism. Then, based on the similarity comparison of the fusion features at two adjacent moments, a step-by-step iterative tracking of the target based on multi-modal optoelectronic information fusion is realized, and the target category is determined by the fusion features to complete the alarm. The present invention combines an improved long short-term memory network and an attention mechanism reinforcement method to utilize multi-modal optoelectronic information fusion of fixed video, radar, and mobile video. At the stage of screening the target to be recognized, multi-modal data is supplemented and coordinated. At the stage of target tracking, the accurate step-by-step iteration of the target position is realized through the comparison of the fusion features at two adjacent moments, and the target category is determined based on the fusion features. Finally, the accuracy and real-time performance of target detection, tracking, and classification can be effectively improved, and high-accuracy recognition and alarm of the target in the agricultural protection scenario can be realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts.

[0020] Figure 1 FIG. 1 is a schematic diagram of an application environment of a multi-modal optoelectronic information fusion-based farmland monitoring method provided in Embodiment 1 of the present invention;

[0021] Figure 2 FIG. 2 is a schematic flowchart of a multi-modal optoelectronic information fusion-based farmland monitoring method provided in Embodiment 1 of the present invention;

[0022] Figure 3 It is a schematic structural diagram of a farmland monitoring device for multimodal optoelectronic information fusion provided in Embodiment 3 of the present invention;

[0023] Figure 4 It is a schematic structural diagram of a computer device provided in Embodiment 4 of the present invention. Detailed implementation manners

[0024] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are presented to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.

[0025] It should be understood that when used in the specification and appended claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0026] It should also be understood that the term "and / or" as used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0027] As used in the specification and appended claims of the present invention, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" depending on the context.

[0028] In addition, in the description of the specification and appended claims of the present invention, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0029] References to "one embodiment" or "some embodiments" etc. described in the specification of the present invention mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0030] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0031] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0032] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0033] In order to illustrate the technical solutions of the present invention, the following will be described through specific embodiments.

[0034] A multi-modal optoelectronic information fusion-based farmland monitoring method provided in Embodiment 1 of the present invention can be applied in an application environment such as Figure 1 where the client communicates with the server. Among them, the client includes but is not limited to terminal devices such as palm computers, desktop computers, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, cloud terminal devices, and personal digital assistants (PDAs). The server can be implemented by an independent server or a server cluster composed of multiple servers.

[0035] The above-mentioned farmland monitoring method for multimodal optoelectronic information fusion can be applied to the server in Figure 1. The computer device corresponding to the server is connected to corresponding databases, rule bases, etc. to obtain corresponding data in the databases. The above computer device is connected to the client to collect data and control instructions sent by the client users, as well as to send relevant monitoring parameter data and video data to the client for retention. The above computer device is also connected to the sensing devices to obtain relevant monitoring data collected by the sensing devices. The sensing devices include, but are not limited to, video sensors on unmanned aerial vehicles, spherical cameras at fixed monitoring points, multi-camera systems at fixed monitoring points, and radars at fixed monitoring points.

[0036] See Figure 2 , which is a schematic flowchart of a farmland monitoring method for multimodal optoelectronic information fusion provided by Embodiment 1 of the present invention. The method may include the following steps:

[0037] Step S101, determine a first effective target outside the preset distance of the monitoring point by the monitoring point radar, determine a second effective target within the preset distance of the monitoring point by the monitoring point multi-camera system, and use the first effective target and the second effective target as tracking targets.

[0038] This farmland monitoring method in this embodiment is executed in an unmanned manner. Therefore, to achieve unmanned monitoring, necessary hardware devices need to be deployed in advance. In this embodiment, a vertical pole is set as the bearing platform and denoted as the monitoring point. A radar, a multi-camera system, and a UAV airport required for monitoring are arranged on the vertical pole, which serves as the hardware basis for implementing this farmland monitoring method in this embodiment.

[0039] For unified coordinate reference, in this embodiment, the vertical pole is taken as the center origin, the west-east direction is the x-axis direction, and the north-south direction is the y-axis direction to form a local coordinate system of the monitoring area. All the following observation coordinates are based on this coordinate system.

[0040] In this embodiment, 4 radars are used to form a 360° observation range, and the aperture of each radar is 90°, which is used to detect the monitoring area, that is, the farmland. Although a 360° observation range can be formed by deploying 4 radars, due to the limitations of the working principle of the radar itself, its recognition accuracy for objects relatively close to itself is poor or it cannot recognize them, that is, there is a recognition blind area. To avoid the tracking targets in the blind area obtained by the radar being misdetected targets, this embodiment sets the use of multi-camera on the monitoring point, that is, the pole, to fuse the radar information to supplement the acquisition of tracking targets. For this reason, this embodiment determines a preset distance according to the distance between the blind area boundary of the radar and the monitoring point, takes the area outside the preset distance from the monitoring point as the radar monitoring area, and takes the area within the preset distance from the monitoring point as the multi-camera monitoring area. The radar identifies and screens the moving objects in the radar monitoring area, and the multi-camera identifies and screens the moving objects in the multi-camera monitoring area.

[0041] Specifically, the radar uses the Doppler effect and beam scanning to detect the speed of the moving object in the radar monitoring area , the echo time t, and the azimuth relative to the center of the radar module . Further, the distance = ct / 2 can be obtained, where c represents the speed of light, and the speed is the radial speed according to the Doppler radar principle. Also, the echo intensity p can be described according to the power of the radar received signal. Therefore, the output parameters are the radial speed , position and echo intensity p, where , .

[0042] For the moving objects in the monitoring area, some belong to categories that are clearly determined to pose no threat, such as operators and small objects performing planting and maintenance operations in the farmland. Therefore, for the moving objects detected by the radar in the radar monitoring area, it is first necessary to preliminarily screen the moving objects according to the moving speed and echo intensity of the moving objects obtained by the radar to obtain the first effective target in the radar monitoring area to exclude such targets as operators.

[0043] Then, for the moving objects in the radar monitoring area, the speed threshold and echo intensity threshold methods are used to screen the first effective target Aobj from them. Here, the speed range is set as the minimum radial speed , the maximum radial speed , the minimum echo intensity pmin = 120, and the maximum echo intensity pmax = 200. Assuming that the moving object detected by the radar scan is obj, the first effective target is as follows: Aobj = obj (pmin < p < pmax and < < ) and one or more of the targets can be selected for tracking subsequently.

[0044] The multi-camera uses its own target recognition algorithm to recognize the objects in the monitoring area of the multi-camera, obtaining the vertical pixel number Hpix and horizontal pixel number Wpix of the target. The target recognition algorithm can adopt any existing feasible solution, such as target recognition based on YOLOv8, and determines the moving speed of the object through the moving object tracking algorithm on the multi-camera. The moving object tracking algorithm on the multi-camera can also adopt any existing feasible solution, and automatically adjusts the viewing angle of the multi-camera through the PTZ (Pan-Tilt-Zoom) function to ensure that the objects in the monitoring area of the multi-camera are always at the center of the field of view of the multi-camera. Then, the orientation of the object relative to the multi-camera can be determined according to the rotation angle of the multi-camera. Based on the parameters of the multi-camera and the position of the object in the field of view of the multi-camera, the distance between the object and the monitoring point can be deduced. Thus, according to the distance and the orientation the position of the object monitored by the multi-camera is determined. where .

[0045] Similarly, since some of the moving objects in the monitoring area belong to categories that can be clearly determined to pose no threat, the objects in the monitoring area of the multi-camera also need to be screened. Since the size S and focal length F of the sensor of the multi-camera are both known, the actual physical size of the objects in the monitoring area of the multi-camera can be calculated first: Hreal = Hpix × ri × S / F, Wreal = Wpix × ri × S / F. According to the actual physical size, the objects in the monitoring area of the multi-camera are screened to determine the second effective target: Bobj = obj (Hreal > 50cm and Wreal > 50cm).

[0046] Thus, the tracking targets that need to be processed and judged for the overall monitoring area can be determined according to Bobj and Aobj. Subsequently, the tracking targets determined according to Bobj and Aobj are required to drive and control the take-off of the UAV, and the UAV is provided with the tracking target positions to assist the UAV in planning the route, approaching for observation and carrying out subsequent tracking and warning operations.

[0047] Step S102: Use the speed and position of the tracking target at the same time as the trajectory at that moment. Based on the trajectory, construct a continuous trajectory. When the trajectory is lost at the current moment, input the trajectory of the previous moment into the trained long short-term memory network with an improved additional gate to output a predicted trajectory as the trajectory at the current moment.

[0048] Before controlling the UAV to start tracking and warning the tracking target, this embodiment first constructs the motion trajectory of the tracking target as data for subsequent recognition of the tracking target category. The construction of the motion trajectory uses existing methods. The following is an example of the motion trajectory construction process in this embodiment.

[0049] Specifically, based on the position and speed of the tracking target monitored by the radar and multi-camera at the same time as the trajectory at that moment:

[0050]

[0051] Among them, represents the trajectory of the tracking target at time k, represents the abscissa of the tracking target at time k, represents the ordinate of the tracking target at time k, represents the lateral speed of the tracking target at time k, , represents the longitudinal speed of the tracking target at time k, , represents the vector speed of the tracking target at time k, represents the azimuth angle of the tracking target relative to the monitoring point.

[0052] Based on the trajectory of the tracking target at time k, the predicted trajectory at the next moment, that is, time k + 1, can be predicted:

[0053]

[0054]

[0055]

[0056] Among them, represents the predicted trajectory at time k + 1, is the noise at time k, F is the mapping matrix, is the sampling interval, , , and are the noise corresponding to the abscissa, the noise corresponding to the ordinate, the noise corresponding to the lateral speed, and the noise corresponding to the longitudinal speed at time k, respectively.

[0057] Thus, we have:

[0058]

[0059] Wherein, represents the predicted value of the abscissa at the (k + 1)-th moment, represents the predicted value of the ordinate at the (k + 1)-th moment, represents the horizontal velocity at the (k + 1)-th moment, represents the vertical velocity at the (k + 1)-th moment.

[0060] If there are actual sampling values at the (k + 1)-th moment, it is easy to understand that at this time there should be:

[0061]

[0062] However, there will be an error between the predicted value and the actual sampling value at the actual (k + 1)-th moment. Therefore, it can be determined that:

[0063]

[0064] And the noise obtained at the (k + 1)-th moment is used for the trajectory prediction at the (k + 2)-th moment, thereby completing the trajectory iteration and finally forming a continuous trajectory.

[0065] Meanwhile, during the construction of the continuous trajectory, it is inevitable that there will be a situation where the radar or the multi-camera misses the recognition of the moving object in its respective monitoring area at a certain or certain moments, that is, the trajectory information at a certain or certain moments is lost. Then, if the trajectory at the current moment is lost, the long short-term memory network (LSTM) improved by the additional gate is used to intelligently predict the trajectory of the moving object at the current moment. The long short-term memory network (LSTM) improved by the additional gate consists of an input gate, a forget gate, an additional gate, and an output gate, and is used to capture the dependence relationship in the continuous trajectory information, so as to reconstruct the lost trajectory information and solve the problem of target loss.

[0066] Specifically, that is, first based on the above trajectory construction method, the known trajectory is used as the training set and input into the long short-term memory network improved by the additional gate for training, that is, through the continuous trajectory known or constructed so far up to the M-th moment , wherein , to train the long short-term memory network model improved by the additional gate and obtain the trained long short-term memory network improved by the additional gate. Then, when the trajectory at the current moment is lost, the trajectory of the previous moment is input into the trained long short-term memory network improved by the additional gate, so that the network outputs the predicted value of the trajectory at the current moment, and the predicted value of the trajectory at the current moment is used as the trajectory at the current moment and the continuous trajectory construction continues.

[0067] In step S103, after the UAV reaches the tracking target according to the continuous trajectory, it tracks and continuously captures the tracking target to obtain a tracking image, extracts a first type of feature from the tracking image, extracts a second type of feature from the continuous trajectory, and fuses the first type of feature and the second type of feature at the same moment into a fused feature by an additive attention mechanism. The position of the tracking target at the latter moment in the adjacent two moments is determined by the similarity of the fused features at the adjacent two moments to achieve the tracking, and the category of the tracking target is determined by the fused feature and then an alarm is issued.

[0068] During the construction of the continuous trajectory, the UAV can determine the latest position of the tracking target to be identified and alarmed based on the latest trajectory in the continuous trajectory, so as to control the UAV in the UAV airport to reach the tracking target to be identified and alarmed, and subsequently complete the tracking and identification and alarm of the tracking target by fusing the features of the tracking image information captured by the video sensor on the UAV and the continuous trajectory information corresponding to the tracking target to construct a fused feature.

[0069] Specifically, after the UAV reaches the position where the tracking target is located, it first captures a tracking image of the tracking target through the video sensor on it at the arrival moment, and it is easy to understand that there is also a continuous trajectory corresponding to the tracking target at the arrival moment.

[0070] Then, in order to realize the fusion and utilization of the tracking image information and the continuous trajectory information in this embodiment, first, the existing CNN network (such as ResNet, but not limited to ResNet) is used to extract features from the tracking image as the first type of feature :

[0071]

[0072] Among them, is the first type of feature extracted by the CNN network from the tracking image, represents the mapping function of the CNN network model, is the weight pre-learned by the CNN network, is the tracking image.

[0073] Then, the continuous trajectory obtained by using the above improved LSTM is subjected to local window mapping to obtain a short window mapping trajectory image, and a second type of feature is extracted from the short window mapping trajectory image by a CNN + attention mechanism (such as Yolov8) network model :

[0074]

[0075] Among them, represents the mapping function of the CNN + attention mechanism network model, is the weight value pre-learned by the CNN network, is the short-window mapping trajectory image.

[0076] Then, cross-modal feature fusion is performed on the extracted type-I and type-II features. Since the type-II features are mainly obtained by radar and the type-I features are obtained by video sensors, and the radar and video feature channels are different, we first use the sigmoid function to perform weighted adjustment on the two features:

[0077]

[0078]

[0079] Among them, and are the weighting of type-I features and the weighting of type-II features respectively, and are the weight values pre-learned by the supervised learning method, is the sigmoid function.

[0080] Then, the spatial attention mechanism is strengthened. Specifically, the additive attention mechanism is used to fuse the type-I features and the type-II features to obtain the fused features :

[0081]

[0082]

[0083] Among them, is the spatial attention weight pre-learned by the supervised learning method, is the fused attention weight matrix, is the normalization exponential function.

[0084] Thus, the fused features corresponding to the arrival time when the UAV reaches the tracking target position can be obtained.

[0085] There are two purposes for obtaining the fused features in this embodiment. On the one hand, based on the fused features obtained by fusing the multi-modal optoelectronic information obtained under radar, fixed video and mobile video, target type recognition is performed to improve the recognition accuracy, and warning and driving away are performed based on the target category. On the other hand, since the UAV warning processing needs to be continuously carried out, the fused features containing multi-modal optoelectronic information are also used to complete the continuous and accurate tracking of the UAV for the tracking target.

[0086] Then, to enable the UAV to track the target, after obtaining the fused features corresponding to the arrival time in this embodiment, at the next moment, the UAV continues to capture the tracking image of the target to be tracked, and the short-window mapped trajectory image corresponding to the next moment is obtained by mapping the continuous trajectory of the target to be tracked corresponding to the next moment through a local window. And as the method for obtaining the fused features corresponding to the arrival time described above, based on the tracking image and the short-window mapped trajectory image at the next moment, the fused features corresponding to the next moment are obtained.

[0087] Thus, the fused features at the arrival time and the next moment after the arrival time, two adjacent moments, are obtained. Here, we regard the process of obtaining the fused features as a function. In the Siamese network, the similarity (cross-correlation) between the multi-modal optoelectronic information (X1, Y1) at the previous moment and the multi-modal optoelectronic information (X2, Y2) at the next moment among two adjacent moments is calculated, that is, the similarity between the fused features at these two adjacent moments is calculated:

[0088]

[0089] Among them, is the similarity score, represents the cross-correlation calculation, represents the tracking image at the previous moment among two adjacent moments, represents the tracking image at the next moment among two adjacent moments, represents the short-window mapped trajectory image at the previous moment among two adjacent moments, represents the short-window mapped trajectory image at the next moment among two adjacent moments, represents the fused features at the previous moment among two adjacent moments, represents the fused features at the next moment among two adjacent moments.

[0090] However, it should be noted that since the position of the target to be tracked at the next moment is random, the content of the tracking image of the next moment captured by the UAV actually has multiple possibilities. That is, before the position of the target to be tracked at the next moment is determined, the tracking image of the next moment can be understood as having multiple candidate images, and each candidate image can obtain a corresponding fused feature at the next moment, so each candidate image corresponds to a similarity score.

[0091] That is, since the position of the tracking target at the latter moment among two adjacent moments is not yet determined for the UAV, there are multiple possible values for the above similarity calculation result. It is easy to understand that when the similarity value is the largest, it means that the similarity between the position of the latter moment and the known position of the tracking target at the previous moment among the corresponding two adjacent moments is the largest, that is, the similarity between the distribution of the tracking target in the images at the two moments is the largest. At this time, the position of the latter moment is most likely the actual position of the tracking target at this latter moment, and the above-mentioned candidate image at this latter moment can be recorded as the tracking image at the latter moment and used as the tracking image at the previous moment in the next tracking iteration. Thus, based on the position of the tracking target at the arrival moment, combined with the similarity value of the fusion features between the arrival moment and the next moment after the arrival moment, the position determination of the tracking target at the next moment after the arrival moment can be completed, that is, the position determination of the tracking target at the latter moment among two adjacent moments:

[0092]

[0093] Among them, represents the position of the latter moment among two adjacent moments corresponding to the maximum similarity value of the tracking target, represents determination the position coordinates corresponding to when taking the maximum value.

[0094] Repeating the above process of determining the position of the tracking target by the UAV at the latter moment among two adjacent moments, the tracking of the tracking target by the UAV can be realized starting from the arrival moment when the UAV reaches the position of the tracking target, and the fusion features at each moment can be continuously obtained during the tracking process. Based on the obtained fusion features, the existing classification model such as YOLOv8 is used to classify the target: a. Harmful targets: people, large animals. b. Harmless targets: medium-sized and smaller animals. Then, the current tracking target is alarmed according to the preset processing strategy, such as light warning, shouting, or pushing information.

[0095] Preferably, in this embodiment, the fusion features corresponding to the arrival moment when the UAV reaches the tracking target, that is, the above arrival moment, are used to complete the target classification. In other embodiments, the operator can use the fusion features corresponding to any moment during the tracking process of the UAV for the tracking target to complete the target classification. At the same time, during the implementation of tracking and alarming the tracking target, the key data videos and data are retained and uploaded to the cloud server.

[0096] The embodiments of the present invention combine an improved long short-term memory network with an attention mechanism reinforcement method to utilize the multi-modal optoelectronic information of fixed videos, radars, and mobile videos. In the stage of screening targets to be recognized, multi-modal data is supplemented and coordinated. In the target tracking stage, the accurate step-by-step iteration of the target position is achieved through the comparison of fusion features at two adjacent moments, and the target category is determined based on the fusion features. The cooperation, supplementation, and optimization of multiple deep learning frameworks are realized in multiple stages, and finally, the accuracy and real-time performance of target detection, tracking, and classification can be effectively improved, and high-accuracy recognition and warning of targets in the agricultural insurance scenario can be achieved.

[0097] In the second embodiment, the trained additional gate-improved long short-term memory network in step S102 specifically includes:

[0098] Forget gate: , where , and are the parameters of the additional gate-improved LSTM network model trained, is the current trajectory state, is the hidden state of the previous time step, and σ is the sigmoid activation function. The overall output range of the forget gate is (0, 1), which determines whether the information before the lost trajectory is useful when the trajectory of the moving object is lost.

[0099] Additional gate: , and the role of this additional gate is to increase 's weight and enhance the recall ability. Where , and are the parameters of the additional gate-improved LSTM network model trained. The content of this additional gate in this embodiment actually introduces an additional gate mechanism for enhancing the weight of the historical memory unit . By adding a stridable historical memory control gate outside the traditional LSTM gating structure, this gate dynamically adjusts 's proportion in the current unit , achieving enhanced memory of remote-dependent information. Different from the existing mainstream LSTM variant Peephole LSTM that only focuses on , this solution supports explicit modeling of the memory influence of longer time intervals, enhances the sensitivity and fidelity to long-term dependence features in sequence modeling, and finally improves the accuracy of predicting the trajectory of moving objects and constructs a more accurate continuous trajectory.

[0100] Input gate: , , , here Determine whether the new information at the current moment is written into the cell state, is the candidate cell state (value range -1 to 1), is the updated cell state, is the cell state at the previous time step, is the historical cell state, and ⊙ represents element-wise multiplication. The input gate mainly combines the current information and the past information to determine which new information needs to be stored. Among them 、 、 、 and 、 are the parameters of the LSTM network model improved by the additional gate trained.

[0101] Output gate: , , controls the influence of the cell state at the current moment on the output. is the hidden state at the current time step. Among them 、 and are the parameters of the LSTM network model improved by the additional gate trained.

[0102] Through the in the foregoing embodiments, where to train the model, and obtain 、 、 、 、 、 、 、 and 、 、 、 , to obtain the trained LSTM network model improved by the additional gate. Further, according to predict , and supplement the lost trajectories that cannot be linearly predicted during the construction of the continuous trajectory.

[0103] In this embodiment of the present invention, the cell state update mechanism is improved through an additional gate, so that the sensitivity and fidelity of the LSTM network to long-term dependence features are enhanced during the trajectory prediction process, and at the same time, the computational amount of the LSTM model during the trajectory prediction process is not significantly increased. Finally, a more efficient and accurate prediction and supplement of the lost trajectories of moving objects can be achieved.

[0104] Corresponding to the method in the above embodiment, Figure 3The structure block diagram of the multi-modal optoelectronic information fusion farmland monitoring device provided in the third embodiment of the present invention is shown. This farmland monitoring device is applied to a computer device, and the computer device is connected to a target database through a preset application programming interface. When the target database is driven to run to execute corresponding tasks, corresponding task logs will be generated, and the above task logs can be collected through the API. For the sake of simplicity, only the parts related to the embodiments of the present invention are shown.

[0105] See Figure 3 , the farmland monitoring device includes:

[0106] A tracking target selection module 31, configured to determine a first effective target outside a preset distance from the monitoring point by a monitoring point radar, determine a second effective target within the preset distance from the monitoring point by a monitoring point multi-camera, and use the first effective target and the second effective target as tracking targets;

[0107] A target trajectory construction module 32, configured to use the speed and position of the tracking target at the same time as the trajectory at that time, construct a continuous trajectory according to the trajectory, and when the trajectory at the current moment is lost, input the trajectory at the previous moment into a trained long short-term memory network with an additional gate improvement to output a predicted trajectory as the trajectory at the current moment;

[0108] A drone tracking and warning module 33, configured to after the drone reaches the tracking target according to the continuous trajectory, track and continuously photograph the tracking target to obtain a tracking image, extract a first type of feature from the tracking image, extract a second type of feature from the continuous trajectory, fuse the first type of feature and the second type of feature at the same time into a fusion feature by an additive attention mechanism, determine the position of the tracking target at the latter moment in the adjacent two moments according to the similarity of the fusion features at the adjacent two moments to achieve the tracking, and perform a warning after determining the category of the tracking target by the fusion feature.

[0109] Optionally, the tracking target selection module 31 includes:

[0110] A preset distance determination unit, configured to limit the preset distance to be determined according to the distance between the blind area boundary of the radar and the monitoring point.

[0111] Optionally, the target trajectory construction module 32 includes:

[0112] An input gate structure limitation unit, configured to limit the update of the cell state in the input gate of the trained long short-term memory network with an additional gate improvement to be:

[0113]

[0114] Wherein, is the output value of the forget gate, is the cell state at the previous time step, is the activation value of the input gate, is the candidate cell state, is the output value of the additional gate, is the historical cell state, represents element-wise multiplication.

[0115] Optionally, the UAV tracking and warning module 33 includes:

[0116] A feature extraction method limiting unit, configured to limit the CNN network to extract the first type of feature from the tracking image;

[0117] Perform local window mapping on the continuous trajectory to obtain a short window mapping trajectory image, and extract the second type of feature from the short window mapping trajectory image by using a CNN + attention mechanism network model.

[0118] Optionally, the UAV tracking and warning module 33 further includes:

[0119] A fused feature acquisition unit, configured to limit fusing the first type of feature and the second type of feature at the same time into a fused feature by using an additive attention mechanism, including:

[0120] Use the sigmoid function to respectively determine the weighting of the first type of feature and the weighting of the second type of feature;

[0121] Based on the weighting of the first type of feature and the weighting of the second type of feature, use the additive attention mechanism to obtain the fused feature.

[0122] Optionally, the UAV tracking and warning module 33 further includes:

[0123] A feature acquisition refinement unit, configured to limit the weighting of the first type of feature and the weighting of the second type of feature are respectively:

[0124]

[0125]

[0126] Wherein, is the first type of feature, is the second type of feature, and are weights pre-learned through a supervised learning method, is the sigmoid function;

[0127] The fused feature is:

[0128]

[0129]

[0130] Among them, is the fused attention weight matrix, is the spatial attention weight pre-learned by the supervised learning method, is the normalized exponential function.

[0131] Optionally, the UAV tracking and warning module 33 further includes:

[0132] A tracking implementation unit, configured to determine the position of the tracking target at the latter moment in the two adjacent moments by the similarity of the fused features at two adjacent moments to implement the tracking, including:

[0133] Taking the position at the latter moment corresponding to the maximum similarity value as the position of the tracking target that the UAV needs to track at the latter moment to implement the tracking.

[0134] It should be noted that for the information interaction, execution process, etc. between the above modules, since they are based on the same concept as the method embodiment of the present invention, the specific functions and the technical effects brought thereby can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0135] Figure 4 This is a schematic structural diagram of a computer device provided in Embodiment 4 of the present invention. As Figure 4 shown, the computer device of this embodiment includes: at least one processor ( Figure 4 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, it implements the steps in any of the above-mentioned embodiments of the multi-modal optoelectronic information fusion-based farmland monitoring method.

[0136] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 4 merely examples of computer devices do not constitute a limitation on computer devices. A computer device may include more or fewer components than shown in the figure, or combine some components, or different components.

[0137] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0138] The memory includes a readable storage medium, internal memory, etc. Among them, the internal memory may be the memory of the computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be the hard disk of the computer device, and in some other embodiments, it may also be an external storage device of the computer device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory may also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders, data, and other programs, such as the program code of computer programs. The memory may also be used to temporarily store the data that has been output or will be output.

[0139] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0140] All or part of the processes in the above method embodiments of the present invention can also be completed by a computer program product. When the computer program product runs on a computer device, it causes the computer device to execute and implement the steps in the above method embodiments.

[0141] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0142] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0143] In the embodiments provided by the present invention, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0144] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0145] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A farmland monitoring method for multimodal optoelectronic information fusion, characterized in that, The method includes: Determining a first effective target outside a preset distance of the monitoring point by a monitoring point radar, determining a second effective target within the preset distance of the monitoring point by a monitoring point multi-camera, and using the first effective target and the second effective target as tracking targets; Taking the speed and position of the tracking target at the same time as the trajectory at that moment, constructing a continuous trajectory according to the trajectory, and when the trajectory is lost at the current moment, inputting the trajectory of the previous moment into a trained long short-term memory network with an additional gate improvement to output a predicted trajectory as the trajectory at the current moment; After the unmanned aerial vehicle reaches the tracking target according to the continuous trajectory, tracking and continuously photographing the tracking target to obtain a tracking image, extracting a first type of feature from the tracking image, extracting a second type of feature from the continuous trajectory, fusing the first type of feature and the second type of feature at the same time by an additive attention mechanism into a fusion feature, determining the position of the tracking target at the latter moment in the adjacent two moments by the similarity of the fusion features at the adjacent two moments to achieve the tracking, and determining the category of the tracking target from the fusion feature and then giving an alarm.

2. The multi-modal optoelectronic information fusion-based farmland monitoring method according to claim 1, characterized in that The preset distance is determined according to the distance between the blind area boundary of the radar and the monitoring point.

3. The farmland monitoring method for multimodal optoelectronic information fusion according to claim 1, wherein Updating the cell state in the input gate of the trained long short-term memory network improved with an additional gate is as follows: , where is the output value of the forget gate, is the cell state at the previous time step, is the activation value of the input gate, is the candidate cell state, is the output value of the additional gate, is the historical cell state, represents element-wise multiplication.

4. The farmland monitoring method for multimodal optoelectronic information fusion according to claim 1, wherein Extracting the first type of feature from the tracking image by a CNN network; Performing local window mapping on the continuous trajectory to obtain a short window mapping trajectory image, and extracting the second type of feature from the short window mapping trajectory image by a CNN + attention mechanism network model.

5. The farmland monitoring method for multimodal optoelectronic information fusion according to claim 1 or 4, characterized in that, The fusing the first type of feature and the second type of feature at the same time by an additive attention mechanism into a fusion feature includes: Using a sigmoid function to respectively determine the weighting of the first type of feature and the weighting of the second type of feature; Based on the weighting of the first type of feature and the weighting of the second type of feature, obtaining the fusion feature by an additive attention mechanism.

6. The farmland monitoring method for multimodal optoelectronic information fusion according to claim 5, characterized in that, The weighting of the first type of features and the weighting of the second type of features are respectively: , , where is a type of feature, is a type of feature, and are weights pre - learned through a supervised learning method, is the sigmoid function; The fused feature is as follows: , , where is the fused attention weight matrix, is the spatial attention weight pre-learned by the supervised learning method, is the normalized exponential function.

7. The farmland monitoring method for multimodal optoelectronic information fusion according to claim 1 or 6, characterized in that, The determining the position of the tracking target at the latter moment in the adjacent two moments by the similarity of the fusion features at the adjacent two moments to achieve the tracking includes: Taking the position at the latter moment corresponding to the maximum value of the similarity as the position of the tracking target that the unmanned aerial vehicle needs to track at the latter moment to achieve the tracking.

8. A farmland monitoring device for multimodal optoelectronic information fusion, characterized in that The device includes: A tracking target selection module, configured to determine a first effective target outside a preset distance of the monitoring point by a monitoring point radar, determine a second effective target within the preset distance of the monitoring point by a monitoring point multi-camera, and use the first effective target and the second effective target as tracking targets; A target trajectory construction module, configured to take the speed and position of the tracking target at the same time as the trajectory at that moment, construct a continuous trajectory according to the trajectory, and when the trajectory is lost at the current moment, input the trajectory of the previous moment into a trained long short-term memory network with an additional gate improvement to output a predicted trajectory as the trajectory at the current moment; The UAV tracking and warning module is used for the UAV to track and continuously photograph the tracking target to obtain a tracking image after reaching the tracking target according to the continuous trajectory, extract a first type of feature from the tracking image, extract a second type of feature from the continuous trajectory, fuse the first type of feature and the second type of feature at the same time by an additive attention mechanism into a fused feature, determine the position of the tracking target at the latter moment in the adjacent two moments according to the similarity of the fused features at the adjacent two moments to realize the tracking, and determine the category of the tracking target from the fused feature and then issue a warning.

9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the multi-modal optoelectronic information fusion-based farmland monitoring method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-modal optoelectronic information fusion-based farmland monitoring method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target object tracking method and device, computer equipment and storage medium

    CN116681730A

  • Radar parameter adaptive method for multi-target tracking

    CN118707510A