An AI vision-based intelligent agricultural environment intelligent monitoring method and system

CN122551174APending Publication Date: 2026-08-11SHANXI ZHONGKE TONGCHANG INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

由于缺少面向视觉与传感联合的时间对齐、异常传感处理与稳定融合评估流程,现有方案难以在光照变化、遮挡与天气干扰以及传感器漂移或缺失等情况下,实现作物视觉信息与环境多传感数据的可靠采集、对齐融合与稳定状态评估

Benefits of technology

本发明通过融合农作物图像与环境多传感数据并构建稳定的时序表征,提升了复杂农业现场下的环境与作物状态自动监测能力。与现有以单一传感或单帧视觉为主的监测方式相比,本发明在传感侧通过前向积分与轨迹旋转方向变化提取可靠点集并进行样条衔接重构,降低传感漂移、缺失与突变对状态评估的影响;在视觉侧采用包含环形级联旁路、可逆折纸重组与随机孔洞注入的改进CSPNet网络,以及包含可逆置换编码与时序剪影对抗的对比预测编码,使模型在光照波动、遮挡与天气干扰条件下仍可输出稳定的帧级潜表示与窗口级上下文表征,从而提高监测结果的鲁棒性与一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551174A_ABST
    Figure CN122551174A_ABST
Patent Text Reader

Abstract

This invention discloses a smart agricultural environment intelligent monitoring method and system based on AI vision, belonging to the field of AI vision technology. The method includes: acquiring crop images and environmental sensing data; preprocessing to form an image frame set and a sensing sampling sequence; performing forward integration and spline concatenation to reconstruct a window sensing feature vector; constructing an improved CSPNet network to obtain a frame-level visual latent representation sequence; outputting a window-level visual context representation based on contrastive predictive coding; performing sliding cross-difference search to generate a joint window state vector; performing bidirectional cumulative arc length difference comparison to form a weighted risk level identifier; and constructing an increasing priority queue to issue early warning items and control instructions. This invention, by integrating an improved CSPNet and contrastive predictive coding, achieves stable assessment and early warning control output for crop environmental conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of AI vision technology, and in particular to a smart agricultural environment intelligent monitoring method and system based on AI vision. Background Technology

[0002] Existing smart agriculture environmental monitoring solutions typically rely on IoT sensors to continuously collect data on parameters such as temperature, humidity, light intensity, carbon dioxide concentration, soil moisture, conductivity, and pH. This data is then uploaded to a platform via gateways or wireless communication for data display, threshold alarms, trend analysis, and report output. Some systems also integrate weather stations, insect traps, and irrigation controllers for more comprehensive environmental and production management. However, current solutions are susceptible to field conditions in actual farmland and greenhouse deployments, leading to long-term data drift, intermittent data loss, short-term abrupt changes, and time synchronization issues across multiple channels. Since the environmental assessment and alarm logic of most systems heavily depend on the numerical stability of a single channel or a few key channels, sensor anomalies often trigger false alarms, missed alarms, or significant fluctuations in status assessments, thus affecting the reliability and continuity of control decisions.

[0003] Visual data acquisition in agricultural settings also faces significant uncertainties. In open-air settings, changes in sunlight angle, cloud cover, backlighting, and shadows cause noticeable fluctuations in image brightness and contrast. In facility agriculture settings, the operation of supplemental lighting equipment, reflection from reflective film, and aging of greenhouse film can all lead to drastic changes in light intensity. Factors such as crop leaves shading each other, swaying branches, rain, fog, dust, lens contamination, and condensation can cause localized occlusion and decreased image quality. Existing visual monitoring methods often employ general feature extraction networks or single-frame-based recognition processes, lacking structural designs for organizing cross-stage information flows, recoverable feature representation, and robust characterization under occlusion or missing conditions. Due to the lack of time alignment, abnormal sensor processing, and stable fusion evaluation processes for the joint use of vision and sensing, existing solutions struggle to reliably acquire, align, fuse, and evaluate the stable state of crop visual information and multi-sensor environmental data under conditions of changing light, occlusion, weather interference, and sensor drift or loss.

[0004] Therefore, how to provide a smart agricultural environment intelligent monitoring method and system based on AI vision is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose an intelligent monitoring method and system for smart agriculture environments based on AI vision. This invention comprehensively utilizes the fusion of crop images and multi-sensor data of the environment, reliable point extraction and spline connection reconstruction of sensor sampling sequences, an improved CSPNet network structure, and innovative temporal representation of contrastive predictive coding. This forms a complete process from simultaneous image and sensor acquisition, preprocessing, sensor feature construction, visual latent representation extraction, temporal context generation, alignment and fusion, risk assessment, to early warning and control issuance. Structurally, this invention innovatively employs a ring-cascaded bypass, reversible origami reconstruction, and random hole injection to construct an improved CSPNet network. It uses reversible permutation coding and temporal silhouette adversarial coding to construct contrastive predictive coding, achieving stable state assessment and risk classification output under conditions of light changes, occlusion, weather interference, and sensor drift or absence. Compared with existing technologies, this invention has the advantages of stable monitoring output, strong robustness, outstanding adaptability to sensor anomalies, and ease of on-site deployment and coordinated control.

[0006] According to an embodiment of the present invention, a smart agricultural environment intelligent monitoring method based on AI vision includes: Simultaneously acquire crop images and environmental sensor data, preprocess the crop images and environmental sensor data to form an image frame set and a sensor sampling sequence; Perform forward integration on the sensing sampling sequence to calculate the local rotation direction change of the trajectory to obtain a reliable point set. Use the reliable point set to perform spline connection and reconstruction to obtain the window sensing feature vector. An improved CSPNet network is constructed to extract features from a set of image frames. Based on the ring-cascaded bypass, cross-stage bypass and closed-loop cascaded transmission are performed in the bypass path. Reversible origami reconstruction is used to fold, rearrange and unfold the image. Random hole injection is introduced to perform random spatial hole-cutting and missing pattern injection on the bypass path features during the training phase to obtain a frame-level visual latent representation sequence. Based on contrastive predictive coding, temporal representation generation processing is performed on the frame-level visual latent representation sequence. Channel permutation is performed through reversible permutation coding. Temporal silhouette adversarial approach is used to construct contrastive predictive target update coding parameters with silhouette future latent representation, and output window-level visual context representation. Calculate the first-order change sequence of the window visual context representation and the window sensing feature vector respectively, perform sliding cross-difference search, and generate the window joint state vector in a chessboard staggered manner; A vector chain is constructed based on the window joint state vector. A bidirectional cumulative arc length difference comparison is performed on the vector chain. The directional entropy is calculated on the distribution of changes in adjacent directions of the vector chain to form a weighted risk level identifier. An increasing priority queue is constructed based on the weighted risk level identifier, and early warning items and control instructions are issued sequentially.

[0007] Optionally, forming the image frame set and the sensing sampling sequence includes: Simultaneously acquire crop images and environmental sensor data. Crop images are acquired by the visual acquisition terminal in the same acquisition cycle to form an original image sequence, and environmental sensor data are acquired by multiple sensor channels in the same acquisition cycle to form an original sensor sequence. Image preprocessing is performed on the original image sequence to obtain a set of image frames. The size of each image frame is normalized, the brightness of each image frame is normalized, and the sharpness of each image frame is determined and image frames that do not meet the sharpness determination conditions are deleted. Sensing preprocessing is performed on the original sensing sequence to obtain a sensing sampling sequence. The sampling values ​​of each sensing channel are standardized in units, abnormal values ​​are replaced, missing values ​​are marked, and the sampling values ​​of each sensing channel are aligned and organized according to time markers to form a sensing sampling sequence corresponding to the image frame set.

[0008] Optionally, obtaining the window sensing feature vector includes: The sensor sampling sequence is divided into channel sampling sequences according to the sensor channel, and the forward integration is performed on each channel sampling sequence in chronological order to obtain the integral sequence. The sampled values ​​of the same channel sampling sequence are used as the first coordinate, and the integral values ​​of the corresponding integral sequence are used as the second coordinate to form a planar trajectory point sequence. Adjacent trajectory points are connected in time order to obtain a trajectory line segment sequence. The cross product sign is calculated using the direction vectors of two adjacent trajectory line segments to determine the rotation direction. The number of times the rotation direction changes along the time order is counted to obtain the rotation count sequence. Based on the rotation count sequence, the sampling points of the trajectory point sequence are sorted and the sampling points with the smallest rotation count are selected to form a reliable point set. The time markers and sampling values ​​in the reliable point set are used as control points to perform spline connection reconstruction to obtain the reconstructed sampling sequence. The window sensing feature vector is calculated on the reconstructed sampling sequence within the time window.

[0009] Optionally, obtaining the frame-level visual latent representation sequence includes: An improved CSPNet network is constructed. The improved CSPNet network sets a main path and a bypass path at each stage of the original CSPNet network. A ring-shaped cascaded bypass and a reversible origami reconstruction are set on the bypass path respectively. Random hole injection is set on the training branch of the bypass path. The CSPNet network is improved by inputting the image frame set frame by frame in chronological order. In each stage, the stage input features are divided into main path input features and side path input features by channel. The main path input features are calculated along the main path to obtain the stage main path output features, and the side path input features are passed along the side path to obtain the stage side path intermediate features. The intermediate features of the stage bypass are processed in a ring-shaped cascade bypass. The intermediate features of the stage bypass are used as the current stage bypass output features and are bypassed to the next stage as the bypass input features of the next stage. The current stage bypass output features and the next stage main road input features are merged to obtain merged features. The merged features are sent back to the bypass path as bypass input features to form a cross-stage closed-loop cascade transmission. The bypass features processed by the ring-cascade bypass are subjected to reversible origami recombination. The bypass features are reversibly folded and rearranged according to the folding and rearrangement rules to obtain folded bypass features. Before entering the final stage of merging, the folded bypass features are reversibly unfolded and restored according to the restoration rules corresponding to the folding and rearrangement rules to obtain restored bypass features. During the training phase, random hole injection is performed on the restored bypass features. A set of hole locations is generated in the spatial location of the restored bypass features, and the corresponding elements of the hole location set are set to zero to obtain the hole bypass features. At the same time, a missing mode label corresponding to the hole location set is generated and synchronously transmitted to the end of the phase to merge with the hole bypass features. The phase main output features and hole bypass features are integrated at the end of the phase to obtain the phase output features. The output features of each phase are output sequentially to form a frame-level visual latent representation sequence corresponding to the image frame. The improved CSPNet network is trained by constructing a training sample set and inputting it into the improved CSPNet network in batches. Forward computation is performed to obtain the corresponding frame-level visual latent representation sequence. The comparative prediction training loss is calculated based on the frame-level visual latent representation sequence. Backpropagation is performed to obtain the parameter gradient. The parameters of the improved CSPNet network are iteratively updated based on the parameter gradient.

[0010] Optionally, the output window-level visual context representation includes: Contrastive predictive coding is constructed and coding paths and prediction paths are set. The frame-level visual latent representation sequence is input into the coding path in chronological order to obtain the context representation sequence. The prediction path generates a prediction representation sequence based on the context representation sequence. Perform reversible permutation coding on the frame-level visual latent representation sequence to generate a channel permutation index, and permutate the channel order of the frame-level visual latent representation at future time steps according to the channel permutation index to obtain the permuted future latent representation; Perform temporal silhouette adversarial processing on the permutation future latent representation, converting each element of the permutation future latent representation into a silhouette future latent representation according to the sign function; Contrastive prediction training sample set is constructed based on the predicted representation sequence and the silhouette future latent representation, and the contrastive prediction training loss is calculated. Based on the contrast prediction training loss, backpropagation is performed to update the encoded path parameters and the predicted path parameters, and a window-level visual context representation is output.

[0011] Optionally, the generation of the window joint state vector includes: Calculate the first-order change sequence of the window-level visual context representation in chronological order, and calculate the first-order change sequence of the window-sensing feature vector in chronological order. Within the sliding range, a sliding cross-difference search is performed on the two first-order change sequences. For each sliding displacement, the element-wise difference between the first-order change sequences is calculated and the absolute values ​​of the differences are summed to obtain the cross-difference value. The sliding displacement with the smallest cross-difference value is selected as the alignment displacement. The starting indices of the two sequences are adjusted to obtain the aligned visual change sequence and the aligned sensor change sequence. In the aligned visual change sequence and the aligned sensor change sequence, the change magnitude of each dimension element is calculated and sorted in descending order of change magnitude. Visual change elements and sensor change elements are selected to form visual candidate vectors and sensor candidate vectors, respectively. The visual candidate vectors and sensor candidate vectors are interleaved in a chessboard pattern while keeping the time index consistent to obtain the window joint state vector.

[0012] Optionally, the formation of the weighted risk level identifier includes: Arrange the window joint state vectors of adjacent time windows in chronological order to obtain a vector chain, and calculate the Euclidean distance sequence between adjacent vectors in the vector chain. A bidirectional cumulative arc length difference comparison is performed on the vector chain to obtain the radian balance sequence. The forward cumulative arc length and the backward cumulative arc length are calculated for each position vector in the vector chain. The radian balance sequence is composed of the absolute value of the difference between the forward cumulative arc length and the backward cumulative arc length of each position vector. The position vector corresponding to the minimum value of the radian balance sequence is selected as the balance center vector. For adjacent vectors in the vector chain, calculate the direction vector and the sequence of angles between adjacent direction vectors. Divide the angle sequence into angle intervals and count the occurrence frequency of each angle interval to obtain the direction distribution. The direction entropy is obtained by taking the logarithm of the proportion of the occurrence frequency of each angle interval to the total frequency, multiplying the proportion, summing the results, and taking the negative value. Calculate the Euclidean distance from each vector in the vector chain to the equilibrium center vector to obtain the center distance sequence. Combine the direction entropy and the mean of the center distance sequence to generate a weighted risk level label.

[0013] Optionally, the sequential issuance of early warning items and control instructions includes: Based on the weighted risk level identifier, risk entries are generated for each time window, and the risk entries are sorted in descending order of weighted risk level identifier to form an ascending priority queue. Read queue records sequentially from the head to the tail of the increasing priority queue, combine the weighted risk level identifier and credibility information corresponding to each queue record to generate an early warning entry, and write the early warning entry into the early warning output channel. For each early warning item, a control instruction is generated and sent to the control execution terminal. The execution receipt returned by the control execution terminal is received, and the execution result identifier and execution completion timestamp in the execution receipt are extracted to form a log record.

[0014] According to an embodiment of the present invention, a smart agricultural environment intelligent monitoring system based on AI vision includes the following modules: The data acquisition and preprocessing module is used to simultaneously acquire crop images and environmental sensor data and complete preprocessing, outputting a set of image frames and a sensor sampling sequence. The sensor feature construction module is used to perform forward integration on the sensor sampling sequence, perform spline connection reconstruction, and output a window sensor feature vector. The visual latent representation generation module is used to construct an improved CSPNet network and extract features from a set of image frames to generate a frame-level visual latent representation sequence. The temporal context generation module is used to generate temporal representations of frame-level visual latent representation sequences based on contrastive predictive coding, and outputs window-level visual context representations. The fusion alignment module is used to calculate the first-order change sequence of the window's visual context representation and the window's sensing feature vector, and generate the window's joint state vector. The risk assessment module is used to construct a vector chain based on the window joint state vector and output a weighted risk level identifier; The early warning and control module is used to construct an incremental priority queue based on the weighted risk level identifier, and to complete the distribution and log recording of monitoring results.

[0015] The beneficial effects of this invention are: This invention enhances the automatic monitoring capability of environment and crop status in complex agricultural fields by fusing crop images and multi-sensor environmental data and constructing stable temporal representations. Compared with existing monitoring methods that rely primarily on single sensors or single-frame vision, this invention extracts reliable point sets and performs spline connection reconstruction on the sensing side through forward integration and trajectory rotation direction changes, reducing the impact of sensor drift, missing data, and abrupt changes on status assessment. On the vision side, it employs an improved CSPNet network that includes ring cascaded bypasses, reversible origami reconstruction, and random hole injection, as well as contrastive predictive coding that includes reversible permutation coding and temporal silhouette adversarial coding. This enables the model to output stable frame-level latent representations and window-level contextual representations even under conditions of light fluctuations, occlusion, and weather interference, thereby improving the robustness and consistency of monitoring results.

[0016] The alignment fusion and risk classification output process proposed in this invention achieves effective alignment of vision and sensing within a time window, construction of joint state vectors, and generation of risk level identifiers based on vector chain arc length difference and direction entropy. Furthermore, it uses an increasing priority queue to sequentially issue warning items and control instructions, and records logs in a closed loop. This overcomes the shortcomings of existing technologies, such as difficulty in synchronizing multi-source data, volatile fusion results, and lack of continuity and executability in risk output. It enhances the engineering feasibility and efficiency of coordinated control under on-site deployment conditions, providing support for continuous monitoring, risk warning, and precision management in smart agriculture. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of an intelligent agricultural environment monitoring method based on AI vision proposed in this invention; Figure 2 This is a structural block diagram of the improved CSPNet network for an AI vision-based intelligent agricultural environment monitoring method proposed in this invention. Figure 3 This is a functional diagram of an AI vision-based intelligent agricultural environment monitoring system proposed in this invention. Detailed Implementation

[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0019] refer to Figure 1 and Figure 2 A smart agricultural environment intelligent monitoring method based on AI vision, comprising: Simultaneously acquire crop images and environmental sensor data, preprocess the crop images and environmental sensor data to form an image frame set and a sensor sampling sequence; Perform forward integration on the sensing sampling sequence to calculate the local rotation direction change of the trajectory to obtain a reliable point set. Use the reliable point set to perform spline connection and reconstruction to obtain the window sensing feature vector. An improved CSPNet network is constructed to extract features from a set of image frames. Based on the ring-cascaded bypass, cross-stage bypass and closed-loop cascaded transmission are performed in the bypass path. Reversible origami reconstruction is used to fold, rearrange and unfold the image. Random hole injection is introduced to perform random spatial hole-cutting and missing pattern injection on the bypass path features during the training phase to obtain a frame-level visual latent representation sequence. Based on contrastive predictive coding, temporal representation generation processing is performed on the frame-level visual latent representation sequence. Channel permutation is performed through reversible permutation coding. Temporal silhouette adversarial approach is used to construct contrastive predictive target update coding parameters with silhouette future latent representation, and output window-level visual context representation. Calculate the first-order change sequence of the window visual context representation and the window sensing feature vector respectively, perform sliding cross-difference search, and generate the window joint state vector in a chessboard staggered manner; A vector chain is constructed based on the window joint state vector. A bidirectional cumulative arc length difference comparison is performed on the vector chain. The directional entropy is calculated on the distribution of changes in adjacent directions of the vector chain to form a weighted risk level identifier. An increasing priority queue is constructed based on the weighted risk level identifier, and early warning items and control instructions are issued sequentially.

[0020] In this embodiment, forming the image frame set and the sensing sampling sequence includes: Simultaneously acquire crop images and environmental sensor data. Crop images are acquired by the visual acquisition terminal in the same acquisition cycle to form an original image sequence, and environmental sensor data are acquired by multiple sensor channels in the same acquisition cycle to form an original sensor sequence. Image preprocessing is performed on the original image sequence to obtain a set of image frames. The size of each image frame is normalized, the brightness of each image frame is normalized, and the sharpness of each image frame is determined and image frames that do not meet the sharpness determination conditions are deleted. Sensing preprocessing is performed on the original sensing sequence to obtain a sensing sampling sequence. The sampling values ​​of each sensing channel are standardized in units, abnormal values ​​are replaced, missing values ​​are marked, and the sampling values ​​of each sensing channel are aligned and organized according to time markers to form a sensing sampling sequence corresponding to the image frame set.

[0021] In this embodiment, obtaining the window sensing feature vector includes: The sensing sampling sequence is divided into channel sampling sequences according to the sensing channels. For each channel sampling sequence, forward integration is performed in chronological order to obtain the integral sequence. Specifically, the process of performing forward integration on each channel sampling sequence in chronological order to obtain the integral sequence is as follows: For each sensing channel, the sampled values ​​are first arranged from early to late according to the time identifier to obtain the corresponding sampling time and sampled value. The first integral value of the integral sequence is set to zero. Starting from the second sampling point, the time difference between the current sampling time and the previous sampling time is calculated point by point. The arithmetic mean of the current sampled value and the previous sampled value is taken as the representative value of the current time period. The representative value is multiplied by the time difference to obtain the integral increment of the time period. The integral increment is added to the previous integral value to obtain the current integral value. The calculation is repeated in time order until the last sampling point. All integral values ​​are arranged into an integral sequence according to the corresponding time. Using the sampled values ​​of the same channel sampling sequence as the first coordinate and the integral values ​​of the corresponding integration sequence as the second coordinate, a planar trajectory point sequence is constructed. Adjacent trajectory points are connected in chronological order to obtain a trajectory segment sequence. The cross product sign is calculated using the direction vectors of two adjacent trajectory segments to determine the rotation direction. The number of changes in adjacent rotation directions along the chronological order is counted to obtain a rotation count sequence. Specifically, the cross product sign is calculated using the direction vectors of two adjacent trajectory segments to determine the rotation direction. The planar trajectory points obtained in chronological order are connected sequentially to form trajectory segments. For any intermediate trajectory point, the direction vector of the previous line segment is taken as the difference between the first and second coordinate differences of the current point and the previous point. The direction vector of the next line segment is taken as the difference between the first and second coordinate differences of the next point and the current point. The cross product value of the two direction vectors is calculated according to the two-dimensional cross product rule. The first coordinate difference of the previous direction vector is multiplied by the second coordinate difference of the next direction vector, and then the second coordinate difference of the previous direction vector is multiplied by the first coordinate difference of the next direction vector. When the cross product value is greater than zero, the rotation direction is determined to be counterclockwise. When the cross product value is less than zero, the rotation direction is determined to be clockwise. When the absolute value of the cross product value is less than or equal to a preset threshold, it is determined to be no rotation. The preset threshold is the product of one-hundredth of the median length of all trajectory line segments in this window and one-hundredth of the median absolute value of the integral sequence in this window, which is used to eliminate false rotation judgments caused by numerical noise. Based on the rotation count sequence, the sampling points of the trajectory point sequence are sorted, and the sampling points with the smallest rotation count are selected to form a reliable point set. Using the time markers and sample values ​​in the reliable point set as control points, spline stitching is performed to reconstruct the reconstructed sampling sequence. The window sensing feature vector is then calculated within a time window on the reconstructed sampling sequence, where: Using the time markers and sampled values ​​in the reliable point set as control points, spline connection reconstruction is performed to obtain the reconstructed sampling sequence. Specifically, the reliable point set is sorted from earliest to latest according to the time markers to form several pairs of adjacent control points. A piecewise cubic spline curve is constructed between each pair of adjacent control points so that the curve passes through the sampled values ​​of the control points at each control point. At the connection of adjacent segments, the curve simultaneously satisfies the continuity of values, the continuity of the first rate of change, and the continuity of the second rate of change. Natural boundary treatment is used at the boundaries. The second rate of change at the earliest and latest control points is set to zero. Using the time markers of the original sensing sampling sequence as the reconstruction time axis, the corresponding curve value is calculated on the spline segment to which each reconstruction time is located, thus obtaining the reconstructed sampling sequence corresponding to the original time marker. The window sensing feature vector is calculated within the time window for the reconstructed sampling sequence. Specifically, all reconstructed sampling values ​​are selected within the time window, and the mean, standard deviation, range, and mean and standard deviation of the difference between adjacent points are calculated. The difference between adjacent points is obtained by subtracting the previous sampling value from the subsequent sampling value. The linear rate of change of the sampling values ​​within the window is calculated. The slope of the linear rate of change is obtained by least-squares linear fitting with the time identifier as the independent variable and the reconstructed sampling value as the dependent variable. At the same time, the number of times the reconstructed sampling value within the window exceeds the mean plus or minus three times the standard deviation is counted and used as an outlier count. The mean, standard deviation, range, mean of the difference, standard deviation of the difference, linear rate of change, and outlier count are concatenated in order to form the window sensing feature vector.

[0022] In this embodiment, obtaining the frame-level visual latent representation sequence includes: An improved CSPNet network is constructed. The improved CSPNet network sets a main path and a bypass path at each stage of the original CSPNet network. A ring-shaped cascaded bypass and a reversible origami reconstruction are set on the bypass path respectively. Random hole injection is set on the training branch of the bypass path. The CSPNet network is improved by inputting image frames sequentially in chronological order. At each stage, the stage input features are divided into main path input features and side path input features by channel. The main path input features are used to calculate the stage main path output features along the main path, and the side path input features are used to propagate along the side path to obtain the stage side path intermediate features. Where: In each stage, the stage input features are divided into main path input features and side path input features by channel. Specifically, the total number of channels of the stage input features is determined, and the channels are fixedly divided into two groups according to the channel index from smallest to largest. The first half of the channels form the main path input features, and the second half of the channels form the side path input features. The ratio of the number of channels in the main path to the number of channels in the side path is set to 1:1. When the total number of channels is odd, the extra channel is incorporated into the main path input features. The channel division is consistent in the training and inference stages and is executed independently once in each stage to obtain the main path input features and side path input features of the stage. The main road input features are calculated along the main road path to obtain the stage main road output features, and the bypass input features are passed along the bypass path to obtain the stage bypass intermediate features. Specifically, the main road input features are sequentially passed through the feature extraction calculation unit configured in the stage to obtain the stage main road output features. The feature extraction calculation unit consists of a linear transformation layer, a normalization layer, and a nonlinear activation layer connected in sequence. Residual superposition is performed within the stage to form the stage main road output features. When the bypass input features are passed along the bypass path, the bypass input features do not enter the main road feature extraction calculation unit within the stage. Instead, only shape alignment processing is performed to make the spatial size consistent with the main road output features before being output as the stage bypass intermediate features. Shape alignment is achieved by downsampling or upsampling the bypass input features. When the main road path performs downsampling in the stage, the bypass path performs downsampling synchronously. When the main road path does not change its spatial size, the bypass path maintains its original spatial size. The intermediate features of the stage bypass are processed in a ring-shaped cascade bypass. The intermediate features of the stage bypass are used as the current stage bypass output features and are bypassed to the next stage as the bypass input features of the next stage. The current stage bypass output features and the next stage main road input features are merged to obtain merged features. The merged features are sent back to the bypass path as bypass input features to form a cross-stage closed-loop cascade transmission. The bypass features processed by the ring-cascade bypass are subjected to reversible origami reconstruction. The bypass features are reversibly folded and rearranged according to the folding and rearrangement rules to obtain folded bypass features. Before merging at the end of the stage, the folded bypass features are reversibly unfolded and restored according to the restoration rules corresponding to the folding and rearrangement rules to obtain restored bypass features, wherein: The bypass features are reversibly folded and rearranged according to the folding and rearrangement rules to obtain folded bypass features. Specifically, the spatial size and number of channels of the bypass features are read, the folding factor is set to two and fixed and does not change with training, the spatial position of the bypass features is divided into two-dimensional blocks according to every two adjacent rows and every two adjacent columns, the values ​​of all channels are taken out sequentially for the four positions in each two-dimensional block, and written into the new channel arrangement in order, so that the values ​​of different spatial positions in the same two-dimensional block are moved into different channel segments. At the same time, the spatial size is reduced by half in both the row direction and the column direction to obtain folded bypass features. When the number of rows or columns of the bypass features is odd, zero padding is used in the last row or the last column to make it even before folding and rearranging. The amplitude threshold of zero padding is set to zero. Before entering the final stage of merging, the folded bypass feature is reversibly unfolded and restored according to the restoration rule corresponding to the folding and rearrangement rule to obtain the restored bypass feature. Specifically, the spatial size and number of channels of the folded bypass feature are read, and the same folding factor of two as during folding and rearrangement is used to divide the channels of the folded bypass feature into four equal segments, which correspond to the values ​​from the four spatial positions of the two-dimensional block during folding and rearrangement. For each spatial position of the folded bypass feature, the corresponding values ​​are taken from the four channel segments in turn and written back to the four spatial positions of the unfolded two-dimensional block in the reverse order of folding and rearrangement, so that the spatial size is doubled in both the row and column directions and restored to the size before folding, thus obtaining the restored bypass feature. During the training phase, random hole injection is performed on the restored bypass features. A set of hole locations is generated in the spatial location of the restored bypass features, and the corresponding elements in the hole location set are set to zero to obtain the hole bypass features. At the same time, a missing pattern label corresponding to the hole location set is generated and synchronously passed to the hole bypass features until the end of the stage for merging. The stage main output features and hole bypass features are integrated at the end of the stage to obtain the stage output features. The output features of each stage are output sequentially to form a frame-level visual latent representation sequence corresponding to the image frame, where: To generate a set of hole locations in the spatial location of the restored bypass features, the following steps are taken: read the number of rows and columns of the restored bypass features, use the spatial grid of the number of rows and columns as the candidate location set, set the hole ratio to 15% and keep it fixed and does not change with training, calculate the number of locations to be hollowed out according to the hole ratio, use a pseudo-random number generator to extract the corresponding number of spatial locations without replacement from the candidate location set as the hole location set, limit the hole location set to non-overlapping square hole blocks, set the side length of the square hole blocks to four spatial units, when the extracted square hole block overlaps with an existing hole block, discard the hole block and extract it again; Generate missing pattern identifiers corresponding to the hole location set. Specifically, create a binary identifier matrix of the same size as the spatial row and column size of the restored bypass feature. Set all initial values ​​of the identifier matrix to zero. For each spatial location covered by the hole location set, set the value of the identifier matrix position to one. Positions not covered by the hole location set are kept to zero. When the hole location set uses square hole blocks, set all positions inside the corresponding square hole block in the identifier matrix to one. Use the binary identifier matrix as the missing pattern identifier and bind it to the hole bypass feature with the same time index and stage index. The improved CSPNet network is trained by constructing a training sample set and inputting it into the improved CSPNet network in batches. Forward computation is performed to obtain the corresponding frame-level visual latent representation sequence. The contrastive prediction training loss is calculated based on the frame-level visual latent representation sequence. Backpropagation is performed to obtain the parameter gradients. The parameters of the improved CSPNet network are iteratively updated based on the parameter gradients. Specifically, the contrastive prediction training loss is calculated based on the frame-level visual latent representation sequence as follows: The frame-level visual latent representations within the same time window in each batch are arranged into a sequence in chronological order and input into the contrast prediction encoding to obtain the corresponding context representation and prediction representation. For each context moment, the future frame-level visual latent representation corresponding to the context moment in time is selected as a positive contrast sample. A preset number of future frame-level visual latent representations are extracted from other time windows in the same batch or other moments in the same time window as negative contrast samples. The preset number is 256. The dot product similarity between the prediction representation and the positive contrast sample and the dot product similarity between the prediction representation and each negative contrast sample are calculated respectively. Each similarity is divided by the temperature coefficient and then an exponential operation is performed. The temperature coefficient is 0.07. The probability value is obtained by dividing the positive contrast similarity index by the sum of the positive contrast similarity index and all negative contrast similarity indexes. The natural logarithm of the probability value is taken and negative as the loss value at the context time step. The contrast prediction training loss of the window is obtained by averaging the loss values ​​of all context time steps within the window. The contrast prediction training loss of the current batch is obtained by averaging the training losses of all windows within the batch.

[0023] In this embodiment, the output window-level visual context representation includes: Contrastive predictive coding is constructed, and coding and prediction paths are set. The frame-level visual latent representation sequence is input into the coding path in chronological order to obtain a context representation sequence. The prediction path then generates a prediction representation sequence based on the context representation sequence, where: Contrastive predictive coding is constructed and coding and prediction paths are set. Specifically, the input is defined as a sequence of frame-level visual latent representations arranged in chronological order. The coding path uses a sequential cumulative structure to process the sequence of frame-level visual latent representations step by step. The context state is initialized as a zero vector. At the first time step, the frame-level visual latent representation is written into the context state. At each of the remaining time steps, the current frame-level visual latent representation is concatenated with the context state of the previous time step and a candidate state is obtained through linear transformation. The candidate state and the context state of the previous time step are updated and synthesized into the current context state by gating, resulting in a context representation sequence composed of context states at each time step. The prediction path generates a prediction representation sequence based on the context representation sequence. Specifically, for each time-time context representation in the context representation sequence, the matching prediction path parameters are first selected according to the prediction target corresponding to the time. The context representation is then linearly transformed to obtain a prediction vector. The prediction vector is then normalized to make the magnitude one. The prediction vectors obtained at each time time are arranged in chronological order to form a prediction representation sequence. A reversible permutation coding process is performed on the frame-level visual latent representation sequence to generate a channel permutation index. The channel order of the future frame-level visual latent representation is then permuted according to this index to obtain the permuted future latent representation. Specifically, the reversible permutation coding process on the frame-level visual latent representation sequence is as follows: The total number of channels in the frame-level visual latent representation is read and a channel index sequence from one to the total number of channels is generated. At the beginning of each training batch, the channel index sequence is shuffled once without replacement using a pseudo-random number generator to obtain the channel permutation index. For each frame-level visual latent representation that needs to be used as a sample for future time moments, the channel order is rearranged according to the channel permutation index, so that the first channel after permutation is taken from the channel pointed to by the original channel permutation index, the second channel after permutation is taken from the second channel pointed to by the original channel permutation index, and so on until all channels are rearranged to obtain the permuted future latent representation. Perform temporal silhouette adversarial processing on the permutation future latent representation, converting each element of the permutation future latent representation into a silhouette future latent representation according to a sign function, where: Temporal silhouette adversarial processing is performed on the permutation future latent representation. Specifically, the permutation future latent representation is used as an amplitude representation, and a silhouette representation of the same dimension is constructed. In the same training iteration, the amplitude representation and the silhouette representation are used to participate in the calculation of the contrast prediction training loss. When calculating the loss, the prediction representation is kept unchanged, so that the prediction representation faces both the contrast target with amplitude information retained and the contrast target with amplitude information removed. The loss values ​​obtained from the two contrast targets are averaged to obtain the adversarial loss value at the future time step, so that the encoding path and the prediction path can take into account both amplitude stability and structural stability during training. The elements of the permutation future latent representation are converted into silhouette future latent representations according to the sign function. Specifically, the value of each element in the permutation future latent representation is judged one by one. When the element value is greater than zero, the silhouette value of the corresponding element is assigned to one. When the element value is equal to zero, the silhouette value of the corresponding element is assigned to zero. When the element value is less than zero, the silhouette value of the corresponding element is assigned to negative one. After replacing all elements, a silhouette future latent representation with the same dimension as the permutation future latent representation is obtained. Contrastive prediction training sample set is constructed based on the predicted representation sequence and the silhouette future latent representation, and the contrastive prediction training loss is calculated. Based on the contrast prediction training loss, backpropagation is performed to update the encoded path parameters and the predicted path parameters, and a window-level visual context representation is output.

[0024] In this embodiment, the generation of the window joint state vector includes: The first-order change sequence is calculated in chronological order for the window-level visual context representation, and the first-order change sequence is calculated in chronological order for the window sensing feature vector. Specifically, the calculation of the first-order change sequence involves: The window-level visual context representations within the same time window are arranged into a vector sequence from early to late according to the time identifier. The window sensing feature vectors are arranged into a vector sequence according to the same time identifier. For any vector sequence, the first-order change vector is calculated moment by moment starting from the second moment. The first-order change vector is obtained by subtracting the vector of the adjacent previous moment from the vector of the current moment element by element. The obtained first-order change vector is bound to the time identifier of the current moment. The element-wise subtraction is repeated for all moments in the sequence to obtain a set of first-order change vectors, which are then arranged in order of time identifier to form the corresponding visual first-order change sequence and sensing first-order change sequence. Within the sliding range, a sliding cross-difference search is performed on the two first-order change sequences. For each sliding displacement, the element-wise difference between the first-order change sequences is calculated, and the absolute values ​​of the differences are summed to obtain the cross-difference value. The sliding displacement with the smallest cross-difference value is selected as the alignment displacement. The starting indices of the two sequences are adjusted to obtain the aligned visual change sequence and the aligned sensor change sequence. Specifically, the sliding cross-difference search is performed on the two first-order change sequences as follows: The sliding range is determined to be no more than three time steps forward and backward. The visual first-order change sequence is used as the reference sequence and the sensor first-order change sequence is used as the alignment sequence. The sliding displacement is traversed from negative three to positive three. A negative displacement indicates that the alignment sequence is shifted forward and a positive displacement indicates that the alignment sequence is shifted backward. After each displacement, the overlapping part of the two sequences is taken and paired time step by time. The first-order change vectors of each pair are subtracted element by element to obtain the difference vector. The absolute values ​​of each element of the difference vector are taken and summed to obtain the time difference. Then, the differences of all overlapping time steps are accumulated to obtain the cross difference value of the displacement. The cross difference values ​​corresponding to all displacements are compared and the displacement with the smallest cross difference value is selected as the alignment displacement. The starting index of the two sequences is extracted and matched one-to-one within the overlapping time range to form the aligned visual change sequence and the aligned sensor change sequence. In the aligned visual change sequence and the aligned sensor change sequence, the change magnitude of each dimension element is calculated and sorted in descending order of change magnitude. Visual change elements and sensor change elements are selected to form visual candidate vectors and sensor candidate vectors, respectively. These candidate vectors are then interleaved in a chessboard pattern while maintaining consistent time indices to obtain the window joint state vector. Specifically, the window joint state vector is as follows: The amplitude of change in both the aligned visual change sequence and the aligned sensor change sequence is statistically analyzed by dimension within the same time window. The amplitude of change is obtained by summing the absolute values ​​of the first-order change values ​​of the current dimension at all times within the window. The amplitude of change in each dimension of the visual side is sorted from largest to smallest and the first sixteen dimensions are selected. The amplitude of change in each dimension of the sensor side is sorted from largest to smallest and the first sixteen dimensions are selected. At each time index within the window, the corresponding element is extracted from the aligned visual change vector according to the selected dimension to form a visual candidate vector. The corresponding element is extracted from the aligned sensor change vector according to the selected dimension to form a sensor candidate vector. A joint vector is constructed according to a fixed interleaving rule. The odd-numbered positions of the joint vector are filled with the elements of the visual candidate vector, and the even-numbered positions are filled with the elements of the sensor candidate vector, so that the visual elements and sensor elements in the joint vector are arranged in an interleaved manner. The interleaved insertion of each time index within the window is repeated and the vectors are spliced ​​in the order of the time index to obtain the window joint state vector corresponding to the time window.

[0025] In this embodiment, the formation of the weighted risk level identifier includes: Arrange the window joint state vectors of adjacent time windows in chronological order to obtain a vector chain, and calculate the Euclidean distance sequence between adjacent vectors in the vector chain. A radian balance sequence is obtained by performing a bidirectional cumulative arc length difference comparison on the vector chain. For each position vector in the vector chain, the forward cumulative arc length and the backward cumulative arc length are calculated. The radian balance sequence is composed of the absolute values ​​of the differences between the forward and backward cumulative arc lengths of each position vector. The position vector corresponding to the minimum value of the radian balance sequence is selected as the balance center vector. Specifically, the radian balance sequence is obtained by performing a bidirectional cumulative arc length difference comparison on the vector chain as follows: The vector chain is numbered sequentially by time, and the Euclidean distance between two adjacent position vectors is calculated as the adjacent arc length. The Euclidean distance is obtained by summing the squares of the differences between corresponding elements of the two vectors and taking the square root. For any position vector in the vector chain, the forward cumulative arc length is obtained by accumulating the adjacent arc length between the current position vector and the previous position vector in reverse time to the head of the chain. The backward cumulative arc length is obtained by accumulating the adjacent arc length between the position vector and the next position vector in forward time to the tail of the chain. The radian balance value is calculated for the position vector. The radian balance value is equal to the absolute value of the difference between the forward cumulative arc length and the backward cumulative arc length. The radian balance value is calculated sequentially for all position vectors in the vector chain and arranged in chronological order to obtain the radian balance sequence. The position vector with the smallest value in the radian balance sequence is selected as the balance center vector. For adjacent vectors in the vector chain, calculate the direction vector and the sequence of angles between adjacent direction vectors. Divide the angle sequence into angle intervals and count the occurrences of each interval to obtain the direction distribution. The direction entropy is obtained by taking the logarithm of the proportion of occurrences of each angle interval to the total number of occurrences, multiplying the proportions, summing the results, and taking the negative value. Calculate the Euclidean distance from each vector in the vector chain to the equilibrium center vector to obtain the center distance sequence. Combine the direction entropy and the mean of the center distance sequence to generate a weighted risk level label. Specifically, the calculation of the direction vector and the sequence of angles between adjacent direction vectors in the vector chain is as follows: Take the difference between two adjacent position vectors in the vector chain in chronological order. Subtract the previous position vector from the next position vector element by element to obtain the direction vector of the adjacent segment. Form a direction vector sequence by taking all the direction vectors in chronological order. Calculate the angle between two adjacent direction vectors in the direction vector sequence. First, calculate the dot product of the two direction vectors. The dot product is obtained by multiplying the corresponding elements and then summing them. Calculate the magnitude of the two direction vectors. The magnitude is obtained by summing the squares of the corresponding elements and then taking the square root. Divide the dot product by the product of the two magnitudes to obtain the cosine value. If the cosine value is greater than one, take one; if the cosine value is less than negative one, take negative one. Take the inverse cosine of the cosine value to obtain the angle value. Form a sequence of angle values ​​by taking all the angle values ​​obtained in chronological order.

[0026] In this embodiment, the sequential issuance of early warning items and control instructions includes: Based on the weighted risk level identifier, risk entries are generated for each time window, and the risk entries are sorted in descending order of weighted risk level identifier to form an ascending priority queue. Read queue records sequentially from the head to the tail of the increasing priority queue, combine the weighted risk level identifier and credibility information corresponding to each queue record to generate an early warning entry, and write the early warning entry into the early warning output channel. For each early warning item, a control instruction is generated and sent to the control execution terminal. The execution receipt returned by the control execution terminal is received, and the execution result identifier and execution completion timestamp in the execution receipt are extracted to form a log record.

[0027] refer to Figure 3 A smart agricultural environment monitoring system based on AI vision includes the following modules: The data acquisition and preprocessing module is used to simultaneously acquire crop images and environmental sensor data and complete preprocessing, outputting a set of image frames and a sensor sampling sequence. The sensor feature construction module is used to perform forward integration on the sensor sampling sequence, perform spline connection reconstruction, and output a window sensor feature vector. The visual latent representation generation module is used to construct an improved CSPNet network and extract features from a set of image frames to generate a frame-level visual latent representation sequence. The temporal context generation module is used to generate temporal representations of frame-level visual latent representation sequences based on contrastive predictive coding, and outputs window-level visual context representations. The fusion alignment module is used to calculate the first-order change sequence of the window's visual context representation and the window's sensing feature vector, and generate the window's joint state vector. The risk assessment module is used to construct a vector chain based on the window joint state vector and output a weighted risk level identifier; The early warning and control module is used to construct an incremental priority queue based on the weighted risk level identifier, and to complete the distribution and log recording of monitoring results.

[0028] Example 1: To verify the feasibility of this invention in practice, it was applied to a continuous agricultural monitoring cycle. The system received crop images and environmental multi-sensor data from a fixed visual acquisition terminal. A total of 960 images were collected, with a resolution of 1920×1080 pixels, covering parts of the crop canopy and leaves. The images exhibited light fluctuations and occlusion, with a maximum-to-minimum ratio of 3.6 for the average brightness and an average contrast ratio of 0.24. Approximately 31% of the images showed occlusion by leaves or supports, and approximately 12% showed localized blurring due to fog or water droplets. The environmental sensor data included six channels: air temperature, air humidity, light intensity, carbon dioxide concentration, soil moisture, and soil conductivity. The sensor sampling sequence totaled 288 sampling points. The soil moisture channel showed one segment of drift with a baseline shift of approximately 9.2%, and two missing segments with lengths of 7 and 11 sampling points respectively. The light intensity channel showed three instantaneous jumps, with jump amplitudes of 4.1 times, 3.7 times, and 4.5 times the normal fluctuation amplitude, respectively.

[0029] After the data enters the processing flow, the image and sensor data are preprocessed simultaneously to form an image frame set and a sensor sampling sequence. On the image side, pixel values ​​are normalized to 0 to 1, and size and brightness are normalized. Sharpness is used as the selection criterion, calculated using the average image gradient amplitude. 286 image frames with a sharpness metric below 0.018 are removed. After preprocessing, the average contrast ratio increases from 0.24 to 0.33, and the standard deviation of the average brightness of low-light frames decreases from 0.19 to 0.11. On the sensor side, the 6-channel data undergoes unit standardization and missing data marking, and is aligned according to time markers to form a sensor sampling sequence. Missing points are only marked with missing data and are not directly filled in.

[0030] Forward integration is performed on the sensor sampling sequence to construct a trajectory point sequence for extracting a reliable point set. For each channel, the integral increment is obtained by multiplying the arithmetic mean of adjacent sample values ​​by the time difference in chronological order, and these increments are accumulated to form an integral sequence. Planar trajectory points are constructed using the sample values ​​and integral values, and connected to form line segments. A two-dimensional cross product is calculated for the direction vectors of adjacent line segments. A cross product greater than 0 indicates counterclockwise rotation, less than 0 indicates clockwise rotation, and the absolute value of the cross product not exceeding a threshold indicates no rotation. The threshold is the product of 1% of the median line segment length and 1% of the median absolute value of the integral. Rotation counts are obtained based on the number of rotation direction changes, and the smallest count is selected to form the reliable point set. The window sensing feature vector consists of the mean, standard deviation, range, mean of adjacent differences, standard deviation of adjacent differences, linear rate of change, and outlier count. The outlier count is determined by adding or subtracting three times the standard deviation from the mean. In this embodiment, the mean outlier count is reduced from 2.7 to 0.9.

[0031] The image frame set is fed into an improved CSPNet to extract frame-level visual latent representation sequences. At each stage, the input features are divided into a main path and a bypass path in a 1:1 channel ratio. The main path completes the stage feature extraction, while the bypass path enters a circular cascaded bypass path and reversible origami reconstruction. The circular cascaded bypass path allows the bypass features to bypass to the next stage and merge with the main path input of the next stage, and the merged features are fed back to form a closed loop. The reversible origami reconstruction uses a folding factor of 2, folding adjacent 2×2 spatial block information into the channel and unfolding it back to its original state before merging according to the reverse rule. In this embodiment, the bypass features are folded from 120×68×128 to 60×34×512, and then restored to 120×68×128. During the training stage, random holes are injected with a hole ratio of 15%, the hole block side length is 4, and the hole blocks do not overlap. During the inference stage, the holes are not set to zero but occlusion robustness is preserved. The frame-level output is a 256-dimensional latent representation. The mean Euclidean distance between the latent representations of adjacent frames in the bright-to-dark transition segment is 0.84 and the standard deviation is 0.19; while the traditional conventional CSPNet has a mean distance of 1.12 and a standard deviation of 0.41.

[0032] Frame-level visual latent representation sequences are fed into contrastive predictive coding to generate window-level visual context representations. Latent representations are input into the coding path in chronological order to obtain a context representation sequence, and the prediction path generates a prediction representation sequence based on the context representation. Reversible permutation coding generates a channel permutation index once per training batch. Latent representations at future time steps are rearranged according to the index to obtain permuted future latent representations, and the inverse permutation index is saved. Temporal silhouette adversarial coding transforms the permuted future latent representations element-wise into silhouette future latent representations and constructs a contrast target with the prediction representation. The contrastive prediction training loss uses 256 negative samples, a temperature coefficient of 0.07, and dot product similarity. Positive samples are silhouette future latent representations corresponding to the same window time, and negative samples are silhouette future latent representations extracted from other windows in the same batch. After training convergence, in windows with an occlusion ratio exceeding 30%, the mean absolute value of the first-order change in the context representation is 0.27; the average value for traditional ordinary contrastive predictive coding is 0.46.

[0033] The system calculates the first-order change sequence for both the window's visual context representation and the window's sensing feature vector. It performs a sliding cross-difference search within three time steps before and after the change, calculating the absolute sum of the element-by-element differences in the overlapping area for each displacement and taking the minimum as the alignment displacement (commonly +2 or +3 in this embodiment). After alignment, the system selects the top 16 dimensions of each change magnitude to form visual and sensing candidate vectors, which are then synthesized into a window joint state vector according to a chessboard staggered rule. The average cross-difference value decreases from 38.4 to 21.7, compared to 35.9 for traditional timestamp alignment. Subsequently, a vector chain is constructed using the window joint state vector. The adjacent Euclidean distance is calculated as the arc length, and the equilibrium center vector is determined by the bidirectional cumulative arc length difference. The angle between adjacent direction vectors is calculated, and the directional entropy is obtained by statistically analyzing the angle interval distribution. Combined with the mean of the center distance sequence, a weighted risk level identifier is generated and written into an increasing priority queue to generate warning entries and control instructions, and log recording is completed.

[0034] To verify the beneficial effects, this embodiment compares the present invention with the traditional method. The training samples are consistent: 800 images for training, 80 for validation, and 80 for testing. The sensor samples are arranged according to windows. In the training set, windows with large illumination fluctuations account for 35%, occluded windows account for 30%, and sensor missing or drifting windows account for 18%. The traditional method uses conventional CSPNet and ordinary contrastive predictive coding, with linear interpolation on the sensor side filling in the missing parts and directly aligning and fusing them. On the same test set, the anomaly event recognition accuracy is 81.6% for the traditional method and 90.4% for the present invention; the recall rate is 74.2% for the traditional method and 88.1% for the present invention; the stable output rate for sensor missing windows is 86.3% for the traditional method and 100% for the present invention; the standard deviation of the strong light to shadow window state evaluation is 0.39 for the traditional method and 0.16 for the present invention; the number of false alarms is 19 for the traditional method and 7 for the present invention; the number of risk level spikes for windows with occlusion exceeding 30% is 14 for the traditional method and 5 for the present invention.

[0035] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A smart agricultural environment intelligent monitoring method based on AI vision, characterized in that, include: Simultaneously acquire crop images and environmental sensor data, preprocess the crop images and environmental sensor data to form an image frame set and a sensor sampling sequence; Perform forward integration on the sensing sampling sequence to calculate the local rotation direction change of the trajectory to obtain a reliable point set. Use the reliable point set to perform spline connection and reconstruction to obtain the window sensing feature vector. An improved CSPNet network is constructed to extract features from a set of image frames. Based on the ring-cascaded bypass, cross-stage bypass and closed-loop cascaded transmission are performed in the bypass path. Reversible origami reconstruction is used to fold, rearrange and unfold the image. Random hole injection is introduced to perform random spatial hole-cutting and missing pattern injection on the bypass path features during the training phase to obtain a frame-level visual latent representation sequence. Based on contrastive predictive coding, temporal representation generation processing is performed on the frame-level visual latent representation sequence. Channel permutation is performed through reversible permutation coding. Temporal silhouette adversarial approach is used to construct contrastive predictive target update coding parameters with silhouette future latent representation, and output window-level visual context representation. Calculate the first-order change sequence of the window visual context representation and the window sensing feature vector respectively, perform sliding cross-difference search, and generate the window joint state vector in a chessboard staggered manner; A vector chain is constructed based on the window joint state vector. A bidirectional cumulative arc length difference comparison is performed on the vector chain. The directional entropy is calculated on the distribution of changes in adjacent directions of the vector chain to form a weighted risk level identifier. An increasing priority queue is constructed based on the weighted risk level identifier, and early warning items and control instructions are issued sequentially.

2. The intelligent agricultural environment monitoring method based on AI vision according to claim 1, characterized in that, The process of forming the image frame set and the sensing sampling sequence includes: Simultaneously acquire crop images and environmental sensor data. Crop images are acquired by the visual acquisition terminal in the same acquisition cycle to form an original image sequence, and environmental sensor data are acquired by multiple sensor channels in the same acquisition cycle to form an original sensor sequence. Image preprocessing is performed on the original image sequence to obtain a set of image frames. The size of each image frame is normalized, the brightness of each image frame is normalized, and the sharpness of each image frame is determined and image frames that do not meet the sharpness determination conditions are deleted. Sensing preprocessing is performed on the original sensing sequence to obtain a sensing sampling sequence. The sampling values ​​of each sensing channel are standardized in units, abnormal values ​​are replaced, missing values ​​are marked, and the sampling values ​​of each sensing channel are aligned and organized according to time markers to form a sensing sampling sequence corresponding to the image frame set.

3. The intelligent monitoring method for smart agricultural environment based on AI vision according to claim 1, characterized in that, The obtained window sensing feature vector includes: The sensor sampling sequence is divided into channel sampling sequences according to the sensor channel, and the forward integration is performed on each channel sampling sequence in chronological order to obtain the integral sequence. The sampled values ​​of the same channel sampling sequence are used as the first coordinate, and the integral values ​​of the corresponding integral sequence are used as the second coordinate to form a planar trajectory point sequence. Adjacent trajectory points are connected in time order to obtain a trajectory line segment sequence. The cross product sign is calculated using the direction vectors of two adjacent trajectory line segments to determine the rotation direction. The number of times the rotation direction changes along the time order is counted to obtain the rotation count sequence. Based on the rotation count sequence, the sampling points of the trajectory point sequence are sorted and the sampling points with the smallest rotation count are selected to form a reliable point set. The time markers and sampling values ​​in the reliable point set are used as control points to perform spline connection reconstruction to obtain the reconstructed sampling sequence. The window sensing feature vector is calculated on the reconstructed sampling sequence within the time window.

4. The intelligent monitoring method for smart agricultural environment based on AI vision according to claim 1, characterized in that, The obtained frame-level visual latent representation sequence includes: An improved CSPNet network is constructed. The improved CSPNet network sets a main path and a bypass path at each stage of the original CSPNet network. A ring-shaped cascaded bypass and a reversible origami reconstruction are set on the bypass path respectively. Random hole injection is set on the training branch of the bypass path. The CSPNet network is improved by inputting the image frame set frame by frame in chronological order. In each stage, the stage input features are divided into main path input features and side path input features by channel. The main path input features are calculated along the main path to obtain the stage main path output features, and the side path input features are passed along the side path to obtain the stage side path intermediate features. The intermediate features of the stage bypass are processed in a ring-shaped cascade bypass. The intermediate features of the stage bypass are used as the current stage bypass output features and are bypassed to the next stage as the bypass input features of the next stage. The current stage bypass output features and the next stage main road input features are merged to obtain merged features. The merged features are sent back to the bypass path as bypass input features to form a cross-stage closed-loop cascade transmission. The bypass features processed by the ring-cascade bypass are subjected to reversible origami recombination. The bypass features are reversibly folded and rearranged according to the folding and rearrangement rules to obtain folded bypass features. Before entering the final stage of merging, the folded bypass features are reversibly unfolded and restored according to the restoration rules corresponding to the folding and rearrangement rules to obtain restored bypass features. During the training phase, random hole injection is performed on the restored bypass features. A set of hole locations is generated in the spatial location of the restored bypass features, and the corresponding elements of the hole location set are set to zero to obtain the hole bypass features. At the same time, a missing mode label corresponding to the hole location set is generated and synchronously transmitted to the end of the phase to merge with the hole bypass features. The phase main output features and hole bypass features are integrated at the end of the phase to obtain the phase output features. The output features of each phase are output sequentially to form a frame-level visual latent representation sequence corresponding to the image frame. The improved CSPNet network is trained by constructing a training sample set and inputting it into the improved CSPNet network in batches. Forward computation is performed to obtain the corresponding frame-level visual latent representation sequence. The comparative prediction training loss is calculated based on the frame-level visual latent representation sequence. Backpropagation is performed to obtain the parameter gradient. The parameters of the improved CSPNet network are iteratively updated based on the parameter gradient.

5. The intelligent monitoring method for smart agricultural environment based on AI vision according to claim 1, characterized in that, The output window-level visual context representation includes: Contrastive predictive coding is constructed and coding paths and prediction paths are set. The frame-level visual latent representation sequence is input into the coding path in chronological order to obtain the context representation sequence. The prediction path generates a prediction representation sequence based on the context representation sequence. Perform reversible permutation coding on the frame-level visual latent representation sequence to generate a channel permutation index, and permutate the channel order of the frame-level visual latent representation at future time steps according to the channel permutation index to obtain the permuted future latent representation; Perform temporal silhouette adversarial processing on the permutation future latent representation, converting each element of the permutation future latent representation into a silhouette future latent representation according to the sign function; Contrastive prediction training sample set is constructed based on the predicted representation sequence and the silhouette future latent representation, and the contrastive prediction training loss is calculated. Based on the contrast prediction training loss, backpropagation is performed to update the encoded path parameters and the predicted path parameters, and a window-level visual context representation is output.

6. The intelligent monitoring method for smart agricultural environment based on AI vision according to claim 1, characterized in that, The generated window joint state vector includes: Calculate the first-order change sequence of the window-level visual context representation in chronological order, and calculate the first-order change sequence of the window-sensing feature vector in chronological order. Within the sliding range, a sliding cross-difference search is performed on the two first-order change sequences. For each sliding displacement, the element-wise difference between the first-order change sequences is calculated and the absolute values ​​of the differences are summed to obtain the cross-difference value. The sliding displacement with the smallest cross-difference value is selected as the alignment displacement. The starting indices of the two sequences are adjusted to obtain the aligned visual change sequence and the aligned sensor change sequence. In the aligned visual change sequence and the aligned sensor change sequence, the change magnitude of each dimension element is calculated and sorted in descending order of change magnitude. Visual change elements and sensor change elements are selected to form visual candidate vectors and sensor candidate vectors, respectively. The visual candidate vectors and sensor candidate vectors are interleaved in a chessboard pattern while keeping the time index consistent to obtain the window joint state vector.

7. The intelligent monitoring method for smart agricultural environment based on AI vision according to claim 1, characterized in that, The formation of the weighted risk level identifier includes: Arrange the window joint state vectors of adjacent time windows in chronological order to obtain a vector chain, and calculate the Euclidean distance sequence between adjacent vectors in the vector chain. A bidirectional cumulative arc length difference comparison is performed on the vector chain to obtain the radian balance sequence. The forward cumulative arc length and the backward cumulative arc length are calculated for each position vector in the vector chain. The radian balance sequence is composed of the absolute value of the difference between the forward cumulative arc length and the backward cumulative arc length of each position vector. The position vector corresponding to the minimum value of the radian balance sequence is selected as the balance center vector. For adjacent vectors in the vector chain, calculate the direction vector and the sequence of angles between adjacent direction vectors. Divide the angle sequence into angle intervals and count the occurrence frequency of each angle interval to obtain the direction distribution. The direction entropy is obtained by taking the logarithm of the proportion of the occurrence frequency of each angle interval to the total frequency, multiplying the proportion, summing the results, and taking the negative value. Calculate the Euclidean distance from each vector in the vector chain to the equilibrium center vector to obtain the center distance sequence. Combine the direction entropy and the mean of the center distance sequence to generate a weighted risk level label.

8. The intelligent monitoring method for smart agricultural environment based on AI vision according to claim 1, characterized in that, The sequential issuance of early warning items and control instructions includes: Based on the weighted risk level identifier, risk entries are generated for each time window, and the risk entries are sorted in descending order of weighted risk level identifier to form an ascending priority queue. Read queue records sequentially from the head to the tail of the increasing priority queue, combine the weighted risk level identifier and credibility information corresponding to each queue record to generate an early warning entry, and write the early warning entry into the early warning output channel. For each early warning item, a control instruction is generated and sent to the control execution terminal. The execution receipt returned by the control execution terminal is received, and the execution result identifier and execution completion timestamp in the execution receipt are extracted to form a log record.

9. A smart agricultural environment intelligent monitoring system based on AI vision, comprising executing the smart agricultural environment intelligent monitoring method based on AI vision as described in any one of claims 1 to 8, characterized in that, Includes the following modules: The data acquisition and preprocessing module is used to simultaneously acquire crop images and environmental sensor data and complete preprocessing, outputting a set of image frames and a sensor sampling sequence. The sensor feature construction module is used to perform forward integration on the sensor sampling sequence, perform spline connection reconstruction, and output a window sensor feature vector. The visual latent representation generation module is used to construct an improved CSPNet network and extract features from a set of image frames to generate a frame-level visual latent representation sequence. The temporal context generation module is used to generate temporal representations of frame-level visual latent representation sequences based on contrastive predictive coding, and outputs window-level visual context representations. The fusion alignment module is used to calculate the first-order change sequence of the window's visual context representation and the window's sensing feature vector, and generate the window's joint state vector. The risk assessment module is used to construct a vector chain based on the window joint state vector and output a weighted risk level identifier; The early warning and control module is used to construct an incremental priority queue based on the weighted risk level identifier, and to complete the distribution and log recording of monitoring results.