Rice disease recognition method based on sparse human visual perception and hypergraph reasoning

CN122821540APending Publication Date: 2026-09-25JILIN AGRI SCI & TECH COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610843406.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

现有方法通常依赖于数据分布中的显式模式或静态先验,难以有效挖掘复杂场景中隐式的结构关系,例如不同病害区域之间的空间语义关联以及通道特征之间的高阶依赖关系

Benefits of technology

通过在骨干网络深层设置仿生视觉感知模块,对浅层细节特征进行仿生视觉感知,生成仿生视觉感知特征,能够模拟人类视觉的注视与周边衰减特性,实现对粗细粒度特征的动态协同提取,有效增强模型对细微病斑、弱特征及小目标的感知能力,降低漏检风险。通过颈部结构中的通道关系超图推理和空间关系超图推理,对仿生视觉感知特征进行多方向空间重排与超图卷积,能够挖掘深层语义特征与浅层纹理特征之间的高阶通道依赖关系以及像素节点之间的空间语义关联,克服传统方法难以建模隐式结构关系的缺陷,提升多尺度特征的表达能力。检测头基于增强后的多尺度特征进行分类预测和位置回归预测,在保持较低模型参数量的同时提高了识别精度与召回率,在复杂农田背景、光照变化及图像退化等严苛条件下仍保持较高的鲁棒性,为水稻病害的精准监测提供了可靠的技术支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821540A_ABST
    Figure CN122821540A_ABST
Patent Text Reader

Abstract

The application discloses a rice disease identification method based on sparse human visual perception and hypergraph reasoning, and relates to the technical field of rice disease identification.In the application, a biomimetic visual perception module is arranged in the deep layer of a backbone network to perform biomimetic visual perception on shallow detail features, and biomimetic visual perception features are generated, so that the fixation and peripheral attenuation characteristics of human vision can be simulated, and the risk of missed detection can be reduced.The biomimetic visual perception features are subjected to multidirectional spatial rearrangement and hypergraph convolution through channel relationship hypergraph reasoning and spatial relationship hypergraph reasoning in the neck structure, so that the high-order channel dependency relationship between deep semantic features and shallow texture features and the spatial semantic association between pixel nodes can be mined, and the expression capability of multi-scale features can be improved.The detection head performs classification prediction and position regression prediction based on the enhanced multi-scale features, the recognition accuracy and recall rate are improved, and reliable technical support is provided for accurate monitoring of rice diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rice disease identification technology, and in particular to a rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning. Background Technology

[0002] Accurate identification of rice diseases is crucial for ensuring food security and sustainable agricultural development. Traditional manual inspection methods are inefficient and highly subjective, making them unsuitable for large-scale rice field monitoring. With the development of computer vision technology and intelligent inspection robots, image-based automatic identification methods have become an important means of achieving precise perception of rice diseases. However, rice diseases in complex field environments are characterized by subtle and unevenly distributed features, and diseased areas often exhibit small targets, weak features, and high similarity to the background, posing significant challenges to high-precision identification. Existing identification models struggle to effectively simulate the human visual system's ability to perceive disease details and perform complex high-order reasoning, resulting in insufficient mining of latent symptom information and a high risk of missed detections and misjudgments.

[0003] In automatic recognition research, object detection is a core tool in image recognition tasks. Previous methods primarily relied on low-level cues such as image texture features, edge information, and frequency domain transformations for preprocessing, and on manually defined features for object classification and localization. These methods are not only inefficient but also limited in performance under complex scenarios. In contrast, current mainstream solutions are mostly based on deep neural networks, which can automatically learn high-order semantic features in images through an end-to-end approach, significantly improving the accuracy of object recognition. For example, Guo Lifeng et al. introduced the advanced single-stage object detection model YOLO into the field of rice disease detection, achieving an average detection accuracy of 90.8% while maintaining a lightweight model. Peng Huilin et al. improved the object detection model by introducing an attention mechanism, increasing various performance indicators by about 15%, verifying the effectiveness of the attention mechanism in this field. Wang et al., starting from the importance of multi-scale feature extraction, effectively improved detection performance by optimizing the model's neck structure and adjusting the loss function. Overall, existing research has made significant progress in improving both detection accuracy and model efficiency.

[0004] While existing methods have made some progress in feature extraction and model structure optimization, most studies are still based on convolutional or Transformer frameworks, focusing on local feature enhancement or global information modeling, lacking the ability to deeply model the potential relationships between targets. Existing methods often rely on explicit patterns or static priors in the data distribution, making it difficult to effectively mine implicit structural relationships in complex scenes, such as spatial semantic associations between different disease areas and high-order dependencies between channel features. In addition, current models differ from the human visual system in their visual perception mechanisms. Human vision has a central high-resolution focus and a gradual decay in peripheral perception when focusing on a target, while existing models struggle to achieve accurate perception and efficient inference of fine-grained disease features. Summary of the Invention

[0005] This invention aims to overcome the shortcomings of the prior art, and specifically provides the following technical solution: 1) In a first aspect, the present invention provides a method for identifying rice diseases based on sparse anthropomorphic visual perception and hypergraph reasoning, the specific technical solution of which is as follows: A rice disease identification model is constructed, comprising a backbone network, a neck structure, and a detection head. The backbone network receives rice sample images, and its shallow layers perform convolution operations on the images to output shallow detail features. A biomimetic visual perception module is located in the deep layers of the backbone network to receive and perform biomimetic visual perception on the shallow detail features, generating biomimetic visual perception features. The neck structure receives these features and performs relational hypergraph inference and spatial relational hypergraph inference to generate multi-scale features. The detection head performs classification prediction and position regression prediction on the multi-scale features to obtain the disease identification results for the rice sample images. The rice disease identification model is then trained. A target rice image is input into the trained model, and the model outputs the disease identification results for the target rice image.

[0006] The beneficial effects of the rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning provided by this invention are as follows: By incorporating a biomimetic visual perception module deep within the backbone network, biomimetic visual perception is achieved on shallow detail features, generating biomimetic visual perception features. This simulates human visual fixation and peripheral attenuation characteristics, enabling dynamic collaborative extraction of coarse and fine-grained features. This effectively enhances the model's ability to perceive subtle lesions, weak features, and small targets, reducing the risk of missed detections. Through channel relationship hypergraph reasoning and spatial relationship hypergraph reasoning in the neck structure, multi-directional spatial rearrangement and hypergraph convolution are performed on the biomimetic visual perception features. This allows for the discovery of high-order channel dependencies between deep semantic features and shallow texture features, as well as spatial semantic associations between pixel nodes. This overcomes the limitations of traditional methods in modeling implicit structural relationships and improves the expressive power of multi-scale features. The detection head performs classification and position regression predictions based on the enhanced multi-scale features. While maintaining a low model parameter count, it improves recognition accuracy and recall. It also maintains high robustness under harsh conditions such as complex farmland backgrounds, varying lighting, and image degradation, providing reliable technical support for the accurate monitoring of rice diseases.

[0007] Based on the above scheme, the rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning of the present invention can be further improved as follows.

[0008] Furthermore, the biomimetic visual perception module is specifically used for: receiving shallow detail features, dividing the shallow detail features into multiple window regions, selecting multiple key regions with the highest relevance for each query region by calculating the correlation between window regions, and establishing a region routing path; extracting query features and key features based on the region routing path, calculating the Euclidean distance between the query position coordinates and the key position coordinates and generating a distance matrix, and weighting and fusing the convolution mapping result of the distance matrix with the attention weights to output the biomimetic visual perception features.

[0009] The beneficial effects of adopting the above-mentioned further scheme are as follows: By dividing the shallow detail features into multiple window regions and calculating the correlation between regions, and selecting the most relevant key regions for each query region to establish a regional routing path, key regions can be adaptively filtered, reducing interference from invalid information. After extracting query features and key features based on the regional routing path, the Euclidean distance between the query location coordinates and the key location coordinates is calculated and a distance matrix is ​​generated. The convolutional mapping result of the distance matrix is ​​then weighted and fused with the attention weights, so that the attention weights are simultaneously constrained by feature similarity and spatial distance, simulating the peripheral attenuation characteristics of human vision. The biomimetic visual perception features output by this process enhance the spatial coordination between local details and global context, effectively improving the ability to identify diseases under conditions of small targets, weak features, and background interference.

[0010] Furthermore, the neck structure includes a channel relationship hypergraph inference module and a spatial relationship hypergraph inference module. The channel relationship hypergraph inference module performs multi-directional spatial rearrangement of deep semantic features in biomimetic visual perception features and shallow texture features in shallow detail features, calculates a similarity matrix, generates a first hyperedge matrix based on the similarity matrix, and performs hypergraph convolution to output channel relationship-enhanced fusion features. The spatial relationship hypergraph inference module flattens the shallow detail features into a sequence of pixel nodes, generates a hyperedge prototype based on the global context of the pixel node sequence, calculates the correlation between each pixel node and the hyperedge prototype to generate a second hyperedge matrix, and performs hypergraph convolution to output spatial relationship-enhanced detail features. The neck structure generates multi-scale features based on the channel relationship-enhanced fusion features and the spatial relationship-enhanced detail features.

[0011] The beneficial effects of adopting the above-mentioned further scheme are as follows: By using the channel relationship hypergraph inference module to perform multi-directional spatial rearrangement of deep semantic features and shallow texture features, calculating the similarity matrix and generating a first hyperedge matrix for hypergraph convolution, it is possible to uncover higher-order dependencies between channels, eliminate redundant information, and output channel-relationship-enhanced fusion features. By using the spatial relationship hypergraph inference module to flatten shallow detail features into pixel node sequences, generating hyperedge prototypes based on the global context, calculating the correlation between pixel nodes and hyperedge prototypes to generate a second hyperedge matrix for hypergraph convolution, it is possible to capture higher-order spatial relationships between pixel nodes and output spatial relationship-enhanced detail features. The neck structure combines the channel-relationship-enhanced fusion features with the spatial relationship-enhanced detail features to generate multi-scale features, realizing joint inference of the channel dimension and spatial dimension, and improving the model's multi-scale expression ability for rice diseases in complex backgrounds.

[0012] Furthermore, the neck structure generates multi-scale features based on the channel relationship enhancement fusion features and the spatial relationship enhancement detail features, including: concatenating the channel relationship enhancement fusion features and the spatial relationship enhancement detail features, performing convolution operations on the concatenated features to aggregate channel information and spatial information, and outputting multi-scale features.

[0013] The beneficial effects of adopting the above-mentioned further scheme are as follows: By concatenating the channel-relationship-enhanced fusion features and the spatial-relationship-enhanced detail features along the channel dimension, and then performing a convolution operation after concatenation, it is possible to simultaneously aggregate semantic relationship information in the channel dimension and detailed positional information in the spatial dimension. This process enables the two types of enhanced features to achieve information complementarity in the same feature map, avoiding the problem of incomplete representation by single-dimensional features. The convolution operation further performs nonlinear recombination on the concatenated features, compressing redundant channels and enhancing effective responses. The final output multi-scale features, while maintaining spatial resolution, incorporate high-order semantics and fine textures, improving the model's consistency and discriminative ability in feature representation of disease targets at different scales.

[0014] Furthermore, the spatial relationship hypergraph inference module generates a hyperedge prototype based on the global context of the pixel node sequence, including: calculating the maximum and average values ​​of the features of each pixel node in the pixel node sequence, concatenating the maximum and average values ​​to obtain global context features, generating an offset from the global context features through a mapping layer, and adding the offset to a preset hyperedge initial prototype vector to obtain the hyperedge prototype.

[0015] The beneficial effects of adopting the above-mentioned further scheme are as follows: By calculating the maximum and average values ​​of the features of each pixel node in the pixel node sequence and concatenating the two, the resulting global context features simultaneously retain the peak response and overall distribution information of the channel dimension, which can more comprehensively represent the global statistical characteristics of the disease scene. The global context features are used to generate offsets through a mapping layer and added to a preset initial hyperedge prototype vector to obtain the hyperedge prototype. This allows the hyperedge prototype to dynamically and adaptively adjust according to the global context of the input image, avoiding the problem of insufficient generalization ability of a fixed prototype for different farmland scenes. The dynamically generated hyperedge prototype provides a more accurate semantic template for subsequent correlation calculations between pixel nodes and hyperedges, thereby enhancing the adaptability of the spatial relationship hypergraph inference module to complex backgrounds, lighting changes, and differences in disease morphology.

[0016] 2) In a second aspect, the present invention also provides a rice disease identification system based on sparse anthropomorphic visual perception and hypergraph reasoning, the specific technical solution of which is as follows: The system comprises a model building module, a model training module, and a model application module. The model building module constructs a rice disease identification model, which includes a backbone network, a neck structure, and a detection head. The backbone network receives rice sample images; its shallow layers perform convolution operations on these images to output shallow detail features. A biomimetic visual perception module is located in the deeper layers of the backbone network. This module receives and performs biomimetic visual perception on the shallow detail features, generating biomimetic visual perception features. The neck structure receives these features and performs hypergraph reasoning and spatial hypergraph reasoning to generate multi-scale features. The detection head performs classification and position regression prediction on these multi-scale features to obtain the disease identification results for the rice sample images. The model training module trains the rice disease identification model. The model application module inputs a target rice image into the trained rice disease identification model and outputs the disease identification results for the target rice image.

[0017] 3) In a third aspect, the present invention also provides an electronic device, the electronic device including a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor, so that the electronic device implements any of the above-mentioned rice disease identification methods based on sparse anthropomorphic visual perception and hypergraph reasoning.

[0018] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements any of the above-mentioned rice disease identification methods based on sparse anthropomorphic visual perception and hypergraph reasoning.

[0019] It should be noted that the beneficial effects of the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementations can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating a rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning according to an embodiment of the present invention. Figure 2 A schematic diagram illustrating data collection and model deployment; Figure 3 This is a schematic diagram of the overall structure of the model; Figure 4 This is a schematic diagram of the backbone network of the model; Figure 5 This is a schematic diagram of the feature extraction process for a sparse biomimetic visual perception module. Figure 6 A schematic diagram of the surrounding visual mechanism; Figure 7 A schematic diagram of the neck structure; Figure 8 A schematic diagram of the channel relationship hypergraph inference module structure based on causal correlation modeling; Figure 9 This is a schematic diagram of the spatial relationship reasoning module structure based on fine-grained hypergraph convolution; Figure 10 This is a schematic diagram of the training curves for model loss and evaluation metrics; Figure 11 This is a schematic diagram of the model's prediction results; Figure 12 This is a schematic diagram of the model's evaluation curve; Figure 13 This is a schematic diagram of the confusion matrix of the model; Figure 14 A schematic diagram visualizing the window-level correlation matrix and Top-k spatial routing relationship of the biomimetic visual perception module; Figure 15 A schematic diagram visualizing the surrounding visual parameters of the biomimetic visual perception module; Figure 16 A schematic diagram visualizing the training results of the peripheral visual parameters of the bionic visual perception module; Figure 17 A schematic diagram visualizing the global distance-weighted response of a biomimetic visual perception module; Figure 18 A schematic diagram of the feature map extracted by the bionic visual perception module; Figure 19 This is a schematic diagram of the channel correlation matrix; Figure 20 This is a schematic diagram of the hyperedge matrix; Figure 21 A schematic diagram visualizing the feature map fusion process; Figure 22 This is a schematic diagram of the correlation matrix between pixels and hyperedges; Figure 23 This is a schematic diagram for robustness verification; Figure 24 This is a schematic diagram of the structure of a rice disease identification system based on sparse anthropomorphic visual perception and hypergraph reasoning according to an embodiment of the present invention. Detailed Implementation

[0021] like Figure 1 As shown in the figure, a rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning according to an embodiment of the present invention includes the following steps: S1. Construct a rice disease identification model, which includes a backbone network, neck structure, and detection head. Specifically: (1) The backbone network is used to receive rice sample images. The shallow layers of the backbone network perform convolution operations on the rice sample images and output shallow detail features.

[0022] The backbone network consists of four stacked processing stages. The first and second stages are defined as shallow layers, and the third and fourth stages as deep layers. At the input of the backbone network, a raw image of a rice sample is received. This image is in three-channel RGB format, and its initial height is denoted as [insert height here]. Width is denoted as The number of channels is 3. ① In the first stage, a CBS module is set at the starting position. This CBS module uses a convolution operation with a stride of 2 to downsample the input image, so that the spatial scale of the output feature map becomes that of the original image. That is, the height is Width is Simultaneously, the number of channels is increased from 3 to 64. Subsequently, the downsampled feature map is fed into the C2f module of this stage. The C2f module performs nonlinear transformations and feature extraction on the feature map through internal convolution operations and residual connections, outputting the first shallow detail features of this stage, whose spatial scale is maintained at [value missing]. The number of channels is 64. ② In the second stage, a CBS module is also set at the starting position. This CBS module receives the first shallow detail features output from the first stage, and performs downsampling using a convolution operation with a stride of 2, reducing the spatial scale of the feature map to... The number of channels was increased from 64 to 128. Subsequently, the downsampled feature map was fed into the second-stage C2f module for feature extraction, outputting the second shallow detail feature of this stage, with a spatial scale of [missing information]. The number of channels is 128. The first and second shallow detail features together constitute the output of the shallow layer of the backbone network, which is then fed to the subsequent biomimetic visual perception module and neck structure. ③ In the third stage, a CBS module is set at the starting position. This CBS module receives the second shallow detail features and performs downsampling using a convolution operation with a stride of 2, reducing the spatial scale of the feature map to [missing information]. The number of channels was increased from 128 to 256. Subsequently, the downsampled feature maps were fed into the third-stage biomimetic visual perception module. This module simulates human visual mechanisms, performing biomimetic visual perception processing on the input shallow detail features, extracting multi-scale representations that combine coarse and fine-grained information, and outputting the third-stage biomimetic visual perception features with a spatial scale of [missing information]. The number of channels is 256. ④ In the fourth stage, a CBS module is set at the starting position. This CBS module receives the biomimetic visual perception features output from the third stage, and performs downsampling using a convolution operation with a stride of 2, reducing the spatial scale of the feature map to [a smaller value]. The number of channels was increased from 256 to 512. Subsequently, the downsampled feature map was fed into the fourth-stage bionic visual perception module for further in-depth anthropomorphic visual feature interaction, outputting the fourth-stage bionic visual perception features with a spatial scale of [missing information]. The number of channels is 512. The biomimetic visual perception features output from the third and fourth stages are used together as the deep output of the backbone network and are fed into the neck structure for subsequent feature fusion and inference.

[0023] The process of outputting shallow detail features is as follows: After the rice disease identification model is started, the input of the backbone network receives an original rice sample image. This image is a three-channel RGB format digital image, denoted as the feature map, and its initial height is denoted as [missing information]. Width is denoted as The number of channels is 3. The rice sample image first enters the first stage of the backbone network. In this stage, the image is first fed into a CBS module. This CBS module processes the input image using a convolution operation with a stride of 2 to achieve downsampling. After this downsampling, the spatial scale of the feature map becomes that of the original image. That is, the height is Width is Simultaneously, the number of channels was increased from 3 to 64. Subsequently, the scale was... The feature map with 64 channels is fed into the C2f module of this stage. The C2f module performs further nonlinear transformations and feature extraction on the feature map through internal convolution operations and residual connections, outputting the shallow detail features of this stage, denoted as the first shallow detail feature, whose spatial scale is maintained at [value missing]. The number of channels is 64. The first shallow detail features are then fed into the second stage of the backbone network. In this stage, they also first pass through a CBS module, which uses a convolution operation with a stride of 2 to downsample, reducing the spatial scale of the feature map to [value missing]. And the number of channels was increased to 128. Afterwards, the scale was... The feature map with 128 channels is fed into the second stage C2f module for feature extraction, outputting the shallow detail features of this stage, denoted as the second shallow detail features, with a spatial scale of [missing information]. The number of channels is 128. At this point, the feature extraction process for the shallow layers (first and second stages) of the backbone network is complete. The generated first and second shallow detail features are output as shallow detail features for subsequent biomimetic visual perception modules and neck structures. The first and second shallow detail features together serve as the shallow output of the backbone network. Rice sample images are the raw input data used to train the rice disease recognition model. These images were collected through fixed-point monitoring equipment (such as high-resolution cameras) in farmland, covering different growth stages, different disease types (such as rice blast and wilt), and different field background environments (such as close-up single leaves and complex grassy backgrounds). The images are typically in three-channel RGB format.

[0024] (2) A biomimetic visual perception module is set up in the deep layer of the backbone network. The biomimetic visual perception module is used to receive and perform biomimetic visual perception on shallow detail features to generate biomimetic visual perception features. Specifically, the biomimetic visual perception module is used to: receive shallow detail features, divide the shallow detail features into multiple window regions, select multiple key regions with the highest relevance for each query region by calculating the correlation between window regions, and establish a region routing path; extract query features and key features according to the region routing path, calculate the Euclidean distance between the query position coordinates and the key position coordinates and generate a distance matrix, and perform weighted fusion of the convolution mapping result of the distance matrix with the attention weights to output biomimetic visual perception features. The biomimetic visual perception module consists of two parts connected in series: the first part is sparse inference attention, and the second part is peripheral visual encoding. Sparse inference attention is used to calculate the sparse correlation between regions and extract the initial attention output; peripheral visual encoding is used to modulate the attention weights according to the spatial distance to simulate the peripheral attenuation characteristics of human vision. The input of the entire module is the shallow detail features output by the shallow layer of the backbone network, and the output is the biomimetic visual perception features. The specific implementation process is as follows: 1) The biomimetic visual perception module receives shallow detail features from the second stage of the backbone network. The spatial scale of this feature map is... The module has 128 channels. It spatially divides the feature map into... There are four window regions of equal size, each with a spatial scale of [missing information]. Let the total number of window areas be . Each window region is considered a potential target region for subsequent regional correlation calculations.

[0025] 2) For each window region, the bionic visual perception module first calculates the query features, key features, and value features in the attention using a 1×1 convolution, denoted as... , and The calculation method is as follows: ,in For the input feature map, the first Feature vectors at each position , , It is a trainable projection parameter matrix. Then, for each window region... and Perform a mean operation to obtain the semantic vector for each window region: ,in, , This is a function to calculate the mean. Next, the correlation matrix between regions is calculated. Its elements Indicates the first The region and the first Correlation index of each region: ,in, and This is a linear projection matrix, used to enhance nonlinear expressive power. Matrix The shape is For each query region From its corresponding number Select the largest value from the row These elements, and the corresponding region indexes, constitute the routing index set for the query region. , shape is The set of route indexes for all query regions together constitutes the region route path.

[0026] 3) Based on the set of route indexes For each query region, from the original full graph and The features corresponding to the routing regions are collected to obtain the sparsified query features. Bond features Their shapes are all ,in The dimension of each feature vector. The specific operation is as follows: ,in, This is a function that extracts features based on an index.

[0027] 4) Calculate the Euclidean distance between the query location coordinates and the key location coordinates and generate a distance matrix. For each query region and its corresponding... Given a routing key region, retrieve the spatial coordinates of the center point of each region. Let the coordinates of the center point of the queried region be denoted as . The coordinates of the center point of the key region are Calculate the Euclidean distance between the two to obtain the distance matrix. It is represented as: ,in, Representative This method measures distance and initializes it with negative Gaussian decay. It is a bias term.

[0028] 5) The convolutional mapping result of the distance matrix is ​​weighted and fused with the attention weights to output the biomimetic visual perception features. The distance matrix... The surrounding visual encoding is obtained by mapping through convolution, followed by ReLU activation and Sigmoid normalization. The calculation method is as follows: ,in, It is a non-linear activation function; It is the Sigmoid normalized activation function; and It is a convolutional layer; It is a mapped peripheral visual encoding. Next, the output of the sparse inference attention is calculated to obtain the final biomimetic visual perception features. : ,in, It is a learnable scalar parameter used to control the strength of the surrounding visual encoding's control over the interaction; The nonlinear decay characteristics of the lateral inhibition effect in human vision were simulated. Here is the activation function used for normalization; This is the feature dimension. The output is the bionic visual perception feature generated by the bionic visual perception module, which is then fed into the subsequent neck structure for further processing.

[0029] The query feature is a feature representation corresponding to the query region extracted from the original feature map based on the region routing path. In the attention calculation, the query feature is used to match with key features to determine the association strength between different regions. Key features are feature representations corresponding to key regions related to the query region extracted from the original feature map based on the region routing path. The dot product of the key features and query features yields the similarity between regions. The distance matrix is ​​a two-dimensional matrix whose elements represent the Euclidean distance between the coordinates of the query region and the coordinates of each key region. This matrix quantifies the spatial proximity between different regions and is the basis for generating peripheral visual encoding. The bionic visual perception feature is a feature map output by the bionic visual perception module after sparse attention calculation and peripheral visual encoding modulation. This feature integrates coarse and fine-grained visual information and reflects the perceptual characteristics of "central fixation and peripheral attenuation" in human vision, providing a more physiologically consistent feature expression for subsequent neck structure analysis.

[0030] (3) The neck structure is used to receive bionic visual perception features and perform relational hypergraph reasoning and spatial relational hypergraph reasoning to generate multi-scale features. Specifically, the neck structure includes a channel relational hypergraph reasoning module and a spatial relational hypergraph reasoning module. The channel relational hypergraph reasoning module performs multi-directional spatial rearrangement of the deep semantic features in the bionic visual perception features and the shallow texture features in the shallow detail features, calculates the similarity matrix, generates the first hyperedge matrix based on the similarity matrix, performs hypergraph convolution, and outputs the channel relation-enhanced fusion features. The spatial relational hypergraph reasoning module flattens the shallow detail features into a sequence of pixel nodes, generates a hyperedge prototype based on the global context of the pixel node sequence, calculates the correlation between each pixel node and the hyperedge prototype to generate the second hyperedge matrix, performs hypergraph convolution, and outputs the spatial relation-enhanced detail features. The channel relationship hypergraph inference module calculates a similarity matrix after multi-directional spatial rearrangement of deep semantic features in biomimetic visual perception features and shallow texture features in shallow detail features. Based on the similarity matrix, it generates the first hyperedge matrix and performs hypergraph convolution to output the channel relationship-enhanced fusion feature. The specific implementation process is as follows: 1) The channel relation hypergraph inference module receives two types of input features: one is deep semantic features from the deep layers of the backbone network's biomimetic visual perception features, and the other is shallow texture features from the shallow layers of the backbone network's shallow detail features. To unify the spatial scale, the module first upsamples the deep semantic features to make their spatial scale consistent with the shallow texture features. Let the upsampled deep semantic features be denoted as... Shallow texture features are Subsequently, a convolutional layer was used to... and Encode the intermediate features separately to obtain intermediate features with the same scale and number of channels. Then, use a 3×3 convolution with a stride of 2 to compress the number of channels and spatial scale of the intermediate features to obtain high-energy features. and ,in Corresponding to deep semantic features, This corresponds to shallow texture features.

[0031] 2) High-energy characteristics Perform a block splitting operation to divide it into There are 3 spatial blocks, each with a spatial scale of 1. ,in and for The height and width. The block operation is represented as: ,in, `<function>` is a tensor transformation function used to rearrange the feature maps into a block sequence. Subsequently, the block-based feature maps are sorted according to five spatial arrangements: left-to-right, top-to-bottom, right-to-left, bottom-to-top, and random. Each sorting operation is represented as: ,in, For the first sorting functions This is the sorted sequence of feature blocks. Perform the same block division and sorting operations to obtain the corresponding feature block sequence. .

[0032] 3) For each spatial arrangement The sorted feature block sequence The query features in the attention layer are obtained through linear mapping using a fully connected layer. The sorted feature block sequence The key features in the attention layer are obtained by mapping using the same fully connected layer. .calculate and The inner product of gives the first Similarity matrix under spatial arrangement : , where the matrix The shape is , This represents the total number of blocks.

[0033] 4) The five similarity matrices obtained under the five spatial arrangements The attention weights for each arrangement are obtained by projecting and pooling the data. : ,in It is a linear projection function. This is the pooling function. It is a scalar, representing the first... The importance of spatial arrangement. The similarity matrices of the five arrangements are weighted and fused according to attention weights to generate the final first hyperedge matrix. : First hyperedge matrix The shape is Each row represents a hyperedge, each column represents a channel node, and the values ​​of the matrix elements represent the correlation strength between the channel node and the hyperedge.

[0034] 5) Based on the first hyperedge matrix Then, perform a hypergraph convolution operation. First, process the input channel features... After flattening and concatenating, the node feature matrix is ​​obtained, denoted as... Its shape is ,in For the number of nodes, The node feature dimension is used. The first stage of hypergraph convolution is feature aggregation from nodes to hyperedges: each hyperedge collects the features of all nodes it connects through a mapping matrix. Perform a linear transformation to obtain the hyperedge features. : ,in, For superedge index, For node indexing, Indicates the relationship with the hyperedge A set of connected nodes. The first hyperedge matrix is ​​the first... Line 1 The element values ​​of the column, Let be the mapping matrix from nodes to hyperedges. The activation function is used. The second stage of hypergraph convolution is feature backpropagation from the hyperedge to the node: transferring the hyperedge features... Through another mapping matrix After projection, the updated node features are backpropagated to each node according to the association strength. : ,in, Represents nodes A set of connected hyperedges This is the mapping matrix from hyperedges to nodes. The updated node features... The original spatial shape is restored, resulting in a fusion feature enhanced with channel relationships. This feature retains information from both the original deep semantic features and shallow texture features, while also incorporating higher-order relationships between channels. As the output of the channel relationship hypergraph inference module, it is fed into the subsequent neck structure for further feature aggregation.

[0035] Among them, deep semantic features are the high-order semantic information contained in the biomimetic visual perception features output by the deep layers (third and fourth stages) of the backbone network. These features have a large number of channels and a small spatial resolution, encoding abstract semantic information such as the category and overall morphology of the disease target. Shallow texture features are the detailed texture information contained in the shallow detail features output by the shallow layers (first and second stages) of the backbone network. These features have a small number of channels and a large spatial resolution, encoding local detail information such as the edges, textures, and color variations of the disease area. Multi-directional spatial rearrangement refers to the operation of dividing the feature map into multiple blocks in a spatial dimension and rearranging these blocks in various different orders. Specifically, it includes five spatial arrangement methods: left to right, top to bottom, right to left, bottom to top, and random sorting. This operation is used to break the fixed dependence of features between channels on spatial position, thereby more comprehensively capturing the non-linear relationships across channels. The similarity matrix is ​​a two-dimensional matrix whose elements represent the inner product of the query feature and the key feature after mapping between different feature blocks. This matrix quantifies the correlation strength between features in different spatial regions and forms the basis for generating the hyperedge matrix. The first hyperedge matrix is ​​generated by weighting and fusing the similarity matrix with attention weights, and it describes the association between hyperedges and nodes in the hypergraph structure. Rows correspond to hyperedge indices, columns to channel node indices, and the values ​​of matrix elements represent the correlation strength between the corresponding channel node and the hyperedge. Hypergraph convolution is a convolution operation extended from graph neural networks, used to process hypergraph structure data. In hypergraph convolution, each hyperedge can connect multiple nodes, and higher-order relationships are modeled through feature aggregation from nodes to hyperedges and feature backpropagation from hyperedges to nodes. Hypergraph convolution can capture complex dependency patterns between channels that go beyond pairwise relationships. The channel relationship-enhanced fusion feature is a feature map output by the channel relationship hypergraph inference module after a series of operations, including multi-directional spatial rearrangement, similarity calculation, hyperedge matrix generation, and hypergraph convolution. This feature incorporates higher-order relationship information between channels on top of the original semantic and texture features, enhancing the semantic expressiveness and information transmission efficiency of cross-layer features.

[0036] The spatial relationship hypergraph inference module generates hyperedge prototypes based on the global context of the pixel node sequence. This includes: calculating the maximum and average values ​​of the features of each pixel node in the pixel node sequence; concatenating the maximum and average values ​​to obtain global context features; generating offsets from the global context features through a mapping layer; and adding the offsets to a preset initial hyperedge prototype vector to obtain the hyperedge prototype. The specific implementation process is as follows: 1) The spatial relation hypergraph inference module receives shallow detail features from the shallow output of the backbone network. Let the input feature map be... Its spatial scale is The number of channels is Feature map Flattening it in the spatial dimension yields a length of... pixel node sequence Each pixel node ( ) is a dimension The eigenvectors of . The flattening operation is represented as: ,in, This is a function that converts a two-dimensional spatial grid into a one-dimensional sequence.

[0037] 2) For pixel node sequences Along the node dimension (i.e. For each feature channel of all pixels, the calculation is performed separately. The maximum value of the feature value for each pixel node in each channel is then calculated to obtain a maximum value vector. , dimension , its first The elements are: Simultaneously, the average value of all pixel node features in each channel is calculated to obtain an average value vector. , dimension , its first The elements are: , the maximum value vector with average vector By concatenating along the channel dimension, global context features are obtained. Its dimensions are : ,in, This is a vector concatenation function.

[0038] 3) Global context features The input is fed into a mapping layer. This mapping layer consists of one or more fully connected layers, coupled with a non-linear activation function. The output dimension of the mapping layer is the same as the dimension of the initial prototype vector of the hyperedge, denoted as . The calculation process of the mapping layer is represented as follows: ,in, This indicates a multi-layer fully connected mapping operation. The generated offset vector has dimensions of Offset This reflects the need for correction of the preset hyperedge prototype based on the global context information of the current input image.

[0039] 4) The preset initial prototype vector of the hyperedge is denoted as... Its dimensions are . It is a learnable parameter that is continuously updated during model training, representing a basis of potential visual correlations between disease categories. The offset... with the initial prototype vector of the hyperedge Adding each element together yields the final dynamic hyperedge prototype. : Hyper-edge prototype The dimension is In actual implementation, the spatial relationship hypergraph reasoning module is configured with... There are 10 super edges, therefore there exist hyperedge prototype vectors For each superedge Each step, from the second to the fourth, is repeated using the same global context feature. The corresponding offset is generated through the mapping layer. Then, with the initial prototype vector of the corresponding hyperedge. Adding them together yields the dynamic prototype of the hyperedge. These hyperedge prototypes This will be used in subsequent steps to calculate the association strength between each pixel node and each hyperedge, thereby constructing the second hyperedge matrix and performing hypergraph convolution.

[0040] The pixel node sequence is a one-dimensional sequence obtained by flattening the input feature map in spatial dimensions. The original shape of the feature map is... After being flattened, it was obtained There are 10 pixel nodes, each node corresponding to a feature vector at a given location, with dimensions 1. A pixel node sequence is a fundamental element in a hypergraph structure, with each node representing a feature at a spatial location. A pixel node feature is the feature vector carried by each node in the pixel node sequence, with dimensions of... This feature vector encodes visual information about the corresponding spatial location, such as texture, edges, or color. In the spatial relation hypergraph inference module, the features of each pixel node are used to calculate its correlation with the hyperedge prototype, thereby determining which hyperedges the node belongs to. Global context features are global statistical information extracted from the entire sequence of pixel nodes, used to guide the dynamic adjustment of the hyperedge prototype. Specifically, the maximum and average values ​​of the feature dimensions of each pixel node are calculated, and then the maximum value vector and the average value vector are concatenated to obtain a dimension... The feature vector represents the global distribution characteristics of the entire feature map. The mapping layer is a trainable fully connected neural network layer used to transform the input global context features into offsets. The mapping layer typically consists of linear transformations and non-linear activation functions, capable of learning complex mappings from the global context to the hyperedge prototype adjustment. The hyperedge initial prototype vector is a pre-defined, learnable parameter vector, denoted as . This vector is randomly initialized at the start of model training and represents the initial template for potential visual correlations between diseases. The hyperedge initial prototype vector is independent of the input image and is a globally shared basis vector. The hyperedge prototype is the final vector obtained by adding the hyperedge initial prototype vector to the offset, denoted as . The hyperedge prototypes are dynamically generated and can adaptively adjust according to the content of the input image. They are used to subsequently calculate the association strength between each pixel node and each hyperedge. Each hyperedge corresponds to a hyperedge prototype vector.

[0041] The spatial relationship hypergraph inference module calculates the correlation between each pixel node and the hyperedge prototype, generates a second hyperedge matrix, performs hypergraph convolution, and outputs detailed features enhanced with spatial relationships. The specific implementation process is as follows: 1) The spatial relation hypergraph inference module receives shallow detail features from the shallow layers of the backbone network and flattens them into a sequence of pixel nodes. Let the pixel node feature matrix be... , shape is ,in This represents the total number of pixel nodes. The feature dimensions for each pixel node. Meanwhile, the module has already generated [data / data] through previous steps. hyperedge prototype vectors Each hyperedge prototype The dimension is In practical implementation, it is usually... Set to with The same, or the pixel node features are mapped to the same space as the hyperedge prototype through linear projection.

[0042] 2) For the first A hyperedge, its hyperedge prototype With each pixel node feature ( Perform inner product calculation to obtain the original correlation score. To enhance expressive power, pixel node features are first... A linear transformation is performed through a mapping layer to obtain projected features of the same dimension as the hyperedge prototype. ,Right now ,in Let be the trainable parameter matrix. Then calculate the inner product: ,in This represents the vector dot product operation. For all... super edge and Repeat the above calculation for each pixel node to obtain the original correlation score matrix. , shape is To ensure comparability of the association strength between each pixel node and each hyperedge, the score of each pixel node is Softmax normalized along the hyperedge dimension to obtain the second hyperedge matrix. : ,in Indicates the first The pixel node belongs to the first The probability of a superedge satisfies .matrix The shape is That is, the second hyperedge matrix.

[0043] 3) Hypergraph convolution consists of two stages: feature aggregation from nodes to hyperedges, and feature backpropagation from hyperedges to nodes.

[0044] ① First stage: Feature aggregation from nodes to hyperedges. For each hyperedge... Collect features of all pixel nodes According to the correlation strength in the second hyperedge matrix Perform weighted aggregation to obtain hyperedge features. During aggregation, the features of each pixel node are first processed through a mapping matrix. Perform a linear transformation, then sum the results using weighted averages: ,in. Let be the mapping matrix from nodes to hyperedges, with dimension . , Dimension of the hyperedge feature; It is a non-linear activation function (such as ReLU).

[0045] ②In the second stage, feature backpropagation from hyperedges to nodes. This involves backpropagating the features of each hyperedge. Through another mapping matrix Perform a linear transformation, then determine the correlation strength according to the second hyperedge matrix. The weighted data is then fed back to each pixel node to obtain the updated pixel node features. : ,in. Let be the mapping matrix from hyperedges to nodes, with dimension . ; This is a non-linear activation function. The updated features of all pixel nodes together constitute a new node feature matrix. , shape is .

[0046] 4) Update the pixel node feature matrix obtained after convolving the hypergraph. To restore to the original spatial shape, that is, from Shape reshaping The feature map is the detailed feature for spatial relation enhancement. This feature is fed into the subsequent processing steps of the neck structure, where it is concatenated and convolved with the fusion feature for channel relation enhancement to aggregate channel and spatial information, ultimately generating multi-scale features for head classification and regression.

[0047] In this context, a pixel node is the basic unit of each spatial location obtained by flattening the input feature map in the spatial relation hypergraph inference module. The spatial scale of the input feature map is... After being flattened, it was obtained There are pixel nodes, each pixel node corresponds to a feature vector at a location, and the dimension is . Pixel nodes are vertices in a hypergraph structure. The second hyperedge matrix is ​​a two-dimensional matrix, denoted as... Its shape is ,in The number of superedges. Number of pixel nodes. Matrix elements. Indicates the first superedge and the first The correlation strength between pixel nodes. The second hyperedge matrix is ​​obtained by calculating the inner product of each pixel node feature and the hyperedge prototype and then normalizing it with Softmax. It is the basis for subsequent hypergraph convolution.

[0048] Specifically, the neck structure generates multi-scale features based on the fusion features enhanced by channel relationships and the detailed features enhanced by spatial relationships. Specifically, the fusion features enhanced by channel relationships and the detailed features enhanced by spatial relationships are concatenated, and the concatenated features are subjected to convolution operations to aggregate channel and spatial information, outputting multi-scale features. The specific implementation process is as follows: 1) In the processing flow of the neck structure, the channel relationship hypergraph inference module outputs channel relationship-enhanced fusion features, denoted as... The spatial scale of this feature is The number of channels is Simultaneously, the spatial relation hypergraph reasoning module outputs enhanced spatial relation details, denoted as... The spatial scale of this feature is The number of channels is The two feature maps have the same spatial scale. This is a prerequisite for subsequent splicing operations.

[0049] 2) Along the channel dimension, and Perform a stitching operation. The stitched result is a new feature map. Its spatial scale remains unchanged. The number of channels remains unchanged. The specific form of the splicing operation is as follows: ,in, This represents the feature map concatenation function along the channel dimension. After concatenation, information from the channel relation enhancement fusion features and the spatial relation enhancement detail features is merged into the same feature map, providing a complete input for subsequent convolutional aggregation.

[0050] 3) The stitched feature map The input is fed into a convolutional layer. This convolutional layer uses multiple convolutional kernels, each typically with a spatial size of 3×3, a stride of 1, and padding of 1 to maintain the output spatial scale. The number of output channels of the convolutional layer is set to an intermediate value as needed, denoted as . This value is usually less than This achieves channel dimension compression and feature recombination. The computational process of convolution operation is represented as follows: ,in, This indicates a convolution operation using a 3×3 convolution kernel. This is the feature map output after the convolution operation. Through the convolution operation, the model can learn the non-linear combination relationship between channels in the concatenated feature map, and at the same time, it can use the spatial receptive field of the convolution kernel to fuse spatial information at adjacent positions, thereby achieving effective aggregation of channel information and spatial information.

[0051] 4) The feature map obtained after the convolution operation This serves as the output of the fused features at the current scale. In the neck structure, the aforementioned concatenation and convolution operations are repeated at multiple different scales and levels. For example, for the backbone network output... , , Feature maps of different scales are processed with corresponding channel relationship enhancement and spatial relationship enhancement techniques, and then the above-mentioned splicing and convolution operations are performed to obtain a set of multi-scale feature maps. These multi-scale feature maps are fed into the detection head and used for classification prediction and location regression prediction of rice disease targets of different sizes.

[0052] The channel relationship-enhanced fusion feature is the feature map output by the channel relationship hypergraph inference module after multi-directional spatial rearrangement, similarity matrix calculation, first hyperedge matrix generation, and hypergraph convolution. This feature incorporates high-order relationship information between channels on top of the original semantic and texture features, enhancing the semantic expressiveness and information transmission efficiency of cross-layer features. The spatial relationship-enhanced detail feature is the feature map output by the spatial relationship hypergraph inference module after calculating the correlation between pixel nodes and hyperedge prototypes, generating the second hyperedge matrix, and hypergraph convolution. This feature incorporates high-order spatial relationship information between pixel nodes on top of the original shallow detail features, enhancing the model's ability to perceive local structures, lesion edges, and texture-similar regions.

[0053] (4) The detection head performs classification prediction and location regression prediction on multi-scale features to obtain the disease identification results of rice sample images. The specific implementation process is as follows: 1) The detection head adopts a decoupled structure, setting independent classification prediction branches and location regression prediction branches for each scale of multi-scale features output by the neck network. Each branch consists of several stacked convolutional layers. Specifically, for an input feature map of one scale... Its spatial scale is The number of channels is The classification branch first compresses the number of channels to an intermediate number using a 1×1 convolutional layer, and then outputs a classification prediction map using a 3×3 convolutional layer. The number of output channels equals the total number of disease categories. (i.e., 10). The location regression branch also first compresses the number of channels through a 1×1 convolutional layer, and then outputs the regression prediction map through a 3×3 convolutional layer. The number of output channels is equal to the number of bounding box coordinate parameters (i.e., 4, corresponding to the horizontal coordinate offset, vertical coordinate offset, width offset, and height offset of the center point, respectively). The detection head repeats the above structure for the feature map at each scale, and the parameters are not shared between different scales.

[0054] 2) The detection head receives a multi-scale feature set from the output of the neck network, denoted as... These correspond to feature maps at large, medium, and small scales, respectively. The spatial scale of each feature map decreases sequentially, while the number of channels increases sequentially. These multi-scale features have had their channel and spatial information enhanced by the aforementioned channel relationship hypergraph inference module and spatial relationship hypergraph inference module.

[0055] 3) For the feature map at each scale This data is then fed into the classification branch of the detection head. The classification branch extracts classification features through convolution operations, ultimately outputting a classification prediction map. Its spatial scale and Same (i.e.) The number of channels equals the total number of categories. For each spatial location in the classification prediction map Its corresponding channel vector This represents the predicted probability of each category at that location. A Sigmoid activation function is used to map the value of each channel to between 0 and 1, representing the confidence level that the corresponding disease category exists at that location. Each location in the classification prediction map corresponds to a specific region in the original image (determined by the feature map downsampling ratio).

[0056] 4) For the feature map at each scale This data is then fed into the location regression branch of the detection head. The location regression branch extracts regression features through convolution operations, ultimately outputting a regression prediction map. Its spatial scale and Same (i.e.) The number of channels is 4. For each spatial location in the regression prediction plot... The four corresponding values ​​are denoted as follows: This represents the offset relative to the preset anchor point. and This represents the offset of the center point of the prediction box relative to the top-left corner of the grid cell. and This represents the scaling factor for the width and height of the prediction box relative to the width and height of the anchor point.

[0057] 5) For each scale of feature map, generate a classification prediction map. Regression Prediction Plot Pairing is performed at the same spatial location. First, based on regression predictions... Using preset anchor point parameters, the actual coordinates of the predicted bounding box in the original image are calculated. For each location, the coordinates of the top-left corner of the grid cell are denoted as... The preset anchor point width and height are The coordinates of the center point of the decoded prediction box and width and height for: Then, combining the confidence scores from the classification prediction map, predictions with confidence scores higher than a preset threshold (e.g., 0.5) are retained. Each retained prediction includes a disease category label (the category with the highest confidence), a confidence score, and bounding box coordinates. The above decoding operation is repeated for feature maps at all scales, and all prediction results are aggregated. Finally, a non-maximum suppression algorithm is used to remove overlapping redundant prediction boxes, yielding the final disease identification result, i.e., the category and location information of each detected rice disease target.

[0058] In this invention, classification prediction is the output of the classification branch in the detection head, representing the probability distribution of each target within a bounding box belonging to various disease categories. Classification prediction typically converts the network output into probability values ​​using a Softmax or Sigmoid function; the category with the highest probability is the predicted disease type. The disease categories in this invention include 10 categories: rice blast, wilt, false smoke disease, sheath disease, leaf spot, brown spot, dead heart, dew disease, rice Dongger virus disease, and normal rice. Position regression prediction is the output of the regression branch in the detection head, representing the offset of each bounding box relative to a preset anchor point or grid position. Regression prediction typically includes four values, corresponding to the horizontal coordinate offset of the bounding box's center point, the vertical coordinate offset of the center point, the width scaling factor, and the height scaling factor. These offsets are converted into the actual coordinates of the final predicted bounding box in the original image through decoding. Disease identification results are the final detection information output by the detection head after processing the input rice sample image, containing the category label and bounding box position for each detected disease target. Disease identification results are typically output in list format, with each element containing the category name, confidence score, and the coordinates of the top-left and bottom-right corners of the bounding box. These results are used to guide practical farmland disease monitoring and control.

[0059] S2. The rice disease identification model is trained, and the specific implementation process is as follows: S20. Collect 5000 images of rice diseases, each image containing one or more disease targets, with precise category labels and bounding box annotations. Randomly divide all images into a training set and a validation set in a 4:1 ratio, with the training set containing 4000 images and the validation set containing 1000 images. The training set is used for learning and updating model parameters, while the validation set is used to evaluate the model's generalization performance during training and does not participate in parameter updates.

[0060] S21. Before starting training, determine the following hyperparameter values: Learning rate is set to 0.01 to control the step size of each parameter update. Momentum is set to 0.937 to accelerate the gradient descent process and reduce oscillations. Weight decay is set to 0.0005 as a regularization term to prevent overfitting. Batch size is set to 16, meaning 16 images are processed and the average gradient is calculated simultaneously in each iteration. Training epochs are set to 300, meaning the training set is fully traversed 300 times. Environment configuration: Python interpreter version 3.7, PyTorch version 1.10, CUDA version 11.3, RTX 8000 GPU, 32GB RAM.

[0061] S22. The loss function of the rice disease identification model consists of two parts: classification loss and regression loss. The classification loss uses binary cross-entropy loss or zoom loss to measure the difference between the class probability output by the detector's classification branch and the true class. The regression loss uses full intersection-over-union loss or smoothed L1 loss to measure the difference between the bounding box coordinates output by the detector's regression branch and the true bounding box coordinates. Total loss... Represented as classification loss Regression loss Weighted sum: ,in, and These are the balancing coefficients, typically set to 1 and 1 respectively. The specific calculation methods for classification loss and regression loss follow the default settings in the YOLOv8 baseline model.

[0062] S23. Input a batch of images (batch size 16) from the training set into the rice disease identification model. The images pass through the backbone network, neck structure, and detection head in sequence, and finally output the classification prediction map and regression prediction map on each scale feature map. The result of forward propagation is the model's prediction output for the current batch of images.

[0063] S24. Compare the model's predicted output with the corresponding ground truth annotations (including class labels and bounding box coordinates) of the image, and calculate the total loss value for the current batch based on the constructed loss function. Then, use the backpropagation algorithm to calculate the gradient of the loss function with respect to each trainable parameter in the model. Specifically, starting from the output layer, calculate the gradient layer by layer according to the chain rule until the input layer.

[0064] S25. Based on the calculated gradient, update all parameters of the model using the stochastic gradient descent optimizer. The parameter update formula is: ,in, The parameters at the current moment, The learning rate is 0.01. The gradient of the loss function with respect to the parameters. The weight decay coefficient is 0.0005. Simultaneously, a momentum mechanism is used to smooth the update direction, specifically implemented using the momentum setting (0.937) in the standard stochastic gradient descent optimizer.

[0065] S26. Repeat steps four through six, with each complete training set (4000 images) processed as one training round. After each training round, evaluate the current model using the validation set, calculating metrics such as mean precision, accuracy, and recall on the validation set to monitor the model's generalization performance. During training, the learning rate is gradually reduced according to a preset cosine annealing or step decay strategy to facilitate fine-tuning of the model later. After 300 training rounds, stop training and save the current model parameters as the final rice disease identification model. During training, the loss function curve and evaluation metric curve should show a stable convergence trend, indicating that the model has learned effective disease feature representations.

[0066] S3. Input the target rice image into the trained rice disease identification model, and output the disease identification result of the target rice image. The specific implementation process is as follows: In actual farmland monitoring scenarios, fixed-point monitoring equipment (such as eagle-eye cameras) deployed in the fields collect images of designated rice-growing areas. The cameras acquire real-time images of the rice plants at preset time intervals or in a continuous acquisition mode. The acquired images are in three-channel RGB format, denoted as [image format notation missing]. Its height is Width is This image is the target rice image. The acquired target rice image is adjusted to the size required for the model input. The trained rice disease identification model requires the spatial scale of the input image to be a fixed value (e.g., ...). (pixels). Therefore, the original image Adjusted by scaling The image is processed by normalizing pixel values ​​to between 0 and 1. The preprocessed image is denoted as... This preprocessing step does not change the number of channels in the image (it remains at 3 channels). The preprocessed target rice image... The input is fed into the trained rice disease identification model. The model performs the following operations sequentially according to the forward propagation sequence: the backbone network receives the input image, the shallow layer outputs shallow detail features, and the deep layer's biomimetic visual perception module receives the shallow detail features and generates biomimetic visual perception features; the neck structure receives the multi-scale features output by the backbone network, the channel relation hypergraph inference module enhances the channel relations of the deep semantic features and shallow texture features, and the spatial relation hypergraph inference module enhances the spatial relations of the shallow detail features. Then, the two types of enhanced features are concatenated and convolved to output multi-scale features; the detection head receives the multi-scale features, the classification branch outputs a classification prediction map, and the regression branch outputs a position regression prediction map. The classification prediction map and regression prediction map output by the detection head need to be decoded to obtain the final disease identification result. For each spatial location on each scale feature map, the offset obtained from the regression prediction map is calculated. Using preset anchor point parameters, the actual coordinates of the predicted bounding boxes in the input image are calculated. Simultaneously, the confidence score for each category is obtained from the classification prediction map, and predictions with a confidence score higher than a preset threshold (e.g., 0.5) are retained. Each retained prediction includes a disease category label, a confidence score, and bounding box coordinates. The predictions decoded from all scale feature maps are aggregated, and non-maximum suppression (NMS) is used to remove overlapping redundant prediction boxes. The NMS threshold is typically set to 0.45; that is, when the intersection-union ratio (IU) of two prediction boxes is greater than 0.45, the box with higher confidence is retained, and the box with lower confidence is suppressed. After NMS processing, the remaining prediction boxes are the final disease identification results. The disease identification results are output in list form, with each element containing the category name (e.g., "rice blast"), a confidence score (e.g., 0.92), and the coordinates of the top-left and bottom-right corners of the bounding box. This output can be used to draw detection boxes and label disease categories on the original image, providing intuitive disease monitoring information for farmland managers.

[0067] The target rice image is an image of the rice plant to be detected, acquired through fixed-point monitoring equipment (such as a high-resolution camera) in an actual farmland monitoring scenario. This image is a three-channel RGB digital image and may contain one or more rice disease targets. The image background is complex and diverse, including close-ups of single leaves, complex grassy backgrounds, backgrounds with strong leaf vein textures, and natural field environment backgrounds. The target rice image serves as input data for the model inference stage.

[0068] This invention proposes a rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning. Specifically, it proposes a biomimetic visual perception module to construct regional relationships and mimic human visual gaze perception of targets, thereby adaptively extracting coarse-to-fine granularity features at a deep core of the backbone network. A channel relationship hypergraph reasoning module is designed, which traverses the spatial domain through multi-directional spatial rearrangement, breaking the insensitivity of feature relationships between channels, and using hypergraph convolution to reason about channel relationship patterns, achieving efficient transmission of semantic information. A spatial relationship hypergraph reasoning module is introduced to reason about spatial relationship patterns of detailed information at the spatial level, guiding the organic integration of detailed and semantic information, and achieving efficient transmission of detailed information. In practical deployment, such as... Figure 2As shown, rice sample images are typically collected through fixed-point monitoring. Specifically, several fixed monitoring points are set up in representative areas of the farmland, each equipped with a high-resolution camera (such as an eagle-eye camera) to continuously observe the designated area. By rationally planning the spatial distribution and field of view coverage of the cameras, effective monitoring of different growth areas can be achieved, thereby obtaining representative and diverse image data, which is then input into the detection model to output disease monitoring results. This method can reduce the cost of manual inspections while ensuring the stability and continuity of data collection, providing reliable data support for the subsequent training and evaluation of rice disease identification models.

[0069] The structure of the rice disease identification model is as follows: Figure 3 As shown, the overall system adopts a single-stage detection framework. Figure 3 This is a schematic diagram of the overall structure of a rice disease identification model based on sparse anthropomorphic visual perception and hypergraph reasoning. The model consists of a backbone network (i.e.,...) Figure 3 The "backbone" and neck structure (i.e.) Figure 3 The system consists of three parts: the "neck" of the rice sample, the detection head, and the input rice sample image. The image first enters the backbone network. In the shallow layers of the backbone, the image undergoes convolution operations sequentially, with output scales of [missing information]. , The feature map, where and Here, represents the height and width of the input image, and 64 and 128 represent the number of channels. BM denotes batch normalization. Subsequently, in the deeper layers of the backbone network, the feature maps continue to undergo convolutional downsampling to obtain a scale of... and The feature maps are then input into biomimetic visual perception (i.e., Figure 3 The "Bionic Visual Perception" module is part of the larger module. The bionic visual perception module is designed for scale... and One feature is set up at each location to simulate the human visual mechanism and extract a mixture of coarse and fine-grained features. The feature extraction process also includes pyramid pooling to enhance multi-scale contextual information.

[0070] The multi-scale features output from the backbone network are fed into the neck structure. The neck contains two fusion paths: a semantic information fusion path and a spatial information fusion path. In the semantic information fusion path, deep features... First, the feature passes through the CBS (convolution + batch normalization + SiLU activation function) module to obtain CBS256×H / 32×W / 32. Then, channel relation inference is performed with the upsampled features (implemented through the channel relation hypergraph inference module), outputting channel relation inference: 256×H / 16×W / 16. This is then concatenated and fused with 512×H / 16×W / 16 to obtain 512×H / 16×W / 16. Finally, C2f outputs 512×H / 16×W / 16. Subsequently, this feature is downsampled by CBS to obtain 128×H / 16×W / 16, and then channel relation inference is performed with shallow features, outputting channel relation inference 128×H / 8×W / 8. This is concatenated to obtain Concat256×H / 8×W / 8, and then passed through CBS and C2f to obtain 128×H / 8×W / 8, completing the transfer of semantic information from deep to shallow. In the spatial information fusion path, the deep feature 512×H / 32×W / 32 first undergoes semantic relation reasoning (implemented through the spatial relation hypergraph reasoning module), outputting the semantic relation reasoning: 512×H / 32×W / 32 (i.e. Figure 3 The feature is first described as "semantic relation inference 512xH / 32xW / 32", then processed by CBS to obtain 256×H / 32×W / 32, and then output by C2f to be 256×H / 16×W / 16. This feature is then subjected to semantic relation inference again, resulting in semantic relation inference: 256×H / 16×W / 16, which is then processed by CBS to obtain 128×H / 16×W / 16, completing the transfer of spatial detail information from shallow to deep. Finally, the neck structure outputs the multi-scale features enhanced by the two paths to three detection heads. Each detection head performs classification prediction and position regression prediction on the feature map of the corresponding scale, and outputs the disease identification results of the rice sample image. This invention uses YOLOv8 as the benchmark model, considering its wide application in multiple fields and good engineering performance. This model has the advantages of structural stability, strong generalization ability and high robustness. Meanwhile, YOLOv8 integrates efficient data augmentation strategies and loss optimization mechanisms, and has a relatively reasonable architecture in terms of backbone network, neck structure, and detection head design, providing a good foundation for this study. The overall detection process is as follows: ① Collected rice sample images are input into a backbone network for feature extraction. In the shallow layers of the backbone network, a BiFormer is introduced to obtain sparse global information; in the deep layers, a biomimetic visual perception module is designed to simulate human visual mechanisms and extract biomimetic visual perception features that combine coarse and fine granular characteristics. Subsequently, the multi-scale features output from the backbone network are input into the neck structure for feature fusion. This fusion process is divided into two paths, respectively realizing the collaborative modeling of semantic information and detailed texture information. ② In the semantic information transmission path, a channel relationship hypergraph inference module is introduced. A cross-scanning mechanism is used to rearrange spatial regions in multiple directions, calculate the relationship strength between deep semantic features and shallow texture features, and construct a hypergraph structure for cross-layer features accordingly. Furthermore, hypergraph convolution is used to model the relationships between channels, thereby achieving effective fusion of semantic and texture features guided by channel relationships. ③ In the detailed information transmission path, shallow detailed features are first downsampled to obtain coarse-grained information that provides guidance. Then, a hypergraph structure is constructed based on regional correlations in the spatial dimension, and spatial relationships are modeled through hypergraph convolution, thereby guiding the transmission and fusion of detailed information between different feature layers. The spatial relationship hypergraph inference module is used to perform the aforementioned spatial relationship modeling. Finally, through joint inference of channel relationships and spatial relationships, a multi-scale feature representation incorporating high-order relationship information is obtained and input into the separate detection head to achieve target classification prediction and location regression prediction.

[0071] like Figure 4 As shown, the biomimetic visual perception feature extraction method based on sparse inference is used in the backbone network of the model. It is mainly composed of a C2f module (the baseline residual convolutional block in the YOLOv8 model) and a biomimetic visual perception module (BVP), and is divided into four stages. The first two stages are regarded as the shallow layers of the backbone network, rich in detailed information. The C2f module, which is dominated by convolution, is used for feature extraction to obtain rich fine-grained features. The latter two stages are rich in semantic information, so the biomimetic visual perception module is used for deep anthropomorphic visual feature interaction. Each stage downsamples the image through the CBS module (convolution + batch normalization + SiLU activation function). The scale ratio coefficients of the four stages are [1 / 4, 1 / 8, 1 / 16, 1 / 32]. At the same time, the three-channel RGB image is converted into a feature map, and its channel number is expanded to [64, 128, 256, 512] to increase the information capacity. Specifically: The input image first undergoes a convolution operation, which corresponds to the CBS module (convolution + batch normalization + SiLU activation function) in this invention, reducing the spatial scale of the input image to its original size. The feature map scale is obtained as follows: ,in and The input image contains the height and width. Then, it enters stage 1, which consists of N stacked C2f modules, outputting a feature map with a scale of H / 4×W / 4×64. Next, it undergoes a convolution operation, downsampling again to reduce the spatial scale. The number of channels is increased to 128, and then it enters stage 2, which also consists of N stacked C2f modules, with an output feature map scale of H / 8×W / 8×128. Afterwards, it undergoes convolutional downsampling to obtain a scale of... The feature map is processed and enters stage 3. In this stage, C2f is no longer used; instead, N BVP modules (Bionic Visual Perception Modules) are stacked, and the output feature map scale is H / 16×W / 16×256. Finally, after convolutional downsampling, a scale of H / 16×W / 16×256 is obtained. The feature map is then processed in stage 4, which also consists of N stacked BVP modules, and the output feature map has a scale of H / 32×W / 32×512. Stages 1 and 2 constitute the shallow layer of the backbone network, using the C2f module to extract detailed features; stages 3 and 4 constitute the deep layer of the backbone network, using the biomimetic visual perception module to extract mixed coarse and fine granular features. Figure 4 The “×N” in the text indicates that N identical modules are repeatedly stacked in each stage (the specific value of N is set according to the model configuration, for example, the value of N is different in different stages in YOLOv8).

[0072] Given that convolutional neural networks have relatively limited ability to model global information, and that conventional Vision Transformers tend to generate significant computational redundancy when acquiring global information and may introduce some invalid features, this invention improves upon the BiFormer structure and integrates a peripheral visual attention mechanism to design a sparse biomimetic visual perception module. For example... Figure 4 As shown, structurally, the biomimetic visual perception module is divided into two parts: sparse inference attention and surrounding visual encoding.

[0073] like Figure 5 As shown in green, sparse inference attention uses a region-related attention mechanism to adaptively learn the relationships between different regions. This operation is similar to graph inference, effectively capturing inter-region dependencies in images and improving the model's performance on complex tasks. Specifically: Sparse inference attention is equivalent to the restricted region sparse inference attention operation; the module divides the feature map into... Each window represents a potential target region. Identification is performed using a coarse-grained method. Count the regions containing the target, and select the most relevant Top- Targets for each region. Each region, construct A directed graph is used to represent the relationships between regions. The most relevant regions are called routes, which are key to capturing spatial semantic associations. The set of routes is used to remove irrelevant regions. First, the query in the attention is computed through a 1×1 convolution. ),key( ) and value ( That is, it can be represented as: ,in, For the input feature map, the first Feature vectors at each position , , It is a trainable projection parameter matrix. Secondly, for In each region and Perform a mean operation to obtain the semantic vector for each region, which is represented as: ,in, It is a function for finding the mean, for the th All within a window area value (or The mean of the values ​​is calculated to obtain the semantic vector of the region. (or Next, in order to enhance the nonlinear expression, [we]... and Performing linear projection followed by inner product gives the similarity matrix a stronger ability to express relationships, which can be represented as: ,in, and The linear projection matrix, the matrix Representing the relationship between regions, the shape is... ,matrix The Middle The line represents the first The region and the rest The relationship index of each region. Based on the relationship matrix, only the most relevant top-... Connect the intervals and obtain the relevant index. Its shape is ,in Indicates the first A query range. Based on ,use The function can extract data with regional correlation. and , represented as: .for and Performing attention operations, i.e., calculating sparse attention, is represented as: ,in, It is the activation function used for normalization. For feature dimensions.

[0074] like Figure 5 As shown, the input feature map is first divided into multiple interval representation window regions and then enters the sparse inference attention branch. In this branch, Query (Q) and Key (K) are obtained through convolutional mapping, and then the most relevant region is selected for each query region by constructing region-related routes. Subsequently, the distance between Q and K of the most relevant regions is calculated, and this distance is fed into the peripheral weight encoding module after Gaussian adjustment, outputting the peripheral view encoding P. At the same time, Value (V) is also obtained through convolutional mapping. Inside the sparse inference attention, Query and Key are inner product and weighted, and the peripheral view encoding is fused as a bias term. Finally, it is multiplied by Value through a multiplication operation (Mul) to obtain the biomimetic visual perception feature. The entire process simulates human visual perception, where the gaze center region corresponds to the query region, the weight of the peripheral regions decays with distance, and the final output is a feature representation with biomimetic characteristics.

[0075] For sparse inference attention, where and The interaction between objects relies entirely on the regional relationship and feature similarity (inner product) between them, which contradicts the human visual pattern. Existing literature shows that referencing the human gaze mechanism can effectively improve the model's feature extraction capabilities. Specifically, humans typically exhibit a unique visual perception pattern when gazing at a target: the central gaze region receives the highest attention, while the perceptual weight of the surrounding regions gradually decreases. This mechanism has significant advantages in feature extraction; it not only strengthens the representation of local details (i.e., the gaze center) but also preserves, to some extent, the contextual information of global semantics (i.e., the surrounding regions), thereby achieving effective synergy between local and global information. In object detection tasks, since the semantic association between targets usually weakens with increasing spatial distance, attention mechanisms that decay linearly with distance are preferred to better align with spatial relationship modeling in real-world scenarios. Inspired by this human visual mechanism, such as... Figure 6 As shown, this invention performs peripheral visual encoding based on sparse inference attention, aiming to simulate the attention distribution pattern in human vision and improve the model's perception ability and feature expression efficiency in multi-scale object detection.

[0076] The Peripheral Vision Transformer (PViT) provides valuable insights into this improvement. For example... Figure 6 As shown, in the original visual mechanism, it simulates the surrounding visual process by constructing attention weights that decay linearly with spatial distance. Each query vector... The key vectors are considered as the gaze center. This represents the surrounding visual area. Therefore, and The interaction strength between components is determined not only by inner product similarity but also modulated by spatial distance decay. Furthermore, considering that the human receptive field is not a completely fixed and regular geometric shape, PViT introduces convolutional projection to learn richer spatial visual patterns, thereby obtaining more expressive peripheral attention weights. However, the weight mapping module in PViT relies on preset peripheral visual initialization parameters, and its weight generation process is not directly driven by input data. Therefore, this module can only fit relatively static positional relationships through parameter updates, making it difficult to obtain sufficient gradient feedback based on input features. Previous research in this invention found that this module is difficult to optimize during training, especially in scenarios with deeper network layers and more complex tasks, such as object detection. The weight mapping module often struggles to learn effectively, thus limiting further improvements in model performance. To alleviate these problems, this invention employs an additional independent classifier for pre-training of this module. That is, usable parameters are obtained first after removing the backbone network, and then loaded into the object detection model and frozen. Although this approach can bring certain performance gains, it still has obvious limitations: on the one hand, the parameters obtained by pre-training are highly dependent on specific training processes and data distributions, and have insufficient generalization ability; on the other hand, frozen loading also limits the module's adaptive optimization ability in downstream detection tasks, thus making it difficult to fully utilize its potential representation ability.

[0077] Figure 6 This demonstrates how the human peripheral visual gaze mechanism inspires the design of attention weights. In the static, dense peripheral visual mechanism (before improvement), query features... Bond features The interaction between them is determined solely by feature similarity, and the attention weights use a preset static distance template when simulating peripheral vision. In the dynamic sparse peripheral vision mechanism (improved), peripheral vision constraints are applied only to the region most relevant to the current gaze center, and the distance matrix is ​​dynamically calculated by generating a dynamic Euclidean distance between Q and K. ,in, and These are the center point coordinates of the query range and the key range, respectively. and These are learnable parameters. Subsequently, a convolutional mapping is used to obtain the surrounding visual encoding. This is used to modulate attention weights, resulting in the strongest response when focusing on the central region, while the weights of the peripheral regions vary with Euclidean distance (i.e., ...). Figure 5 The text "Euclidean distance" in the image increases and then decreases, thus more realistically simulating the peripheral attenuation characteristics of human vision.

[0078] To address the aforementioned limitations, such as Figure 6As shown, this invention improves the visual weight generation method in PViT by proposing a dynamically sparse peripheral visual attention mechanism. This method applies peripheral visual constraints only to the region most relevant to the current gaze center. Due to the query vector... With key vector The regional correlation between them can adaptively change with the input content, therefore for each gaze center Only its relationship with the most relevant region needs to be calculated. Distance matrix between This allows the construction of sparse distance constraints dynamically driven by the input, specifically represented as: ,in and yes and The coordinates of the positions can be used to calculate the Euclidean distance between the two. Representative This method measures distance and initializes it with negative Gaussian decay. It is a bias term. For the distance matrix... By performing a mapping through convolution, we obtain the attention weights based on relational content awareness, which are expressed as: ,in It is a non-linear activation function; It is the Sigmoid normalized activation function; and It is a convolutional layer; It is a mapped peripheral visual encoding. Finally, the output of the bionic visual perception module is represented as: ,in, Used to control peripheral visual encoding pairs The intensity of interactive control is similar to the lateral inhibition effect in human vision; simultaneously, the modulation effect of the peripheral region on the central region exhibits a non-linear decay. The function is used to better fit this characteristic. The smaller the value, the more negative it is, thus reducing the corresponding attention weight.

[0079] Multi-scale features extracted by the backbone network need to be transmitted through the neck structure for spatial and semantic information. For example... Figure 7As shown, the method of this invention is based on the PANet architecture. The neck structure contains two paths: a semantic path that transmits high-order features from deep layers to shallow layers, and a spatial path that transmits texture details from shallow layers to deep layers. These two paths work together to effectively promote the organic fusion of multi-scale features. However, the above methods lack consideration for channel and spatial relationships, leading to the introduction of invalid information and the neglect of hidden relationship patterns during the fusion process. Therefore, this invention proposes a feature fusion method based on spatial semantic relationship hypergraph inference to construct the neck structure of the model. The core modules are the channel relationship hypergraph inference module (CI) and the spatial relationship hypergraph inference module (PI). Both operate on neighborhood features in different paths, considering efficient feature fusion paths from both channel and spatial dimensions.

[0080] Figure 7 The backbone network, whose output multi-scale backbone network features are denoted as P3, P4, and P5, correspond to shallow detail features (high resolution) and deep semantic features (low resolution), respectively. These features are fed into the neck structure (i.e., Figure 7 Feature fusion is performed on the "neck network" in the model. The neck structure contains two paths: one for semantic information transmission and the other for spatial information transmission. In the semantic information transmission path, deep features are enhanced with shallow features through the CI module to improve channel relationships; in the spatial information transmission path, shallow features are modeled with pixel-level spatial relationships through the PI module. Figure 7 In this context, CI and PI represent the specific application locations of these two types of hypergraph inference modules, respectively. After multiple rounds of feature interaction and fusion, the neck structure outputs enhanced multi-scale feature maps, which are then fed into the corresponding detection heads for classification prediction and location regression prediction.

[0081] In semantic fusion, the goal is to efficiently convey high-order semantic information. Since different channels play different roles in feature extraction, channel features often exhibit significant differences. Traditional methods typically fuse features by directly concatenating and compressing them, but this is insufficient for effectively modeling the differences and relationships between channels. Therefore, this invention proposes a channel relationship hypergraph inference module to enhance information interaction between channels.

[0082] Modeling channel relationships is not easy. Inner product operations in attention mechanisms are better suited for depicting spatial relationships, but in the channel dimension, features need to be flattened first, which destroys the original spatial information. Furthermore, using inner products on redundant channels may amplify noise correlation, leading to information redundancy. To address this issue, this invention suppresses noise interference by modeling causal relationships between channels and combines this with hypergraph inference to achieve more accurate information delivery. Compared to traditional graph convolution, hypergraph convolution can depict higher-order channel relationships, reducing computational redundancy while avoiding the omission of key channel information, thereby improving feature representation capabilities and optimizing semantic fusion effects.

[0083] Specifically, such as Figure 8 As shown, this invention first unifies neighborhood features to the same scale through upsampling. Then, convolutional layers are used to encode the features, obtaining features of the same scale and number of channels. and Then, a 3×3 convolution with a stride of 2 is used to compress the number of channels and the size to obtain high-energy features. and .for It is then divided into blocks. Each scale is The feature block is represented as: ,in, The function is a Reshape operation used for block division. Subsequently, the divided feature maps are sorted in five spatial arrangements: left-to-right, top-to-bottom, right-to-left, bottom-to-top, and random, as follows: ,in, It is a function used for sorting; Representing the The five spatial arrangement feature blocks are sorted; for each sorted block, a fully connected layer is used to map them, treating them as query features in the attention mechanism. .for Perform the same partitioning and mapping operations, treating it as a key feature. , represented as: The two perform inner product calculations to obtain similarity matrices for five spatial arrangements; for each of the five spatial arrangements, projection and pooling operations are performed to obtain attention weights, which are then weighted and fused to obtain the hyperedge matrix, represented as: ,in, It is a pooling function; These are attention weights, which are used to fuse the similarity matrix through weighted averaging, and are expressed as follows: Based on hyperedge matrix This allows the execution of hypergraph convolution operations. Each hyperedge collects features from all vertices and forms hyperedge features through linear projection. These hyperedge features are then inversely restored to the nodes, represented as: ,in and These are the node and the hyperedge index, respectively. Represents the hyperedge feature. Represents the channel features after hypergraph inference; and These are the mapping matrices from nodes to hyperedges and from hyperedges to nodes, respectively. Represents the activation function; These are the channel features of the input, derived from... and It is obtained by flattening and piecing together.

[0084] Figure 8 In this module, the input includes small-scale and large-scale feature maps, corresponding to deep semantic features and shallow texture features, respectively. The input is first encoded by a convolutional layer, and then upsampled to enlarge the small-scale feature map to the same spatial size as the large-scale feature map, resulting in a cross-scale channel feature map. Subsequently, the feature map enters the channel relationship pattern construction process: first, it undergoes a convolutional downsampling projection layer to compress the number of channels and spatial size; then, a block-based operation divides the feature map into multiple blocks, and spatial rearrangement is performed using five sorting paths under a multi-phase path scanning mechanism. The rearranged feature blocks are mapped to Q and K through projection layers respectively to obtain query features and key features, and then cosine similarity is calculated to obtain a similarity matrix. The similarity matrix is ​​then aggregated through the relevance matrix path and input into a multilayer perceptron to generate attention weights, which are then weighted and fused to obtain a hyperedge matrix. Simultaneously, the original features are organized into a node set. The hyperedge matrix and the node set enter the hypergraph convolution module together. After two mappings from node to hyperedge and from hyperedge to node, the output channel relationship aggregation feature is added to the original input through residual connection to enhance the stability of feature propagation.

[0085] The explanation of the spatial relation hypergraph reasoning module is as follows: For each vector in the features, this invention dynamically models the correlation to generate a second hyperedge matrix. For example... Figure 9 As shown, firstly, the flattened feature map, i.e. ,in , and These represent the height and width of the feature map, respectively. This represents the number of channels. Then, the maximum and average values ​​of the flattened pixel node features are calculated and concatenated to obtain the global context features. Next, the offset is obtained from the global context features through a mapping layer. This offset is related to the initial prototype vector of the hyperedge. Adding them together yields the hyperedge primitive, which is equivalent to the potential visual correlation between diseases, specifically represented as: ,in. This indicates the global context features. The offset matrix obtained after linear mapping has dimensions of ; Represents the number of superedges; The initial prototype vector of the hyperedge is preset, and its dimension is also . To calculate the contribution value of each pixel node, the pixel node features are... Perform linear mapping and connect with the hyperedge prototype Perform the inner product operation to obtain the final second hyperedge matrix. : ,in, This indicates the pixel node features The feature matrix obtained after linear projection has dimension . ; For hyperedge prototype matrix The transpose of, with dimension ; The inner product result has a dimension of ; The scaling factor (usually taken as the feature dimension) ); This is a temperature coefficient used to control the smoothness of the Softmax distribution; As a normalized activation function, the score of each pixel node is normalized along the hyperedge dimension to obtain the second hyperedge matrix. Its shape is Subsequently, using the second hyperedge matrix... Based on this, hypergraph convolution is performed, specifically as shown in the formula. As shown (i.e., feature aggregation from node to hyperedge and feature backpropagation from hyperedge to node). This process generates semantically related guiding features, which are then concatenated with small-scale features and further aggregated through convolution operations to finally obtain fused features of adjacent scales.

[0086] Figure 9 In the hypergraph, the large-scale input features are first downsampled via convolution to obtain node set features, which then enter the hyperedge matrix construction branch. In this branch, the input features are first globally represented using global average pooling and global max pooling to obtain global context features, which are then projected to generate dynamic offsets. These offsets are added to the initial hyperedge prototype vector in the trainable parameters to obtain the hyperedge prototype. Subsequently, the node set features and the hyperedge prototype are multiplied by a matrix, the inner product is calculated, and then normalized using Softmax to obtain the node-hyperedge matrix. This matrix describes the association strength between each pixel node and each hyperedge; the figure shows M hyperedges, from hyperedge 1 to hyperedge M. Simultaneously, the small-scale features are also downsampled via convolution and then enter the hypergraph convolution module along with the node set features. The hypergraph convolution process consists of two stages: node-to-hyperedge and hyperedge-to-node. The first stage aggregates node features to hyperedges to obtain hyperedge features; the second stage feeds the hyperedge features back to the nodes to update their features. The updated node features are concatenated and aggregated with the original features to obtain relation pattern-guided features. After further convolution processing, the final output is the relation-guided fused features, which are detailed features that enhance spatial relationships.

[0087] The technical effects of this invention are illustrated through the following experiments. The rice sample images used in this experiment comprise 5000 images covering 10 categories, including rice blast, wilt, false smoke, sheath disease, leaf spot, brown spot, dead heart, dew disease, rice Dongger virus disease, and normal rice. Although this dataset contains a large number of categories and can comprehensively reflect the diversity of rice disease types, the small overall sample size and limited number of samples per class increase the difficulty for the model to perform effective feature learning and accurate classification. The rice disease dataset was split into training and validation sets in a 4:1 ratio. The hyperparameters were set as follows: learning rate 0.01, momentum 0.937, weight decay 0.0005, batch size 16, and 300 training epochs. The environment configuration was as follows: Python interpreter version 3.7, PyTorch version 1.10, CUDA version 11.3, RTX 8000 GPU, and 32GB of RAM. In object detection tasks, there are two fundamental metrics: precision (P) and recall (R). Precision represents the proportion of true positive samples out of all predicted positive samples; recall represents the proportion of true positive samples correctly detected by the model. Since these two metrics are inversely proportional, the mean precision (mAP) is further derived. mAP is a comprehensive metric used to measure the detection effectiveness across all categories. Generally, the closer mAP is to 1, the better the model's performance. Therefore, this invention uses P, R, and mAP as a composite metric to comprehensively and objectively evaluate the model's detection capabilities.

[0088] 1) To verify the stability and convergence of the model training process, this invention visualized the changes in the loss function and detection performance metrics during the training and validation phases. The results are as follows: Figure 10As shown in the visualization, the model exhibits a stable convergence trend over 300 training epochs. The losses on the training set continuously decrease with each iteration, indicating that the model's ability to predict target location regression, classification, and bounding box distribution gradually improves. The validation set loss curve generally follows the same trend as the training loss, without significant rebounds or sharp oscillations, indicating good generalization ability and no significant overfitting during training. Meanwhile, the evaluation metrics on the validation set rapidly improve in the early stages of training and then gradually stabilize, indicating that the model has learned the main target features after a few iterations and further optimizes detection accuracy and localization quality in subsequent training. In particular, the continuous improvement of the mAP50-95 metric shows that the model maintains good detection robustness at different IoU thresholds. The learning rate curve decays smoothly according to the preset strategy, which helps the model to fine-tune parameters in the later stages of training, thereby improving the final convergence stability. In summary, the training process shows sufficient convergence, stable performance improvement, and good consistency between the validation metrics and the loss curve, proving that the constructed model has good target detection performance and reliable generalization ability.

[0089] Meanwhile, the model's prediction results are as follows Figure 11 As shown in the figure, the prediction results demonstrate that the model can locate and identify leaf disease targets to a certain extent under different imaging backgrounds. The sample backgrounds include close-up single leaves, complex grassy backgrounds, backgrounds with strong leaf vein textures, and natural field environments, indicating that the test images exhibit strong scene diversity. The model demonstrates relatively stable performance under complex backgrounds, proving its generalization and robustness in implementation. From the perspective of target scale, Figure 11 The disease scale varies considerably. The model is relatively stable in detecting medium-scale and well-defined disease areas, and also performs well on small-scale lesions with blurred edges or those highly similar to leaf texture, indicating that the model has a good response capability to disease characteristics at various scales. Furthermore, different diseases differ in color, morphology, and spatial distribution; some diseases manifest as linear, stripe-like, or irregular patches, and multiple similar disease areas may appear in the same image. The model is also able to provide effective predictions in most samples.

[0090] To further verify the detection performance of the model, such as Figure 12As shown, the performance evaluation curves of the model in this invention are visualized. The results show that the model has good overall detection performance. The Precision-Confidence curve generally increases with increasing confidence, indicating that increasing the confidence threshold can effectively reduce false detections; the Recall-Confidence curve gradually decreases with increasing confidence, reflecting that a high threshold will lead to some missed detections. The F1-Confidence curve reaches a high level in the medium confidence range, indicating that this range can well balance precision and recall. The Precision-Recall curve generally remains in the high region, but there are certain differences between different categories, indicating that the detection stability of some categories is good, while a few categories are still difficult to identify due to target scale, background interference, or category similarity. In summary, the model's detection results are relatively stable, and in practical applications, a medium confidence threshold should be selected to balance false detections and missed detections.

[0091] Combination Figure 13 The confusion matrix shown further reveals that the model exhibits relatively stable recognition capabilities across some categories, indicating that the network can extract effective disease area features. However, some samples still show missed detections and false detections. This is mainly because the sample size of the data in this invention is relatively limited, and the number of disease categories is large, leading to insufficient feature learning and fitting capabilities for some categories. Future research will focus on optimizing features under small sample conditions, class balancing, and improving the model's generalization ability.

[0092] 2) To verify the independence and synergy of the proposed modules, an ablation experiment was conducted to verify the contribution of the proposed method to performance improvement. The results are shown in Table 1. The experimental results show that the P, R, and mAP50 of the baseline model are 0.768, 0.630, and 0.723, respectively, indicating that the original model has a certain detection capability, but still suffers from missed detections in complex field scenarios. Specifically: ① After introducing the biomimetic visual perception module (BVP module), although the model accuracy decreased slightly, the recall and mean precision were improved, indicating that the biomimetic visual perception mechanism can enhance the model's ability to represent the features of diseased areas, helping to capture weak features and fine-grained lesion information. ② After introducing the spatial relation hypergraph reasoning module (PI module), the overall model performance was further improved, indicating that spatial relation reasoning can effectively promote the transmission of detailed information and enhance the model's perception ability of small-scale targets and diseased areas with similar textures. ③ After introducing the channel relation hypergraph reasoning module (CI module), the model showed a similar trend to the biomimetic visual perception module, namely, improved recall and mean precision while slightly decreased accuracy. This demonstrates that channel relation hypergraph reasoning enhances semantic information interaction and multi-scale feature fusion while improving the model's response capability to potential disease areas. Considering that disease monitoring is usually based on continuous frame images, practical applications focus more on the full detection of disease targets; therefore, improving recall is crucial for reducing the risk of missed detections. ④ When the three modules—bionic visual perception module, spatial relation hypergraph reasoning module, and channel relation hypergraph reasoning module—are introduced simultaneously, the model achieves optimal detection results, with precision, recall, and mean precision reaching 0.785, 0.697, and 0.769, respectively. This shows that the three modules complement each other from the perspectives of bionic visual feature extraction, spatial relation modeling, and channel semantic reasoning, effectively improving the model's ability to detect rice disease targets in complex backgrounds, verifying the effectiveness of the method and the rationality of the module design.

[0093] Table 1: 3) In order to intuitively visualize the working mechanism and performance improvement reasons of each module, this invention visualizes the processing flow and extracted features of each module.

[0094] ① Visualization Experiment of Bionic Visual Perception Module: To analyze the working mechanism of the sparse routing mechanism in the bionic visual perception module, this invention... Figure 14 (a) visualizes the window-level correlation matrix and its Top-k spatial index relationship for different input images, as shown in the figure. Figure 14As shown in (b), different images exhibit significant differences in the window-level correlation matrix, indicating that the relationships between diseased image regions are not simple local neighborhood relationships, but rather complex cross-regional feature associations. Therefore, the model can adaptively select the Top-k candidate regions most relevant to the current query region based on the response intensity between regions, thereby forming sparse regional interaction paths. Figure 14 (c) Visualization further illustrates the specific workings of this process. Centered on the selected query window Q, the model connects to multiple K regions distributed in different spatial locations based on feature relevance. This demonstrates that the biomimetic visual perception module can adaptively establish cross-regional associations within the global candidate regions. Therefore, the sparse routing mechanism can transform complex spatial relationships into learnable adjacency indices, providing a more effective basis for subsequent peripheral visual constraints and pixel-level attention calculations, thereby enhancing the model's ability to express fine-grained disease features in complex contexts.

[0095] After determining the Top-k candidate regions through sparse spatial routing, the biomimetic vision perception module further initializes their surrounding vision to construct biomimetic spatial constraints. Figure 15 It is evident that the weight responses in the untrained state are mainly concentrated in the Top-k routing region, rather than the entire feature map, indicating that this mechanism can limit irrelevant regions from participating in attention calculations during the initialization phase. Simultaneously, the weights within the candidate region exhibit a continuous and gradual distribution, reflecting a certain distance decay prior, giving the model a perceptual tendency similar to human vision's "center-focused, periphery-faded" tendency. These results demonstrate that the biomimetic visual perception module can construct an effective biomimetic visual spatial prior based on sparsely correlated regions, providing structural support for subsequent fine-grained disease feature modeling.

[0096] The results after training are as follows Figure 16 As shown, the peripheral visual weights of the bionic vision perception module exhibit a more clearly defined structured distribution within the Top-k candidate regions. Compared to the initialization results, the weights no longer simply exhibit distance decay but instead form differentiated responses based on feature content and spatial relationships, indicating that the model can adaptively adjust attention constraints within sparse routing regions. High-response regions are mainly concentrated in local areas strongly correlated with the query location, suggesting that the peripheral visual encoding has learned a spatial perception pattern that better matches the distribution of disease features during training. This result demonstrates that the bionic vision perception module not only provides bionic spatial priors during the initialization phase but also further optimizes inter-regional interactions during training, thereby enhancing the model's ability to express fine-grained disease features. Therefore, the aforementioned method of constructing dynamic distance inputs for peripheral visual learning effectively solves the problem of difficult visual parameter training in our previous work.

[0097] To further analyze the global spatial perception characteristics of the trained biomimetic visual perception module, this invention visualizes the distance-weight response relationship in the peripheral visual encoding, as shown in the following figure. Figure 16 As shown. Figure 17 (a) shows the 3D distance-weight curves of different attention heads. It can be seen that each attention head has a different response to spatial distance, focusing on local detail perception or mid-to-long-distance context modeling, respectively, which reflects the complementarity of multi-head mechanism in spatial relationship modeling. Figure 17 (b) shows the 3D distance-weight curves under different query windows. The weight distributions differ for different windows, indicating that the peripheral visual encoding is not a fixed distance decay template, but can dynamically adjust spatial constraints according to the query area and its Top-k routing objects. Figure 17 (c) shows the distance-weighted response heatmap for the attention head dimension, further validating that different attention heads exhibit differentiated response patterns across different distance ranges. These results demonstrate that the biomimetic visual perception module can learn diverse spatial perception patterns based on sparse routing, thus balancing the representation of local lesion details with the modeling of global region relationships.

[0098] Finally, this invention provides a visual analysis of the feature maps extracted by the biomimetic visual perception module. For example... Figure 18 As shown, after processing by the biomimetic visual perception module, the model exhibits a relatively clear response in the diseased area and its adjacent texture locations, while also suppressing interference from irrelevant background areas to a certain extent. Strong activation is observed in the feature maps for leaf edges, lesion textures, and local abnormal areas, indicating that the biomimetic visual perception module enhances the model's ability to perceive fine-grained disease features. However, some feature maps also show indistinct features, indicating the presence of redundant feature channels. Future work in this invention will focus on optimizing and improving the compression of redundant features in the model.

[0099] ② Visualization Experiment of Channel Relationship Hypergraph Inference Module: To intuitively analyze the channel interaction relationships of cross-layer features in the channel relationship hypergraph inference module, this invention first visualizes the channel correlation matrix after multi-path fusion. Figure 19 As can be seen, the matrix exhibits a clear blocky and striped distribution, indicating that the different channels are not uniformly correlated, but rather exhibit grouped dependencies, varying strengths, and potential redundancy. This suggests that simple channel concatenation or convolutional fusion is insufficient to fully characterize the complex semantic relationships between cross-layer features. Therefore, this invention further constructs a channel-hyperedge association structure based on this correlation matrix, transforming dense pairwise channel relationships into higher-order channel relationship representations, providing a basis for subsequent hypergraph inference.

[0100] Furthermore, this invention visualizes the hyperedge correlation matrix generated from the channel correlation matrix. The horizontal axis of this matrix represents the hyperedge index, the vertical axis represents the channel node index, and the color intensity indicates the correlation strength between different channels and their corresponding hyperedges. Figure 20 It is evident that the responses of different hyperedges to channel nodes are not uniform, with some hyperedges exhibiting strong activation within specific channel ranges. This indicates that the model can adaptively aggregate channels with similar semantics or related responses into the same hyperedge. Furthermore, significant differences exist in the response patterns among different hyperedges, suggesting that the channel relationship hypergraph inference module models higher-order dependencies between channels through hyperedge structure.

[0101] Finally, this invention provides a visual analysis of the feature map output by the channel relationship hypergraph inference module. Figure 21 As can be seen, after processing by the channel relation hypergraph inference module, the feature responses are more concentrated, while the background region is weaker compared to the original feature map. This indicates that the module can effectively enhance disease-related channel information and suppress interference from irrelevant features. Simultaneously, there are significant differences in the activation regions and response intensities of different channels, suggesting that the channel relation hypergraph inference module, through channel correlation modeling and hyperedge association aggregation, can extract key features such as lesion morphology, edges, and local textures from multiple semantic perspectives. These results demonstrate that the channel relation hypergraph inference module helps strengthen the effective interaction between cross-layer features and improves the model's ability to express fine-grained disease regions.

[0102] ③ Visualization Experiment of Spatial Relationship Hypergraph Reasoning Module Figure 22 The image shows the association matrix between pixel nodes and hyperedges generated by the spatial relation hypergraph inference module. It can be seen that the association strength between different pixel nodes and each hyperedge exhibits certain response differences within the local node range, indicating that the model can assign different hyperedge participation weights to different nodes based on pixel features.

[0103] The response regions and intensities of different hyperedges in pixel space show significant differences, indicating that the spatial relationship hypergraph inference module can adaptively aggregate pixel nodes with different locations, textures, and semantic relationships through multiple hyperedges to achieve diverse high-order spatial relationship modeling. The pixel node features, after being aggregated to hyperedges, form differentiated hyperedge representations, demonstrating that the spatial relationship hypergraph inference module can compress potentially related local pixel information into more compact high-order semantic features.

[0104] After completing the high-order feature aggregation from nodes to hyperedges, the spatial relation hypergraph inference module further transmits the hyperedge information back to the pixel nodes. Compared with the input pixel node features, the returned node response is stronger and the distribution is more continuous, indicating that hypergraph inference can enhance the information representation of related pixel regions and improve the consistency of spatial features.

[0105] The spatial relation hypergraph inference module showed strong responses to lesion edges, local textures, and abnormal areas in multiple channels, while suppressing some background interference. This indicates that pixel-level hypergraph inference can enhance the spatial representation and fine-grained feature extraction capabilities of key disease areas.

[0106] 4) To evaluate the advantages of the proposed model, its detection performance was compared with that of seven mainstream object detection models. The seven models were trained and evaluated on the same dataset. As shown in Table 2, under the same conditions, each detection model had varying degrees of missed detections. The model of this invention missed the fewest targets and achieved better detection results.

[0107] Table 2: Specifically, the model of this invention achieves an average accuracy of 0.769 with only 8.0M parameters, outperforming YOLOv8 and Faster R-CNN, demonstrating the superiority of sparse perception and hypergraph inference in mining latent features of rice diseases. Although its inference speed (22 FPS) is lower than that of the highly optimized YOLO series, this is mainly due to the introduction of higher-order inference operators such as region-related attention and hypergraph convolution, which increases the computational density and memory access overhead of feature routing, and it has not yet undergone hardware-level acceleration optimization such as TensorRT. However, in agricultural monitoring scenarios, this speed is sufficient to meet real-time requirements, and its high-precision recognition capability with low parameter count has significant practical value. Furthermore, since different detection models use different feature extraction networks and fusion methods, the information learned by different networks regarding key targets also differs. Due to differences in experimental configuration and reproduction details, the above metrics are for reference only.

[0108] 5) Some key hyperparameters in the model directly affect the feature modeling effect and the final detection performance. To further verify the rationality of the parameter settings of each module, this invention conducted debugging experiments on the relevant hyperparameters.

[0109] ① Hyperparameter tuning experiment of bionic visual perception module: Since the bionic visual perception module is built based on the two-layer routing attention idea of ​​BiFormer, its region partitioning method and Top-k sparse routing parameters follow the settings in BiFormer. Therefore, this invention does not repeatedly search the number of blocks and Top-k of the bionic visual perception module separately, but focuses on analyzing the introduction method of the bionic visual perception module in the features of different scales at the bottom layer of the backbone network.

[0110] Table 3: To verify the role of the biomimetic visual perception module at different low-level feature scales, this invention introduced the module into the last two feature layers of the backbone network for comparative experiments. As shown in Table 3, introducing the biomimetic visual perception module at only a single scale improved the model's average performance, but slightly reduced accuracy. The third-stage feature layer, with its higher spatial resolution, was more sensitive to lesion edges and detailed textures, resulting in a slightly higher performance improvement than the fourth stage. The model achieved optimal performance when the biomimetic visual perception module was introduced at both scales simultaneously. Although accuracy decreased slightly, recall significantly improved, a crucial metric for agricultural inspection.

[0111] ② Impact of the Number of Blocks in the Spatial Relationship Hypergraph Inference Module on Performance: The spatial relationship hypergraph inference module constructs cross-layer channel correlations through multi-path spatial scanning. The number of blocks determines the spatial granularity of cross-layer channel relationship modeling. If the number of blocks is too small, it is difficult to fully characterize the channel response differences between different regions; if the number of blocks is too large, it may introduce local noise and increase the difficulty of relationship modeling. Therefore, this invention conducts a comparative experiment on the number of blocks in the spatial relationship hypergraph inference module.

[0112] Table 4: To analyze the impact of the number of blocks in the spatial relationship hypergraph inference module on cross-layer channel relationship modeling, this invention conducted experiments with different numbers of blocks. As shown in Table 4, when the number of blocks is small, the spatial division is coarse, making it difficult to fully reflect the differences in channel responses between different regions, resulting in low model performance. With the increase in the number of blocks, the spatial relationship hypergraph inference module can construct cross-layer channel correlations in finer-grained spatial regions, and the detection performance gradually improves. The model achieves the best results when the number of blocks is 64. Further increasing to 100 may introduce local noise due to overly fine regional division, making the channel relationship modeling more fragmented, leading to a slight decrease in performance. Therefore, this invention sets the number of blocks in the spatial relationship hypergraph inference module to 64.

[0113] ③ Impact of Hyperedge Count on Performance: Both the channel relation hypergraph inference module and the spatial relation hypergraph inference module rely on hypergraph convolution for inference, and the key parameter of this technology is the number of hyperedges. Therefore, this invention conducted ablation and comparative experiments on the number of hyperedges, and adjusted the number of hyperedges in both modules. The results are shown in Table 5.

[0114] Table 5: Clearly, too few or too many hyperedges will result in insufficient information to express relationships or information redundancy. Therefore, we set the hyperparameters to be consistent with methods such as YOLOv13, setting the number of hyperedges to 16.

[0115] 6) Robustness Validation Experiments: In real-world detection scenarios, various environmental factors such as rainfall, fog, and changes in lighting can degrade image quality, thereby affecting model performance. To evaluate the robustness of the proposed method under these harsh conditions, we simulated four typical types of interference: noise, overexposure, backlighting, and blurring. These interferences were applied to the test dataset, generating four data subsets with degraded quality. The model performance was then evaluated on each of these affected datasets to verify its stability and generalization ability under adverse imaging conditions.

[0116] Depend on Figure 23 As can be seen, the model achieves the highest mAP (0.769) under normal conditions. Under noise and backlighting conditions, the mAP is 0.750 and 0.757 respectively, showing only a small performance drop, indicating that the model has good adaptability to certain levels of background interference and illumination changes. However, under overexposure and blurring conditions, the mAP drops to 0.723 and 0.701 respectively, indicating that the model's detection performance is more significantly affected when image details and textures are weakened or blemish edges become unclear. Overall, the model maintains a certain detection capability under various degradation scenarios, but it is more sensitive to blurring and strong exposure anomalies, suggesting that further improvements in degradation image enhancement, deblurring preprocessing, or adaptive illumination enhancement can further enhance the model's robustness in complex real-world scenarios.

[0117] This invention constructs a feature inference model for rice disease identification based on sparse anthropomorphic visual perception. In the feature extraction stage, it achieves dynamic coordination of coarse and fine-grained information, and in the feature fusion stage, it completes the modeling of channel and spatial high-order relationships, thereby effectively improving the ability to identify small targets, weak features, and latent lesions in complex field environments. Experimental results show that this method achieves stable improvements in precision, recall, and mAP, with an overall performance improvement of approximately 4.6%. It also demonstrates stronger robustness and generalization ability in complex backgrounds, multi-scale targets, and multi-class scenarios. Furthermore, while reducing the risk of missed detections, it verifies the effectiveness and synergistic advantages of anthropomorphic visual mechanisms and hypergraph relationship modeling in fine-grained target detection.

[0118] In the above embodiments, although the steps are numbered S1, S2, etc., they are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation. The scheme after adjusting the order is also within the protection scope of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.

[0119] like Figure 24 As shown, an embodiment of the present invention provides a rice disease identification system based on sparse anthropomorphic visual perception and hypergraph reasoning, which includes a model building module, a model training module, and a model application module. The model building module is used to: construct a rice disease identification model, which includes a backbone network, a neck structure, and a detection head; the backbone network receives rice sample images, and its shallow layers perform convolution operations on the rice sample images to output shallow detail features; the deep layers of the backbone network contain a biomimetic visual perception module, which receives and performs biomimetic visual perception on the shallow detail features to generate biomimetic visual perception features; the neck structure receives the biomimetic visual perception features and performs relational hypergraph reasoning and spatial relational hypergraph reasoning to generate multi-scale features; the detection head performs classification prediction and position regression prediction on the multi-scale features to obtain the disease identification results of the rice sample images; The model training module is used to train the rice disease identification model; The model application module is used to: input the target rice image into the trained rice disease identification model, and output the disease identification results of the target rice image.

[0120] Optionally, in the above technical solution, the bionic visual perception module is specifically used to: receive shallow detail features, divide the shallow detail features into multiple window regions, select multiple key regions with the highest relevance for each query region by calculating the correlation between window regions, and establish a region routing path; extract query features and key features according to the region routing path, calculate the Euclidean distance between the query position coordinates and the key position coordinates and generate a distance matrix, and perform weighted fusion of the convolution mapping result of the distance matrix with the attention weights to output the bionic visual perception features.

[0121] Optionally, in the above technical solution, the neck structure includes a channel relationship hypergraph inference module and a spatial relationship hypergraph inference module. The channel relationship hypergraph inference module performs multi-directional spatial rearrangement of the deep semantic features in the biomimetic visual perception features and the shallow texture features in the shallow detail features, calculates the similarity matrix, generates a first hyperedge matrix based on the similarity matrix, and performs hypergraph convolution to output channel relationship-enhanced fusion features. The spatial relationship hypergraph inference module flattens the shallow detail features into a sequence of pixel nodes, generates a hyperedge prototype based on the global context of the pixel node sequence, calculates the correlation between each pixel node and the hyperedge prototype to generate a second hyperedge matrix, and performs hypergraph convolution to output spatial relationship-enhanced detail features. The neck structure generates multi-scale features based on the channel relationship-enhanced fusion features and the spatial relationship-enhanced detail features.

[0122] Optionally, in the above technical solution, the neck structure generates multi-scale features based on the channel relationship enhancement fusion features and the spatial relationship enhancement detail features, including: concatenating the channel relationship enhancement fusion features and the spatial relationship enhancement detail features, performing convolution operations on the concatenated features to aggregate channel information and spatial information, and outputting multi-scale features.

[0123] Optionally, in the above technical solution, the spatial relationship hypergraph inference module generates a hyperedge prototype based on the global context of the pixel node sequence, including: calculating the maximum value and average value of the features of each pixel node in the pixel node sequence, concatenating the maximum value and average value to obtain the global context features, generating an offset from the global context features through a mapping layer, and adding the offset to a preset hyperedge initial prototype vector to obtain the hyperedge prototype.

[0124] It should be noted that the beneficial effects of the rice disease identification system based on sparse anthropomorphic visual perception and hypergraph reasoning provided in the above embodiments are the same as those of the rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.

[0125] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned rice disease identification methods based on sparse anthropomorphic visual perception and hypergraph reasoning.

[0126] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-mentioned rice disease identification methods based on sparse anthropomorphic visual perception and hypergraph reasoning.

[0127] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for identifying rice diseases based on sparse anthropomorphic visual perception and hypergraph reasoning, characterized in that, include: A rice disease identification model is constructed, comprising a backbone network, a neck structure, and a detection head. The backbone network receives rice sample images, and its shallow layers perform convolution operations on the images to output shallow detail features. A biomimetic visual perception module is located in the deep layers of the backbone network to receive and perform biomimetic visual perception on the shallow detail features, generating biomimetic visual perception features. The neck structure receives these features and performs relational hypergraph reasoning and spatial relational hypergraph reasoning to generate multi-scale features. The detection head performs classification prediction and position regression prediction on these multi-scale features to obtain the disease identification results for the rice sample images. The rice disease identification model is trained. The target rice image is input into the trained rice disease identification model, and the disease identification result of the target rice image is output.

2. The rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning according to claim 1, characterized in that, The biomimetic visual perception module is specifically used for: receiving the shallow detail features, dividing the shallow detail features into multiple window regions, selecting multiple key regions with the highest relevance for each query region by calculating the correlation between the window regions, and establishing a region routing path; extracting query features and key features according to the region routing path, calculating the Euclidean distance between the query position coordinates and the key position coordinates and generating a distance matrix, and weighting and fusing the convolution mapping result of the distance matrix with the attention weights to output the biomimetic visual perception features.

3. The rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning according to claim 2, characterized in that, The neck structure includes a channel relationship hypergraph inference module and a spatial relationship hypergraph inference module. The channel relationship hypergraph inference module performs multi-directional spatial rearrangement of the deep semantic features in the biomimetic visual perception features and the shallow texture features in the shallow detail features, calculates a similarity matrix, generates a first hyperedge matrix based on the similarity matrix, performs hypergraph convolution, and outputs a fusion feature with enhanced channel relationships. The spatial relationship hypergraph inference module flattens the shallow detail features into a sequence of pixel nodes, generates a hyperedge prototype based on the global context of the pixel node sequence, calculates the correlation between each pixel node and the hyperedge prototype to generate a second hyperedge matrix, performs hypergraph convolution, and outputs a detail feature with enhanced spatial relationships. The neck structure generates multi-scale features based on the fusion features enhanced by the channel relationships and the detail features enhanced by the spatial relationships.

4. The rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning according to claim 3, characterized in that, The neck structure generates multi-scale features based on the channel relationship enhancement fusion features and the spatial relationship enhancement detail features, including: concatenating the channel relationship enhancement fusion features and the spatial relationship enhancement detail features, performing a convolution operation on the concatenated features to aggregate channel information and spatial information, and outputting multi-scale features.

5. A rice disease identification system based on sparse anthropomorphic visual perception and hypergraph reasoning, characterized in that, It includes a model building module, a model training module, and a model application module; The model building module is used to: construct a rice disease identification model, which includes a backbone network, a neck structure, and a detection head; the backbone network is used to receive rice sample images, and the shallow layers of the backbone network perform convolution operations on the rice sample images to output shallow detail features; the deep layers of the backbone network are equipped with a biomimetic visual perception module, which is used to receive and perform biomimetic visual perception on the shallow detail features to generate biomimetic visual perception features; the neck structure is used to receive the biomimetic visual perception features and perform relational hypergraph reasoning and spatial relational hypergraph reasoning to generate multi-scale features; the detection head performs classification prediction and position regression prediction on the multi-scale features to obtain the disease identification results of the rice sample images; The model training module is used to train the rice disease identification model. The model application module is used to: input the target rice image into the trained rice disease identification model, and output the disease identification result of the target rice image.

6. The rice disease identification system based on sparse anthropomorphic visual perception and hypergraph reasoning according to claim 5, characterized in that, The biomimetic visual perception module is specifically used for: receiving the shallow detail features, dividing the shallow detail features into multiple window regions, selecting multiple key regions with the highest relevance for each query region by calculating the correlation between the window regions, and establishing a region routing path; extracting query features and key features according to the region routing path, calculating the Euclidean distance between the query position coordinates and the key position coordinates and generating a distance matrix, and weighting and fusing the convolution mapping result of the distance matrix with the attention weights to output the biomimetic visual perception features.

7. A rice disease identification system based on sparse anthropomorphic visual perception and hypergraph reasoning according to claim 6, characterized in that, The neck structure includes a channel relationship hypergraph inference module and a spatial relationship hypergraph inference module. The channel relationship hypergraph inference module performs multi-directional spatial rearrangement of the deep semantic features in the biomimetic visual perception features and the shallow texture features in the shallow detail features, calculates a similarity matrix, generates a first hyperedge matrix based on the similarity matrix, performs hypergraph convolution, and outputs a fusion feature with enhanced channel relationships. The spatial relationship hypergraph inference module flattens the shallow detail features into a sequence of pixel nodes, generates a hyperedge prototype based on the global context of the pixel node sequence, calculates the correlation between each pixel node and the hyperedge prototype to generate a second hyperedge matrix, performs hypergraph convolution, and outputs a detail feature with enhanced spatial relationships. The neck structure generates multi-scale features based on the fusion features enhanced by the channel relationships and the detail features enhanced by the spatial relationships.

8. A rice disease identification system based on sparse anthropomorphic visual perception and hypergraph reasoning according to claim 7, characterized in that, The neck structure generates multi-scale features based on the channel relationship enhancement fusion features and the spatial relationship enhancement detail features, including: concatenating the channel relationship enhancement fusion features and the spatial relationship enhancement detail features, performing a convolution operation on the concatenated features to aggregate channel information and spatial information, and outputting multi-scale features.

9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning as described in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the rice disease identification method based on sparse anthropomorphic visual perception and hypergraph reasoning as described in any one of claims 1 to 4.