A method and system for autonomous picking of agaricus bisporus based on vla

By constructing multi-scale visual feature maps and cross-attention calculations, and combining them with the picking task description information, an object-level visual label sequence is generated. This solves the problem that existing VLA models have difficulty distinguishing between foreground and background in the mushroom cultivation environment, and achieves high accuracy and robustness in generating picking actions.

CN121464885BActive Publication Date: 2026-03-17SHANGHAI HENGZE FUHUI INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610013450.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-03-17
Estimated Expiration
2046-01-07

AI Technical Summary

Technical Problem

Existing general VLA models have difficulty distinguishing between foreground objects and background noise when dealing with button mushroom cultivation environments. This results in redundant information in the visual feature sequence, diluting key mushroom features and making it difficult to achieve accurate harvesting in complex and dense scenes.

Method used

By constructing multi-scale visual feature maps and performing cross-attention calculations, an object-level visual tag sequence is generated. The picking task description information is encoded into a text tag sequence to form a joint semantic sequence. This sequence is then input into the pre-trained VLA model backbone for autoregressive prediction to generate an action tag sequence.

Benefits of technology

In complex and dense cultivation rack scenarios, the accuracy and robustness of motion generation are improved, the risks of mis-collection, missed collection and trajectory instability are reduced, and end-to-end multimodal reasoning and motion generation are realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121464885B_ABST
    Figure CN121464885B_ABST
Patent Text Reader

Abstract

The application provides an autonomous picking method and system for Agaricus bisporus based on VLA, and relates to the technical field of cultivation. The method comprises the following steps: acquiring picking task description information and visual observation data, encoding the visual observation data to obtain a visual feature sequence; constructing a multi-scale visual feature map based on the visual feature sequence, and generating an object-level visual label sequence; encoding the picking task description information into a text label sequence, mapping the object-level visual label sequence to a unified semantic space, and splicing to form a joint semantic sequence; inputting the joint semantic sequence into a pre-trained VLA model, and generating an action label sequence through autoregressive prediction; decoding the action label sequence to obtain picking control instructions, and driving a picking device to complete picking, so as to improve the key visual information extraction capability in a complex and unstructured cultivation scene, and enhance the accuracy and robustness of action generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cultivation technology, and more specifically, to a method and system for autonomous harvesting of button mushrooms based on VLA. Background Technology

[0002] As an important edible fungus, the industrial cultivation of button mushrooms has gradually shifted towards high-density, three-dimensional models. In the automated harvesting of button mushrooms, accurately identifying mature mushrooms and generating precise robotic arm control commands is crucial for achieving unmanned operation. In recent years, with the rise of embodied AI technology, the Vision-Language-Action (VLA) large-scale model has become a research hotspot in the field of agricultural robot control due to its ability to understand natural language commands and directly generate action sequences.

[0003] Existing general-purpose VLA models, such as RT-2, typically employ a standard Vision Transformer (VIT) architecture when processing visual information. This architecture directly segments the input visually observed image into fixed-size patches and linearly maps these patches into long sequences of visual tokens. These tokens are then concatenated with text instruction tokens and input into the model backbone for inference. However, this general visual encoding method proves insufficient when dealing with the cultivation environment of Agaricus bisporus.

[0004] The growth environment of button mushrooms is highly unstructured and dense, with cultivation racks often containing complex background structures such as mushroom bags, covering soil, and supports. Furthermore, mushrooms at different growth stages vary significantly in size and may occlude each other. Existing direct image patching methods treat foreground objects and background noise indiscriminately, resulting in visual feature sequences containing a large amount of redundant information representing soil or supports, diluting crucial mushroom features. Simultaneously, single-scale image patch encoding struggles to simultaneously capture tiny young mushrooms and represent large, mature mushrooms holistically. These issues directly cause VLA models to struggle to accurately focus on the target object from noisy visual input during inference, leading to biased picking action generation, inaccurate positioning, and even accidental damage to surrounding strains.

[0005] Therefore, how to improve the VLA model's ability to extract key visual information in complex and dense unstructured scenarios, so as to enhance the accuracy and robustness of action generation, has become a technical problem that VLA-based agricultural harvesting robots urgently need to solve. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this application provides a method and system for autonomous harvesting of button mushrooms based on VLA.

[0007] Firstly, this application provides a method for autonomous harvesting of button mushrooms based on VLA, including:

[0008] Obtain the picking task description information and visual observation data of the target scene; encode the visual observation data to obtain a visual feature sequence;

[0009] Based on the visual feature sequence, a multi-scale visual feature map is constructed, and the cross-attention between the object query vector and the multi-scale visual feature map is used to generate an object-level visual tag sequence.

[0010] The picking task description information is encoded into a text tag sequence, and the text tag sequence is mapped to a unified semantic space and concatenated with the object-level visual tag sequence to obtain a joint semantic sequence; the joint semantic sequence is input into the backbone of a pre-trained VLA model, and an action tag sequence is generated through autoregressive prediction.

[0011] The harvesting control command is obtained by decoding the action marker sequence, and the harvesting equipment is controlled to perform the button mushroom harvesting action based on the harvesting control command.

[0012] Optionally, encoding the visual observation data to obtain a visual feature sequence includes:

[0013] Within a preset time window, acquire multiple consecutive frames of visual observation images, divide the multiple frames of visual observation images into multiple image blocks, and combine the image blocks corresponding to spatial locations in different time frames to form a spatiotemporal unit.

[0014] Linear mapping and spatiotemporal location encoding are performed on each of the aforementioned spatiotemporal units to generate a temporal visual marker sequence, which serves as the visual feature sequence.

[0015] Optionally, performing linear mapping and spatiotemporal location encoding on each of the spatiotemporal units includes:

[0016] Based on the pre-defined geometric layout information of the button mushroom cultivation rack, a corresponding hierarchical index, row index, and column index are assigned to each spatiotemporal unit. The hierarchical index, row index, and column index are embedded and mapped to obtain a structural position code. The structural position code is then superimposed with the spatial position code for the image block to serve as the spatiotemporal position code, so as to distinguish spatiotemporal units of different levels, rows, and columns in the temporal visual marker sequence.

[0017] Optionally, before performing linear mapping and spatiotemporal location encoding on each of the aforementioned spatiotemporal units, the method further includes:

[0018] Based on the mushroom cultivation images, the cap edge region in each frame of visual observation image is labeled, and the spatiotemporal units falling into the cap edge region are marked as high-sensitivity regions, while the spatiotemporal units falling into the background structure region are marked as low-sensitivity regions.

[0019] The background structure region includes a bag region, a support region, and a cultivation substrate region that does not contain the cap of Agaricus bisporus. When encoding the spatiotemporal position of the spatiotemporal unit of the high-sensitivity region, a first time weight is used, and when encoding the spatiotemporal position of the spatiotemporal unit of the low-sensitivity region, a second time weight different from the first time weight is used, so as to generate the temporal visual marker sequence that is sensitive to the local deformation of the cap edge.

[0020] Optionally, the calibration of the cap edge region in each frame of visual observation image includes:

[0021] Using a segmentation or detection network trained on Agaricus bisporus, reasoning is performed on the visually observed image to obtain a segmentation mask or target box characterizing the cap region.

[0022] Edge detection and region growing are performed on the visually observed image based on the geometric shape features and gray-scale distribution features of the cap's outer contour, and the region where the cap edge is located is extracted as the cap edge region.

[0023] Optionally, constructing a multi-scale visual feature map based on the visual feature sequence includes:

[0024] At least two visual feature maps with different spatial resolutions are obtained based on the visual feature sequence;

[0025] For each of the aforementioned visual feature maps, a geometric response index is calculated within a preset local neighborhood. The geometric response index reflects the edge curvature or second-order gradient intensity within the local neighborhood.

[0026] For the same spatial location, geometric saliency weights are determined based on the geometric response indices of that spatial location at different spatial resolutions. Then, the visual feature maps are weighted and fused at their corresponding spatial locations according to the geometric saliency weights to obtain the multi-scale visual feature maps.

[0027] Optionally, the calculation of the geometric response index for each of the visual feature maps within a preset local neighborhood includes:

[0028] Based on the hierarchical index, row index, and column index of each spatiotemporal unit, a local neighborhood aligned with the structure of the button mushroom cultivation rack is defined for each spatial location in the visual feature map. The local neighborhood extends along the row and column directions of the cultivation rack and is limited to the spatial location range corresponding to the same and / or adjacent levels.

[0029] In the local neighborhood, first-order derivative convolutions are performed on each of the visual feature maps in the horizontal and vertical directions to obtain gradient components. A local structure tensor is constructed based on the gradient components, and the gradient intensity of the principal direction of the structure tensor is calculated. The gradient intensity of the principal direction is used as the geometric response index.

[0030] Optionally, the generation of the object-level visual marker sequence includes:

[0031] Based on the hierarchical index, row index, and column index of each spatiotemporal unit, the structural anchor point corresponding to the mushroom cultivation rack structural unit is determined, and the structural anchor point vector obtained by embedding mapping of each structural anchor point is combined with the preset learnable query basis vector to generate an object query vector with cultivation rack structural location information.

[0032] Based on the corresponding spatial position of each of the structural anchor points in the multi-scale visual feature map, an attention mask region is constructed that is aligned with the corresponding cultivation rack level and row and column unit. The attention mask region is limited to a preset row and column neighborhood that is at the same level and / or adjacent level as the structural anchor point.

[0033] Within each attention mask region, cross-attention calculation is performed using the corresponding object query vector as the query vector and the feature vectors in the multi-scale visual feature map that fall into the attention mask region as the key vector and value vector. The cross-attention output is used as the object-level visual label of the corresponding structural anchor point to obtain the object-level visual label sequence.

[0034] Optionally, obtaining the joint semantic sequence includes:

[0035] Based on a preset semantic dictionary, the text tag sequence and the object-level visual tag sequence are input into a unified semantic mapping module and mapped to a first semantic representation and a second semantic representation located in a common semantic representation space, respectively. The semantic dictionary includes multiple semantic atoms used to characterize the level, row and column positions of the mushroom cultivation rack, as well as the maturity, size, and relative positional relationship of the target mushroom with the support structure of the cultivation rack.

[0036] Within the common semantic representation space, a sparse coefficient vector with respect to the semantic dictionary is solved for the first semantic representation to obtain the first semantic sparse coefficients; a sparse coefficient vector with respect to the semantic dictionary is solved for each of the second semantic representations to obtain a set of second semantic sparse coefficients that correspond one-to-one with each object-level visual tag.

[0037] The first semantic sparse coefficient and the coefficients of each of the second semantic sparse coefficients that correspond to the semantic atoms used to characterize the cultivation rack level, row and column position, target mushroom maturity, size, and the relative positional relationship between the target mushroom and the cultivation rack support structure are attached as semantic constraint coefficients to the corresponding object-level visual tags and arranged together with the text tag sequence to form the joint semantic sequence.

[0038] Secondly, this application provides a VLA-based autonomous harvesting system for button mushrooms, comprising:

[0039] The acquisition module is used to acquire the picking task description information and visual observation data of the target scene; and to encode the visual observation data to obtain a visual feature sequence.

[0040] The first generation module constructs a multi-scale visual feature map based on the visual feature sequence, and generates an object-level visual tag sequence by calculating the cross attention between the object query vector and the multi-scale visual feature map.

[0041] The second generation module is used to encode the picking task description information into a text tag sequence, map the text tag sequence and the object-level visual tag sequence to a unified semantic space and concatenate them to obtain a joint semantic sequence; input the joint semantic sequence into the backbone of a pre-trained VLA model, and generate an action tag sequence through autoregressive prediction;

[0042] The control module is used to decode the action marker sequence to obtain the picking control command, and control the picking equipment to perform the button mushroom picking action based on the picking control command.

[0043] Compared with existing technologies, this application constructs visual and semantic representations for harvesting tasks at the VLA front end, forming a unified decision-making chain from visual feature sequences, multi-scale visual feature maps, object-level visual labels to joint semantic sequences and action label sequences. This enables the model to specifically enhance key visual information related to harvesting decisions in complex and dense cultivation rack scenarios. On the one hand, multi-scale visual feature maps are constructed based on visual feature sequences, and cross-attention calculation is introduced between object query vectors and multi-scale visual feature maps. Visual label sequences are extracted at the object-level granularity, allowing the VLA model to focus on specific button mushroom targets and their spatial relationships, rather than just global coarse-grained image features. Thus, even under conditions of severe occlusion and complex background structures, it can still obtain object-level visual representations with high discriminative power.

[0044] On the other hand, this application encodes the harvesting task description information into a text tag sequence, maps it to a unified semantic space with the object-level visual tag sequence, and concatenates them to construct a joint semantic sequence. Then, the pre-trained VLA model backbone directly generates the action tag sequence in an autoregressive manner, realizing end-to-end multimodal reasoning from natural language task description and visual scene to harvesting action sequence. Compared with the existing segmented architecture of detection results, rule planning, and control instructions, this application jointly models task semantics, object-level visual information, and action decisions within the same model. This allows the action generation process to more closely rely on key visual cues and task constraints, reducing information loss and engineering coupling problems caused by intermediate representation transformation. This is beneficial for improving the accuracy and robustness of action generation in complex and dense unstructured cultivation scenarios, and reducing the risks of false sampling, missed sampling, and trajectory instability. Attached Figure Description

[0045] Figure 1 A flowchart illustrating a VLA-based autonomous harvesting method for button mushrooms provided in this application embodiment;

[0046] Figure 2 A flowchart illustrating a method for obtaining a visual feature sequence provided in this application embodiment;

[0047] Figure 3 A flowchart illustrating a method for constructing multi-scale visual feature maps provided in this application embodiment;

[0048] Figure 4 This is a schematic diagram of a VLA-based autonomous harvesting system for button mushrooms provided in an embodiment of this application. Detailed Implementation

[0049] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0050] See Figure 1 The diagram shows a flowchart of a VLA-based autonomous harvesting method for button mushrooms provided in this application, including steps S101 to S104, wherein:

[0051] S101: Obtain the picking task description information and visual observation data of the target scene; encode the visual observation data to obtain a visual feature sequence;

[0052] S102: Construct a multi-scale visual feature map based on the visual feature sequence, and generate an object-level visual tag sequence by calculating the cross-attention between the object query vector and the multi-scale visual feature map.

[0053] S103: Encode the picking task description information into a text tag sequence, and map the text tag sequence and the object-level visual tag sequence to a unified semantic space and concatenate them to obtain a joint semantic sequence; input the joint semantic sequence into the backbone of a pre-trained VLA model, and generate an action tag sequence through autoregressive prediction;

[0054] S104: Decode the harvesting control command according to the action marker sequence, and control the harvesting equipment to perform button mushroom harvesting action based on the harvesting control command.

[0055] Regarding the above S101:

[0056] In one implementation, S101 is used to acquire picking task description information and visual observation data, and to encode the visual observation data to obtain a visual feature sequence for subsequent VLA inference.

[0057] In this embodiment, the harvesting task description information can be obtained in various forms, not limited to a single input channel. For example, operators can select preset harvesting strategy templates in the graphical user interface via a touchscreen, host computer interface, or mobile terminal app, such as "prioritize harvesting mature mushrooms," "harvest only the second layer of mushrooms," and "avoid mushrooms near the supports." Natural language commands can also be collected through a voice input device and transcribed into text form by a speech recognition module. Structured task commands can also be issued through the production management system. Regardless of the method used, this embodiment uniformly represents the final task content as harvesting task description information oriented towards the VLA model, so that text encoding and semantic alignment can be performed in subsequent steps.

[0058] Visual observation data can be obtained by one or more visual acquisition devices deployed in the button mushroom cultivation environment. Specifically, the visual acquisition devices may include an industrial camera fixedly mounted in front of the cultivation rack, a sliding rail camera that can move along the cultivation channel, a top-view camera mounted at the end of a robotic arm, a depth camera, or a composite camera with multispectral and infrared imaging capabilities, or any combination of the above devices. In order to cover multiple cultivation racks and different perspectives, multiple cameras can be configured as needed in this embodiment to achieve joint observation in the horizontal, vertical, or top-view manner; the camera resolution, frame rate, exposure parameters, etc., can be adjusted according to factors such as scene brightness, mushroom size, and movement speed, and are not limited to specific values.

[0059] During implementation, the control unit can trigger the vision acquisition device to acquire images of the target cultivation area according to a preset sampling period or based on the movement state of the robotic arm. The visual observation data can be a single frame image, a sequence of multiple frames acquired within a preset time window, or other visual modalities such as depth maps or point cloud data.

[0060] In one alternative implementation, in order to better reflect the changes in the relationship between the movement of the robotic arm and the occlusion of the mushroom, multiple frames of images can be continuously acquired as the robotic arm approaches the target area, forming a temporally continuous visual observation sequence.

[0061] To facilitate subsequent VLA model processing, the visual observation data can undergo necessary preprocessing before being input into the visual encoding network.

[0062] Preprocessing operations may include, but are not limited to: image denoising, adaptive adjustment of brightness and contrast, color space conversion, geometric distortion correction, image cropping and scaling, and distortion region masking. For example, in some industrialized mushroom houses with unstable lighting conditions, histogram equalization or brightness normalization can be performed on the original image to reduce the impact of light flicker on feature extraction; similarly, for wide-angle lenses mounted at the end of robotic arms, distortion correction can be performed first to reduce geometric deformation in edge regions.

[0063] In the visual feature encoding stage, this implementation can use various types of visual encoding networks to extract features from the preprocessed visual observation data, and is not limited to a specific network structure.

[0064] In one implementation, the visual coding network can be a multi-layer convolutional neural network that extracts feature maps at different semantic levels through cascaded convolution and downsampling operations. In another implementation, the visual coding network can adopt a visual transformer (ViT) architecture, which divides the image into several image blocks, and inputs them into a multi-layer self-attention network after linear mapping and position encoding. Alternatively, a hybrid structure combining convolutional networks and transformers can be used, where the convolutional layers are responsible for local texture modeling and the transformer layers are responsible for global relation modeling under a larger receptive field.

[0065] When processing multi-frame visual observation data, each frame image can be input into a visual coding network to obtain frame-level features. Then, the features of each frame can be stacked or concatenated in a preset manner according to the time order to obtain a visual feature sequence that reflects the temporal changes of the target scene. Alternatively, multiple frames can be input into the visual coding network in a channel or batch manner at the input end, and the network can learn the temporal correlation in the spatial-temporal joint domain.

[0066] This implementation method is not limited to a specific multi-frame fusion method. The key is to convert the visual observation data containing information such as mushroom clusters, bags, and supports in the mushroom cultivation scene into a one-dimensional or multi-dimensional visual feature sequence through an adapted encoding strategy, so that the feature sequence can be used as the input source of the "visual label sequence" by subsequent modules.

[0067] In a typical implementation, the visual encoding network can run on edge computing devices with GPUs or dedicated acceleration chips, such as industrial computers or embedded controllers equipped with deep learning inference engines. The system can deploy pre-trained or scene-fine-tuned visual encoding models based on existing deep learning frameworks, and use batch processing or pipelined methods to encode continuously acquired visual observation data in real time or near real time, so as to ensure that the generated visual feature sequences can be provided to subsequent multi-scale feature construction and object query decoding modules in a timely manner.

[0068] Regarding S102 above:

[0069] In one implementation, the visual feature sequence obtained in step S101 is used to drive the downstream multi-scale feature construction and object-level feature decoding process.

[0070] In this embodiment, the visual feature sequence can be a one-dimensional feature vector sequence output by a visual coding network, or a two-dimensional or three-dimensional feature tensor arranged by frame or by channel. To facilitate detailed modeling of the structure of the button mushroom clusters, bags, and cultivation racks in the spatial domain, the control terminal can first restore or map the visual feature sequence into a basic feature map with a clear spatial arrangement relationship according to the generation method of the visual feature sequence. For example, when the visual feature sequence comes from the labeled sequence obtained by linear mapping after dividing the input image into image blocks, the feature sequence can be rearranged into a two-dimensional feature map according to the row and column positions of the image blocks in the original image; when the visual feature sequence comes from the intermediate layer features output by the convolutional network, the intermediate layer features can be directly regarded as the basic feature map.

[0071] After obtaining the basic feature map, to take into account the different scales of the button mushroom targets and the overall structure of the cultivation rack, this implementation method can construct multi-scale visual feature maps through multi-layer convolution, pooling, convolution with stride, upsampling, or feature pyramid fusion. For example, several feature maps with gradually decreasing resolution can be constructed from the basic feature map to capture the overall layout of larger mature mushroom clusters and the cultivation rack; alternatively, higher resolution feature maps can be obtained by upsampling and fusing with shallow features to enhance the ability to express cap edges, young mushrooms, and local occlusion relationships.

[0072] In different implementations, the number of layers, resolution combinations, and number of channels of the multi-scale visual feature map can be flexibly configured according to hardware computing power, camera resolution, and target mushroom size distribution, and are not limited to a specific hierarchical structure.

[0073] In order to extract object-level representations related to the picking task from multi-scale visual feature maps, this implementation introduces an object query vector as a query signal at the decoding end.

[0074] The object query vectors can be pre-set as several learnable vectors, with initial values ​​randomized or initialized based on prior tasks. They can also be configured according to information such as the geometric layout of the mushroom cultivation rack, the number of cultivation rows and columns, and the expected maximum number of targets. In a simple implementation, a fixed number of object query vectors can be set to cover multiple mushroom targets that may exist in the current field of view.

[0075] When performing object-level feature decoding, a multi-layered cross-attention module can be used to enable information interaction between the object query vector and multi-scale visual feature maps. Specifically, within a cross-attention layer, the object query vector of the current layer is used as the query end, and the feature vectors of one or more scales of visual feature maps after spatial expansion are used as the key and value ends. The cross-attention module assigns attention weights to features at different locations based on the correlation between the query vector and the features at each spatial location, and performs weighted aggregation of the features to obtain the updated object query representation.

[0076] The aforementioned cross-attention operation can be performed on a single-scale feature map, or it can be performed sequentially or in parallel on feature maps of multiple scales. For example, the approximate location can be obtained first on a lower-resolution feature map, and then the edges and local structures can be refined on a higher-resolution feature map.

[0077] To better adapt to dense mushroom cluster scenarios, the specific implementation of cross-attention in this embodiment can combine multi-head attention, residual connections, normalization, and feedforward networks to enhance the model's ability to model complex spatial relationships. In some embodiments, a self-attention layer can be inserted before or after cross-attention, enabling object query vectors to exchange information about mutual occlusion, cluster-level distribution, etc., further improving the object-level representation's ability to characterize mushroom cluster relationships.

[0078] After several layers of cross-attention and optional self-attention updates, each object query vector will aggregate contextual information from multi-scale visual feature maps to form an object-level visual representation corresponding to a candidate mushroom target or target cluster in the scene. In this implementation, the updated object query vector can be regarded as an object-level visual marker, and arranged in a predetermined order to form an object-level visual marker sequence.

[0079] This implementation method does not limit the specific scoring or screening mechanism. As long as it can provide a set of object-level visual tags that can represent the key targets and their relationships in the scene for the subsequent joint semantic modeling stage, under the premise of ensuring that the model complexity is controllable.

[0080] In this way, the visual features that were originally mixed with foreground mushrooms, mushroom bags, supports and other background textures are transformed into an object-level visual label sequence oriented towards the specific picking target. This lays the visual foundation for subsequent alignment with the picking task description information in a unified semantic space and for driving the VLA model to generate action sequences.

[0081] Regarding the above S103:

[0082] Based on the aforementioned implementation method, step S103 can be used to fuse the picking task description information and the object-level visual tag sequence in a unified semantic space, and drive the pre-trained VLA model backbone to generate action tag sequences, thereby completing the reasoning from natural language and scene perception to action sequences.

[0083] In one implementation, the harvesting task description information is first encoded into a text tag sequence. The harvesting task description information can be natural language text input by the operator, such as "harvest the larger mature mushrooms in the first two rows of the second layer" or "skip the mushrooms near the support and only pick the mushrooms on the outer side that have not yet opened their caps," or it can be a semi-structured description generated by combining options in a graphical interface, such as "target layer = 2, target row = 1~2, size = large, status = mature."

[0084] For natural language text, basic preprocessing can be performed first, including character set normalization, removal of irrelevant control characters, and standardization of number and unit expressions as needed. Then, tokenization can be performed using word segmentation or sub-word segmentation tools. For example, word-level segmentation, sub-word-based encoding methods (such as Byte-Pair Encoding or SentencePiece), or the built-in tokenizer of a specific model can be used to split the entire task description into a string of discrete text tokens.

[0085] For structured or semi-structured descriptions, they can be converted into a sequence of tags consistent with natural language using fixed templates or simple mapping rules, so that task information from different sources maintains a consistent input format during the encoding stage.

[0086] After tokenization, this implementation can map the text token sequence into a vector representation suitable for VLA model processing using a text encoding module. The text encoding module can be a standalone lightweight language encoder, or it can share some embedding layers or the first few layers with the VLA model backbone to improve parameter reusability and semantic consistency. Specifically, for each text token, a word vector or sub-word vector can be retrieved by looking up a table, and positional representation information can be superimposed according to the token order. For example, absolute position indexes, relative position labels, or other positional information suitable for sequence modeling can be used to encode the text token vector sequence representing the semantics of the picking task.

[0087] Secondly, the text tag sequence and the object-level visual tag sequence are mapped to a unified semantic space and concatenated to obtain a joint semantic sequence. To this end, this implementation can set up a set of cross-modal mapping layers to project the object-level visual tags and text tags to representation domains with consistent dimensions and compatible semantic spaces.

[0088] Object-level visual tags can be transformed and normalized through linear transformation, multilayer perceptron, or lightweight attention layers, converting the original visual feature-based tags into high-level semantic vectors that simultaneously contain target category, spatial relationship, and scene context; text tags are then processed through corresponding projection layers to match the scale and distribution of the object-level visual tags in the vector space.

[0089] In some implementations, modality identifiers or modality embedding vectors can be reserved for different modalities to explicitly distinguish between tags “from task description” and tags “from scene awareness” in a unified semantic space.

[0090] After completing the cross-modal mapping, a joint semantic sequence can be generated according to a predetermined concatenation strategy. The concatenation order can be configured according to task requirements. For example, text tags describing the global task intent can be placed at the beginning of the sequence, while object-level visual tags can be arranged at the end of the sequence according to certain rules (such as by their level, row, or importance score predicted by the model). Alternatively, one or more separator tags or global task tags can be inserted between text tags and visual tags to explicitly identify the semantic scope of different regions.

[0091] This implementation does not limit the specific splicing order and the number of markers, as long as the joint semantic sequence can simultaneously carry the semantics of the picking task and the object-level visual information in the same sequence structure.

[0092] Subsequently, the joint semantic sequence is input into the pre-trained VLA model backbone, and action tag sequences are generated through autoregressive prediction. The VLA model backbone can use a multi-layer transformer network as the basic structure, or it can use other sequence modeling architectures that include self-attention layers, cross-attention layers, and feedforward networks.

[0093] In one implementation, the VLA model backbone can be viewed as a conditional sequence generation model, which outputs action tags in an autoregressive manner under the given joint semantic sequence.

[0094] Specifically, during the training phase, common strategies can be used to enable the model to predict the next action tag in the current step based on the joint semantic sequence and the previously generated action tags at each step. During the inference phase, a greedy search, bundle search, or random sampling strategy can be used to select specific action tags from the predicted action tag probability distribution and gradually expand the action tag sequence until an end tag is generated or a preset length is reached.

[0095] Motion marker sequences are used to describe the motion planning results of robotic arms or harvesting devices in discrete form. The semantics of motion markers may differ in different implementations.

[0096] For example, in one implementation, each action marker can correspond to an operation instruction for a specific object-level visual marker, such as "select target i," "move above the target," "press down to contact," "close the gripper," or "retract to a safe height." In another implementation, the action markers can be further refined into end-effector pose discrete units, posture adjustment instructions, gripping force levels, or safety check instructions, forming a hierarchical action sequence from high-level action intentions to low-level action primitives. By appropriately designing the action marker vocabulary, it is possible to cover the common multi-step operation process in button mushroom harvesting while ensuring that the sequence length is controllable.

[0097] In practical deployment, the VLA model backbone can use a large model pre-trained on a general robot operation dataset as a foundation, and further fine-tune it on data from a button mushroom harvesting scenario so that the model can better adapt to the geometry of the cultivation rack, mushroom growth characteristics, and harvesting strategy constraints. This implementation can choose to completely freeze the pre-trained backbone and only perform minor fine-tuning on the cross-modal mapping layer and action output head, or it can perform joint fine-tuning on some backbone layers to achieve a balance between model capacity and adaptation performance.

[0098] Regarding S104 above:

[0099] Step S104 is used to convert the action marker sequence generated by the VLA model backbone into specific picking control instructions, and drive the picking equipment to perform picking actions in the actual button mushroom cultivation environment.

[0100] For ease of understanding, the implementation of step S104 will be described below with reference to a typical robotic arm actuator, but this application is not limited to a specific type of actuator or control system.

[0101] In one implementation, the action marker sequence can be viewed as a discretized description of the harvesting task execution process. Each action marker corresponds to a high-level action primitive or motion stage, such as "select target i," "move above the target," "adjust posture," "approach and contact," "close gripper," "retreat to a safe posture," and "place the mushroom in the designated collection container." During the actual decoding process, the control unit first parses the action marker sequence, identifying the action type, target object index, and possible parameter placeholder information, converting the original marker sequence into a series of structured action instruction descriptions. For different periods or projects, the action marker vocabulary can be flexibly configured according to the specific harvesting process, for example, adding markers such as "detect successful grasp," "replan path," and "perform obstacle avoidance" to meet different process requirements.

[0102] After completing the motion tag parsing, it is necessary to map each high-level motion primitive into low-level control commands that can be executed by the robotic arm control system. In this embodiment, the control layer can be implemented in various ways, such as a kinematics-based trajectory planning module, a motion reuse module based on a motion primitive library, or a discrete control sequence generation module based on an interpolation algorithm.

[0103] In one alternative implementation, a set of motion templates corresponding one-to-one with motion markers can be predefined, each template containing a relative description of the end effector pose change, execution order, and constraints.

[0104] For example, "move above the target" can correspond to generating a trajectory that smoothly transitions from the current end pose to a safe height above the target, with the target mushroom's three-dimensional position in the working coordinate system as a reference; "approach and contact" can correspond to slowly pressing down in the vertical direction to a preset contact depth, and limiting speed and acceleration during the approach to reduce the impact on the cap and bag.

[0105] To perform the specific calculations of the aforementioned trajectory and control commands, the control system can invoke the kinematics solving and path planning module to convert the desired end pose relative to the target into a time series of joint angles or linear axis displacements. This path planning module can be implemented based on conventional inverse kinematics solving, gradient descent, sampling planning algorithms, or planning tools provided in existing robot control frameworks, such as point-to-point trajectory generation based on interpolation, spline trajectory interpolation, etc.

[0106] In some implementations, in order to improve the safety of movement in narrow gaps between cultivation racks, path planning can explicitly consider the geometric models or simplified collision bodies of the cultivation rack columns, beams and mushroom bag clusters, perform collision detection and safety distance constraints on the planned trajectory, and, if necessary, make appropriate adjustments to certain high-risk actions in the action markers or insert transition actions.

[0107] During the decoding of the action marker sequence, the control system can also adjust the control commands online based on sensor feedback.

[0108] For a complete sequence of actions consisting of multiple action markers, the control system can generate and issue control commands sequentially according to the order of the markers, or merge or rearrange some actions while ensuring safety, in order to improve execution efficiency.

[0109] For example, when multiple targets are located in the same cultivation layer or adjacent rows, "retreat to a safe position away from the cultivation rack" can be replaced with "move above the next target" during decoding, thereby reducing unnecessary back-and-forth movements and improving the cycle time of the entire harvesting task. In scenarios involving multi-robotic arm collaboration or robotic arm and mobile chassis collaboration, the action marking sequence can be extended to include multiple subject control fields such as "robotic arm number" and "chassis movement command," which are then coordinated by the advanced scheduling module based on the marking content.

[0110] The harvesting control commands generated through the above decoding and planning process can be issued to the actuators in various forms, including target trajectories, speed and acceleration commands in joint space or Cartesian space, gripper opening and closing control signals, and control signals from auxiliary equipment (such as lighting, spraying, and conveying devices). These control commands can be transmitted to servo drives, PLCs, or motion control cards via fieldbus, Ethernet, or other industrial communication protocols, where the underlying control loop completes position, speed, or force control closed-loop operations.

[0111] This embodiment does not limit the specific hardware architecture. As long as the corresponding picking control command can be obtained by decoding the action tag sequence and the picking equipment can be driven to complete picking operations such as grasping, transporting and placing in the button mushroom cultivation environment, it can be considered to fall within the protection scope of this invention.

[0112] In this way, the discrete action tag sequence output by the VLA model is transformed into continuous control commands applicable to specific harvesting equipment, enabling the high-level action intent generated based on unified semantic modeling to be reliably executed in real button mushroom cultivation scenarios. This allows for autonomous harvesting control in complex and dense cultivation environments while ensuring the safety of equipment and mushrooms.

[0113] Optional, see Figure 2 The flowchart of a method for obtaining a visual feature sequence provided in this application embodiment includes steps S201 to S202, wherein:

[0114] S201: Acquire multiple consecutive frames of visual observation images within a preset time window, divide the multiple frames of visual observation images into several image blocks, and combine the image blocks corresponding to the spatial positions in different time frames to form a spatiotemporal unit.

[0115] S202: Perform linear mapping and spatiotemporal position encoding on each of the spatiotemporal units to generate a temporal visual marker sequence, which serves as the visual feature sequence.

[0116] Based on the above implementation, the encoding of visual observation data can be further designed in conjunction with the time dimension so as to explicitly express the dynamic information of the mushroom cultivation scene changing over time in the visual feature sequence, thereby providing richer temporal clues for subsequent multi-scale feature construction and action generation.

[0117] In this optional implementation, a preset time window can be set for the target cultivation area during the robotic arm's harvesting or inspection operations. Within this time window, the control unit continuously acquires multiple frames of visual observation images according to a preset sampling frequency or based on an event-triggered mechanism. The length of the preset time window and the sampling frequency can be configured according to the robotic arm's movement speed, the rate of change in mushroom occlusion, and computational resource constraints. For example, a window length of tens to hundreds of milliseconds and a sampling number of several to more than ten frames can be used; the specific values ​​are not limited.

[0118] For each frame of image acquired within the time window, a uniform spatial partitioning strategy can be used to divide it into several image blocks. The shape of the image blocks can typically be a regular grid, or it can be a custom irregular region based on the actual cultivation rack structure, such as a rectangular region aligned with the arrangement of mushroom bags along the row and column directions. Each image block obtained can be regarded as a local receptive field at the corresponding spatial location, used to capture local texture and edge information of the cap, mushroom bag, or support at that location.

[0119] The size and number of image patches can be flexibly set according to the camera resolution, the size of the target mushroom, and the desired level of detail. For example, smaller image patches can be used at a closer view of the cultivation rack to obtain more detailed cap edge information, while larger image patches can be used to reduce the computational load when observing from a distance.

[0120] After dividing the images into frames, spatiotemporal units can be constructed based on the image blocks. Specifically, image blocks corresponding to spatial positions in different time frames can be grouped together according to their row and column positions in the image to form spatiotemporal units that unfold along the time dimension. Each spatiotemporal unit corresponds to a fixed line of sight and spatial region in the mushroom cultivation scene, and the changes in the content of that region are recorded within a time window, such as a mushroom cap gradually becoming visible from partial obscuration, or the end effector of a robotic arm entering the area from outside the frame.

[0121] Furthermore, depending on the specific implementation requirements, the construction of spatiotemporal units may also allow for the introduction of a certain neighborhood extension in the spatial dimension, so that a small number of adjacent grids together constitute a spatiotemporal unit, thereby enhancing the expressive ability of locally connected regions. This application does not limit this.

[0122] For each spatiotemporal unit, it can be converted into a vector-based symbolic representation through linear mapping.

[0123] Specifically, the image block pixels within a spatiotemporal unit or the features extracted by a lightweight convolutional network can be concatenated or aggregated in a predetermined order to obtain a fixed-dimensional intermediate representation. Then, through one or more linear transformations and nonlinear activations, this intermediate representation can be mapped into a feature vector suitable for use as a visual label.

[0124] To explicitly represent the spatiotemporal structure in subsequent encoding, this embodiment can also overlay spatiotemporal location encoding information onto each spatiotemporal unit. The spatiotemporal location encoding can include at least two types of components: one type represents the spatial location corresponding to the spatiotemporal unit, such as its row index, column index, derived polar coordinates, or correspondence with the geometric coordinates of the cultivation rack in the image; the other type represents the temporal order of each frame within the spatiotemporal unit, such as the index of the time frame, relative time offset, or normalized timestamp. The aforementioned spatial and temporal location information can be obtained in vector form by looking up a table, and then added to or concatenated with the linearly mapped feature vector to obtain the spatiotemporal location encoding result containing both spatial and temporal location information.

[0125] After completing the linear mapping and spatiotemporal location encoding of all spatiotemporal units, the feature vectors corresponding to each spatiotemporal unit can be arranged into a temporal visual label sequence according to certain arrangement rules.

[0126] The arrangement rules can be flexibly determined according to task requirements. For example, they can be arranged sequentially according to the time dimension, forming multiple time segments, and then the time segments can be concatenated according to time indices. Alternatively, they can be arranged adjacently according to the spatial location, so as to highlight the change trend of a local area over time. This implementation does not limit the specific sequence organization method, as long as the final temporal visual marker sequence can reflect the main spatiotemporal change patterns of the target cultivation area within the preset time window.

[0127] In a typical implementation, the generated temporal visual label sequence can be directly used as a specific form of the aforementioned visual feature sequence for subsequent multi-scale visual feature construction and object query-driven cross-attention decoding. This spatiotemporal unit-based encoding strategy enables the model to automatically learn and utilize dynamic features such as occlusion changes during robotic arm approach, subtle deformations exhibited at different mushroom growth stages, and light fluctuations across multiple frames. This provides the VLA model with richer and more discriminative key visual information input than single-frame static encoding in complex and dense unstructured Agaricus bisporus cultivation scenarios.

[0128] Based on the above-mentioned spatiotemporal unit-based encoding, in an optional implementation, when performing linear mapping and spatiotemporal position encoding on each spatiotemporal unit, the geometric layout information of the button mushroom cultivation rack can also be combined to add a structured position identifier to each spatiotemporal unit, so as to explicitly distinguish units of different cultivation levels and row and column positions in the temporal visual label sequence, thereby reducing the interference of multi-layer and multi-row repetitive structures on the model's discrimination ability.

[0129] Specifically, before actual deployment, the mushroom cultivation rack can be calibrated once or periodically to obtain its geometric layout information within the robot's workspace. This geometric layout information can include the total number of layers, the height of each layer, the number and spacing of cultivation rows within each layer, the number and spacing of columns of mushroom bags or cultivation troughs within each row, and the reference positions of supporting columns and beams. Calibration can be achieved using conventional industrial camera calibration combined with manual measurement, or with the assistance of automated calibration processes with markers, laser rangefinders, or 3D sensors, as long as a mapping relationship can be established between the image plane coordinates and the cultivation rack's layer and row / column structure coordinates.

[0130] Given the known geometric layout of the cultivation rack and the camera's extrinsic and intrinsic parameters, this implementation method can determine the corresponding cultivation rack structure position for each spatiotemporal unit.

[0131] Specifically, based on the center coordinates or coverage area of ​​a spatiotemporal unit in an image, the cultivation level (e.g., layer 1, layer 2, etc.), the cultivation row (e.g., front row, middle row, back row, etc., or numbered by row number), and the approximate column index within that row (e.g., column 1, column 2, etc.) can be inferred through back projection, geometric transformation, or by consulting a pre-established mapping table. For cameras with different viewing angles, corresponding mapping relationships can be established separately. When using a moving camera or a robotic arm end-effector camera, the mapping process can be dynamically corrected by combining the robotic arm joint angles and end-effector pose.

[0132] After obtaining the hierarchical index, row index, and column index, these discrete structural location information can be converted into vector representations suitable for neural network processing.

[0133] To this end, this implementation can configure an embedding mapping module for each index dimension, mapping the hierarchical number, row number, and column number to fixed-length vector representations, such as hierarchical embedding vectors, row embedding vectors, and column embedding vectors. The mapping method can be to read the corresponding vector from a pre-trained or randomly initialized embedding matrix using a lookup table, or it can be calculated based on the index value using a small feedforward network.

[0134] In some implementations, the hierarchy and row / column indexes can be normalized or encoded, for example, by using segmented encoding or group encoding, to enhance the adaptability of structural position encoding to cultivation racks of different scales.

[0135] Subsequently, the above structural positional coding can be combined with the basic spatial positional coding for the image block to form the final spatiotemporal positional coding for that spatiotemporal unit.

[0136] One common approach is to add the structural position code and the basic spatial position code element by element, so that the encoding vector simultaneously reflects the pixel position of the spatiotemporal unit on the image plane and its hierarchy, row and column identity in the cultivation rack structure; another approach is to concatenate the structural position code and the basic position code in the channel dimension, and then unify the dimensions through linear transformation.

[0137] The specific combination method can be flexibly adjusted according to the downstream network structure and performance requirements, and this application does not impose any restrictions on it.

[0138] This enables the downstream multi-scale feature construction and object query decoding modules to distinguish different levels and rows of units in the same image view when faced with multi-layered, multi-rowed, and highly repetitive cultivation rack textures. This helps to avoid confusing mushroom targets from the upper layer or adjacent rows with the current target, thereby improving the ability to extract key visual information and the stability of target localization in complex and dense cultivation rack structures.

[0139] In an alternative implementation, to make the temporal visual marker sequence more sensitive to subtle deformations of the cap edge of Agaricus bisporus and more robust to slow changes and noise in the background region, the cap edge region and background structure region in the cultivation image can be pre-calibrated before performing linear mapping and spatiotemporal location encoding on each spatiotemporal unit, and the encoding weight of different spatiotemporal units in the time dimension can be distinguished based on the calibration results.

[0140] Specifically, the cap edge region in each frame of visually observed images can be labeled based on images of Agaricus bisporus cultivation. Labeling of the cap edge region can be achieved using various technical approaches. In one implementation, a segmentation or detection network trained on Agaricus bisporus can be used to infer the probability map, segmentation mask, or target box representing the cap region for each frame of the image. Then, by dilating and thinning the mask boundary or box edge, the pixel band where the cap edge is located can be approximated. In another implementation, traditional image processing methods can be used to perform edge detection, threshold segmentation, and region growing on the image. Combining the approximately circular or oval outer contour of the cap and its grayscale distribution characteristics, edge regions that better match the geometric shape of Agaricus bisporus can be selected from the candidate regions.

[0141] For scenes with significant lighting variations or complex backgrounds, a combination of learning-based and rule-based methods can be used. For example, a learning model can be used to first provide a rough outline of the cap region, and then edge detection and geometric constraints can be used to fine-tune the edge positions. This application does not limit the specific edge labeling algorithm, as long as it can provide a sufficient division on the image plane to distinguish between the cap edge region and non-edge regions.

[0142] After completing the labeling of the cap edge region, the background structure region can be further defined. The background structure region can include the surface area of ​​the mushroom bag, the exposed area of ​​the cultivation substrate, the area of ​​the cultivation rack support, and other areas that do not contain the cap of Agaricus bisporus. For example, by performing an inverse masking operation on the segmentation network output, the non-cap area can be divided into "mushroom bag area," "support area," and "cultivation substrate area" according to its geometric location. Alternatively, based on the pre-labeled geometric layout of the cultivation rack and the mushroom bag placement rules, a raster mapping method can be used to determine whether certain areas mainly correspond to mushroom bags or supports.

[0143] For areas that are difficult to classify clearly, such as areas with strong reflections or severe shadows, the system strategy can be used to temporarily exclude them from the high-sensitivity or low-sensitivity areas, and they can be processed by subsequent modules with normal weights.

[0144] Based on the known cap edge region and background structure region in each frame of the image, the corresponding labels can be mapped onto the aforementioned spatiotemporal units.

[0145] Specifically, for a certain spatiotemporal unit, if the overlap ratio between the image patch it covers and the edge region of the cap in a certain frame exceeds a preset threshold, the spatiotemporal unit can be marked as a high-sensitivity region; if the coverage area of ​​the spatiotemporal unit mainly falls in the bag area, the support area, and the cultivation substrate area that does not contain the cap of Agaricus bisporus and has limited overlap with the edge region of the cap, it can be marked as a low-sensitivity region.

[0146] For spatiotemporal units that simultaneously cover both the cap edge and the background, strategic judgments can be made based on the overlap ratio, edge confidence, or contextual information. If necessary, these areas can be marked as highly sensitive regions to avoid missing crucial deformation information. This implementation does not limit specific thresholds or judgment logic; developers can adjust relevant parameters based on actual image resolution, cap size, and noise levels.

[0147] After marking high-sensitivity and low-sensitivity regions, a differentiated temporal weighting process can be introduced when performing spatiotemporal location encoding on spatiotemporal units. To this end, this embodiment can assign a first temporal weight to spatiotemporal units in high-sensitivity regions and a second temporal weight, different from the first, to spatiotemporal units in low-sensitivity regions. The first and second temporal weights can be embodied in various forms during the encoding process. For example, during superimposed temporal location encoding, high-sensitivity spatiotemporal units can be given a higher temporal resolution or a larger temporal step weight, making subtle changes in the cap edges in recent frames more significant in the feature space. For low-sensitivity spatiotemporal units, a stronger temporal smoothing or a smaller temporal variation weight can be used to relatively weaken the impact of slowly changing bag textures, scaffold shadows, or slight illumination fluctuations on the final temporal visual label sequence. In some embodiments, the first temporal weight can correspond to a smaller temporal decay rate, allowing the deformation trajectory of the high-sensitivity region to be preserved over a period of time; while the second temporal weight can correspond to a larger temporal decay rate, allowing short-term fluctuations in the background region to be smoothed out more quickly.

[0148] The specific weighting method and its numerical relationship can be flexibly configured according to the target sensitivity and noise level of the system, and this application does not limit it.

[0149] In this way, encoding the cap edge region and background structure region with different time weights in the same temporal visual label sequence allows for a more complete expression of the changes in key visual information such as the local deformation of the cap edge of Agaricus bisporus, the degree of opening, and the occlusion relationship with surrounding strains in the temporal dimension. Meanwhile, factors with less impact on harvesting decisions, such as the slow changes in the texture of the mushroom bag, the slight brightness fluctuations of the cultivation substrate, and the reflection of the support, are relatively suppressed. This improves the model's ability to extract key visual information and the robustness of action generation in complex and dense cultivation environments during the subsequent multi-scale feature construction and object-level feature decoding stages.

[0150] Based on the above implementation method, in an optional implementation method, when calibrating the cap edge region in each frame of visual observation image, a learning-based segmentation / detection network and a traditional image processing method based on geometric and grayscale features can be combined. The two methods can be used alone or in combination to improve the adaptability under different lighting conditions and cultivation environments.

[0151] In one implementation, a segmentation network or detection network trained on Agaricus bisporus can be used to reason about visually observed images to obtain a segmentation mask or target box characterizing the cap region.

[0152] To this end, a certain number of images of button mushroom cultivation scenes can be collected in advance, covering different cultivation levels, row and column positions, lighting conditions and growth stages. The cap areas can be labeled manually or by semi-automatic labeling tools. The labeling format can be pixel-level area masking, or it can be a rectangular box, a polygonal box or a combination of key points.

[0153] Based on the above data, a network architecture suitable for the target scale and deployment computing power can be selected for training, such as an encoder-decoder segmentation network, an instance segmentation network, or a lightweight detection network. During training, loss functions such as cross-entropy, boundary consistency, or region overlap can be used to iteratively optimize the network. During the deployment phase, the network can be pruned, quantized, or distilled according to the inference latency requirements of the actual pipeline.

[0154] During the inference phase, the control unit inputs the preprocessed visual observation image into the segmentation network or detection network to obtain a segmentation mask or bounding box for the cap region in each frame. For the segmentation mask, a simple thresholding process can be used to convert the probability map into a binary mask, and operations such as opening / closing, hole filling, or small region filtering can be performed on the mask to remove isolated noise points or incomplete regions. For the bounding box output by the detection network, local thresholding or edge detection can be performed again inside the box to obtain a region that more closely matches the actual cap contour.

[0155] Finally, the mask processed above is regarded as the cap region, and by performing operations such as dilation, erosion and contour extraction on the mask boundary, a pixel band approximating the cap edge is obtained. The image area covered by this pixel band is regarded as the cap edge region, which is used for subsequent correspondence with spatiotemporal units and high-sensitivity region marking.

[0156] In another implementation, edge detection and region growing can be performed on the visually observed image based on the geometric shape features and gray-scale distribution features of the cap's outer contour, without relying on or with a weak reliance on the learning network, to extract the region where the cap's edge is located.

[0157] Specifically, the image can first be smoothed, denoised, and contrast-enhanced. Then, common edge detection operators can be used to extract initial edges from the image, such as gradient operators or Canny operators, to obtain a set of candidate edge points. Subsequently, the geometric characteristics of the approximately circular or oblate cap of Agaricus bisporus can be taken into account to cluster or trace the edge points, identify edge segments with relatively rounded and closed contours and areas within a reasonable range, and remove edges that are too thin, irregular, or obviously correspond to the edges of supports or bags.

[0158] For the selected candidate cap outlines, the gray distribution or color differences between the interior and surrounding areas can be further analyzed. For example, the difference in brightness and texture between the cap area and the surface of the covering soil or mushroom bag can be analyzed to improve the accuracy of identifying the real cap area.

[0159] Based on the geometric and grayscale feature screening described above, the candidate cap outline can be used as the seed region. Region growing or merging strategies can be employed to expand from inside and outside the outline to the entire cap region. Pixel bands within a certain width near the outline are uniformly labeled and considered as the area where the cap edge is located. For complex lighting or local occlusion, multi-scale edge detection, region growing with different threshold combinations, or morphological reconstruction can be used to supplement and correct edges that are difficult to identify.

[0160] Learning-based methods and geometric rule-based methods can also be used in combination. For example, a detection network can be used to provide the approximate location of the cap, and then edge detection and region growth can be used to refine the edge range, so as to improve the accuracy and robustness of cap edge region labeling while ensuring real-time performance.

[0161] By using any one or more of the above methods, a relatively reliable cap edge region can be obtained in each frame of visual observation image, providing a clear spatial reference for subsequent high-sensitivity region labeling and time-weighted differential coding based on spatiotemporal units, thereby better highlighting cap edge deformation information directly related to harvesting decisions in complex and dense cultivation scenarios.

[0162] In one alternative implementation, when constructing a multi-scale visual feature map based on the visual feature sequence, local geometric information can be further used to weight and fuse features at different scales in order to highlight key structures such as the edge of the mushroom cap in the multi-scale space, and suppress high-frequency noise such as the texture of the mushroom bag, the edge of the support, and reflective points, which have a smaller impact on the picking decision.

[0163] Optional, see Figure 3 The flowchart of a method for constructing multi-scale visual feature maps provided in this application embodiment includes steps S301 to S303, wherein:

[0164] S301: Obtain at least two visual feature maps with different spatial resolutions based on the visual feature sequence;

[0165] S302: Calculate the geometric response index for each of the visual feature maps in a preset local neighborhood, wherein the geometric response index reflects the edge curvature or second-order gradient intensity in the local neighborhood.

[0166] S303: For the same spatial location, determine the geometric saliency weight based on the geometric response index of the spatial location at different spatial resolutions, and perform weighted fusion of each visual feature map at the corresponding spatial location according to the geometric saliency weight to obtain the multi-scale visual feature map.

[0167] Specifically, at least two visual feature maps with different spatial resolutions can be obtained first based on the visual feature sequence. The visual feature sequence can be derived from the output of the aforementioned visual coding network, and multi-scale feature maps can be recovered or derived in different ways: for example, feature maps from different levels in the coding network can be directly selected as visual feature maps of different scales; alternatively, a low-resolution feature map can be constructed on a certain basic feature map through downsampling, and a high-resolution feature map can be constructed by upsampling and fusing with shallow features, thereby aggregating the global context at a lower resolution and preserving local details such as the cap edge at a higher resolution.

[0168] Feature maps of different scales can vary in spatial size, receptive field size, and semantic abstraction level. The specific number of layers and resolution combinations can be flexibly configured according to computing power, camera resolution, and target size distribution.

[0169] After obtaining multi-scale visual feature maps, geometric response indices can be calculated for each scale feature map within a preset local neighborhood. These indices are used to measure the edge curvature or second-order gradient intensity within that neighborhood. The preset local neighborhood can be a window region centered on the current spatial location and covering a number of pixels or feature units. Its size can be set according to the image resolution and target size.

[0170] The calculation of geometric response metrics can be implemented in various ways. For example, gradient components can be obtained by performing first-order difference convolutions on the feature map in the horizontal and vertical directions within the local neighborhood. Then, a local tensor reflecting the edge direction and intensity distribution can be constructed based on these gradient components, and the gradient intensity of the principal direction can be extracted from it. Alternatively, second-order differences, filter kernels with second-order sensitivity, or multi-scale difference operators can be used to measure feature changes within the local neighborhood to characterize the curvature of the edge and the intensity of high-frequency changes. For different scenarios and implementation requirements, different forms of geometric response metrics, such as structural tensors, approximate curvature responses, and second-order gradient magnitudes, can be selected. This application does not limit itself to a specific algorithm.

[0171] After completing the geometric response calculation on feature maps at various scales, a correspondence can be established between geometric response indices at different spatial resolutions for the same spatial location.

[0172] In general, feature maps at different scales can be mapped or aligned to a unified reference coordinate system to ensure that features from different resolutions have comparable geometric response values ​​at the same physical location.

[0173] For example, for a location in a low-resolution feature map, it can be mapped to several adjacent locations in a high-resolution feature map by scaling or interpolation, and the geometric responses of these locations can be summarized; alternatively, alignment can be performed directly using a pre-established scale mapping relationship. For each reference spatial location, geometric response indices at various scales can be collected to form a cross-scale geometric response set.

[0174] Based on the aforementioned set of cross-scale geometric responses, a geometric saliency weight can be determined for each spatial location. The geometric saliency weight is used to reflect whether the location consistently exhibits strong edge or high curvature characteristics at different scales, thereby distinguishing stable cap edges from isolated noise responses.

[0175] For example, the geometric response at each scale can be normalized, and weights can be calculated based on the magnitude, consistency, or simultaneous occurrence of the response at multiple scales. For instance, if a location exhibits a high and consistent edge response across multiple scales, it can be assigned a larger geometric saliency weight. Conversely, if a location shows a sharp response only at a single scale and a flat response at other scales, it is more likely to correspond to a highlight, noise texture, or fragmented background structure, and its corresponding geometric saliency weight can be set smaller. The specific weighting function can be designed empirically or through a data-driven approach; this application does not limit the form of weight calculation.

[0176] After determining the geometric saliency weights, the visual feature maps at each scale can be weighted and fused at their corresponding spatial locations based on these weights to obtain a multi-scale visual feature map. The fusion method can be to sum the feature vectors from different scales according to their geometric saliency weights at the same spatial location, or to first concatenate the multi-scale features along the channel dimension and then compress them into a unified dimension through a weighted linear transformation.

[0177] For locations with higher geometric saliency weights, the fused features will retain more edge and shape information that is stable across multiple scales, such as the cap outline and contact boundaries with surrounding strains. For locations with lower weights, the fusion result will relatively suppress high-frequency noise that is prominent only at a single scale, thereby reducing the interference of bag texture, support edges, or reflective points on subsequent object query decoding in the feature space. The multi-scale visual feature map generated by fusion can maintain the same spatial resolution as a reference scale, or it can be retained in a multi-scale form for subsequent object query vectors to selectively perform cross-attention calculations at different scales.

[0178] In this way, structural information at different scales is integrated in a targeted manner at the visual representation level, so that geometric structures such as the edge of the mushroom cap, which are highly related to picking decisions, are more prominently represented in the fused feature map, while the perturbation of features by background texture and local noise is effectively suppressed, thus providing clearer and more reliable key visual information for subsequent object-level visual label generation and action generation of the VLA model.

[0179] In one optional implementation, when calculating the geometric response index for each visual feature map within a preset local neighborhood, the hierarchical index, row index, and column index of the spatiotemporal unit can be further combined to align the local neighborhood with the structural grid of the button mushroom cultivation rack. This allows for explicit differentiation of features at different levels, rows, and columns during the geometric analysis stage, reducing the interference of multi-layered and multi-row repetitive structures on the geometric response.

[0180] Specifically, based on the aforementioned allocation of hierarchical, row, and column indices to spatiotemporal units according to the geometric layout of the cultivation rack, this structural location information can be mapped to the spatial coordinates of visual feature maps at various scales. For any spatial location in the visual feature map, the corresponding cultivation level, cultivation row, and approximate column number can be determined based on the structural index of its source image patch or corresponding spatiotemporal unit. For stationary cameras, this mapping relationship can be fixed after a single calibration; for end-effector cameras that move with the robotic arm, the correspondence between the structural index and the feature map coordinates can be dynamically updated by combining the current end-effector pose and a pre-established scene model.

[0181] After determining the structural index, a local neighborhood aligned with the mushroom cultivation rack structure can be defined for each spatial location in the visual feature map. Unlike conventional square neighborhoods centered on a pixel grid, the local neighborhoods in this embodiment preferentially extend along the row and column directions of the cultivation rack and are limited to positions within the same and / or adjacent layers.

[0182] For example, for a specific feature location, a strip-shaped neighborhood stretched along the row direction can be constructed within the current row and one or two adjacent rows of the same level, or a rectangular neighborhood can be constructed within several adjacent columns of the same level, thus making the geometric analysis closer to the actual layout of the cultivation rack. When it is necessary to consider the occlusion relationship between adjacent levels, the neighborhood range can also be appropriately extended at the corresponding row and column positions of adjacent upper and lower levels, but generally not across multiple levels, to avoid erroneously merging targets that are far apart. This application does not limit the specific shape and size of the local neighborhood; developers can adjust the neighborhood morphology according to the cultivation rack structure, target density, and computing resources.

[0183] After obtaining the structurally aligned local neighborhood, first-order derivative convolutions in the horizontal and vertical directions can be performed on the feature values ​​within this neighborhood on each visual feature map to obtain local gradient components. These gradient components reflect the trend and intensity of feature changes along the row and column directions, and can be implemented using simple difference operators, edge detection operators, or predefined convolution kernels. Based on this, a local structure tensor reflecting the distribution of edge directions and intensity can be constructed using the local gradient components. The local structure tensor can be understood as a statistical description of the combination relationship of gradient components, comprehensively characterizing the dominant edge directions and overall gradient energy distribution within the local neighborhood, rather than being limited to the gradient magnitude at a single point.

[0184] Subsequently, feature analysis can be performed on the local structure tensor to obtain the gradient intensity corresponding to the principal direction, which is then used as a geometric response index. The gradient intensity of the principal direction can be understood as the overall change intensity along the most significant edge direction within the local neighborhood, reflecting geometric information such as "whether there is a stable edge" and "whether the edge is continuous along a certain direction".

[0185] In the cultivation of button mushrooms, the edge of the cap usually appears as a relatively stable curve with a relatively continuous outline in a local area, while the edges of the support or bag have different directional patterns in the structural coordinate system. By calculating the gradient intensity of the main direction in the local neighborhood aligned with the cultivation rack structure, the edge response related to the cap outline and the local high-frequency changes caused by the background structure or noise can be distinguished more accurately.

[0186] It should be noted that although this embodiment uses the construction of a local structural tensor through gradient components and the calculation of the gradient intensity of the principal direction within the structural alignment neighborhood as an example to illustrate the calculation method of the geometric response index, in practical applications, other indices that can reflect the curvature of the local edge or the degree of high-frequency changes can be used to replace or supplement the above method. For example, the geometric response index can be defined using the second-order difference response in different directions, the response intensity of the multi-scale edge filter, or other statistics based on directional consistency. This application does not limit the specific calculation formula of the geometric response index. As long as the selected index can effectively characterize the geometric features related to key structures such as the edge of the Agaricus bisporus cap within the preset local neighborhood and is used to determine the geometric saliency weight in the multi-scale feature fusion stage, it can be considered to fall within the protection scope of this embodiment.

[0187] In one alternative implementation, when generating the object-level visual tag sequence, the structural information of the mushroom cultivation rack can be explicitly utilized. The interaction range between the object query vector and the multi-scale visual feature map is constrained by structural anchor points and attention mask regions, so that the object-level feature decoding process corresponds one-to-one with the actual cultivation level and row and column units.

[0188] Specifically, after assigning hierarchical, row, and column indices to the spatiotemporal units, these structural indices can be used as discrete identifiers for the structural units of the cultivation rack. Based on the actual layout of the cultivation rack, the combination of "hierarchical index - row index - column index" can be considered a structural anchor point, used to identify the spatial unit corresponding to a certain layer, row, or column on the cultivation rack. For example, in a scenario with 4 cultivation racks, 6 rows per layer, and 10 mushroom bag positions per row, (layer=2, row=3, column=5) can be considered the structural anchor point for the 5th mushroom bag position in the 3rd row of the 2nd layer. For each structural anchor point, its discrete index can be converted into a fixed-length vector representation through the embedding mapping module, forming a structural anchor point vector. For unified processing, embedding matrices can be set for the hierarchical index, row index, and column index respectively, and the corresponding embedding vectors of the three can be summed or concatenated before undergoing a linear transformation to obtain the final structural anchor point vector.

[0189] After obtaining the structural anchor vector, it can be combined with a pre-defined learnable query basis vector to generate an object query vector containing the structural location information of the cultivation rack. The learnable query basis vector can be understood as a general query template without specific location semantics, used to carry priors about the target category, shape, and context, while the structural anchor vector injects location information "where on the cultivation rack".

[0190] For example, the combination of the two can be vector summation or concatenation followed by transformation through a linear layer. For instance, an object query vector can be generated for each possible structural anchor point to specifically focus on potential button mushroom targets near that level and row / column unit; alternatively, multiple adjacent rows / columns can be grouped into a single object query vector based on the size of a local area of ​​the cultivation rack, providing a unified model for targets within a small area. This application does not limit the number or specific structure of the object query vectors, as long as they contain the level and row / column position information of the corresponding cultivation rack structural unit during initialization.

[0191] While constructing the object query vector, an attention mask region aligned with the cultivation rack structure can be defined for each object query based on the corresponding spatial position of the structural anchor point in the multi-scale visual feature map. To this end, the aforementioned mapping relationship from image patches to cultivation rack coordinates can be used to convert the "layer, row, column" indexes into spatial regions on the visual feature maps at each scale. For example, for a given structural anchor point (layer=2, row=3, column=5), the feature position corresponding to the vicinity of the third row and fifth column of the second layer can be marked in the multi-scale visual feature map, and a finite row and column neighborhood around this position can be selected as the attention mask region. This neighborhood can be limited to the current row and its adjacent row or two rows at the same level, or it can include the current column and several columns before and after it. The specific neighborhood range can be set according to the row spacing, column spacing of the cultivation rack, and the desired search area size. For example, in the row direction, it can be set to the current row and the row before and after, and in the column direction, it can be set to the current column and the column to the left and right, thus forming a 3×3 structural unit neighborhood; when it is necessary to consider the occlusion effect of the upper and lower layers, the corresponding row and column positions of the adjacent layers can also be included in the mask area, but it generally does not cross multiple non-adjacent layers.

[0192] After determining the attention mask region, cross-attention computation can be performed within each mask region, using the corresponding object query vector as the query vector and the feature vectors in the multi-scale visual feature map that fall within that mask region as the key and value vectors. In other words, the object query only interacts with features near the corresponding structural anchor point, without globally focusing on the entire feature map, thus explicitly reflecting the structural prior of "the target near a certain cultivation grid" during attention computation.

[0193] The cross-attention module can employ a multi-head attention structure, enabling object queries to select useful information from features of different scales and directions within the same neighborhood. When needed, short self-attention layers can be inserted between the cross-attention layers to allow object queries at different structural anchors to exchange information about occlusion, cluster relationships, etc. After completing several layers of cross-attention updates, the final object query vector corresponding to each structural anchor can be regarded as an object-level visual label, recording the presence, morphological features, and local context of targets within or near that cultivation structural unit.

[0194] Arrange the object-level visual tags corresponding to all structural anchor points in a predetermined order to obtain the object-level visual tag sequence, which is then aligned with the picking task description information in a unified semantic space and drives the VLA model to generate an action tag sequence.

[0195] In one alternative implementation, the process of obtaining the joint semantic sequence can be combined with a preset semantic dictionary and a sparse representation mechanism to align and constrain the picking task description information and object-level visual tags at the semantic level, so that the joint semantic sequence has explicitly encoded the structural semantics and target attribute semantics related to the button mushroom cultivation environment before entering the VLA model backbone.

[0196] Specifically, a semantic dictionary tailored to the cultivation scenario of button mushrooms can be pre-constructed. This semantic dictionary can contain several semantic atoms, each representing a basic semantic dimension or typical state in the cultivation scenario. For example, semantic atoms used to characterize the cultivation rack level (such as "first layer", "second layer", etc.), semantic atoms used to characterize the range of cultivation row and column positions (such as "front row", "middle row", "near the aisle side", "near the support side", etc.), semantic atoms used to characterize the maturity range of the target mushroom (such as "nearly mature", "suitable for harvesting", "overripe", etc.), semantic atoms used to characterize the size range of the target mushroom (such as "small", "medium", "large", etc.), and semantic atoms used to characterize the relative positional relationship between the target mushroom and the cultivation rack support structure (such as "adjacent to the column", "near the beam", "located in the middle of the row", etc.).

[0197] The specific division of semantic atoms can be flexibly set according to production management requirements and harvesting strategies. For example, when it is necessary to distinguish the maturity stage more finely, more maturity atoms can be introduced. When the safety risks of "being close to the support" are more sensitive, the number of atoms with relative position to the support can be refined.

[0198] After obtaining the semantic dictionary, a unified semantic mapping module can be set up to map text tag sequences and object-level visual tag sequences to a common semantic representation space. This mapping module can include its own embedding layers and several lightweight transformation networks to project the original vectors of text tags and visual tags into a common semantic representation space with consistent dimensions, so as to facilitate subsequent semantic alignment and comparison.

[0199] For example, a text tag sequence can be processed through a text embedding layer and 1-2 layers of feedforward network or attention layer to obtain a first semantic representation; object-level visual tags can be processed through linear transformation or a small multilayer perceptron to obtain a second semantic representation. Here, both the "first semantic representation" and the "second semantic representation" should be located in the same semantic vector space so that subsequent sparse decomposition can be performed around the same semantic dictionary.

[0200] Within a common semantic representation space, sparse coefficient vectors with respect to the semantic dictionary can be obtained for the first and second semantic representations respectively, based on a pre-defined semantic dictionary.

[0201] In some implementations, the semantic dictionary can be viewed as a set of semantic basis vectors / semantic basis consisting of multiple semantic atom vectors, and an attempt can be made to approximate the current semantic vector using a linear combination of a small number of semantic atoms. For the first semantic representation on the text side, a set of first semantic sparsity coefficients can be solved, such that the first semantic representation can be characterized by a linear combination of the semantic dictionary; similarly, a set of second semantic sparsity coefficients can be solved for the second semantic representation corresponding to each object-level visual tag.

[0202] Sparsity can be achieved by introducing sparse constraints, regularization terms, or by directly using pre-trained sparse coding networks during the solution process. The goal is to enable each semantic representation to activate only a few semantic atoms in the semantic dictionary, thereby clearly reflecting the main values ​​of the text instruction or visual object in semantic dimensions such as "hierarchy, row and column position, maturity, size, and relative positional relationship with the supporting structure".

[0203] For those skilled in the art, appropriate sparse solution strategies can be selected based on the actual scenario and computing power conditions, without being limited to a specific algorithm.

[0204] After obtaining the first semantic sparsity coefficient on the text side and the second semantic sparsity coefficient on each object side, coefficient terms of corresponding semantic atoms closely related to the mushroom cultivation environment can be selected and used as the basis for subsequent object constraints.

[0205] For example, coefficient terms corresponding to semantic atoms used to characterize the cultivation rack hierarchy, row and column positions, target mushroom maturity, target mushroom size, and the relative positional relationship between the target mushroom and the cultivation rack support structure can be extracted from the first and second semantic sparse coefficients, and these coefficient terms are determined as semantic constraint coefficients. These semantic constraint coefficients are used to explicitly carry structured semantic information related to the picking decision in the joint semantic sequence, enabling subsequent action generation to simultaneously consider task semantic constraints and object-level visual information.

[0206] In some implementations, the coefficient terms corresponding to the semantic atoms in the first semantic sparse coefficient are used to characterize the scope of attention or preference of the picking task description information in the semantic dimension; the coefficient terms corresponding to the semantic atoms in the second semantic sparse coefficient are used to characterize the inference result or matching degree of the object-level visual tags in the semantic dimension. Based on the semantic constraint coefficients, object-level visual tags can be enhanced, filtered, or sorted to preferentially retain object-level visual tags that are more consistent with the picking task description information in dimensions such as cultivation rack level, row and column position, maturity, size, and relative positional relationship.

[0207] In some implementations, the "relative positional relationship between the target mushroom and the support structure of the cultivation rack" can be determined by the spatial position corresponding to the pre-calibrated geometric layout of the cultivation rack and the object-level visual markers. For example, one or more semantic atoms of relative positional relationships such as "close to the column," "close to the beam," "close to the side of the support," and "located in the middle of the row" can be determined based on the relative distance, relative orientation, or preset neighborhood that the object's center point or object mask falls into with the support structure such as the cultivation rack's columns and beams in the image or working coordinate system. The size of the preset neighborhood can be configured according to factors such as the camera's field of view, the scale of the cultivation rack, and the safety gap for picking.

[0208] In one implementation, the coefficients of the corresponding semantic atoms on the text side and the object side can be compared or combined. For example, the coefficients of the corresponding components on the text side can be directly appended to the object tag, or the similarity or difference between the text side and the object side on these components can be calculated, and the result can be used as an additional constraint feature of the object tag to indicate the degree of matching between the object and the semantic requirements of the current task.

[0209] Finally, semantic constraint coefficients can be appended to the corresponding object-level visual tags to form an enhanced object-level semantic tag sequence. This enhanced sequence is then arranged with the original text tag sequence in a predetermined order to obtain a joint semantic sequence. The object-level visual tags and their semantic constraint coefficients can be combined through vector concatenation or additive fusion, ensuring that each object-level tag, before entering the VLA model backbone, not only contains geometric and appearance features obtained from scene perception but also carries structural hierarchy, row and column positions, attribute state information extracted from the semantic dictionary, and the degree of semantic association with the current picking task.

[0210] The concatenation order of the joint semantic sequence can be text tags first and object-level semantic tags second, or it can be adjusted appropriately according to the model structure and task requirements, as long as the VLA model backbone can simultaneously access the task description and the object-level visual information enhanced by semantic constraints.

[0211] In this way, not only is the target of button mushroom modeled at the object level in geometric space, but semantic atoms related to the cultivation rack structure and mushroom attributes are also explicitly introduced in semantic space. This allows the joint semantic sequence to complete a sparse correspondence and screening of task semantics and object attributes before entering the VLA model. In this way, the model’s attention to key targets and important constraints is improved in the complex and dense button mushroom cultivation environment, providing a clearer semantic basis for the generation of subsequent action label sequences.

[0212] Based on the same inventive concept, this application also provides a VLA-based autonomous harvesting system for button mushrooms, corresponding to a VLA-based autonomous harvesting method for button mushrooms. Since the principle of the system in this application is similar to the VLA-based autonomous harvesting method for button mushrooms described above, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be described again.

[0213] Reference Figure 4 The diagram shown is a schematic of a VLA-based autonomous harvesting system for button mushrooms provided in an embodiment of this application. The system includes:

[0214] The acquisition module 10 is used to acquire the picking task description information and visual observation data of the target scene; and to encode the visual observation data to obtain a visual feature sequence.

[0215] The first generation module 20 constructs a multi-scale visual feature map based on the visual feature sequence, and generates an object-level visual tag sequence by calculating the cross attention between the object query vector and the multi-scale visual feature map.

[0216] The second generation module 30 is used to encode the picking task description information into a text tag sequence, and map the text tag sequence and the object-level visual tag sequence to a unified semantic space and concatenate them to obtain a joint semantic sequence; input the joint semantic sequence into the pre-trained VLA model backbone, and generate an action tag sequence through autoregressive prediction;

[0217] The control module 40 is used to decode the action marker sequence to obtain the picking control command, and control the picking equipment to perform the button mushroom picking action based on the picking control command.

[0218] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A VLA-based autonomous Agaricus bisporus picking method, characterized in that, The method comprises the following steps: obtaining picking task description information and visual observation data of a target scene; encoding the visual observation data to obtain a visual feature sequence; constructing a multi-scale visual feature map based on the visual feature sequence, and generating an object-level visual mark sequence by cross-attention calculation between an object query vector and the multi-scale visual feature map; encoding the picking task description information into a text mark sequence, mapping the text mark sequence and the object-level visual mark sequence to a unified semantic space, and splicing to obtain a joint semantic sequence; inputting the joint semantic sequence into a pre-trained VLA model backbone to generate an action mark sequence through autoregressive prediction; decoding the action mark sequence to obtain a picking control instruction, and controlling a picking device to perform a double-spore mushroom picking action based on the picking control instruction; the visual feature sequence is obtained by acquiring continuous multiple frames of visual observation images within a preset time window, dividing the multiple frames of visual observation images into multiple image blocks, and combining the image blocks corresponding to the spatial positions in different time frames to form a space-time unit; performing linear mapping and space-time position encoding on each space-time unit to generate a time-series visual mark sequence as the visual feature sequence; the linear mapping and space-time position encoding of each space-time unit includes: according to the pre-marked double-spore mushroom cultivation rack geometric layout information, assigning a corresponding hierarchical index, row index and column index to each space-time unit, obtaining a structure position encoding by embedding mapping the hierarchical index, row index and column index, and superimposing the structure position encoding and the spatial position encoding for the image block as the space-time position encoding to distinguish the space-time units of different levels, rows and columns in the time-series visual mark sequence; the multi-scale visual feature map is constructed based on the visual feature sequence, which includes obtaining at least two visual feature maps with different spatial resolutions based on the visual feature sequence; calculating a geometric response index in a preset local neighborhood for each visual feature map, which reflects the edge curvature or second-order gradient intensity in the local neighborhood; for the same spatial position, the geometric saliency weight is determined according to the geometric response index of the spatial position at different spatial resolutions, and the corresponding spatial positions of each visual feature map are weighted and fused according to the geometric saliency weight to obtain the multi-scale visual feature map.

2. The VLA-based autonomous Agaricus bisporus picking method according to claim 1, characterized in that, Before performing linear mapping and space-time position encoding on each space-time unit, it also includes: based on the double-spore mushroom cultivation image, marking the cap edge region in each frame of visual observation image, marking the space-time unit falling into the cap edge region as a high-sensitive region, and marking the space-time unit falling into the background structure region as a low-sensitive region; The background structure region includes a bag region, a support region, and a cultivation substrate region without a double mushroom cap, and a first time weight is used when the spatio-temporal position coding is performed on the spatio-temporal unit of the high-sensitive region, and a second time weight different from the first time weight is used when the spatio-temporal position coding is performed on the spatio-temporal unit of the low-sensitive region, to generate the time sequence visual marker sequence.

3. The VLA-based autonomous Agaricus bisporus picking method according to claim 2, characterized in that, The calibration of the cap edge region in each frame of visual observation image comprises: using a segmentation network or a detection network trained for double mushroom to perform inference on the visual observation image to obtain a segmentation mask or a target box representing a cap region; based on the geometric shape features and the gray distribution features of the cap outer contour, performing edge detection and region growing on the visual observation image to extract a region where the cap edge is located as the cap edge region.

4. The VLA-based autonomous Agaricus bisporus picking method according to claim 1, wherein, The calculation of the geometric response index in each of the visual feature maps within a preset local neighborhood comprises: based on the hierarchical index, the row index and the column index of each spatio-temporal unit, defining a local neighborhood aligned with the double mushroom cultivation shelf structure for each spatial position in the visual feature map, the local neighborhood extends along the cultivation shelf row direction and the column direction, and is limited within the spatial position range corresponding to the same level and / or adjacent levels; performing first derivative convolution in the horizontal direction and the vertical direction on each of the visual feature maps within the local neighborhood to obtain gradient components, and constructing a local structure tensor based on the gradient components and calculating the main direction gradient strength of the structure tensor as the geometric response index.

5. The VLA-based autonomous Agaricus bisporus picking method according to claim 4, characterized in that, The generation of the object-level visual marker sequence comprises: based on the hierarchical index, the row index and the column index of each spatio-temporal unit, determining a structure anchor point corresponding to a double mushroom cultivation shelf structure unit, and combining a structure anchor point vector obtained by embedding mapping each structure anchor point with a preset learnable query basis vector to generate an object query vector with cultivation shelf structure position information; based on the corresponding spatial position of each structure anchor point in the multi-scale visual feature map, constructing an attention mask region aligned with the corresponding cultivation shelf level and row-column unit, the attention mask region is limited within a preset row-column neighborhood at the same level and / or adjacent levels with the structure anchor point; in each of the attention mask regions, performing cross-attention calculation with the corresponding object query vector as the query vector, and the feature vector in the multi-scale visual feature map falling within the attention mask region as the key vector and the value vector, taking the cross-attention output as the object-level visual marker of the corresponding structure anchor point to obtain the object-level visual marker sequence.

6. The VLA-based autonomous Agaricus bisporus picking method according to claim 1, wherein, The joint semantic sequence comprises: based on a preset semantic dictionary, inputting the text marker sequence and the object-level visual marker sequence into a unified semantic mapping module to map them into a first semantic representation and a second semantic representation in a common semantic representation space, respectively, the semantic dictionary comprising a plurality of semantic atoms for representing double mushroom cultivation shelf levels, row-column positions, and target mushroom maturity, size, and relative position relationship with the cultivation shelf support structure. In the common semantic representation space, a sparse coefficient vector about the semantic dictionary is solved for the first semantic representation, obtaining a first semantic sparse coefficient; a sparse coefficient vector about the semantic dictionary is solved for each of the second semantic representations, obtaining a second semantic sparse coefficient set corresponding to each object-level visual marker one by one; The first semantic sparse coefficient and the coefficient items in each of the second semantic sparse coefficients corresponding to the semantic atoms used to represent the cultivation shelf level, row and column positions, target mushroom maturity, size and relative position relationship between the target mushroom and the cultivation shelf support structure are added to the corresponding object-level visual marker as semantic constraint coefficients, and are arranged together with the text marker sequence to form the joint semantic sequence.

7. A VLA-based autonomous Agaricus bisporus picking system, characterized in that, Comprise: The acquisition module is used for acquiring the picking task description information and visual observation data of the target scene; The visual observation data is encoded to obtain a visual feature sequence; The visual feature sequence includes: acquiring a plurality of continuous frames of visual observation images within a preset time window, dividing the plurality of frames of visual observation images into a plurality of image blocks, and combining the image blocks corresponding to the spatial positions in different time frames to form a space-time unit; Linear mapping and space-time position coding are performed on each space-time unit to generate a time-series visual marker sequence as the visual feature sequence; The linear mapping and space-time position coding of each space-time unit include: according to the pre-marked double-spore mushroom cultivation shelf geometric layout information, each space-time unit is assigned a corresponding level index, row index and column index, the level index, row index and column index are embedded and mapped to obtain a structure position code, and the structure position code is superimposed with the spatial position code of the image block as the space-time position code to distinguish the space-time units of different levels, rows and columns in the time-series visual marker sequence; The first generation module constructs a multi-scale visual feature map based on the visual feature sequence, and generates an object-level visual marker sequence by cross-attention calculation between the object query vector and the multi-scale visual feature map; The multi-scale visual feature map is constructed based on the visual feature sequence, which includes obtaining at least two visual feature maps with different spatial resolutions based on the visual feature sequence; A geometric response indicator is calculated in a preset local neighborhood for each visual feature map, and the geometric response indicator reflects the edge curvature or second-order gradient strength in the local neighborhood; For the same spatial position, the geometric saliency weight is determined according to the geometric response indicators of the spatial position at different spatial resolutions, and the corresponding spatial positions of each visual feature map are weighted and fused according to the geometric saliency weight to obtain the multi-scale visual feature map; The second generation module is used for encoding the picking task description information into a text marker sequence, and mapping the text marker sequence and the object-level visual marker sequence to a unified semantic space and splicing to obtain a joint semantic sequence; the joint semantic sequence is input into a pre-trained VLA model trunk to generate an action marker sequence through autoregressive prediction; A control module is configured to decode picking control instructions from the action label sequence and control the picking device to perform the double-spore mushroom picking action based on the picking control instructions.

Citation Information

Patent Citations

  • Adhesive mushroom visual identification and measurement method based on feature point detection

    CN110059663A

  • Feature extraction model training method and device, equipment and storage medium

    CN116089651A