Multimodal perception and prescient decision large supermarket robot operation and maintenance method and system

By combining multimodal perception and predictive decision-making methods with unified token embedding of 3D Gaussian splashing, 2D vision and dexterous hand tactile senses, the problem of missing perception modalities and short-sighted decision-making in robots in large supermarket environments is solved, and the efficiency, safety and adaptability of dexterous hand operation are achieved.

CN122378690APending Publication Date: 2026-07-14HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU DIANZI UNIV
Filing Date
2026-04-09
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing multimodal large models of robots have low success rates in highly dynamic and uncertain real-world environments such as large supermarkets. Their perceptual modalities lack semantics and are short-sighted in decision-making, unable to understand complex physical interactions and future scene states. Furthermore, they lack adaptive capabilities, making it difficult for them to perform unseen tasks and cope with environmental changes.

Method used

Employing a multimodal perception and predictive decision-making approach, this method uses a unified token embedding of 3D Gaussian splashing, 2D vision, dexterous hand tactile feedback, and natural language, combined with a predictive decision-making model, to generate dynamically adjusted dexterous hand action commands. This enables deep perception of the environment and prediction of future states, thereby optimizing task sequences and path planning.

Benefits of technology

It improves the robot's perception accuracy and operational compliance in complex environments, significantly enhances the success rate and safety of handling irregularly shaped and fragile goods, shortens task execution time, and strengthens its adaptability in dynamic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122378690A_ABST
    Figure CN122378690A_ABST
Patent Text Reader

Abstract

The application discloses a multimodal perception and foresight decision large supermarket robot operation and maintenance method and system, and the method comprises the following steps: collecting multimodal perception data including 3D Gaussian splash original data, 2D environment pictures, natural language operation and maintenance instruction data and dexterous hand tactile data; constructing an operation and maintenance task optimization mathematical model to generate an initial operation and maintenance task sequence; generating a multimodal unified feature vector based on the multimodal perception data; predicting potential risks and optimal operation timing by using a foresight decision model according to the multimodal unified feature vector to generate a dynamically adjusted operation and maintenance task sequence; and outputting an optimal operation and maintenance scheme by optimizing action timing and path through an adaptive neighborhood search strategy based on the dynamically adjusted operation and maintenance task sequence. The application realizes efficient cooperation of cross-modal information by using a unified token embedding method, generates dexterous hand action instructions adapted to the scene in combination with a foresight decision mechanism, and enables a robot to accurately perceive an environment, dexterously execute an operation and avoid risks in advance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent robot technology, and specifically relates to a method and system for the operation and maintenance of large supermarket robots with multimodal perception and predictive decision-making. Background Technology

[0002] Existing multimodal large model (VLA) robotics techniques, after being trained on large datasets, can perform a range of complex manipulation tasks in specific testing environments. However, when these models are deployed to highly dynamic, uncertain, and real-world environments such as large supermarkets where prior knowledge is difficult to fully cover, their task success rate drops significantly. This is mainly due to two core bottlenecks:

[0003] 1. Semantic gaps and short-sighted decision-making in real-world perceptual modalities. While current advanced VLA models (such as MLA) can integrate vision, language, and action, their visual perception largely relies on raw 2D images or 3D point clouds lacking structured information. They cannot understand the crucial geometric details (such as the precise shape of the object) and physical contact relationships in "hand-object" interactions. Furthermore, the model's decisions are largely based on current observations, lacking predictive reasoning about future scene states after the action is performed. This results in stiff movements, low success rates, and an inability to handle complex situations such as dynamic occlusion when handling fragile, soft, or other items requiring precise force control.

[0004] 2. Insufficient Task Generalization and Dynamic Scene Adaptation Capabilities. An ideal supermarket robot should be able to proactively discover and execute multiple tasks (such as restocking, organizing, and picking). However, existing systems are mostly passive "following commands" models, whose decision-making logic heavily relies on combinations of objects and instructions seen in the training data, making it difficult to generalize to unseen products or zero-sample task instructions. When new anomalies occur in the environment (such as products being accidentally misplaced by customers), the model cannot make autonomous judgments and plans based on a deep understanding of the scene. The core reason for this is the lack of a perceptual foundation capable of extracting general physical and interactive semantics from raw observations, thus limiting its ability to efficiently adapt in unknown environments.

[0005] Therefore, researching how to enable robot systems to have a deep perception of complex physical interactions and predictive decision-making capabilities for future states is the key to breaking through the current bottlenecks in the application of large robot models and realizing their autonomous, precise, and reliable operation in the real, open world. Summary of the Invention

[0006] To address the issues of semantic gaps in perception, short-sighted decision-making, and insufficient scene adaptation capabilities mentioned in the background technology, this invention provides a multimodal perception and predictive decision-making method and system for the operation and maintenance of robots in large supermarkets. Through an innovative unified token embedding method, efficient cross-modal information collaboration is achieved. Combined with a predictive decision-making mechanism, dexterous hand action instructions adapted to the scene are generated, enabling the robot to accurately perceive the environment, dexterously execute operations, and proactively avoid risks, thus optimizing the overall operational performance of robot dexterous hand operations in the complex scenarios of large supermarkets. This method aims to enable the robot system to possess autonomous intelligence throughout the entire process of "observation-thinking-action-reflection," similar to human store clerks, to cope with complex and dynamic retail environments.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0008] The first aspect is a multimodal perception and predictive decision-making method for the operation and maintenance of robots in large supermarkets, which includes the following steps:

[0009] S1. Collect multimodal perception data including 3D Gaussian splash raw data, 2D environmental images, natural language operation and maintenance command data, and dexterous hand tactile data; construct an operation and maintenance task optimization mathematical model and generate an initial operation and maintenance task sequence;

[0010] S2. Generate a unified multimodal feature vector based on multimodal sensing data;

[0011] S3. Based on the multimodal unified feature vector, use the predictive decision model to predict potential risks and optimal operation timing, and generate a dynamically adjusted operation and maintenance task sequence.

[0012] S4. Based on dynamically adjusting the operation and maintenance task sequence, optimize the action timing and path through an adaptive neighborhood search strategy, and output the optimal operation and maintenance solution.

[0013] Preferably, S1 includes:

[0014] S11. Construct a set of operation and maintenance task points and a set of dexterous hand operation modes; the robot moves to the task area and simultaneously collects multimodal perception data.

[0015] S12. Perform constraint processing on operation and maintenance task points and multimodal perception data;

[0016] S13. Construct the objective function of the mathematical model for optimizing operation and maintenance tasks.

[0017] As a preferred embodiment, S12 includes: applying compatibility constraints to the dexterous hand operation mode; and applying validity constraints to the multimodal perception data.

[0018] As a preferred option, the objective function is expressed as:

[0019] in, Operation and maintenance task sequence A single task element; A unified feature vector for multiple modalities; This is a set of constraint parameters for robot motion and dexterous hand manipulation. A function to calculate the execution time of a single task element; The predictive risk cost function quantifies the potential losses caused by environmental changes and dexterity errors. This is the risk weighting coefficient.

[0020] Preferably, S2 includes:

[0021] Density clustering and feature extraction are performed on the raw 3D Gaussian splash data to obtain geometric feature vectors, which are then mapped to standardized tokens. Dexterity hand tactile data are fused to obtain tactile feature vectors, which are then mapped to standardized tokens. Natural language operation and maintenance command data are input into a pre-trained language model to extract semantic feature vectors, which are then mapped to standardized tokens. Preprocessed 2D environment images are input into a convolutional neural network to extract visual feature vectors, which are then mapped to standardized tokens. Using the MLA cross-modal feature alignment approach, the standardized tokens are fused through a multi-head cross-attention mechanism to generate a multimodal unified feature vector.

[0022] Preferably, S3 includes:

[0023] S31. Input the multimodal unified feature vector into the temporal fusion Transformer model to predict the environmental state changes, task execution risks and adaptive dexterity hand movement trends in the next few time steps, and output the predicted feature vector and the corresponding dexterity hand movement parameters.

[0024] S32. Based on the predicted feature vector and dexterous hand motion parameters, the tasks in the initial operation and maintenance task sequence are reordered using the priority adjustment coefficient formula to generate a dynamically adjusted operation and maintenance task sequence.

[0025] Preferably, S4 includes:

[0026] S41. Initialize search parameters;

[0027] S42. Generate neighborhood solutions within the neighborhood of the current optimal solution, which includes dynamically adjusting the operation and maintenance task sequence and the dexterous hand action parameters; adjust the dynamically adjusted operation and maintenance task sequence and the dexterous hand action parameters to generate a neighborhood solution set;

[0028] S43. Evaluate neighborhood solutions using the optimization objective function as the fitness function;

[0029] S44. Update the search radius according to the number of iterations in the search parameters;

[0030] S45. After the iteration terminates, the output includes the optimal operation and maintenance plan, which includes the optimal movement path, the optimal operation and maintenance task sequence, and the optimal dexterous hand motion parameters.

[0031] As a preferred embodiment, S42 includes: performing swap, insertion, and split operations on tasks in the dynamically adjusted operation and maintenance task sequence, and simultaneously performing joint angle fine-tuning, operation force adaptation, and action timing optimization on the dexterous hand motion parameters.

[0032] As a preferred option, S43 also includes evaluating the neighborhood solution by comprehensively optimizing the objective function and the dexterity hand action adaptation function:

[0033]

[0034] in, This represents the dexterity of hand movements, where n represents the number of tasks in the current task sequence. This represents the dexterity hand movement parameters planned for the i-th task. ActMatch represents the multimodal fusion feature vector corresponding to the i-th task. ) is the action-feature matching degree function, and σ(·) is the Sigmoid activation function.

[0035] Secondly, a multimodal perception and predictive decision-making system for the operation and maintenance of large supermarket robots includes:

[0036] The data acquisition and model initialization module is used to collect multimodal perception data, including raw 3D Gaussian splash data, 2D environmental images, natural language operation and maintenance command data, and dexterous hand tactile data; construct an operation and maintenance task optimization mathematical model, and generate an initial operation and maintenance task sequence.

[0037] The feature unification module is used to generate multimodal unified feature vectors based on multimodal sensing data;

[0038] The dynamic adjustment module is used to predict potential risks and optimal operation timing based on multimodal unified feature vectors and predictive decision-making models, and generate dynamically adjusted operation and maintenance task sequences.

[0039] The optimization module is used to optimize the timing and path of actions based on dynamically adjusted operation and maintenance task sequences, and output the optimal operation and maintenance solution through an adaptive neighborhood search strategy.

[0040] The multimodal perception and predictive decision-making large supermarket robot operation and maintenance system is used to implement the multimodal perception and predictive decision-making large supermarket robot operation and maintenance method and steps as described in the first aspect.

[0041] Compared with the prior art, the beneficial effects of the present invention are reflected in:

[0042] 1. Combining 3D Gaussian splashing technology to achieve complementary 3D spatial modeling and 2D visual details, and using a tokenization method based on clustering and activation functions, provides precise environmental and cargo feature support for dexterous hand motion planning, improves motion adaptation accuracy, and reduces scene representation and motion planning deviations;

[0043] 2. Introducing a multi-dimensional tactile sensor for dexterous hands, drawing on the MLA multimodal-action mapping approach, and innovating a four-modal unified token embedding method, the method achieves deep alignment of cross-modal information through multi-head cross-attention and feature-action association mechanisms. This enables image-assisted accurate identification of goods and real-time tactile feedback of action effects, significantly improving the success rate and safety of dexterous hand operations on irregularly shaped, fragile, and similar-looking goods.

[0044] 3. The constructed predictive decision-making model integrates four modal features to predict environmental changes and risks in advance, simultaneously predict the trend of adaptable dexterous hand movements, dynamically adjust task sequences and action parameters, avoid maintenance interruptions caused by action errors, and improve the adaptability of dexterous hand operations in dynamic scenarios.

[0045] 4. The adaptive neighborhood search strategy simultaneously optimizes the task path and dexterous hand motion parameters, achieving a balance between global exploration and local adaptation. Four-modal data provides comprehensive constraint support, improving the efficiency of solving the optimal solution, shortening the total task execution time, and enhancing the engineering practicality of dexterous hand operations. Attached Figure Description

[0046] Figure 1 This is a flowchart illustrating the method steps of Embodiment 1 of the present invention. Detailed Implementation

[0047] To make the technical means, inventive features, objectives, and effects of the invention readily understandable, the invention is further elaborated here with reference to specific implementation examples. However, the invention is not limited to the following implementation examples.

[0048] It should be noted that the processes, steps, parameters, etc. described in this specification are only used to complement the content disclosed in the specification, so that those skilled in the art can understand and read them. They are not intended to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any adjustment of the logical order or change of parameter settings should still fall within the scope of the technical content disclosed in the present invention, provided that it does not affect the effects and objectives that the present invention can produce.

[0049] Example 1:

[0050] A multimodal perception and predictive decision-making method for the operation and maintenance of robots in large supermarkets, including the following steps:

[0051] S1. Collect multimodal sensing data, construct a mathematical model for optimizing operation and maintenance tasks, and generate an initial sequence of operation and maintenance tasks;

[0052] S11, Operation and Maintenance Task Initialization Phase; Assume the set of supermarket robot operation and maintenance task points is C={c1,c2,…,cn} (e.g., c1 is the goods storage location, c2 is the target shelf location, c3 is the inventory scanning location), and the set of robot end effector operation modes is H={h1,h2,…,hk} (e.g., h1 is "strong grip" for rice, flour, and cooking oil, h2 is "adaptive pinch" for fruits and vegetables, and h3 is "light touch" for fragile items). Collect multimodal perception data. When the robot moves to the task area, the multi-sensor data acquisition process is simultaneously initiated:

[0053] (1) The target shelf and surrounding aisle are scanned in three dimensions by using a top binocular vision sensor and a lidar to generate 3D Gaussian splash raw data describing spatial geometry and semantics.

[0054] (2) Take 2D high-definition pictures of the target goods (such as a glass bottle of juice) and its surrounding environment at the task point through the high-definition camera at the front of the robot.

[0055] (3) Receive the backend command "Replenish XX brand orange juice on the third shelf of shelf A03" through the voice module and convert it into structured text data.

[0056] (4) When the dexterous hand approaches the goods, its surface array tactile sensors are activated to prepare to collect the pressure contact surface distribution signal during the grasping process.

[0057] The multimodal sensing data source set is M={m1,m2,m3,m4} (m1 is 3D Gaussian splash scene data, m2 is dexterous hand tactile data, m3 is natural language operation and maintenance command data, and m4 is captured image data). The corresponding datasets are: M1 (the set of 3D Gaussian wave field scene data), M2 (the set of dexterous hand tactile data), M3 (the set of natural language operation and maintenance command data), and M4 (the set of captured image data). The dexterous hand motion parameter set is A={a1,a2,⋯,am} (including joint angles, operating force, motion timing, etc.).

[0058] S12. Constraint handling, the constraints in constraint handling include:

[0059] (1) Operation mode compatibility constraints:

[0060] in, For dexterous hand operation mode h, the task point and The compatibility function, a non-zero value indicates that the operation mode h allows execution Execute after task The task (or vice versa) must be within the physical constraints, ensuring that the task sequence matches the dexterity's operational characteristics. For example, after performing a "forceful grasp" of a heavy object, the dexterity's joints need to relax briefly, and it is not advisable to immediately perform a "light touch" task that requires high precision.

[0061] (2) Constraints on the validity of multimodal data:

[0062]

[0063] Where s(·) is the matching degree function between the four-modal data and the task point, and λ is the preset matching degree threshold, to ensure that each task point has effective and reliable multimodal data support and avoid action planning failure due to lack of perception information.

[0064] S13. Generating the Optimization Objective Function: The objective function for optimizing the total task time and risk cost in the objective processing is defined as follows:

[0065] Where: q is the initial operation and maintenance task sequence A single task element; F is the multimodal unified feature vector to be generated; b is the set of robot motion and dexterous hand operation constraint parameters (including dexterous hand joint angle range, operation force threshold, robot movement speed limit, etc.). The function for calculating the execution time of a single task element is determined by the path distance, the complexity of the dexterous hand operation, and the multimodal feature matching degree. This is a predictive risk cost function used to quantify potential losses caused by dynamic environmental changes and dexterity errors. This is a risk weighting coefficient used to balance task execution efficiency and dexterity operation safety.

[0066] S2. Generate a unified multimodal feature vector based on multimodal sensing data;

[0067] S21. Perform density clustering on the 3D Gaussian splash data to extract geometric feature vectors G, including shelf height and cargo gaps, and map them to geometric feature tokens using the following formula:

[0068]

[0069] in:

[0070] This represents the geometric feature token vector extracted and transformed from 3D Gaussian splash data;

[0071] This represents the geometric feature vector extracted from the original 3D Gaussian splash data through density clustering;

[0072] and These are the weight matrix and bias vector used for feature projection, respectively, obtained through model training;

[0073] (·) represents a non-linear activation function (such as ReLU), used to introduce the non-linear expressive power of the model.

[0074] S22. Input the captured 2D images into a convolutional neural network (CNN) to extract visual feature vectors such as the color of the goods, the trademark, and the orientation of the bottle cap. The following formula maps to visual feature tokens:

[0075]

[0076] in:

[0077] This represents a visual feature token vector extracted and transformed from a 2D image.

[0078] This represents the visual feature vector extracted from a preprocessed 2D image using a convolutional neural network (CNN).

[0079] and These are the projection weight matrix and bias vector of the visual features, respectively;

[0080] S23. Input the structured language instructions into the pre-trained model and extract semantic feature vectors such as "supplement", "A03 third layer", and "orange juice". It is mapped to a semantic instruction token using the following formula:

[0081]

[0082] in: For projection matrix, This is the bias vector.

[0083] S24. The dexterous hand performs a pre-grasping light touch action, and the tactile sensor collects the pressure distribution feature vector of the initial contact. The following formula maps to tactile feature tokens:

[0084]

[0085] in:

[0086] This represents the tactile feature token vector converted from the tactile signal.

[0087] This represents the tactile feature vector obtained by fusing pressure, temperature, and contact area signals collected by the surface array tactile sensors of the dexterous hand.

[0088] and These are the projection weight matrix and bias vector of the tactile features, respectively.

[0089] It is a layer normalization function used to stabilize the dimensions and distribution of multi-dimensional tactile signals.

[0090] The weighting coefficients for the normalization term are used to balance the original projection and the normalization effect.

[0091] S25. Drawing on the MLA cross-modal feature alignment approach, a multimodal unified feature vector is obtained:

[0092]

[0093] in:

[0094] This represents a unified feature token vector that integrates information from four modalities, serving as the core input for subsequent decision-making models; This represents a multi-head cross-attention mechanism used to calculate the correlation between tokens of different modalities and to achieve special alignment and deep fusion.

[0095] and These are the weight matrix and bias vector of the fusion layer, respectively;

[0096] This represents the feature context related to action planning extracted from the current multimodal features;

[0097] This is a feature-action correlation function that outputs a scalar to quantify the degree of matching between current environmental features and preset dexterity hand action patterns.

[0098] The association weight coefficient is used to control the contribution strength of feature-action association information to the final unified token.

[0099] S3. Based on the multimodal unified feature vector, use the predictive decision model to predict potential risks and optimal operation timing, and generate a dynamically adjusted operation and maintenance task sequence.

[0100] S31, Predictive Decision-Making Model Reasoning: [The sentence is incomplete and requires more context to be translated accurately.] Input a temporal fusion Transformer model to predict environmental state changes, task execution risks, and adaptive dexterity hand movement trends at several future time steps, and output a predicted feature vector. and corresponding dexterous hand motion parameters ;

[0101] S32. Dynamic adjustment of task sequences and action preparameters: based on The initial operation and maintenance task sequence is adjusted using a priority adjustment coefficient formula. Reorder the tasks in the process:

[0102]

[0103] in:

[0104] This represents the priority adjustment coefficient for task q. The larger the value, the more priority the task needs to be processed or adjusted.

[0105] β is a balancing coefficient (0≤β≤1), used to adjust the relative importance of execution time cost and predictable risk cost in decision-making;

[0106] This is a function for calculating the execution time of task q.

[0107] Let be the fusion risk-cost function for task q. It quantifies not only environmental risk but also the risk caused by the dexterous hand's movement parameters. Compared with current predicted features Potential operational risks caused by fit deviation (such as slippage, crushing);

[0108] and These are the maximum values ​​when normalizing all tasks.

[0109] For high-risk task points, the coefficient Due to risk items As the number of tasks increases, the system may postpone their execution or allocate more time for execution and adjustment, generating a dynamically adjusted task sequence Q. Aadj and the corrected set of dexterous hand motion parameters A adj .

[0110] S4. Based on dynamically adjusting the operation and maintenance task sequence, optimize the action timing and path through an adaptive neighborhood search strategy, and output the optimal operation and maintenance solution.

[0111] S41. Initialize search parameters: Set the initial search radius r0 (for the task sequence perturbation range), iteration decay coefficient η, maximum number of iterations, and dexterity hand motion parameter optimization threshold ε. S42. Generate neighborhood solutions: In the current optimal solution (Q... Aadj A adjWithin the neighborhood of the target object, new solutions are generated synchronously. On one hand, the task sequence is "exchanged" (e.g., the order of two inventory tasks is changed) and "inserted" (a temporary obstacle avoidance stop is inserted into the transport path); on the other hand, the dexterous hand motion parameters are "fine-tuned" (the initial angle of the finger joints is fine-tuned based on the bottle cap orientation identified by the visual token). S43, Neighborhood Solution Evaluation: To optimize the objective function. The overall time consumption and risk of neighborhood solutions are evaluated. Simultaneously, a dexterous hand action fitness function is introduced for specific evaluation.

[0112]

[0113] in:

[0114] This represents the dexterity of hand movements. The higher the value, the better the match between the currently planned movement parameters and the environmental features perceived by the multimodal sensor.

[0115] n represents the number of tasks in the current task sequence.

[0116] This represents the dexterity hand motion parameters planned for the i-th task.

[0117] Let represent the multimodal fusion feature vector corresponding to the i-th task.

[0118] ActMatch () is the action-feature matching function, used to calculate the matching degree of action-features given features. Below, motion parameters The probability of successful execution or security score.

[0119] σ(·) is the Sigmoid activation function, which is used to map the average matching degree to the (0,1) interval, making it convenient to set the target threshold.

[0120] S44. Dynamically Adjust the Search Strategy: As iterations proceed, dynamically reduce the search radius according to the following formula:

[0121] iter

[0122] in:

[0123] This represents the search radius of the current iteration, which determines the range of neighborhood solutions that can be generated.

[0124] This is the initial search radius.

[0125] The iterative decay coefficient (usually) ).

[0126] iter This represents the current iteration number.

[0127] This allows the algorithm to focus on global path exploration and action pattern optimization in the early stages, and on fine-tuning local paths and adapting dexterous hand action parameters in the later stages.

[0128] S45. Determine the optimal solution: After the iteration terminates, output the fitness. Optimal and motion adaptation The optimal solution that meets the requirements. This solution includes the robot's optimal movement path, task execution sequence, and precisely generated dexterous hand operation instructions for each operation step (e.g., open with α degrees, contact the anti-slip texture in the middle of the bottle with β Newtons of pressure, hold for γ seconds and then lift), forming a complete operation and maintenance solution that can be directly executed.

[0129] The operation and maintenance method described in this invention constructs an optimized model that integrates multimodal perception and predictive decision-making. It innovatively employs a unified token embedding method to achieve deep alignment and fusion of 3D scenes, 2D vision, tactile feedback, and language commands. Furthermore, it utilizes a temporal prediction model to anticipate risks and pre-adjust actions. Finally, an adaptive neighborhood search strategy is used to achieve collaborative optimization of the task path and dexterous hand action parameters. In typical scenarios such as "replenishing fragile goods," this method significantly improves the robot's perception accuracy in complex environments and irregularly shaped goods, making dexterous hand operations smoother, safer, and more efficient. It effectively overcomes the problems of perception-action disconnect and decision-making lag in traditional technologies, thereby enhancing the overall intelligence level of robot operation and maintenance in large supermarkets.

[0130] Example 2:

[0131] A multimodal perception and predictive decision-making robot operation and maintenance system for large supermarkets, including:

[0132] The data acquisition and model initialization module is used to collect multimodal perception data, including raw 3D Gaussian splash data, 2D environmental images, natural language operation and maintenance command data, and dexterous hand tactile data; construct an operation and maintenance task optimization mathematical model, and generate an initial operation and maintenance task sequence.

[0133] The feature unification module is used to generate multimodal unified feature vectors based on multimodal sensing data;

[0134] The dynamic adjustment module is used to predict potential risks and optimal operation timing based on multimodal unified feature vectors and predictive decision-making models, and generate dynamically adjusted operation and maintenance task sequences.

[0135] The optimization module is used to optimize the timing and path of actions based on dynamically adjusted operation and maintenance task sequences, and output the optimal operation and maintenance solution through an adaptive neighborhood search strategy.

Claims

1. A multimodal perception and predictive decision-making method for the operation and maintenance of robots in large supermarkets, characterized in that: Includes the following steps: S1. Collect multimodal perception data including 3D Gaussian splash raw data, 2D environmental images, natural language operation and maintenance command data, and dexterous hand tactile data; construct an operation and maintenance task optimization mathematical model and generate an initial operation and maintenance task sequence; S2. Generate a unified multimodal feature vector based on multimodal sensing data; S3. Based on the multimodal unified feature vector, use the predictive decision model to predict potential risks and optimal operation timing, and generate a dynamically adjusted operation and maintenance task sequence. S4. Based on dynamically adjusting the operation and maintenance task sequence, optimize the action timing and path through an adaptive neighborhood search strategy, and output the optimal operation and maintenance solution.

2. The multimodal perception and predictive decision-making method for the operation and maintenance of large supermarket robots according to claim 1, characterized in that, S1 includes: S11. Construct a set of operation and maintenance task points and a set of dexterous hand operation modes; The robot moves to the task area and simultaneously collects multimodal perception data; S12. Perform constraint processing on operation and maintenance task points and multimodal perception data; S13. Construct the objective function of the mathematical model for optimizing operation and maintenance tasks.

3. The multimodal perception and predictive decision-making method for the operation and maintenance of large supermarket robots according to claim 2, characterized in that, S12 includes: compatibility constraints on dexterous hand operation modes; and validity constraints on multimodal perception data.

4. The multimodal perception and predictive decision-making method for the operation and maintenance of large supermarket robots according to claim 2, characterized in that, The objective function is expressed as: in, Operation and maintenance task sequence A single task element; A unified feature vector for multiple modalities; This is a set of constraint parameters for robot motion and dexterous hand manipulation. A function to calculate the execution time of a single task element; The predictive risk cost function quantifies the potential losses caused by environmental changes and dexterity errors. This is the risk weighting coefficient.

5. The multimodal perception and predictive decision-making method for the operation and maintenance of large supermarket robots according to claim 1, characterized in that, S2 include: Density clustering and feature extraction are performed on the raw 3D Gaussian splash data to obtain geometric feature vectors, which are then mapped to standardized tokens. Dexterity hand tactile data are fused to obtain tactile feature vectors, which are then mapped to standardized tokens. Natural language operation and maintenance command data are input into a pre-trained language model to extract semantic feature vectors, which are then mapped to standardized tokens. Preprocessed 2D environment images are input into a convolutional neural network to extract visual feature vectors, which are then mapped to standardized tokens. Using the MLA cross-modal feature alignment approach, the standardized tokens are fused through a multi-head cross-attention mechanism to generate a multimodal unified feature vector.

6. The multimodal perception and predictive decision-making method for the operation and maintenance of large supermarket robots according to claim 1, characterized in that, S3 include: S31. Input the multimodal unified feature vector into the temporal fusion Transformer model to predict the environmental state changes, task execution risks and adaptive dexterity hand movement trends in the next few time steps, and output the predicted feature vector and the corresponding dexterity hand movement parameters. S32. Based on the predicted feature vector and dexterous hand motion parameters, the tasks in the initial operation and maintenance task sequence are reordered using the priority adjustment coefficient formula to generate a dynamically adjusted operation and maintenance task sequence.

7. The multimodal perception and predictive decision-making method for the operation and maintenance of large supermarket robots according to claim 1, characterized in that, S4 include: S41. Initialize search parameters; S42. Generate neighborhood solutions within the neighborhood of the current optimal solution, which includes dynamically adjusting the operation and maintenance task sequence and the dexterous hand action parameters; adjust the dynamically adjusted operation and maintenance task sequence and the dexterous hand action parameters to generate a neighborhood solution set; S43. Evaluate neighborhood solutions using the optimization objective function as the fitness function; S44. Update the search radius according to the number of iterations in the search parameters; S45. After the iteration terminates, the output includes the optimal operation and maintenance plan, which includes the optimal movement path, the optimal operation and maintenance task sequence, and the optimal dexterous hand motion parameters.

8. The multimodal perception and predictive decision-making method for the operation and maintenance of large supermarket robots according to claim 7, characterized in that, S42 includes: performing swap, insertion, and split operations on tasks in the dynamically adjusted operation and maintenance task sequence, and simultaneously performing joint angle fine-tuning, operation force adaptation, and action timing optimization on the dexterous hand motion parameters.

9. The multimodal perception and predictive decision-making method for the operation and maintenance of large supermarket robots according to claim 7, characterized in that, S43 also includes evaluating neighborhood solutions by comprehensively optimizing the objective function and the dexterity hand action adaptation function: ; in, This represents the dexterity of hand movements, where n represents the number of tasks in the current task sequence. This represents the dexterity hand movement parameters planned for the i-th task. ActMatch represents the multimodal fusion feature vector corresponding to the i-th task. ) is the action-feature matching degree function, and σ(·) is the Sigmoid activation function.

10. A multimodal perception and predictive decision-making robot operation and maintenance system for large supermarkets, characterized in that: include: The data acquisition and model initialization module is used to collect multimodal perception data, including raw 3D Gaussian splash data, 2D environmental images, natural language operation and maintenance command data, and dexterous hand tactile data; construct an operation and maintenance task optimization mathematical model, and generate an initial operation and maintenance task sequence. The feature unification module is used to generate multimodal unified feature vectors based on multimodal sensing data; The dynamic adjustment module is used to predict potential risks and optimal operation timing based on multimodal unified feature vectors and predictive decision-making models, and generate dynamically adjusted operation and maintenance task sequences. The optimization module is used to optimize the timing and path of actions based on dynamically adjusted operation and maintenance task sequences, and output the optimal operation and maintenance solution through an adaptive neighborhood search strategy. The multimodal perception and predictive decision-making large supermarket robot operation and maintenance system is used to implement the multimodal perception and predictive decision-making large supermarket robot operation and maintenance method and steps as described in claim 1.