Multi-modal input intelligent agent decision interaction method and system
By acquiring and deeply analyzing multimodal inputs, flexible decision-making strategies are generated and reliable data transmission is ensured. This overcomes the limitations of single-modal input in traditional intelligent agent decision-making interaction systems and improves the system's collaborative efficiency and stability in complex environments.
Patent Information
- Application Number
- CN202511446053.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Traditional intelligent agent decision-making and interaction systems rely on single-modal input, which cannot fully reflect the environmental state and lack the ability to deeply analyze multi-source data. This results in inflexible decision-making strategy generation, untimely and unreliable transmission of interaction results, and difficulty in meeting the actual needs of complex environments.
A multimodal input acquisition unit is used to acquire visual, audio, and sensor signal data. Multimodal feature data is generated through multi-scale feature extraction, spectrum analysis, and time series analysis. This data is then combined with machine learning for environmental analysis and decision-making strategy generation. Finally, the interactive results are transmitted via wireless communication.
It achieves comprehensive environmental perception, generates flexible decision-making strategies, ensures accurate execution of interactive actions and real-time reliability of data transmission, and improves the system's collaborative work efficiency and stability in complex environments.
Smart Images

Figure CN120930073B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent interaction technology, and in particular to a method and system for intelligent agent decision-making interaction with multimodal input. Background Technology
[0002] In current intelligent system applications, whether in industrial control, smart homes, service robots, or autonomous driving, increasingly higher demands are being placed on the decision-making and interaction capabilities of intelligent agents. Traditional intelligent agent decision-making and interaction systems mostly rely on single-modal input signals for data acquisition and processing, such as acquiring image information solely through a visual camera or acquiring sound signals solely through a voice acquisition device. This single-modal input approach has significant limitations; when the target environment undergoes complex changes, single-modal data often fails to fully reflect the true state of the environment.
[0003] Taking service robots as an example, if a robot relies solely on visual signals to identify its surroundings, the accuracy of the visual data will significantly decrease in situations such as dim lighting or obstructed views. This can lead to the robot being unable to accurately determine the location of obstacles, user needs, and other crucial information, thus affecting the generation of decision-making strategies and ultimately resulting in interactions that fail to meet actual requirements. In the field of industrial control, relying solely on sensor data such as temperature and pressure is insufficient to fully grasp the operating status of equipment. Once data deviations or omissions occur, the intelligent agent may make incorrect decisions, leading to equipment malfunctions or even safety accidents.
[0004] Traditional systems often employ simplistic input data processing methods, lacking the ability to effectively fuse and deeply analyze multi-source data. Even when some systems attempt to introduce multiple modalities, they typically only process and superimpose data from different modalities individually, failing to fully uncover the inherent relationships between them. This results in processed feature data that cannot accurately and comprehensively represent the target environment. Furthermore, in the decision-making strategy generation stage, traditional systems are mostly based on preset rules or simple algorithm models, lacking flexibility and adaptability, and struggling to cope with complex and changing environmental conditions. Regarding the transmission of interactive result data, traditional systems generally suffer from low transmission efficiency and poor security, failing to provide timely and reliable feedback of interactive results to external devices, thus impacting the overall collaborative efficiency and application effectiveness of the system.
[0005] With the rapid development of artificial intelligence technology, various fields have put forward higher requirements for the real-time performance, accuracy, and reliability of intelligent agent decision-making interaction. Traditional intelligent agent decision-making interaction systems with single-modal input, simple data processing, fixed decision-making patterns, and inefficient data transmission are gradually failing to meet the needs of practical application scenarios. There is a need for an intelligent agent decision-making interaction system that can achieve multimodal input, efficient data processing, flexible decision generation, and reliable data transmission. Summary of the Invention
[0006] The main objective of this invention is to provide a method and system for intelligent agent decision-making interaction with multimodal input, aiming to solve the technical problems in the prior art.
[0007] This invention proposes a multimodal input intelligent agent decision-making interaction system, comprising:
[0008] A multimodal input acquisition unit is used to acquire input signal data of multiple modes from the target environment;
[0009] The input processing unit is used to process the input signal data of the multiple modalities to obtain multimodal feature data;
[0010] A decision strategy generation unit is used to generate a decision strategy for the intelligent agent based on the multimodal feature data;
[0011] An interactive execution unit is used to execute interactive actions according to the decision-making strategy;
[0012] The data transmission unit is used to transmit the interaction result data to an external device.
[0013] Preferably, the input processing unit processes the input signal data of the multiple modalities to obtain multimodal feature data, specifically as follows:
[0014] Acquire visual modal input signal data, audio modal input signal data, and sensor modal input signal data of the target environment;
[0015] Multi-scale feature extraction is performed on the visual modality input signal data to obtain visual feature data;
[0016] Perform spectrum analysis on the audio modal input signal data to obtain audio feature data;
[0017] Time series analysis is performed on the sensor modal input signal data to obtain sensor feature data;
[0018] The visual feature data, audio feature data, and sensor feature data are fused to obtain multimodal feature data.
[0019] Preferably, the decision strategy generation unit generates a decision strategy for the agent based on the multimodal feature data, specifically as follows:
[0020] Environmental scenario analysis is performed based on the multimodal feature data. An environmental feature tensor is constructed through pattern recognition. Object recognition is performed on the multimodal feature data based on a machine learning classifier to generate a target environmental state model that includes environmental objects, scenario labels, and environmental states.
[0021] The target environment state model is mapped to the agent's objective to construct an agent behavior decision model;
[0022] The agent's behavior strategy is determined based on the agent's behavior decision-making model, thus obtaining the decision strategy.
[0023] Preferably, the decision strategy generation unit determines the agent's behavior strategy based on the agent's behavior decision model, specifically as follows:
[0024] Acquire the initial position information and target position information of the intelligent agent, and determine the obstacle information in the environment based on the target environment state model;
[0025] Based on the path planning algorithm, path planning is performed on the initial position information and the target position information, and the obstacle information is used as planning constraints to output the initial behavior path of the agent;
[0026] Real-time acquisition of agent state change data during the execution of the agent according to the initial behavior path, and determination of the environmental dynamics of the agent's real-time position based on the state change data;
[0027] Acquire behavioral stability data of the intelligent agent, wherein the behavioral stability data includes data on the intelligent agent's ability to adapt to different environmental changes;
[0028] Acquire the agent's reaction time data to dynamic changes in the environment, and determine the agent's behavior path correction lag based on the reaction time data.
[0029] Based on the adaptive capability data and the behavioral path correction lag, the environmental dynamics of the agent's real-time position are analyzed to determine the cumulative behavioral path deviation of the agent within the recognition reaction time.
[0030] If the cumulative amount of the behavior path offset is less than a preset value, the behavior offset direction and behavior offset distance executed by the agent according to the initial behavior path are determined based on the cumulative amount of the behavior path offset, and the correction direction and correction distance of the agent are determined based on the behavior offset direction and behavior offset distance to obtain correction data;
[0031] Based on the correction data, the initial behavior path of the agent during real-time execution is corrected to obtain the first behavior strategy.
[0032] If the cumulative offset of the behavior path is greater than a preset value, obtain the environmental dynamic change data of the real-time execution path of the agent, construct an environmental change map based on the environmental dynamic change data of the real-time execution path, and determine the environmental dynamic change trend in the target environment based on the environmental change map.
[0033] Interpolation is performed on the dynamic change trend of the environment using an interpolation algorithm to determine the dynamic information of the environment within the preset range of the initial behavior path. The behavioral stability of the agent within the preset range of the initial behavior path is determined based on the dynamic information of the environment and the adaptability data. The initial behavior path segments with a cumulative offset greater than a preset value are adjusted based on the behavioral stability to obtain an updated behavior path. The agent executes the behavior according to the updated behavior path to obtain a second behavior strategy.
[0034] Preferably, the input processing unit further includes a multimodal feature matching analysis subunit, used for:
[0035] Based on the multimodal feature data, the feature distribution, feature mean, and feature peak value of each mode are calculated to obtain the feature spectrum;
[0036] The feature matching degree is determined based on the feature spectrum, and a feature matching parameter set is generated.
[0037] The feature fusion parameters are adjusted based on the feature matching parameter set to optimize the multimodal feature data.
[0038] Preferably, the decision strategy generation unit further adjusts the agent's behavioral decision model based on the feature matching parameter set, specifically as follows:
[0039] The reliability of the environmental state model is analyzed based on the feature matching parameter set, and the parameters in the environmental state model are adjusted accordingly.
[0040] The agent's behavior decision-making model is reconstructed based on the adjusted environmental state model, and an updated decision-making strategy is generated.
[0041] Preferably, the interaction execution unit executes the interaction action according to the decision strategy, specifically as follows:
[0042] Obtain standard data of different objects in the target environment, and annotate the standard data to obtain annotated data;
[0043] An object recognition model is constructed based on a deep learning model, and the labeled data is imported into the object recognition model for training.
[0044] The agent obtains interaction data of the target environment according to the decision-making strategy, imports the interaction data into the trained object recognition model for object recognition, and counts the number of each object type to obtain interaction result data.
[0045] Preferably, the interactive execution unit further includes an exception handling subunit, used for:
[0046] Real-time monitoring of state fluctuations and data offsets during the execution of interactive actions by intelligent agents, and analysis of the impact of fluctuation range on interaction results;
[0047] Dynamically adjust the behavior parameters and data acquisition paths within the target range to generate an anomaly handling dataset;
[0048] Adjust the object recognition model parameters based on the anomaly handling dataset to optimize the interaction result data.
[0049] Preferably, the data transmission unit transmits the interaction result data to an external device, specifically as follows:
[0050] A hierarchical communication protocol stack is constructed based on wireless communication technology, and the interaction result data is modulated to construct a transmission signal.
[0051] The transmission signal is sent to an external device according to the layered communication protocol stack, and the transmission signal is decoded to obtain interactive decision data of the target environment.
[0052] Preferably, the present invention also includes a multimodal input intelligent agent decision-making interaction method, the method comprising the following steps:
[0053] The multimodal input acquisition unit acquires input signal data of multiple modes from the target environment.
[0054] The input signal data of the multiple modalities are processed by the input processing unit to obtain multimodal feature data;
[0055] The decision strategy generation unit generates a decision strategy for the intelligent agent based on the multimodal feature data;
[0056] The interactive execution unit executes interactive actions according to the decision-making strategy;
[0057] The interaction result data is transmitted to external devices based on the data transmission unit.
[0058] The beneficial effects of this invention are as follows:
[0059] By setting up a multimodal input acquisition unit, the system can acquire input signal data from multiple modalities of the target environment, breaking the limitation of traditional intelligent agent decision-making and interaction systems that rely on a single modal input. In practical applications, signal data from vision, hearing, touch, or other modalities can be effectively captured by the system, enabling it to comprehensively perceive the state of the target environment from multiple dimensions. For example, in a smart home scenario, the system can acquire indoor image information through cameras to determine the activity status of family members, acquire voice commands through sound acquisition devices, and obtain indoor environmental parameters through temperature and humidity sensors. The acquisition of this multimodal data makes the system's understanding of the home environment more comprehensive, avoiding cognitive biases caused by missing or inaccurate single-modal data.
[0060] The input processing unit processes input signal data from multiple modalities to obtain multimodal feature data. This unit does not simply process different modal data individually, but rather achieves deep fusion and analysis of multi-source data. Through effective data processing methods, the system can uncover the inherent correlations between different modal data, transforming scattered and fragmented input signals into multimodal feature data with unified representational meaning. This processing approach enables the feature data to more accurately and comprehensively reflect key information of the target environment, providing a high-quality data foundation for the generation of subsequent decision-making strategies. For example, in autonomous driving scenarios, the input processing unit can fuse road image data acquired by cameras, distance data acquired by radar, and the vehicle's own speed and position data to generate feature data that clearly represents road conditions, surrounding obstacle information, and the vehicle's own state, allowing the system to gain a deeper understanding of the driving environment.
[0061] The decision-making strategy generation unit generates decision-making strategies for the intelligent agent based on multimodal feature data. Because the multimodal feature data it relies on is comprehensive and accurate, this unit can break free from the constraints of traditional systems that generate decisions based on fixed rules or simple algorithms, and flexibly adjust the decision logic according to real-time changes in the target environment. During the interaction between the service robot and the user, the decision-making strategy generation unit can accurately determine the user's true needs based on information such as facial expressions, tone of voice, and gestures reflected in the multimodal feature data, and then generate a matching interaction decision strategy. For example, when a user shows anxiety and their voice contains a request for help, the system can quickly generate a decision to prioritize responding to the request for help, improving the targeting and effectiveness of the interaction.
[0062] The interactive execution unit executes interactive actions based on the decision-making strategy. This unit works efficiently with the decision-making strategy generation unit to ensure that the generated decision-making strategy can be translated into actual interactive actions in a timely and accurate manner. In industrial robot operation scenarios, once the decision-making strategy generation unit determines the specific work process and action parameters, the interactive execution unit can accurately execute actions such as grasping, handling, and assembly according to the strategy. This avoids action deviations caused by the disconnect between the execution and decision-making stages, ensuring the smooth completion of the task.
[0063] The data transmission unit transmits the interaction results to external devices, ensuring the real-time nature and reliability of this data transmission. In scenarios involving multi-device collaboration, such as the collaboration between multiple intelligent agents and a central control system in a smart factory, the data transmission unit can quickly transmit the interaction results from each intelligent agent to the central control system. This allows the central control system to promptly grasp the working status and task completion of each intelligent agent, facilitating overall scheduling and coordination. Simultaneously, reliable data transmission prevents data loss, damage, or leakage during transmission, ensuring the security of data interaction throughout the system, promoting efficient collaboration between the system and external devices, and improving the overall system's efficiency and stability in practical applications. Attached Figure Description
[0064] Figure 1 This is a timing diagram of the multimodal input intelligent agent decision-making interaction system described in this invention;
[0065] Figure 2 A flowchart for generating decision strategies for the decision strategy generation unit.
[0066] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0067] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0068] like Figure 1 As shown, this application provides a multimodal input intelligent agent decision-making interaction system, including: a multimodal input acquisition unit, an input processing unit, a decision strategy generation unit, an interaction execution unit, and a data transmission unit. The multimodal input acquisition unit acquires visual modal input signal data, audio modal input signal data, and sensor modal input signal data of the target environment through configured visual sensors, audio acquisition devices, and various environmental sensors. The input processing unit receives the aforementioned multimodal input signal data and performs multi-scale feature extraction, spectrum analysis, and time series analysis operations, transmitting the processed multimodal feature data to the decision strategy generation unit. The decision strategy generation unit constructs an environmental feature tensor based on the multimodal feature data, performs object recognition and scenario analysis through a machine learning classifier, generates a target environment state model, maps it to the intelligent agent target to form an intelligent agent behavior decision model, and outputs a decision strategy. The interaction execution unit drives the intelligent agent to perform interactive actions according to the decision strategy, collects interactive data, and generates interactive result data. The data transmission unit transmits the interactive result data to an external device after modulation and demodulation via a wireless communication module.
[0069] In one embodiment, Example 1: See Figure 2The system collects visual modal input signal data of the target environment through high-definition cameras deployed on the agent. This image data is transmitted to the input processing unit at a rate of 30 frames per second. Simultaneously, a microphone array mounted on the agent's head continuously captures audio modal input signal data of the environment, with a sampling frequency set to 44.1 kHz. Multi-axis inertial measurement units and infrared sensors distributed on the agent's chassis generate sensor modal input signal data in real time, including acceleration, angular velocity, and distance measurements. The input processing unit initiates a parallel processing flow to perform multi-scale feature extraction operations on the visual data stream. It employs a convolutional neural network with a residual connection structure to extract spatial features and local texture features at three different resolution levels. The first level processes the original resolution image to obtain macroscopic contour features, the second level analyzes intermediate semantic features from the downsampled image, and the third level focuses on high-resolution regions to extract microscopic detail features. Finally, the output feature maps from the three levels are concatenated to generate visual feature data. The audio data stream enters the spectrum analysis process. First, Hanning windows are applied for frame segmentation, with each frame containing 1024 sampling points. A Fast Fourier Transform (FFT) is used to convert the time-domain signal to a frequency-domain representation. The energy of the Mel filter bank is calculated, and its logarithm is taken to obtain the log-Mel spectrum. Based on this, a Discrete Cosine Transform (DCT) is performed to extract the Mel frequency cepstral coefficients. Finally, the first 20 dimensions of coefficients are selected as the audio feature data. The sensor data stream undergoes time-series analysis, using a sliding window mechanism to process the real-time data stream. The window length is set to 2 seconds. Within each window, time-domain statistics, including mean, variance, and zero-crossing rate, are calculated. Simultaneously, a Fast Fourier Transform is performed to obtain the amplitudes of the main frequency components. The time-domain and frequency-domain features are combined to form the sensor feature data. The feature fusion module receives three types of modal feature data, performs global average pooling to reduce the dimensionality of the visual feature data, performs principal component analysis to compress the audio feature data, and standardizes the sensor feature data. It then uses a feature weighted concatenation strategy to connect the processed feature vectors into a unified multimodal feature vector, where the weight of the visual feature is set to 0.5, the weight of the audio feature is 0.3, and the weight of the sensor feature is 0.2. The final output is multimodal feature data with a dimension of 512.
[0070] After receiving multimodal feature data, the decision-making strategy generation unit initiates the environmental scenario analysis process. First, it constructs an environmental feature tensor, reshaping the multimodal feature data into a three-dimensional tensor structure. The first dimension corresponds to spatial features, the second to temporal features, and the third retains modal features. Spatiotemporal feature patterns are extracted through three-dimensional convolution operations. Object recognition is performed based on a support vector machine classifier. During the training phase, a labeled dataset is used to learn the mapping relationship between multimodal features and object categories. During the deployment phase, real-time feature data is input, and the output recognition results include object categories and location coordinates. The scenario label generation module analyzes the spatial relationships and temporal series patterns between objects, determining whether the current environmental state belongs to a scenario type such as "congested," "empty," or "dangerous" based on a predefined scenario rule base. The environmental state assessment module calculates the environmental stability index by comprehensively considering object distribution density, movement trends, and acoustic features, ultimately generating a target environmental state model containing a list of environmental objects, scenario labels, and an environmental state index. The agent target mapping module reads the preset task target. When the target is "rapid traversal," it activates the path optimization strategy; when the target is "safe obstacle avoidance," it adopts the conservative decision-making strategy. A value function is constructed using a reinforcement learning framework, with environmental state model parameters as state input and agent actions as the action space. The Q-learning algorithm iteratively optimizes the action selection strategy to form the agent's behavior decision model. The behavior policy generation module outputs specific instruction sequences based on the decision model. For mobile agents, it generates motion parameters including speed and turning angle; for robotic arm agents, it outputs joint angle control sequences. Finally, these are packaged into a structured decision policy data package and transmitted to the interactive execution unit.
[0071] In the visual feature extraction stage, the convolutional neural network adopts a variant of the VGG architecture, containing 13 convolutional layers with 5 convolutional blocks. Each convolutional layer is followed by batch normalization and ReLU activation functions. Max pooling is set between every two convolutional blocks, and the final fully connected layer outputs a 1024-dimensional feature vector. For multi-scale processing, the original image resolution is maintained at 1920×1080. The second layer downsamples to 960×540, and the third layer crops key regions while maintaining a resolution of 480×270. Feature maps from each layer are uniformly scaled to the same spatial size using bilinear interpolation before channel concatenation. In audio feature processing, the Mel filter bank is designed with 40 triangular filters, covering a frequency range of 80Hz to 8kHz. The 20-dimensional coefficients selected after discrete cosine transform include static features and their first-order differences. For sensor feature processing, the sliding window step size is set to 0.5 seconds. 12-dimensional statistics are calculated for time-domain features, and the amplitudes of the first 5 main frequency components are extracted for frequency-domain features, ultimately forming a 17-dimensional feature vector. In the dimensionality reduction operation before feature fusion, visual features are compressed to 256 dimensions using global average pooling, audio features are reduced to 15 dimensions by retaining 95% of the variance through principal component analysis, and sensor features remain at their original 17 dimensions. During weighted concatenation, feature-level weighting is used instead of modality-level weighting, with each feature dimension assigned a weight coefficient based on its classification importance in the training set. When constructing the environmental feature tensor, the 512-dimensional feature vector is reshaped into a 16×16×2 three-dimensional structure, the 3D convolutional kernel size is set to 3×3×2, and the number of output channels is expanded to 64 dimensions. The support vector machine classifier uses a radial basis function kernel, with the penalty parameter C set to 1.0 and the kernel function parameter gamma configured to 0.01. Multi-class classification is implemented using a one-to-many strategy. The scenario rule base contains 32 logical rules, such as "trigger a congestion label when the number of moving objects > 5 and the average speed < 0.5 m / s". In the reinforcement learning framework, the discount factor γ is set to 0.9, the initial exploration rate ε is 0.5 and decays over time, the Q-value table dimension is the number of environment states multiplied by the action space size, the environment states are discretized into 100 state intervals, and the action space includes 8 basic movement directions. The decision policy data package is encapsulated in JSON format, containing an action type field, a parameter array, and execution timestamp information.
[0072] In one embodiment, Example 2: When the decision-making strategy generation unit initiates the path planning process, it first receives the initial position coordinates and target position coordinates of the agent from the navigation system. The initial position is obtained through the Global Positioning System module, and the target position is specified by the task planning system. Environmental obstacle information comes from the object recognition results in the target environment state model. The model output includes static obstacle position coordinates and dynamic obstacle trajectory prediction data. The A* path planning algorithm is used to calculate the optimal path. During the algorithm implementation, the environment is gridded into 1m × 1m cells, and cells occupied by obstacles are marked as impassable areas. The heuristic function uses Manhattan distance calculation. An open list priority queue stores nodes to be expanded. In each iteration, the node with the smallest evaluation function value is selected for expansion. When the target position node is added to the closed list, the initial behavioral path is generated by backtracking. This path consists of a series of continuous path points, each containing three-dimensional coordinates and orientation angle information.
[0073] The real-time state monitoring system tracks the agent's motion state through multi-sensor fusion. The inertial measurement unit provides raw acceleration and angular velocity data, the visual odometry system processes camera images to calculate pose changes, and the wheel encoder records the number of wheel rotations. The state change data processing flow includes sensor data time synchronization, coordinate system transformation, and kinematic calculation, ultimately outputting an agent state vector containing position, velocity, and acceleration. The environmental dynamic change detection module compares two consecutive frames of LiDAR point cloud data, calculates the scene displacement vector through an iterative nearest-point algorithm, and identifies newly added or disappeared obstacles using visual background modeling technology. The dynamic change results are represented as a list of obstacle position changes.
[0074] The behavior stability database stores historical execution records, including path tracking error statistics and control parameter adjustment records under different road conditions. The adaptive data calculation module analyzes historical data; when the agent travels on gravel roads, the average lateral deviation is 0.2 meters, increasing to 0.35 meters on slippery roads. The control system response delay varies from 0.1 seconds to 0.3 seconds. These quantitative indicators constitute the adaptive dataset. The recognition reaction time measurement system records the end-to-end delay from sensor data acquisition to control command output, including a 15-millisecond image transmission delay, an 80-millisecond data processing time, and a 35-millisecond control algorithm calculation time, for a cumulative recognition reaction time of 130 milliseconds. This time parameter is used to calculate the behavior path correction lag.
[0075] The offset accumulation analysis system establishes a kinematic model to predict the agent's trajectory. The model input includes the current velocity vector, control delay parameters, and environmental change gradients. A Kalman filter algorithm fuses the predicted trajectory with actual pathpoint data, integrating the position deviation within a 130-millisecond time window to output the cumulative offset parameter. A preset threshold of 0.5 meters is set. When the cumulative offset falls below this value, a correction mode is triggered. The geometry calculation module establishes a reference coordinate system based on the direction of the line connecting the nearest pathpoints, decomposing the offset vector into lateral and longitudinal deviation components. When the lateral deviation exceeds 0.3 meters, a steering correction command is generated; when the longitudinal deviation exceeds 0.4 meters, the velocity parameters are adjusted. The correction data is encapsulated into a control command package. When the cumulative offset exceeds 0.5 meters, the path replanning process is initiated. The environmental dynamic change monitoring system constructs a real-time obstacle distribution map within 5 meters in front of the agent, with LiDAR scan data updating obstacle coordinates every 0.1 seconds. The environmental change trend analysis module calculates the obstacle movement velocity vector and establishes a position-time change matrix. The cubic spline interpolation algorithm generates dense sampling points within a 1.5-meter radius on both sides of the initial path to predict the obstacle distribution within the next 3 seconds. The behavior stability assessment module combines the road surface friction coefficient and the agent's maximum steering angular velocity parameters to calculate the feasible passage speed for different path segments. When a dynamic obstacle intersection risk is predicted for a path segment, avoidance waypoints are inserted into the path point sequence to generate an updated behavior path.
[0076] The first behavioral strategy employs an incremental adjustment scheme, updating control commands every 100 milliseconds. Lateral correction commands are converted into front wheel steering angle adjustments, and longitudinal correction commands into motor speed changes, with adjustments not exceeding 15% of the original values. During the second behavioral strategy implementation, the path replanning module retains 80% of the waypoints from the original path and inserts circular transition path segments in conflict areas. The radius of the circular arc is set to 1.2 meters based on the agent's minimum turning radius, and the maximum length of the transition path segment is limited to 3 meters. The updated pathpoint sequence undergoes smoothing to eliminate abrupt transitions, and Bezier curves are used to connect adjacent waypoints to ensure continuous path curvature. The final output behavioral strategy data package includes the pathpoint coordinate sequence, preset speed values for each point, and suggested steering angle values, which are transmitted to the underlying motion control system for execution.
[0077] In one embodiment, Example 3: The multimodal feature matching and analysis subunit initiates the feature spectrum construction process, receiving visual feature data, audio feature data, and sensor feature data from the input processing unit. The feature distribution calculation module estimates the probability density for each type of modality feature, using the kernel density method to generate a distribution surface in the feature space. Visual features are calculated in 256-dimensional space, audio features are analyzed in 15-dimensional space, and sensor features are processed in 17-dimensional space. Feature mean calculation uses the arithmetic mean method, averaging each dimension of the visual feature vector to generate 256 mean points, calculating 15 mean parameters for audio features, and outputting 17 mean indices for sensor features. Feature peak detection is achieved by finding local maxima on the distribution surface. Visual features detect 20-30 significant peak points, audio features identify 3-5 main peaks, and sensor features locate 4-6 key peaks. The final generated feature spectrum data structure includes three parts: distribution surface parameters, mean vector, and peak position coordinates.
[0078] The feature matching degree analysis module parses the feature spectrum data. The matching degree between visual and audio features is achieved by calculating the spatial distribution overlap rate, and the correlation coefficient is calculated by aligning the feature curves in the temporal dimension. Sensor and visual feature matching analysis uses a similarity measurement method to compare the distance distribution of feature peak positions. The generated feature matching parameter set contains three dimensions: inter-modal feature similarity is represented by a value in the range of 0-1, the feature consistency coefficient records the difference in distribution pattern, and the feature association strength quantifies the peak correspondence. The parameter set is stored in matrix form, with rows corresponding to different modal combinations and columns containing the three matching index values. The feature fusion parameter adjustment module dynamically configures the fusion weights according to the matching parameter set. When the visual-audio similarity is below 0.6, the visual feature weight increases from the baseline value of 0.5 to 0.55, and the audio weight decreases from 0.3 to 0.25. Sensor feature weights are adjusted according to the consistency coefficient; when the coefficient exceeds 0.8, the weight increases by 0.05. The fusion process adopts a feature-level weighting strategy, assigning an independent weight coefficient to each feature dimension. The optimized eigenvalues are calculated using the following formula:
[0079]
[0080] in: This represents the optimized eigenvalues. Indicates the first dimensional original eigenvalues, For the corresponding weighting coefficients, This represents the total dimension of the features. The weight coefficients are dynamically updated based on the importance of each feature in the classification task; importance is calculated by the feature dimension's contribution to the classification of the training set. The optimized multimodal feature data maintains a 512-dimensional dimension, but the feature distribution better reflects the characteristics of the current environment.
[0081] After receiving the feature matching parameter set, the decision-making strategy generation unit initiates a model reliability assessment. The confidence score of the environmental state model is calculated using a weighted method based on matching parameters, with similarity accounting for 60% of the weight, consistency coefficient for 30%, and association strength for 10%. When the confidence score falls below 0.7, a parameter adjustment mechanism is triggered: the object recognition module's recognition threshold decreases from 0.8 to 0.75, and the scenario judgment delay of the scenario label generation module increases by 0.2 seconds. Parameter adjustment employs a Bayesian update method, where the probability of object existence in the environmental state model is... Updated to:
[0082]
[0083] in: This indicates the probability that the updated object exists. This represents the reliability factor calculated from the feature matching parameter set. These are the original probability values. The updated environment state model output includes the adjusted object probability distribution and scenario confidence.
[0084] The agent behavior decision model reconstruction module reads the updated environmental state parameters and recalculates the state transition probability matrix using a dynamic programming algorithm. The action selection strategy in the decision model is updated using a value iteration method. The state space is discretized into 100 intervals, and the action space contains 8 basic actions. The state reward value in the value function update formula is adjusted according to the environmental state confidence level; for every 0.1 decrease in confidence, the reward value decreases by 5%. The reconstruction process generates a new Q-value table, overwriting the original decision model parameters. The final output updated decision strategy includes an action sequence and execution time sequence. A reliability flag is added to the action type field, and the parameter array is supplemented with confidence reference values. The strategy data packet is transmitted to the interactive execution unit via a dedicated interface. The transmission frequency is dynamically adjusted according to the rate of environmental change, increasing to a 10Hz update rate in dynamic environments and decreasing to 1Hz in static environments.
[0085] In the feature spectrum construction stage, the bandwidth parameter for kernel density estimation was set to 1.06 times the feature variance. A Gaussian kernel function was used for visual features, an Epanechnikov kernel for audio features, and a triangular kernel for sensor features. During feature matching analysis, the spatial distribution overlap rate was calculated using the Jaccard similarity coefficient, and the temporal correlation coefficient was calculated using the Pearson product-moment correlation coefficient. The distance distribution comparison in the similarity metric employed EarthMover's Distance algorithm. The feature fusion weight adjustment step size was set to 0.01, with the maximum adjustment range limited to ±0.1. The weight coefficients in the environmental state model confidence calculation formula were dynamically configured based on the modal data quality; when visual data was affected by illumination, the visual correlation weights were reduced by 20%. The reliability factor in the Bayesian update... The similarity coefficient is calculated using a linear combination of feature matching parameters, resulting in a similarity index of 0.6, a consistency coefficient of 0.3, and a correlation strength of 0.1. The discount factor for the dynamic programming algorithm is maintained at 0.9, the number of iterations is set to 100, and the convergence threshold is set to 0.001. The updated decision strategy data package adds a reliability field, containing an environmental state confidence value and a feature matching quality score. The underlying control system adjusts the action execution intensity based on the score; when the score is below 0.6, the action intensity is reduced by 20%.
[0086] In one embodiment, Example 4: The interactive execution unit initiates a workflow in a warehouse goods identification scenario. The standard data acquisition system scans typical warehouse goods to generate a basic dataset containing physical parameters for five common goods types. See Table 1 for the standard goods parameter table.
[0087] Table 1: Standard Goods Physical Parameters Table
[0088]
[0089] The data annotation process involves capturing 200 multi-angle images of each type of goods. The annotation system marks the bounding boxes of the goods in the images and associates them with category labels, while simultaneously recording ambient lighting conditions and shooting distance parameters. The object recognition model adopts an improved YOLOv5 architecture, replacing the backbone network with an EfficientNet-B3 structure, and adding a bidirectional feature fusion module to the feature pyramid. During the training phase, the input size is set to 640×640 pixels, the batch size is configured to 32, the optimizer uses the AdamW algorithm, and the initial learning rate is 0.001 adjusted by cosine annealing. The training process lasts for 150 epochs, with each epoch traversing the entire dataset three times.
[0090] When the agent performs recognition tasks in the shelving aisle, the stereo vision system acquires depth images at a rate of 15 frames per second, and the point cloud generation module converts the RGB-D data into 3D coordinate information. The real-time recognition process normalizes each frame, converting pixel values to the 0-1 range, and adjusts them to the model input size using bilinear interpolation. The forward inference process is completed on the embedded GPU, with a single frame processing time of 65 milliseconds. The output layer parses and generates bounding box coordinates, class confidence, and segmentation masks. The non-maximum suppression algorithm sets the intersection-union ratio threshold to 0.45 and the confidence threshold to 0.6, filtering overlapping detection results. The goods statistics module establishes a class counting dictionary, accumulating the count of goods that appear stably within 10 consecutive frames. The anomaly handling subunit deploys a status monitoring thread, and the inertial measurement unit provides real-time feedback of vibration intensity data. When a forklift passes by, causing shelving vibration, the monitoring system records X-axis acceleration fluctuations exceeding 0.5g. The data offset detection module analyzes the recognition results of 5 consecutive frames and finds a 15cm horizontal drift in the recognition position of the B-class wooden crate. The fluctuation impact assessment model calculates the statistical error caused by the positional offset. When the offset exceeds 20% of the cargo size, it is marked as a high-risk anomaly.
[0091] The system dynamically adjusts its response to abnormal events. The motion control system reduces the agent's speed from 0.8 m / s to 0.3 m / s and increases the image acquisition frequency from 15 Hz to 25 Hz. The data acquisition path is redesigned as a zigzag trajectory, and three additional sampling points are added in the abnormal region. A parameter adjustment record table generates a structured dataset containing timestamps, original parameter values, and new parameter values, stored in a circular buffer for the most recent 100 records. The model optimization module loads the latest abnormal dataset, updates the weights of the fully connected layers of the identification model using stochastic gradient descent with momentum, sets the momentum coefficient to 0.9, and adjusts the learning rate to one-tenth of the initial value. The online learning process lasts for 20 iterations, with 32 sets of sample data containing vibration interference input each time. The optimized model output layer adds an uncertainty estimation branch, automatically increasing the confidence threshold to 0.75 when the environmental vibration intensity exceeds a threshold. In standard data acquisition, Class A cardboard box samples include scenes with different stacking heights, while Class B wooden boxes collect surface conditions with varying degrees of newness. The annotation system uses a semi-automatic tool; after manual annotation of the first frame, the optical flow algorithm tracks the target position in subsequent frames. Data augmentation strategies during model training include random rotation (±15 degrees), brightness adjustment (±20%), and adding Gaussian noise (σ=0.05). During the inference phase, point cloud data processing uses voxel grid downsampling with a leaf size set to 0.02 meters. The anomaly monitoring system employs a multi-level early warning mechanism: a yellow alert is triggered when vibration lasts longer than 0.5 seconds, and a red alert is activated after 1 second. The path replanning algorithm uses A* search to generate new trajectories within feasible regions; the amplitude of the zigzag path is set to 30% of the channel width, and the wavelength is set to 1.5 meters based on the agent's minimum turning radius. During online learning, backbone network parameters are frozen, with only the fully connected layers of the detection head being fine-tuned; each update takes approximately 120 milliseconds. The optimized interactive result data includes a reliability marker field; when environmental interference exists, the output results include a confidence score for reference by subsequent processing modules.
[0092] In one embodiment, Example 5: The data transmission unit constructs a layered communication protocol stack in the intelligent inspection robot system. The physical layer uses the IEEE 802.11ac standard for wireless transmission, with the operating frequency configured in the 5GHz band and the channel bandwidth set to 80MHz. The data link layer implements MAC address filtering and frame check sequence verification, adding a 16-bit cyclic redundancy check code to each data packet. The network layer deploys an IPv6 protocol stack, using stateless addresses to automatically generate link-local addresses, and employs OLSR for multi-hop routing. The application layer defines a dedicated data exchange format, designing a compact binary message structure based on Protocol Buffers. During the data preprocessing stage of the interaction results, a data compression process is initiated. Differential coding technology is used for the cargo identification results, and placeholders are used to replace identical data fields between consecutive frames. Statistical table data uses a columnar storage format, with numerical fields compressed using Delta encoding and text fields converted using dictionary encoding. The compressed data stream undergoes forward error correction coding, adding Reed-Solomon error correction codes, which can correct up to 10% of data packet loss. The modulation module maps the binary data stream to orthogonal frequency division multiplexing subcarriers, using a 64-QAM modulation scheme where each symbol carries 6 bits of information, and the guard interval is set to one-quarter of the symbol length.
[0093] The signal encapsulation process adds layered protocol headers. The physical layer preamble contains 8 short training symbols and 2 long training symbols for signal synchronization and channel estimation. The data link layer frame header contains 6 bytes each for the destination MAC address and source MAC address, and a 2-byte frame type field identifies the protocol version. The network layer packet header contains a 40-byte IPv6 header, a flow label field indicating data priority, and a hop count limit of 64. The application layer message header contains a timestamp (8 bytes), data length (4 bytes), and checksum (4 bytes) field, for a total header length of 158 bytes. The signal transmission mechanism employs an adaptive rate control algorithm, monitoring the received signal strength indicator and signal-to-noise ratio (SNR) parameters in real time. When the RSSI is below -70dBm, it switches to 16-QAM modulation mode; when the SNR is below 20dB, it enables spatial-temporal block coding diversity. The transmission power is dynamically adjusted based on link quality, with a base power of 15dBm and a maximum increase to 20dBm. The external device connection establishes a TCP three-way handshake. The initial sequence number is generated using encrypted pseudo-random numbers, and the window size is initially 16KB, dynamically scaling according to network congestion. Decoding is performed synchronously at the receiving end. The physical layer uses preambles for symbol timing synchronization and carrier frequency offset compensation, and employs the least mean square algorithm for channel equalization. The error correction decoding module uses the Berkeley-Massey algorithm for Reed-Solomon decoding to recover the original data stream. The protocol parser peels off the protocol header layer by layer, verifying the check fields of each layer. The IPv6 header checks for version number and destination address matching, and the transport layer verifies the TCP checksum and sequence number continuity. The application layer parser deserializes Protocol Buffers messages to reconstruct the complete interaction result data structure.
[0094] In warehouse environments, the transmission system faces multipath fading challenges. The solution employs multi-antenna receive diversity technology, deploying two receiving antennas with a spacing of half a wavelength. Channel estimation uses a least-squares algorithm, with a pilot symbol interval of four data symbols. The data transmission cycle is linked to the agent's motion state; when the agent is moving at high speed (exceeding 1 m / s), the transmission interval is shortened to 100 milliseconds, and extended to 500 milliseconds when stationary. External device interfaces support multiple connection methods, including Wi-Fi direct connection, infrastructure mode via a router, and cellular network backup connection. A network switching algorithm monitors the current connection quality, automatically switching to a 4G LTE network when the Wi-Fi signal strength remains below -75 dBm for 3 consecutive seconds. Data encryption uses the AES-256-GCM algorithm, and key exchange is implemented using the elliptic curve Diffie-Hellman protocol, generating a unique symmetric encryption key for each session. The transmission quality monitoring system records the round-trip time, packet loss rate, and throughput for each data packet, triggering a retransmission mechanism when the loss rate exceeds 15% for five consecutive packets. The retransmission strategy employs selective retransmission, retransmitting only lost data segments rather than the entire data packet. Flow control uses a sliding window protocol, with the window size dynamically adjusted based on network conditions, expanding to a maximum of 64KB. The final decoded interactive decision data includes cargo identification results, environmental status assessments, and action recommendations, outputting to an external display system in structured JSON format.
[0095] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A multimodal input intelligent agent decision-making interaction system, characterized in that, include: A multimodal input acquisition unit is used to acquire input signal data of multiple modes from the target environment; The input processing unit is used to process the input signal data of the multiple modalities to obtain multimodal feature data; A decision strategy generation unit is used to generate a decision strategy for the intelligent agent based on the multimodal feature data; An interactive execution unit is used to execute interactive actions according to the decision-making strategy; The data transmission unit is used to transmit the interaction result data to an external device; The decision strategy generation unit determines the agent's behavior strategy based on the agent's behavior decision model, specifically as follows: Acquire the agent's initial position information and target position information, and determine obstacle information in the environment based on the target environment state model; Based on the path planning algorithm, path planning is performed on the initial position information and the target position information, and the obstacle information is used as planning constraints to output the initial behavior path of the agent; Real-time acquisition of agent state change data during the execution of the agent according to the initial behavior path, and determination of the environmental dynamics of the agent's real-time position based on the state change data; Acquire behavioral stability data of the intelligent agent, wherein the behavioral stability data includes data on the intelligent agent's ability to adapt to different environmental changes; Acquire the agent's reaction time data to dynamic changes in the environment, and determine the agent's behavior path correction lag based on the reaction time data. Based on the adaptive capability data and the behavioral path correction lag, the environmental dynamics of the agent's real-time position are analyzed to determine the cumulative behavioral path deviation of the agent within the recognition reaction time. If the cumulative amount of the behavior path offset is less than a preset value, the behavior offset direction and behavior offset distance executed by the agent according to the initial behavior path are determined based on the cumulative amount of the behavior path offset, and the correction direction and correction distance of the agent are determined based on the behavior offset direction and behavior offset distance to obtain correction data; Based on the correction data, the initial behavior path of the agent during real-time execution is corrected to obtain the first behavior strategy. If the cumulative offset of the behavior path is greater than a preset value, obtain the environmental dynamic change data of the real-time execution path of the agent, construct an environmental change map based on the environmental dynamic change data of the real-time execution path, and determine the environmental dynamic change trend in the target environment based on the environmental change map. Interpolation is performed on the dynamic change trend of the environment using an interpolation algorithm to determine the dynamic information of the environment within the preset range of the initial behavior path. The behavioral stability of the agent within the preset range of the initial behavior path is determined based on the dynamic information of the environment and the adaptability data. The initial behavior path segments with a cumulative offset greater than a preset value are adjusted based on the behavioral stability to obtain an updated behavior path. The agent executes the behavior according to the updated behavior path to obtain a second behavior strategy.
2. The multimodal input intelligent agent decision-making interaction system according to claim 1, characterized in that, The input processing unit processes the input signal data of the multiple modalities to obtain multimodal feature data, specifically as follows: Acquire visual modal input signal data, audio modal input signal data, and sensor modal input signal data of the target environment; Multi-scale feature extraction is performed on the visual modality input signal data to obtain visual feature data; Perform spectrum analysis on the audio modal input signal data to obtain audio feature data; Time series analysis is performed on the sensor modal input signal data to obtain sensor feature data; The visual feature data, audio feature data, and sensor feature data are fused to obtain multimodal feature data.
3. The multimodal input intelligent agent decision-making interaction system according to claim 2, characterized in that, The decision strategy generation unit generates a decision strategy for the agent based on the multimodal feature data, specifically: Environmental scenario analysis is performed based on the multimodal feature data. An environmental feature tensor is constructed through pattern recognition. Object recognition is performed on the multimodal feature data based on a machine learning classifier to generate a target environmental state model that includes environmental objects, scenario labels, and environmental states. The target environment state model is mapped to the agent's objective to construct an agent behavior decision model; The agent's behavior strategy is determined based on the agent's behavior decision-making model, thus obtaining the decision strategy.
4. The multimodal input intelligent agent decision-making interaction system according to claim 2, characterized in that, The input processing unit further includes a multimodal feature matching and analysis subunit, used for: Based on the multimodal feature data, the feature distribution, feature mean, and feature peak value of each mode are calculated to obtain the feature spectrum; The feature matching degree is determined based on the feature spectrum, and a feature matching parameter set is generated. The feature fusion parameters are adjusted based on the feature matching parameter set to optimize the multimodal feature data.
5. The multimodal input intelligent agent decision-making interaction system according to claim 4, characterized in that, The decision strategy generation unit further adjusts the agent's behavioral decision model based on the feature matching parameter set, specifically as follows: The reliability of the environmental state model is analyzed based on the feature matching parameter set, and the parameters in the environmental state model are adjusted accordingly. The agent's behavior decision-making model is reconstructed based on the adjusted environmental state model, and an updated decision-making strategy is generated.
6. The multimodal input intelligent agent decision-making interaction system according to claim 1, characterized in that, The interactive execution unit executes interactive actions according to the decision-making strategy, specifically: Obtain standard data of different objects in the target environment, and annotate the standard data to obtain annotated data; An object recognition model is constructed based on a deep learning model, and the labeled data is imported into the object recognition model for training. The agent obtains interaction data of the target environment according to the decision-making strategy, imports the interaction data into the trained object recognition model for object recognition, and counts the number of each object type to obtain interaction result data.
7. The multimodal input intelligent agent decision-making interaction system according to claim 6, characterized in that, The interactive execution unit further includes an exception handling subunit, used for: Real-time monitoring of state fluctuations and data offsets during the execution of interactive actions by intelligent agents, and analysis of the impact of fluctuation range on interaction results; Dynamically adjust the behavior parameters and data acquisition paths within the target range to generate an anomaly handling dataset; Adjust the object recognition model parameters based on the anomaly handling dataset to optimize the interaction result data.
8. The multimodal input intelligent agent decision-making interaction system according to claim 1, characterized in that, The data transmission unit transmits the interaction result data to an external device, specifically as follows: A hierarchical communication protocol stack is constructed based on wireless communication technology, and the interaction result data is modulated to construct a transmission signal. The transmission signal is sent to an external device according to the layered communication protocol stack, and the transmission signal is decoded to obtain interactive decision data of the target environment.
9. A multimodal input intelligent agent decision-making interaction method, characterized in that, Includes the following steps: The multimodal input acquisition unit acquires input signal data of multiple modes from the target environment. The input signal data of the multiple modalities are processed by the input processing unit to obtain multimodal feature data; The decision strategy generation unit generates a decision strategy for the intelligent agent based on the multimodal feature data; The interactive execution unit executes interactive actions according to the decision-making strategy; The interaction result data is transmitted to an external device based on the data transmission unit; The decision strategy generation unit determines the agent's behavior strategy based on the agent's behavior decision model, specifically as follows: Acquire the agent's initial position information and target position information, and determine obstacle information in the environment based on the target environment state model; Based on the path planning algorithm, path planning is performed on the initial position information and the target position information, and the obstacle information is used as planning constraints to output the initial behavior path of the agent; Real-time acquisition of agent state change data during the execution of the agent according to the initial behavior path, and determination of the environmental dynamics of the agent's real-time position based on the state change data; Acquire behavioral stability data of the intelligent agent, wherein the behavioral stability data includes data on the intelligent agent's ability to adapt to different environmental changes; Acquire the agent's reaction time data to dynamic changes in the environment, and determine the agent's behavior path correction lag based on the reaction time data. Based on the adaptive capability data and the behavioral path correction lag, the environmental dynamics of the agent's real-time position are analyzed to determine the cumulative behavioral path deviation of the agent within the recognition reaction time. If the cumulative amount of the behavior path offset is less than a preset value, the behavior offset direction and behavior offset distance executed by the agent according to the initial behavior path are determined based on the cumulative amount of the behavior path offset, and the correction direction and correction distance of the agent are determined based on the behavior offset direction and behavior offset distance to obtain correction data; Based on the correction data, the initial behavior path of the agent during real-time execution is corrected to obtain the first behavior strategy. If the cumulative offset of the behavior path is greater than a preset value, obtain the environmental dynamic change data of the real-time execution path of the agent, construct an environmental change map based on the environmental dynamic change data of the real-time execution path, and determine the environmental dynamic change trend in the target environment based on the environmental change map. Interpolation is performed on the dynamic change trend of the environment using an interpolation algorithm to determine the dynamic information of the environment within the preset range of the initial behavior path. The behavioral stability of the agent within the preset range of the initial behavior path is determined based on the dynamic information of the environment and the adaptability data. The initial behavior path segments with a cumulative offset greater than a preset value are adjusted based on the behavioral stability to obtain an updated behavior path. The agent executes the behavior according to the updated behavior path to obtain a second behavior strategy.
Citation Information
Patent Citations
Multi-task intelligent coordination execution method and system based on health care accompanying robot
CN120494765A
AI interactive data processing system based on multi-modal perception and dynamic decision
CN120688015A