Task execution strategy generation and adjustment method and device, equipment and medium
By acquiring and fusing visual, audio, and linguistic information, and combining reinforcement learning models to dynamically adjust task execution strategies, the problem of insufficient multi-dimensional information integration in existing technologies is solved, thereby improving the accuracy and adaptability of robots in complex environments.
Patent Information
- Application Number
- CN202511051733.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies cannot effectively integrate multi-dimensional information such as vision, audio, and language, resulting in low accuracy of robots in performing tasks in complex environments and poor environmental adaptability.
By acquiring visual, audio, and linguistic information from the environment, encoding and fusing it to generate comprehensive features, and then combining these features with a reinforcement learning model to dynamically adjust the task execution strategy.
It enables accurate optimization and flexible adaptation of task execution strategies in complex environments, thereby enhancing the comprehensive perception and autonomous decision-making capabilities of embodied intelligent devices.
Smart Images

Figure CN120951241A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for generating and adjusting task execution strategies. Background Technology
[0002] In the current development of embodied intelligence technology, traditional task processing methods still have many limitations in complex environments. In existing technologies, robot task execution heavily relies on single types of environmental information, such as visual information or data from a single sensor for task perception and decision-making. In this mode, when facing dynamic and changing real-world environments, robots lack comprehensive information integration and analysis capabilities, which can easily lead to delayed task response, strategy execution errors, and affect the overall task performance.
[0003] In the fintech sector, smart devices often rely solely on single image recognition or voice interaction modules when performing tasks such as lobby assistance, intelligent guidance, and risk assessment, lacking the ability to collaboratively process multimodal information. For example, when intelligent robots guide customers in bank branches, they cannot simultaneously combine visual information, ambient audio information, and customer voice commands for comprehensive analysis. This leads to a significant decrease in task efficiency and accuracy in situations involving dense customer traffic, complex noise levels, or ambiguous language communication, making it difficult to meet the high standards of intelligent service requirements.
[0004] In the healthcare sector, embodied intelligent devices are widely used in patient care, intelligent inspections, and medical assistance tasks. However, in existing technologies, the perception modules function relatively independently, and there is a lack of effective integration between visual perception, audio perception, and language command information processing, making it difficult to support intelligent collaboration of devices in medical environments. For example, when nursing robots perform rounds in hospitals, they cannot simultaneously perceive patients' verbal feedback, abnormal sounds and visual information from the surrounding environment, making it difficult to identify sudden patient conditions or environmental risks in a timely and accurate manner, thus affecting patient safety and nursing efficiency.
[0005] Overall, existing embodied intelligence task processing methods lack a unified information encoding and fusion mechanism when facing multimodal complex environments. They cannot effectively integrate multidimensional information such as vision, audio, and language, resulting in insufficient environmental understanding and limited strategy generation effects in actual task execution, which restricts the flexibility, accuracy, and intelligence level of task execution. Summary of the Invention
[0006] The main objective of this invention is to provide a method, apparatus, device, and storage medium for generating and adjusting task execution strategies, aiming to solve the technical problem that existing technologies cannot integrate visual, audio, and linguistic multimodal information for real-time task decision-making and adaptive adjustment in complex dynamic environments, resulting in low task execution accuracy and poor environmental adaptability.
[0007] To achieve the above objectives, the present invention provides a method for generating and adjusting task execution strategies, comprising:
[0008] Acquire visual information, audio information, and task-related language instructions from the environment;
[0009] The visual information, the audio information, and the language instruction information are encoded to obtain visual features, audio features, and language features, respectively.
[0010] By fusing the visual features, the audio features, and the language features, a comprehensive feature is generated.
[0011] An initial task execution strategy is generated based on the comprehensive features, and the actions corresponding to the initial task execution strategy are executed.
[0012] During the execution of the initial task execution strategy, a reinforcement learning model is used to dynamically adjust the initial task execution strategy based on real-time feedback information from the environment, resulting in an updated task execution strategy.
[0013] Furthermore, to achieve the above objectives, the present invention provides a task execution strategy generation and adjustment apparatus, comprising:
[0014] The multimodal perception module is used to acquire visual information, audio information, and task-related language instructions from the environment.
[0015] The feature encoding module is used to encode the visual information, the audio information, and the language instruction information to obtain visual features, audio features, and language features, respectively.
[0016] A multimodal fusion module is used to fuse the visual features, the audio features, and the language features to generate a comprehensive feature;
[0017] The task decision execution module is used to generate an initial task execution strategy based on the comprehensive features and execute the actions corresponding to the initial task execution strategy.
[0018] The strategy adaptive optimization module is used to dynamically adjust the initial task execution strategy based on real-time feedback information from the environment during the execution of the initial task execution strategy, using a reinforcement learning model to obtain an updated task execution strategy.
[0019] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a task execution strategy generation and adjustment program stored in the memory and executable on the processor, wherein when the task execution strategy generation and adjustment program is executed by the processor, it implements the steps of the task execution strategy generation and adjustment method as described above.
[0020] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a task execution strategy generation and adjustment program, wherein when the task execution strategy generation and adjustment program is executed by a processor, it implements the steps of the task execution strategy generation and adjustment method described above.
[0021] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as embodied intelligence, fintech, and healthcare. It discloses a method, apparatus, device, and medium for generating and adjusting task execution strategies, including: acquiring visual information, audio information, and task-related language instruction information from the environment; encoding the visual information, audio information, and language instruction information to obtain visual features, audio features, and language features respectively; fusing the visual features, audio features, and language features to generate comprehensive features; generating an initial task execution strategy based on the comprehensive features; and executing the actions corresponding to the initial task execution strategy. During the execution of the initial task execution strategy, based on real-time feedback information from the environment, a reinforcement learning model is used to dynamically adjust the initial task execution strategy to obtain an updated task execution strategy. This invention, through the joint acquisition and fusion of multimodal information and the dynamic adjustment of task execution strategies using a reinforcement learning model, achieves accurate optimization and flexible adaptation of task execution strategies in complex environments, improving the comprehensive perception and autonomous decision-making capabilities of embodied intelligent devices in dynamic environments. Attached Figure Description
[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0023] Figure 1 This is a schematic diagram of an application environment for the task execution strategy generation and adjustment method in one embodiment of the present invention;
[0024] Figure 2 This is a flowchart illustrating an embodiment of the task execution strategy generation and adjustment method of the present invention;
[0025] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the task execution strategy generation and adjustment device of the present invention;
[0026] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0027] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0028] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0029] The task execution strategy generation and adjustment method provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can obtain visual information, audio information, and task-related language instructions from the environment through the user terminal. It encodes the visual, audio, and language instructions to obtain visual, audio, and language features respectively. These features are then fused to generate a comprehensive feature. Based on this comprehensive feature, an initial task execution strategy is generated, and the corresponding actions are executed. During the execution of the initial task execution strategy, a reinforcement learning model is used to dynamically adjust the strategy based on real-time feedback from the environment, resulting in an updated strategy. This invention achieves accurate optimization and flexible adaptation of task execution strategies in complex environments through the joint acquisition and fusion of multimodal information, combined with dynamic adjustment of task execution strategies using a reinforcement learning model. This enhances the comprehensive perception and autonomous decision-making capabilities of embodied intelligent devices in dynamic environments. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0030] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the task execution strategy generation and adjustment method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0031] like Figure 2 As shown, the task execution strategy generation and adjustment method proposed in this invention includes the following steps:
[0032] S10: Acquire visual information, audio information, and task-related language instructions from the environment;
[0033] In this embodiment, acquiring visual information from the environment specifically refers to acquiring image data related to the environment in which the device is located through an imaging acquisition component configured on the embodied smart device. The visual information covers image data within the visible light range and may also include infrared, depth map, or 3D point cloud data. The imaging acquisition component can employ a high-definition camera based on a CMOS sensor or CCD sensor. The frame rate, resolution, and dynamic range of the acquired image information can be adjusted according to the specific application scenario. The acquisition method can be either single-frame static acquisition or continuous multi-frame dynamic acquisition.
[0034] Audio information in the environment refers to the raw audio data acquired by a sound pickup module within a spatial range and converted into electrical signals. The pickup module can be a single-channel microphone or a multi-channel microphone array structure. The array structure enables directional perception and localization of spatial sound sources. Audio information includes various data types such as ambient background noise, human voices, equipment operating sounds, and environmental event trigger sounds. Timestamp information can be recorded synchronously during audio information acquisition for subsequent multimodal information fusion processing.
[0035] Task-related language instruction information refers to natural language input information with human-computer interaction or task control attributes acquired in response to the device's current task objective. The source of language instruction information can be commands input by the user in real-time via voice, or text data acquired through text input terminals, remote communication systems, graphical interfaces, etc. Language instruction information can include control commands, descriptive information, environmental explanations, or other text data related to task planning and execution. The language instruction information acquisition module may include a voice receiving terminal, a text input module, a network information receiving unit, etc., with the specific implementation flexibly adjusted based on different hardware structures.
[0036] In practical applications, the acquisition of visual information, audio information and language command information requires data synchronization and time alignment. It is necessary to combine a unified clock signal or a high-precision synchronization module within the system to ensure the consistency of multi-source information in the time dimension, so as to support subsequent information fusion and decision-making processes.
[0037] During execution, image acquisition devices can be fixed to mobile robots, robotic arms, or other intelligent devices, and combined with real-time data from the device's position sensors, to achieve continuous and dynamic acquisition of environmental visual information. The image acquisition devices can employ wide-angle lenses, fisheye lenses, or zoom lenses to adapt to different spatial ranges and detail capture requirements.
[0038] During audio information acquisition, microphone arrays with beamforming capabilities can be used to enhance the signal strength of sound from specific directions and reduce interference from ambient background noise. For complex indoor reflective environments, reverberation cancellation and noise suppression modules can be configured to further optimize audio data quality.
[0039] For acquiring language command information, a real-time speech recognition system based on natural language processing algorithms can be used to quickly convert speech input into structured text information. For non-real-time control scenarios, task-related text information can be input through handheld terminals or remote management systems. The information format can include instruction sets, keyword combinations, or complete natural language sentences, which can be flexibly set according to the task execution environment and user needs.
[0040] While collecting multimodal information, the unified data management module inside the system records the timestamps, data sources, and spatial location parameters of various types of information in real time, ensuring high-precision matching of different types of information in spatial and temporal dimensions, and improving the accuracy and robustness of subsequent information processing and fusion.
[0041] Example Description: In the field of embodied intelligence, mobile inspection robots, during autonomous inspection tasks in complex industrial environments, utilize visual information acquisition to collect real-time data on the surrounding structure, equipment markings, warning signs, and personnel distribution, ensuring the system has a comprehensive understanding of the physical space structure and dynamic environmental changes. Through audio information acquisition, the robot can capture equipment operating sounds, abnormal noises, or environmental alerts, effectively identifying potential faults or risks and compensating for visual blind spots caused by obstructions, low light, or confined spaces. Combined with a language command information acquisition module, the system receives inspection commands, fault descriptions, or safety prompts input by on-site personnel via voice, and can also receive text task information from a remote control center via network access, enabling human-machine collaborative task execution. This real-time acquisition and synchronous processing of multi-source information allows embodied intelligence systems to perform autonomous path planning, dynamic decision adjustments, and efficient task execution based on comprehensive and accurate environmental and task information when facing complex and changing industrial environments, significantly improving the reliability and intelligence level of inspection operations.
[0042] In the field of healthcare, mobile nursing robots acquire visual information in the ward through high-definition cameras, monitor the patient's condition and surrounding environment in real time, and acquire voice information from patients or medical staff through microphone arrays. Combined with voice input of nursing needs or emergency call instructions, the device can accurately determine the patient's needs, adjust nursing behavior in a timely manner, and ensure the accuracy and safety of medical services.
[0043] In the fintech business, intelligent banking service robots use visual information to identify customer location and identity, and audio information to detect customer voice content. By combining the acquired language instructions, they can understand business needs and support face-to-face business consultation, voice interaction guidance, and identity verification processes, thereby improving service efficiency and customer experience.
[0044] This embodiment integrates visual information, audio information, and task-related language instructions from the environment, enabling the device to comprehensively perceive the environmental state and task requirements. This avoids incomplete perception caused by isolated information processing, improves the accuracy of environmental understanding, and enhances the system's autonomous decision-making and response capabilities in dynamic and complex environments.
[0045] S20, the visual information, the audio information, and the language instruction information are encoded to obtain visual features, audio features, and language features respectively;
[0046] In this embodiment, visual information typically refers to environmental image data acquired through an imaging device, including but not limited to color images, depth images, or multispectral images. The data format of visual information can be a standard two-dimensional image matrix, or it can include depth information or point cloud data, depending on the type and configuration of the imaging device. Audio information refers to environmental sound signals acquired through acoustic sensors, including various sound source data such as environmental noise, speech information, and mechanical operating sounds, and is typically stored in the form of audio waveforms or spectrograms. Language instruction information refers to task-related text or voice instructions, which can be either verbal instructions input by the user or structured text information transmitted through an interface, including operational requirements, scene descriptions, or interactive commands.
[0047] Encoding transforms raw information into structured, low-dimensional, and highly expressive feature data, facilitating subsequent computation and fusion processing. Visual information encoding utilizes structures such as convolutional neural networks to extract spatial and semantic features from images, outputting fixed-dimensional visual feature vectors. This process includes specific steps such as feature extraction, scale normalization, and dimensionality compression. Audio information encoding employs spectral analysis, temporal modeling, and feature extraction to generate audio features reflecting environmental sound patterns, speech components, or anomalous sound source characteristics. The encoding of language instruction information first performs text standardization and semantic parsing, then combines contextual information to construct semantic embedding representations, generating structured language feature representations to support downstream multimodal fusion and task decision-making.
[0048] Visual information encoding can employ convolutional networks or visual transform structures to perform multi-scale convolution operations on the original image sequence, extracting local and global features. Then, pooling or projection operations are used to reduce the data dimensionality, yielding visual feature vectors. Audio information encoding can be based on short-time Fourier transform, Mel-frequency cepstral coefficients, or deep audio coding models to perform noise reduction, feature extraction, and temporal modeling on the acquired audio signals, outputting stable audio feature representations. Language command information encoding can combine text segmentation, embedding mapping, and multi-head attention mechanisms to construct high-dimensional semantic features, outputting language features through semantic compression and structural optimization. In different implementations, visual information sources can be monocular, binocular, or multimodal imaging systems; audio information sources can be omnidirectional microphones, array microphones, or directional pickup devices; and language command information can originate from a local voice input module or a remote text transmission system. Specific configurations can be flexibly selected and adapted according to application requirements.
[0049] Example Description: In the field of embodied intelligence, when an indoor inspection robot performs tasks in complex scenarios, the system acquires visual information through image sensors to perceive the indoor structure, equipment layout, and obstacle distribution in real time. It collects environmental audio information through audio sensors to monitor equipment operation status or abnormal sounds. It obtains task-related language command information through a language module and receives user instructions or environmental announcements. It performs encoding operations on the above information to generate visual features, audio features, and language features in a unified format. This ensures that the robot has multimodal information understanding capabilities during inspection, path planning, and task decision-making, thereby improving its autonomous obstacle avoidance and dynamic adjustment capabilities.
[0050] In the field of healthcare, ward service robots collect visual information through camera devices to identify the patient's location and bed environment during patient care, collect audio information through sound pickup devices to perceive physiological audio characteristics such as the patient's breathing and coughing sounds, and obtain language command information through voice interaction modules to receive patient needs or nursing instructions. The acquired information is encoded to generate visual features, audio features, and language features, ensuring that the robot has the ability to perceive patient status, understand speech, and assist in the execution of nursing tasks, thereby improving the level of intelligent services in medical scenarios.
[0051] In the field of fintech, bank smart guidance devices acquire visual information about customers and their environment through cameras during customer service, monitor customer behavior and counter layout, collect environmental audio information through a sound pickup system, monitor abnormal sound sources or peak business alerts in real time, and acquire language command information through a voice system to receive customer inquiries, business requests or security warnings. They encode multi-source information to generate visual features, audio features and language features, assisting the system in customer identification, business guidance and emergency response, and enhancing the intelligent interaction and service capabilities in the financial business environment.
[0052] This embodiment enables the system to achieve structured expression of multi-source information by independently encoding visual information, audio information, and language command information. This improves information processing efficiency and fusion accuracy, enhances the ability to fully perceive complex environments, and ensures the stability and reliability of downstream multimodal fusion and task strategy generation.
[0053] S30, the visual features, the audio features, and the language features are fused to generate a comprehensive feature;
[0054] In this embodiment, visual features refer to data representations with spatial structure and image feature description capabilities obtained through processing visual information. These visual features typically include color distribution, shape contours, spatial coordinates, edge textures, and image depth information, which can be obtained through convolution operations, image pyramid processing, feature point extraction, and other methods. The input data for visual features comes from digital images acquired by image sensors. These digital images have undergone illumination correction, feature extraction, and compression processing during the encoding stage, and possess numerical features of fixed dimensions.
[0055] Audio features refer to the temporal feature vectors obtained by denoising, extracting temporal features, and modeling temporal sequences from audio information. These features include the spectral distribution, temporal variations, and energy intensity of the environmental audio signal. The input data for audio features comes from digital audio collected by an audio sensor array, which is then processed through multi-channel synchronization, temporal alignment, and spectral analysis to form a standardized feature format.
[0056] Language features refer to semantic feature representations generated after processing language instruction information through word segmentation, context association, and semantic encoding. The input data for language features comes from text information obtained by transcribing speech instructions. After natural language processing, embedding vector mapping and semantic association operations are performed to obtain a fixed-length vector representation that can be processed subsequently.
[0057] Visual features, audio features, and linguistic features may all exhibit inconsistencies in dimensionality, time step mismatches, and differences in data volume during the input phase. Therefore, it is necessary to perform spatial dimensionality normalization on visual features separately, ensuring that the numerical distribution of each channel of the visual features meets the unified feature space constraints, and guaranteeing that the computational stability of visual features is not affected by scale differences during subsequent processing. Audio features undergo temporal alignment processing, using interpolation or pruning to ensure that the time step of the audio features is consistent with the target time window, guaranteeing that audio information has a synchronous temporal reference with visual and linguistic instruction information within the processing cycle. Linguistic features undergo semantic attention weighting processing, adjusting the weight ratio of each word in the fusion process by analyzing the semantic contribution of different lexical units in the linguistic features, ensuring that the semantic weight of key instructions or important descriptions in the linguistic instruction information is fully preserved in subsequent processing.
[0058] The visual, audio, and linguistic features, processed separately above, are combined into a single data input. A feature concatenation operation connects the three feature sequences along either the feature dimension or the time dimension to generate a concatenated feature. This concatenated feature contains the original data structure of all modalities. The concatenated feature is then input into the cross-modal association modeling process. Through matrix operations, interactive attention operations, or multi-channel information mapping, the association relationships between the visual, audio, and linguistic features are established, generating multimodal association features that characterize the semantic interaction and temporal coupling between modalities. The multimodal association features are then compressed using linear mapping or low-dimensional projection to convert them into a unified format of comprehensive features. This comprehensive feature serves as the direct input for subsequent task execution strategy generation, possessing the characteristics of including all modal information, balanced semantic weights, and a unified structural dimension.
[0059] In practical applications, spatial normalization of visual features can be achieved by adjusting the mean and variance of each channel through batch normalization or layer normalization operations, or by unifying the scale of visual features through unit vector normalization. Temporal alignment of audio features can be achieved through linear interpolation, resampling, or dynamic time warping to align the audio timeline with the visual sampling period, or by adjusting the audio data time frame synchronously through sliding window processing. Semantic attention weighting of language features can be based on self-attention mechanisms, or by adjusting the feature values of key words through word frequency statistics, or by assigning higher attention weights to domain-specific words in language features through external knowledge graphs. During feature concatenation, visual, audio, and language features can be directly connected along a fixed-dimensional order, or the length of each feature can be adjusted through structural mapping before concatenation. Cross-modal association modeling can be achieved through multi-head interactive attention, multi-layer fusion networks, or tensor-based methods, or by constructing intermodal mapping relationships based on matrix factorization or low-rank approximation. Feature compression can be achieved by using fully connected mapping, reducing feature dimensionality through autoencoder compression, or selecting the main feature dimension through singular value decomposition to output comprehensive features.
[0060] Example Description: In the field of embodied intelligence, indoor guidance robots perform path guidance tasks in complex office environments. By normalizing visual features, temporally aligning audio features, and semantically weighting language features, the robot can simultaneously understand user language commands, environmental audio cues, and visual obstacle information. The comprehensive features generated through feature splicing and cross-modal modeling provide an information basis for guidance path planning, helping the robot dynamically adjust the guidance path, avoid collisions, and respond to user requests in real time.
[0061] In the field of healthcare, nursing robots collect multimodal information in the patient's living environment. They achieve accurate recognition of bed layout and patient movements by processing visual features through spatial normalization, synchronously monitor changes in patient breathing and equipment alarm sounds by processing audio features through temporal alignment, recognize emergency commands issued by patients by processing language features through semantic attention weighting, and generate comprehensive features through multimodal fusion to support the robot in quickly determining the priority of nursing tasks and adjusting nursing strategies.
[0062] In the field of fintech, smart service equipment in bank branches uses spatial normalization to process visual features to identify customer behavior and queuing status, temporal alignment to process audio features to identify sudden sound events, semantic attention weighting to process language features to understand customer service requests and security broadcast content, and the integrated features support the service equipment to dynamically adjust customer guidance paths and service priorities during peak periods, thereby improving service efficiency and customer experience in the financial business environment.
[0063] This embodiment effectively unifies the data scale, time window, and semantic weight of different modal features by using spatial normalization of visual features, temporal alignment of audio features, and semantic attention weighting of linguistic features. It establishes a deep interactive relationship between visual, audio, and linguistic features through feature splicing and cross-modal association modeling. Furthermore, it generates comprehensive features containing multimodal information through feature compression, which helps the subsequent task execution strategy generation process to fully understand and accurately analyze environmental information, thereby improving the decision-making accuracy and response flexibility of task execution.
[0064] S40, Generate an initial task execution strategy based on the comprehensive features, and execute the actions corresponding to the initial task execution strategy;
[0065] In this embodiment, the comprehensive feature is a multimodal information set formed by fusing visual, audio, and linguistic features. This comprehensive feature has a unified data structure, expressing the spatial information, temporal changes, and semantic descriptions of the environment. The process of generating an initial task execution strategy refers to inferring an appropriate operation plan for the current scenario based on the environmental information reflected by the comprehensive feature through a strategy generation model. The initial task execution strategy includes specific action arrangements and execution flows. The comprehensive feature is input to the task planner, which can be a deep neural network, decision tree model, or other structure that infers task plans based on environmental features. The task planner analyzes the spatial relationships, temporal information, and semantic instructions in the comprehensive feature, combined with the task objective and device status, to generate an initial task execution strategy with execution logic. The initial task execution strategy is an operation plan oriented towards a specific device, typically including a continuous sequence of actions, execution order, temporal arrangement, and parameter configuration. This strategy is parsed and converted into action instructions that the device can recognize. The action instructions are represented as a structured data set, including target location, movement path, operation mode, and control parameters. After receiving an action command, the device converts it into low-level control parameters, including position parameters, speed parameters, attitude adjustment parameters, and other parameters that affect the device's execution state. These control parameters drive the execution device, such as a robotic arm, mobile platform, or intelligent execution module, to complete physical operations in the environment according to the action requirements of the initial task execution strategy, thus achieving the initial execution of the task objective.
[0066] In practical applications, the comprehensive features are output as a fixed-dimensional feature vector through a multimodal fusion module, which is then input to the task planner. The task planner can employ a policy reasoning structure based on a fully connected network, combining historical data and task templates to output an initial task execution strategy adapted to the current environmental state. The structure of the task execution strategy can employ sequence labeling, a tree structure, or a state machine model. The strategy contains multiple actions to be executed, with actions including spatial location, operation mode, and execution parameters. The process of parsing the initial task execution strategy involves a sequence decoder or structure parsing module, converting the strategy content into device action instructions. These action instructions encapsulate control targets and parameter information. Control parameters can be generated through a dynamics calculation module, inverse kinematics algorithm, or direct mapping, and are then sent to the execution device. The execution device can be a multi-degree-of-freedom robotic arm, an automated navigation mobile chassis, an intelligent interactive device, or other devices with execution functions. The execution device adjusts its own state according to the control parameters to complete actions such as target position movement, attitude adjustment, and execution of operation instructions.
[0067] Example: In the field of embodied intelligence, intelligent handling robots in a warehouse environment generate an initial task execution strategy by inputting comprehensive features into a task planner. The strategy includes handling paths, obstacle avoidance arrangements, and grasping operations. The robot analyzes the strategy and generates control parameters to drive the robotic arm and mobile platform to work together to execute the handling task, ensuring the accuracy and efficiency of warehouse operations.
[0068] In the healthcare business, ward inspection robots generate initial task execution strategies based on comprehensive characteristics. The strategies include inspection paths, equipment inspection sequences, and anomaly monitoring actions. Based on the strategies, the robots generate action instructions and control parameters to perform bed inspections, instrument readings, and environmental detection tasks, ensuring the safety and responsiveness of the medical environment.
[0069] In the field of fintech business, bank intelligent guidance devices generate initial task execution strategies based on comprehensive features. The strategies include customer guidance paths, information display locations, and interactive actions. The devices analyze the strategies to generate control parameters, driving the display screen, indicator lights, and mechanical components to work together, thereby improving the efficiency and accuracy of customer guidance and information delivery in financial service scenarios.
[0070] This embodiment generates an initial task execution strategy based on comprehensive features and executes corresponding actions, thereby effectively supporting task decision-making with multimodal environmental information and improving the relevance and environmental adaptability of the task execution strategy. The generation of action commands and the dynamic mapping of equipment control parameters ensure the stability and accuracy of the task execution process, contributing to equipment coordination and efficient operation driven by multimodal information in complex environments.
[0071] S50, during the execution of the initial task execution strategy, the initial task execution strategy is dynamically adjusted using a reinforcement learning model based on the real-time feedback information from the environment to obtain an updated task execution strategy.
[0072] In this embodiment, during the execution of the initial task execution strategy, the execution device gradually completes various operations according to the previously generated task execution strategy. During the operation, the environmental state changes in real time, requiring real-time acquisition of environmental feedback information. Real-time environmental feedback information refers to the state change data acquired by sensors, vision systems, audio systems, or other information acquisition modules in the environment after the device performs an action. This data may include location coordinates, spatial layout, dynamic obstacles, audio events, semantic changes, or other parameter information reflecting the environmental state. The real-time feedback information is input into the reinforcement learning model. The reinforcement learning model uses structured data of state, action, and reward to dynamically evaluate the rationality and environmental adaptability of the current task execution strategy.
[0073] Reinforcement learning models are decision-making models that continuously optimize policy output based on interactive feedback. Model structures can include value function-based methods, policy gradient methods, Actor-Critic structures, or other reinforcement learning algorithm frameworks. The model calculates the deviation between the task objective and the current environment state based on environmental feedback and the task state under the current policy, generating an immediate reward signal. This immediate reward signal is used to evaluate the effectiveness of the current policy in the environment; a higher reward signal indicates a more adaptable policy, while a lower signal suggests the policy needs optimization. Based on the reward signal, reinforcement learning models employ a parameter update mechanism to optimize the parameter structure of the policy network. The updated policy exhibits higher environmental adaptability and execution efficiency.
[0074] By dynamically adjusting, an updated task execution strategy is generated. The updated strategy reflects the ability to respond instantly to changes in the environment. Compared with the initial strategy, the updated strategy optimizes the configuration for real-time changes in the environment, including adjusting the action sequence, optimizing execution parameters, or updating path planning, to ensure the continuity, accuracy, and safety of the task execution process.
[0075] In the specific implementation process, real-time feedback information from the environment can be collected through a multi-source sensor network, including LiDAR, vision cameras, depth sensors, microphone arrays, etc. The collected data is structured into an environmental state vector through a preprocessing module. The difference between the environmental state vector and the current task objective is calculated. The calculation method can be Euclidean distance, cosine similarity, spatial overlap ratio, or other metrics to obtain the task state deviation parameter.
[0076] The reinforcement learning model generates an instant reward signal based on task state deviation parameters. This reward signal is dynamically adjusted according to the deviation magnitude, a custom reward function, or historical experience data, forming the basis for policy optimization. Parameter updates in the reinforcement learning model employ gradient descent, policy iteration, proximal policy optimization, or other algorithmic mechanisms to dynamically optimize the policy network structure. The updated task execution policy is output through the policy network, containing new action sequences, parameter adjustments, and operational arrangements. The policy is then distributed to the execution device through action command hierarchies. The execution device adjusts its actions according to the updated policy to adapt to environmental changes.
[0077] Example: In the field of embodied intelligence, intelligent inspection robots perform inspection tasks in complex factory areas. During the process, they acquire environmental feedback information in real time, such as the location of obstacles and changes in equipment status. The reinforcement learning model dynamically adjusts the inspection path and operation strategy based on the feedback information to ensure that the inspection task is executed efficiently and safely.
[0078] In the field of healthcare, rehabilitation assistive robots, during patient rehabilitation training, acquire real-time information on patient posture, muscle activity, and changes in environmental factors. This information is then used to enhance the learning model, dynamically optimize training strategies and assistive movements, thereby improving the personalization and effectiveness of rehabilitation training.
[0079] In the fintech business, intelligent security equipment performs security patrols at bank branches. By combining real-time environmental feedback information, such as crowd density, abnormal sounds, or visual anomalies, reinforcement learning models dynamically adjust patrol strategies and response actions, enhancing the real-time nature and intelligence of security measures.
[0080] This embodiment utilizes a reinforcement learning model to dynamically adjust task execution strategies based on real-time feedback information from the environment, achieving adaptive optimization of the task execution process. The reinforcement learning mechanism enhances the real-time adaptability of task execution strategies in complex and dynamic environments, effectively addressing issues such as strategy failure, action errors, or task interruptions caused by environmental changes, thereby improving the stability, continuity, and intelligence level of task execution.
[0081] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as embodied intelligence, fintech, and healthcare. It discloses a method, apparatus, device, and medium for generating and adjusting task execution strategies, including: acquiring visual information, audio information, and task-related language instruction information from the environment; encoding the visual information, audio information, and language instruction information to obtain visual features, audio features, and language features respectively; fusing the visual features, audio features, and language features to generate comprehensive features; generating an initial task execution strategy based on the comprehensive features; and executing the actions corresponding to the initial task execution strategy. During the execution of the initial task execution strategy, based on real-time feedback information from the environment, a reinforcement learning model is used to dynamically adjust the initial task execution strategy to obtain an updated task execution strategy. This invention, through the joint acquisition and fusion of multimodal information and the dynamic adjustment of task execution strategies using a reinforcement learning model, achieves accurate optimization and flexible adaptation of task execution strategies in complex environments, improving the comprehensive perception and autonomous decision-making capabilities of embodied intelligent devices in dynamic environments.
[0082] In one embodiment, step S10 above includes:
[0083] S101, acquires ambient optical signals through an optical sensor and converts the ambient optical signals into visual information in digital image format;
[0084] S102, Acquire ambient sound wave signals through an acoustic sensor array, and convert the ambient sound wave signals into digital audio signal audio information;
[0085] S103: Obtain user voice commands through a voice receiving device and convert the user voice commands into text-formatted language command information.
[0086] In this embodiment, visual information, audio information, and task-related language instructions in the environment are acquired, specifically including independent acquisition processes for multiple information categories. First, environmental optical signals are acquired using optical sensors. These sensors can be CCD image sensors, CMOS image sensors, or other image acquisition devices with photoelectric conversion capabilities. The sensor installation position is configured according to the environmental layout to ensure complete coverage of optical information within the target area. The environmental optical signal is a continuous light intensity signal formed after external light is reflected and scattered. During signal acquisition, parameter optimization measures such as filtering, exposure control, and resolution adjustment are combined to improve signal quality and image clarity. The acquired environmental optical signal undergoes format conversion through a signal conversion module. The conversion process includes analog-to-digital conversion, signal encoding, color space mapping, and data format encapsulation, ultimately generating visual information in a digital image format that meets image processing requirements. This visual information serves as input data for subsequent processing, preserving environmental spatial structure, object shape, and texture information.
[0087] Secondly, ambient sound signals are acquired through an acoustic sensor array, which consists of multiple distributed microphones. The spatial arrangement of the microphones meets the requirements for beamforming and sound source localization. Ambient sound signals include various audio information such as ambient noise, conversations, and equipment operation sounds. The sensor array synchronously acquires these signals. The raw sound signals undergo noise reduction, echo suppression, signal separation, and gain control through a signal processing module, improving the signal-to-noise ratio and localization accuracy. The processed sound signals are then converted into digital audio signals through analog-to-digital conversion, encoding compression, and format encapsulation. This audio information possesses temporal continuity, spectral distribution, and intensity characteristics, serving as an important input source for multimodal information fusion.
[0088] Finally, user voice commands are acquired through a voice receiving device, which can be a near-field microphone, an array-type voice input device, or other hardware systems with voice acquisition capabilities. User voice commands are acquired in real time by the voice receiving device. The voice signal undergoes noise reduction, gain adjustment, speech segmentation, and preprocessing to reduce environmental noise interference and voice distortion. The acquired voice signal is input to the speech recognition module, which internally performs acoustic model inference, language model matching, and text decoding operations, converting the continuous voice signal into structured, standardized text-formatted language command information. This language command information retains the user's expressed intent, semantic content, and command logic, facilitating subsequent semantic understanding and task decision-making.
[0089] Through the above operations, the independent acquisition and standardized output of visual information, audio information and language command information are fully realized, ensuring that the input data source for multimodal information fusion processing is accurate, clear and of high quality.
[0090] This embodiment converts optical signals, sound signals, and voice signals in the environment into visual information in digital image format, audio information in digital audio signal format, and language command information in text format, respectively. This ensures that the sources of multimodal information are comprehensive, the data format is unified, and the processing structure is clear. It improves the adaptability and reliability of multi-source information in the downstream fusion, perception, and decision-making process, and enhances the device's perception capability and multi-information collaborative processing level in complex environments.
[0091] In one embodiment, step S20 above includes:
[0092] S201, Perform illumination correction processing on the visual information to generate illumination-corrected visual information;
[0093] S202, extract the spatial and semantic features of the illumination correction visual information, and perform convolutional compression processing on the spatial and semantic features to generate visual features;
[0094] S203, perform noise reduction processing on the audio information to generate noise-reduced audio information;
[0095] S204, extract the time-domain and frequency-domain features of the noise-reduced audio information, and perform time-series modeling processing on the time-domain and frequency-domain features to generate audio features;
[0096] S205, perform word segmentation processing on the language instruction information to generate a word sequence;
[0097] S206, perform context association processing on the word sequence to generate a context association representation, and perform semantic encoding processing on the context association representation to generate language features.
[0098] In this embodiment, visual information, audio information, and language command information are encoded to obtain visual features, audio features, and language features, respectively. This specifically includes multi-level feature extraction and encoding operations. First, illumination correction processing is performed on the visual information, which is digital image data generated based on ambient optical signals. Illumination correction processing addresses issues such as uneven ambient lighting, local shadows, or brightness variations by employing techniques such as image enhancement, histogram equalization, or illumination estimation compensation to correct the overall brightness and contrast distribution of the image, improving the stability and usability of the image under various lighting conditions, and generating illumination-corrected visual information. This illumination-corrected visual information retains key visual elements such as environmental spatial structure, object boundaries, and color texture, serving as the input basis for subsequent visual feature extraction.
[0099] Building upon this foundation, spatial and semantic features of illumination-corrected visual information are extracted. Spatial features reflect the geometric structure, location distribution, and edge information in the image, while semantic features reflect the object categories, scene attributes, and semantic identifiers present in the image. The extraction process incorporates a multi-layered convolutional neural network, with the front-end network extracting low-level spatial structure information and the deep network capturing high-level semantic information. Convolutional compression is then applied to both spatial and semantic features, including feature downsampling, channel integration, and feature compression operations. This reduces redundant information and computational overhead, while preserving highly expressive and representative visual features, resulting in a structured visual feature output that facilitates multimodal fusion and downstream task processing.
[0100] Noise reduction processing is performed on the audio information, which is digital audio data generated based on the conversion of ambient sound wave signals. The noise reduction process includes techniques such as temporal filtering, frequency domain enhancement, and deep denoising network inference to eliminate background noise, equipment operating noise, and non-target sound interference, while retaining effective information in speech, key sound sources, and ambient sounds to generate noise-reduced audio information. The noise-reduced audio information serves as input data for time-series modeling, possessing temporal continuity and spectral clarity.
[0101] The process extracts temporal and frequency domain features from the noise-reduced audio signal. Temporal features include amplitude variations, energy distribution, and temporal structure, while frequency domain features include spectral envelope, band energy, and acoustic signature characteristics. Extraction methods can include short-time Fourier transform, Mel-frequency cepstral coefficient extraction, and multi-scale feature analysis. Temporal modeling is then applied to these features. This modeling utilizes recurrent neural networks, temporal convolutional networks, or self-attention mechanisms to capture the temporal dependence, sequence characteristics, and global structure of the audio signal, enhancing its adaptability to dynamic changes in complex environments and ultimately outputting the audio features.
[0102] The language instruction information, which is in text format, undergoes lexical segmentation. Lexical segmentation involves splitting continuous text into semantic units based on pre-trained word segmentation models or adaptive dictionary rules, generating lexical sequences that accurately represent the structure, semantic boundaries, and lexical combinations within the language instruction information. Contextual association processing is then performed on the lexical sequences. This process uses multi-layered bidirectional neural networks or self-attention networks to capture the dependencies, long-distance semantic connections, and syntactic structures within the lexical sequences, generating contextual representations. Further semantic encoding is then performed, extracting deeper semantic meanings, sentiment tendencies, and semantic consistency from the contextual information to generate structured and expressive language features, facilitating multimodal information fusion and task decision-making.
[0103] This embodiment extracts and encodes multi-level features from visual information, audio information, and language instructions to obtain visual, audio, and language features that express the spatial structure of the environment, dynamic audio information, and semantic instructions. This enhances the expressive power, robustness, and adaptability of multimodal environmental information, providing clearly structured and comprehensive input data for multi-information fusion and task decision-making, thereby improving the system's understanding of complex environments and task execution efficiency.
[0104] In one embodiment, step S30 above includes:
[0105] S301, Perform spatial dimension normalization processing on the visual features to generate normalized visual features;
[0106] S302, Perform temporal alignment processing on the audio features to generate aligned audio features;
[0107] S303, Perform semantic attention weighting processing on the language features to generate weighted language features;
[0108] S304, the normalized visual features, the aligned audio features, and the weighted language features are concatenated to generate concatenated features;
[0109] S305, Perform cross-modal association modeling on the spliced features to generate multimodal association features;
[0110] S306, Perform feature compression processing on the multimodal association features to generate comprehensive features.
[0111] In this embodiment, visual features, audio features, and linguistic features are fused to generate comprehensive features. This process includes multi-step feature standardization, alignment, weight adjustment, structural integration, and correlation modeling. First, spatial dimension normalization is performed on the visual features. Visual features typically possess multi-dimensional spatial information, involving the geometric structure of the image, object distribution, and local details. Spatial dimension normalization uses techniques such as scale adjustment, numerical normalization, and spatial mapping transformation to unify visual features from different sources and scales to a standard spatial expression range, eliminating scale inconsistencies caused by differences in viewpoint, distance, or imaging conditions, thus generating normalized visual features. These normalized visual features possess the characteristics of a unified spatial expression structure and stable numerical distribution, facilitating fusion with other modal information.
[0112] Temporal alignment is performed on audio features, which reflect the time-series structure of acoustic signals in the environment. Temporal alignment unifies audio features with different sampling times or temporal deviations under a standard time axis through time window adjustment, dynamic time warping, or synchronization operations based on reference signals. This ensures that audio information maintains temporal consistency during multimodal information fusion, generates aligned audio features, and improves the temporal coordination and semantic correspondence of audio data in fusion computation.
[0113] Semantic attention weighting is applied to language features, which are high-dimensional feature representations after semantic encoding, including word semantics, contextual information, and language structural relationships. Semantic attention weighting dynamically adjusts the weight distribution of each language feature according to the importance of different language segments through an attention mechanism, enhancing the ability to express key semantic regions and key information of instructions, reducing interference from irrelevant or redundant language content, generating weighted language features, and improving the accuracy and effectiveness of semantic expression of language instruction information in multimodal fusion.
[0114] Normalized visual features, aligned audio features, and weighted language features are concatenated. The feature concatenation operation integrates multimodal features from different sources and in different forms into a unified data structure through dimensional expansion, feature stacking, and spatial alignment, generating concatenated features. The concatenated features retain the independence of each modal information in structure, while also possessing the basic structure for joint expression of cross-modal information, providing data support for subsequent cross-modal association modeling.
[0115] Cross-modal association modeling is performed on the spliced features. Cross-modal association modeling captures the inherent correlation, semantic coupling features and spatiotemporal coordination patterns between visual, audio and language instruction information through multi-layer neural networks, interactive attention structures or graph structure analysis methods. It strengthens the interactive expression of different modal information in spatial structure, time series and semantic content, and generates multimodal association features. The multimodal association features have a unified, coordinated and complete multimodal fusion expression form, which improves the overall information understanding ability of the system.
[0116] Feature compression is performed on the multimodal association features. Feature compression uses fully connected networks, dimensionality reduction algorithms, or information filtering mechanisms to compress the dimensionality of the multimodal association features, reduce redundant information, retain the expressive information useful for downstream tasks, and generate comprehensive features. The comprehensive features are compact structures that express the results of multimodal information fusion, and have the characteristics of high expression efficiency, clear information structure, and suitability for downstream task decision-making and control.
[0117] This embodiment generates comprehensive features with unified expression structure, coordinated information, and clear semantics by standardizing, temporally aligning, semantically weighting, jointly splicing, associative modeling, and compressing visual, audio, and linguistic features. This enhances the ability of multimodal information to be fused and expressed in complex environments, as well as the overall perception, understanding, and decision-making level of the system. It effectively improves the accuracy and reliability of the device in performing tasks in dynamic, changing, and information-complex environments.
[0118] In one embodiment, step S40 above includes:
[0119] S401, The comprehensive features are input into the task planner, and the task planner generates an initial task execution strategy.
[0120] S402, parse the action sequence in the initial task execution strategy, and generate device-executable action instructions based on the action sequence;
[0121] S403, convert the device's executable action instructions into actuator control parameters;
[0122] S404, Drive the actuator to complete the corresponding action according to the actuator control parameters.
[0123] In this embodiment, an initial task execution strategy is generated based on comprehensive features, and the actions corresponding to the initial task execution strategy are executed. Specifically, this includes multi-stage feature input, strategy generation, action parsing, and control execution. First, the comprehensive features are input into the task planner. These comprehensive features are high-dimensional features that are compact and coordinated, obtained by fusing visual, audio, and linguistic multimodal information, and possess a complete and unified ability to express environmental information. The task planner is a multi-layered decision-making module, which may include deep neural networks, graph structure models, or logical reasoning units. It accepts comprehensive feature input and, based on environmental representation, task requirements, and preset model parameters, generates an initial task execution strategy suitable for the current environment and task. The initial task execution strategy describes the device's action path, operation sequence, and behavioral logic in the environment, and possesses the ability to respond to real-time environmental conditions.
[0124] The initial task execution strategy is parsed to extract the action sequence. The initial task execution strategy contains structured action sequence information. The action sequence is a set of parameters describing the device's operation behavior, including the device's movement path, attitude adjustment, operation commands, and task execution order. The action sequence parsing decomposes the action information in the initial task execution strategy through structure decoding, parameter extraction, and sequence restoration to generate executable action instructions for the device. The executable action instructions are a low-level, clearly structured set of commands suitable for direct device control, ensuring the accuracy and stability of the device's actions.
[0125] The device executes action commands into actuator control parameters. The actuator control parameters are physical control signals designed for the specific device execution structure, including position parameters, speed parameters, angle adjustment information, and operation execution signals. The control parameter conversion is achieved through command mapping, parameter recombination, and control signal encoding. This transforms abstract action commands into physical control commands that the device execution layer can directly recognize and respond to, ensuring the compatibility and accuracy of control information between the task decision-making and device execution layers.
[0126] The actuator is driven to complete the corresponding action according to the control parameters of the actuator. The actuator can be a robotic arm, a mobile chassis, a sensor unit, or an operating module of other embodied intelligent devices. The driving operation is achieved through control signal input, power system adjustment, and coordination of the actuator components to complete the action corresponding to the initial task execution strategy. This ensures that the equipment completes the specific operation task in the environment according to the task planning requirements, forming a complete perception-decision-execution closed loop.
[0127] This embodiment improves the device's task response efficiency and execution accuracy in complex environments by using task planning, strategy generation, action sequence parsing, and control parameter conversion based on comprehensive features. It forms a dynamic task execution capability supported by multimodal information, effectively enhancing the device's autonomous decision-making level and operational reliability in changing environments.
[0128] In one embodiment, step S50 above includes:
[0129] S501, Collect environmental state change data after executing the action corresponding to the initial task execution strategy;
[0130] S502, determine the difference between the current environmental state and the target state based on the preset task objective;
[0131] S503, Generate an instant reward signal based on the difference value;
[0132] S504, input the environmental state change data and the instantaneous reward signal into the reinforcement learning model, and update the policy parameters through the reinforcement learning model;
[0133] S505 generates an updated task execution strategy based on the updated strategy parameters.
[0134] In this embodiment, during the execution of the actions corresponding to the initial task execution strategy, a reinforcement learning model is used to dynamically adjust the initial task execution strategy based on real-time feedback information from the environment. Specifically, this includes state perception based on the action execution results, difference calculation, reward generation, and strategy optimization. First, environmental state change data is collected after the actions corresponding to the initial task execution strategy are executed. This environmental state change data reflects the environmental feedback triggered by the action execution, covering spatial location information, visual changes in the environment, audio changes, language command updates, device operating status, and external interference factors. State data collection can be based on visual sensors, audio sensors, language interaction modules, and internal position and state monitoring devices to form comprehensive, real-time, and clearly structured environmental state change information.
[0135] Based on the preset task objectives, the difference between the current environmental state and the target state is determined. The preset task objectives are the expected environmental parameters and operating states set before task execution based on task requirements, environmental conditions and equipment capabilities. The difference value reflects the degree of deviation between the actual execution effect and the task objective. The difference value is calculated by comparing environmental state data, matching parameters and quantifying deviations, and outputting basic reference indicators for strategy adjustment.
[0136] Real-time reward signals are generated based on the difference values. These signals quantify the effectiveness of the current strategy and operation, reflecting the positive or negative contribution of the device's actions to the environmental state. The reward signal generation is based on numerical mapping, function transformation, or structural mapping mechanisms of the difference values, outputting reward information that is highly real-time and suitable for reinforcement learning optimization processes.
[0137] Environmental state change data and real-time reward signals are input into the reinforcement learning model. The reinforcement learning model is a decision model with adaptive optimization capabilities. It supports a policy update mechanism based on the state-action-reward relationship. By inputting environmental state and reward information and combining historical policy parameters and environmental feedback, the policy parameters are dynamically optimized. The policy parameters are a set of structured parameters that control the logic of generating the initial task execution policy. Adjusting the policy parameters makes the policy generated by the task planner more in line with the real-time environmental requirements and task objectives.
[0138] An updated task execution strategy is generated based on the updated strategy parameters. The updated task execution strategy reflects the results of strategy optimization and has higher environmental adaptability and task execution efficiency. It forms a dynamic strategy adjustment process based on multiple rounds of environmental feedback, continuously improving the device's autonomous operation capability and task response level in complex environments.
[0139] This embodiment constructs a dynamic closed-loop task execution and strategy adjustment system by collecting environmental state change data in real time, quantifying differences, generating rewards, and optimizing through reinforcement learning. This improves the task completion efficiency and decision-making accuracy of the device in complex and ever-changing environments, enabling the embodied intelligent device to perform adaptive tasks based on environmental feedback, and effectively enhancing the system's operational flexibility and environmental adaptability.
[0140] In one embodiment, after step S40 above, the method further includes:
[0141] S601, Input the visual features into the 3D reconstruction engine to generate spatial geometric topology;
[0142] S602, determine the coordinates of the sound source by processing the audio features through sound source localization;
[0143] S603, Add the language features to the spatial geometric topology to generate a semantically annotated spatial topology model.
[0144] S604, Spatial alignment is performed on the semantically labeled spatial topology model and the sound source coordinates to generate a multimodal spatial dataset;
[0145] S605, Based on the multimodal spatial dataset, voxelization processing is performed to generate a comprehensive environment model with multimodal attributes;
[0146] S606, Extract the coordinates of dynamic obstacles from the integrated environment model;
[0147] S607, detect the spatial conflict between the planned path of the initial task execution strategy and the coordinates of the dynamic obstacle, and generate a spatial conflict detection result;
[0148] S608, Generate a path replanning instruction based on the spatial conflict detection result, and map the path replanning instruction into action control parameters;
[0149] S609, drive the execution device to complete the replanning action according to the action control parameters.
[0150] In this embodiment, after generating the initial task execution strategy based on comprehensive features and executing the actions, multimodal environmental information is further acquired to improve the environmental understanding and dynamic replanning process. First, visual features are input into the 3D reconstruction engine to generate spatial geometric topology. The visual features are the representation of environmental visual information obtained through the encoding of the preceding steps. The 3D reconstruction engine constructs the spatial geometric structure of the environmental scene based on the visual features and outputs a spatial geometric topology that reflects the position, spatial layout, and geometric shape of objects in the environment. The spatial geometric topology is the basic data structure used for subsequent multimodal fusion and spatial perception, and has the ability to map spatial positions and express topological structures.
[0151] Sound source coordinates are determined by processing audio features through sound source localization. The audio features are derived from the encoding results of environmental sound information. Sound source localization processing includes time difference method based on multi-sensor array, spatial array method, or direction estimation method based on acoustic model. It outputs the spatial coordinate information of the sound source in the environment. The sound source coordinates are an important parameter for multimodal spatial perception and support subsequent spatial alignment and semantic fusion.
[0152] Linguistic features are added to the spatial geometric topology to generate a spatial topology model with semantic annotation. The linguistic features are derived from the encoded expression of user commands and environmental linguistic command information. The semantic annotation process is based on the linguistic feature content to locate target objects, semantic regions and functional divisions in the spatial geometric topology, and outputs a spatial topology model containing spatial structure and semantic information. The spatial topology model supports the unified expression of environmental structure and semantic information.
[0153] Spatial alignment of semantically labeled spatial topology models with sound source coordinates is performed to generate a multimodal spatial dataset. The spatial alignment process ensures the fusion of visual, audio, and language command information under a unified spatial coordinate system through coordinate system mapping, spatial transformation matrix calculation, and synchronous integration of multi-source data. The output is a multimodal spatial dataset with complete structure and comprehensive information. The multimodal spatial dataset is a basic data resource to support high-precision environmental modeling and dynamic replanning.
[0154] Based on the multimodal spatial dataset, voxelization is performed to generate a comprehensive environment model with multimodal attributes. Voxelization divides the continuous space into regular voxel units. Through voxel information aggregation, spatial sparse representation and multimodal feature mapping, a comprehensive environment model with efficient spatial representation and multimodal information fusion is generated. The comprehensive environment model has the ability to express the environment in real time, with rich information and clear structure.
[0155] The system extracts the coordinates of dynamic obstacles from the integrated environment model. These coordinates reflect the positions of obstacles in the environment that change or move over time. The extraction process is based on the real-time updates of the integrated environment model and dynamic object recognition, and outputs high-precision, real-time spatial location data of dynamic obstacles.
[0156] The system detects spatial conflicts between the planned path of the initial task execution strategy and the coordinates of dynamic obstacles, and generates spatial conflict detection results. The spatial conflict detection is based on the spatial relationship analysis of path and obstacle position data, collision prediction and path feasibility assessment, and outputs the detection results of conflict existence, conflict location and conflict type.
[0157] Path replanning instructions are generated based on spatial conflict detection results and mapped to action control parameters. The path replanning instructions are a set of instructions to avoid spatial conflicts and optimize path planning, while the action control parameters are the physical parameters required for the device to perform path adjustment. The two achieve dynamic task adjustment through parameter mapping and instruction conversion.
[0158] The actuator is driven by motion control parameters to complete the replanning action. The actuator is a physical operating component of the embodied intelligent device. The driving process adjusts the device's actions based on the control parameters to complete path changes, obstacle avoidance operations, and continuous task execution, ensuring the continuity and safety of the device's task execution in complex environments.
[0159] Example Description: In the autonomous navigation and task execution application of service-oriented embodied intelligent robots, for cleaning tasks in complex indoor environments, the embodied intelligent robot first collects optical signals from the environment in real time using its configured optical sensors. Specifically, this includes visible light information from the ground, walls, and furniture surfaces. The collected optical signals are processed into digital image format visual information by a data conversion module for use by the subsequent visual processing module. Simultaneously, the robot uses a distributed array of built-in acoustic sensors to collect sound wave signals from the environment, covering information such as people walking, equipment operation, and sudden abnormal sounds. The obtained sound wave signals are converted into digital audio signals for subsequent audio feature extraction. Furthermore, the robot receives verbal instructions from the user regarding the cleaning task through a voice receiver, such as "Please clean the tea room first" or "Avoid the meeting area." These verbal instructions are converted into text format by a processing module, supporting semantic understanding.
[0160] The robot performs illumination correction processing on the acquired visual information to eliminate differences in light intensity and glare interference, generating illumination-corrected visual information to ensure the stability and accuracy of subsequent image feature extraction. Next, the system extracts spatial structure features and semantic attribute features from the illumination-corrected visual information, uses a convolutional compression algorithm to reduce feature dimensionality, retains key information, and ultimately forms a clearly structured visual feature representation. Audio information undergoes noise reduction processing to remove background noise and irrelevant audio interference, generating noise-reduced audio information. Further, temporal and frequency domain features are extracted from the noise-reduced audio information to reflect the temporal variation patterns and spectral energy distribution of the sound signal. Then, a temporal modeling method is used to capture the dynamic correlations in the audio features, ultimately generating stable and reliable audio features. Language instruction information is parsed into semantic units by a word segmentation module to generate word sequences. In the context association module, combined with the overall semantic logic, a context association representation is formed. The semantic encoding module further abstracts and compresses the association representation, ultimately forming language features that express the instruction intent and environmental requirements.
[0161] To effectively integrate multimodal information, the system performs spatial dimension normalization on visual features to ensure that visual features from different sources have a unified scale standard. Audio features are aligned using temporal alignment technology to eliminate differences in acquisition time, forming synchronized audio representations. Language features are weighted with semantic attention to enhance the expressive power of important semantic content. These normalized visual features, aligned audio features, and weighted language features are then fused into a unified spliced feature. Further, a cross-modal association modeling method is used to capture the deep interdependencies between multimodal features, forming a multimodal association feature representation. Finally, feature compression reduces the overall feature dimensionality, outputting a comprehensive feature that possesses a complete and unified multimodal information expression capability.
[0162] Based on comprehensive features, the system inputs data into the task planner. The task planner generates an initial task execution strategy based on multimodal information. The strategy includes specific action sequence information, specifying the robot's travel path, cleaning area, and obstacle avoidance actions. The action sequence is converted into executable action instructions by the parsing module. These instructions are further mapped into specific actuator control parameters. The control parameters drive the robot chassis and actuators to work together to complete ground travel, path adjustment, and cleaning operations.
[0163] During the initial task execution strategy of the robot, the environmental state continuously changes. The system collects environmental state change data in real time after the action is performed, including the distribution of ground obstacles, personnel movement and sound source changes. Combined with the preset cleaning task target, the system calculates the difference between the current environmental state and the target state, and generates an instant reward signal based on the difference. The reward signal and environmental state data are input into the reinforcement learning model. The reinforcement learning model dynamically updates the strategy parameters and generates an updated task execution strategy, enabling the robot to adapt to tasks in complex environments.
[0164] As environmental information becomes richer, the system inputs the aforementioned visual features into the 3D reconstruction engine. Based on multi-angle image data, it generates spatial geometric topology information reflecting the spatial structure of the environment. Audio features are processed using a sound source localization method to accurately calculate the coordinates of the sound source in space. Language features are added to the spatial geometric topology, forming a spatial topology model that includes both physical structure and semantic information. The spatial topology model is spatially aligned with the sound source coordinates to generate a multimodal spatial dataset under a unified coordinate system. This multimodal spatial dataset is then voxelized to form a comprehensive environmental model with efficient spatial representation capabilities. Based on this comprehensive environmental model, the system extracts the coordinate information of dynamic obstacles in real time, detects spatial conflicts between the planned path of the initial task execution strategy and dynamic obstacles, and generates spatial conflict detection results. If path conflicts exist, the system generates path replanning instructions based on the detection results and maps these instructions to new motion control parameters, driving the robot to dynamically adjust its path and actions to ensure efficient cleaning task execution in dynamic and complex environments.
[0165] In the application of mobile nursing robots in hospital environments, for tasks such as clinical nursing supply delivery and ward rounds, the mobile nursing robot first uses its onboard optical sensors to collect optical signals from environmental locations such as hospital corridors, treatment areas, and ward entrances in real time. The collected optical signals are processed by a conversion module into visual information in digital image format, ensuring that the acquired visual information has clear structure and stable color representation capabilities. Simultaneously, the robot uses a distributed array of acoustic sensors to collect sound wave signals from the environment, covering information such as conversations among medical staff, operation of medical equipment, and abnormal sound alarms. The sound wave signals are converted into digital audio signals for use by the subsequent audio processing module. Furthermore, medical staff can issue voice commands through a voice receiving device, such as "Please deliver IV supplies to ward 1205" or "Check if there are any abnormal personnel gathering at the nurses' station entrance." The system converts the acquired voice commands into text-based language command information, facilitating subsequent parsing and understanding of the command content.
[0166] The robotic system performs illumination correction processing on visual information to eliminate illumination interference caused by different lighting conditions and screen reflections in the medical environment, generating illumination-corrected visual information to ensure the accuracy and consistency of subsequent image feature extraction. Based on the illumination-corrected visual information, the system extracts spatial structural features of the hospital environment and semantic information such as medical signs and door numbers, and uses convolutional compression methods to reduce the dimensionality of these features, outputting visual features. Audio information undergoes noise reduction processing to eliminate background noise and irrelevant interference, generating noise-reduced audio information. The system further extracts the temporal and frequency domain features of the audio, reflecting the specific manifestations of sounds such as alarms and abnormal medical equipment sounds. Temporal modeling methods enhance the dynamic expression of audio information, outputting stable and reliable audio features. Language command information is segmented into a word sequence, and combined with the semantics of the medical task, a semantic expression of the command is constructed through contextual association analysis. Semantic encoding methods further abstract and compress the expressed information to form language features.
[0167] In the multimodal information fusion stage, visual features undergo spatial dimension normalization to ensure that different visual information in a medical environment has a uniform scale of expression. Audio features eliminate temporal differences in signal acquisition through temporal alignment. Language features are weighted by semantic attention to enhance key semantic information relevant to the task. The three types of information are fused through a splicing operation to form spliced features. The system performs cross-modal association modeling on the spliced features to capture the deep interdependencies between visual, audio, and language command information, outputting multimodal association features. Further feature compression processing generates comprehensive and unified integrated features.
[0168] The system inputs data into the task planner based on comprehensive features. The task planner generates an initial task execution strategy, which includes the robot's travel path, obstacle avoidance strategy, and task action sequence. The action sequence is parsed and converted into executable action instructions, which are further mapped into control parameters for the execution device. These parameters control the robot chassis and manipulator structure to complete specific actions such as path navigation, material delivery, and voice broadcasting.
[0169] During execution, the robot system collects real-time data on changes in environmental conditions, such as changes in personnel positions, the distribution of obstacles in corridors, and changes in alarm information. Based on the hospital's task objectives, the system calculates the difference between the environmental state and the target state and generates an immediate reward signal. The environmental state data and reward signal are input into a reinforcement learning model, which dynamically updates the strategy parameters. The system then generates an updated task execution strategy based on this update, ensuring the robot's ability to flexibly adjust in dynamic and complex medical environments.
[0170] In response to emergencies such as crowds gathering around nurses' stations or equipment malfunctions, the system further inputs visual features into a 3D reconstruction engine to generate a spatial geometric topology. This is combined with sound source localization to determine the sound source coordinates, and linguistic features are embedded into the spatial geometric topology to generate a spatial topology model with medical semantic annotations. This spatial topology model is aligned with the sound source coordinate space to form a multimodal spatial dataset. This dataset is then voxelized to generate a comprehensive environmental model expressing multimodal information. Based on this environmental model, the system extracts the coordinates of dynamic obstacles, detects spatial conflicts between the path planning of the task execution strategy and obstacles, and generates conflict detection results. If path risks exist, the system generates path replanning instructions, which are mapped to motion control parameters to drive the robot to perform path adjustments and obstacle avoidance operations, ensuring the safe and efficient completion of material delivery and inspection tasks in a medical environment.
[0171] In the fintech business, customer service support systems for smart bank branches first utilize optical sensors deployed in the lobby and self-service areas to continuously collect optical signals from environmental locations such as counter areas, waiting areas, and ticket dispensers. These signals are then converted into digital image formats to ensure the environmental images possess clear spatial structure and information about human activity. Simultaneously, the system uses a distributed array of acoustic sensors to collect ambient sound signals, including customer inquiries, device voice prompts, and noise from self-service equipment, generating digital audio signals for subsequent audio processing. Bank staff or system operators issue verbal commands via a voice receiver, such as "Guide the customer to counter 3" or "Check for any anomalies in the self-service equipment area." The system automatically converts the voice information into structured text format for subsequent semantic understanding and task assignment.
[0172] Based on visual information, the system performs illumination correction processing to eliminate image deviations caused by light source reflections and display screen interference in the business hall, outputting illumination-corrected visual information. The system further extracts the spatial structural features and environmental semantic features of the illumination-corrected visual information, including personnel distribution, signage, and area division information. Convolutional compression is used to reduce the dimensionality of the feature data, generating stable and complete visual features. Audio information undergoes noise reduction processing to remove background noise and irrelevant interference, generating noise-reduced audio information. The system extracts the temporal and frequency domain features of the audio data, reflecting specific manifestations such as voice dialogues and abnormal sounds. Temporal modeling methods enhance the dynamic characteristics of the audio, outputting audio features. Language instruction information is segmented into word sequences. The system combines banking business semantics with contextual analysis to generate accurate contextual representations. After semantic encoding and compression, language features are output.
[0173] In the multimodal fusion stage, the system performs spatial dimension normalization on visual features to ensure consistent expression of visual information from different sources. Audio features are aligned in the temporal domain to solve the problem of time synchronization between audio data and environmental images. Language features are weighted by semantic attention to enhance the expression of key information such as customer instructions and business requests. The three types of features are spliced and fused to generate spliced features. The system extracts the deep-level connections between visual, audio, and language instruction information through cross-modal association modeling to generate multimodal association features. Furthermore, feature compression is used to form comprehensive features.
[0174] The system inputs comprehensive features into the task planner to generate an initial task execution strategy. This strategy includes an intelligent guidance path, information prompting flow, and risk prevention action sequence. The action sequence is parsed and converted into executable action instructions for the device, which are then mapped to control parameters of the execution device. This drives the guide screen, guide light strip, voice broadcast module, and other components to perform intelligent guidance, information dissemination, and environmental prompting operations.
[0175] During system task execution, real-time environmental status change data is collected, including changes in customer location, equipment operating status, and on-site sound changes. Combined with pre-set service procedures and risk warning requirements, the system calculates the difference between the current environmental state and the business objective state, generating an immediate reward signal. Environmental data and reward signals are input into a reinforcement learning model to dynamically update strategy parameters. Based on the latest parameters, the system generates an updated task execution strategy, ensuring flexible adjustment of task planning during peak business periods, abnormal events, and temporary changes, thereby improving customer experience and business continuity.
[0176] In response to abnormal situations, such as queuing issues, customer congestion, or sudden noises in the self-service area, the system generates a spatial geometric topology based on visual features, determines the coordinates of the sound source through sound source localization, and embeds linguistic features into the spatial geometric topology to generate a semantically annotated spatial topology model. The system performs spatial alignment, generates a multimodal spatial dataset, and after voxelization, forms a comprehensive environmental model. It extracts the coordinates of dynamic obstacles, detects path-obstacle conflicts in the task execution strategy, and generates spatial conflict detection results. Based on these results, the system generates path replanning instructions, maps them to action control parameters, and drives the system to perform actions such as path adjustment, voice prompts, and crowd control guidance. This effectively addresses the complex environmental changes in bank branches, ensuring intelligent, secure, and efficient customer service.
[0177] This embodiment integrates multimodal information from visual, audio, and linguistic features, combined with spatial geometric reconstruction, sound source localization, semantic annotation, and spatial alignment, to form a structurally complete and information-rich multimodal spatial data representation. Based on dynamic obstacle detection and path conflict analysis using a comprehensive environmental model, it generates path replanning instructions and control parameters in real time, driving the execution device to dynamically adjust the task path. This effectively enhances the device's autonomous obstacle avoidance capability, path optimization capability, and task continuity in complex dynamic environments, realizing the embodied intelligent system's efficient integration of multimodal information and adaptive operation capability in dynamic environments.
[0178] In one embodiment, a task execution strategy generation and adjustment apparatus is provided, which corresponds one-to-one with the task execution strategy generation and adjustment method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the task execution strategy generation and adjustment device of the present invention. The modules include a multimodal perception module 10, a feature encoding module 20, a multimodal fusion module 30, a task decision execution module 40, and a strategy adaptive optimization module 50. Detailed descriptions of each functional module are as follows:
[0179] The multimodal perception module 10 is used to acquire visual information, audio information, and task-related language instruction information in the environment;
[0180] Feature encoding module 20 is used to encode the visual information, the audio information and the language instruction information to obtain visual features, audio features and language features respectively;
[0181] The multimodal fusion module 30 is used to fuse the visual features, the audio features, and the language features to generate a comprehensive feature;
[0182] The task decision execution module 40 is used to generate an initial task execution strategy based on the comprehensive features and execute the actions corresponding to the initial task execution strategy.
[0183] The strategy adaptive optimization module 50 is used to dynamically adjust the initial task execution strategy based on real-time feedback information from the environment during the execution of the initial task execution strategy, using a reinforcement learning model to obtain an updated task execution strategy.
[0184] In one embodiment, the multimodal sensing module 10 is specifically used for:
[0185] An ambient optical signal is acquired through an optical sensor, and the ambient optical signal is converted into visual information in digital image format;
[0186] Ambient sound wave signals are acquired through an acoustic sensor array, and the ambient sound wave signals are converted into digital audio signals.
[0187] The system acquires user voice commands through a voice receiving device and converts the user voice commands into text-formatted language command information.
[0188] In one embodiment, the feature encoding module 20 is specifically used for:
[0189] Perform illumination correction processing on the visual information to generate illumination-corrected visual information;
[0190] The spatial and semantic features of the illumination correction visual information are extracted, and the spatial and semantic features are convolutionally compressed to generate visual features;
[0191] The audio information is subjected to noise reduction processing to generate noise-reduced audio information;
[0192] Extract the time-domain and frequency-domain features of the noise-reduced audio information, and perform time-series modeling on the time-domain and frequency-domain features to generate audio features;
[0193] The language instruction information is subjected to lexical segmentation processing to generate a lexical sequence;
[0194] Context association processing is performed on the lexical sequence to generate a context association representation, and semantic encoding processing is performed on the context association representation to generate language features.
[0195] In one embodiment, the multimodal fusion module 30 is specifically used for:
[0196] The visual features are subjected to spatial dimension normalization processing to generate normalized visual features;
[0197] Perform temporal alignment processing on the audio features to generate aligned audio features;
[0198] Semantic attention weighting processing is performed on the language features to generate weighted language features;
[0199] The normalized visual features, the aligned audio features, and the weighted language features are concatenated to generate concatenated features;
[0200] Perform cross-modal association modeling on the spliced features to generate multimodal association features;
[0201] The multimodal correlation features are subjected to feature compression processing to generate comprehensive features.
[0202] In one embodiment, the task decision execution module 40 is specifically used for:
[0203] The comprehensive features are input into the task planner, which then generates an initial task execution strategy.
[0204] The action sequence in the initial task execution strategy is parsed, and device-executable action instructions are generated based on the action sequence;
[0205] Convert the executable action commands of the device into control parameters of the execution device;
[0206] The actuator is driven to complete the corresponding action according to the control parameters of the actuator.
[0207] In one embodiment, the policy adaptive optimization module 50 is specifically used for:
[0208] Collect environmental state change data after executing the actions corresponding to the initial task execution strategy;
[0209] Determine the difference between the current environmental state and the target state based on the preset task objective;
[0210] An instant reward signal is generated based on the difference value;
[0211] The environmental state change data and the instantaneous reward signal are input into the reinforcement learning model, and the policy parameters are updated through the reinforcement learning model.
[0212] An updated task execution strategy is generated based on the updated strategy parameters.
[0213] In one embodiment, the task decision execution module 40 is specifically used for:
[0214] The visual features are input into a 3D reconstruction engine to generate spatial geometric topology;
[0215] The sound source coordinates are determined by processing the audio features to locate the sound source.
[0216] The linguistic features are added to the spatial geometric topology to generate a semantically labeled spatial topology model.
[0217] Spatial alignment is performed between the semantically labeled spatial topology model and the sound source coordinates to generate a multimodal spatial dataset;
[0218] Based on the multimodal spatial dataset, voxelization is performed to generate a comprehensive environment model with multimodal attributes;
[0219] Extract the coordinates of dynamic obstacles from the integrated environment model;
[0220] Detect spatial conflicts between the planned path of the initial task execution strategy and the coordinates of the dynamic obstacles, and generate spatial conflict detection results;
[0221] Based on the spatial conflict detection results, a path replanning instruction is generated, and the path replanning instruction is mapped to action control parameters.
[0222] The actuator is driven to complete the replanning action according to the motion control parameters.
[0223] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a task execution strategy generation and adjustment method on the server side.
[0224] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the user-side functions or steps of a task execution strategy generation and adjustment method.
[0225] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0226] Acquire visual information, audio information, and task-related language instructions from the environment;
[0227] The visual information, the audio information, and the language instruction information are encoded to obtain visual features, audio features, and language features, respectively.
[0228] By fusing the visual features, the audio features, and the language features, a comprehensive feature is generated.
[0229] An initial task execution strategy is generated based on the comprehensive features, and the actions corresponding to the initial task execution strategy are executed.
[0230] During the execution of the initial task execution strategy, a reinforcement learning model is used to dynamically adjust the initial task execution strategy based on real-time feedback information from the environment, resulting in an updated task execution strategy.
[0231] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0232] Acquire visual information, audio information, and task-related language instructions from the environment;
[0233] The visual information, the audio information, and the language instruction information are encoded to obtain visual features, audio features, and language features, respectively.
[0234] By fusing the visual features, the audio features, and the language features, a comprehensive feature is generated.
[0235] An initial task execution strategy is generated based on the comprehensive features, and the actions corresponding to the initial task execution strategy are executed.
[0236] During the execution of the initial task execution strategy, a reinforcement learning model is used to dynamically adjust the initial task execution strategy based on real-time feedback information from the environment, resulting in an updated task execution strategy.
[0237] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0238] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0239] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0240] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for generating and adjusting task execution strategies, characterized in that, Includes the following steps: Acquire visual information, audio information, and task-related language instructions from the environment; The visual information, the audio information, and the language instruction information are encoded to obtain visual features, audio features, and language features, respectively. By fusing the visual features, the audio features, and the language features, a comprehensive feature is generated. An initial task execution strategy is generated based on the comprehensive features, and the actions corresponding to the initial task execution strategy are executed. During the execution of the initial task execution strategy, a reinforcement learning model is used to dynamically adjust the initial task execution strategy based on real-time feedback information from the environment, resulting in an updated task execution strategy.
2. The task execution strategy generation and adjustment method as described in claim 1, characterized in that, Acquire visual information, audio information, and task-related language instructions from the environment, including: An ambient optical signal is acquired through an optical sensor, and the ambient optical signal is converted into visual information in digital image format; Ambient sound wave signals are acquired through an acoustic sensor array, and the ambient sound wave signals are converted into digital audio signals. The system acquires user voice commands through a voice receiving device and converts the user voice commands into text-formatted language command information.
3. The task execution strategy generation and adjustment method as described in claim 1, characterized in that, Encoding the visual information, the audio information, and the language instruction information to obtain visual features, audio features, and language features respectively includes: Perform illumination correction processing on the visual information to generate illumination-corrected visual information; The spatial and semantic features of the illumination correction visual information are extracted, and the spatial and semantic features are convolutionally compressed to generate visual features; The audio information is subjected to noise reduction processing to generate noise-reduced audio information; Extract the time-domain and frequency-domain features of the noise-reduced audio information, and perform time-series modeling on the time-domain and frequency-domain features to generate audio features; The language instruction information is subjected to lexical segmentation processing to generate a lexical sequence; Context association processing is performed on the lexical sequence to generate a context association representation, and semantic encoding processing is performed on the context association representation to generate language features.
4. The task execution strategy generation and adjustment method as described in claim 1, characterized in that, By fusing the visual features, the audio features, and the language features, a comprehensive feature is generated, including: The visual features are subjected to spatial dimension normalization processing to generate normalized visual features; Perform temporal alignment processing on the audio features to generate aligned audio features; Semantic attention weighting processing is performed on the language features to generate weighted language features; The normalized visual features, the aligned audio features, and the weighted language features are concatenated to generate concatenated features; Perform cross-modal association modeling on the spliced features to generate multimodal association features; The multimodal correlation features are subjected to feature compression processing to generate comprehensive features.
5. The task execution strategy generation and adjustment method as described in claim 1, characterized in that, An initial task execution strategy is generated based on the comprehensive features, and the actions corresponding to the initial task execution strategy are executed, including: The comprehensive features are input into the task planner, which then generates an initial task execution strategy. The action sequence in the initial task execution strategy is parsed, and device-executable action instructions are generated based on the action sequence; Convert the executable action commands of the device into control parameters of the execution device; The actuator is driven to complete the corresponding action according to the control parameters of the actuator.
6. The task execution strategy generation and adjustment method as described in claim 1, characterized in that, During the execution of the initial task execution strategy, based on real-time feedback information from the environment, a reinforcement learning model is used to dynamically adjust the initial task execution strategy to obtain an updated task execution strategy, including: Collect environmental state change data after executing the actions corresponding to the initial task execution strategy; Determine the difference between the current environmental state and the target state based on the preset task objective; An instant reward signal is generated based on the difference value; The environmental state change data and the instantaneous reward signal are input into the reinforcement learning model, and the policy parameters are updated through the reinforcement learning model. An updated task execution strategy is generated based on the updated strategy parameters.
7. The task execution strategy generation and adjustment method as described in claim 1, characterized in that, After generating an initial task execution strategy based on the comprehensive features and executing the actions corresponding to the initial task execution strategy, the method further includes: The visual features are input into a 3D reconstruction engine to generate spatial geometric topology; The sound source coordinates are determined by processing the audio features to locate the sound source. The linguistic features are added to the spatial geometric topology to generate a semantically labeled spatial topology model. Spatial alignment is performed between the semantically labeled spatial topology model and the sound source coordinates to generate a multimodal spatial dataset; Based on the multimodal spatial dataset, voxelization is performed to generate a comprehensive environment model with multimodal attributes; Extract the coordinates of dynamic obstacles from the integrated environment model; Detect spatial conflicts between the planned path of the initial task execution strategy and the coordinates of the dynamic obstacles, and generate spatial conflict detection results; Based on the spatial conflict detection results, a path replanning instruction is generated, and the path replanning instruction is mapped to action control parameters. The actuator is driven to complete the replanning action according to the motion control parameters.
8. A task execution strategy generation and adjustment device, characterized in that, The task execution strategy generation and adjustment device includes: The multimodal perception module is used to acquire visual information, audio information, and task-related language instructions from the environment. The feature encoding module is used to encode the visual information, the audio information, and the language instruction information to obtain visual features, audio features, and language features, respectively. A multimodal fusion module is used to fuse the visual features, the audio features, and the language features to generate a comprehensive feature; The task decision execution module is used to generate an initial task execution strategy based on the comprehensive features and execute the actions corresponding to the initial task execution strategy. The strategy adaptive optimization module is used to dynamically adjust the initial task execution strategy based on real-time feedback information from the environment during the execution of the initial task execution strategy, using a reinforcement learning model to obtain an updated task execution strategy.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a task execution strategy generation and adjustment program stored in the memory and executable on the processor. When the task execution strategy generation and adjustment program is executed by the processor, it implements the steps of the task execution strategy generation and adjustment method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a task execution strategy generation and adjustment program, which, when executed by a processor, implements the steps of the task execution strategy generation and adjustment method as described in any one of claims 1-7.
Citation Information
Cited By
Data processing method oriented to customer relationship management
CN121581875A
Industrial autonomous mobile robot control method, device, equipment and medium
CN121733591A
An industrial autonomous mobile robot control method, apparatus, device and medium
CN121733591B
Data processing method and device, equipment, storage medium and program product
CN121882092A