Cross-scenario adaptive online optimization method for embodied robot perception parameters

By acquiring historical and current scene data of the embodied robot, and combining visual and voice sensor analysis, the degree of environmental change is calculated and perception parameters are optimized. This solves the problem of rigid robot perception parameters, realizes online optimization of perception parameters that is adaptive across scenes, and improves the robot's perception robustness and task execution reliability in complex environments.

CN121613757BActive Publication Date: 2026-04-17BEIJING MIANBI INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING MIANBI INTELLIGENT TECH CO LTD
Filing Date
2026-02-03
Publication Date
2026-04-17

Smart Images

  • Figure CN121613757B_ABST
    Figure CN121613757B_ABST
Patent Text Reader

Abstract

The application discloses a cross-scene adaptive embodied robot perception parameter online optimization method and relates to the technical field of robot perception control. The method comprises the following steps: obtaining historical scene types and environment data of an embodied robot, identifying a current scene type, collecting voice data and calling a special environment analysis model to determine current environment data if the scene types are the same, directly matching standard perception parameters if the scene types are different, calculating an environment change degree based on the historical and current environment data, constructing a perception parameter optimization strategy, including determining an optimization depth and a parameter adjustment space, finally taking minimization of a perception deviation as a target, performing iterative optimization, evaluating parameter performance by using an integrated prediction plug-in, and outputting optimal adaptive parameters. The application realizes adaptive switching of perception parameters between different scenes and dynamic optimization in the same scene, and improves the perception robustness and task execution reliability of the robot in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot perception and control technology, specifically to an online optimization method for the perception parameters of an embodied robot that is adaptable across different scenarios. Background Technology

[0002] An embodied robot is an intelligent agent that interacts with the physical environment in real time and needs to operate in diverse real-world scenarios. However, different scenarios have significantly different environmental characteristics. Lighting conditions, spatial layout, and acoustic characteristics are not static and unchanging. Their dynamic changes can directly interfere with the normal operation of the onboard perception module, leading to a decrease in visual recognition accuracy or failure of voice interaction functions.

[0003] Currently, the configuration of robot perception parameters mostly adopts fixed patterns or relies on manual adjustments, making it difficult to adapt to dynamic environmental changes. When the robot enters a new scene or encounters sudden environmental changes, its perception performance is prone to significant degradation. Although there are methods in existing technologies that switch preset parameter sets through scene recognition, they fail to fully consider the continuous changing characteristics of the environment within the same scene and require the support of a large amount of pre-configured data, exhibiting insufficient flexibility in practical applications. In addition, traditional parameter optimization methods are mainly based on offline training modes, which are difficult to effectively cope with unknown environmental states that occur during actual operation, thus affecting the accuracy and reliability of robots in complex application scenarios. Summary of the Invention

[0004] This invention addresses the technical problems of insufficient adaptability of robot perception systems due to dynamic environmental changes, rigid parameter configuration, and lack of real-time performance caused by offline optimization modes in existing technologies. It provides an online optimization method for embodied robot perception parameters that is adaptive across different scenarios.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0006] This invention provides an online optimization method for the perception parameters of an embodied robot that is adaptive across different scenarios, including:

[0007] Obtain the historical scene type and historical environment data corresponding to the current perception parameters of the embodied robot;

[0008] The current scene type is determined by acquiring images of the current scene through the onboard image sensor of the embodied robot;

[0009] If the current scene type and the historical scene type are the same, the current voice data is collected by the onboard voice sensor of the embodied robot, and the current environmental data is determined by analyzing the current scene type and the current voice data.

[0010] The degree of change in the current environment is calculated based on the historical environmental data and the current environmental data, and a perception parameter optimization strategy is configured based on the degree of change in the current environment and the current perception parameters.

[0011] With the goal of minimizing perception bias, based on the current scene type and current environment data, perception parameters are optimized online according to the perception parameter optimization strategy to obtain suitable perception parameters for task recognition.

[0012] The beneficial effects of this invention are:

[0013] Compared to existing technologies, this invention first achieves dynamic adaptive adjustment of perception parameters based on environmental changes. By calculating the degree of environmental change in real time and constructing a parameter optimization strategy, it effectively overcomes the performance degradation problem of fixed parameter configurations during sudden environmental changes. Second, by integrating visual and speech multimodal perception data, it comprehensively captures changes in scene features, improving the accuracy of environmental state analysis. Third, through an online optimization mechanism combined with a perception deviation prediction model, it achieves intelligent and efficient parameter optimization while ensuring real-time performance. Finally, it can achieve fine-tuning of parameters in the same scene type and quickly match standard parameters in new scenes, comprehensively improving the perception robustness and task execution reliability of the embodied robot in complex application scenarios. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating the online optimization method for the perception parameters of an embodied robot for cross-scenario adaptation provided by the present invention.

[0015] Figure 2 This is a schematic diagram of the online iterative optimization process for sensing parameters provided by the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0018] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0019] Example 1, as Figure 1 As shown, this embodiment of the invention provides an online optimization method for the perception parameters of an embodied robot that is adaptive across different scenarios, including:

[0020] S10: Obtain the historical scene type and historical environment data corresponding to the current perception parameters of the embodied robot;

[0021] Specifically, the historical scene types and historical environmental data corresponding to the current perception parameters of the embodied robot are obtained. The embodied robot is a general-purpose handling robot. The perception parameters include visual OCR parameters and speech recognition parameters. The scene types include at least residential indoor environments, laboratories, logistics sorting centers, and industrial warehouses. The environmental data includes visual environmental data and acoustic environmental data. The environmental parameters of the visual environmental data include at least light intensity characteristics, dominant color spectrum, spatial structure complexity, and a list of key objects. The environmental parameters of the acoustic environmental data include at least background noise decibels, noise type characteristics, and reverberation time.

[0022] First, historical scene types and historical environment data associated with the embodied robot's current perception parameters are acquired. The embodied robot is a general-purpose handling robot. The robot's current perception parameters are a set of configuration parameters used when perceiving its environment, representing the core settings controlling the operation of the visual and speech recognition modules, including visual OCR parameters and speech recognition parameters. Specifically, the visual OCR parameters control the threshold and feature extraction rules for the optical character recognition process; the speech recognition parameters adjust the sensitivity and filtering coefficients for speech signal processing and semantic parsing.

[0023] Acquire historical scene types and historical environment data corresponding to the current perception parameters. Specifically, scene types include at least residential indoor environments, laboratories, logistics sorting centers, and industrial warehouses, representing typical environmental classifications in which the robot successfully performed tasks historically. Environmental data includes visual environment data and acoustic environment data, representing a set of quantifiable physical features in the historical environment. Among them, the environmental parameters of the visual environment data include at least light intensity characteristics, dominant color spectrum, spatial structure complexity, and a list of key objects; the environmental parameters of the acoustic environment data include at least background noise decibels, noise type characteristics, and reverberation time.

[0024] Specifically, acquiring historical scene types and environmental data corresponding to the current perception parameters is essentially a process of establishing an environmental state benchmark. Historical operation records can be retrieved from local storage units or cloud databases for quantifying the degree of subsequent environmental changes and formulating perception parameter optimization strategies. The historical period corresponding to the historical scene types and environmental data refers to the time period during which the embodied robot's perception parameters were successfully adapted, set according to the embodied robot's working cycle, such as typical environmental sampling data from the past 7 days.

[0025] S20: Acquire images of the current scene using the onboard image sensor of the embodied robot, and identify and determine the current scene type based on the current scene images;

[0026] Secondly, the embodied robot acquires images of the current scene using its onboard image sensor. This onboard image sensor is a visual perception device integrated into the robot's body, mounted on the embodied robot's mobile platform or gimbal structure. During the robot's operation, this onboard image sensor actively captures real-time visual information of its environment, forming a current scene image representing the environmental state. This onboard image sensor can be a color camera or a depth camera, capable of acquiring two-dimensional planar visual data of the environment or three-dimensional spatial visual data including distance information, respectively.

[0027] After acquiring the current scene image, it is necessary to perform scene type recognition analysis to determine the current scene type. Specifically, this recognition process can be achieved through a pre-trained scene classification model. This scene classification model can extract deep features from the image, such as spatial layout, distribution of typical objects, architectural structural features, and material texture information, and output the scene category to which the current environment belongs.

[0028] The scene types are categorized based on actual application needs. Typical categories include residential indoor environments, laboratory environments, logistics sorting center environments, and industrial warehouse environments. Each scene type has visual characteristic patterns that distinguish it from other categories. For example, residential indoor environments typically include household items and relatively soft lighting; industrial warehouse environments exhibit well-organized shelving structures and specific industrial lighting characteristics.

[0029] S30: If the current scene type and the historical scene type are the same, collect the current voice data through the onboard voice sensor of the embodied robot, and analyze and determine the current environmental data based on the current scene type and the current voice data;

[0030] Furthermore, by comparing the current scene type with the previously acquired historical scene types, when the current scene type identified by the onboard image sensor is inconsistent with the stored historical scene types, it indicates that the embodied robot has entered a new working environment. At this point, the fundamental characteristics of the environment have changed drastically. Under these conditions, if the parameter optimization method based on the original historical environmental data is continued, it will not only lead to a significant reduction in optimization efficiency, but may also cause the optimization process to fail due to the inherent mismatch of the environmental model.

[0031] Therefore, a standard parameter matching process needs to be initiated immediately. This process aims to provide the embodied robot with validated standard perception parameters based on prior knowledge of the new scenario, thereby ensuring its basic perception capabilities and task execution reliability in unfamiliar environments.

[0032] Specifically, if the current scene type and the historical scene type are different, the environmental complexity is evaluated based on the current scene image and the current voice data to obtain the current environmental complexity.

[0033] Call the preset scene type-perception parameter lookup table, and obtain standard perception parameters based on the current scene type and current environment complexity to perform task recognition;

[0034] The environmental complexity evaluation based on the current scene image and current voice data includes:

[0035] Convolutional neural networks are used to construct image feature recognizers and speech feature recognizers;

[0036] The image feature recognizer extracts features from the current scene image and evaluates the visual complexity based on the visual features, including illumination uniformity, background clutter, and text recognition difficulty.

[0037] The speech feature recognizer extracts features from the current speech data and evaluates the speech complexity based on the speech features. The speech features include the signal-to-noise ratio and the proportion of burst noise.

[0038] The current environment complexity is calculated by weighting the visual complexity and the speech complexity.

[0039] When the current scenario type is different from the historical scenario type, the first step is to quantitatively assess the new working environment, that is, to obtain the complexity of the current environment. This complexity is a comprehensive quantitative indicator used to accurately characterize the degree of challenge that the current environment poses to the robot's perception system.

[0040] Specifically, to evaluate environmental complexity based on the current scene image and current speech data, it is first necessary to construct an image feature recognizer and a speech feature recognizer using convolutional neural networks (CNNs). A CNN is a deep learning model with a multi-layered perceptual structure, which can automatically extract hierarchical features from raw data through convolutional kernel operations. The constructed image feature recognizer is a CNN model specifically designed to process visual signals, used to extract visual features related to perceptual difficulty from the current scene image; the speech feature recognizer is a CNN model specifically designed to process acoustic signals, used to extract acoustic features related to recognition challenges from the current speech data.

[0041] Specifically, the image feature recognizer performs multi-level feature extraction on the current scene image, evaluates the complexity based on the extracted visual features, and outputs the visual complexity. Specifically, the visual features used include illumination uniformity, background clutter, and text recognition difficulty. Illumination uniformity characterizes the stability of the ambient light distribution, background clutter reflects the degree of interference from non-target objects in the scene on visual attention, and text recognition difficulty specifically evaluates the clarity and regularity of text regions in optical character recognition tasks.

[0042] Simultaneously, the speech feature recognizer extracts acoustic features from the current speech data, evaluates the complexity based on the extracted features, and outputs the speech complexity. Specifically, the speech features used include the signal-to-noise ratio (SNR) and the proportion of burst noise. The SNR quantifies the relative relationship between the target speech signal strength and the background noise strength, while the proportion of burst noise statistically analyzes the frequency of instantaneous high-intensity noise within a given time period.

[0043] Then, after obtaining the visual complexity and speech complexity separately, the visual complexity and speech complexity are weighted according to a preset weight ratio to obtain the current environmental complexity value that comprehensively represents the difficulty of environmental perception. The preset weights are set according to the degree of dependence of the specific task on visual and speech perception. For example, in a scenario where speech interaction is the main task, the weight of speech complexity can be set to 0.7, and the weight of visual complexity can be set to 0.3 accordingly; while in a scenario where visual navigation is the main task, the weight of visual complexity can be set to 0.8, and the weight of speech complexity can be set to 0.2.

[0044] Furthermore, after obtaining the current environment complexity, a pre-defined scene type-perception parameter lookup table is invoked. This lookup table pre-stores multiple sets of standard perception parameter configurations for each scene type, and each set of standard perception parameter configurations corresponds to a specific range of environment complexity. The matching operation uses the current scene type as the first-level query index and the calculated current environment complexity as the second-level query index to perform a precise search in the scene type-perception parameter lookup table, thereby matching and obtaining the standard perception parameters that best suit the current environment conditions.

[0045] In summary, by directly matching the new scene type and the predefined standard perception parameters under this environmental complexity, the embodied robot can quickly acquire perception performance that is basically adapted to the new scene without complex online parameter tuning. This approach effectively ensures the continuity of task execution when the robot faces scene switching, providing reliable technical support for maintaining a high success rate of basic operations in cross-scene operations.

[0046] Furthermore, if the current scene type is the same as the historical scene type, it indicates that the embodied robot is still operating within the same environmental framework, and the main structure and macroscopic features of the environment remain stable. At this point, the main factors affecting perception performance become the dynamically changing physical parameters within the environment, such as instantaneous fluctuations in light intensity, variations in background noise levels, and changes in air reverberation characteristics. While these parameter changes are insufficient to fundamentally alter the scene type, they continuously interfere with the accuracy of the visual and speech perception modules.

[0047] Therefore, a more refined environmental monitoring process needs to be initiated. Specifically, airborne voice sensors should be used to collect and analyze current environmental data, thereby providing real-time data support for precise fine-tuning of perception parameters within the same scenario framework.

[0048] Specifically, the onboard voice sensor of the embodied robot collects current voice data, and the current environmental data is determined based on the current scene type and the current voice data analysis, including:

[0049] The onboard voice sensor of the embodied robot collects the voice data sequence of the current area within a preset time zone;

[0050] Within the preset environment analysis model library, an appropriate environment analysis model is obtained by matching the current scene type.

[0051] Using the adaptive environment analysis model, the current environment data is determined based on the analysis of the speech data sequence.

[0052] First, the onboard voice sensor of the embodied robot is activated to continuously collect audio signals from the current operating area within a preset time zone, forming a speech data sequence with temporal characteristics. The collected speech data sequence includes mixed audio information such as ambient background sound and operation commands. The preset time zone is set according to the real-time requirements of the specific application scenario, for example, it is set to a continuous sampling period of the most recent 30 seconds to ensure that the acquired acoustic features can accurately reflect the current state of the environment.

[0053] Secondly, model matching is performed within a pre-defined environmental analysis model library. This library stores dedicated environmental analysis models corresponding to different scene types. Each dedicated environmental analysis model is specifically trained for the acoustic characteristics of a particular scene. Based on the identified current scene type, a perfectly matching adaptive environmental analysis model is selected from the library. This adaptive environmental analysis model is a machine learning model fully trained with acoustic data from a specific scene, capable of accurately analyzing the speech data features of the corresponding scene, and used to extract key acoustic parameters representing the environmental state from mixed audio information.

[0054] Specifically, the training method for the adaptive environment analysis model includes:

[0055] Using the current scene type as a constraint, a sample speech data sequence set is collected, and historical environmental data corresponding to different sample speech data sequences are collected as sample environmental data to obtain a sample environmental dataset;

[0056] Using the sample speech data sequence set as input and the sample environment dataset as supervision, a convolutional neural network is trained until convergence to obtain an adaptive environment analysis model.

[0057] Specifically, the training method for the adaptive environment analysis model includes the following steps:

[0058] For example, since there is a complex nonlinear mapping relationship between acoustic environmental features and environmental parameters, and convolutional neural networks have outstanding advantages in time-series signal feature extraction and pattern recognition, a convolutional neural network architecture is chosen to construct this environmental analysis model.

[0059] Specifically, this adaptive environment analysis model mainly consists of a feature extraction layer, a temporal analysis layer, and a regression output layer. The feature extraction layer employs a one-dimensional convolutional kernel structure to extract local acoustic features from the speech data sequence; the kernel size is configured according to the sampling frequency of the speech signal. The temporal analysis layer includes pooling operations and a fully connected network, which progressively abstracts higher-level acoustic pattern representations by downsampling and nonlinearly transforming the feature sequence. The regression output layer uses a linear activation function to map the learned features to specific environmental parameter values.

[0060] During training, key hyperparameters included a learning rate of 0.001, a training epoch count of 200, and a batch size of 32. The learning rate setting ensured the stability of the gradient descent process, the number of training epochs guaranteed sufficient learning of the acoustic features by the model, and the batch size balanced memory usage with training efficiency.

[0061] Specifically, in the training data preparation phase, a set of sample speech data sequences is collected under the constraint of the current specific scene type. Each sample speech data sequence contains acoustic signals from multiple consecutive sampling points. At the same time, historical environmental data corresponding to different sample speech data sequences are collected as sample environmental data to form a sample environmental dataset containing dimensions such as background noise decibels, noise type characteristics, and reverberation time.

[0062] Furthermore, the sample speech data sequence set is used as the model input, and the corresponding sample environment dataset is used as the supervision signal, divided into training, validation, and test sets in a 6:2:2 ratio. Mean squared error is used as the loss function, and the network parameters are iteratively updated using the backpropagation algorithm in conjunction with the Adam optimizer. The training process is terminated when the validation set loss no longer decreases for several consecutive training epochs and the model prediction error is below a preset threshold, such as 5%, resulting in a converged adapted environment analysis model. This training method ensures that the adapted environment analysis model can accurately establish the mapping relationship from speech sequences to environmental parameters, thus providing reliable technical support for environmental state analysis.

[0063] Finally, the selected adaptive environment analysis model is invoked, and the collected speech data sequence is input into the adaptive environment analysis model for analysis and processing. This adaptive environment analysis model has learned the correlation between acoustic features and visual environment parameters in a specific scenario during training. Therefore, it can comprehensively analyze and output complete current environment data, including both visual and acoustic environment data, based on the speech data sequence, providing quantified current environment data.

[0064] Finally, the current environmental data obtained through the above process includes visual environmental data and acoustic environmental data. The environmental parameters of the visual environmental data include at least illumination intensity characteristics, dominant color spectrum, spatial structure complexity, and a list of key objects. The environmental parameters of the acoustic environmental data include at least background noise in decibels, noise type characteristics, and reverberation time. This current environmental data can provide accurate environmental state input for subsequent optimization of perception parameters.

[0065] S40: Calculate the current environmental change degree based on the historical environmental data and the current environmental data, and configure a perception parameter optimization strategy based on the current environmental change degree and the current perception parameters;

[0066] Specifically, the degree of change in the current environment is calculated based on the historical and current environmental data, including:

[0067] Multidimensional deviation calculation is performed on the historical environmental data and the current environmental data. The ratio of the environmental deviation value to the corresponding historical environmental parameter value in the historical environmental data is set as the environmental parameter change coefficient, and multiple environmental parameter change coefficients are obtained.

[0068] Based on the current scene type matching, the weight ratio of the adaptive environment parameters is obtained, and the change coefficients of the multiple environment parameters are weighted and calculated to obtain the current environment change degree.

[0069] First, a multidimensional deviation calculation is performed on historical and current environmental data. This calculation is performed separately for each corresponding environmental parameter, subtracting the value of a parameter in the current environmental data from the value of the same parameter in the historical environmental data to obtain the environmental deviation value for that parameter. Second, the environmental deviation value is compared with the corresponding historical environmental parameter value to obtain the environmental parameter variation coefficient for that parameter. The calculated environmental parameter variation coefficient is a dimensionless relative change quantity used to characterize the relative magnitude and direction of change of a specific environmental parameter relative to its historical baseline state. By repeating this calculation process for all environmental parameters, a set of environmental parameter variation coefficients reflecting the relative degree of change of each parameter can be obtained.

[0070] For example, suppose in an industrial warehouse scenario, the background noise recorded in historical environmental data is 55 dB. The background noise in the current environmental data is 65 dB. Then the environmental deviation value of the background noise = 65 dB - 55 dB = 10 dB, and the environmental parameter variation coefficient of the background noise = 10 dB / 55 dB ≈ 0.182.

[0071] Furthermore, the appropriate environmental parameter weights are obtained based on the current scene type. Since the impact of various environmental parameters on the perception system differs across scene types, a pre-defined environmental parameter weight allocation scheme matching the characteristics of each scene type is required. Key environmental parameters with a greater impact on the perception performance of that scene are assigned higher weights. The weight of each environmental parameter is set based on the sensitivity analysis results of the impact of each environmental parameter on the performance of the perception module in a specific scene. For example, in an industrial warehouse scene, where there is usually significant mechanical noise interference and relatively stable lighting conditions, acoustic environmental parameters have a more critical impact on the performance of the speech recognition module. Therefore, the weight of the background noise decibel parameter can be set to 0.6, the weight of the light intensity feature can be set to 0.3, and the weight of the spatial structure complexity can be set to 0.1.

[0072] Specifically, after obtaining the change coefficients of multiple environmental parameters, the coefficients are weighted and summed according to the matched weight proportions to finally output the current environmental change degree. This current environmental change degree is a comprehensive quantitative indicator used to reflect the overall change magnitude of the current environment relative to the historical environmental state, providing an important basis for the formulation of subsequent sensing parameter optimization strategies.

[0073] Furthermore, based on the current environmental variability and the current perception parameters, a perception parameter optimization strategy is configured, including:

[0074] Obtain the adaptation environment change step size corresponding to the current scene type, and set the ratio of the current environment change size to the adaptation environment change step size as the adaptation optimization depth;

[0075] The adaptation and optimization depth is multiplied by the initial up and down adjustment ratio of the preset perception parameters to construct the adaptation and perception parameter adjustment space.

[0076] The adaptation optimization depth and the adaptation perception parameter adjustment space are used as perception parameter optimization strategies.

[0077] First, obtain the adaptive environmental change step size corresponding to the current scene type. This adaptive environmental change step size is a preset baseline change amount used to standardize the magnitude of environmental changes. It is set according to the sensitivity requirements of different scene types to environmental changes. For example, in a laboratory environment requiring high-precision sensing, this step size can be set to 0.05, while in an industrial warehouse scene with large environmental fluctuations, this step size can be set to 0.1. Divide the calculated current environmental change degree by this adaptive environmental change step size; the ratio is defined as the adaptation optimization depth. This adaptation optimization depth characterizes the multiple relationship between the current environmental change and the baseline step size, determining the intensity of parameter optimization.

[0078] Secondly, based on the calculated adaptation optimization depth and the initial upward and downward adjustment ratios of the preset perception parameters, an adaptation perception parameter adjustment space is constructed. Specifically, the initial upward and downward adjustment ratios are multiplied by the adaptation optimization depth to obtain the actually usable parameter adjustment range. The preset perception parameters are a set of configurable variables that control the behavior of the robot's perception module, including parameters such as the visual OCR recognition threshold and the speech recognition filtering coefficient. The initial upward and downward adjustment ratios are set according to the parameter sensitivity and system stability requirements, for example, an initial value of 5%. The final adaptation perception parameter adjustment space defines the boundaries that each perception parameter is allowed to be adjusted during the optimization process.

[0079] Finally, the adaptation optimization depth and the adaptation perception parameter adjustment space are considered as the core components of the perception parameter optimization strategy. This strategy clarifies the intensity and range limits of parameter optimization, providing specific execution guidance for the subsequent online optimization process.

[0080] S50: With the goal of minimizing perception bias, based on the current scene type and current environment data, perform online optimization of perception parameters according to the perception parameter optimization strategy, and obtain suitable perception parameters for task recognition.

[0081] Specifically, such as Figure 2 As shown, with the goal of minimizing perception bias, based on the current scene type and current environment data, perception parameters are optimized online according to the perception parameter optimization strategy to obtain adaptive perception parameters, including:

[0082] Within the adaptive perception parameter adjustment space, a first initial perception parameter is randomly selected, and the first initial perception parameter is combined with the current scene type and the current environment data to obtain a first perception scheme;

[0083] Based on the aforementioned adaptation and optimization depth call perception deviation prediction plugin, the first predicted perception deviation is obtained according to the first perception scheme.

[0084] Continue iteratively selecting initial sensing parameters and iteratively predicting sensing deviations within the adaptation sensing parameter adjustment space until the adaptation optimization convergence number is reached. Output the initial sensing parameter corresponding to the minimum predicted sensing deviation during the optimization process and set it as the adaptation sensing parameter. The adaptation optimization convergence number is the product of the adaptation optimization depth and the standard optimization convergence number, and the standard optimization convergence number is no greater than 100.

[0085] First, within the established adaptive perception parameter adjustment space, a set of parameter values ​​is randomly selected as the first initial perception parameters. These first initial perception parameters are then combined with the current scene type and current environmental data to form a complete and evaluable first perception scheme, which serves as the starting point for subsequent iterative optimization processes.

[0086] Secondly, based on the adaptation optimization depth calculated above, a dedicated perception bias prediction plugin is invoked. This plugin is an integrated component of pre-trained machine learning models. By analyzing the matching relationship between the perception scheme and the scene environment, it predicts the perception performance of a specific parameter configuration under given environmental conditions. The first perception scheme is input into the perception bias prediction plugin, and through the model's forward computation, the corresponding first predicted perception bias is obtained. This first predicted perception bias is a quantified numerical indicator used to evaluate the expected recognition error of the first perception scheme, characterizing the theoretical performance level of this set of parameter configurations in the current scene environment.

[0087] Specifically, based on the adaptation and optimization depth, the perception deviation prediction plugin is invoked to predict the first predicted perception deviation according to the first perception scheme, including:

[0088] Based on the historical operation records of similar embodied robots, a sample perception scheme set and a sample perception deviation set were collected as a sample dataset.

[0089] The sample dataset is divided into K equal parts, and the first training set is obtained by selecting K times with replacement. The first training set is obtained by iteratively selecting K times, where K is greater than or equal to 20.

[0090] The deep learning models are trained to convergence using the K training sets respectively, resulting in K perceptual bias prediction branches. These branches are then integrated and constructed using a mean fusion strategy to build a perceptual bias prediction plugin.

[0091] The number of adaptation branches selected is P, which is obtained by rounding down the product of the adaptation optimization depth and the number of standard branches selected. The number of standard branches selected is 3, and P is greater than or equal to 3 and less than or equal to K.

[0092] P perception deviation prediction branches are randomly selected from the K perception deviation prediction branches to predict the first perception scheme and output the first predicted perception deviation.

[0093] First, the foundational data for training the perception bias prediction plugin is constructed. Specifically, based on the historical operational records of similar embodied robots, a set of sample perception schemes and their corresponding actual perception bias sets are collected to form a sample dataset. The sample perception schemes include historically used combinations of perception parameters, corresponding scene types, and environmental data; the sample perception bias sets record the actual perception errors generated by the corresponding perception schemes during actual operation.

[0094] Secondly, the complete sample dataset is divided into K subsets, where K is an integer not less than 20. Then, K random samplings with replacement are performed, each forming a new training set. This process is repeated K times to obtain K distinct training sets.

[0095] Furthermore, K deep learning models with identical structures are trained on K training sets until convergence, resulting in K perceptual bias prediction branches. The deep learning models are regression models based on multilayer perceptrons, including an input layer, multiple fully connected hidden layers, and an output layer. The convergence criterion is set based on the change in the loss function during training; for example, the model is considered convergent when the mean squared error loss on the validation set decreases by no more than 0.5% over 20 consecutive training epochs. The resulting K perceptual bias prediction branches are then integrated using a mean fusion strategy to form a perceptual bias prediction plugin. Specifically, the mean fusion strategy combines the outputs of multiple perceptual bias prediction branches and takes the average value during prediction to improve the stability and accuracy of the prediction.

[0096] Furthermore, based on the current required computational accuracy and efficiency, the number of branches participating in the prediction is dynamically determined. The number of adapted branches selected (P) is obtained by multiplying the adaptation optimization depth by the number of standard branches selected and rounding the result. Here, the number of standard branches selected is set to 3, therefore P is an integer not less than 3 and not exceeding K. A greater adaptation optimization depth indicates more significant environmental changes, and the number of adapted branches (P) participating in the prediction increases accordingly to provide more accurate bias predictions.

[0097] Finally, P perception deviation prediction branches are randomly selected from the K perception deviation prediction branches. The first perception scheme is simultaneously input into these P perception deviation prediction branches, and each perception deviation prediction branch outputs its predicted perception deviation. The average value of all predicted perception deviations is calculated, and this average value is used as the final first predicted perception deviation output.

[0098] Furthermore, within the aforementioned space for adjusting the adaptive perception parameters, the initial perception parameters are randomly selected repeatedly and combined with the current scene information to form a new perception scheme. Then, the perception deviation prediction plugin is invoked to obtain a new predicted perception deviation. This selection and prediction process will be executed iteratively.

[0099] Specifically, this iterative process will continue until the adaptive optimization convergence number is reached. This adaptive optimization convergence number is a dynamically determined integer value, representing the upper limit of the maximum number of iterations in the parameter optimization process. It is determined by multiplying the adaptive optimization depth by the preset standard optimization convergence number, where the standard optimization convergence number is a fixed value, representing the typical number of iterations required under the degree of change in the baseline environment, and its value is set to be no greater than 100.

[0100] Finally, after all iterations are completed, the predicted perception deviation values ​​obtained in all iterations are compared, and the initial perception parameter corresponding to the smallest predicted perception deviation is selected and formally determined as the final adapted perception parameter. Specifically, since the predicted perception deviation represents the expected recognition error of the perception scheme under given environmental conditions, the larger the predicted perception deviation, the worse the matching degree between the parameter configuration and the current environment, and the less ideal the expected perception performance; the smaller the predicted perception deviation, the better the matching degree between the parameter configuration and the current environment, and the better the expected perception performance. Therefore, the initial perception parameter corresponding to the smallest predicted perception deviation is selected as the adapted perception parameter, which is the optimal solution sought in the online optimization process.

[0101] In summary, the embodiments of this application have at least the following technical effects:

[0102] Compared to existing technologies, this application firstly achieves rapid adaptation and precise parameter tuning across different scenarios by constructing a scenario-type-driven hierarchical decision-making mechanism. When the scenario type changes, it can immediately match offline-optimized standard parameters to ensure basic perception performance; when in the same scenario, it initiates online fine-tuning based on multimodal environmental data to effectively cope with dynamic changes within the environment. Secondly, it introduces quantified environmental variability as an optimization basis. Through weighted fusion calculation of multidimensional environmental parameters, it transforms complex environmental states into quantifiable adjustment indicators, giving the parameter optimization process a clear guiding direction and adjustment range, overcoming the blindness of relying on experience for parameter tuning in traditional methods.

[0103] Furthermore, an integrated perception bias prediction model and a dynamic branch selection strategy are employed to significantly improve computational efficiency while maintaining prediction accuracy. Computational resources can be adaptively configured based on the degree of environmental change, achieving a good balance between optimization performance and operational overhead. Finally, by constructing a perception parameter adjustment space and an iterative optimization mechanism, efficient parameter search is performed within a limited range, ensuring rapid convergence to a near-optimal solution within a finite number of iterations. This comprehensively enhances the perception robustness and task execution reliability of the embodied robot in complex and ever-changing environments.

[0104] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0105] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0106] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A method for online optimization of perception parameters of an embodied robot for cross-scenario adaptation, characterized in that the method... include: Obtain the historical scene type and historical environment data corresponding to the current perception parameters of the embodied robot; The current scene type is determined by acquiring images of the current scene through the onboard image sensor of the embodied robot; If the current scene type and the historical scene type are the same, the current voice data is collected by the onboard voice sensor of the embodied robot, and the current environmental data is determined by analyzing the current scene type and the current voice data. The degree of change in the current environment is calculated based on the historical environmental data and the current environmental data, and a perception parameter optimization strategy is configured based on the degree of change in the current environment and the current perception parameters. With the goal of minimizing perception bias, based on the current scene type and current environment data, perception parameters are optimized online according to the perception parameter optimization strategy to obtain suitable perception parameters for task recognition. If the current scene type and the historical scene type are different, the environmental complexity is evaluated based on the current scene image and the current voice data to obtain the current environmental complexity. Call the preset scene type-perception parameter lookup table, and obtain standard perception parameters based on the current scene type and current environment complexity to perform task recognition; The environmental complexity evaluation based on the current scene image and current voice data includes: Convolutional neural networks are used to construct image feature recognizers and speech feature recognizers; The image feature recognizer extracts features from the current scene image and evaluates the visual complexity based on the visual features, including illumination uniformity, background clutter, and text recognition difficulty. The speech feature recognizer extracts features from the current speech data and evaluates the speech complexity based on the speech features. The speech features include the signal-to-noise ratio and the proportion of burst noise. The current environment complexity is calculated by weighting the visual complexity and the speech complexity.

2. The method of claim 1, wherein, The historical scene type and historical environment data corresponding to the current perception parameters of the embodied robot are obtained. The embodied robot is a general-purpose handling robot. The perception parameters include visual OCR parameters and speech recognition parameters. The scene type includes at least residential indoor environment, laboratory, logistics sorting center and industrial warehouse. The environment data includes visual environment data and acoustic environment data. The environmental parameters of the visual environment data include at least light intensity characteristics, dominant color spectrum, spatial structure complexity and key object list. The environmental parameters of the acoustic environment data include at least background noise decibels, noise type characteristics and reverberation time.

3. The online optimization method for embodied robot perception parameters for cross-scene adaptation according to claim 1, characterized in that, The robot's onboard voice sensor collects current voice data, and the current environmental data is determined based on the current scene type and the current voice data analysis, including: The onboard voice sensor of the embodied robot collects the voice data sequence of the current area within a preset time zone; Within the preset environment analysis model library, an appropriate environment analysis model is obtained by matching the current scene type. Using the adaptive environment analysis model, the current environment data is determined based on the analysis of the speech data sequence.

4. The online optimization method for perception parameters of an embodied robot oriented towards cross-scene adaptation as described in claim 3, characterized in that, The training method for the adaptive environment analysis model includes: Using the current scene type as a constraint, a sample speech data sequence set is collected, and historical environmental data corresponding to different sample speech data sequences are collected as sample environmental data to obtain a sample environmental dataset; Using the sample speech data sequence set as input and the sample environment dataset as supervision, a convolutional neural network is trained until convergence to obtain an adaptive environment analysis model.

5. The online optimization method for perception parameters of an embodied robot oriented towards cross-scene adaptation according to claim 1, characterized in that, The degree of change in the current environment is calculated based on the historical and current environmental data, including: Multidimensional deviation calculation is performed on the historical environmental data and the current environmental data. The ratio of the environmental deviation value to the corresponding historical environmental parameter value in the historical environmental data is set as the environmental parameter change coefficient, and multiple environmental parameter change coefficients are obtained. Based on the current scene type matching, the weight ratio of the adaptive environment parameters is obtained, and the change coefficients of the multiple environment parameters are weighted and calculated to obtain the current environment change degree.

6. The online optimization method for perception parameters of an embodied robot oriented towards cross-scene adaptation according to claim 1, characterized in that, Based on the current environmental variability and current sensing parameters, a sensing parameter optimization strategy is configured, including: Obtain the adaptation environment change step size corresponding to the current scene type, and set the ratio of the current environment change size to the adaptation environment change step size as the adaptation optimization depth; The adaptation and optimization depth is multiplied by the initial up and down adjustment ratio of the preset perception parameters to construct the adaptation and perception parameter adjustment space. The adaptation optimization depth and the adaptation perception parameter adjustment space are used as perception parameter optimization strategies.

7. The online optimization method for perception parameters of an embodied robot oriented towards cross-scene adaptation according to claim 6, characterized in that, With the goal of minimizing perception bias, based on the current scene type and current environment data, perception parameters are optimized online according to the perception parameter optimization strategy to obtain adaptive perception parameters, including: Within the adaptive perception parameter adjustment space, a first initial perception parameter is randomly selected, and the first initial perception parameter is combined with the current scene type and the current environment data to obtain a first perception scheme; Based on the aforementioned adaptation and optimization depth call perception deviation prediction plugin, the first predicted perception deviation is obtained according to the first perception scheme. Continue iteratively selecting initial sensing parameters and iteratively predicting sensing deviations within the adaptation sensing parameter adjustment space until the adaptation optimization convergence number is reached. Output the initial sensing parameter corresponding to the minimum predicted sensing deviation during the optimization process and set it as the adaptation sensing parameter. The adaptation optimization convergence number is the product of the adaptation optimization depth and the standard optimization convergence number, and the standard optimization convergence number is no greater than 100.

8. The online optimization method for perception parameters of an embodied robot oriented towards cross-scene adaptation according to claim 7, characterized in that, Based on the aforementioned adaptation and optimization depth call perception deviation prediction plugin, a first predicted perception deviation is obtained according to the first perception scheme, including: Based on the historical operation records of similar embodied robots, a sample perception scheme set and a sample perception deviation set were collected as a sample dataset. The sample dataset is divided into K equal parts, and the first training set is obtained by selecting K times with replacement. The first training set is obtained by iteratively selecting K times, where K is greater than or equal to 20. The deep learning models are trained to convergence using the K training sets respectively, resulting in K perceptual bias prediction branches. These branches are then integrated and constructed using a mean fusion strategy to build a perceptual bias prediction plugin. The number of adaptation branches selected is P, which is obtained by rounding down the product of the adaptation optimization depth and the number of standard branches selected. The number of standard branches selected is 3, and P is greater than or equal to 3 and less than or equal to K. P perception deviation prediction branches are randomly selected from the K perception deviation prediction branches to predict the first perception scheme and output the first predicted perception deviation.

Citation Information

Patent Citations

  • Task driven dynamic adaptive environment sensing mobile robot, system and method

    CN109445313A

  • Coal mine subsurface environment sensing method and apparatus, and storage medium

    WO2024234503A1