Industrial equipment action parameter generation method and device, equipment and storage medium

By acquiring natural language instructions and visual data, and using target industrial models and visual-physical inversion functions to generate motion parameters adapted to industrial equipment, the problems of long development cycles and poor flexibility in existing technologies are solved, and efficient automation and flexible production of industrial equipment are realized.

CN122132819APending Publication Date: 2026-06-02INST OF ADVANCED TECH UNIV OF SCI & TECH OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF ADVANCED TECH UNIV OF SCI & TECH OF CHINA
Filing Date
2026-05-08
Publication Date
2026-06-02

Smart Images

  • Figure CN122132819A_ABST
    Figure CN122132819A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and storage medium for generating motion parameters for industrial equipment, belonging to the field of artificial intelligence technology. The method includes: acquiring industrial instructions described in natural language for an industrial scenario; obtaining multimodal data based on visual data of the industrial scenario and state data of the industrial equipment; extracting features from the multimodal data using a target industrial model to obtain multimodal feature vectors; predicting the position data of the target object corresponding to the industrial instructions from the visual data based on the multimodal feature vectors; mapping the position data to three-dimensional pose data of the target object using a visual-physical inversion function; and generating motion parameters corresponding to the industrial instructions based on the three-dimensional pose data, so that the industrial equipment can execute the industrial instructions according to the motion parameters. This application achieves automatic conversion from natural language industrial instructions to industrial equipment motion parameters, which helps to improve the automation level of industrial equipment executing industrial instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and in particular relates to a method, apparatus, device and storage medium for generating motion parameters of industrial equipment. Background Technology

[0002] In fields such as electronics manufacturing, precision assembly and appearance inspection are core components of production line automation. With the rapid pace of product upgrades in consumer electronics, the demand for flexible manufacturing capabilities is increasing daily. Therefore, the ability to quickly determine the actions of industrial equipment and execute control commands based on different task requirements has become a crucial direction for the development of industrial automation.

[0003] Currently, the level of intelligence in automated production lines is limited, still heavily reliant on engineers manually teaching points, configuring parameters, and writing control logic. In recent years, with the development of artificial intelligence technology, model-driven industrial task processing solutions have emerged. These solutions utilize models to analyze natural language task descriptions, visual perception information, or on-site status information to assist in generating task decision results during equipment control.

[0004] However, traditional methods relying on manual teaching and fixed programming suffer from long development cycles, high debugging costs, and extremely poor flexibility. When product dimensions are adjusted or process standards change, the original program becomes ineffective, often requiring parameter resetting and program modification, making it difficult to meet the demands of flexible production. While existing model-driven solutions can improve natural language understanding to some extent, they lack a realistic understanding of the industrial physical world, failing to accurately convert natural language commands into executable device parameters, and further hindering the generation of standardized engineering documents that can be directly parsed and executed by industrial control systems. Summary of the Invention

[0005] This application aims to address at least one of the technical problems existing in the related art. To this end, this application proposes a method, apparatus, device, and storage medium for generating motion parameters of industrial equipment, which realizes the automatic conversion from natural language industrial instructions to motion parameters of industrial equipment, and helps to improve the automation level of industrial equipment in executing industrial instructions.

[0006] In a first aspect, this application provides a method for generating motion parameters of industrial equipment, the method comprising: Obtain industrial instructions described in natural language for industrial scenarios, and obtain multimodal data corresponding to the industrial instructions based on visual data of the industrial scenario and status data of industrial equipment in the industrial scenario. Feature extraction is performed on the multimodal data using the target industrial model to obtain a multimodal feature vector; Based on the multimodal feature vector, predict the position data of the target object corresponding to the industrial command from the visual data; The position data is mapped to the three-dimensional pose data of the target object using a visual physics inversion function, and motion parameters corresponding to the industrial command are generated based on the three-dimensional pose data, so that the industrial equipment can execute the industrial command according to the motion parameters.

[0007] In the above technical solution, multimodal data is generated by integrating industrial scene visual data and industrial equipment status data with natural language industrial commands as input. Multimodal features are extracted through the target industrial model, and the position data of the target object is accurately predicted from the visual data. Then, the position data is mapped to three-dimensional pose data through the visual-physical inversion function to generate motion parameters adapted to the industrial equipment. This realizes an automatic conversion process from natural language commands to industrial equipment motion parameters, improves the shortcomings of general large models in industrial scenarios, and realizes model-driven industrial automation (generation of industrial equipment motion parameters). The visual-physical inversion function realizes accurate inversion from semantic space to physical motion space. Even if there are phenomena such as product size adjustment or process standard change, the motion parameters generated based on three-dimensional pose data and industrial command requirements can automatically and accurately adapt to the operating status of industrial equipment and process requirements of industrial scenarios without extensive manual intervention for program rewriting. This meets the needs of flexible industrial production and improves the efficiency and automation of industrial equipment in executing industrial commands.

[0008] Secondly, this application provides an industrial equipment motion parameter generation device, the device comprising: The acquisition module is used to acquire industrial instructions described in natural language for industrial scenarios, and to obtain multimodal data corresponding to the industrial instructions based on the visual data of the industrial scenario and the status data of the industrial equipment in the industrial scenario. The extraction module is used to extract features from the multimodal data using the target industrial model to obtain a multimodal feature vector; The prediction module is used to predict the position data of the target object corresponding to the industrial command from the visual data based on the multimodal feature vector; The generation module is used to map the position data into the three-dimensional pose data of the target object through a visual physics inversion function, and generate motion parameters corresponding to the industrial command based on the three-dimensional pose data, so that the industrial equipment can execute the industrial command according to the motion parameters.

[0009] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the industrial equipment motion parameter generation method as described in the first aspect above.

[0010] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the industrial equipment motion parameter generation method as described in the first aspect above.

[0011] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the industrial equipment motion parameter generation method as described in the first aspect.

[0012] In a sixth aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the industrial equipment motion parameter generation method as described in the first aspect above.

[0013] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0014] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is one of the flowcharts illustrating the method for generating motion parameters of industrial equipment provided in some embodiments of this application; Figure 2 This is a second schematic flowchart of a method for generating motion parameters of industrial equipment provided in some embodiments of this application; Figure 3 This is a schematic diagram of the structure of an industrial equipment motion parameter generation system provided in some embodiments of this application; Figure 4 This is a schematic diagram of the structure of an industrial equipment motion parameter generation device provided in some embodiments of this application; Figure 5 These are schematic diagrams of the structure of electronic devices provided in some embodiments of this application.

[0015] Explanation of reference numerals in the attached figures: 400: Industrial equipment motion parameter generation device; 401: Acquisition module; 402: Extraction module; 403: Prediction module; 404: Generation module; 500: Electronic device; 501: Processor; 502: Memory. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0017] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0018] The following description, in conjunction with the accompanying drawings, details the industrial equipment motion parameter generation method, apparatus, equipment, and storage medium provided in this application through specific embodiments and application scenarios.

[0019] The method for generating motion parameters for industrial equipment can be applied to a terminal, specifically executed by the hardware or software within the terminal.

[0020] The industrial equipment motion parameter generation method provided in this application embodiment can be executed by an electronic device or a functional module or functional entity in an electronic device that can implement the industrial equipment motion parameter generation method. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras and wearable devices. The industrial equipment motion parameter generation method provided in this application embodiment is described below using an electronic device as the execution subject as an example.

[0021] Figure 1 This is one of the flowcharts illustrating a method for generating motion parameters of industrial equipment provided in some embodiments of this application. For example... Figure 1 As shown, the method for generating motion parameters of industrial equipment includes steps 110, 120, 130 and 140.

[0022] Step 110: Obtain industrial instructions described in natural language for industrial scenarios, and obtain multimodal data corresponding to the industrial instructions based on visual data of industrial scenarios and status data of industrial equipment in industrial scenarios.

[0023] In this embodiment of the application, industrial instructions are operation instructions described in natural language and designed for industrial scenarios (such as electronic equipment assembly, appearance inspection, and parts sorting).

[0024] Visual data refers to visual data used to characterize the state of an industrial scene (including industrial equipment, target objects, and the working environment of the equipment within the scene). In some embodiments, visual data includes image data (such as two-dimensional color images of the scene taken by an industrial camera) and point cloud data.

[0025] The status data of industrial equipment includes the operating status data of industrial equipment (such as robotic arms, conveyor belts, cylinders, etc.), which may include the posture data, operation data, load data, etc., so that the subsequent generated motion parameters can be adapted to the current status of the industrial equipment.

[0026] Multimodal data refers to a data set obtained by integrating one or more of industrial instructions, visual data, and industrial equipment status data. It can comprehensively represent industrial instruction requirements, scene environment, and equipment status, providing a comprehensive data foundation for subsequent feature extraction and location prediction.

[0027] For example, industrial instructions could be "affix component labels to the target surface of product model A" or "obtain the physical location of the lower left corner of the target surface of product model A." Industrial instructions can be received through the human-machine interface of an industrial control platform in an industrial setting, or through a voice recognition module, etc. This application does not specifically limit the form of obtaining industrial instructions. In some embodiments, visual data (image data and point cloud data) of the industrial scene are simultaneously collected using 3D vision devices such as cameras, and status data of the industrial equipment (such as the current pose and operating parameters of the robotic arm) are collected using equipment sensors. A data fusion algorithm is then used to integrate these three types of data into multimodal data.

[0028] The following is a specific example of multimodal data: Multimodal data include:

[0029] in, Industrial instructions described using natural language; Visual data for industrial scenarios, among which For color image data, Point cloud data containing depth information and geometric features; It represents the real-time state vector (state data) of industrial equipment, which may include, for example, the joint angle of the robotic arm, the pose of the end effector, and the input and output states of the programmable logic controller.

[0030] Step 120: Extract features from the multimodal data using the target industrial model to obtain multimodal feature vectors.

[0031] A target industrial model refers to a model trained on industrial scenario data and adapted to specific industrial scenarios or needs. In some embodiments, the target industrial model is used for feature extraction from multimodal data, possesses natural language understanding capabilities, and has a realistic perception of the industrial physical world.

[0032] For example, the target industrial model can be pre-tuned based on a general model using data relevant to industrial scenarios to enhance its adaptability to those scenarios. Feature extraction from the target industrial model transforms multi-source heterogeneous data of different modalities into standardized feature vectors, eliminating modal differences. Furthermore, as the target industrial model is adapted to industrial scenarios, its adaptability can be integrated during feature extraction, ensuring that the extracted multi-modal feature vectors accurately represent industrial command requirements, target object features, and equipment status, providing accurate data support for subsequent location data prediction and motion parameter generation.

[0033] The following is a specific example of feature extraction:

[0034] in, For multimodal feature vectors, Industrial instructions described using natural language. Visual data for industrial scenarios, Represents the real-time state vector (state data) of industrial equipment; the target industrial model uses an encoder ( Multimodal feature vectors are obtained by extracting features in the form of ().

[0035] Step 130: Based on multimodal feature vectors, predict the position data of the target object corresponding to the industrial command from the visual data.

[0036] Location data is the spatial representation of a target object in visual data. Based on multimodal feature vectors, the location data of the target object corresponding to an industrial command can be predicted from visual data. For example, multimodal features can be input into the location prediction branch of the target industrial model. Combined with the spatial features of the visual data, a coordinate prediction algorithm can be used to predict the two-dimensional pixel coordinates of the target object in the image data. At the same time, a depth estimation algorithm can be used to predict the depth data of the target object from the point cloud data. The integrated data are then used to obtain the location data of the target object.

[0037] Prediction based on multimodal feature vectors, combined with the correlation between industrial instructions, equipment status and visual data, can improve the prediction accuracy of position data compared to prediction based on single visual data. The predicted position data provides accurate input for subsequent visual-physical inversion and 3D pose mapping.

[0038] Step 140: Map the position data to the three-dimensional pose data of the target object through the visual physics inversion function, and generate motion parameters corresponding to the industrial commands based on the three-dimensional pose data, so that the industrial equipment can execute the industrial commands according to the motion parameters.

[0039] The visual-physical inversion function is used to establish a mapping relationship from the visual feature space to the Euclidean physical space. In some embodiments, position data is input into the visual-physical inversion function, and the mapping from visual position data to 3D pose data is completed through a preset coordinate transformation formula. The 3D pose data can be optimized by combining the process constraints of industrial scenarios (such as the tolerance requirements of precision assembly), and then the 3D pose data can be transformed into action parameters that can be executed by industrial equipment through a parameter mapping algorithm.

[0040] By establishing an inversion mechanism between visual and physical action parameters (visual-physical inversion function), the mapping barrier between semantic space and physical space is broken through, transforming high-dimensional abstract natural language instructions into physical pose parameters in three-dimensional space. This solves the pain point that general large models cannot perceive the physical world and cannot realize semantic-to-physical pose transformation.

[0041] In some embodiments, the generated motion parameters can be further injected into industrial operator templates and generate files for industrial control systems to parse and execute. This eliminates the need for engineers to manually write code. Even if product dimensions are slightly adjusted or process standards are changed, the motion parameters generated based on 3D pose data and industrial instruction requirements can accurately adapt to the operating status of industrial equipment and the process requirements of industrial scenarios, enabling industrial equipment to accurately execute industrial instructions and improve the accuracy of industrial operations.

[0042] The industrial equipment motion parameter generation method provided in this application takes natural language industrial commands as input, integrates industrial scene visual data and industrial equipment status data to generate multimodal data, encodes it based on the target industrial model to obtain multimodal feature vectors, accurately predicts the position data of the target object from the visual data, and then maps the position data into three-dimensional pose data through a visual-physical inversion function to generate motion parameters adapted to the industrial equipment. This realizes an automatic conversion process from natural language commands to industrial equipment motion parameters, improves the shortcomings of general large models in industrial scenarios, and realizes model-driven industrial automation (industrial equipment motion parameter generation). The visual-physical inversion function achieves accurate inversion from semantic space to physical motion space. Even if there are phenomena such as product size adjustments or changes in process standards, the motion parameters generated based on three-dimensional pose data and industrial command requirements can automatically and accurately adapt to the operating status of industrial equipment and the process requirements of industrial scenarios without extensive manual intervention for program rewriting. This meets the needs of flexible industrial production and improves the efficiency and automation of industrial equipment in executing industrial commands.

[0043] In some embodiments of this application, the target industrial model is trained through the following steps: Based on industrial knowledge, the initial model is trained in the first stage through instruction fine-tuning to enhance its ability to understand industrial instructions and obtain a general industrial model. Based on multimodal data samples and corresponding action parameter samples, a second-stage training of the general industrial model is performed using the direct preference optimization algorithm to obtain a target industrial model adapted to the industrial scenario.

[0044] It is understood that the initial model refers to a basic model (such as a general multimodal model, language model, etc.) that has not been trained with industrial knowledge or fine-tuned for industrial scenarios. It has basic feature extraction and data processing capabilities. In this embodiment, it needs to be trained in two stages to adapt to industrial needs and industrial scenarios.

[0045] In the first phase of training, the initial model was fine-tuned based on industrial knowledge to enable it to understand industrial instructions. This established a preliminary mapping between industrial instructions and semantic understanding, laying the foundation for the second phase of multimodal data adaptation and action parameter generation, and solving the problem that the initial model did not understand industrial scenarios and instructions.

[0046] Industrial knowledge can include industrial production-related content such as mechanical structure principles, electrical control logic, and industrial task knowledge graphs. This is used to enhance the model's ability to understand industrial semantic information (industrial instructions) and guide the model to meet actual industrial needs.

[0047] For example, in the first stage of training, based on the initial model, industrial knowledge (such as electronic assembly process specifications and equipment operation logic) is injected, and typical industrial instruction texts are selected as training samples. The semantic understanding module of the model is adjusted through instruction fine-tuning to enhance the model's ability to recognize industrial instructions. Alternatively, common instruction types in the industrial field (such as assembly, testing, and debugging related instructions) can be integrated to build an instruction sample library. The initial model can then be fine-tuned in combination with industrial knowledge to optimize the model's ability to semantically parse industrial instructions.

[0048] The second stage of training, building upon the first stage, combines multimodal data samples and action parameter samples. It further optimizes the general industrial model using the Direct Preference Optimization (DPO) algorithm, enabling the general industrial model to accurately associate industrial instructions with multimodal data. Then, action parameters are generated from the multimodal data, achieving adaptation to industrial scenarios and solving the problem that the general industrial model obtained in the first stage of training cannot be combined with scenario data.

[0049] Multimodal data refers to the fusion of industrial instruction samples, visual data samples, and state data samples, used as input for the second-stage training. Action parameter samples refer to the action parameters corresponding to the industrial instruction samples, which can be used as label references for model training. For example, in the second-stage training, multimodal data samples and corresponding action parameter samples can be acquired to construct a mapping sample library; the mapping sample library is input into a general industrial model, and a direct preference optimization algorithm is used to compare high-quality and low-quality action parameter samples, adjust the weights of the general industrial model, and make the model prioritize outputting high-quality action parameters that conform to industrial standards. After training, the target industrial model is obtained.

[0050] By employing a progressive, layered fine-tuning approach that incorporates industrial knowledge injection and adapts to industrial scenarios, the large model is first infused with professional knowledge of industrial processes, mechanical and electrical systems, and then optimized for industrial scenarios by combining physical consistency constraints. This addresses the shortcomings of general large models that do not understand industrial processes and lack physical rule constraints.

[0051] The industrial equipment motion parameter generation method provided in this application, based on industrial knowledge, fine-tunes the initial model through instruction fine-tuning in the first stage of training to enable it to understand industrial instructions. Based on the general industrial model obtained in the first stage of training, the model is further optimized by combining multimodal data samples and motion parameter samples through a direct preference optimization algorithm to achieve accurate adaptation to the industrial scenario and obtain the target industrial model. Through the two-stage training combining basic adaptation and precise optimization, the trained target industrial model not only has the semantic understanding ability of industrial instructions, but also can generate compliant motion parameters by combining multimodal scenario data. This solves the problem that the general model does not understand the industrial scenario and the parameter output is inaccurate, while reducing the cost of manual intervention and achieving accurate connection between industrial instructions and equipment execution, providing reliable model support for subsequent motion parameter generation and engineering file output.

[0052] In some embodiments of this application, the multimodal data samples include industrial instruction samples, visual data samples, and state data samples; Based on multimodal data samples and corresponding action parameter samples, a second stage of training is performed on the general industrial model using the direct preference optimization algorithm to obtain a target industrial model adapted to the industrial scenario, including: For each multimodal data sample, the preference reward value for each corresponding action parameter sample is determined based on the state data sample. The preference reward value is used to characterize the degree to which the action parameter sample conforms to the physical consistency constraints of the industrial scenario. Based on the preference reward value of each action parameter sample, select the preferred action parameter samples and the unpreferred action parameter samples from the action parameter samples; Industrial instruction samples and visual data samples are used as input data samples for a general industrial model, and combined with corresponding preferred action parameter samples and undesirable action parameter samples to construct preference training sample pairs. Based on the preference training sample pairs, the general industrial model is iteratively updated using the direct preference optimization algorithm to ensure that the motion parameter samples generated from the multimodal feature vectors output by the general industrial model conform to the physical consistency constraints of the industrial scenario, thus obtaining the target industrial model.

[0053] Physical consistency constraints in industrial scenarios refer to the industrial physical laws, equipment operating limits, and process specifications that motion parameters must conform to. For example, the motion parameters of industrial equipment robotic arms must not exceed the equipment's range of motion, and precision assembly parameters must meet tolerance requirements to avoid equipment failures or operational errors.

[0054] The preference reward value is used to quantify the degree to which the action parameter sample is adapted to the industrial scenario. Generally, the higher the preference reward value, the more the action parameter sample conforms to the operating logic and process requirements of the industrial equipment. Conversely, the lower the value, the worse the adaptability. It is used as the data basis for screening the preferred action parameter sample and the inferior action parameter sample.

[0055] Based on the state data samples, the preferred reward value for each action parameter sample is determined. Specifically, this involves combining the physical consistency constraints of the industrial scenario to quantitatively evaluate the suitability of each action parameter sample and assigning a corresponding preferred reward value, thus providing a clear standard for subsequent sample selection. For example, this can be based on a pre-defined preferred reward value evaluation mechanism, combining equipment operating parameters (such as robotic arm load and range of motion) and industrial scenario physical consistency constraints (such as tolerance requirements) from the state data samples, setting multiple evaluation indicators (such as parameter suitability, equipment safety, and process compliance), assigning corresponding weights to each indicator, comparing the action parameter samples with the state data samples, and calculating the preferred reward value for each action parameter sample through weighted summation; using the state data samples and physical consistency constraints as input features, and "whether the action parameter sample is suitable for the industrial scenario" as a label, a reward value prediction model is trained. The action parameter sample to be evaluated and the corresponding state data sample are input into this model, and the corresponding preferred reward value is directly output.

[0056] By using preference reward values, physical consistency constraints are transformed into quantifiable values, making the judgment of the quality of action parameter samples more objective and providing accurate input data (preference differences) for the direct preference optimization algorithm.

[0057] Understandably, the preferred action parameter samples are those with higher preference reward values ​​(such as reaching a preset threshold), which conform to the physical consistency constraints of the industrial scenario and can be used as action parameter samples to guide industrial equipment to perform operations, and can be used as positive examples for training the target industrial model. On the other hand, the undesirable action parameter samples are those with lower preference reward values ​​(such as below a preset threshold), which do not fully conform to the physical consistency constraints and have problems such as poor equipment adaptability and operation exceeding the limits, and can be used as negative examples for model training.

[0058] For example, based on the preference reward value of each action parameter sample, the preferred action parameter samples and the unpreferred action parameter samples can be selected. This can be done by comparing the preference reward value with a threshold preset according to the process requirements of the industrial scenario or equipment parameters, and classifying them according to the comparison results. Alternatively, it can be done by combining the preference reward value distribution of all action parameter samples, classifying the action parameter samples at the top of the distribution as preferred action parameter samples, and classifying the samples at the bottom of the distribution as unpreferred action parameter samples.

[0059] In some embodiments, each set of preference training sample pairs includes a set of input data samples, a preferred action parameter sample, and a disadvantaged action parameter sample. In some embodiments, each set of input data samples may also be paired with multiple preferred action parameter samples and multiple disadvantaged action parameter samples. By enriching the comparison dimensions of the sample pairs, the model can more comprehensively learn the differences between high-quality and low-quality samples.

[0060] The general industrial model is iteratively updated using the direct preference optimization algorithm. Specifically, by inputting preference training sample pairs into the general industrial model, the direct preference optimization algorithm compares the differences between the optimal and unoptimized samples, continuously adjusting the model parameters so that the model's output action parameters gradually conform to the physical consistency constraints of the industrial scenario, ultimately completing the training from the general industrial model to the target industrial model. For example, the constructed preference training sample pairs are input into the general industrial model in batches. The direct preference optimization algorithm is used to calculate the deviation between the model's output action parameters and the optimal and unoptimized samples, iteratively updating the model to strengthen its tendency to output high-quality parameters. An iteration termination condition is set (such as the average preference reward value of the model's output parameters reaching a preset threshold or the number of iterations reaching an upper limit). After iteration, the target industrial model is obtained.

[0061] The following is a specific example of performing the second stage training of a general industrial model using the direct preference optimization algorithm: Set motion parameters Status data of industrial equipment The physical feedback reward is The alignment loss function is minimized using the direct preference optimization algorithm. :

[0062] in, This is the general industrial model that needs to be trained. Indicates from dataset The expected value of the sampled data; The input data sample consists of industrial instruction samples and visual data samples; This represents the dataset used for preference optimization training (consisting of preference training sample pairs); This indicates the preferred action parameter sample; This represents a sample of parameters for the undesirable action. This is the current model strategy; As a reference model strategy; For activation functions; The temperature parameter is used. Through the above training mechanism, the model output can meet the physical consistency constraints of industrial tasks.

[0063] The industrial equipment motion parameter generation method provided in this application is based on multimodal data samples and motion parameter samples. It selects preferred and unpredictable samples by quantifying preference reward values, constructs preference training sample pairs, and then iteratively updates the general industrial model through a direct preference optimization algorithm. Finally, it obtains a target industrial model that meets the physical consistency constraints of industrial scenarios and adapts to industrial needs. This ensures that the motion parameters output by the target industrial model strictly conform to the physical consistency constraints of industrial scenarios, improving the adaptability and accuracy of the motion parameters. It associates industrial instructions, visual data, equipment status, and motion parameters, enabling the target industrial model to not only understand industrial instructions but also output compliant parameters in combination with scenario data. It effectively overcomes the problem of general large models lacking physical constraints and process cognition in precision industrial operations, and significantly improves the generalization ability and robustness in dealing with complex unstructured environments while ensuring execution accuracy.

[0064] In some embodiments of this application, visual data includes image data and point cloud data; positional data includes two-dimensional pixel coordinates in the image data and depth data in the point cloud data; The position data is mapped to the 3D pose data of the target object using a visual-physical inversion function, and motion parameters corresponding to industrial commands are generated based on the 3D pose data, including: Using the visual physics inversion function, the three-dimensional pose data of the target object in the world coordinate system is determined based on the two-dimensional pixel coordinates and depth data of the target object. Based on 3D pose data and combined with prior spatial information of the industrial scene, motion parameters corresponding to industrial commands are generated.

[0065] Visual data includes image data and point cloud data. Image data can specifically be two-dimensional color images, which include information such as the planar position, shape, and texture of the target object. Point cloud data consists of a large number of three-dimensional points, each of which includes coordinate information in three dimensions. It can be used to characterize the spatial depth and three-dimensional shape of the target object, and is used to provide depth data of the target object.

[0066] Two-dimensional pixel coordinates correspond to the pixel positions of the target object in the image data, and are used to locate the specific position of the target object in the two-dimensional image; depth data corresponds to the spatial depth information in the point cloud data, which represents the actual distance from a point on the surface of the target object to the image acquisition device, and is used to supplement the spatial dimension information of the target object, thereby realizing the subsequent conversion from two-dimensional position to three-dimensional pose.

[0067] In some embodiments, location data (two-dimensional pixel coordinates and depth data) is used as input, and with the help of a visual physics inversion function, combined with image acquisition device parameters and coordinate transformation rules, two-dimensional visual information is transformed into three-dimensional pose data in the world coordinate system, thereby realizing the transformation from visual perception to physical spatial positioning.

[0068] 3D pose data is used to characterize the complete spatial state of a target object in physical space, and can include 3D position (such as the target object's coordinates in the world coordinate system) and attitude angles (the target object's pitch angle, yaw angle, roll angle, etc.). For example, 2D pixel coordinates and depth data can be directly input into a visual physics inversion function, and the corresponding coordinate transformation formula of the function can output the 3D position coordinates of the target object in the world coordinate system. Through the visual physics inversion function, the mapping from 2D position data to 3D pose data is realized, overcoming the shortcoming of general models that cannot perceive the industrial physical world. In addition, using the world coordinate system as a reference ensures the uniformity of 3D pose data, providing a foundation for subsequent integration with prior information of industrial scene space to generate motion parameters for adapted equipment.

[0069] Spatial prior information refers to pre-set spatial information in an industrial scenario, which may include the range of motion of industrial equipment, the standard dimensions and placement of target objects, the location of obstacles in the scenario, and process tolerance requirements. Based on 3D pose data and combined with the spatial prior information of the industrial scenario, motion parameters corresponding to industrial commands are generated. This can be achieved by extracting the 3D position coordinates and posture angles from the 3D pose data, combining them with the spatial prior information of the industrial scenario (such as the range of motion of the robotic arm and the assembly tolerance requirements of the target object), and using a parameter mapping algorithm to transform the 3D pose data into motion parameters of the industrial equipment (such as the movement path, rotation angle, and gripping force of the robotic arm); or by constructing a motion parameter generation model, taking 3D pose data, spatial prior information, and industrial commands as input, and using a pre-set rule base (such as motion parameter templates corresponding to different commands) and machine learning algorithms to automatically generate appropriate motion parameters.

[0070] The industrial equipment motion parameter generation method provided in this application uses a visual physics inversion function to map the two-dimensional pixel coordinates and depth data of a target object into three-dimensional pose data in the world coordinate system. Finally, it combines the spatial prior information of the industrial scene to generate motion parameters corresponding to industrial commands, realizing the mapping from two-dimensional position data to three-dimensional pose data and solving the problem that general models cannot perceive the industrial physical world. By combining the spatial prior information of the industrial scene, it avoids the motion parameters from being out of touch with the industrial scene or the operation of industrial equipment, making the generated motion parameters conform to the actual operating requirements of the industrial scene and improving the accuracy of motion parameter generation. In addition, the conversion process of motion parameters is completed automatically, reducing manual intervention and solving the problem of time-consuming and laborious manual positioning and coordinate recording in traditional industrial scenes, thus improving the efficiency of motion parameter generation.

[0071] In some embodiments of this application, the visual-physical inversion function includes an intrinsic parameter matrix of the image acquisition device for acquiring image data, a hand-eye calibration matrix for converting the coordinate system of the image acquisition device to the coordinate system of the industrial equipment, and a depth correction function for correcting depth data. Using the visual physics inversion function, the 3D pose data of the target object in the world coordinate system is determined based on the 2D pixel coordinates and depth data of the target object, including: Based on the intrinsic parameter matrix, the normalized coordinates of the target object in the coordinate system of the image acquisition device are determined according to the two-dimensional pixel coordinates. Based on the depth correction function, the depth correction coefficient of the target object is determined according to the depth data; Based on the normalized coordinates, depth data, and depth correction coefficient of the target object, determine the three-dimensional coordinates of the target object in the coordinate system of the image acquisition device; Based on the hand-eye calibration matrix, coordinate transformation is performed on the three-dimensional coordinates in the coordinate system of the image acquisition device to obtain the three-dimensional coordinates of the target object in the coordinate system of the industrial equipment. The three-dimensional coordinates of the target object in the coordinate system of the industrial equipment are used to determine the three-dimensional pose data of the target object in the world coordinate system.

[0072] Image acquisition equipment refers to devices used to acquire visual data (image data, point cloud data) from industrial scenes, such as industrial cameras and LiDAR. Its intrinsic parameter matrix directly affects the accuracy of position data extraction and the inversion effect of 3D pose. The intrinsic parameter matrix is ​​the inherent parameter matrix of the image acquisition equipment, which may include information such as focal length, pixel size, and principal point coordinates. It is used to eliminate image distortion and transform 2D pixel coordinates into normalized coordinates in the coordinate system of the image acquisition equipment.

[0073] Depth correction functions are used to correct errors in depth data within location data. They eliminate noise and depth measurement biases (such as light interference and equipment accuracy errors) in point cloud data, ensuring the accuracy of the depth data. For example, depth data can be input into the depth correction function, and abnormal noise data can be identified and removed through statistical analysis (such as calculating the variance and mean of the depth data). Then, combined with the measurement accuracy of the image acquisition equipment and the lighting conditions of the industrial scene, an error model is established, and the depth correction coefficient is calculated. The depth correction coefficient can effectively eliminate noise or measurement bias in depth data, avoiding large 3D coordinate errors that may result from inaccurate depth data.

[0074] Based on the normalized coordinates, depth data, and depth correction coefficient of the target object, the three-dimensional coordinates in the coordinate system of the image acquisition device are determined. Through coordinate calculation, the two-dimensional normalized coordinates are transformed into three-dimensional coordinates in the coordinate system of the image acquisition device, thus realizing the fusion of two-dimensional pixel coordinates and depth data.

[0075] The hand-eye calibration matrix is ​​used to realize coordinate transformation between the coordinate system of the image acquisition device and the coordinate system of the industrial equipment. For example, it can be obtained through hand-eye calibration experiments. It is used to accurately map the three-dimensional coordinates in the image acquisition device coordinate system to the industrial equipment coordinate system, ensuring that the coordinate system is compatible with the operation of the industrial equipment. The image acquisition device coordinate system refers to a three-dimensional coordinate system established with the image acquisition device (such as an industrial camera) as the origin, which can be used to represent the spatial position of the target object relative to the acquisition device. The industrial equipment coordinate system refers to a three-dimensional coordinate system established with the industrial equipment as the origin, which is used to represent the spatial position of the target object relative to the industrial equipment. This application does not specifically limit the setting method of the image acquisition device coordinate system and the industrial equipment coordinate system. The hand-eye calibration matrix solves the problem of inconsistent three-dimensional coordinates between the image acquisition device coordinate system and the industrial equipment coordinate system, which leads to the inability of the three-dimensional coordinates to adapt to the execution of equipment actions.

[0076] The 3D coordinates in the industrial equipment coordinate system are used to determine the 3D pose data of the target object in the world coordinate system. For example, it can be based on a pre-set transformation relationship between the industrial equipment coordinate system and the world coordinate system. The 3D coordinates in the industrial equipment coordinate system are substituted into the transformation formula to obtain the 3D position coordinates in the world coordinate system. Then, the pose angle of the target object is calculated by combining the stereo contour information of the point cloud data with the pose estimation algorithm. After integration, the complete 3D pose data is obtained. Alternatively, the 3D coordinates in the industrial equipment coordinate system and the point cloud data can be used as input. Through the pre-set world coordinate system mapping rules and the combination of machine learning algorithms, the coordinate transformation and pose calculation can be completed automatically, and the 3D pose data in the world coordinate system can be directly output.

[0077] Here is a specific example:

[0078] in, The three-dimensional coordinates of the target object in the world coordinate system. For hand-eye calibration matrix, For depth correction function, This is the intrinsic parameter matrix of the image acquisition device. These are the two-dimensional pixel coordinates of the target object in the image data. This refers to the depth data of the target object within the point cloud data.

[0079] The industrial equipment motion parameter generation method provided in this application includes a visual-physical inversion function that comprises at least an intrinsic parameter matrix, a hand-eye calibration matrix, and a depth correction function. Through error correction and coordinate transformation (hand-eye calibration matrix), the inversion error of 3D pose data is minimized, and the accuracy of 3D pose data is improved. This method is suitable for industrial scenarios with high positioning accuracy requirements, such as electronic precision assembly and parts sorting, and helps to solve problems such as large positioning errors in general models. By constructing a visual-physical inversion function, the mapping between semantic space and physical space is realized, mapping high-dimensional abstract natural language instructions into physical pose parameters in 3D space. This solves the problems of time-consuming, laborious, and error-prone manual positioning and coordinate recording in traditional industrial scenarios.

[0080] In some embodiments of this application, the method further includes: Construct a syntax rule tree corresponding to the industrial scenario. The syntax rule tree is used to represent the syntax constraints for generating action parameters corresponding to industrial instructions in a tree form. Generate a corresponding mask matrix based on the syntax rule tree; the mask matrix is ​​used to filter the action parameters during the generation process so that the action parameters conform to the syntax constraints.

[0081] Understandably, a syntax rule tree is used to represent the grammatical constraints for generating action parameters in industrial scenarios in a tree-like form. In some embodiments, the root node of the syntax rule tree is the overall grammatical constraint target, the intermediate nodes are classification constraints (such as device type and instruction type constraints), and the leaf nodes are specific constraint clauses. This is used to transform abstract grammatical constraints into structured and parsable rules, thereby ensuring that the generation of action parameters conforms to industrial scenario specifications.

[0082] The grammatical constraints corresponding to industrial scenarios are action parameter constraint rules formulated based on the process requirements, equipment operation logic, and safety specifications of the industrial scenario. Specifically, they are used to limit the format, value range, and logical relationships between action parameters.

[0083] In some embodiments, an initial syntax rule tree is generated based on the equipment parameters and process specifications of the industrial scenario; and the constraint level and constraint clauses can be adjusted according to actual production needs.

[0084] A mask matrix is ​​a two-dimensional matrix generated from a syntax rule tree. It is used to perform filtering and selection operations during the generation of action parameters. The matrix elements can correspond to different dimensions of the action parameters. By using a mask (to block out parameter dimensions that do not meet the constraints and to filter out illegal values), the output action parameters are ensured to meet the syntax constraints.

[0085] The corresponding mask matrix is ​​generated based on the syntax rule tree. This can be done by parsing the node constraint clauses of the syntax rule tree to determine the dimensions of the action parameters (such as movement distance, rotation angle, gripping force, etc.) and constructing a two-dimensional mask matrix (where rows correspond to the dimensions of the action parameters and columns correspond to the constraint types). The matrix element values ​​are set according to the constraint clauses (e.g., the element corresponding to the dimension that meets the constraint is 1, and the element corresponding to the dimension that does not meet the constraint is 0). During the parameter generation process, only the parameter values ​​and dimensions with matrix elements of 1 (or whose weights meet the requirements) are retained.

[0086] The industrial equipment motion parameter generation method provided in this application constructs a syntax rule tree corresponding to the industrial scenario, transforming abstract motion parameter syntax constraints into structured, hierarchical tree rules. The syntax rule tree is then converted into a mask matrix to filter parameters during the motion parameter generation process. The syntax rule tree matches the actual needs of the industrial scenario, integrating constraints such as the operating limits of industrial equipment and safety regulations of the industrial scenario. This ensures that the subsequently generated motion parameters are adaptable to industrial equipment and meet production requirements, guaranteeing the compliance of the motion parameters. Real-time filtering via the mask matrix during motion parameter generation can promptly block illegal parameters, preventing them from being output to the industrial equipment, reducing the risk of industrial equipment malfunctions or operational errors, and improving the reliability of motion parameter generation.

[0087] In some embodiments of this application, the method further includes: The action parameters are injected into an industrial operator template that conforms to the syntax rule tree, and a structured engineering file is generated by combining it with a mask matrix. This allows the industrial equipment to execute industrial instructions according to the action parameters by calling the structured engineering file.

[0088] In some embodiments, the method further includes: constructing multiple sets of industrial operator templates (e.g., grabbing instruction templates, placement instruction templates, and detection instruction templates) adapted to different industrial instructions based on the industrial scenario and the syntax rule tree. Industrial operators are standardized, reusable basic operation units in an industrial scenario that implement specific industrial functions (such as equipment control, data processing, and process execution). The industrial operator template can serve as a link between action parameters and structured engineering files. It contains built-in operators recognizable by industrial equipment (such as movement operators, grabbing operators, and stop operators), parameter placeholders, and syntax formats, and pre-sets constraint logic corresponding to the syntax rule tree to ensure that the injected action parameters conform to the equipment execution specifications.

[0089] For example, a structured engineering file is a standardized file that includes action parameters, industrial operators, syntax constraint information, etc. It usually has a fixed format (such as being compatible with industrial equipment control protocols) and can be directly called by industrial control platforms to convert action parameters into instruction files that can be recognized and executed by industrial equipment.

[0090] In some embodiments, structured engineering files are used to invoke industrial operators to drive various industrial devices to execute action parameters in order to carry out industrial instructions.

[0091] Injecting action parameters into an industrial operator template that conforms to the syntax rule tree can be done in two ways: first, by injecting the parameters one by one into the corresponding positions according to the template placeholders, while verifying the compatibility between the parameters and the template syntax, thus forming the injected industrial operator file; second, by automatically selecting suitable industrial operator templates based on the industrial instruction type and syntax rule tree constraints. Through a parameter mapping mechanism, action parameters are automatically injected into the corresponding placeholders according to the parameter format and order of the template, eliminating the need for manual matching.

[0092] The mask matrix is ​​used to generate structured engineering files. For example, the industrial operator template after injecting action parameters can be compared and verified with the mask matrix. The mask matrix can be used to filter out content in the template that does not conform to the syntax constraints, so that the template content fully conforms to the constraints.

[0093] Here is a specific example:

[0094] in, This is a structured engineering file containing a sequence of industrial operator calls; It is a multimodal feature vector; It is a mask matrix; For action parameters; This refers to injecting action parameters into an industrial operator template that conforms to a syntax rule tree.

[0095] The industrial equipment action parameter generation method provided in this application injects the action parameters filtered according to the syntax rule tree into an industrial operator template that conforms to the syntax rule tree constraints, and verifies them through a mask matrix, ultimately generating a standardized structured engineering file. This realizes the conversion from natural language instructions to executable action parameters of the equipment, and further generates standardized engineering files that can be directly parsed and executed by the industrial control system. It automates parameter injection, engineering file generation and calling, eliminating the need for manual writing of equipment control code, greatly reducing the cost of manual intervention, and effectively improving the efficiency and automation of industrial equipment in executing industrial instructions.

[0096] In some embodiments of this application, the method for generating motion parameters of industrial equipment is applied to the actual engineering scenario of "attaching specific component labels to the exterior parts of a certain model of product" in an electronic manufacturing production line. The method for generating motion parameters of industrial equipment also includes: Step 1: Acquisition of multi-source heterogeneous input data and state modeling: Acquiring industrial instructions in natural language form For example: "Attach one component label to the target surface of product A" and "Obtain the physical location of the lower left corner of the target surface of product A". Obtain formatted real-time device state vectors through the Manufacturing Execution System (MES) interface. It also analyzes the process attributes of the target product, including the model identifier modelName="Model_A", the label part number labelPn="Label_PN_01", the target mounting surface cover="Target_Surface", and the theoretical layout coordinates; Use visual platforms to acquire observation data It includes initial images (image data) captured by the upper and lower cameras and corresponding point cloud data, used to describe the spatial geometric features of the label and product appearance components.

[0097] Step 2: Feature extraction of the industrial model based on two-stage fine-tuning: The industrial instructions and features are input into the target industrial model after two stages of fine-tuning. Based on internalized mechanical structure knowledge, electrical control logic, and industrial process rules, the target industrial model identifies the temporal dependency between "acquiring the physical location of corner points" and "attaching labels," and generates a process semantic execution graph that includes core sub-tasks such as "suction cup labeling," "camera positioning," and "robot point-to-point (PTP) movement," thereby achieving structured parsing of industrial instructions in natural language form.

[0098] Step 3: Inversion of physical action parameters of visual semantic features: For the task of "obtaining the physical location of corner points", a visual processing link is planned based on the target industrial model; The image processing operators in the intelligent vision platform, including modules such as "image source," "line search," and "line measurement," are invoked to extract the position data of the target object, namely the two-dimensional pixel coordinates of the intersection of the target boundary in the image. And depth data in point cloud data ; After acquiring the location data, the camera calibration parameter file is read from the visual operator configuration file. The camera intrinsic parameter matrix, hand-eye calibration matrix, and depth correction function are extracted from the visual physical inversion function and substituted into the visual physical inversion function.

[0099] in, The three-dimensional coordinates of the target object in the world coordinate system. For hand-eye calibration matrix, For depth correction function, This is the intrinsic parameter matrix of the image acquisition device. These are the two-dimensional pixel coordinates of the target object in the image data. The depth data of the target object in the point cloud data; Through the above inversion calculation, the visual pixel features are converted into a set of precise physical pose parameters in the world coordinate system. This parameter set includes the three-dimensional spatial coordinates of the target's grab position. and end effector attitude parameters The six degrees of freedom pose information is used to guide the robotic arm in absolute coordinate positioning.

[0100] Step 4: Constrained Decoding of Engineering Instructions Based on Formal Syntax Trees: In the process of generating the underlying control logic and outputting motion parameters, a standardized JavaScript Object Notation (JSON) format syntax rule tree for the motion control platform and the Smart Vision Platform (SVP) is introduced. The syntax rule tree describes the legal instruction structure that the industrial control platform is allowed to generate in a tree structure, including action nodes, parameter nodes, and control nodes. During the decoding phase, according to the syntax rule tree Generate the corresponding mask matrix ; The mask matrix dynamically applies to the output probability distribution of the target industrial model, filtering the generated action parameters so that the model can only generate legal fields allowed by the current syntax node. For example, when generating robotic arm motion commands, the mask matrix filters out illegal semantic outputs that do not conform to the industrial control interface specifications and generates platform-compatible action type key values ​​and their corresponding parameter fields, thereby ensuring that the generated action parameters conform to the syntax constraints of the industrial scenario and the industrial control system interface specifications.

[0101] Step 5: Standardized Engineering Document Generation and Equipment Implementation: The physical pose parameters obtained in step three are converted into compliant action parameters and injected into the industrial operator template to generate a structured engineering file that can be directly parsed by the industrial control system. The generated intelligent vision platform file encapsulates the connection topology of the vision operators to ensure that the vision processing flow conforms to syntax constraints and industry standards. At the same time, the motion control platform flow file is generated to transform the high-level task semantics into specific mechanical action sequences and programmable logic controller variable reading and writing, so as to achieve a precise combination of action parameters and industrial operators. The aforementioned structured engineering document was then sent to the edge controller of the production line. The industrial control platform called the document to drive each industrial device to perform physical actions such as belt start and stop, camera triggering, robotic arm transfer, and cylinder extension and retraction in sequence, thereby completing the label grabbing and attaching operation.

[0102] The industrial equipment motion parameter generation method provided in this application realizes the automated conversion from natural language industrial instructions to industrial equipment execution engineering files, solves the problem that general models cannot adapt to industrial scenarios, reduces the need for manual coordinate preset and low-level control program writing, improves the flexibility of industrial automation systems, and ensures the accuracy, compliance and executability of motion parameters, adapting to the high-precision and automated production needs of electronic manufacturing.

[0103] Figure 2 This is a second schematic flowchart of a method for generating motion parameters of industrial equipment provided in some embodiments of this application. For example... Figure 2 As shown, the method for generating motion parameters of industrial equipment also includes: Multimodal input: Using natural language commands ("grab the memory bar on the conveyor belt"), environmental aggregate information (depth camera data / point cloud data), and real-time device status (current robotic arm position, input and output signals) as inputs, multimodal data is generated by fusion. Core elements of inverse reasoning: Semantic feature extraction utilizes large models to identify target objects and predicts the location data of target objects from visual data; By combining environmental data, the position data is transformed into three-dimensional pose data in the world coordinate system through the visual physics inversion function; Motion trajectory planning is performed in conjunction with equipment constraints to generate motion parameters that meet the requirements; Project document generation stage: Using the physical parameter set as an intermediate format, the motion parameters are combined with the syntax rule tree and mask matrix through the operator interface adapter (corresponding to the industrial operator template) to generate two types of structured project files: an Extensible Markup Language (XML) format configuration file adapted to the motion control platform and a JavaScript Object Notation (JSON) format configuration file adapted to the intelligent vision platform.

[0104] Figure 3 This is a schematic diagram of the structure of an industrial equipment motion parameter generation system provided in some embodiments of this application. For example... Figure 3 As shown, the industrial equipment motion parameter generation system includes: a data base layer, a model building and fine-tuning layer, an inversion and generation layer, and an application execution layer. Among them: The data foundation layer includes three types of data: general corpus (mechanical principles, electrical logic, etc.), scenario data (integrated labeling and inspection, kitting management images, etc.), and equipment status data (workstation layout, equipment constraints).

[0105] The model building and fine-tuning layers are used to achieve two-stage training of the target industrial model, including: In the first phase of training, general industrial fine-tuning is performed. By injecting industrial mechanisms and fine-tuning instructions, a general model with basic industrial knowledge is produced. In the second stage of training scenario adaptation and fine-tuning, through reinforcement learning and preference optimization, a target industrial model adapted to specific industrial scenarios is produced.

[0106] The inversion and generation layer is used to complete the transformation from semantics to action parameters, including: By probing engineering and logical reasoning, natural language industrial instructions are parsed to generate process semantic execution diagrams; The visual-physical inversion module maps the position data into three-dimensional pose data. 3D pose data is converted into action parameters that can be executed by industrial equipment through parameter mapping. The action parameters are injected into the industrial operator template that conforms to the syntax rule tree through the operator interface matching; Generate via XML / JSON: Combine a mask matrix to generate structured project files in Extensible Markup Language (XML) or JavaScript Object Notation (JSON) format.

[0107] The application execution layer is used to execute industrial instructions, including: The motion control platform and intelligent vision platform call up structured engineering files and issue control commands. The equipment receives instructions from physical devices (industrial equipment), executes the industrial instructions according to the action parameters, and completes the production task.

[0108] The industrial equipment motion parameter generation method provided in this application can be executed by an industrial equipment motion parameter generation device. This application uses the example of an industrial equipment motion parameter generation device executing the method to illustrate the industrial equipment motion parameter generation device provided in this application.

[0109] Figure 4 This is a schematic diagram of the structure of an industrial equipment motion parameter generation device provided in some embodiments of this application. For example... Figure 4 As shown, the industrial equipment motion parameter generation device 400 includes: The acquisition module 401 is used to acquire industrial instructions described in natural language for industrial scenarios, and to obtain multimodal data corresponding to the industrial instructions based on the visual data of the industrial scenario and the status data of industrial equipment in the industrial scenario. The extraction module 402 is used to extract features from multimodal data through the target industrial model to obtain multimodal feature vectors; Prediction module 403 is used to predict the position data of the target object corresponding to the industrial command from visual data based on multimodal feature vectors; The generation module 404 is used to map position data into three-dimensional pose data of the target object through a visual-physical inversion function, and generate motion parameters corresponding to industrial commands based on the three-dimensional pose data, so that industrial equipment can execute industrial commands according to the motion parameters.

[0110] In some embodiments, the target industrial model is trained through the following steps: Based on industrial knowledge, the initial model is trained in the first stage through instruction fine-tuning to enhance its ability to understand industrial instructions and obtain a general industrial model. Based on multimodal data samples and corresponding action parameter samples, a second-stage training of the general industrial model is performed using the direct preference optimization algorithm to obtain a target industrial model adapted to the industrial scenario.

[0111] In some embodiments, the multimodal data samples include industrial instruction samples, visual data samples, and state data samples; Based on multimodal data samples and corresponding action parameter samples, a second stage of training is performed on the general industrial model using the direct preference optimization algorithm to obtain a target industrial model adapted to the industrial scenario, including: For each multimodal data sample, the preference reward value for each corresponding action parameter sample is determined based on the state data sample. The preference reward value is used to characterize the degree to which the action parameter sample conforms to the physical consistency constraints of the industrial scenario. Based on the preference reward value of each action parameter sample, select the preferred action parameter samples and the unpreferred action parameter samples from the action parameter samples; Industrial instruction samples and visual data samples are used as input data samples for a general industrial model, and combined with corresponding preferred action parameter samples and undesirable action parameter samples to construct preference training sample pairs. Based on the preference training sample pairs, the general industrial model is iteratively updated using the direct preference optimization algorithm to ensure that the motion parameter samples generated from the multimodal feature vectors output by the general industrial model conform to the physical consistency constraints of the industrial scenario, thus obtaining the target industrial model.

[0112] In some embodiments, the visual data includes image data and point cloud data; the location data includes two-dimensional pixel coordinates in the image data and depth data in the point cloud data; Module 404 is used for: Using the visual physics inversion function, the three-dimensional pose data of the target object in the world coordinate system is determined based on the two-dimensional pixel coordinates and depth data of the target object. Based on 3D pose data and combined with prior spatial information of the industrial scene, motion parameters corresponding to industrial commands are generated.

[0113] In some embodiments, the visual-physical inversion function includes an intrinsic parameter matrix of the image acquisition device for acquiring image data, a hand-eye calibration matrix for converting the coordinate system of the image acquisition device to the coordinate system of the industrial equipment, and a depth correction function for correcting depth data. Using the visual physics inversion function, the 3D pose data of the target object in the world coordinate system is determined based on the 2D pixel coordinates and depth data of the target object, including: Based on the intrinsic parameter matrix, the normalized coordinates of the target object in the coordinate system of the image acquisition device are determined according to the two-dimensional pixel coordinates. Based on the depth correction function, the depth correction coefficient of the target object is determined according to the depth data; Based on the normalized coordinates, depth data, and depth correction coefficient of the target object, determine the three-dimensional coordinates of the target object in the coordinate system of the image acquisition device; Based on the hand-eye calibration matrix, coordinate transformation is performed on the three-dimensional coordinates in the coordinate system of the image acquisition device to obtain the three-dimensional coordinates of the target object in the coordinate system of the industrial equipment. The three-dimensional coordinates of the target object in the coordinate system of the industrial equipment are used to determine the three-dimensional pose data of the target object in the world coordinate system.

[0114] In some embodiments, the industrial equipment motion parameter generation device 400 further includes a constraint module, which is used for: Construct a syntax rule tree corresponding to the industrial scenario. The syntax rule tree is used to represent the syntax constraints for generating action parameters corresponding to industrial instructions in a tree form. Generate a corresponding mask matrix based on the syntax rule tree; the mask matrix is ​​used to filter the action parameters during the generation process so that the action parameters conform to the syntax constraints.

[0115] In some embodiments, the industrial equipment motion parameter generation device 400 further includes a file generation module, which is used for: The action parameters are injected into an industrial operator template that conforms to the syntax rule tree, and a structured engineering file is generated by combining it with a mask matrix. This allows the industrial equipment to execute industrial instructions according to the action parameters by calling the structured engineering file.

[0116] The industrial equipment motion parameter generation device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific type of device.

[0117] The industrial equipment motion parameter generation device in this application embodiment can be a device with an operating system. This operating system can be a Microsoft (Windows) operating system, an Android operating system, an iOS operating system, or other possible operating systems; this application embodiment does not specifically limit it.

[0118] The industrial equipment motion parameter generation device provided in this application embodiment can realize all the processes implemented in the above-described industrial equipment motion parameter generation method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0119] Figure 5 These are schematic diagrams of the structure of an electronic device provided in some embodiments of this application. In some embodiments, such as Figure 5 As shown, this application embodiment also provides an electronic device 500, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the program is executed by the processor 501, it implements the various processes of the above-described industrial equipment action parameter generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0120] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.

[0121] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described industrial equipment action parameter generation method embodiment and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0122] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0123] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method for generating action parameters of industrial equipment.

[0124] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0125] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described industrial equipment action parameter generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0126] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0127] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0129] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0130] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0131] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for generating motion parameters of industrial equipment, characterized in that, include: Obtain industrial instructions described in natural language for industrial scenarios, and obtain multimodal data corresponding to the industrial instructions based on visual data of the industrial scenario and status data of industrial equipment in the industrial scenario. Feature extraction is performed on the multimodal data using the target industrial model to obtain a multimodal feature vector; Based on the multimodal feature vector, predict the position data of the target object corresponding to the industrial command from the visual data; The position data is mapped to the three-dimensional pose data of the target object using a visual physics inversion function, and motion parameters corresponding to the industrial command are generated based on the three-dimensional pose data, so that the industrial equipment can execute the industrial command according to the motion parameters.

2. The method according to claim 1, characterized in that, The target industrial model is obtained through the following steps: Based on industrial knowledge, the initial model is trained in the first stage through instruction fine-tuning to enhance the initial model's ability to understand the industrial instructions, thereby obtaining a general industrial model. Based on multimodal data samples and corresponding action parameter samples, the general industrial model is trained in the second stage using the direct preference optimization algorithm to obtain the target industrial model adapted to the industrial scenario.

3. The method according to claim 2, characterized in that, The multimodal data samples include industrial instruction samples, visual data samples, and state data samples; The second stage of training, based on multimodal data samples and corresponding action parameter samples, uses a direct preference optimization algorithm to train the general industrial model, resulting in a target industrial model adapted to the industrial scenario. This includes: For each of the multimodal data samples, a preference reward value is determined for each of the corresponding action parameter samples based on the state data samples. The preference reward value is used to characterize the degree to which the action parameter samples conform to the physical consistency constraints of the industrial scenario. Based on the preference reward value of each action parameter sample, select the preferred action parameter samples and the unpreferred action parameter samples from the action parameter samples; The industrial instruction samples and the visual data samples are used as input data samples for the general industrial model, and combined with the corresponding preferred action parameter samples and undesirable action parameter samples, preference training sample pairs are constructed. Based on the preference training sample pairs, the general industrial model is iteratively updated using the direct preference optimization algorithm so that the action parameter samples generated from the multimodal feature vectors output by the general industrial model conform to the physical consistency constraints of the industrial scenario, thereby obtaining the target industrial model.

4. The method according to claim 1, characterized in that, The visual data includes image data and point cloud data; the position data includes two-dimensional pixel coordinates in the image data and depth data in the point cloud data; The process of mapping the position data to the three-dimensional pose data of the target object using a visual-physical inversion function, and generating motion parameters corresponding to the industrial command based on the three-dimensional pose data, includes: Using the visual physics inversion function, the three-dimensional pose data of the target object in the world coordinate system is determined based on the two-dimensional pixel coordinates of the target object and the depth data. Based on the three-dimensional pose data and combined with the spatial prior information of the industrial scene, motion parameters corresponding to the industrial command are generated.

5. The method according to claim 4, characterized in that, The visual-physical inversion function includes an intrinsic parameter matrix of the image acquisition device used to acquire the image data, a hand-eye calibration matrix used to convert the coordinate system of the image acquisition device to the coordinate system of the industrial equipment, and a depth correction function used to correct the depth data. The step of determining the three-dimensional pose data of the target object in the world coordinate system using the visual physics inversion function, based on the two-dimensional pixel coordinates of the target object and the depth data, includes: Based on the intrinsic parameter matrix, the normalized coordinates of the target object in the coordinate system of the image acquisition device are determined according to the two-dimensional pixel coordinates. Based on the depth correction function, the depth correction coefficient of the target object is determined according to the depth data; The three-dimensional coordinates of the target object in the coordinate system of the image acquisition device are determined based on the normalized coordinates of the target object, the depth data, and the depth correction coefficient. Based on the hand-eye calibration matrix, coordinate transformation is performed on the three-dimensional coordinates in the coordinate system of the image acquisition device to obtain the three-dimensional coordinates of the target object in the coordinate system of the industrial equipment. The three-dimensional coordinates of the target object in the coordinate system of the industrial equipment are used to determine the three-dimensional pose data of the target object in the world coordinate system.

6. The method according to claim 1, characterized in that, The method further includes: Construct a syntax rule tree corresponding to the industrial scenario. The syntax rule tree is used to represent the syntax constraints for generating action parameters corresponding to the industrial instructions in a tree form. A corresponding mask matrix is ​​generated based on the syntax rule tree; wherein the mask matrix is ​​used to filter the action parameters during the generation of the action parameters so that the action parameters conform to the syntax constraints.

7. The method according to claim 6, characterized in that, The method further includes: The action parameters are injected into an industrial operator template that conforms to the syntax rule tree, and a structured engineering file is generated by combining the mask matrix. The structured engineering file is then called to drive each industrial device to execute the industrial instructions according to the action parameters.

8. A device for generating motion parameters for industrial equipment, characterized in that, The acquisition module is used to acquire industrial instructions described in natural language for industrial scenarios, and to obtain multimodal data corresponding to the industrial instructions based on the visual data of the industrial scenario and the status data of the industrial equipment in the industrial scenario. The extraction module is used to extract features from the multimodal data using the target industrial model to obtain a multimodal feature vector; The prediction module is used to predict the position data of the target object corresponding to the industrial command from the visual data based on the multimodal feature vector; The generation module is used to map the position data into the three-dimensional pose data of the target object through a visual physics inversion function, and generate motion parameters corresponding to the industrial command based on the three-dimensional pose data, so that the industrial equipment can execute the industrial command according to the motion parameters.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for generating industrial equipment motion parameters as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for generating industrial equipment motion parameters as described in any one of claims 1 to 7.