A multimodal large model training method and system for electronic equipment assembly scenarios

By collecting and processing text and image data of the assembly process, combining the Transformer model and manual guidance data set, the generalization performance of multimodal large models is optimized, solving the efficiency problem of pipeline redesign in electronic equipment assembly scenarios, and achieving efficient assembly tasks.

CN119089978BActive Publication Date: 2025-05-16TONGJI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411204561.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-05-16
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

The prior art requires a lot of time and effort to redesign and configure the assembly line in electronic equipment assembly scenarios, making it difficult to efficiently realize the generalization performance of multimodal large models.

Method used

Collect text and image data of the assembly process, build process guidance data sets and associated information data sets, train multimodal large models through Transformer pre-training models, and perform incremental learning and interactive teaching through manual guidance data sets to optimize model performance.

Benefits of technology

It improves the generalization performance of multimodal large models, saves assembly time, and improves the efficiency and accuracy of assembly tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119089978B_ABST
    Figure CN119089978B_ABST
Patent Text Reader

Abstract

The present invention relates to a multimodal large model training method and system for electronic equipment assembly scenes, the method comprising the following steps: collecting data required for the electronic equipment assembly process, constructing a process guidance data set; obtaining physical images in the actual assembly process, constructing a related information data set; inputting the process guidance data set and the related information data set into a pre-trained model based on Transformer for training, and preliminarily obtaining a multimodal large model; obtaining action information of manually performing assembly tasks in the same task, and constructing a multimodal data set for manual guidance; inputting into the multimodal large model, fine-tuning the large model, updating the assembly details to improve the model performance, and obtaining a multimodal large model for electronic equipment assembly scenes; for parts or assembly details that have not been learned, the generalization of the model is improved through interactive learning of physical teaching. Compared with the prior art, the present invention improves the generalization performance of the multimodal large model, saves assembly time, and improves the efficiency of assembly tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model training, and in particular to a multimodal large model training method and system for electronic equipment assembly scenarios. Background Art

[0002] Large language models have shown significant advantages in various fields with their powerful learning ability and wide applicability. Large models can not only extract effective information from natural language, but also capture complex data patterns and perform well in multiple tasks. Therefore, when it comes to multimodal large language models, their advantages are more prominent. Large models with multimodal capabilities can process multiple types of inputs such as images and sounds while processing text data, achieving cross-modal understanding and generation, enabling the model to more comprehensively understand and respond to complex situations in the real world, further expanding its scope of application and practicality.

[0003] As a representative assembly scenario, electronic equipment assembly has a standardized assembly process, in which each link involves assembly robots, and parts are assembled according to fixed standards. Since it involves the assembly process of electronic equipment, it has high requirements for refinement. At present, there are many assembly line operations with preset fixed postures. For example, Chinese patent CN 111132535B discloses a circuit board installation module, which involves the field of product processing technology. The circuit board installation module includes a placer, a storage, a detector, a sensor, a buffer and a transporter. The placer is used to place two products, the storage is used to place the circuit board, the detector is suitable for detecting whether the circuit board and the product are qualified, the sensor is suitable for monitoring whether the buffer stores the circuit board, and the transporter is suitable for installing the circuit board to the product. However, for new links and processes, it often takes a lot of time and effort to redesign and configure the assembly line. Summary of the invention

[0004] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide a multimodal large model training method and system for electronic equipment assembly scenarios, which improves the generalization performance of the multimodal large model, saves assembly time, and greatly improves the efficiency of assembly tasks.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] A multimodal large model training method for electronic equipment assembly scenarios includes the following steps:

[0007] Collect the fixed professional assembly text and the corresponding assembly model data required for the electronic equipment assembly process, and build the process guidance dataset I data ;

[0008] Obtain the physical image of the operating table during the actual assembly process of the assembly robot, record the state data of the assembly robot, extract the association information between multiple parts, and construct the association information dataset C data ;

[0009] The process guides the data set I data and the associated information dataset C data Perform preprocessing and input into the Transformer-based pre-trained model for training, and initially obtain a multi-modal large model suitable for assembly scenarios;

[0010] Obtain visual data, hand and arm motion details data, and tactile dexterity data from manually performing assembly tasks in the same task, and construct a multimodal dataset H of manual guidance;

[0011] The manually guided multimodal dataset H is input into the multimodal large model, and the large model is fine-tuned through an incremental learning method, and the assembly details are updated to improve the model performance, thereby obtaining a multimodal large model for electronic equipment assembly scenarios.

[0012] Furthermore, for the parts or assembly details that have not been learned, the multimodal large model of the electronic device assembly scene performs interactive learning through teaching.

[0013] Furthermore, the assembly process includes various links in mainboard assembly, screen assembly, battery assembly, shell assembly and overall assembly; the assembly model data includes assembly schematics, circuit structure diagrams, parts appearance diagrams, parts installation diagrams, and various picture forms such as top views, side views, enlarged views and omitted views of parts and installation areas.

[0014] Furthermore, the process guides the data set I data It is a collection of pictures and texts describing the fixed process, and its specific form is as follows:

[0015] I data =[s i ,t i ,(p1,p2,...,p m ),l i ],(i=1,2,...,j)

[0016] In the formula, j is the number of assembly steps, i is the step number, and s i is the index of the current specific assembly step, t i is the text form of the assembly method description corresponding to this step, p1, p2, ..., p m are the multiple assembly models corresponding to the text assembly method description, l i It is a remark-style language prompt artificially added in this assembly step.

[0017] Furthermore, the association information includes a cross-domain association C1 and a local domain association C2; the cross-domain association C1 is a process guidance data set I data The connection with the physical image information collected in this round, the domain association C2 is the connection between the physical image information collected in this round at different time sequences.

[0018] Furthermore, the preprocessing includes unimodal preprocessing and multimodal data alignment and fusion. The unimodal preprocessing includes word segmentation, stop word removal and stem extraction of text data, and scaling, cropping and normalization of image data; the multimodal data alignment and fusion steps are to convert the data of each modality into high-level semantic features through a specific feature extraction method; and use alignment technology to map the high-level semantic features into a shared feature space to achieve alignment between modalities.

[0019] Furthermore, the steps of constructing the manually guided multimodal dataset are as follows:

[0020] Get every action information of workers during the assembly process

[0021] The plurality of action information Combination, constructed under the corresponding assembly steps s i A human-guided multimodal dataset in

[0022] The human-guided multimodal dataset combines multiple assembly steps Then a manually guided multimodal dataset H is constructed.

[0023] Furthermore, the action information Including visual data Manual dexterity data Key point tactile information and arm movement information Right now

[0024] Furthermore, the specific form of the incremental learning is: loading a pre-trained multimodal large model, initializing model parameters, setting a learning rate, and processing the manually guided multimodal data set H in batches and then putting it into the model for fine-tuning. The objective function of the incremental learning is as follows:

[0025]

[0026] In the formula, θ is the model parameter, D t is the current batch data, D r is the playback data, L(·) represents the loss function, Corresponding assembly steps i Human-guided multimodal data, f(s i ; θ) is the model output, and λ is the weight factor of the playback data.

[0027] According to another aspect of the present invention, a multimodal large model training system for electronic equipment assembly scenarios is provided, comprising:

[0028] The process guidance data collection module is used to collect the necessary assembly manual information during the electronic equipment assembly process, ensure the correspondence between the text and image information of each assembly step, and construct a process guidance data set;

[0029] The associated information building module is used to obtain the physical images of the electronic devices assembled by the assembly robot during the actual assembly process of the electronic devices, and to find the association between the acquired images and the external data set, as well as the connection with the same type of data in time series, so as to build the associated information data set;

[0030] The manual guidance data acquisition module is used to acquire data of workers wearing eye trackers, tactile force feedback data gloves, and motion capture points under specified assembly tasks and performing assembly tasks in a motion capture environment, thereby constructing a manual guidance multimodal dataset;

[0031] An incremental learning fine-tuning module is used to establish a specific scheme of the incremental learning method through the collected artificial guidance multimodal data set, and propose a fine-tuning method suitable for obtaining a multimodal large model for electronic equipment assembly scenarios on the basis of ensuring the preliminary model;

[0032] A multimodal large model training module is used to put the data into the Transformer-based pre-training model for training in different time periods, and after obtaining a multimodal large model that is initially suitable for the assembly scene, fine-tune it through the incremental learning method to obtain a multimodal large model for the electronic equipment assembly scene;

[0033] The physical interaction correction module is used to achieve interactive learning of multimodal large models of electronic equipment assembly scenarios and correct motion trajectories through physical teaching for parts or assembly details that have not been learned.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] 1. The present invention collects text and assembly model data in the assembly manual to construct a process guidance data set, and obtains physical images to obtain a multi-part and multi-dimensional associated information data set, thereby improving the availability of data, highlighting the characteristics of assembly movements, and effectively improving the degree of specialization of training data while reducing learning costs.

[0036] 2. The present invention collects visual data, hand and arm movement detail data, and tactile dexterity data in manually performed assembly tasks, constructs a detailed and precise manually guided multimodal data set, forms a professional assembly process standard, and puts it into a large multimodal model. Through incremental learning, there is no need to retrain the entire model, which greatly improves the execution efficiency and accuracy of multimodality in specific scenarios, and provides a new solution for assembly work in electronic equipment scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A flowchart of a multimodal large model training method, system, device and medium for electronic device assembly scenarios according to an embodiment of the present invention;

[0038] Figure 2 This is a structural diagram of an associated information data set of a multimodal large model training method, system, device and medium for an electronic device assembly scenario according to an embodiment of the present invention;

[0039] Figure 3 This is a schematic diagram of the structure of a multi-modal large model training system for electronic equipment assembly scenarios according to an embodiment of the present invention;

[0040] Figure 4 A schematic diagram of a multimodal large model interactive learning for electronic device assembly scenarios according to an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0042] Example 1

[0043] This embodiment provides a multi-modal large model training method for electronic equipment assembly scenarios. Figure 1 As shown, the following steps are included:

[0044] S1. Collect the fixed professional assembly text and the corresponding assembly model data required for the electronic equipment assembly process, and build the process guidance data set I data .

[0045] Obtain the necessary assembly manual information during the assembly of electronic equipment. The assembly process is mainly divided into main links such as motherboard assembly, screen assembly, battery assembly, shell assembly, and overall assembly. It is assumed that there are j assembly steps in each link, i is the step number, s is the step number i Represents the current specific assembly step index. The assembly manual information includes each link in the assembly process and the steps in each link. iThe fixed text form of the process includes the language description of the assembly sequence in the assembly steps, the explanation of the subject and object of the assembly, the basic connection principle between the parts and the matters that need to be paid attention to in this link, etc., recorded as t i At the same time, in a text content t i The description will correspond to a single or multiple assembly models, which may be in the form of assembly schematics, circuit structure diagrams, parts appearance diagrams, parts installation diagrams, and multi-part and multi-angle pictures of parts and installation areas, such as top views, side views, enlarged views, and omitted views, so they are recorded as p1, p2, ..., p m If you need to add special instructions or installation tips for some steps, you need to add comment language prompts. i , thereby constructing the guidance dataset I data , expressed as I data =[s i ,t i ,(p1,p2,...,p m ),l i ],(i=1,2,...,j).

[0046] S2. Obtain the physical image of the operating table during the actual assembly process of the assembly robot, record the state data of the assembly robot, extract the association information between multiple parts, and construct the association information dataset C data .

[0047] Assembly robots are generally composed of multi-degree-of-freedom industrial robots and dexterous hands. The multi-degree-of-freedom industrial robots move over a large range in the workspace, and the dexterous hands perform precision assembly tasks between parts. Physical images of electronic equipment include electronic motherboards, electronic motherboards with installed electronic components, etc. Physical images of electronic equipment contain more available information and can construct multi-dimensional, multi-part related information. Figure 2 As shown, the association information is divided into cross-domain association and local domain association. Specifically, the guidance dataset I data The part text information and part schematic diagram in the assembly model are compared with the part information collected from the current round of assembly model, and the physical picture and schematic diagram are compared and the text information relationship is mapped, so as to realize cross-domain association, which is denoted as C1; the local domain association refers to the connection between the physical image information collected in this round in the same data domain, which mainly includes the physical position relationship between parts in the physical picture, such as up and down, embedded, and covered, the overall position relationship between the local part and the assembly as a whole, and the dynamic change of the regional state at different times t, including various information such as the assembly installation direction and mode, so as to construct the local domain association, which is denoted as C2. The above cross-domain association and local domain association jointly construct the association information dataset C data .

[0048] S3. Process guidance dataset I data and the associated information dataset C data After preprocessing, the model is input into the Transformer-based pre-trained model for training, and a large multimodal model suitable for assembly scenarios is initially obtained.

[0049] The collected guidance dataset I data and the associated information dataset C data After preprocessing, the training data for targeted vertical scenes is used. The preprocessing is divided into single-modal preprocessing and multi-modal data alignment and fusion. For each modal data, a separate preprocessing step is performed according to its characteristics. For example, for text data, word segmentation, stop word removal, stem extraction and other processing are performed. For image data, scaling, cropping, normalization and other operations are performed to make the data of each modality in a format and state suitable for subsequent processing and analysis. At the same time, after these steps, feature extraction and alignment techniques are performed to convert the data of each modality into high-level semantic features through specific feature extraction methods, and alignment techniques such as registration and alignment transformation are used to map these features to a shared feature space, thereby achieving alignment between modalities. The preprocessed data is put into a Transformer-based pre-training model for pre-training, and several supervisory information in the form of <input, output> is added for preliminary fine-tuning to initially obtain a multi-modal large model m suitable for assembly scenarios. φ (φ is a parameter).

[0050] S4. Obtain visual data, hand and arm motion detail data, and tactile dexterity data from manually performed assembly tasks in the same task, and construct a manually guided multimodal dataset H.

[0051] Under the same assembly task, workers will wear eye trackers, tactile force feedback data gloves and motion capture points, and perform the same assembly task in a motion capture environment surrounded by multiple motion capture devices, so as to construct multimodal data into a manually guided multimodal data set. The eye tracker can track and measure the position of the eyeball and specific videos, and record the actual position information of the visual attention point in the scene and convert it into relevant parameters; the tactile force feedback data gloves are WiseGlove, which can capture the movement information of the human hand in real time and with high precision, including the bending, stretching, palm-to-palm, rubbing and other movements of the fingers, and output the dexterous data of the task operation, so as to control the hand movements of the dexterous hand and complete complex and delicate tasks; the use of motion capture devices and motion capture points is to collect specific movements in detail. The motion capture device records and restores the movement posture of the object by capturing the position information of the motion capture points attached to the key parts of the human body, such as reflective markers or LED markers. Among them, the motion capture points are the key elements in the motion capture technology as the target detection points, which are tracked and measured in real time by the motion capture device, so as to achieve accurate capture and recording of the movement of the object.

[0052] The data collected in the above process together constitute the artificial guidance multimodal dataset H, where specifically, each action information Consists of multimodal data, including visual data Manual dexterity data Key point tactile information and arm movement information Right now The information of multiple actions constructs the corresponding assembly step s i A human-guided multimodal dataset in Therefore, a multimodal dataset with human guidance combining multiple assembly steps Then a manually guided multimodal dataset H is constructed.

[0053] S5. Input the manually guided multimodal dataset H into the multimodal large model, fine-tune the large model through an incremental learning method, update the assembly details to improve the model performance, and obtain a multimodal large model for electronic equipment assembly scenarios.

[0054] Manual guidance of multimodal datasets into the multimodal large model m φ In the process, the incremental learning method is used to keep the initial performance of the model, and update the assembly details through new specialized data to improve the model performance. The incremental learning method is a continuous learning paradigm that allows multimodal large models to φ Based on existing knowledge, new tasks or new knowledge can be gradually learned without retraining the entire model.

[0055] The specific form of incremental learning is to load the pre-trained multimodal large model, initialize the model parameters, set the learning rate, and process the manually guided multimodal data set H in batches and then put it into the model for fine-tuning. This incremental learning freezes some layers such as the underlying feature extraction layer, and only fine-tunes the top layer or the layer related to a specific task. At the same time, the idea of ​​"replaying" data is introduced. Before each incremental learning, a part of the "replay" data is randomly extracted from the old data and mixed with the current batch data for training, so as to avoid the risk of catastrophic forgetting. Therefore, the objective function of the incremental learning method is as follows:

[0056]

[0057] In the formula, θ is the model parameter, D t is the current batch data, D r is the playback data, L(·) represents the loss function, Corresponding assembly steps i Human-guided multimodal data, f(s i ; θ) is the model output, λ is the weight factor of the playback data to balance the influence of new and old data, thereby updating the assembly detail information to improve the model performance, and thus obtaining a multimodal large model for electronic equipment assembly scenarios.

[0058] In another preferred embodiment, the steps are also included:

[0059] S6. For parts or assembly details that have not been learned, interactive learning in which professionals drag the assembly robot arm for a limited number of times to teach can effectively improve the generalization of the multimodal large model for electronic equipment assembly scenarios.

[0060] like Figure 4 As shown, in this embodiment, when encountering unlearned parts or assembly details during the assembly process, the assembly robot executes a deviation trajectory in the face of unfamiliar parts and unfamiliar assembly environments. The professional can drag the assembly robot arm a limited number of times for physical teaching to generate the target trajectory. object , to achieve interactive learning, so that the deviation trajectory ξ w Become the corrected trajectory ξ c The corrected trajectory does not necessarily completely overlap the target trajectory. Instead, the reward function in the correction process is defined to satisfy the positive reward at each time step as much as possible, so as to obtain a corrected trajectory that achieves the overall motion task while ensuring accuracy and smoothness. The corrected relevant data information is then transmitted back to the multimodal large model for fine-tuning and generalization improvement, so that the multimodal large model for electronic assembly scenarios can achieve the generalization of more demanding assembly tasks.

[0061] Define the robot's motion state during the motion process as x, and the motion trajectory ξ consists of a series of robot motion states x, ξ = x 0:t , and introduce the robot's environment E, including the changes between different states and the relationship with the assembly equipment, to better describe the robot's motion space. At the same time, the idea of ​​reward mechanism is introduced, and the reward function is set as R, where the reward function R should be a relationship between the trajectory and the environment in the form of φ(ξ,E)∈[0,1] k Linear combination of features, and introduce a parameter vector θ∈R k , then the reward function is expressed as follows within a certain time t:

[0062]

[0063] Where θ is the parameter vector, ξ is the motion trajectory, and E is the robot's environment.

[0064] The reward function R needs to ensure that the target trajectory ξ at each time step is taught by the expert drag object The reward value needs to be higher than the deviation trajectory ξ w Departure trajectory ξ that is being corrected i ,satisfy Thus, the corrected trajectory ξ is obtained c .

[0065] Example 2

[0066] This embodiment provides a multi-modal large model training system for electronic equipment assembly scenarios. Figure 3 As shown, including:

[0067] The process guidance data collection module is used to collect the necessary assembly manual information during the electronic equipment assembly process, ensure the correspondence between the text and image information of each assembly step, and construct a process guidance data set;

[0068] The associated information building module is used to obtain the physical images of the electronic devices assembled by the assembly robot during the actual assembly process of the electronic devices, and to find the association between the acquired images and the external data set, as well as the connection with the same type of data in time series, so as to build the associated information data set;

[0069] The manual guidance data acquisition module is used to acquire data of workers wearing eye trackers, tactile force feedback data gloves, and motion capture points under specified assembly tasks and performing assembly tasks in a motion capture environment, thereby constructing a manual guidance multimodal dataset;

[0070] An incremental learning fine-tuning module is used to establish a specific scheme of the incremental learning method through the collected artificial guidance multimodal data set, and propose a fine-tuning method suitable for obtaining a multimodal large model for electronic equipment assembly scenarios on the basis of ensuring the preliminary model;

[0071] A multimodal large model training module is used to put the data into the Transformer-based pre-training model for training in different time periods, and after obtaining a multimodal large model that is initially suitable for the assembly scene, fine-tune it through the incremental learning method to obtain a multimodal large model for the electronic equipment assembly scene;

[0072] The physical interaction correction module is used to achieve interactive learning of multimodal large models of electronic equipment assembly scenarios and correct motion trajectories through physical teaching for parts or assembly details that have not been learned.

[0073] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A multimodal large model training method for electronic equipment assembly scenarios, characterized in that: The following steps are involved: Collect the fixed professional assembly text and the corresponding assembly model data required for the electronic equipment assembly process, and build the process guidance dataset I data ; Obtain the physical image of the operating table during the actual assembly process of the assembly robot, record the state data of the assembly robot, extract the association information between multiple parts, and construct the association information dataset C data The association information includes cross-domain association C1 and local domain association C2; cross-domain association C1 is the process guidance data set I data The connection with the physical image information collected in this round. The domain association C2 is the connection between the physical image information collected in this round at different time sequences; The process guides the data set I data and the associated information dataset C data Perform preprocessing and input into the Transformer-based pre-trained model for training, and initially obtain a multi-modal large model suitable for assembly scenarios; Obtain visual data, hand and arm motion detail data, and tactile dexterity data in the same task of manually performing an assembly task, and construct a manually guided multimodal dataset H. The steps for constructing the manually guided multimodal dataset H are as follows: Get information about every action of workers during the assembly process The action information Including visual data Manual dexterity data Key point tactile information and arm movement information Right now The plurality of action information Combination, constructed under the corresponding assembly steps s i A human-guided multimodal dataset in The human-guided multimodal dataset combines multiple assembly steps Then, a manually guided multimodal dataset H is constructed; The manually guided multimodal dataset H is input into the multimodal large model, and the large model is fine-tuned through an incremental learning method, and the assembly details are updated to improve the model performance, thereby obtaining a multimodal large model for electronic equipment assembly scenarios.

2. The multimodal large model training method for electronic device assembly scenarios according to claim 1 is characterized in that: For parts or assembly details that have not been learned, the multimodal large model of the electronic device assembly scene performs interactive learning through teaching.

3. The multimodal large model training method for electronic device assembly scenarios according to claim 1 is characterized in that: The assembly process includes multiple links in mainboard assembly, screen assembly, battery assembly, shell assembly and overall assembly; The assembly model data includes assembly schematics, circuit structure diagrams, parts outline diagrams, parts installation diagrams, and various picture forms such as top views, side views, enlarged views, and omitted views of parts and installation areas.

4. The multimodal large model training method for electronic device assembly scenarios according to claim 1 is characterized in that: The process guides the data set I data It is a collection of pictures and texts describing the fixed process, and its specific form is as follows: I data =[s i ,t i ,(p1,p2,...,p n ),l i ],(i=1,2,...,j) In the formula, j is the number of assembly steps, i is the step number, and s i is the index of the current specific assembly step, t i is the text form of the assembly method description corresponding to this step, p1, p2, ..., p n are multiple assembly models corresponding to the assembly method description in text form, l i It is a remark-style language prompt artificially added in this assembly step.

5. The multimodal large model training method for electronic device assembly scenarios according to claim 1 is characterized in that: The preprocessing includes single-modal preprocessing and multi-modal data alignment and fusion. The single-modal preprocessing includes word segmentation, stop word removal and stem extraction of text data, and scaling, cropping and normalization of image data. The multi-modal data alignment and fusion step is to convert the data of each modality into high-level semantic features through a specific feature extraction method. Alignment technology is used to map the high-level semantic features into a shared feature space to achieve alignment between modalities.

6. The multimodal large model training method for electronic device assembly scenarios according to claim 1 is characterized in that: The specific form of the incremental learning is: loading a pre-trained multimodal large model, initializing model parameters, setting the learning rate, and processing the manually guided multimodal data set H in batches and then putting it into the model for fine-tuning. The objective function of the incremental learning is as follows: In the formula, θ is the model parameter, D t is the current batch data, D r is the playback data, L(·) represents the loss function, Corresponding assembly steps i Human-guided multimodal data, f(s i ; θ) is the model output, and λ is the weight factor of the playback data.

7. A multimodal large model training system for electronic equipment assembly scenarios, characterized in that: include: The process guidance data collection module is used to collect the necessary assembly manual information during the electronic equipment assembly process, ensure the correspondence between the text and image information of each assembly step, and construct a process guidance data set; The association information construction module is used to obtain the physical image of the electronic device assembled by the assembly robot in the actual assembly process of the electronic device, and find the association between the collected image and the external data set, as well as the connection with the same type of data in time series, so as to construct the association information data set. The association information includes cross-domain association C1 and local domain association C2; the cross-domain association C1 is the process guidance data set I data The connection with the physical image information collected in this round. The domain association C2 is the connection between the physical image information collected in this round at different time sequences; The manual guidance data acquisition module is used to acquire data of workers wearing eye trackers, tactile force feedback data gloves and motion capture points under a specified assembly task and performing assembly tasks in a motion capture environment, thereby constructing a manual guidance multimodal dataset. The steps for constructing the manual guidance multimodal dataset are as follows: Get information about every action of workers during the assembly process The action information Including visual data Manual dexterity data Key point tactile information and arm movement information Right now The plurality of action information Combination, constructed under the corresponding assembly steps s i A human-guided multimodal dataset in The human-guided multimodal dataset combines multiple assembly steps Then a manually guided multimodal dataset was constructed; An incremental learning fine-tuning module is used to establish a specific scheme of the incremental learning method through the collected artificial guidance multimodal data set, and propose a fine-tuning method suitable for obtaining a multimodal large model for electronic equipment assembly scenarios on the basis of ensuring the preliminary model; A multimodal large model training module is used to put the process guidance data set and the associated information data set into a Transformer-based pre-training model for training, and after obtaining a multimodal large model that is initially suitable for the assembly scene, fine-tune it through the incremental learning method to obtain a multimodal large model for the electronic device assembly scene; The physical interaction correction module is used to achieve interactive learning of multimodal large models of electronic equipment assembly scenarios and correct motion trajectories through physical teaching for parts or assembly details that have not been learned.

Citation Information

Patent Citations

  • Circuit board mounting modules, electronic device assembly systems and methods

    CN111132535B

  • Intelligent flexible assembly method for 3C products

    CN118013838A

  • Autonomous construction method and system for intelligent bulldozing robot with body

    CN118238153A