Industrial imaging pose adjusting system based on large language model

Through the industrial imaging posture adjustment system based on the large language model, using the robot and model training module, efficient and stable imaging posture adjustment is achieved, which solves the problems of reliance on manual experience and insufficient evaluation functions in the existing technology, and improves the imaging quality and adjustment efficiency.

CN120635191AActive Publication Date: 2025-09-12BEIJING FOCUSIGHT TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511120361.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-12
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing industrial imaging posture adjustment methods rely on manual experience or evaluation functions, resulting in a cumbersome, inefficient and unstable adjustment process, making it difficult to achieve consistency and optimization in different scenarios.

Method used

An industrial imaging posture adjustment system based on a large language model is adopted. The camera and light source are grasped by a robot arm. Combined with the data acquisition module and the model training module, the large language model and the text classification model are used to extract user input information, train the imaging level and output the reference posture, and combine manual fine-tuning to achieve the optimal effect.

Benefits of technology

It reduces the workload of optoelectronic engineers, lowers the experience and ability requirements, ensures the stability and consistency of imaging effects, and improves adjustment efficiency and imaging quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635191A_ABST
    Figure CN120635191A_ABST
Patent Text Reader

Abstract

The invention relates to an industrial imaging pose adjusting system based on a large language model. The industrial imaging pose adjusting system comprises a fixing device, a manipulator device, a data acquisition module, a model training module and an application module. Efficient and stable adjustment of the industrial imaging pose is realized by building a device, planning an imaging path, collecting imaging data, training a large model and finally utilizing the model to output a reference pose in combination with manual fine adjustment. According to the method, the large imaging pose model is obtained through training, and the reference pose is given by inputting the detected object and defect information, so that the time consumption and instability of manual adjustment are greatly reduced; by obtaining the reference pose, the time consumption of manual adjustment is reduced, the time cost is saved, the requirements for operators are reduced, the labor cost is saved, the stability of the effect is improved, and the imaging quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual imaging technology, and in particular to an industrial imaging posture adjustment system based on a large language model. Background Art

[0002] In the industrial imaging process, adjusting the position and angle (i.e., posture) of the camera and light source relative to the object being measured is crucial and directly affects the image quality. Currently, there are two main methods for posture adjustment: 1. Manual adjustment Photoelectric engineers adjust the position and angle (i.e., posture, hereafter referred to as posture) of the light source and camera relative to the object being measured based on their experience; The steps are: (1) Select an initial pose based on experience; (2) Manually adjust the position of the light source and camera, observe the imaging effect according to the human eye, observe the feedback, and adjust the position of the light source and camera based on the feedback and experience; (3) Continuously adjust until a posture that satisfies the optoelectronic engineer is obtained; 2. Automatic adjustment Currently, there are some intelligent imaging devices that use a robot to drive the light source and camera to adjust the posture. During the adjustment process, the evaluation function is used to measure the quality of the imaging effect. The adjustment is performed by traversing or planning the robot path, and the posture with the highest evaluation function score is selected as the final posture.

[0003] However, both of the above methods have defects: 1. Manual adjustment requires experienced engineers to make adjustments based on their own experience, which places high demands on the operator's ability and increases labor costs. The adjustment process is cumbersome and manual, which is inefficient. Different optoelectronic engineers have different experiences and abilities, so the adjustment results will vary greatly, resulting in unstable effects. 2. Automatic adjustment relies entirely on the evaluation function to measure the quality of imaging effects. The score of the imaging function is difficult to fully match the quality of the actual imaging effect, resulting in poor results. Moreover, the degree of correlation between the score of the imaging function and the actual effect varies greatly in different scenarios, making it difficult to achieve universal application. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an industrial imaging posture adjustment system based on a large language model, which solves the problem that a large number of repeated manual adjustments are required during the posture adjustment stage of industrial imaging, which is cumbersome and has unstable effects.

[0005] The technical solution adopted by the present invention to solve the technical problem is: an industrial imaging posture adjustment system based on a large language model, comprising: A fixing device for fixing the object to be measured; The robot device includes two manipulators, which respectively grasp the camera and the light source, can record the current posture state, and can adjust to the specified posture according to the posture adjustment instruction; The data acquisition module is used to traverse the postures according to the set imaging path, perform imaging at each posture, record the current posture state, process the imaging data to obtain pixel-level annotations and scoring indicators, and determine the imaging level; The model training module includes a large language model and a text classification model. The large language model is used to extract information from user input, and the text classification model is trained with a description text containing the extracted information and pose parameters as input and imaging levels as labels. The application module is used to receive the measured object and defect information input by the user and output the reference pose through the trained model.

[0006] Furthermore, the imaging modes of the present invention include a reflection mode and a transmission mode. In the reflection mode, the camera and the light source are located on the same side of the object to be measured, and in the transmission mode, the camera and the light source are located on both sides of the object to be measured.

[0007] Furthermore, the posture parameters described in the present invention include the vertical distance from the optical center of the camera lens to the object to be measured, the angle between the optical axis of the camera lens and the vertical direction, the vertical distance from the center of the light source to the object to be measured, the angle between the optical axis of the light source and the vertical direction, and the horizontal distance from the center of the camera lens to the center of the light source.

[0008] Furthermore, the scoring indicators described in the present invention include contrast, integrity and discrimination. Contrast is the average grayscale difference between the marked area and the surrounding area, integrity is the intersection-over-union ratio of the connected domain and the annotation after binarization, and discrimination is the difference ratio between the average grayscale of the relevant connected domain and the average grayscale of the marked area.

[0009] Furthermore, the imaging grades of the present invention are divided into A, B, C and D, where A corresponds to contrast > 50, B corresponds to 30≤contrast≤50, C corresponds to 15≤contrast≤30, and D corresponds to contrast <15.

[0010] Furthermore, the large language model of the present invention is a deepseek14b model, and the text classification model is based on a Bert model.

[0011] Furthermore, the present invention includes an imaging posture adjustment method, which includes the following steps: S1. Build the device: Fix the object to be measured, use two manipulators to grab the camera and light source respectively, and determine the initial pose according to the imaging mode; S2. Plan the imaging path: set the starting value, ending value and traversal step length of the posture parameters; S3. Collect imaging data: traverse the posture along the path, record the posture status, process the imaging data to obtain pixel-level annotation and scoring indicators, and determine the imaging level; S4. Training model: Use the large language model to extract information, construct a description text containing information and posture parameters, and train the text classification model; S5. Application model: Extract user demand information, plan the path and construct description text, predict the imaging level through the model, select the optimal posture and output the reference posture through the large language model, and manually fine-tune to the best effect.

[0012] Furthermore, the starting value and ending value of the posture parameters described in the present invention are determined based on the values ​​under the optimal posture, and the step size is set according to the actual scenario.

[0013] Furthermore, when processing imaging data according to the present invention, pixel-level annotations at the current posture are obtained through affine transformation.

[0014] Furthermore, when the present invention applies the model, three poses corresponding to the optimal level are randomly sampled as reference poses.

[0015] The beneficial effect of the present invention is to solve the defects existing in the background technology. 1. By using a large amount of manual adjustment data and designing prompt words as samples, the large model is fine-tuned, allowing the large model to learn the changing rules between posture adjustment and imaging effects in various scenarios. This enables the large model to design imaging postures according to specific input scenarios. The large model will give the corresponding posture through input, and the output posture is the optimal posture, which can be used as the initial posture. Then, manual adjustments can be made to this initial posture. This reduces the workload of optoelectronic engineers, lowers the experience and ability requirements of optoelectronic engineers, saves time and labor costs, and different engineers can make small adjustments under the same reference posture. The final result is also stable, making the imaging effect stable and easy to achieve the optimal effect. 2. Using the trained AI model to predict the imaging effect is more powerful than the traditional effect evaluation function fitting function and is closer to human judgment. It can reduce the workload of optoelectronic engineers by manually making small adjustments. During the adjustment process, it is manually judged. Compared with the evaluation function, the effect judgment is more reasonable and more universal, ensuring the final imaging effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is the overall flow chart of the training and application of the large-scale model for industrial imaging posture adjustment of the present invention; Figure 2This is a side view posture relationship diagram of the light source, camera, and object under reflection in the present invention; Figure 3 It is a diagram showing the relationship between the light source, the camera, and the object under test in one of the front-view poses under the reflection condition of the present invention; Figure 4 This is a side view posture relationship diagram of the transmission mode light source, camera, and object under test of the present invention; Figure 5 It is a schematic diagram of the model architecture of the present invention. DETAILED DESCRIPTION

[0017] The present invention will now be described in further detail with reference to the accompanying drawings and preferred embodiments. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0018] like Figure 1-Figure 5 The overall implementation steps of the industrial imaging posture adjustment system based on the large language model are as follows: 1. Device Construction Fix the object to be measured, and two manipulators grab the camera and light source respectively to adjust the posture. The manipulator can record the current posture state and adjust to the specified posture according to the posture adjustment instruction. The construction method of the device is divided into reflection mode and transmission mode according to the different imaging modes, as shown below: Figure 2 and Figure 4 As shown in the figure, they are the side view pose diagrams of the reflection mode and the transmission mode, respectively. The plane where the object is located is the plane reference system, where: hc1 is the vertical distance from the optical center of the camera lens to the object being measured; αc1 is the angle between the optical axis of the camera lens and the vertical direction; hs1 is the vertical distance from the center of the light source to the object being measured; αs1 is the angle between the optical axis of the light source and the vertical direction; d1 is the horizontal distance from the center of the camera lens to the center of the light source; The same is true for the front view posture diagram. When looking at the front view posture, αc2 is the angle between the optical axis of the camera lens and the vertical direction; αs2 is the angle between the optical axis of the light source and the vertical direction; d2 is the horizontal distance from the center of the camera lens to the center of the light source; (the effect of the front view is similar to that of the side view and is not drawn).

[0019] It should be noted that the front view also has the vertical distance between the camera and the light source relative to the plane of the object being measured, such as Figure 3 As shown, that is, hc2 and hs2, except that hc2=hc1, hs2=hs1.

[0020] 2. Imaging Path Planning 1. Artificial optimal imaging posture The imaging is manually adjusted until the best imaging effect is achieved, and defects are annotated at the pixel level. The pose state at this time is regarded as the initial pose state; Graded imaging Adjust the system posture state to obtain images at each level of imaging effect grading. The imaging effect level depends on the specific application scenario. The imaging effect level used in this invention is divided into 4 levels: Grade A: The contrast between the defect and surrounding pixels exceeds 50 (50 is a numerical value); Grade B: The contrast between the defect and surrounding pixels is between 30 and 50; Grade C: The contrast ratio between the defect and surrounding pixels is between 15 and 30; Level D: The contrast between the defect and surrounding pixels is less than 15; One image of each of the above grades is obtained, which serves as a reference standard for subsequent grading of batch-collected images and is called imaging grade data (corresponding to the content of step (1) under the third sub-step in the third major step).

[0021] 2. Imaging path: Let the starting and ending values ​​of hc1, αc1, αc2, hs1, αs1, αs2, d1, d2, hc2, and hs2, the adjustment step size, and the values ​​under the optimal posture be: hc1:hc1_s,hc1_e, hc1_m, hc1_b; αc1:αc1_s,αc1_e,αc1_m,αc1_b; αc2:αc2_s,αc2_e,αc2_m,αc2_b; hs1:hs1_s,hs1_e,hs1_m,hs1_b; αs1:αs1_s,αs1_e,αs1_m,αs1_b; αs2:αs2_s,αs2_e,αs2_m,αs2_b; d1:d1_s,d1_e,d1_m,d1_b; d2:d2_s,d2_e,d2_m,d2_b; hc2:hc2_s,hc2_e, hc2_m, hc2_b; hs2:hs2_s,hs2_e,hs2_m,hs2_b; For example: hc1_s is the starting value of hc1, hc1_e is the end value of hc1, hc1_m is the traversal step length of hc1, hc1_b; is the value of hc1 under the best posture; Others are similar.

[0022] With the initial pose as the center, an area is formed around it and traversed with a certain step length. The area range and step length are determined by multiplying the initial pose by a coefficient; this coefficient is designed and adjusted according to the specific application scenario. like: hc1_s=hc1_b*0.5; hc2_s=hc2_b*0.5; hs1_s=hs1_b*0.5; hs2_s=hc2_b*0.5; hc1_e=hc1_b*1.5; hc2_e=hc2_b*1.5; hs1_e=hs1_b*1.5; hs2_e=hc2_b*1.5; hc1_m=hc1_b*0.05; hc2_m=hc2_b*0.05; hs1_m=hs1_b*0.05; hs2_m=hc2_b*05 αc1_s=0;αc2_s=0;αs1_s=0;αs2_s=0; αc1_e=90;αc2_e=90;αs1_e=90;αs2_e=90; αc1_m=5°;αc2_m=5°;αs1_m=5°;αs1_m=5°; d1_s=d1_b*0.5;d2_s=d2_b*0.5;d1_e=d1_e*1.5;d2_e=d2_e*1.5; d1_m=d1_b*0.05; d2_m=d2_b*0.05; The above specific values ​​are the values ​​adopted by the present invention and can be adjusted according to the actual application scenario requirements. The formula can be transformed, such as hc1_b = hc1_s / 0.5.

[0023] 3. Imaging Data Acquisition Traverse according to the value set in step 2, perform imaging at each traversed posture, and record the current posture state. And according to the posture change between the optimal posture state and the current posture state, perform affine transformation on the pixel-level annotation at the optimal posture state to obtain the pixel-level annotation of the imaging data at the current posture state, that is, the imaging data at each posture has a pixel-level annotation; Using each imaging data and its corresponding pixel-level annotation, the score index of all imaging data is calculated as follows: 1. Calculate contrast: Contrast refers to the contrast between the grayscale of pixels within the pixel-level annotated area of ​​the imaging data and the surrounding pixels in the pixel-level annotated area. The calculation method is as follows: (1) Obtain the average grayscale m_g_d of the corresponding image position at the pixel level annotation; (2) Obtain the average grayscale m_g_n of the image in the pixel area (i.e., n_p pixels are expanded outward, and n_p is 3 in this invention); (3) Contrast = |m_g_d - m_g_n|; 2. Computational integrity: Binarize the image with a threshold of |m_g_d - m_g_n| / 2, and the intersection-and-union ratio of the connected domain that intersects with the pixel-level annotation is the completeness. 3. Calculate discrimination Calculate the average grayscale of all connected domains obtained by binarization. Let the average grayscale closest to m_g_d be m_g_x, then the discrimination = |m_g_d-m_g_x| / m_g_d; After obtaining the scoring indicators composed of contrast, integrity, and discrimination, the effect level of each imaging data is judged by the scoring indicators. The judgment method is: (1) Calculation difference The function for calculating the difference is: f(x,y)=|xy| / y The difference between the data calculated by function f and the imaging grade data obtained in step 2 is calculated. That is, the contrast, integrity, and difference of the current data are respectively taken as x, and the contrast, integrity, and difference of the imaging grade data are correspondingly substituted into the function f to obtain the difference values ​​of contrast, integrity, and difference. The largest of these three difference values ​​is taken as the difference value between the current data and the imaging grade data. The smallest difference between the current data and the imaging grade data is taken as the grade difference value of the current data; (2) Classification If the grade difference value is less than 0.1, the grade of the data belongs to the grade corresponding to the grade data with the smallest value difference; For data with a difference of more than 0.1 from all imaging grade data, manual re-judgment is performed to determine its grade; This step obtains a data set, each data in the data set includes: hc1_idx: the vertical distance between the side view camera and the plane of the object being measured; hs1_idx: the vertical distance between the side view light source and the plane of the object being measured; αc1_idx: vertical angle of the side view camera; αc2_idx: vertical angle of the orthographic camera; αs1_idx: vertical angle of the side view light source; αs2_idx: vertical angle of the front view light source; d1_idx: the horizontal distance between the side view camera center and the light source center; d2_idx: the horizontal distance between the center of the orthographic camera and the center of the light source; Where: idx=1,2,3,.....N is the sequence number of the current data, and N is the number of data in the data set; The above operations obtain the imaging level of all imaging data; 4. Training Model The model operation principle diagram is as follows Figure 5 As shown: 1. Large language model information extraction (participates in training and inference, not parameter updates) Select a large language model, denoted as LLM, and deepseek14b is used in this invention; Set the system prompt words to: "You are an expert in designing appearance inspection imaging systems. Your task is to extract product and material information from user input, retrieve area, defect type, and severity descriptions, and form a description of the inspection requirements. For example, when the input is: (1) Industry: 3C; (2) Product: iPhone; (3) Components: Cover glass: (4) Detection area: curved surface: (5) Defects to be inspected: minor scratches; At this point, the extraction result is: slight scratches on the curved surface of the mobile phone cover glass;" Input the user's question into the large language model and obtain the corresponding text output. In the user input of the training phase, the present invention uses the format of the above system prompt word input example; 2. Construct data description text Add the following to the output of the large language model in the above step: "The vertical distance between the side view camera and the plane of the object being measured = {hc1_} The vertical distance between the side view light source and the plane of the object being measured = {hs1_}; The vertical angle of the side view camera = {αc1_}; The vertical angle of the front view camera = {αc2_}; The vertical angle of the side view light source = {αs1_}; The vertical angle of the front view light source = {αs2_}; The horizontal distance between the side view camera center and the light source center = {d1_}; The horizontal distance between the center of the front view camera and the center of the light source = {d2_};”, The variable part is in {}, which is filled with the value corresponding to each data, so that each data has a description text; 3. Training text classification models (participating in training and inference, and participating in parameter updates) The description text is input into a prediction model that uses text for prediction. The predicted label is the imaging grade of the data, thereby training the text classification model. The base model used in this invention is the Bert model, which is used for grade prediction tasks (i.e., classification tasks). 4. Large language model output (only involved in inference, not parameter update) The large language model still uses the large language model selected above in extracting user input information, but the prompt words are different from the above. The prompt words are designed and the corresponding answers are generated using the large language model. The prompt words used in the present invention are described in detail in the model application below.

[0024] 5. Application Model 1. Information Extraction Similar to the training process, the required information is obtained from the user input and converted into text output; 2. Imaging path planning and prediction Similar to the training phase, the starting range and traversal step size of the pose parameters in the imaging data acquisition are customized. The data description text is constructed according to the method of the training phase. The constructed text is input into the imaging prediction network for prediction, and the predicted output of all data description texts is obtained. 3. Imaging pose selection Among all the predicted outputs, the output result is the data corresponding to the optimal level, and n_i poses are obtained by random sampling among them. In this invention, n_i=3; 4. Large language model output The system prompt is: "Answer the user's questions based on the results output by the prediction model;" as well as, Format the posture selected by the prediction model into user prompt words in the following format: For the side view camera, the vertical distances between the light source and the plane of the object being measured are {30 mm, 25 mm} respectively; for the side view camera, the angles between the light source and the vertical direction are {30° and 60°} respectively; and the horizontal distance between them is {50 mm}; for the front view camera, the angles between the light source and the vertical direction are {0° and 0}° respectively; and the horizontal distance between them is {0 mm}; The {} is a variable part, which changes the specific value according to the result value output by the prediction model; Input the large language model, which will understand the sentence, organize the language, and then output the result.

[0025] Additional optional actions: Develop a drawing program based on the pose parameters in advance, and call the drawing program based on the predicted pose parameters to obtain the light path diagram; this can be considered an intelligent agent; 5. Manual adjustment verification The posture parameters output by the large language model are used as the initial posture to build the imaging system, and fine-tuning is performed on this basis until the imaging effect reaches the best.

[0026] The above description only describes specific embodiments of the present invention. Various examples do not limit the essential content of the present invention. After reading the description, ordinary technicians in the relevant technical field can make modifications or variations to the specific embodiments described above without departing from the essence and scope of the invention.

Claims

1. An industrial imaging posture adjustment system based on a large language model, characterized by: include, A fixing device for fixing the object to be measured; The robot device includes two manipulators, which respectively grasp the camera and the light source, can record the current posture state, and can adjust to the specified posture according to the posture adjustment instruction; The data acquisition module is used to traverse the postures according to the set imaging path, perform imaging at each posture, record the current posture state, process the imaging data to obtain pixel-level annotations and scoring indicators, and determine the imaging level; The model training module includes a large language model and a text classification model. The large language model is used to extract information from user input, and the text classification model is trained with a description text containing the extracted information and pose parameters as input and imaging levels as labels. The application module is used to receive the measured object and defect information input by the user and output the reference pose through the trained model.

2. The industrial imaging posture adjustment system based on a large language model according to claim 1, characterized in that: The imaging posture adjustment method includes the following steps: S1. Build the device: Fix the object to be measured, use two manipulators to grab the camera and light source respectively, and determine the initial pose according to the imaging mode; S2. Plan the imaging path: set the starting value, ending value and traversal step length of the posture parameters; S3. Collect imaging data: traverse the posture according to the path, record the posture status, process the imaging data to obtain pixel-level annotation and scoring indicators, and determine the imaging level; S4. Training model: Use the large language model to extract information, construct a description text containing information and posture parameters, and train the text classification model; S5. Application model: Extract user demand information, plan the path and construct description text, predict the imaging level through the model, select the optimal posture and output the reference posture through the large language model, and manually fine-tune to the best effect.

3. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: The imaging mode includes a reflection mode and a transmission mode. In the reflection mode, the camera and the light source are located on the same side of the object to be measured. In the transmission mode, the camera and the light source are located on both sides of the object to be measured.

4. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: The posture parameters include the vertical distance from the center of the camera lens to the object to be measured, the angle between the optical axis of the camera lens and the vertical direction, the vertical distance from the center of the light source to the object to be measured, the angle between the optical axis of the light source and the vertical direction, and the horizontal distance from the center of the camera lens to the center of the light source.

5. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: The scoring indicators include contrast, completeness and discrimination. Contrast is the average grayscale difference between the marked area and the surrounding area. Completeness is the intersection-over-union ratio of the connected domain and the annotation after binarization. Discrimination is the difference ratio between the average grayscale of the relevant connected domain and the average grayscale of the marked area.

6. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: The imaging grades are divided into A, B, C and D. Grade A corresponds to contrast > 50, grade B corresponds to 30≤contrast≤50, grade C corresponds to 15≤contrast≤30, and grade D corresponds to contrast < 15.

7. The industrial imaging posture adjustment system based on a large language model according to claim 1 or 2, characterized in that: The large language model is the deepseek14b model, and the text classification model is based on the Bert model.

8. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: The starting value and ending value of the posture parameter are determined based on the value under the optimal posture, and the step size is set according to the actual scene.

9. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: When processing the imaging data, pixel-level annotations at the current position are obtained through affine transformation.

10. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: When applying the model, three poses corresponding to the optimal level are randomly sampled as reference poses.

Citation Information

Patent Citations

  • Automatic calibration and feedback method and system

    CN118135026A

  • Industrial robot assembly method and system based on multi-modal large model

    CN118744425A

  • Apparatus and method for detecting user intention for image capture or video recording

    CN119422382A

  • Self-adaptive system for imaging of industrial line-scan digital camera

    CN119450189A

  • LLM model-based defect detection system

    CN119688686A