Industrial imaging pose adjustment system based on large language model

By using an industrial imaging pose adjustment system based on a large language model, the optimal pose is predicted by a robotic arm and a model, solving the problems of cumbersome manual adjustment and unstable results in existing technologies, and achieving efficient and stable imaging results.

CN120635191BActive Publication Date: 2025-10-17BEIJING FOCUSIGHT TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511120361.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-10-17
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing industrial imaging pose adjustment methods rely on human experience or evaluation functions, resulting in cumbersome operation, low efficiency, and unstable results, making it difficult to achieve the best imaging effect in different scenarios.

Method used

An industrial imaging pose adjustment system based on a large language model is adopted. The pose of the camera and light source is adjusted by a robotic arm device. Combined with a data acquisition module and a model training module, the optimal pose is predicted by a large language model and a text classification model, and the best imaging effect is achieved through manual fine-tuning.

Benefits of technology

It reduces the workload and experience requirements of optoelectronic engineers, improves the stability and consistency of imaging effects, reduces labor costs, and the imaging effect is closer to human evaluation, adapting to the imaging needs of different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635191B_ABST
    Figure CN120635191B_ABST
Patent Text Reader

Abstract

The present invention relates to an industrial imaging posture adjustment system based on a large language model, comprising a fixing device, a manipulator device, a data acquisition module, a model training module, and an application module. By constructing a device, planning an imaging path, acquiring imaging data, and training a large model, the model is finally used to output a reference posture, and combined with manual fine-tuning, efficient and stable adjustment of the industrial imaging posture is achieved. The present invention obtains a large imaging posture model through training, and by inputting information about the object to be measured and defects, a reference posture is given, which greatly reduces the time consumption and instability of manual adjustment; by obtaining a reference posture, the time consumption of manual adjustment is reduced, time costs are saved, the requirements for operators are reduced, labor costs are saved, the stability of the effect is improved, and the imaging quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of visual imaging technology, and in particular to an industrial imaging pose adjustment system based on a large language model. BACKGROUND

[0002] In the industrial imaging process, the position and angle (i.e., pose) of the camera and light source relative to the measured object is crucial and directly affects the imaging quality. Currently, there are two main ways to adjust the pose:

[0003] 1. Manual adjustment

[0004] The position and angle (i.e., pose) of the light source and camera relative to the measured object are adjusted by the photoelectric engineer according to experience;

[0005] The steps are:

[0006] (1) According to experience, select an initial pose;

[0007] (2) Manually adjust the pose of the light source and camera, and according to the observation of the imaging effect by the human eye, adjust the pose of the light source and camera according to the feedback and experience;

[0008] (3) Keep adjusting until a pose that the photoelectric engineer considers satisfactory is obtained;

[0009] 2. Automatic adjustment

[0010] There are some intelligent imaging devices that use a mechanical hand to adjust the pose of the light source and camera. During the adjustment process, an evaluation function is used to measure the quality of the imaging effect, and the mechanical hand path is traversed or planned to adjust the pose. The pose with the highest evaluation function score is selected as the final pose.

[0011] However, both of the above methods have defects:

[0012] 1. Manual adjustment requires experienced engineers to adjust based on their own experience, which requires high operating skills and increases labor costs. The adjustment process is tedious and manual, which is inefficient. Different photoelectric engineers have different experiences and abilities, and the adjustment results will vary greatly, resulting in unstable effects;

[0013] 2. Automatic adjustment relies entirely on the evaluation function to measure the quality of the imaging effect, and it is difficult to completely match the evaluation function score with the actual imaging effect, resulting in poor results. Moreover, the correlation between the evaluation function score and the true effect varies greatly in different scenarios, making it difficult to be universally applicable. SUMMARY

[0014] The technical problem solved by the present application is to provide an industrial imaging pose adjustment system based on a large language model, which solves the problem of the need for a large number of repeated manual adjustments, complexity and unstable effect in the pose adjustment stage of industrial imaging.

[0015] The technical solution adopted by the present application to solve its technical problem is: an industrial imaging pose adjustment system based on a large language model, comprising,

[0016] A fixing device for fixing a measured object;

[0017] A mechanical hand device comprising two mechanical hands for grabbing a camera and a light source respectively, capable of recording the current pose state and adjusting to the specified pose according to the pose adjustment instruction;

[0018] A data acquisition module for traversing the pose according to the set imaging path, imaging at each pose, recording the current pose state, processing the imaging data to obtain pixel-level labeling and score indicators, and determining the imaging level;

[0019] A model training module comprising a large language model and a text classification model, the large language model being used to extract information from user input, and the text classification model being trained with a description text containing the extracted information and pose parameters as input and the imaging level as label;

[0020] An application module for receiving user input of the measured object and defect information, and outputting a reference pose through the trained model.

[0021] Further, the imaging mode of the present application includes a reflection mode and a transmission mode, in the reflection mode the camera and the light source are located on the same side of the measured object, and in the transmission mode the camera and the light source are located on the two sides of the measured object.

[0022] Further, the pose parameters of the present application include the vertical distance from the camera lens optical center to the measured object, the angle between the camera lens optical axis and the vertical direction, the vertical distance from the light source center to the measured object, the angle between the light source optical axis and the vertical direction, and the horizontal distance from the camera lens center to the light source center.

[0023] Further, the score indicators of the present application include contrast, integrity and discrimination, the contrast is the average gray difference between the labeled area and the surrounding area, the integrity is the intersection over union of the connected domain after binarization and the label, and the discrimination is the difference ratio of the average gray of the relevant connected domain and the average gray of the labeled area.

[0024] Further, the imaging level of the present application is divided into A, B, C and D levels, A level corresponds to contrast > 50, B level corresponds to 30≤contrast≤50, C level corresponds to 15≤contrast≤30, and D level corresponds to contrast<15.

[0025] Further, the large language model of the present application is a deepseek14b model, and the text classification model is based on a Bert model.

[0026] Still further, the present application includes an imaging pose adjustment method, which comprises the following steps:

[0027] S1, building a device: fixing the measured object, grabbing the camera and the light source by two mechanical hands, and determining the initial pose according to the imaging mode;

[0028] S2, planning an imaging path: setting the initial value, the terminal value and the traversal step of the pose parameter;

[0029] S3, collecting imaging data: traversing the pose according to the path, recording the pose state, processing the imaging data to obtain pixel-level labeling and score indicators, and judging the imaging level;

[0030] S4, training a model: extracting information using a large language model, constructing a description text containing information and pose parameters, and training a text classification model;

[0031] S5, applying a model: extracting user demand information, planning a path and constructing a description text, predicting the imaging level through the model, selecting the optimal pose and outputting the reference pose through the large language model, and manually fine-tuning to the best effect.

[0032] Further, the initial value and the terminal value of the pose parameter are determined based on the value of the optimal pose, and the step is set according to the actual scene.

[0033] Further, when processing the imaging data, the pixel-level labeling under the current pose is obtained through affine transformation.

[0034] Further, when applying the model, 3 reference poses are randomly sampled from the pose corresponding to the optimal level.

[0035] The present application has the advantages of solving the defects in the background art,

[0036] 1. By using a large amount of manual adjustment data and designing prompt words as samples, the large model is fine-tuned, allowing the large model to learn the changing rules between posture adjustment and imaging effects in various scenarios. This enables the large model to design imaging postures according to specific input scenarios. The large model will give the corresponding posture through input, and the output posture is the optimal posture, which can be used as the initial posture. Then, manual adjustments can be made to this initial posture. This reduces the workload of optoelectronic engineers, lowers the experience and ability requirements of optoelectronic engineers, saves time and labor costs, and different engineers can make small adjustments under the same reference posture. The final result is also stable, making the imaging effect stable and easy to achieve the optimal effect.

[0037] 2. Using the trained AI model to predict the imaging effect is more powerful than the traditional effect evaluation function fitting function and is closer to human judgment. It can reduce the workload of optoelectronic engineers by manually making small adjustments. During the adjustment process, it is manually judged. Compared with the evaluation function, the effect judgment is more reasonable and more universal, ensuring the final imaging effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is the overall flow chart of the training and application of the large-scale model for industrial imaging posture adjustment of the present invention;

[0039] Figure 2 This is a side view posture relationship diagram of the light source, camera, and object under reflection in the present invention;

[0040] Figure 3 It is a diagram showing the relationship between the light source, the camera, and the object under test in one of the front-view poses under the reflection condition of the present invention;

[0041] Figure 4 This is a side view posture relationship diagram of the transmission mode light source, camera, and object under test of the present invention;

[0042] Figure 5 It is a schematic diagram of the model architecture of the present invention. DETAILED DESCRIPTION

[0043] The present invention will now be described in further detail with reference to the accompanying drawings and preferred embodiments. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0044] like Figures 1-5 The overall implementation steps of the industrial imaging posture adjustment system based on the large language model are as follows:

[0045] 1. Device Construction

[0046] The measured object is fixed, and two mechanical hands respectively grab the camera and the light source to adjust the pose, the mechanical hand can record the current pose state, and can be adjusted to the specified pose according to the pose adjustment instruction, the construction method of the device is divided into reflection mode and transmission mode according to the difference of the imaging mode, respectively as shown in Figure 2 and Figure 4 The side view pose diagram of the reflection mode and the transmission mode is shown, and the plane where the measured object is located is taken as the plane reference system, wherein:

[0047] hc1 is the vertical distance from the optical center of the camera lens to the measured object;

[0048] αc1 is the angle between the camera lens optical axis and the vertical direction;

[0049] hs1 is the vertical distance from the light source center to the measured object;

[0050] αs1 is the angle between the light source optical axis and the vertical direction;

[0051] d1 is the horizontal distance from the camera lens center to the light source center;

[0052] The front view pose diagram is the same, and when the front view pose is,

[0053] αc2 is the angle between the camera lens optical axis and the vertical direction;

[0054] αs2 is the angle between the light source optical axis and the vertical direction;

[0055] d2 is the horizontal distance from the camera lens center to the light source center; (the front view effect is similar to the side view, not drawn).

[0056] It should be noted that the front view also has the vertical distance of the camera and the light source relative to the plane of the measured object, as shown in Figure 3 , that is, hc2 and hs2, only hc2=hc1, hs2=hs1.

[0057] II. Imaging path planning

[0058] 1. Artificial optimal imaging pose

[0059] The imaging is adjusted by artificial adjustment until the best imaging effect is obtained, and the pixel-level labeling of defects is performed, and the pose state at this time is taken as the initial pose state;

[0060] Graded imaging

[0061] Adjust the system pose state to obtain images under each level of imaging effect grading, and the imaging effect level is determined by the specific application scene, and the imaging effect level adopted in the present application is divided into four levels:

[0062] A level: the contrast between the defect and the surrounding pixels exceeds 50 (50 is a numerical value);

[0063] B: the contrast between the defect and the surrounding pixels is between 30 and 50;

[0064] C: the contrast between the defect and the surrounding pixels is between 15 and 30;

[0065] D: the contrast between the defect and the surrounding pixels is less than 15;

[0066] One image of each of the above grades is obtained as a reference standard for subsequent grading of images obtained in batch acquisition, referred to as imaging grade data (corresponding to the content in step (1) under the 3rd sub-step in the 3rd major step).

[0067] 2. Imaging path:

[0068] Let hc1, ac1, ac2, hsl, asl, as2, dl, d2, hc2, hs2 correspond to the starting and ending values, the adjustment step size and the values under the optimal pose respectively:

[0069] hc1: hc1_s, hc1_e, hc1_m, hc1_b;

[0070] ac1: ac1_s, ac1_e, ac1_m, ac1_b;

[0071] ac2: ac2_s, ac2_e, ac2_m, ac2_b;

[0072] hsl: hsl_s, hsl_e, hsl_m, hsl_b;

[0073] asl: asl_s, asl_e, asl_m, asl_b;

[0074] as2: as2_s, as2_e, as2_m, as2_b;

[0075] dl: dl_s, dl_e, dl_m, dl_b;

[0076] d2: d2_s, d2_e, d2_m, d2_b;

[0077] hc2: hc2_s, hc2_e, hc2_m, hc2_b;

[0078] hs2: hs2_s, hs2_e, hs2_m, hs2_b;

[0079] For example, hc1_s is the starting value of hc1,

[0080] hc1_e is the ending value of hc1,

[0081] hc1_m is the traversal step of hc1,

[0082] hc1_b is the value of hc1 in the best pose;

[0083] Other similar.

[0084] Centered on the initial pose, a region is formed around it, and traversal is performed at a certain step; the region range and step are determined by multiplying the initial pose by a coefficient; this coefficient is designed and adjusted by the specific application scenario;

[0085] For example:

[0086] hc1_s=hc1_b*0.5; hc2_s=hc2_b*0.5; hs1_s=hs1_b*0.5; hs2_s=hc2_b*0.5;

[0087] hc1_e=hc1_b*1.5; hc2_e=hc2_b*1.5; hs1_e=hs1_b*1.5; hs2_e=hc2_b*1.5;

[0088] hc1_m=hc1_b*0.05; hc2_m=hc2_b*0.05;hs1_m=hs1_b*0.05; hs2_m=hc2_b*05

[0089] αc1_s=0;αc2_s=0;αs1_s=0;αs2_s=0;

[0090] αc1_e=90;αc2_e=90;αs1_e=90;αs2_e=90;

[0091] αc1_m=5°;αc2_m=5°;αs1_m=5°;αs1_m=5°;

[0092] d1_s=d1_b*0.5;d2_s=d2_b*0.5;d1_e=d1_e*1.5;d2_e=d2_e*1.5;

[0093] d1_m=d1_b*0.05; d2_m=d2_b*0.05;

[0094] The above specific values are the values taken by the present application, which can be adjusted according to the actual application scenario requirements. The formula can be transformed, such as hc1_b= hc1_s / 0.5.

[0095] III. Image data acquisition

[0096] According to the value set in step two, traversal is performed, imaging is performed at each pose in the traversal, and the current pose state is recorded, and the pixel-level label of the best pose state is subjected to affine transformation according to the pose change amount of the best pose state and the current pose state, to obtain the pixel-level label of the imaging data in the current pose state, that is, the imaging data in each pose has a pixel-level label;

[0097] Using each imaging data and its corresponding pixel-level label, the score index of all imaging data is calculated, and the calculation method is:

[0098] 1. Calculate the contrast:

[0099] The contrast refers to the contrast of the pixel gray scale in the pixel-level label region of the imaging data and the pixels around the pixel-level label region, and the calculation method is as follows:

[0100] (1) Obtain the average gray scale m_g_d of the image position corresponding to the pixel-level label;

[0101] (2) Obtain the average gray scale m_g_n of the image in the pixel-level label region (i.e. the label is expanded by n_p pixels, and n_p is 3 in the present application);

[0102] (3) Contrast = |m_g_d - m_g_n|;

[0103] 2. Calculate the integrity:

[0104] According to the threshold |m_g_d - m_g_n| / 2, the image is binarized, and the intersection ratio of the connected domain obtained and the pixel-level label is the integrity of the pixel-level label;

[0105] 3. Calculate the distinction

[0106] The average gray scale of all connected domains obtained by binarization is calculated, and the average gray scale closest to m_g_d is m_g_x, and the distinction = |m_g_d - m_g_x| / m_g_d;

[0107] After obtaining the score index composed of contrast, integrity and distinction, the effect level of each imaging data is judged by the score index, and the judgment method is:

[0108] (1) Calculate the difference

[0109] The function for calculating the difference is: f(x,y)=|x-y| / y

[0110] The difference between the data calculated by the function f and the imaging level data obtained in step two is calculated,

[0111] That is, respectively, the contrast, integrity, difference of the current data as x, the contrast, integrity, difference of the imaging level data as y, the difference value of the contrast, integrity, difference is obtained by substituting into the function f, and the maximum of the three difference values is taken as the difference value between the current data and the imaging level data;

[0112] The minimum difference value between the current data and the imaging level data is taken as the level difference value of the current data;

[0113] (2) Level classification

[0114] If the level difference value is less than 0.1, the level of the data belongs to the level corresponding to the minimum value difference level data;

[0115] For data with a difference higher than 0.1 from all imaging level data, manual rejudgment is performed to determine its level;

[0116] This step obtains a data set, each data in the data set includes:

[0117] hc1_idx: the vertical distance between the side view camera and the measured object plane;

[0118] hs1_idx: the vertical distance between the side view light source and the measured object plane;

[0119] αc1_idx: the vertical angle of the side view camera;

[0120] αc2_idx: the vertical angle of the front view camera;

[0121] αs1_idx: the vertical angle of the side view light source;

[0122] αs2_idx: the vertical angle of the front view light source;

[0123] d1_idx: the horizontal distance between the center of the side view camera and the center of the light source;

[0124] d2_idx: the horizontal distance between the center of the front view camera and the center of the light source;

[0125] Wherein: idx=1, 2, 3,..., N is the serial number of the current data, and N is the number of data in the data set;

[0126] The above operation obtains the imaging level of all imaging data;

[0127] Four, train the model

[0128] The model running principle diagram is as shown in Figure 5 ;

[0129] 1. Large language model extracts information (participates in training and reasoning, does not participate in parameter update)

[0130] Select a large language model, denoted as LLM, which is deepseek14b in this invention;

[0131] Set the system prompt word as:

[0132] "You are an appearance detection imaging system design expert, your task is to extract product and material from user input, search area, defect type and severity description information, and form a description of detection requirements;

[0133] For example, when the input is:

[0134] (1) Industry: 3C;

[0135] (2) Product: Apple phone;

[0136] (3) Component: Cover glass:

[0137] (4) Detection area: Curved part:

[0138] (5) Defects to be detected: Slight scratch;

[0139] At this time, the extracted result is: slight scratch on the curved area of the cover glass of the mobile phone; "

[0140] Input the user's question into the large language model to get the corresponding text output, and in the training stage, the user input uses the format of the above system prompt word input example;

[0141] 2. Construct data description text

[0142] Add the following to the output of the large language model in the above step:

[0143] "The vertical distance between the side view camera and the measured object plane = {hc1_}

[0144] The vertical distance between the side view light source and the measured object plane = {hs1_};

[0145] The vertical angle of the side view camera = {αc1_};

[0146] The vertical angle of the front view camera = {αc2_};

[0147] The vertical angle of the side view light source = {αs1_};

[0148] The vertical angle of the front view light source = {αs2_};

[0149] The horizontal distance between the center of the side view camera and the center of the light source = {d1_};

[0150] Horizontal distance between the center of the front view camera and the center of the light source = {d2_} ;

[0151] Where {} is a variable part, fill in the value corresponding to each data, so that each data has a description text;

[0152] 3. Training text classification model (participate in training and inference, participate in parameter update)

[0153] The description text is input into a text prediction model to make predictions, and the predicted label is the imaging level of the data to train the text classification model. The base model used in the present application is the Bert model, which is used for level prediction tasks (i.e. classification tasks);

[0154] 4. Large language model output (only participates in inference, does not participate in parameter update)

[0155] The large language model is also the large language model selected in the above extraction of user input information, but the prompt words are different from the above. The prompt words are designed to generate corresponding answers using the large language model. The prompt words used in the present application are described in detail in the following model application.

[0156] Five, application model

[0157] 1. Information extraction

[0158] Similar to the training process, the requirement information obtained from the user input is converted into text output;

[0159] 2. Imaging path planning and prediction

[0160] Similar to the training stage, the starting range and traversal step of the pose parameter in the imaging data collection are customized. The data description text is constructed according to the method of the training stage, and the constructed text is input into the imaging prediction network to make predictions, and the prediction output of all data description texts is obtained;

[0161] 3. Imaging pose selection

[0162] Among all the prediction outputs, the output result is the data corresponding to the optimal level. Random sampling is performed among them to obtain n_i poses. In the present application, n_i = 3;

[0163] 4. Large language model output

[0164] The system prompt word is: "According to the output of the prediction model, answer the question raised by the user;"

[0165] And,

[0166] The pose selected by the prediction model is formatted into a user prompt word, and the format is:

[0167] Side view camera, the vertical distance between light source and measured object plane is {30mm, 25mm} respectively; Side view camera, the angle between light source and vertical direction is {30° and 60°} respectively; The horizontal distance between them is {50mm}; Front view camera, the angle between light source and vertical direction is {0° and 0°} respectively; The horizontal distance between them is {0mm};

[0168] Wherein {} is a variable part, the specific value changes according to the result value output by the prediction model;

[0169] Inputting a large language model, outputting a result after the large language model understands a sentence and organizes language.

[0170] Additional optional operations:

[0171] Develop a drawing program based on pose parameters in advance, and call the drawing program to obtain the light path diagram based on the obtained predicted pose parameters; Thus, it can be calculated as an intelligent agent;

[0172] 5. Artificial adjustment verification

[0173] Taking the output pose parameter of the large language model as the initial pose, an imaging system is built, and on this basis, fine tuning is performed until the imaging effect reaches the best.

[0174] The above description is only a specific embodiment of the present application, and various examples do not limit the essential content of the present application, and those skilled in the art can modify or deform the previously described specific embodiments without departing from the essence and scope of the present application after reading the specification.

Claims

1. An industrial imaging posture adjustment system based on a large language model, characterized by: include, A fixing device for fixing the object to be measured; The robot device includes two manipulators, which respectively grasp the camera and the light source, can record the current posture state, and can adjust to the specified posture according to the posture adjustment instruction; The data acquisition module is used to traverse the postures according to the set imaging path, perform imaging at each posture, record the current posture state, process the imaging data to obtain pixel-level annotations and scoring indicators, and determine the imaging level; The model training module includes a large language model and a text classification model. The large language model is used to extract information from user input, and the text classification model is trained with a description text containing the extracted information and pose parameters as input and imaging levels as labels. The application module is used to receive the measured object and defect information input by the user and output the reference pose through the trained model.

2. The industrial imaging posture adjustment system based on a large language model according to claim 1, characterized in that: The imaging posture adjustment method includes the following steps: S1. Build the device: Fix the object to be measured, use two manipulators to grab the camera and light source respectively, and determine the initial pose according to the imaging mode; S2. Plan the imaging path: set the starting value, ending value and traversal step length of the posture parameters; S3. Collect imaging data: traverse the posture according to the path, record the posture status, process the imaging data to obtain pixel-level annotation and scoring indicators, and determine the imaging level; S4. Training model: Use the large language model to extract information, construct a description text containing information and posture parameters, and train the text classification model; S5. Application model: Extract user demand information, plan the path and construct description text, predict the imaging level through the model, select the optimal posture and output the reference posture through the large language model, and manually fine-tune to the best effect.

3. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: The imaging mode includes a reflection mode and a transmission mode. In the reflection mode, the camera and the light source are located on the same side of the object to be measured. In the transmission mode, the camera and the light source are located on both sides of the object to be measured.

4. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: The posture parameters include the vertical distance from the center of the camera lens to the object to be measured, the angle between the optical axis of the camera lens and the vertical direction, the vertical distance from the center of the light source to the object to be measured, the angle between the optical axis of the light source and the vertical direction, and the horizontal distance from the center of the camera lens to the center of the light source.

5. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: The scoring indicators include contrast, completeness and discrimination. Contrast is the average grayscale difference between the marked area and the surrounding area. Completeness is the intersection-over-union ratio of the connected domain and the annotation after binarization. Discrimination is the difference ratio between the average grayscale of the relevant connected domain and the average grayscale of the marked area.

6. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: The imaging grades are divided into A, B, C and D. Grade A corresponds to contrast > 50, grade B corresponds to 30≤contrast≤50, grade C corresponds to 15≤contrast≤30, and grade D corresponds to contrast < 15.

7. The industrial imaging posture adjustment system based on a large language model according to claim 1 or 2, characterized in that: The large language model is the deepseek14b model, and the text classification model is based on the Bert model.

8. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: The starting value and ending value of the posture parameter are determined based on the value under the optimal posture, and the step size is set according to the actual scene.

9. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: When processing the imaging data, pixel-level annotations at the current position are obtained through affine transformation.

10. The industrial imaging posture adjustment system based on a large language model according to claim 2, characterized in that: When applying the model, three poses corresponding to the optimal level are randomly sampled as reference poses.

Citation Information

Patent Citations

  • Self-adaptive system for imaging of industrial line-scan digital camera

    CN119450189A

  • Using language models in autonomous and semi-autonomous systems and applications

    US20240420418A1