Fine-tuning and Inference Methods for Large Model of Drawing Recognition

By fine-tuning and reasoning the drawing recognition large model, using data augmentation and model parameter adjustment, the problem of accuracy and low efficiency of complex engineering drawing recognition is solved, and higher generalization ability and robustness are achieved.

CN119740672BActive Publication Date: 2025-07-01WUXI XUELANG DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510228475.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-01
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

In the prior art, when dealing with complex engineering drawings, there are problems of poor identification accuracy and low efficiency.

Method used

The fine-tuning and inference method of drawing recognition large models are used. By obtaining original sample data, labeling information and task statements, data augmentation and model parameter adjustments are carried out, and the initial drawing recognition model is iteratively fine-tuned, and the target drawing recognition model is generated.

Benefits of technology

Improves the generalization and robustness of the target drawing recognition model, allowing it to accurately locate and understand the information required by users in complex engineering drawings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119740672B_ABST
    Figure CN119740672B_ABST
Patent Text Reader

Abstract

The present application provides a method for fine-tuning and inference of a large model for drawing recognition. Among them, the method includes: obtaining original sample data; performing augmentation processing according to the original sample data and annotation information to generate training sample data; in the current training round, obtaining an image sequence to be processed and a text sequence to be processed; adjusting by a feature processing module and performing feature processing to obtain a feature processing result; adjusting by a multi-modal recognition module and performing multi-modal recognition to obtain a recognition result; determining a loss result according to the recognition result, the true value, and a preset loss function, and iteratively fine-tuning an initial drawing recognition model according to the loss result to obtain a target drawing recognition model, and using the target drawing recognition model for inference. The present application enables the target drawing recognition model to accurately locate and understand the information required by the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology. Specifically, it relates to a method for fine-tuning and inference of a large model for drawing recognition. Background Art

[0002] With the rapid development of mechanical design and manufacturing, engineering drawings play an increasingly important role in technical communication and product manufacturing. It not only ensures the accurate transmission of design concepts and the precise manufacturing of products, but also promotes effective communication and collaboration among team members.

[0003] The existing methods for drawing parameter recognition in the prior art can be based on rule engines, deep learning-based image recognition methods, and the use of Optical Character Recognition (OCR) technology to parse engineering drawings.

[0004] However, when faced with complex drawings, these methods can usually only process single-modal drawing information, and there are double challenges of accuracy and efficiency when dealing with drawings involving various complex information such as geometric shapes, specifications, tolerances, materials, and technical requirements. Summary of the Invention

[0005] The purpose of this application is to provide a method for fine-tuning and inference of a large model for drawing recognition to solve the problems of poor accuracy and low efficiency in drawing recognition in the prior art in view of the above deficiencies in the prior art.

[0006] To achieve the above purpose, the technical solutions adopted in the embodiments of this application are as follows:

[0007] In a first aspect, an embodiment of this application provides a method for fine-tuning and inference of a large model for drawing recognition, and the method includes:

[0008] Obtain original sample data, annotation information, and a task statement. The original sample data includes: a plurality of original engineering drawing images, and the task statement is used to indicate the task type of at least one task to be recognized in the original sample data;

[0009] According to the original sample data and the annotation information, perform augmentation processing on the original sample data to generate the training sample data;

[0010] In the current training round, input the training sample data into the image processing module in the initial drawing recognition model to obtain a sequence of images to be processed, and input the task statement into the text processing module in the initial drawing recognition model to obtain a sequence of texts to be processed. The initial drawing recognition model includes: an image processing module, a text processing module, a feature processing module, and a multi-modal recognition module;

[0011] The feature processing module adjusts the parameters of the feature processing module based on a preset first enhancement factor and the loss result of the previous training round, and the feature processing module processes the to-be-processed image sequence and the to-be-processed text sequence to obtain a feature processing result;

[0012] The multimodal recognition module adjusts the parameters of the multimodal recognition module based on a preset second enhancement factor, a preset third enhancement factor, and the loss result of the previous training round, and the multimodal recognition module performs multimodal recognition on the feature processing result to obtain a recognition result;

[0013] Based on the recognition result, the true value, and a preset loss function, a loss result is determined, and the initial drawing recognition model is iteratively fine-tuned according to the loss result to obtain a target drawing recognition model, and the target drawing recognition model is used for inference.

[0014] In a second aspect, another embodiment of the present application provides a storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the method according to any one of the above first aspects.

[0015] The beneficial effects of the present application are as follows: By obtaining original sample data, annotation information, and task statements, and performing augmentation processing on the original sample data according to the original sample data and the annotation information to generate training sample data, the diversity of the training sample data can be improved. In the current training round, the training sample data is input into the image processing module in the initial drawing recognition model to obtain a to-be-processed image sequence, and the task statement is input into the text processing module in the initial drawing recognition model to obtain a to-be-processed text sequence. The feature processing module adjusts the parameters of the feature processing module based on a preset first enhancement factor and the loss result of the previous training round, and the feature processing module processes the to-be-processed image sequence and the to-be-processed text sequence to obtain a feature processing result. The multimodal recognition module adjusts the parameters of the multimodal recognition module based on a preset second enhancement factor, a preset third enhancement factor, and the loss result of the previous training round, and the multimodal recognition module performs multimodal recognition on the feature processing result to obtain a recognition result. Thus, based on the recognition result, the true value, and a preset loss function, a loss result can be determined, and the initial drawing recognition model is iteratively fine-tuned according to the loss result to obtain a target drawing recognition model, and the target drawing recognition model is used for inference. This can improve the generalization ability and robustness of the target drawing recognition model, enabling the target drawing recognition model to accurately locate and understand the information required by the user in complex engineering drawing images.

[0016] At the same time, by introducing the adapter, cross-task or cross-domain transfer learning can be achieved while keeping most of the structures of the feature processing module and the multimodal recognition module unchanged, so that the recognition effect of the target drawing recognition model is better. It also reduces the complexity and computational complexity of the feature processing module and the multimodal recognition module without sacrificing too much performance, thereby improving training efficiency.

[0017] In addition, by controlling the enhancement amplitude of the middle layer of the feature processing module and the multimodal recognition module through the enhancement factor, the output of the middle layer of the feature processing module and the multimodal recognition module can be adjusted more flexibly, thereby generating more diverse feature representations, which helps the feature processing module and the multimodal recognition module to learn a wider feature space during the training process, thereby improving the generalization ability of the model, and preventing overfitting during the training process, thereby improving the robustness of the feature processing module and the multimodal recognition module. It can also enable the feature processing module and the multimodal recognition module to better process sample data carrying noise, thereby improving the anti-interference ability of the feature processing module and the multimodal recognition module. Moreover, by introducing the enhancement factor, the training effect of the feature processing module and the multimodal recognition module can be optimized by adjusting the value of the enhancement factor, which helps to steadily improve the performance of the feature processing module and the multimodal recognition module. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 A schematic flow chart of a method for fine-tuning and reasoning a large model for drawing recognition provided in an embodiment of the present application;

[0020] Figure 2 A schematic diagram of the structure of the initial drawing recognition model provided in an embodiment of the present application;

[0021] Figure 3 A schematic diagram of a process for generating training sample data in the fine-tuning and reasoning method of the large model for drawing recognition provided in an embodiment of the present application;

[0022] Figure 4 A schematic diagram of a process for obtaining target sample data in the fine-tuning and reasoning method of the drawing recognition large model provided in an embodiment of the present application;

[0023] Figure 5It is a schematic flowchart when obtaining training sample data in the fine-tuning and inference method of the drawing recognition large model provided by the embodiment of the present application;

[0024] Figure 6 It is a schematic structural diagram of the feature processing module in the initial drawing recognition model provided by the embodiment of the present application;

[0025] Figure 7 It is a schematic flowchart when adjusting the parameters of the feature processing module in the fine-tuning and inference method of the drawing recognition large model provided by the embodiment of the present application;

[0026] Figure 8 It is a schematic flowchart when determining the first parameter adjustment information of each first fully connected layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiment of the present application;

[0027] Figure 9 It is a schematic diagram when determining the first parameter adjustment information of the first fully connected layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiment of the present application;

[0028] Figure 10 It is a schematic structural diagram of the multi-modal recognition module in the initial drawing recognition model provided by the embodiment of the present application;

[0029] Figure 11 It is a schematic flowchart when adjusting the parameters of the multi-modal recognition module in the fine-tuning and inference method of the drawing recognition large model provided by the embodiment of the present application;

[0030] Figure 12 It is a schematic flowchart when determining the second parameter adjustment information of each self-attention layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiment of the present application;

[0031] Figure 13 It is a schematic diagram when determining the query parameter adjustment information of the self-attention layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiment of the present application;

[0032] Figure 14 It is a schematic diagram when determining the key parameter adjustment information of the self-attention layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiment of the present application;

[0033] Figure 15 It is a schematic diagram when determining the value parameter adjustment information of the self-attention layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiment of the present application. Detailed implementation manners

[0034] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. It should be understood that the accompanying drawings in this application only serve the purpose of illustration and description, and are not used to limit the protection scope of this application. Additionally, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of this application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. Moreover, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application.

[0035] In addition, the described embodiments are only some embodiments of this application, not all of the embodiments. The components of the embodiments of this application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of this application claimed, but merely represents the selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts fall within the protection scope of this application.

[0036] It should be noted that the term "including" will be used in the embodiments of this application to indicate the existence of the features stated thereafter, but does not exclude the addition of other features.

[0037] The existing methods for identifying drawing parameters can be rule-based, deep learning-based image recognition methods, and methods for parsing engineering drawings using OCR technology.

[0038] Specifically, the rule-based drawing parsing method identifies drawing parameters through predefined rules. However, when dealing with complex or diverse drawings, it is difficult to adapt to different types of parameter identification, and its scalability is poor. The deep learning-based image recognition method uses deep learning models such as convolutional neural networks (CNNs) to identify geometric shapes, but it is still limited to processing image data and difficult to combine text or symbol information. The method of combining OCR technology with drawing content parsing can extract text information, but it has insufficient accuracy in parsing symbol annotations, tolerance information, and processing techniques, and requires frequent manual intervention.

[0039] In summary, the existing methods for identifying drawing parameters usually can only process single-modal drawing information when facing complex drawings, and there are dual challenges of accuracy and efficiency when dealing with drawings involving various complex information such as geometric shapes, specifications, tolerances, materials, and technical requirements.

[0040] The fine-tuning and inference method of the drawing recognition large model provided by the embodiments of the present application will be described in detail below in combination with multiple embodiments.

[0041] Figure 1 It is a schematic flowchart of a method for fine-tuning and inference of the drawing recognition large model provided by the embodiments of the present application. Referring to Figure 1 as shown, the execution subject of this method may include any electronic device with processing capabilities. This method includes:

[0042] S101. Obtain the original sample data, annotation information, and task statements.

[0043] Among them, the original sample data includes: multiple original engineering drawing images, and the task statement is used to indicate the task type of at least one task to be recognized in the original sample data.

[0044] Optionally, the original engineering drawing images can be industrial drawings common in the actual production process. The data sources of the original engineering drawing images can specifically include: scanned process drawings and construction drawings from the actual factory production line, high-precision CAD export drawings provided by professional design departments or third-party engineering companies, and mixed-format engineering drawings provided by different industry standards and upstream and downstream enterprises in the supply chain. The original engineering drawing images can include die drawings, structural part drawings, assembly drawings, processing process drawings, standard part drawings, etc.

[0045] Optionally, the task type of the task to be recognized is used to indicate the operations performed when recognizing the original engineering drawing images, such as dimensional specifications, tolerances, geometric shapes, and materials, etc.

[0046] Optionally, the annotation information can be obtained through manual annotation, and specifically can include: geometric shape information, such as: contour features of parts, hole positions, surface features, sectional views, etc.; specification dimensions and tolerances, such as: parameters such as length, width, thickness, hole diameter, angle, arc radius, and minimum and maximum tolerance ranges, etc.; material information, such as: types of part materials (steel, aluminum alloy, plastic, etc.), material grades, and corresponding standard codes; technical requirements and notes, such as: specific surface treatment processes, roughness requirements, assembly requirements, heat treatment conditions, execution standards, or special markings, etc.

[0047] Exemplarily, for different original engineering drawing images, the task statements are respectively: extract geometric shape information from the original engineering drawing images, extract specification dimension and tolerance information from the original engineering drawing images, extract material information from the original engineering drawing images, and extract technical requirements and notes from the original engineering drawing images, etc.

[0048] S102. Augment the original sample data according to the original sample data and the annotation information to generate training sample data.

[0049] Optionally, after obtaining the original sample data and annotation information, the original sample data and the annotation information can be superimposed, and the superimposed sample data can be augmented to generate training sample data.

[0050] Exemplarily, the superimposed sample data can be augmented by operations such as mirror flipping, rotation, translation, scaling, cropping, and adding noise to generate training sample data.

[0051] S103. In the current training round, input the training sample data into the image processing module in the initial drawing recognition model to obtain a sequence of images to be processed, and input the task statement into the text processing module in the initial drawing recognition model to obtain a sequence of texts to be processed.

[0052] Optionally, Figure 2 is a schematic structural diagram of the initial drawing recognition model provided by the embodiments of the present application. Refer to Figure 2 As shown, the initial drawing recognition model includes: an image processing module, a text processing module, a feature processing module, and a multi-modal recognition module. Among them, the output ends of the image processing module and the text processing module are both connected to the input end of the feature processing module, the output end of the feature processing module is connected to the input end of the multi-modal recognition module, and the output end of the multi-modal recognition module is used to output the recognition result.

[0053] Optionally, continue to refer to Figure 2 As shown, the image processing module may include a first sub-image processing module and a second sub-image processing module. The first sub-image processing module is used to extract image features from the training sample data, and the second sub-image processing module is used to perform local detail restoration on the extracted image features to obtain a sequence of images to be processed.

[0054] Optionally, continue to refer to Figure 2 As shown, the text processing module is used to perform word segmentation and vector generation processing on the input task statement to generate a sequence of texts to be processed.

[0055] Exemplarily, the first sub-image processing module can be implemented based on a vision-based processing model. For example, the vision-based processing model can be InternViT-6B, the second sub-image processing module can be implemented based on PixelUnshuffle, and the text processing module can be implemented based on a text tokenizer.

[0056] S104. The feature processing module adjusts the parameters of the feature processing module based on a preset first enhancement factor and the loss result of the previous training round, and the feature processing module performs feature processing on the image sequence to be processed and the text sequence to be processed to obtain a feature processing result.

[0057] Optionally, the feature processing module may include at least one first adapter. After receiving the image sequence to be processed and the text sequence to be processed, the feature processing module can adjust the parameters of the feature processing module based on a preset first enhancement factor and the loss result of the previous training round, so that the adjusted feature processing module can perform feature processing such as linear transformation and activation on the image sequence to be processed and the text sequence to be processed to obtain a feature processing result.

[0058] Among them, the parameters of the feature processing module at least include: the parameters of the first adapter.

[0059] Exemplarily, the feature processing module can be implemented based on a multi-layer perceptron (MLP for short), and a first adapter is introduced into the multi-layer perceptron.

[0060] Exemplarily, after receiving the image sequence to be processed and the text sequence to be processed, the feature processing module can adjust the parameters of the first adapter in the feature processing module and the other parameters in the feature processing module except the first adapter based on a preset first enhancement factor and the loss result of the previous training round, so that the adjusted feature processing module can perform feature processing such as linear transformation and activation on the image sequence to be processed and the text sequence to be processed to obtain a feature processing result. Among them, the first enhancement factor can be used to control the enhancement amplitude of the intermediate layer in the feature processing module.

[0061] Exemplarily, the first adapter can be a residual adapter network (Residual adapter_1s), which contains a small number of parameters and is connected to the original network layer in the feature processing module through a residual connection to achieve fine-tuning of the network output.

[0062] S105. The multi-modal recognition module adjusts the parameters of the multi-modal recognition module based on a preset second enhancement factor, a preset third enhancement factor, and the loss result of the previous training round, and the multi-modal recognition module performs multi-modal recognition on the feature processing result to obtain a recognition result.

[0063] Optionally, at least one second adapter may be included in the multi-modal recognition module. After the multi-modal recognition module receives the feature processing result output by the feature processing module, it may adjust the parameters of the multi-modal recognition module based on a preset second enhancement factor, a preset third enhancement factor, and the loss result of the previous training round, so that the adjusted multi-modal recognition module can perform multi-modal recognition on the feature recognition result to obtain the recognition result of the current training round.

[0064] Among them, the parameters of the multi-modal recognition module at least include: the parameters of the second adapter. Both the preset second enhancement factor and the preset third enhancement factor are used to control the enhancement amplitude of the intermediate layer in the multi-modal recognition module. Specifically, the second enhancement factor and the third enhancement factor are used to control the enhancement amplitude of different intermediate layers.

[0065] Exemplarily, the multi-modal recognition module may be implemented based on the Transformer architecture, and a second adapter network is introduced into the Transformer architecture.

[0066] Exemplarily, after the multi-modal recognition module receives the feature processing result output by the feature processing module, it may adjust the parameters of the second adapter in the multi-modal recognition module and other parameters in the multi-modal recognition module except the second adapter based on a preset second enhancement factor, a preset third enhancement factor, and the loss result of the previous training round, so that the adjusted multi-modal recognition module can perform multi-modal recognition on the feature recognition result to obtain the recognition result of the current training round.

[0067] Exemplarily, the second enhancement factor and the third enhancement factor are used to control the enhancement amplitude of different intermediate layers in the Transformer architecture.

[0068] S106. Determine the loss result according to the recognition result, the ground truth, and a preset loss function, and iteratively fine-tune the initial drawing recognition model according to the loss result to obtain a target drawing recognition model, and use the target drawing recognition model for inference.

[0069] Optionally, after obtaining the recognition result of the current round, the loss result of the current round may be calculated according to the recognition result of the current round, the ground truth, and a preset loss function.

[0070] Optionally, compare the loss result of the current round with a preset iteration end condition. If the loss result of the current round meets the preset iteration end condition, end the training, and use the initial drawing recognition model of the current round as the target drawing recognition model, so that the target drawing recognition model can be applied for inference.

[0071] Optionally, if the loss result of the current round does not meet the preset iteration end condition, perform the steps of S104 - S105 above, and fine-tune the initial drawing recognition model according to the loss result of the current round for the next round of training.

[0072] In this embodiment, by obtaining the original sample data, annotation information, and task statements, and augmenting the original sample data according to the original sample data and annotation information to generate training sample data, the diversity of the training sample data can be improved. In the current training round, the training sample data is input into the image processing module in the initial drawing recognition model to obtain the image sequence to be processed, and the task statement is input into the text processing module in the initial drawing recognition model to obtain the text sequence to be processed. The feature processing module adjusts the parameters of the feature processing module based on the preset first enhancement factor and the loss result of the previous training round, and the feature processing module performs feature processing on the image sequence to be processed and the text sequence to be processed to obtain the feature processing result. The multi-modal recognition module adjusts the parameters of the multi-modal recognition module based on the preset second enhancement factor, preset third enhancement factor, and the loss result of the previous training round, and the multi-modal recognition module performs multi-modal recognition on the feature processing result to obtain the recognition result. Thus, the loss result can be determined according to the recognition result, the true value, and the preset loss function, and the initial drawing recognition model can be iteratively fine-tuned according to the loss result to obtain the target drawing recognition model, and the target drawing recognition model is used for inference, which can improve the generalization ability and robustness of the target drawing recognition model, so that the target drawing recognition model can accurately locate and understand the information required by the user in complex engineering drawing images.

[0073] At the same time, by introducing the adapter, cross-task or cross-domain transfer learning can be realized while keeping most of the structures of the feature processing module and the multi-modal recognition module unchanged, so that the recognition effect of the obtained target drawing recognition model is better. It also reduces the complexity and computational amount of the feature processing module and the multi-modal recognition module without sacrificing too much performance, and improves the training efficiency.

[0074] In addition, by controlling the enhancement amplitude of the middle layer of the feature processing module and the multimodal recognition module through the enhancement factor, the output of the middle layer of the feature processing module and the multimodal recognition module can be adjusted more flexibly, thereby generating more diverse feature representations, which helps the feature processing module and the multimodal recognition module to learn a wider feature space during the training process, thereby improving the generalization ability of the model, and preventing overfitting during the training process, thereby improving the robustness of the feature processing module and the multimodal recognition module. It can also enable the feature processing module and the multimodal recognition module to better process sample data carrying noise, thereby improving the anti-interference ability of the feature processing module and the multimodal recognition module. Moreover, by introducing the enhancement factor, the training effect of the feature processing module and the multimodal recognition module can be optimized by adjusting the value of the enhancement factor, which helps to steadily improve the performance of the feature processing module and the multimodal recognition module.

[0075] In one possible implementation, Figure 3 A flow chart of generating training sample data in the fine-tuning and reasoning method of the drawing recognition large model provided in the embodiment of the present application, referring to Figure 3 As shown, the above S102 performs augmentation processing on the original sample data according to the original sample data and the annotation information to generate training sample data, including:

[0076] S301: Add annotation information to the corresponding position in the original sample data to obtain annotated sample data.

[0077] Optionally, the annotation information is added to the corresponding position in the original sample data to obtain the annotated sample data.

[0078] S302 : Based on a preset rotation matrix, a preset scaling matrix, and a preset cropping matrix, the labeled sample data is randomly rotated, randomly scaled, and randomly cropped in sequence to obtain randomly processed labeled sample data.

[0079] Optionally, after obtaining the labeled sample data, the labeled sample data may be randomly rotated based on a preset rotation matrix to obtain the randomly rotated labeled sample data.

[0080] Optionally, after obtaining the labeled sample data, the labeled sample data may be randomly scaled based on a preset scaling matrix to obtain the randomly scaled labeled sample data.

[0081] Optionally, after obtaining the labeled sample data, the labeled sample data may be randomly cropped based on a preset cropping matrix to obtain the randomly cropped labeled sample data.

[0082] Optionally, the annotation sample data after random rotation, the annotation sample data after random scaling, and the annotation sample data after random cropping are all used as the annotation sample data after random processing.

[0083] Optionally, the annotation sample data after random rotation, the annotation sample data after random scaling, and the annotation sample data after random cropping can also be randomly combined to obtain the annotation sample data after random processing.

[0084] Exemplarily, assuming that the original coordinates of the image in the annotation sample data are (x, y), and the rotation angle is θ (random selection range is [-θ_max, θ_max]), the coordinates (x', y') of the image in the annotation sample data after random rotation can be expressed by the following formula:

[0085] x' = cos(θ) * x - sin(θ) * y

[0086] y' = sin(θ) * x + cos(θ) * y

[0087] Among them, θ is the rotation angle, and the random selection range is [-θ_max, θ_max], which can be selected through the random selection function Uniform(), specifically θ = Uniform(-θ_max, θ_max).

[0088] Exemplarily, assuming that the original coordinates of the image in the annotation sample data are (x, y), and the image scaling factor is s (random selection range is [s_min, s_max]), the coordinates (x', y') of the image in the annotation sample data after random scaling can be expressed by the following formula:

[0089] x' = s * x

[0090] y' = s * y

[0091] Among them, s is the scaling factor, and the random selection range is [s_min, s_max], which can be selected through the random selection function Uniform(), specifically s = Uniform(s_min, s_max).

[0092] Exemplarily, assuming that the original area of the image in the annotation sample data is W x H, a sub-rectangle area W_c x H_c is cropped from the original area W x H of the image in the annotation sample data according to a random offset, and the starting point of the cropping is represented by the offset (Δx, Δy) as follows:

[0093] Δx = Uniform(0, W - W_c)

[0094] Δy = Uniform(0, H - H_c)

[0095] The cropped image region is (x, y) ∈ [Δx, Δx + W_c] × [Δy, Δy + H_c]

[0096] Where W and H are the original width and height of the image in the labeled sample data, W_c and H_c are the cropped width and height of the image in the labeled sample data, and Δx and Δy are the cropping offsets.

[0097] S303. Perform consistency processing on the annotation information in the randomly processed labeled sample data to obtain target sample data.

[0098] Optionally, after obtaining the randomly processed labeled sample data, the randomly processed labeled sample data can be checked for consistency to obtain target sample data, so as to ensure the consistency between the randomly processed labeled sample data and the annotation information.

[0099] S304. Adjust the image quality of the target sample data to obtain training sample data.

[0100] Optionally, after performing data augmentation and enhancement processing on the target sample data, the target sample data is augmented to obtain training sample data.

[0101] By adding the annotation information to the corresponding position in the original sample data, the labeled sample data is obtained, and based on a preset rotation matrix, a preset scaling matrix, and a preset cropping matrix, the labeled sample data is sequentially subjected to random rotation, random scaling, and random cropping processing to obtain the randomly processed labeled sample data, and the annotation information in the randomly processed labeled sample data is subjected to consistency processing to obtain target sample data, and then the image quality of the target sample data is adjusted to obtain training sample data, which can increase the diversity of the training sample data, help the model learn richer features, reduce the risk of overfitting, and improve the generalization performance of the model. Moreover, it can also save the data collection cost and improve the applicability and flexibility of the model.

[0102] In a possible implementation manner Figure 4 is a schematic flowchart of a process for obtaining target sample data in the fine-tuning and inference method of the drawing recognition large model provided by the embodiment of the present application. Refer to Figure 4 as shown, the above S303 performs consistency processing on the annotation information in the randomly processed labeled sample data to obtain target sample data, including:

[0103] S401. Determine whether the first annotation text in the annotation information of the randomly processed annotation sample data is complete.

[0104] Optionally, traverse each annotation text in the annotation information of the randomly processed annotation sample data. For the currently traversed annotation text, that is, the first annotation text, determine whether the first annotation text is complete. Here, the first annotation text is any one of the annotation texts in the annotation information.

[0105] Optionally, the position of the text box of the first annotation text can be compared with the area of the image in the annotation sample data to determine whether the first annotation text is complete.

[0106] S402. If so, retain the first annotation text in the randomly processed annotation sample data; if not, delete the first annotation text in the randomly processed annotation sample data.

[0107] Optionally, if the first annotation text is complete, retain the first annotation text; if the first annotation text is incomplete, delete the first annotation text.

[0108] Exemplarily, it can be determined whether to delete by judging whether the first annotation text is completely located within the image area of the randomly processed annotation sample data.

[0109] Exemplarily, assume that the image area of the randomly processed annotation sample data is [x_min, x_max] × [y_min, y_max], and the bounding box of the first annotation text is [x1, x2] × [y1, y2]. The condition for deleting the annotation is: x1 >= x_max or x2 <= x_min or y1 >= y_max or y2 <= y_min. If this condition is met, delete the first annotation text.

[0110] Here, [x_min, x_max] is the range of the image area of the randomly processed annotation sample data on the x-axis, [y_min, y_max] is the range of the image area of the randomly processed annotation sample data on the y-axis, [x1, x2] is the range of the bounding box of the first annotation text on the x-axis, and [y1, y2] is the range of the bounding box of the first annotation text on the y-axis.

[0111] Optionally, if the above condition is not met, the first annotation text can be deleted.

[0112] Optionally, if the above conditions are not met, the first annotation text can also be repositioned based on the position of the first annotation text and the randomly processed annotation sample data to obtain the target position of the first annotation text, and in the randomly processed annotation sample data, update the position of the first annotation text to the target position.

[0113] Exemplarily, when it is necessary to retain the incomplete first annotation text, the bounding box of the first annotation text can be repositioned to adapt to the randomly processed annotation sample data.

[0114] Exemplarily, taking the random processing as only random cropping as an example, assume that the cropped area of the image in the annotation sample data is [x_min, x_max]×[y_min, y_max], the bounding box of the first annotation text before cropping is [x1, x2]×[y1,y2], and the new bounding box after cropping is [x1', x2']×[y1', y2']. Then the area of the first annotation text after cropping can be represented by the following formula:

[0115] x1' = max(x1, x_min) x2' = min(x2, x_max) y1' = max(y1, y_min) y2' =min(y2, y_max)

[0116] Where, [x1', x2'] is the new range of the first annotation text after cropping on the x-axis, [y1', y2'] is the new range of the first annotation text after cropping on the y-axis, max(a, b) is to take the larger value of a and b, and min(a, b) is to take the smaller value of a and b.

[0117] Exemplarily, and in the randomly processed annotation sample data, update the position of the bounding box of the first annotation text to the target position [x1', x2']×[y1', y2'].

[0118] Optionally, after all the first annotation texts are updated, the annotation sample data can also be checked for consistency to obtain the target sample data.

[0119] Exemplarily, taking the random processing as only random cropping as an example, if the area of the new bounding box [x1', x2']×[y1', y2'] of the first annotation text after cropping is 0, then delete the first annotation text, so as to ensure the consistency between the annotation information and the image in the target sample data.

[0120] S403. After all the randomly processed annotation sample data are processed, the target sample data is obtained.

[0121] Optionally, after all the labeled sample data after random processing have been executed according to the steps of S401 - S402 above, that is, after the consistency processing is completed, the target sample data can be obtained.

[0122] In a possible implementation, Figure 5 is a schematic flowchart of a process for obtaining training sample data in the method for fine - tuning and inference of the drawing recognition large model provided by the embodiments of the present application. Refer to Figure 5 as shown, the above - mentioned S304 performs image quality adjustment on the target sample data to obtain training sample data, including:

[0123] S501. Perform illumination adjustment on the target sample data to obtain the target sample data with adjusted pixels.

[0124] Optionally, the pixel values of each sample data in the target sample data can be adjusted to achieve illumination adjustment, so as to obtain the target sample data with adjusted pixels.

[0125] Exemplarily, assume that the original pixel value of a certain sample data is I(x, y), and the adjusted pixel value is I'(x, y), then the adjusted pixel value can be expressed by the following formula:

[0126] I'(x, y) = α * I(x, y) + β

[0127] Wherein, I(x, y) is the original pixel value of a certain sample data at the position (x, y), I'(x, y) is the pixel value of a certain sample data at the position (x, y) after illumination change, α is the brightness scaling coefficient for controlling the overall brightness change, and β is the brightness offset for simulating different illumination conditions.

[0128] S502. Perform contrast adjustment on the target sample data to obtain the target sample data with adjusted contrast.

[0129] Optionally, the pixel values of each sample data in the target sample data can be stretched or compressed through linear transformation to achieve contrast adjustment, so as to obtain the target sample data with adjusted contrast.

[0130] Exemplarily, assume that the original pixel value of a certain sample data is I(x, y), and the adjusted pixel value is I'(x, y), then the adjusted pixel value can be expressed by the following formula:

[0131] I'(x, y) = (I(x, y) - μ) * c + μ

[0132] Among them, I(x, y) is the original pixel value of a certain sample data at the position (x, y), I'(x, y) is the pixel value of a certain sample data at the position (x, y) after contrast adjustment, μ is the average pixel value of a certain sample data, defined as μ = (1 / N) * Σ I(x, y), where N is the total number of pixels of the sample data, and c is the contrast adjustment coefficient. c > 1 represents enhanced contrast, and 0 < c < 1 represents reduced contrast.

[0133] S503. Perform blurring adjustment on the target sample data to obtain the target sample data after blurring adjustment.

[0134] Optionally, image blurring adjustment can be performed on each sample data in the target sample data through a convolution operation to obtain the target sample data after blurring adjustment.

[0135] Exemplarily, assuming that the original pixel value of a certain sample data is I(x, y) and the adjusted pixel value is I'(x, y), the adjusted pixel value can be expressed by the following formula:

[0136] I'(x, y) = ΣI(x - u, y - v) * K(u, v)

[0137] Among them, the range of u and v is from -k to k, I(x, y) is the original pixel value of a certain sample data at the position (x, y), I'(x, y) is the pixel value of a certain sample data at the position (x, y) after blurring adjustment, K(u, v) is the convolution kernel used to define the degree of blurring, such as a Gaussian blurring kernel, and k is the radius of the convolution kernel to control the blurring range.

[0138] S504. Generate training sample data according to the target sample data after pixel adjustment, the target sample data after contrast adjustment, and the target sample data after blurring adjustment.

[0139] Optionally, after obtaining the target sample data after pixel adjustment, the target sample data after contrast adjustment, and the target sample data after blurring adjustment, the target sample data after pixel adjustment, the target sample data after contrast adjustment, and the target sample data after blurring adjustment can be used as training sample data.

[0140] By performing illumination adjustment on the target sample data, the target sample data after pixel adjustment is obtained. By performing contrast adjustment on the target sample data, the target sample data after contrast adjustment is obtained. By performing blur adjustment on the target sample data, the target sample data after blur adjustment is obtained. Thus, the training sample data can be generated based on the target sample data after pixel adjustment, the target sample data after contrast adjustment, and the target sample data after blur adjustment, so that the obtained training sample data can cover drawings with various resolutions, orientations, and clarity levels as inputs, enabling the obtained target drawing recognition model to have stronger robustness and the ability to process diverse inputs when facing complex real application scenarios.

[0141] It can be understood that the above has given an exemplary description of the process of obtaining the training sample data in the fine-tuning and inference method of the drawing recognition large model provided in this application embodiment. The following gives an exemplary description of the process of training the initial drawing recognition model based on the training sample data.

[0142] It is worth noting that during the training process of the initial drawing recognition model, the initial drawing recognition model can be trained in a fine-tuning manner. However, when using general fine-tuning methods, such as low-rank adaptation methods like LORA (Low-Rank Adaptation), the weight updates of different layers show obvious skewness. That is, the bottom layer (such as the embedding layer) and the top layer (such as the head layer of the language model) will occupy most of the weight updates during fine-tuning, while the weight updates of the middle layer are relatively small. This causes performance deviations in the initial drawing recognition model after fine-tuning because the middle layer is equally important for the overall performance of the model.

[0143] To narrow the gap in weight updates between the middle layer and the bottom and top layers during fine-tuning, enhance the importance of the middle layer during the fine-tuning process, and narrow the performance gap with full-scale fine-tuning, the following uses multiple embodiments to elaborate in detail on the structure and training process of the initial drawing recognition model in the fine-tuning and inference method of the drawing recognition large model provided in this application.

[0144] In one possible implementation, Figure 6 is a schematic structural diagram of the feature processing module in the initial drawing recognition model provided in the embodiment of this application. Referring to Figure 6 as shown, the feature processing module in the initial drawing recognition model includes: a splicing layer and at least one perception module connected in sequence; the perception module includes: a first fully connected layer and an activation layer; a first adapter is respectively deployed in each first fully connected layer of each perception module, and each first fully connected layer has a weight matrix.

[0145] Optionally, the splicing layer is used to splice the image sequence to be processed and the text sequence to be processed of the input feature processing module, and input the spliced feature sequence into each perception module for feature processing.

[0146] Optionally, the feature processing module includes multiple perception modules, which are connected in sequence. A perception module can be a layer in a Multi-Layer Perceptron (MLP for short). Each perception module includes a first fully connected layer and a first adapter.

[0147] Among them, the first adapter can be a Residual adapters, which contains a small number of parameters and is connected to the first fully connected layer in the feature processing module through a residual connection to fine-tune the output of the perception module.

[0148] In a possible implementation, Figure 7 is a schematic flowchart of a process for adjusting the parameters of the feature processing module in the method for fine-tuning and inferring the drawing recognition large model provided by the embodiments of this application. Refer to Figure 6 and Figure 7 As shown, in step S104 above, based on the preset first enhancement factor and the loss result of the previous training round, adjusting the parameters of the feature processing module includes:

[0149] S701. Determine the first parameter adjustment information of each first fully connected layer according to the first enhancement factor, the loss result of the previous training round, the weight matrix of each first fully connected layer in the previous training round, and the weight matrix of the first adapter of each first fully connected layer in the previous training round.

[0150] It can be understood that the first adapters in each first fully connected layer are connected to each first fully connected layer through a residual connection. When adjusting each first fully connected layer, that is, adjusting the weight matrix of each first fully connected layer, then the weight matrix of the first adapter of each first fully connected layer in the previous round can be adjusted, and the weight matrix of each first fully connected layer in the previous training round can be adjusted, so as to realize the adjustment of each first fully connected layer, so that the input data can be processed through the adjusted weight matrix of each first fully connected layer.

[0151] Optionally, when adjusting the parameters of the feature processing module, the first adapter parameter adjustment information corresponding to the first adapter of each first fully connected layer can be obtained according to the loss result of the previous training round. On this basis, the first parameter adjustment information of each first fully connected layer is obtained according to the first enhancement factor, the loss result of the previous training round, and the first adapter parameter adjustment information.

[0152] Optionally, when adjusting the parameters of the feature processing module, the first adapter parameter adjustment information corresponding to the first adapter of each first fully connected layer can be obtained according to the first enhancement factor and the loss result of the previous training round. On this basis, according to the loss result of the previous training round and the first adapter parameter adjustment information, the first parameter adjustment information of each first fully connected layer is obtained.

[0153] Among them, the first parameter adjustment information is used to indicate the adjustment information of the weight matrix of the first fully connected layer and each parameter in the first adapter in the first fully connected layer.

[0154] S702. According to each first parameter adjustment information, adjust the weight matrix of each first fully connected layer to obtain the weight matrix of each first fully connected layer in the current training round.

[0155] Optionally, after obtaining the first parameter adjustment information, the weight matrix of each first fully connected layer can be adjusted respectively according to the first parameter adjustment information of each first fully connected layer, so as to obtain the weight matrix of each first fully connected layer in the current training round, so that the input data can be processed by the weight matrix of each first fully connected layer in the current training round to obtain the output of each first fully connected layer.

[0156] Determine the first parameter adjustment information of each first fully connected layer through the first enhancement factor, the loss result of the previous training round, the weight matrix of each first fully connected layer in the previous training round, and the weight matrix of the first adapter of each first fully connected layer in the previous training round, and adjust the weight matrix of each first fully connected layer through each first parameter adjustment information to obtain the weight matrix of each first fully connected layer in the current training round, so as to be able to finely adjust the weight matrix of each first fully connected layer, thereby being able to more accurately locate the parameters to be adjusted, narrowing the weight update gap between the middle layer and the bottom layer and the top layer during fine-tuning, enhancing the importance of the middle layer during the fine-tuning process, and narrowing the performance gap with full-scale fine-tuning. At the same time, unnecessary computational overhead is avoided, computing resources are utilized more efficiently, and overfitting of each first fully connected layer to the input data can also be reduced, improving the generalization ability of each first fully connected layer on unseen data.

[0157] In a possible implementation manner, Figure 8 is a schematic flowchart of a process for determining the first parameter adjustment information of each first fully connected layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiments of the present application. Refer to Figure 8As shown, the above S701 determines the first parameter adjustment information of each first fully connected layer according to the first enhancement factor, the loss result of the previous training round, the weight matrix of each first fully connected layer in the previous training round, and the weight matrix of the first adapter of each first fully connected layer in the previous training round, including:

[0158] S801. Determine the first low-rank transformation matrix and the second low-rank transformation matrix according to the image dimension in the training sample data.

[0159] Optionally, the first low-rank transformation matrix and the second low-rank transformation matrix can be obtained by decomposing according to the image dimension in the training sample data through a low-rank decomposition method.

[0160] Among them, the dimension after multiplying the first low-rank transformation matrix and the second low-rank transformation matrix matches the image dimension. The low-rank decomposition method can include methods such as singular value decomposition (SVD) and non-negative matrix factorization (NMF).

[0161] Exemplarily, in the initial situation, the image dimension in the training sample data can be decomposed using a pre-configured target rank to obtain the first low-rank transformation matrix and the second low-rank transformation matrix, and during the training process, the target rank can be optimized through methods such as singular value decomposition (SVD) and non-negative matrix factorization (NMF) to achieve the optimization of the first low-rank transformation matrix and the second low-rank transformation matrix.

[0162] Exemplarily, assume that the image dimension in the training sample data is , where is the height, is the width, the first low-rank transformation matrix U1 can be obtained as and the second low-rank transformation matrix V1^T as , where is the current target rank.

[0163] S802. Determine the weight matrix of the first adapter in the current training round according to the loss result of the previous training round and the weight matrix of the first adapter of the first fully connected layer in the previous training round.

[0164] Optionally, taking the Lth first fully connected layer as an example, Figure 9 This is a schematic diagram when determining the first parameter adjustment information of the first fully connected layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiments of the present application. Refer to Figure 9As shown, x1_L is the input of the L-th first fully-connected layer, W1_L is the original trainable weight of the L-th first fully-connected layer, U1_L is the first low-rank transformation matrix of the L-th first fully-connected layer, V1^T_L is the second low-rank transformation matrix of the L-th first fully-connected layer. The first low-rank transformation matrix U1_L can be initialized to 0, and the second low-rank transformation matrix V1^T_L can be initialized to N1_L(0, σ²), that is, randomly initialized from a normal distribution with a mean of 0 and a variance of σ². adapter1_L is the first adapter of the L-th first fully-connected layer, and h1_L is a differentiable function in the L-th first fully-connected layer, which is used to implement regularization or transformation to help the model better adapt to new tasks or datasets. α1 is the first enhancement factor.

[0165] Among them, the first enhancement factor α1 is a parameter used to control the enhancement amplitude. For example, when α1 = 1.5, the enhancement factor of the intermediate layer is increased to 1.5 times.

[0166] Optionally, continue to refer to Figure 9 As shown, the weight matrix of the first adapter in the current training round can be determined according to the loss result of the previous training round and the weight matrix of the first adapter of the first fully-connected layer in the previous training round.

[0167] Exemplarily, the gradient can be calculated according to the loss result of the previous training round and the weight matrix of the first adapter of the first fully-connected layer in the previous training round, so as to adjust the weight matrices W1_1 and W1_2 of the first adapter through the backpropagation algorithm.

[0168] S803. Determine the first adapter in the current training round according to the weight matrix of the first adapter in the current training round and the second low-rank transformation matrix.

[0169] Optionally, taking the L-th first fully-connected layer as an example, the first adapter in the current training round can be calculated according to the weight matrices W1_1 and W1_2 of the first adapter in the current training round and the second low-rank transformation matrix V1_L^T of the L-th first fully-connected layer.

[0170] Exemplarily, the first adapter adapter1_L(V1_L^T) of the L-th first fully-connected layer can be expressed by the following formula:

[0171] adapter1_L(V1_L^T) = σ(W1_2 * ReLU1(W1_1 * V1_L^T))

[0172] Exemplarily, after obtaining the weight matrices W1_1 and W1_2 of the first adapter in the current training round, W1_1, W1_2, and the second low-rank transformation matrix V1_L^T of the L-th first fully connected layer can be input into the above formula to obtain the first adapter adapter1_L(V1_L^T) of the L-th first fully connected layer in the current training round.

[0173] Among them, V1_L^T is the second low-rank transformation matrix of the L-th first fully connected layer, W1_1 and W1_2 are the weight matrices of the first adapter, ReLU1 is the rectified linear unit activation function corresponding to the first fully connected layer, defined as ReLU1(x) = max(0, x), and σ is the activation function (such as Sigmoid or other non-linear functions).

[0174] S804. Determine the first parameter adjustment information of the first fully connected layer according to the first enhancement factor, the first adapter in the current training round, and the first low-rank transformation matrix.

[0175] Exemplarily, taking the L-th first fully connected layer as an example, after obtaining the first adapter adapter1_L(V1_L^T) in the current training round, the first parameter adjustment information ΔW1_L of the L-th first fully connected layer can be obtained according to the first enhancement factor α1, the first adapter adapter1_L(V1_L^T) in the current training round, and the first low-rank transformation matrix U1_L.

[0176] Exemplarily, the first parameter adjustment information ΔW1_L of the L-th first fully connected layer can be expressed by the following formula:

[0177] ΔW1_L = α1 * U1_L * adapter_1(V1_L^T)

[0178] Exemplarily, taking the L-th first fully connected layer as an example, after obtaining the first parameter adjustment information ΔW1_L of the L-th first fully connected layer, the first parameter adjustment information ΔW1_L can be superimposed on the weight matrix W1_L of the L-th first fully connected layer to obtain the weight matrix W1'_L of the L-th first fully connected layer in the current training round, where W1'_L = W1_L + ΔW1_L.

[0179] Determine the first low-rank transformation matrix and the second low-rank transformation matrix based on the image dimensions in the training sample data, and determine the weight matrix of the first adapter in the current training round according to the loss result of the previous training round and the weight matrix of the first adapter in the first fully-connected layer in the previous training round. Then, determine the first adapter in the current training round according to the weight matrix of the first adapter in the current training round and the second low-rank transformation matrix, so that the first parameter adjustment information of the first fully-connected layer can be determined according to the first enhancement factor, the first adapter in the current training round, and the first low-rank transformation matrix, enabling the first fully-connected layer to be continuously updated by fine-tuning during the training process. Moreover, when the first fully-connected layer is updated, it can update the intermediate layer more evenly, which helps reduce the performance deviation during the fine-tuning process, improve the generalization ability and overall performance of the obtained target drawing recognition model. Additionally, it can also enable the initial drawing recognition model to converge faster, thus obtaining the target drawing recognition model more efficiently.

[0180] In a possible implementation, Figure 10 is a schematic structural diagram of the multi-modal recognition module in the initial drawing recognition model provided by the embodiments of this application. Refer to Figure 10 as shown, the multi-modal recognition module includes: a plurality of sequentially connected feature recognition modules, and each feature recognition module includes: a self-attention layer and a second fully-connected layer connected in sequence; a plurality of second adapters are respectively deployed in each self-attention layer, and a third adapter is respectively deployed in each second fully-connected layer.

[0181] Among them, the second adapter can be a Residual adapters, which contains a small number of parameters and is connected to the self-attention layer through a residual connection to achieve fine-tuning of the output of the self-attention layer.

[0182] Among them, the third adapter can be a Residual adapters, which contains a small number of parameters and is connected to each second fully-connected layer through a residual connection to achieve fine-tuning of the output of each second fully-connected layer.

[0183] In a possible implementation, Figure 11 is a schematic flowchart of a process for adjusting the parameters of the multi-modal recognition module in the method for fine-tuning and inference of the large drawing recognition model provided by the embodiments of this application. Refer to Figure 10 and Figure 11 as shown, the above S105 adjusts the parameters of the multi-modal recognition module based on a preset second enhancement factor, a preset third enhancement factor, and the loss result of the previous training round, including:

[0184] S1101. Determine the second parameter adjustment information of each self-attention layer according to the second enhancement factor, the loss result of the previous training round, the weight matrix of each self-attention layer in the previous round, and the weight matrix of each second adapter of each self-attention layer in the previous training round.

[0185] It can be understood that each second adapter in each self-attention layer is connected based on the self-attention layer through a residual connection. When adjusting each self-attention layer, that is, adjusting the weight matrix of each self-attention layer, then the weight matrix of each second adapter of each self-attention layer in the previous round can be adjusted, and the weight matrix of each self-attention layer in the previous training round can be adjusted, so as to realize the adjustment of each self-attention layer, so that the input data can be processed through the adjusted weight matrix of each self-attention layer.

[0186] Optionally, when adjusting the parameters of the multi-modal recognition module, the second adapter parameter adjustment information corresponding to each second adapter of each self-attention layer can be obtained according to the loss result of the previous training round. On this basis, according to the second enhancement factor, the loss result of the previous training round, and the second adapter parameter adjustment information, the second parameter adjustment information of each self-attention layer is obtained.

[0187] Optionally, when adjusting the parameters of the multi-modal recognition module, the second adapter parameter adjustment information corresponding to each second adapter of each self-attention layer can be obtained according to the second enhancement factor and the loss result of the previous training round. On this basis, according to the loss result of the previous training round and the second adapter parameter adjustment information, the second parameter adjustment information of each self-attention layer is obtained.

[0188] Among them, the second parameter adjustment information is used to indicate the adjustment information of the weight matrix of the self-attention layer and each parameter in each second adapter of the self-attention layer.

[0189] S1102. Adjust the weight matrix of each self-attention layer according to each second parameter adjustment information to obtain the weight matrix of each self-attention layer in the current training round.

[0190] Optionally, after obtaining the second parameter adjustment information, the weight matrices of the respective attention layers can be adjusted according to the second parameter adjustment information of the respective attention layers, so as to obtain the weight matrices of the respective attention layers in the current training round, enabling the input data to be feature-processed through the weight matrices of the respective attention layers in the current training round to obtain the outputs of the respective attention layers, finely adjusting the weight matrices of the respective attention layers, thus being able to more accurately locate the parameters to be adjusted, narrowing the weight update gap between the middle layer and the bottom layer and the top layer during fine-tuning, enhancing the importance of the middle layer during the fine-tuning process, and narrowing the performance gap with full-scale fine-tuning. At the same time, unnecessary computational overhead is avoided, computational resources are utilized more efficiently, and overfitting of the respective attention layers to the input data can also be reduced, improving the generalization ability of the respective attention layers on unseen data.

[0191] S1103. Determine the third parameter adjustment information of each second fully-connected layer according to the third enhancement factor, the loss result of the previous training round, the weight matrices of each second fully-connected layer in the previous training round, and the weight matrices of the third adapters of each second fully-connected layer in the previous training round.

[0192] Optionally, when adjusting the parameters of the multi-modal recognition module, the third adapter parameter adjustment information corresponding to the third adapter of each second fully-connected layer can be obtained according to the loss result of the previous training round. On this basis, the first parameter adjustment information of each second fully-connected layer can be obtained according to the third enhancement factor, the loss result of the previous training round, and the third adapter parameter adjustment information.

[0193] Optionally, when adjusting the parameters of the multi-modal recognition module, the third adapter parameter adjustment information corresponding to the third adapter of each second fully-connected layer can be obtained according to the third enhancement factor and the loss result of the previous training round. On this basis, the first parameter adjustment information of each second fully-connected layer can be obtained according to the loss result of the previous training round and the third adapter parameter adjustment information.

[0194] S1104. Adjust the weight matrices of each second fully-connected layer according to the respective third parameter adjustment information to obtain the weight matrices of each second fully-connected layer in the current training round.

[0195] Optionally, after obtaining the third parameter adjustment information, the weight matrices of the respective second fully-connected layers can be adjusted according to the third parameter adjustment information of each second fully-connected layer, so as to obtain the weight matrices of the respective second fully-connected layers in the current training round, enabling the input data to be feature-processed through the weight matrices of the respective second fully-connected layers in the current training round to obtain the outputs of the respective second fully-connected layers, being able to finely adjust the weight matrices of the respective second fully-connected layers, thereby being able to more accurately locate the parameters to be adjusted, narrowing the weight update gap between the intermediate layer and the bottom layer and the top layer during fine-tuning, enhancing the importance of the intermediate layer during the fine-tuning process, and narrowing the performance gap with full-scale fine-tuning. At the same time, unnecessary computational overhead is avoided, computational resources are utilized more efficiently, and overfitting of the respective second fully-connected layers to the input data can also be reduced, improving the generalization ability of the respective second fully-connected layers on unseen data.

[0196] In a possible implementation manner, the weight matrix of the self-attention layer includes a query matrix, a key matrix, and a value matrix, the second enhancement factor includes a query enhancement factor, a key enhancement factor, and a value enhancement factor, and each second adapter of the self-attention layer includes a second adapter corresponding to the query matrix, a second adapter corresponding to the key matrix, and a second adapter corresponding to the value matrix. Figure 12 This is a schematic flowchart for determining the second parameter adjustment information of the self-attention layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiments of the present application. Refer to Figure 12 As shown, the above S1101 determines the second parameter adjustment information of the self-attention layer according to the second enhancement factor, the loss result of the previous training round, the weight matrix of the self-attention layer in the previous round, and the weight matrix of each second adapter of the self-attention layer in the previous training round, including:

[0197] S1201. Determine the query parameter adjustment information of the self-attention layer according to the query enhancement factor, the loss result of the previous training round, the query matrix of the self-attention layer in the previous round, and the weight matrix of the second adapter corresponding to the query matrix of the self-attention layer in the previous round.

[0198] Optionally, when adjusting the parameters of the multi-modal recognition module, the respective second adapter parameter adjustment information corresponding to the query matrix of each second adapter of the self-attention layer can be obtained according to the loss result of the previous training round. On this basis, the query parameter adjustment information of the self-attention layer can be obtained according to the query enhancement factor, the loss result of the previous training round, and the respective second adapter parameter adjustment information corresponding to the query matrix of each second adapter.

[0199] Optionally, when adjusting the parameters of the multi-modal recognition module, the adjustment information of the parameters of each second adapter corresponding to the query matrix of each second adapter in each attention layer can be obtained according to the query enhancement factor and the loss result of the previous training round. On this basis, according to the loss result of the previous training round and the adjustment information of the parameters of each second adapter corresponding to the query matrix of each second adapter, the adjustment information of the query parameters of each attention layer can be obtained.

[0200] S1202. Determine the adjustment information of the key parameters of each attention layer according to the key enhancement factor, the loss result of the previous training round, the key matrix of each attention layer in the previous round, and the weight matrix of the second adapter corresponding to the key matrix of each attention layer in the previous round.

[0201] Optionally, when adjusting the parameters of the multi-modal recognition module, the adjustment information of the parameters of each second adapter corresponding to the key matrix of each second adapter in each attention layer can be obtained according to the loss result of the previous training round. On this basis, according to the key enhancement factor, the loss result of the previous training round, and the adjustment information of the parameters of each second adapter corresponding to the key matrix of each second adapter, the adjustment information of the key parameters of each attention layer can be obtained.

[0202] Optionally, when adjusting the parameters of the multi-modal recognition module, the adjustment information of the parameters of each second adapter corresponding to the key matrix of each second adapter in each attention layer can be obtained according to the key enhancement factor and the loss result of the previous training round. On this basis, according to the loss result of the previous training round and the adjustment information of the parameters of each second adapter corresponding to the key matrix of each second adapter, the adjustment information of the key parameters of each attention layer can be obtained.

[0203] S1203. Determine the adjustment information of the value parameters of each attention layer according to the value enhancement factor, the loss result of the previous training round, the value matrix of each attention layer in the previous round, and the weight matrix of the second adapter corresponding to the value matrix of each attention layer in the previous round.

[0204] Optionally, when adjusting the parameters of the multi-modal recognition module, the adjustment information of the parameters of each second adapter corresponding to the value matrix of each second adapter in each attention layer can be obtained according to the loss result of the previous training round. On this basis, according to the value enhancement factor, the loss result of the previous training round, and the adjustment information of the parameters of each second adapter corresponding to the value matrix of each second adapter, the adjustment information of the value parameters of each attention layer can be obtained.

[0205] Optionally, when adjusting the parameters of the multi-modal recognition module, the adjustment information of the second adapter parameters corresponding to the value matrices of the respective second adapters of each attention layer can be obtained according to the value enhancement factor and the loss result of the previous training round. On this basis, according to the loss result of the previous training round and the adjustment information of the second adapter parameters corresponding to the value matrices of the respective second adapters, the adjustment information of the value parameters of each attention layer can be obtained.

[0206] S1204. Determine the second parameter adjustment information of each attention layer according to the query parameter adjustment information, the key parameter adjustment information, and the value parameter adjustment information of each attention layer.

[0207] Optionally, after obtaining the query parameter adjustment information, the key parameter adjustment information, and the value parameter adjustment information of each attention layer, the query parameter adjustment information, the key parameter adjustment information, and the value parameter adjustment information of each attention layer can be integrated to obtain the second parameter adjustment information of each attention layer.

[0208] Exemplarily, the query parameter adjustment information, the key parameter adjustment information, and the value parameter adjustment information of each attention layer are weighted averaged, concatenated, or combined in other forms to obtain the second parameter adjustment information of each attention layer.

[0209] In a possible implementation manner, the above S1201 determines the query parameter adjustment information of each attention layer according to the query enhancement factor, the loss result of the previous training round, the query matrix of each attention layer in the previous round, and the weight matrix of the second adapter corresponding to the query matrix of each attention layer in the previous round, including:

[0210] S1301. Determine the third low-rank transformation matrix and the fourth low-rank transformation matrix according to the image dimension in the training sample data.

[0211] Among them, the dimension after multiplying the third low-rank transformation matrix and the fourth low-rank transformation matrix matches the image dimension.

[0212] Optionally, the third low-rank transformation matrix and the fourth low-rank transformation matrix can be obtained by decomposing according to the image dimension in the training sample data through the low-rank decomposition method.

[0213] Among them, the dimension after multiplying the third low-rank transformation matrix and the fourth low-rank transformation matrix matches the image dimension. The low-rank decomposition method can include methods such as singular value decomposition (SVD) and non-negative matrix factorization (NMF).

[0214] Exemplarily, in the initial situation, the image dimension in the training sample data can be decomposed using a pre-configured target rank to obtain a third low-rank transformation matrix and a fourth low-rank transformation matrix, and during the training process, methods such as singular value decomposition (SVD) and non-negative matrix factorization (NMF) can be used to optimize the target rank to achieve the optimization of the third low-rank transformation matrix and the fourth low-rank transformation matrix.

[0215] Exemplarily, assume that the image dimension in the training sample data is , where is the height, is the width, and the third low-rank transformation matrix U3 can be obtained as and the fourth low-rank transformation matrix V4^T as , where, is the current target rank.

[0216] S1302. Determine the weight matrix of the second adapter corresponding to the query matrix in the current training round according to the loss result of the previous training round and the weight matrix of the second adapter corresponding to the query matrix of the self-attention layer in the previous training round.

[0217] Optionally, taking the L-th self-attention layer as an example, Figure 13 This is a schematic diagram when determining the query parameter adjustment information of the self-attention layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiments of the present application. Referring to Figure 13 as shown, x2q_L is the input corresponding to the query matrix of the L-th self-attention layer, W2q_L is the original trainable weight corresponding to the query matrix of the L-th self-attention layer, U3_L is the third low-rank transformation matrix corresponding to the query matrix of the L-th self-attention layer, V4^T_L is the fourth low-rank transformation matrix corresponding to the query matrix of the L-th self-attention layer. The third low-rank transformation matrix U3_L can be initialized to 0, and the fourth low-rank transformation matrix V4^T_L can be initialized to N2q_L(0, σ²), that is, randomly initialized from a normal distribution with a mean of 0 and a variance of σ². adapter2q_L is the second adapter corresponding to the query matrix of the L-th self-attention layer, and h2q_L is the differentiable function corresponding to the query matrix in the L-th self-attention layer, used to implement regularization or transformation to help the model better adapt to new tasks or data sets. αq is the query enhancement factor.

[0218] Among them, the query enhancement factor αq is a parameter used to control the enhancement amplitude. For example, when αq = 1.5, the enhancement factor of the intermediate layer is increased to 1.5 times.

[0219] Exemplarily, continue to refer to Figure 13As shown, the gradient can be calculated based on the loss result of the previous training round and the weight matrix of the second adapter corresponding to the query matrix of the self-attention layer in the previous training round, so as to adjust the weight matrices Wq_1 and Wq_2 of the second adapter corresponding to the query matrix in the current training round through the backpropagation algorithm.

[0220] S1303. Determine the second adapter corresponding to the query matrix in the current training round according to the weight matrix of the second adapter corresponding to the query matrix in the current training round and the fourth low-rank transformation matrix.

[0221] Optionally, taking the L-th self-attention layer as an example, the second adapter corresponding to the query matrix in the current training round can be calculated according to the weight matrices Wq_1 and Wq_2 of the second adapter corresponding to the query matrix and the fourth low-rank transformation matrix V4^T_L corresponding to the query matrix of the L-th self-attention layer.

[0222] Exemplarily, the second adapter adapter2q_L corresponding to the query matrix of the L-th self-attention layer can be expressed by the following formula:

[0223] adapter2q _L(V4_L^T) = σ(Wq_2* ReLU2q(Wq_1*V4^T_L))

[0224] Where V4^T_L is the fourth low-rank transformation matrix corresponding to the query matrix of the L-th self-attention layer, Wq_1 and Wq_2 are the weight matrices of the second adapter corresponding to the query matrix in the current training round, ReLU2q is the rectified linear unit activation function corresponding to the query matrix, defined as ReLU2q (x) = max(0, x), and σ is the activation function (such as Sigmoid or other non-linear functions).

[0225] Exemplarily, after obtaining the weight matrices Wq_1 and Wq_2 of the second adapter corresponding to the query matrix in the current training round, Wq_1 and Wq_2 and the fourth low-rank transformation matrix V4^T_L corresponding to the query matrix of the L-th self-attention layer can be input into the above formula to obtain the second adapter adapter2q_L(V4^T_L) corresponding to the query matrix of the L-th self-attention layer.

[0226] S1304. Determine the query parameter adjustment information of the self-attention layer according to the query enhancement factor, the second adapter corresponding to the query matrix in the current training round, and the third low-rank transformation matrix.

[0227] Exemplarily, taking the L-th self-attention layer as an example, after obtaining the second adapter adapter2q_L (V4^T_L) corresponding to the query matrix in the current training round, the query parameter adjustment information ΔWq_L of the L-th self-attention layer can be obtained according to the query enhancement factor αq, the second adapter adapter2q_L (V4^T_L) corresponding to the query matrix in the current training round, and the third low-rank transformation matrix U3_L corresponding to the query matrix.

[0228] Exemplarily, the query parameter adjustment information ΔWq_L of the L-th self-attention layer can be expressed by the following formula:

[0229] ΔWq_L = αq * U3_L * adapter2q_1(V4_L^T)

[0230] Exemplarily, taking the L-th self-attention layer as an example, after obtaining the query parameter adjustment information ΔWq_L of the L-th self-attention layer, the query parameter adjustment information ΔWq_L can be superimposed on the weight matrix Wq_L corresponding to the query matrix of the L-th self-attention layer to obtain the weight matrix Wq'_L of the query matrix of the L-th self-attention layer in the current training round, where Wq'_L = Wq_L + ΔWq_L.

[0231] In a possible implementation manner, the above S1202 determines the key parameter adjustment information of each self-attention layer according to the key enhancement factor, the loss result of the previous training round, the key matrix of each self-attention layer in the previous round, and the weight matrix of the second adapter corresponding to the key matrix of each self-attention layer in the previous round, including:

[0232] S1401. Determine the fifth low-rank transformation matrix and the sixth low-rank transformation matrix according to the image dimension in the training sample data, and the dimension after multiplying the fifth low-rank transformation matrix and the sixth low-rank transformation matrix matches the image dimension.

[0233] Wherein, the dimension after multiplying the fifth low-rank transformation matrix and the sixth low-rank transformation matrix matches the image dimension.

[0234] Optionally, the fifth low-rank transformation matrix and the sixth low-rank transformation matrix can be obtained by decomposing according to the image dimension in the training sample data through a low-rank decomposition method.

[0235] Wherein, the dimension after multiplying the fifth low-rank transformation matrix and the sixth low-rank transformation matrix matches the image dimension. The low-rank decomposition method can include methods such as singular value decomposition (SVD) and non-negative matrix decomposition (NMF).

[0236] Exemplarily, in the initial situation, the image dimension in the training sample data can be decomposed using a pre-configured target rank to obtain a fifth low-rank transformation matrix and a sixth low-rank transformation matrix, and during the training process, methods such as singular value decomposition (SVD) and non-negative matrix factorization (NMF) can be used to optimize the target rank to achieve the optimization of the fifth low-rank transformation matrix and the sixth low-rank transformation matrix.

[0237] Exemplarily, assume that the image dimension in the training sample data is , where is the height, is the width, the fifth low-rank transformation matrix U5 can be obtained as and the sixth low-rank transformation matrix V6^T as , where, is the current target rank.

[0238] S1402. Determine the weight matrix of the second adapter corresponding to the key matrix in the current training round according to the loss result of the previous training round and the weight matrix of the second adapter corresponding to the key matrix of the self-attention layer in the previous training round.

[0239] Optionally, taking the L-th self-attention layer as an example, Figure 14 is a schematic diagram when determining the key parameter adjustment information of the self-attention layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiments of the present application. Referring to Figure 14 as shown, x2k_L is the input corresponding to the key matrix of the L-th self-attention layer, W2k_L is the original trainable weight corresponding to the key matrix of the L-th self-attention layer, U5_L is the fifth low-rank transformation matrix corresponding to the key matrix of the L-th self-attention layer, V6^T_L is the sixth low-rank transformation matrix corresponding to the key matrix of the L-th self-attention layer. The fifth low-rank transformation matrix U5_L can be initialized to 0, and the sixth low-rank transformation matrix V6^T_L can be initialized to N2k_L(0, σ²), that is, randomly initialized from a normal distribution with a mean of 0 and a variance of σ². adapter2k_L is the second adapter corresponding to the key matrix of the L-th self-attention layer, h2k_L is a differentiable function corresponding to the key matrix in the L-th self-attention layer, used to implement regularization or transformation to help the model better adapt to new tasks or data sets, and αk is the key enhancement factor.

[0240] Among them, the key enhancement factor αk is a parameter used to control the enhancement amplitude. For example, when αk = 1.5, the enhancement factor of the intermediate layer increases to 1.5 times.

[0241] Exemplarily, continue to refer to Figure 14As shown, the gradient can be calculated based on the loss result of the previous training round and the weight matrix of the second adapter corresponding to the key matrix of the self-attention layer in the previous training round, so as to adjust the weight matrices Wk_1 and Wk_2 of the second adapter corresponding to the key matrix in the current training round through the backpropagation algorithm.

[0242] S1403. Determine the second adapter corresponding to the key matrix in the current training round according to the weight matrix of the second adapter corresponding to the key matrix in the current training round and the sixth low-rank transformation matrix.

[0243] Optionally, taking the L-th self-attention layer as an example, the second adapter corresponding to the key matrix in the current training round can be calculated according to the weight matrices Wk_1 and Wk_2 of the second adapter corresponding to the key matrix and the sixth low-rank transformation matrix V6^T_L corresponding to the key matrix of the L-th self-attention layer.

[0244] Exemplarily, the second adapter adapter2k_L corresponding to the key matrix of the L-th self-attention layer can be expressed by the following formula:

[0245] adapter2k _L(V6_L^T) = σ(Wk_2* ReLU2k(Wk_1*V6^T_L))

[0246] Where V6^T_L is the sixth low-rank transformation matrix corresponding to the key matrix of the L-th self-attention layer, Wk_1 and Wk_2 are the weight matrices of the second adapter corresponding to the key matrix in the current training round, ReLU2k is the rectified linear unit activation function corresponding to the key matrix, defined as ReLU2k (x) = max(0, x), and σ is the activation function (such as Sigmoid or other non-linear functions).

[0247] Exemplarily, after obtaining the weight matrices Wk_1 and Wk_2 of the second adapter corresponding to the key matrix in the current training round, Wk_1 and Wk_2 and the sixth low-rank transformation matrix V6^T_L corresponding to the key matrix of the L-th self-attention layer can be input into the above formula to obtain the second adapter adapter2k_L (V6^T_L) corresponding to the key matrix of the L-th self-attention layer.

[0248] S1404. Determine the key parameter adjustment information of the self-attention layer according to the key enhancement factor, the second adapter corresponding to the key matrix in the current training round, and the fifth low-rank transformation matrix.

[0249] Exemplarily, taking the L-th self-attention layer as an example, after obtaining the second adapter adapter2k_L (V6^T_L) corresponding to the key matrix in the current training round, the key parameter adjustment information ΔWk_L of the L-th self-attention layer can be obtained according to the key enhancement factor αk, the second adapter adapter2k_L (V6^T_L) corresponding to the key matrix in the current training round, and the fifth low-rank transformation matrix U5_L corresponding to the key matrix.

[0250] Exemplarily, the key parameter adjustment information ΔWk_L of the L-th self-attention layer can be expressed by the following formula:

[0251] ΔWk_L = αk * U5_L * adapter2k_1(V6_L^T)

[0252] Exemplarily, taking the L-th self-attention layer as an example, after obtaining the key parameter adjustment information ΔWk_L of the L-th self-attention layer, the key parameter adjustment information ΔWk_L can be superimposed on the weight matrix Wk_L corresponding to the key matrix of the L-th self-attention layer to obtain the weight matrix Wk'_L of the key matrix of the L-th self-attention layer in the current training round, where Wk'_L = Wk_L + ΔWk_L.

[0253] In a possible implementation manner, the above S1203 determines the value parameter adjustment information of each self-attention layer according to the value enhancement factor, the loss result of the previous training round, the value matrix of each self-attention layer in the previous round, and the weight matrix of the second adapter corresponding to the value matrix of each self-attention layer in the previous round, including:

[0254] S1501. Determine the seventh low-rank transformation matrix and the eighth low-rank transformation matrix according to the image dimension in the training sample data, and the dimension after multiplying the seventh low-rank transformation matrix and the eighth low-rank transformation matrix matches the image dimension.

[0255] Among them, the dimension after multiplying the seventh low-rank transformation matrix and the eighth low-rank transformation matrix matches the image dimension.

[0256] Optionally, the seventh low-rank transformation matrix and the eighth low-rank transformation matrix can be obtained by decomposing according to the image dimension in the training sample data through a low-rank decomposition method.

[0257] Among them, the dimension after multiplying the seventh low-rank transformation matrix and the eighth low-rank transformation matrix matches the image dimension. The low-rank decomposition method can include methods such as singular value decomposition (SVD) and non-negative matrix decomposition (NMF).

[0258] Exemplarily, in the initial situation, the image dimension in the training sample data can be decomposed using a pre-configured target rank to obtain a seventh low-rank transformation matrix and an eighth low-rank transformation matrix, and during the training process, methods such as singular value decomposition (SVD) and non-negative matrix factorization (NMF) can be used to optimize the target rank to achieve the optimization of the seventh low-rank transformation matrix and the eighth low-rank transformation matrix.

[0259] Exemplarily, assume that the image dimension in the training sample data is , where is the height, is the width, the seventh low-rank transformation matrix U7 can be obtained as and the eighth low-rank transformation matrix V8^T as , where, is the current target rank.

[0260] S1502. Determine the weight matrix of the second adapter corresponding to the value matrix in the current training round according to the loss result of the previous training round and the weight matrix of the second adapter corresponding to the value matrix of the self-attention layer in the previous training round.

[0261] Optionally, taking the L-th self-attention layer as an example, Figure 15 is a schematic diagram when determining the value parameter adjustment information of the self-attention layer in the fine-tuning and inference method of the drawing recognition large model provided by the embodiments of the present application. Referring to Figure 15 as shown, x2v_L is the input corresponding to the value matrix of the L-th self-attention layer, W2 v _L is the original trainable weight corresponding to the value matrix of the L-th self-attention layer, U7_L is the seventh low-rank transformation matrix corresponding to the value matrix of the L-th self-attention layer, V8^T_L is the eighth low-rank transformation matrix corresponding to the value matrix of the L-th self-attention layer. The seventh low-rank transformation matrix U7_L can be initialized to 0, and the eighth low-rank transformation matrix V8^T_L can be initialized to N2v_L(0, σ²), that is, randomly initialized from a normal distribution with a mean of 0 and a variance of σ². adapter2v_L is the second adapter corresponding to the value matrix of the L-th self-attention layer, h2v_L is the differentiable function corresponding to the value matrix in the L-th self-attention layer, used to implement regularization or transformation to help the model better adapt to new tasks or datasets, and αv is the value enhancement factor.

[0262] Among them, the value enhancement factor αv is a parameter used to control the enhancement amplitude. For example, when αv = 1.5, the enhancement factor of the middle layer increases to 1.5 times.

[0263] Exemplarily, continue to refer to Figure 15As shown, the gradient can be calculated based on the loss result of the previous training round and the weight matrix of the second adapter corresponding to the value matrix of the self-attention layer in the previous training round, so as to adjust the weight matrices Wv_1 and Wv_2 of the second adapter corresponding to the value matrix in the current training round through the backpropagation algorithm.

[0264] S1503. Determine the second adapter corresponding to the value matrix in the current training round according to the weight matrix of the second adapter corresponding to the value matrix in the current training round and the eighth low-rank transformation matrix.

[0265] Optionally, taking the L-th self-attention layer as an example, the second adapter corresponding to the value matrix in the current training round can be calculated according to the weight matrices Wv_1 and Wv_2 of the second adapter corresponding to the value matrix and the eighth low-rank transformation matrix V8^T_L corresponding to the value matrix of the L-th self-attention layer.

[0266] Exemplarily, the second adapter adapter2v_L corresponding to the value matrix of the L-th self-attention layer can be expressed by the following formula:

[0267] adapter2v _L(V8_L^T) = σ(Wv_2* ReLU2v(Wv_1*V8^T_L))

[0268] Where V8^T_L is the eighth low-rank transformation matrix corresponding to the value matrix of the L-th self-attention layer, Wv_1 and Wv_2 are the weight matrices of the second adapter corresponding to the value matrix in the current training round, ReLU2v is the rectified linear unit activation function corresponding to the value matrix, defined as ReLU2v (x) = max(0, x), and σ is the activation function (such as Sigmoid or other non-linear functions).

[0269] Exemplarily, after obtaining the weight matrices Wv_1 and Wv_2 of the second adapter corresponding to the value matrix in the current training round, Wv_1 and Wv_2 and the eighth low-rank transformation matrix V8^T_L corresponding to the value matrix of the L-th self-attention layer can be input into the above formula to obtain the second adapter adapter2v_L (V8^T_L) corresponding to the value matrix of the L-th self-attention layer.

[0270] S1504. Determine the value parameter adjustment information of the self-attention layer according to the value enhancement factor, the second adapter corresponding to the value matrix in the current training round, and the seventh low-rank transformation matrix.

[0271] Exemplarily, taking the L-th self-attention layer as an example, after obtaining the second adapter adapter2v_L (V8^T_L) corresponding to the value matrix in the current training round, the value parameter adjustment information ΔWv_L of the L-th self-attention layer can be obtained according to the value enhancement factor αv, the second adapter adapter2v_L (V8^T_L) corresponding to the value matrix in the current training round, and the seventh low-rank transformation matrix U7_L corresponding to the value matrix.

[0272] Exemplarily, the value parameter adjustment information ΔWv_L of the L-th self-attention layer can be represented by the following formula:

[0273] ΔWv_L = αv * U7_L * adapter2v_1(V8_L^T)

[0274] Exemplarily, taking the L-th self-attention layer as an example, after obtaining the value parameter adjustment information ΔWv_L of the L-th self-attention layer, the value parameter adjustment information ΔWv_L can be superimposed on the weight matrix Wv_L corresponding to the value matrix of the L-th self-attention layer to obtain the weight matrix Wv'_L of the value matrix of the L-th self-attention layer in the current training round, where Wv'_L = Wv_L + ΔWv_L.

[0275] Through the above processing, after freezing the weights of the pre-trained model, the trainable low-rank decomposition matrices can be injected into each layer of the Transformer architecture, thus greatly reducing the number of trainable parameters in the downstream tasks, and enabling the self-attention layer to update the intermediate layers in the self-attention layer more evenly during the update, which can reduce the performance deviation in the fine-tuning process, improve the generalization ability and overall performance of the obtained target drawing recognition model, and also enable the initial drawing recognition model to converge faster, so as to obtain the target drawing recognition model more efficiently.

[0276] In a possible implementation manner, the above S1103 determines the third parameter adjustment information of each second fully connected layer according to the third enhancement factor, the loss result of the previous training round, the weight matrices of each second fully connected layer in the previous training round, and the weight matrices of the third adapters of each second fully connected layer in the previous training round, including:

[0277] S1901. Determine the ninth low-rank transformation matrix and the tenth low-rank transformation matrix according to the image dimension in the training sample data.

[0278] Optionally, the specific processing process of S1901 can be implemented with reference to the processing process of the above S801. Specifically, the ninth low-rank transformation matrix can be regarded as the first low-rank transformation matrix in the above S801, and the tenth low-rank transformation matrix can be regarded as the second low-rank transformation matrix in the above S801.

[0279] S1902. Determine the weight matrix of the third adapter in the current training round according to the loss result of the previous training round and the weight matrix of the third adapter in the second fully connected layer in the previous training round.

[0280] Optionally, the specific processing procedure of S1902 can be implemented with reference to the processing procedure of S802 above. Specifically, the second fully connected layer can be regarded as the first fully connected layer in S802 above, and the third adapter can be regarded as the first adapter in S802 above, so as to determine the weight matrix of the third adapter in the current training round. Among them, the second fully connected layer and the first fully connected layer can have the same or different original trainable weights, adapters, differentiable functions, and weight matrices, and the third enhancement factor α3 can be the same as or different from the first enhancement factor α1.

[0281] S1903. Determine the third adapter in the current training round according to the weight matrix of the third adapter in the current training round and the ninth low-rank transformation matrix.

[0282] Optionally, the specific processing procedure of S1903 can be implemented with reference to the processing procedure of S803 above, so as to calculate the third adapter adapter3_L(V10_L^T) in the current training round.

[0283] S1904. Determine the third parameter adjustment information of the second fully connected layer according to the third enhancement factor, the third adapter in the current training round, and the tenth low-rank transformation matrix.

[0284] Optionally, the specific processing procedure of S1904 can be implemented with reference to the processing procedure of S804 above, so as to obtain the third parameter adjustment information ΔW3_L of the L-th second fully connected layer.

[0285] Exemplarily, taking the L-th second fully connected layer as an example, after obtaining the third parameter adjustment information ΔW3_L of the L-th second fully connected layer, the third parameter adjustment information ΔW3_L can be superimposed on the weight matrix W3_L of the L-th second fully connected layer to obtain the weight matrix W3'_L of the L-th second fully connected layer in the current training round, where W3'_L = W3_L + ΔW3_L.

[0286] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the fine-tuning and inference method of the above-mentioned drawing recognition large model.

[0287] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. A fine-tuning and reasoning method for a large model of drawing recognition, characterized in that: include: Acquire original sample data, annotation information, and a task statement, wherein the original sample data includes: a plurality of original engineering drawing images, and the task statement is used to indicate a task type of at least one task to be identified in the original sample data; According to the original sample data and the annotation information, the original sample data is augmented to generate training sample data; In the current training round, the training sample data is input into the image processing module in the initial drawing recognition model to obtain an image sequence to be processed, and the task statement is input into the text processing module in the initial drawing recognition model to obtain a text sequence to be processed, wherein the initial drawing recognition model includes: an image processing module, a text processing module, a feature processing module and a multimodal recognition module; The feature processing module adjusts the parameters of the feature processing module based on a preset first enhancement factor and a loss result of a previous training round, and performs feature processing on the image sequence to be processed and the text sequence to be processed to obtain a feature processing result, wherein the first enhancement factor is used to control the enhancement amplitude of the intermediate layer in the feature processing module; The multimodal recognition module adjusts the parameters of the multimodal recognition module based on a preset second enhancement factor, a preset third enhancement factor, and a loss result of a previous training round, and the multimodal recognition module performs multimodal recognition on the feature processing result to obtain a recognition result, wherein the second enhancement factor is used to adjust the self-attention layer in the multimodal recognition module, and the third enhancement factor is used to adjust the fully connected layer in the multimodal recognition module; According to the recognition result, the true value and the preset loss function, the loss result is determined, and the initial drawing recognition model is iteratively fine-tuned according to the loss result to obtain the target drawing recognition model, and the target drawing recognition model is used for reasoning.

2. The fine-tuning and reasoning method of the large model for drawing recognition according to claim 1 is characterized in that: The augmenting the original sample data according to the original sample data and the annotation information to generate the training sample data includes: Adding the annotation information to the corresponding position in the original sample data to obtain the annotated sample data; Based on a preset rotation matrix, a preset scaling matrix, and a preset cropping matrix, the labeled sample data is randomly rotated, randomly scaled, and randomly cropped in sequence to obtain randomly processed labeled sample data; Performing consistency processing on the labeling information in the randomly processed labeling sample data to obtain target sample data; Image quality adjustment is performed on the target sample data to obtain the training sample data.

3. The fine-tuning and reasoning method of the large model for drawing recognition according to claim 2 is characterized in that: The performing consistency processing on the label information in the randomly processed label sample data to obtain target sample data includes: Determine whether a first annotation text in the annotation information in the randomly processed annotation sample data is complete, wherein the first annotation text is any one of the annotation texts in the annotation information; If yes, retaining the first annotation text in the randomly processed annotation sample data; if no, deleting the first annotation text in the randomly processed annotation sample data; After all randomly processed labeled sample data are processed, the target sample data are obtained.

4. The fine-tuning and reasoning method of the large model for drawing recognition according to claim 2 is characterized in that: The step of adjusting the image quality of the target sample data to obtain the training sample data includes: Performing illumination adjustment on the target sample data to obtain pixel-adjusted target sample data; Performing comparison and adjustment on the target sample data to obtain the compared and adjusted target sample data; Performing fuzzy adjustment on the target sample data to obtain fuzzy adjusted target sample data; The training sample data is generated according to the pixel-adjusted target sample data, the contrast-adjusted target sample data and the blur-adjusted target sample data.

5. The fine-tuning and reasoning method of the large model for drawing recognition according to claim 1 is characterized in that: The feature processing module in the initial drawing recognition model includes: a splicing layer and at least one perception module connected in sequence; the perception module includes: a first fully connected layer and an activation layer; a first adapter is respectively deployed in each first fully connected layer of each perception module, and each first fully connected layer has a weight matrix; The adjusting the parameters of the feature processing module based on the preset first enhancement factor and the loss result of the previous training round includes: Determine first parameter adjustment information of each first fully connected layer according to the first enhancement factor, the loss result of the previous training round, the weight matrix of each first fully connected layer in the previous training round, and the weight matrix of the first adapter of each first fully connected layer in the previous training round; According to each of the first parameter adjustment information, the weight matrix of each of the first fully connected layers is adjusted to obtain the weight matrix of each of the first fully connected layers in the current training round.

6. The fine-tuning and reasoning method of the large model for drawing recognition according to claim 5 is characterized in that: The determining, according to the first enhancement factor, the loss result of the previous training round, the weight matrix of each first fully connected layer in the previous training round, and the weight matrix of the first adapter of each first fully connected layer in the previous training round, the first parameter adjustment information of each first fully connected layer includes: Determine a first low-rank transformation matrix and a second low-rank transformation matrix according to the image dimension in the training sample data, wherein the dimension after multiplying the first low-rank transformation matrix by the second low-rank transformation matrix matches the image dimension; Determine a weight matrix of the first adapter in a current training round according to a loss result of a previous training round and a weight matrix of the first adapter of the first fully connected layer in the previous training round; Determining the first adapter in the current training round according to the weight matrix of the first adapter in the current training round and the second low-rank transformation matrix; Determine first parameter adjustment information of the first fully connected layer according to the first enhancement factor, the first adapter in the current training round, and the first low-rank transformation matrix.

7. The fine-tuning and reasoning method of the large model for drawing recognition according to claim 1 is characterized in that: The multimodal recognition module includes: a plurality of feature recognition modules connected in sequence, each feature recognition module includes: a self-attention layer and a second fully connected layer connected in sequence; a plurality of second adapters are respectively deployed in each attention layer, and a third adapter is respectively deployed in each second fully connected layer; The adjusting the parameters of the multimodal recognition module based on the preset second enhancement factor, the preset third enhancement factor and the loss result of the previous training round includes: Determine second parameter adjustment information of the respective attention layers according to the second enhancement factor, the loss result of the previous training round, the weight matrix of the respective attention layers in the previous training round, and the weight matrix of each second adapter of the respective attention layers in the previous training round; According to each of the second parameter adjustment information, the weight matrix of each attention layer is adjusted to obtain the weight matrix of each attention layer in the current training round; Determining third parameter adjustment information of each second fully connected layer according to the third enhancement factor, the loss result of the previous training round, the weight matrix of each second fully connected layer in the previous training round, and the weight matrix of the third adapter of each second fully connected layer in the previous training round; According to each of the third parameter adjustment information, the weight matrix of each second fully connected layer is adjusted to obtain the weight matrix of each second fully connected layer in the current training round.

8. The fine-tuning and reasoning method of the large model for drawing recognition according to claim 7 is characterized in that: The weight matrix of each attention layer includes a query matrix, a key matrix and a value matrix, the second enhancement factor includes a query enhancement factor, a key enhancement factor and a value enhancement factor, and each second adapter of each attention layer includes a second adapter corresponding to the query matrix, a second adapter corresponding to the key matrix and a second adapter corresponding to the value matrix; The determining the second parameter adjustment information of each attention layer according to the second enhancement factor, the loss result of the previous training round, the weight matrix of each attention layer in the previous training round, and the weight matrix of each second adapter of each attention layer in the previous training round includes: Determine query parameter adjustment information of each attention layer according to the query enhancement factor, the loss result of the previous training round, the query matrix of each attention layer in the previous round, and the weight matrix of the second adapter corresponding to the query matrix of each attention layer in the previous round; Determine key parameter adjustment information of each attention layer according to the key enhancement factor, the loss result of the previous training round, the key matrix of each attention layer in the previous round, and the weight matrix of the second adapter corresponding to the key matrix of each attention layer in the previous round; Determine value parameter adjustment information of each attention layer according to the value enhancement factor, the loss result of the previous training round, the value matrix of each attention layer in the previous round, and the weight matrix of the second adapter corresponding to the value matrix of each attention layer in the previous round; Determine the second parameter adjustment information of each attention layer according to the query parameter adjustment information of each attention layer, the key parameter adjustment information of each attention layer and the value parameter adjustment information of each attention layer.

9. The fine-tuning and reasoning method of the large model for drawing recognition according to claim 8 is characterized in that: Determining query parameter adjustment information of each attention layer according to the query enhancement factor, the loss result of the previous training round, the query matrix of each attention layer in the previous round, and the weight matrix of the second adapter corresponding to the query matrix of each attention layer in the previous round includes: Determine a third low-rank transformation matrix and a fourth low-rank transformation matrix according to the image dimension in the training sample data, wherein the dimension after multiplying the third low-rank transformation matrix by the fourth low-rank transformation matrix matches the image dimension; Determine a weight matrix of the second adapter in the current training round according to the loss result of the previous training round and the weight matrix of the second adapter in the previous training round corresponding to the query matrix of the self-attention layer; Determining the second adapter in the current training round according to the weight matrix of the second adapter in the current training round and the third low-rank transformation matrix; Determine query parameter adjustment information of the self-attention layer according to the query enhancement factor, the second adapter in the current training round, and the fourth low-rank transformation matrix.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the method for fine-tuning and reasoning a large model for drawing recognition as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Model training method, article identification method and device, electronic equipment and medium

    CN115310547A

  • Universal multi-modal learning method based on deep interactive adaptation network model

    CN116882477A