SOP file generation method and system and electronic equipment
By using a large language model to predict working time categories and integrate SOP association information, the problem of low efficiency and accuracy of SOP file production in the existing technology is solved, and efficient and accurate SOP file generation is achieved.
Patent Information
- Application Number
- CN202510100510.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the production efficiency of SOP files is low and the accuracy is low, which is heavily dependent on the production personnel's process experience and subjective understanding.
By obtaining the basic data, including the process name and step name in the SOP file to be generated, input it into the preset large language model for working class prediction, obtain SOP association information based on the combined working class of working class and the preset working library, and finally generate the target SOP file by integrating this information.
It realizes efficient generation of SOP files, improves file accuracy, reduces costs, and improves the consistency and controllability of the production process.
Smart Images

Figure CN119990076A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a SOP file generation method, system and electronic equipment. Background Art
[0002] SOP (Standard Operating Procedure) is an important document to ensure the efficiency, standardization and consistency of the production process, especially in the field of automobile manufacturing. It can standardize the business and operation process of automobile manufacturing, and describe in detail the standard operating steps of each process link to ensure the consistency and controllability of the production process, thereby reducing unreasonable process operations, reducing operating costs and improving product quality.
[0003] At present, many automobile manufacturers still rely on manual methods to prepare SOP documents. However, this method has many disadvantages. Understandably, for a new product, such as a new model, it usually involves more than 30,000 steps. If you want to prepare the SOP document for the new model, it will take a lot of time and manpower, and the efficiency is low. In addition, this method relies heavily on the process experience and subjective understanding of the production staff, resulting in low accuracy of the produced SOP documents. Summary of the invention
[0004] The present invention provides a SOP file generation method, system and electronic device to solve the problems of low production efficiency and low accuracy of SOP files in the prior art.
[0005] The present invention provides a method for generating an SOP file, the method comprising: obtaining basic data, the basic data comprising: a process name and a step name in the SOP file to be generated;
[0006] Input the basic data into a preset large language model to predict the working time category, and obtain the working time category corresponding to the process name and the work step name in the basic data, wherein the working time category refers to the category of the action or operation to be performed corresponding to the work step;
[0007] Based on the working time category and a preset combined working time library, obtaining SOP association information corresponding to the working time category in the combined working time library, wherein the SOP association information includes operation content;
[0008] The target SOP file is obtained by integrating the basic data, the working time category, and the SOP related information.
[0009] In one embodiment of the present invention, before obtaining basic data, the method further includes:
[0010] Acquire a first data set, the first data set includes a plurality of first sample groups, the first sample group includes a historical SOP file of any product, the historical SOP file includes: a process name, a process step name, and a work time category;
[0011] Preprocessing the first data set to obtain sample data, wherein the preprocessing includes: error data processing, null value data processing, and format conversion, wherein the format conversion refers to converting the historical SOP file that has undergone error data processing and null value data processing into the sample data according to a preset format, wherein the preset format includes: an instruction name, an input item, and an output item, wherein the instruction name is used to instruct the large language model to perform a work hour category prediction based on the input item;
[0012] The sample data is input into the large language model for back propagation to complete model training.
[0013] In one embodiment of the present invention, the first sample group further includes a process flow file corresponding to the historical SOP file, and the process flow file includes: a process name, a step name, and size information of parts involved in each step;
[0014] The error data processing includes:
[0015] Based on the process name and the work step name, the historical SOP file and the process flow file are associated to obtain a target data table, wherein the primary key of the target data table is the process name and the work step name, and the value in the target data table is the work time category and the size information of the parts involved in the work step;
[0016] Extracting the part names in the process step names in the target data table, if the part names in two process step names are the same and the corresponding size information is the same, then determining that the process steps corresponding to the two current process step names are the same, and determining the tuples where the two current process step names are located as the tuples to be compared respectively;
[0017] If the working time categories in the two tuples to be compared are different, the two tuples to be compared are currently determined to be erroneous data, and based on the erroneous data, an erroneous data prompt is issued or the erroneous data is deleted.
[0018] In one embodiment of the present invention, the null value data processing includes:
[0019] If any tuple in the target data table includes a process name and a step name, and lacks a working time category, the current tuple is determined to be null value data, and the null value data is deleted.
[0020] In one embodiment of the present invention, the historical SOP file further includes: work content corresponding to the work time category, and description information of the work time category, the value in the target data table also includes the work content and the description information, and the sample data includes: first sample data, second sample data, and third sample data;
[0021] The input items of the first sample data are process name and step name, and the output item is the work time category; the input item of the second sample data is the work content, and the output item is the work time category; the input item of the third sample data is description information, and the input item is the work time category;
[0022] The sample data is input into the large language model for back propagation to complete model training, including: inputting the first sample data, the second sample data, and the third sample data into the large language model respectively to perform work hour category prediction to complete back propagation and obtain a trained large language model.
[0023] In one embodiment of the present invention, the sample data is input into the large language model to perform back propagation to complete model training, including:
[0024] splicing the instruction name and the input item in the sample data to obtain spliced data;
[0025] Inputting the concatenated data into the input layer of the large language model to input the data to be predicted of fixed dimension;
[0026] Inputting the data to be predicted into the pre-trained language layer and the low-rank adaptation layer in the large language model respectively, performing work hour category prediction, and obtaining first prediction data output by the pre-trained language layer and second prediction data output by the low-rank adaptation layer, wherein the low-rank adaptation layer includes a dimensionality reduction matrix and a dimensionality increase matrix connected in sequence;
[0027] Determine the sum of the first prediction data and the second prediction data as final prediction data;
[0028] Based on the final prediction data and the corresponding output items in the sample data, the low-rank adaptation layer is back-propagated and updated until the model converges or the number of iterations reaches a preset number.
[0029] In one embodiment of the present invention, after completing the training of the large language model, the method further includes:
[0030] Acquire a second data set, where the second data set includes a plurality of second sample groups, and the products corresponding to the second sample groups are different from the products corresponding to the first sample groups;
[0031] Performing the preprocessing on the second data set to obtain test data, wherein the format of the test data is the same as the format of the sample data;
[0032] The test data is input into the large language model, and the large language model evaluation is completed based on the prediction result output by the large language model.
[0033] In one embodiment of the present invention, the method further includes:
[0034] Based on a preset calibration rule, the target SOP file is calibrated, and the calibrated file is determined as the final SOP file;
[0035] Inputting the work content in the final SOP file into the large language model, performing work hour category prediction, and obtaining a target work hour category;
[0036] If the target working time category is missing in the combined working time library, the target working time category and the work content corresponding to the target working time category are added to the combined working time library for the next SOP file generation.
[0037] The present invention also provides a SOP file generation system, the system comprising: a basic data acquisition module, used to acquire basic data, the basic data comprising: a process name and a work step name in the SOP file to be generated;
[0038] A working time category prediction module is used to input the basic data into a preset large language model to perform working time category prediction, and obtain the working time category corresponding to the process name and the work step name in the basic data, wherein the working time category refers to the category of the action or operation to be performed corresponding to the work step;
[0039] A correlation information acquisition module, for obtaining SOP correlation information corresponding to the working time category in the combined working time library based on the working time category and a preset combined working time library, wherein the SOP correlation information includes operation content;
[0040] The integration module is used to obtain a target SOP file by integrating the basic data, the working time category, and the SOP related information.
[0041] The present invention also provides an electronic device, comprising a processor, a memory and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute a computer program stored in the memory to implement a SOP file generation method provided in any one of the above embodiments.
[0042] Beneficial effects of the present invention: The SOP file generation method, system and electronic device provided by the present invention obtain basic data, the basic data include: the process name and the work step name in the SOP file to be generated; input the basic data into a preset large language model, perform work time category prediction, obtain the work time category corresponding to the process name and the work step name in the basic data, the work time category refers to the category of the action or operation to be performed corresponding to the work step; based on the work time category and the preset combined work time library, obtain the SOP related information corresponding to the work time category in the combined work time library, the SOP related information includes the operation content; by integrating the basic data, the work time category, and the SOP related information, obtain the target SOP file. The method can automatically generate the target SOP file with high efficiency, high accuracy and low cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic diagram of a flow chart of a method for generating a SOP file provided by an embodiment of the present invention;
[0044] Figure 2 An exemplary structural diagram of a large language model in a SOP file generation method provided in one embodiment of the present invention;
[0045] Figure 3 A schematic diagram of a modeling process in a SOP file generation method provided in one embodiment of the present invention;
[0046] Figure 4 A schematic diagram of the structure of a SOP file generation system provided by an embodiment of the present invention;
[0047] Figure 5 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0049] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0050] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0051] In order to facilitate understanding of the SOP file generation method, system, electronic device and storage medium provided by the present invention, some technical terms involved in the present invention are explained below.
[0052] Process: A process refers to a major production task or a group of interrelated tasks completed in accordance with certain process requirements and operating sequences during the entire production process, such as welding processes and assembly processes.
[0053] Work step: A work step refers to a specific, separate operation or task in a process. It is usually the smallest execution unit of a process, and each work step usually refers to a specific operation link. For example, in a welding process, the following work steps may be included: preparing welding materials, welding operations, and cleaning after welding.
[0054] Combine the following Figures 1 to 5 , the SOP file generation method, system and electronic device provided by the present invention are explained.
[0055] See also Figure 1 , Figure 1 A flow chart of a method for generating a SOP file according to an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method includes:
[0056] S110: Obtain basic data, where the basic data includes: process name and step name in the SOP file to be generated.
[0057] S120: Input the basic data into a preset large language model to predict the working time category, and obtain the working time category corresponding to the process name and the working step name in the basic data, wherein the working time category refers to the category of the action or operation to be performed corresponding to the working step.
[0058] It should be noted that the working time category can be a single-level category or a multi-level category, such as a first-level working time category, a second-level working time category, etc. The working time category of each level is gradually refined, such as the first-level working time category is parts installation, the second-level working time category is placing parts to the installation position, etc.
[0059] It is understandable that a process may include multiple steps. Therefore, a process name usually corresponds to multiple step names, and each step name corresponds to a corresponding working time category.
[0060] It should also be noted that by setting up a large language model and using the large language model to predict the working time category, the accuracy of the target SOP file generated subsequently can be improved to a certain extent. It is understandable that if a large language model is used to generate the entire SOP file directly based on the process name and the work step name, on the one hand, it will increase the difficulty of training the large language model, and on the other hand, it will also cause the accuracy of the generated SOP file to be low. Therefore, this embodiment uses a large language model to predict the working time category, and based on the predicted working time category, performs subsequent SOP file generation, which can not only improve the degree of automation of SOP file generation, but also better ensure the accuracy of the target SOP file generated subsequently.
[0061] S130: Based on the working time category and a preset combined working time library, obtain SOP association information corresponding to the working time category in the combined working time library, wherein the SOP association information includes work content.
[0062] It should be noted that the combined man-hour library includes man-hour categories (such as first-level man-hour categories, second-level man-hour categories, etc.), work content (man-hour work content corresponding to the man-hour categories), man-hour codes, coding time, and man-hour value types (such as value-added man-hours and non-value-added man-hours), etc. Based on the man-hour categories predicted in step S120, the SOP association information corresponding to the man-hour categories is searched in the combined man-hour library, so as to facilitate the subsequent generation of the target SOP file.
[0063] In some embodiments, the SOP-related information also includes work hour codes, coding time, work hour value types, and description information of work hour categories.
[0064] S140: Obtain a target SOP file by integrating the basic data, the work hour category, and the SOP related information.
[0065] It can be understood that by integrating basic data, working time categories, and SOP related information, a target SOP file with higher accuracy can be obtained.
[0066] The SOP file generation method in the above embodiment can improve the generation efficiency of the SOP file while effectively improving the accuracy of the generated target SOP file, with low cost and high real-time performance.
[0067] In some embodiments, before obtaining basic data, the method further includes:
[0068] 1. Obtain a first data set, wherein the first data set includes a plurality of first sample groups, wherein the first sample group includes a historical SOP file of any product, wherein the historical SOP file includes: a process name, a process step name, and a work time category.
[0069] It is understandable that the product may be a historical product, such as a historical vehicle model, etc. By obtaining historical SOP files of multiple historical vehicle models, subsequent model training can be facilitated.
[0070] 2. Preprocess the first data set to obtain sample data, wherein the preprocessing includes: error data processing, null value data processing, and format conversion. The format conversion refers to converting the historical SOP file that has undergone error data processing and null value data processing into the sample data according to a preset format. The preset format includes: instruction name, input items, and output items. The instruction name is used to instruct the large language model to perform working hour category prediction based on the input items.
[0071] It should be noted that performing the above preprocessing on the first data set can help improve the accuracy of subsequent model training.
[0072] It is understandable that, since the role of the large language model is to predict the work time category, the input item can be any one or more items remaining in the historical SOP file except the work time category, and the output item is the work time category.
[0073] 3. Input the sample data into the large language model for back propagation to complete model training.
[0074] It should be noted that by adopting the above training method, the accuracy of the large language model can be improved significantly.
[0075] In some embodiments, the first sample group also includes a process flow (BOP) file corresponding to the historical SOP file, and the process flow file includes: a process name, a step name, and size information of parts involved in each step.
[0076] In some embodiments, the error data processing includes:
[0077] 1. Based on the process name and the work step name, the historical SOP file and the process flow file are associated to obtain a target data table, wherein the primary key of the target data table is the process name and the work step name, and the value in the target data table is the working time category and the size information of the parts involved in the work step.
[0078] It can be understood that, assuming that the target data table is distributed in rows, each row corresponds to a tuple, and each tuple includes the process name, step name, working time category, and size information of the parts involved in the step in the row.
[0079] It should also be mentioned that the values in the target data table can be the remaining contents in the historical SOP file except the process name and the work step name. Therefore, in addition to the working time category described in the above steps and the size information of the parts involved in the work step, its values can also include information such as the work content.
[0080] 2. Extract the part names in the process step names in the target data table. If the part names in two process step names are the same and the corresponding size information is the same, it is determined that the process steps corresponding to the two current process step names are the same, and the tuples where the two current process step names are located are respectively determined as tuples to be compared.
[0081] It is understandable that the BOP file records the size of the parts in each work step (work step name). Therefore, by comparing the part names in the work step name and the size of the hardware corresponding to the work step in the BOP file, it is possible to more accurately determine whether the two current work step names belong to the same work step, that is, whether the tuples containing the two current work step names belong to the same work step. If the part names in the two work step names are the same and the corresponding size information is the same, it can be determined that the work steps corresponding to the two current work step names are the same, that is, they belong to the same work step.
[0082] 3. If the working time categories in the two tuples to be compared are different, the two tuples to be compared are currently determined to be erroneous data, and based on the erroneous data, an erroneous data prompt is issued or the erroneous data is deleted.
[0083] It is understandable that if the working time categories in the two tuples to be compared are different, it means that there are errors in the two tuples to be compared. Therefore, based on the above error data, an error data prompt is issued or the error data is deleted. The error data prompt is used to instruct the user to modify or delete the error data. In addition, a corresponding error data processing program can also be set, such as automatically deleting the two tuples to be compared, or deleting one of the tuples to be compared.
[0084] In some embodiments, the null value data processing includes:
[0085] If any tuple in the target data table includes a process name and a step name, and lacks a working time category, the current tuple is determined to be null value data, and the null value data is deleted.
[0086] It is understandable that the above-mentioned processing of null value data can help improve the efficiency and accuracy of subsequent model training. In addition, if any tuple in the target data table includes a process name and a step name, and lacks work content, the current tuple can also be determined as null value data and the null value data can be deleted.
[0087] In some embodiments, the historical SOP file also includes: work content corresponding to the working time category, and descriptive information of the working time category, the values in the target data table also include the work content and the descriptive information, and the sample data includes: first sample data, second sample data, and third sample data.
[0088] It can be understood that, assuming that the working time category is visual inspection, the description information of the working time category can be manual observation, etc.
[0089] The input items of the first sample data are process name and step name, and the output item is working time category; the input item of the second sample data is work content, and the output item is working time category; the input item of the third sample data is description information, and the input item is working time category.
[0090] In some embodiments, the sample data is input into the large language model for back propagation to complete model training, including: inputting the first sample data, the second sample data, and the third sample data into the large language model respectively to perform work hour category prediction to complete back propagation and obtain a trained large language model.
[0091] It should be noted that by generating the above-mentioned first sample data, it is possible to facilitate the large language model to better learn the forward prediction content of the working time category, and improve the forward working time category prediction ability of the large language model based on the process name and the step name. By generating the above-mentioned second sample data, the working time category summary ability of the large language model can be improved. By generating the above-mentioned third sample data, it is possible to assist the large language model in understanding the meaning of the working time category, thereby improving the prediction accuracy of the large language model.
[0092] In some embodiments, the sample data is input into the large language model to perform back propagation to complete model training, including:
[0093] 1. Concatenate the instruction name and input item in the sample data to obtain concatenated data.
[0094] 2. Input the concatenated data into the input layer of the large language model to input the data to be predicted with a fixed dimension.
[0095] 3. Input the data to be predicted into the pre-trained language layer and the low-rank adaptation (LoRA) layer in the large language model respectively to perform working hour category prediction, and obtain the first prediction data output by the pre-trained language layer and the second prediction data output by the low-rank adaptation layer, wherein the low-rank adaptation layer includes a dimensionality reduction matrix and a dimensionality increase matrix connected in sequence.
[0096] It should be mentioned that the pre-trained language layer in this embodiment can be an intermediate parameter layer in a PLM (Pre-trained Language Model). In this embodiment, the low-rank adaptation layer is embedded in the PLM model to achieve the construction of a large language model structure. On this basis, the large language model is fine-tuned by fine-tuning and iterating the dimensionality reduction matrix and dimensionality increase matrix in the low-rank adaptation layer.
[0097] 4. Determine the sum of the first prediction data and the second prediction data as the final prediction data. It can be understood that the first prediction data and the second prediction data are added to obtain the final prediction data.
[0098] 5. Based on the final prediction data and the corresponding output items in the sample data, backpropagate and update the low-rank adaptation layer until the model converges or the number of iterations reaches a preset number.
[0099] It can be understood that during the model training process, the parameters of the pre-trained language layer are fixed, and only the dimension reduction matrix and the dimension increase matrix are trained and updated, so as to achieve fine-tuning of the large language model.
[0100] Figure 2 For an exemplary structural diagram of a large language model in a SOP file generation method provided in an embodiment of the present invention, please refer to Figure 2 , Figure 2 The x in the formula represents the model input, i.e., the concatenated data, and d represents the dimension of the concatenated data. W0 represents the original parameters of the pre-trained language layer, R represents the real number domain, and d×k represents the dimension of the real number domain. A represents the reduced dimension matrix, and r represents the dimension after the reduced dimension. B represents the increased dimension matrix. A∈N(0,σ 2 ) indicates that A belongs to normal distribution. The original value of B matrix is 0. In the subsequent iterations, it is gradually updated based on the original value. h indicates the final predicted data.
[0101] Since the subsequent applications of the large language model are mostly for new products, such as the generation of SOP files for new car models, it is necessary to perform model evaluation on the trained large language model to understand the generalization ability of the large language model for new products, that is, future applications. In some embodiments, a second data set different from the above-mentioned first data set is input into the large language model to achieve model evaluation, that is, a data set that has not been recognized by the large language model is input into the large language model for model evaluation. Specifically, after completing the training of the large language model, it also includes:
[0102] 1. Obtain a second data set, where the second data set includes a plurality of second sample groups, and the products corresponding to the second sample groups are different from the products corresponding to the first sample groups.
[0103] It is understandable that the second data set also includes historical SOP files of a certain product, etc., but the product is different from the product corresponding to the first data set.
[0104] 2. Perform the preprocessing on the second data set to obtain test data, wherein the format of the test data is the same as that of the sample data.
[0105] 3. Input the test data into the large language model, and complete the large language model evaluation based on the prediction result output by the large language model.
[0106] It can be understood that based on the prediction results output by the large language model and the output items in the test data, the gap between the two can be obtained. Based on this gap, the accuracy of the large language model in predicting the working hour category of the new product can be known, thereby completing the evaluation of the large language model.
[0107] The following is an example of a specific embodiment to illustrate the modeling process in the SOP file generation method in the above embodiment. Please refer to Figure 3 :
[0108] First, obtain the first data set.
[0109] Secondly, data preprocessing is performed to obtain sample data. Preprocessing includes: error data processing, null value data processing, and format conversion.
[0110] Then, the large language model is trained using the sample data. Specifically, it is determined whether the number of iterations reaches the preset number. If not, error calculation, back propagation, and parameter update are performed; if reached, model evaluation is performed.
[0111] Finally, save the model.
[0112] In some embodiments, the method further comprises:
[0113] 1. Based on a preset calibration rule, the target SOP file is calibrated, and the calibrated file is determined as the final SOP file.
[0114] It should be noted that the calibration rules can be set according to actual needs, such as correction of incoherent sentences, etc. Manual calibration can also be used for calibration.
[0115] 2. Input the work content in the final SOP file into the large language model, perform work hour category prediction, and obtain the target work hour category.
[0116] It should be noted that in the above embodiment, by generating the second sample data and training the large language model based on the second sample data, the large language model can be equipped with the ability to summarize the work time categories. Therefore, by inputting the work content in the final SOP file into the large language model and predicting the work time category, a target work time category with high accuracy can be obtained.
[0117] It is understandable that since the large language model is applied to the generation of the SOP file of the new product, by predicting the work time category based on the job content in the final SOP file after calibration, it is easy to obtain the target work time category corresponding to the job content. If the target work time category does not exist in the combined work time library, the target work time category can be added to the combined work time library to facilitate the subsequent generation of the SOP file and avoid information gaps caused by the failure to update the combined work time library in a timely manner.
[0118] 3. If the target working time category is missing in the combined working time library, the target working time category and the work content corresponding to the target working time category are added to the combined working time library for the next SOP file generation.
[0119] The following is an illustrative example of the training process of the large language model in the SOP file generation method in the above embodiment.
[0120] First, obtain the first data set, such as obtaining the historical SOP files and BOP files of five models (model A, model B, model C, model D, model E). Of course, you can also obtain the historical combined working time library. The data in the first data set can be presented in the form of a table, as shown in Table 1 below:
[0121] Table 1 Data examples in the first dataset
[0122]
[0123] Table 1 exemplarily shows some data in the first data set. In a specific implementation process, the first data set may also include other types of data, such as serial number, position, and number of actions (the number of actions corresponding to working hours).
[0124] It should be mentioned that the working time codes in Table 1 are exemplary codes and have no actual meaning.
[0125] Secondly, data preprocessing is performed. Data preprocessing includes: error data processing, null value data processing, and format conversion. The following is an exemplary description of error data processing: 1. Based on the process name and step name, the data in the first data set are counted and associated to obtain data such as the work content under the same process name and step name. 2. The part name in the step name of the data is extracted using a preset general large model to obtain the data under each part name, for example:
[0126] Part names in the work step names: right front door glass rear guide rail assembly, door;
[0127] Under this part name, it is assumed that two tuple data are obtained, namely:
[0128] Tuple 1: Levels (hours category): Parts installation (level 1 hour category) + Door glass guide groove (level 2 hour category); process_name (operation name): Install the rear guide rail assembly of the right front door glass to the door, operation_name (operation step name): Install the rear guide rail assembly of the right front door glass to the door; Data source (from the historical SOP file of a specific model)......
[0129] Tuple 2: Levels: Parts installation (first-level labor time category) + placing parts to the installation position (second-level labor time category); process_name: Install the rear guide rail assembly of the right front door glass to the door, operation_name: Install the rear guide rail assembly of the right front door glass to the door; Data source......
[0130] As can be seen from the above, the process names and step names in tuple 1 and tuple 2 are the same, and the part names extracted from the step names are also the same, both of which are the right front door glass rear guide rail assembly and the door. However, there are differences in the corresponding working time categories, which means that there is a contradiction between the two data, that is, there is at least one wrong data, which needs to be deleted or modified. By modifying tuple 1 or tuple 2, the processing of wrong data can be better achieved.
[0131] After processing the error data and the null value data, format conversion can be performed to obtain first sample data including instruction (instruction name), input (input item), and Output (output item). The first sample data generated after format conversion is exemplarily shown below.
[0132] Example 1:
[0133] instruction: You are a senior engineer in the final assembly workshop of automobile manufacturing. You are proficient in writing SOP documents and are good at splitting the names of the steps under the process into one or more action groups.
[0134] Input: There is a process named Install rear side panel rubber plugging cover, and the work step name under it is: Install 2 side panel rubber plugging covers to the left rear side panel. How many dynamic element groups are included in the work steps under this process, and what are they?
[0135] Output: Contains 1 dynamic element group, which is part installation + installation plug.
[0136] It can be understood that "parts installation" is a first-level labor time category, and "installation blocking" is a second-level labor time category. In addition, the action element group represents a combination of multiple actions.
[0137] Example 2:
[0138] instruction: You are a senior engineer in the final assembly workshop of automobile manufacturing. You are proficient in writing SOP documents and are good at splitting the names of the steps under the process into one or more action groups.
[0139] Input: There is a process named Install rear side panel rubber plugging cover, and the work step name under it is: Install 1 side panel rubber plugging cover to the left rear side panel. How many dynamic element groups are included in the work steps under this process, and what are they?
[0140] Output: Contains 1 dynamic element group, which is part installation + installation plug.
[0141] Then, the model is trained. The large language model is trained using the sample data of model A, model B, model C, and model D converted in the above format. The sample data corresponding to model E is used as the data for subsequent model evaluation.
[0142] The relevant hyperparameters of the large language model can be set, where the batch size can be 3, the learning rate can be 0.0001, and the number of iterations can be 5. By using back propagation to continuously adjust the parameters in the low-rank adaptive sub-model, when the number of iterations reaches the preset number, the model training ends and the model is saved.
[0143] Finally, the model is evaluated and saved. The sample data corresponding to vehicle type E is input into the large language model to obtain the model prediction results. The prediction accuracy of the large language model is obtained by statistically analyzing the gap between the model prediction results and the actual working hours category, so as to understand the generalization ability of the large language model.
[0144] The following is an exemplary display of the model prediction results obtained in the SOP file generation method:
[0145] {process_name: "FOB (Frequency Operated Button, remote control key) key scan";
[0146] operation_name: "Confirm the qualified result of software flashing";
[0147] batch_no (batch number): 2;
[0148] llm_output (model prediction result): "Contains 1 action element group, which is inspection + visual inspection";
[0149] version: sft_v21}
[0150] In addition, in the actual application process of the large language model, its output is only the predicted working time category. Therefore, it is also necessary to export SOP related information such as job content, working time coding, coding time, etc. from the combined working time library based on the working time category. By integrating the basic data, working time category, and SOP related information, the target SOP file is obtained. After obtaining the target SOP file, manual calibration can be performed, and the calibrated file is determined as the final SOP file. The job content in the final SOP file is input into the large language model, and the working time category is predicted to obtain the target working time category. The target working time category that does not appear in the combined working time library is updated to the combined working time library to facilitate the subsequent generation of SOP files.
[0151] The following Table 2 shows some of the contents in the combined working time library. See Table 2 for details:
[0152] Table 2 Example of some data in the combined working time database
[0153]
[0154]
[0155] In summary, the SOP file generation method in the above embodiment has the following advantages:
[0156] 1. When training the model, this method uses the process name and step name as the model input, and the work time category (the work time category can include the first-level work time category, the second-level work time category, etc.) as the model output. Compared with directly using the detailed SOP file content (such as detailed work time content, work time code, coding time, and work time category, etc.) as the model output, this method can improve the accuracy of the generated target SOP file to a certain extent.
[0157] 2. After generating the target SOP file for each new product (such as a new car model), the file is calibrated, and the work content in the calibrated final SOP file is input into the large language model to predict the working time category and obtain the target working time category; and the target working time category and its corresponding work content that are missing in the combined working time library are added to the combined working time library to form a closed loop of the SOP file generation process.
[0158] 3. By generating data in three formats, namely the first sample data, the second sample data, and the third sample data, before model training, and using these three data for model training, it is possible to improve the diversity of sample data while also improving the large language model's ability to predict positive working time categories, summarize working time categories, and understand the meaning of working time categories. It is understandable that by generating the first sample data, it is possible to facilitate the large language model to better learn the forward prediction content of working time categories, and improve the large language model's ability to predict positive working time categories based on process names and work step names. By generating the second sample data, the large language model's ability to summarize working time categories can be improved. By generating the third sample data, the large language model can be assisted in understanding the meaning of working time categories, thereby improving the prediction accuracy of the large language model.
[0159] The SOP file generation system provided by the present invention is described below. The SOP file generation system described below and the SOP file generation method described above can be referenced to each other.
[0160] Please refer to Figure 4 , the SOP file generation system provided in this embodiment includes:
[0161] A basic data acquisition module 410 is used to acquire basic data, wherein the basic data includes: a process name and a process step name in the SOP file to be generated;
[0162] A working time category prediction module 420 is used to input the basic data into a preset large language model to perform working time category prediction, and obtain the working time category corresponding to the process name and the work step name in the basic data, wherein the working time category refers to the category of the action or operation to be performed corresponding to the work step;
[0163] The associated information acquisition module 430 is used to obtain SOP associated information corresponding to the working time category in the combined working time library based on the working time category and the preset combined working time library, wherein the SOP associated information includes operation content;
[0164] The integration module 440 is used to obtain the target SOP file by integrating the basic data, the working time category, and the SOP related information. The basic data acquisition module 410, the working time category prediction module 420, the related information acquisition module 430, and the integration module 440 are connected. The SOP file generation system in this embodiment can achieve the technical effect achieved by the SOP file generation method in the above embodiment, which will not be repeated here.
[0165] In some embodiments, the method further comprises: a model training module for acquiring a first data set, wherein the first data set comprises a plurality of first sample groups, wherein the first sample groups comprise a historical SOP file of any product, wherein the historical SOP file comprises: a process name, a process step name, and a work time category;
[0166] Preprocessing the first data set to obtain sample data, wherein the preprocessing includes: error data processing, null value data processing, and format conversion, wherein the format conversion refers to converting the historical SOP file that has undergone error data processing and null value data processing into the sample data according to a preset format, wherein the preset format includes: an instruction name, an input item, and an output item, wherein the instruction name is used to instruct the large language model to perform a work hour category prediction based on the input item;
[0167] The sample data is input into the large language model for back propagation to complete model training.
[0168] In some embodiments, the model training module is specifically used to associate the historical SOP file with the process flow file based on the process name and the work step name to obtain a target data table, wherein the primary key of the target data table is the process name and the work step name, and the value in the target data table is the work time category and the size information of the parts involved in the work step;
[0169] Extracting the part names in the process step names in the target data table, if the part names in two process step names are the same and the corresponding size information is the same, then determining that the process steps corresponding to the two current process step names are the same, and determining the tuples where the two current process step names are located as the tuples to be compared respectively;
[0170] If the working time categories in the two tuples to be compared are different, the two tuples to be compared are currently determined to be erroneous data, and based on the erroneous data, an erroneous data prompt is issued or the erroneous data is deleted.
[0171] In some embodiments, the model training module is also specifically used to determine the current tuple as null value data and delete the null value data if any tuple in the target data table includes a process name and a step name and lacks a working time category.
[0172] In some embodiments, the model training module is also specifically used to input the first sample data, the second sample data, and the third sample data into the large language model respectively to perform work hour category prediction to complete back propagation and obtain a trained large language model.
[0173] In some embodiments, the model training module is further specifically used to concatenate the instruction name and the input item in the sample data to obtain concatenated data;
[0174] Inputting the concatenated data into the input layer of the large language model to input the data to be predicted of fixed dimension;
[0175] Inputting the data to be predicted into the pre-trained language layer and the low-rank adaptation layer in the large language model respectively, performing work hour category prediction, and obtaining first prediction data output by the pre-trained language layer and second prediction data output by the low-rank adaptation layer, wherein the low-rank adaptation layer includes a dimensionality reduction matrix and a dimensionality increase matrix connected in sequence;
[0176] Determine the sum of the first prediction data and the second prediction data as final prediction data;
[0177] Based on the final prediction data and the corresponding output items in the sample data, the low-rank adaptation layer is back-propagated and updated until the model converges or the number of iterations reaches a preset number.
[0178] In some embodiments, the invention further includes: a model evaluation module, configured to obtain a second data set, wherein the second data set includes a plurality of second sample groups, and the products corresponding to the second sample groups are different from the products corresponding to the first sample groups;
[0179] Performing the preprocessing on the second data set to obtain test data, wherein the format of the test data is the same as the format of the sample data;
[0180] The test data is input into the large language model, and the large language model evaluation is completed based on the prediction result output by the large language model.
[0181] In some embodiments, it further includes: a combined work time library update module, which is used to calibrate the target SOP file based on a preset calibration rule, and determine the calibrated file as the final SOP file;
[0182] Inputting the work content in the final SOP file into the large language model, performing work hour category prediction, and obtaining a target work hour category;
[0183] If the target working time category is missing in the combined working time library, the target working time category and the work content corresponding to the target working time category are added to the combined working time library for the next SOP file generation.
[0184] In some embodiments, an electronic device is also provided, which may be a server, and its internal structure is shown in FIG. Figure 5 As shown. The electronic device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, the functions or steps on the server side of the above method are implemented.
[0185] In some embodiments, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program: obtaining basic data, the basic data including: the process name and the work step name in the SOP file to be generated; inputting the basic data into a preset large language model, performing work time category prediction, and obtaining the work time category corresponding to the process name and the work step name in the basic data, the work time category referring to the category of the action or operation to be performed corresponding to the work step; based on the work time category and a preset combined work time library, obtaining SOP associated information corresponding to the work time category in the combined work time library, the SOP associated information including the job content; obtaining the target SOP file by integrating the basic data, the work time category, and the SOP associated information.
[0186] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented: obtaining basic data, the basic data including: the process name and the work step name in the SOP file to be generated; inputting the basic data into a preset large language model, performing work time category prediction, and obtaining the work time category corresponding to the process name and the work step name in the basic data, the work time category refers to the category of the action or operation to be performed corresponding to the work step; based on the work time category and the preset combined work time library, obtaining the SOP associated information corresponding to the work time category in the combined work time library, the SOP associated information including the job content; obtaining the target SOP file by integrating the basic data, the work time category, and the SOP associated information.
[0187] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or electronic device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0188] The flow chart and block diagram in the accompanying drawings illustrate the possible implementation architecture, function and operation of the method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0189] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed by the present invention shall still be covered by the claims of the present invention.
Claims
1. A method for generating a SOP file, characterized in that: include: Obtain basic data, the basic data including: process name and step name in the SOP file to be generated; Input the basic data into a preset large language model to predict the working time category, and obtain the working time category corresponding to the process name and the work step name in the basic data, wherein the working time category refers to the category of the action or operation to be performed corresponding to the work step; Based on the working time category and a preset combined working time library, obtaining SOP association information corresponding to the working time category in the combined working time library, wherein the SOP association information includes operation content; The target SOP file is obtained by integrating the basic data, the working time category, and the SOP related information.
2. The SOP file generation method according to claim 1, characterized in that, Before obtaining basic data, it also includes: Acquire a first data set, the first data set includes a plurality of first sample groups, the first sample group includes a historical SOP file of any product, the historical SOP file includes: a process name, a process step name, and a work time category; Preprocessing the first data set to obtain sample data, wherein the preprocessing includes: error data processing, null value data processing, and format conversion, wherein the format conversion refers to converting the historical SOP file that has undergone error data processing and null value data processing into the sample data according to a preset format, wherein the preset format includes: an instruction name, an input item, and an output item, wherein the instruction name is used to instruct the large language model to perform a work hour category prediction based on the input item; The sample data is input into the large language model for back propagation to complete model training.
3. The SOP file generation method according to claim 2, characterized in that, The first sample group also includes a process flow file corresponding to the historical SOP file, wherein the process flow file includes: a process name, a step name, and size information of parts involved in each step; The error data processing includes: Based on the process name and the work step name, the historical SOP file and the process flow file are associated to obtain a target data table, wherein the primary key of the target data table is the process name and the work step name, and the value in the target data table is the work time category and the size information of the parts involved in the work step; Extracting the part names in the process step names in the target data table, if the part names in two process step names are the same and the corresponding size information is the same, then determining that the process steps corresponding to the two current process step names are the same, and determining the tuples where the two current process step names are located as the tuples to be compared respectively; If the working time categories in the two tuples to be compared are different, the two tuples to be compared are currently determined to be erroneous data, and based on the erroneous data, an erroneous data prompt is issued or the erroneous data is deleted.
4. The SOP file generation method according to claim 3, characterized in that, The null value data processing includes: If any tuple in the target data table includes a process name and a step name, and lacks a working time category, the current tuple is determined to be null value data, and the null value data is deleted.
5. The SOP file generation method according to claim 3 or 4, characterized in that, The historical SOP file further includes: work content corresponding to the work time category, and description information of the work time category, the value in the target data table also includes the work content and the description information, and the sample data includes: first sample data, second sample data, and third sample data; The input items of the first sample data are process name and step name, and the output item is the work time category; the input item of the second sample data is the work content, and the output item is the work time category; the input item of the third sample data is description information, and the input item is the work time category; The sample data is input into the large language model for back propagation to complete model training, including: inputting the first sample data, the second sample data, and the third sample data into the large language model respectively to perform work hour category prediction to complete back propagation and obtain a trained large language model.
6. The SOP file generation method according to any one of claims 2 to 4, characterized in that: Inputting the sample data into the large language model for back propagation to complete model training, including: splicing the instruction name and the input item in the sample data to obtain spliced data; Inputting the concatenated data into the input layer of the large language model to input the data to be predicted of fixed dimension; Inputting the data to be predicted into the pre-trained language layer and the low-rank adaptation layer in the large language model respectively, performing work hour category prediction, and obtaining first prediction data output by the pre-trained language layer and second prediction data output by the low-rank adaptation layer, wherein the low-rank adaptation layer includes a dimensionality reduction matrix and a dimensionality increase matrix connected in sequence; Determine the sum of the first prediction data and the second prediction data as final prediction data; Based on the final prediction data and the corresponding output items in the sample data, the low-rank adaptation layer is back-propagated and updated until the model converges or the number of iterations reaches a preset number.
7. The SOP file generation method according to any one of claims 2 to 4, characterized in that: After completing the training of the large language model, the method further includes: Acquire a second data set, where the second data set includes a plurality of second sample groups, and the products corresponding to the second sample groups are different from the products corresponding to the first sample groups; Performing the preprocessing on the second data set to obtain test data, wherein the format of the test data is the same as the format of the sample data; The test data is input into the large language model, and the large language model evaluation is completed based on the prediction result output by the large language model.
8. The SOP file generation method according to any one of claims 1 to 4, characterized in that: Also includes: Based on a preset calibration rule, the target SOP file is calibrated, and the calibrated file is determined as the final SOP file; Inputting the work content in the final SOP file into the large language model, performing work hour category prediction, and obtaining a target work hour category; If the target working time category is missing in the combined working time library, the target working time category and the work content corresponding to the target working time category are added to the combined working time library for the next SOP file generation.
9. A SOP file generation system, characterized in that: include: A basic data acquisition module is used to acquire basic data, wherein the basic data includes: a process name and a work step name in the SOP file to be generated; A working time category prediction module is used to input the basic data into a preset large language model to perform working time category prediction, and obtain the working time category corresponding to the process name and the work step name in the basic data, wherein the working time category refers to the category of the action or operation to be performed corresponding to the work step; A correlation information acquisition module, for obtaining SOP correlation information corresponding to the working time category in the combined working time library based on the working time category and a preset combined working time library, wherein the SOP correlation information includes operation content; The integration module is used to obtain a target SOP file by integrating the basic data, the working time category, and the SOP related information.
10. An electronic device, characterized in that: It comprises a processor, a memory and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute a computer program stored in the memory to implement the SOP file generation method as described in any one of claims 1 to 8.