A text content cleaning and annotation method and system based on a large model
By obtaining the identification and data processing mapping responsibility area of the big model server, determining the adaptive value sorting table and capability quantization information, and planning data processing scheduling strategies, the load balancing problem of distributed big model server under the influence of multiple factors is solved, and data processing efficiency and quality are improved.
Patent Information
- Application Number
- CN202510593037.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-09
AI Technical Summary
When using distributed large-scale model servers for text data cleaning and labeling, the load balancing method does not adapt to the influence of multiple factors, resulting in a decline in data processing efficiency and quality. Especially in different regional areas, the requirements of production data output units and unbalanced processing capabilities of large-scale model servers for specific fields have not been comprehensively considered.
By obtaining the server identification and data processing mapping responsibility area of the large model server, determining the adaptive value sorting table and capability quantization information in the data processing field, predicting production data output information and data transmission quality, and planning data processing scheduling strategies to achieve load balancing and improving data processing efficiency and quality.
Under the influence of multiple factors, the load balancing of large-model servers is achieved, which improves the efficiency and quality of data processing, adapts to the needs of production data output units within different regions, and improves the accuracy and real-timeness of data cleaning and labeling.
Smart Images

Figure CN120104979B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method and system for cleaning and annotating text content based on a large model. Background Art
[0002] Data cleaning and annotation are the core links of data preprocessing, and play a key role in improving data quality and mining data value. With the rapid development of artificial intelligence technology, the cleaning and annotation of text data have become increasingly important. Traditional text cleaning and annotation methods often rely on manual operations, which are not only inefficient but also prone to errors. With the continuous growth of data volume, manual cleaning and annotation can no longer meet the actual needs.
[0003] Currently, there are some platforms that can use large models to clean and annotate the text data provided by users and return the cleaned and annotated data to users. However, this solution still has the following limitations in practical applications: First, by using several large model servers distributed at the backend of the platform, the cleaning and annotation requirements of data within the scope of each large model server can be met. However, due to the different data processing requirements (including data processing volume and data processing type) of different production data output units within each scope at different times, and the different processing capabilities of each large model server, there may be peak pressures on data processing loads at different times. Traditional server load balancing methods cannot achieve satisfactory results in scenarios with the above multiple influencing factors. Second, since each large model server mainly serves the data processing tasks of several production data output units within its scope, each large model server needs to have stronger professionalism in data processing in a specific field, rather than in the entire field (usually, large language models trained in multiple fields have weak generalization ability. For example, sentiment words in e-commerce "negative reviews" may be misinterpreted according to medical semantics). For such scenarios, existing server load balancing solutions do not comprehensively consider the domain relevance of large models. Third, since the production task types of different production data output units may be different, their requirements for data processing latency will also be different (for example, data processing requests sent by artificial intelligence customer service in the e-commerce industry require high real-time performance, while the real-time performance required for analyzing and summarizing commercial data in the business field is not high). Such requirements will also have a greater impact on the implementation of server load balancing, thereby affecting the efficiency and quality of data cleaning and annotation in the overall area.
[0004] Therefore, how to improve the quality of data cleaning and annotation services provided by a distributed large model data processing architecture for a large number of production data output units in different areas, while achieving load balancing of large model servers under the influence of multiple factors in the adaptation scenario and improving the efficiency and quality of data processing, is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] The main object of the present invention is to provide a method and system for cleaning and annotating text content based on a large model, aiming to solve at least one of the above technical problems.
[0006] To achieve the above object, the present invention provides a method for cleaning and annotating text content based on a large model, the method comprising the following steps:
[0007] Obtain the server association information of each large model server in the distributed large model data processing architecture, and extract the server identifier and the data processing mapping responsibility area in the server association information;
[0008] Use the server identifier to determine the data processing domain adaptation value ranking table of each large model server and the data processing capacity quantization information of each processing cycle in the target processing period; wherein, the data processing capacity quantization information includes a first data processing capacity quantization value for text content cleaning and a second data processing capacity quantization value for text content annotation;
[0009] Use the data processing mapping responsibility area to access the production unit deployment location database, and query a number of production data output units that have an initial data processing matching relationship with each large model server;
[0010] According to the historical production database of each production data output unit, predict and generate the production data output information of each production data output unit in the target processing period; according to the historical network condition information of each production data output unit, predict the data transmission quality between each production data output unit and each large model server in each processing cycle in the target processing period;
[0011] Based on the production data output information of each production data output unit in the target processing period, the data processing domain adaptation value ranking table of each large model server, and the data processing capacity quantization information of each processing cycle in the target processing period, plan the production data output units assigned to each large model server in each processing cycle in the target processing period, and generate a data processing scheduling strategy;
[0012] According to the data processing scheduling strategy, control each large model server to perform the text content cleaning and annotation actions of the production data transmitted by the assigned production data output unit in each processing cycle in the target processing period.
[0013] Optionally, the step of using the server identifier to determine the data processing domain adaptation value ranking table of each large model server and the data processing capacity quantization information of each processing cycle in the target processing period specifically includes:
[0014] Using the server identifier, query the task running list and the server historical processing data information repository of each large model server, and determine the data processing capacity quantification information of each large model server in each processing cycle during the target processing period;
[0015] Using the server identifier, query the training process information of each large model server, extract the model training sample data in the training process information, analyze the model training sample data, and generate a sorting table of the processing data domain adaptation values of each large model server.
[0016] Optionally, the step of using the server identifier to query the task running list and the server historical processing data information repository of each large model server and determining the data processing capacity quantification information of each large model server in each processing cycle during the target processing period specifically includes:
[0017] Using the server identifier, query the task running list and the server historical processing data information repository of each large model server;
[0018] Extract the task execution period and the task execution data volume of each to-be-executed task in the task running list, divide the task execution period into each processing cycle in the target processing period, and determine the to-be-processed data volume of each large model server in each processing cycle during the target processing period;
[0019] Extract the processing data volume of each large model server in each processing cycle during the historical data processing process recorded in the server historical processing data information repository, and determine the standard processing data volume of each large model server in each processing cycle;
[0020] According to the difference between the to-be-processed data volume of each large model server in each processing cycle during the target processing period and the standard processing data volume in each processing cycle, determine the data processing capacity quantification information of each large model server in each processing cycle during the target processing period.
[0021] Optionally, the to-be-executed tasks include text content cleaning tasks and text content annotation tasks, the to-be-processed data volume includes the text content to-be-cleaned data volume and the text content to-be-annotated data volume, and the standard processing data volume includes the text content standard cleaning data volume and the text content standard annotation data volume.
[0022] Optionally, the step of determining the data processing capacity quantification information of each large model server in each processing cycle during the target processing period according to the difference between the to-be-processed data volume of each large model server in each processing cycle during the target processing period and the standard processing data volume in each processing cycle specifically includes:
[0023] Calculate the first data processing capacity quantization value for text content cleaning based on the difference between the amount of data to be cleaned and the standard amount of cleaned data for text content in each processing cycle of each large model server during the target processing period;
[0024] Calculate the second data processing capacity quantization value for text content cleaning based on the difference between the amount of data to be annotated and the standard amount of annotated data for text content in each processing cycle of each large model server during the target processing period;
[0025] Use the first data processing capacity quantization value for text content cleaning and the second data processing capacity quantization value for text content cleaning to determine the data processing capacity quantization information of each large model server in each processing cycle during the target processing period.
[0026] Optionally, the steps of querying the training process information of each large model server using the server identifier, extracting the model training sample data in the training process information, and analyzing the model training sample data to generate a sorting table of the processing data domain adaptation values of each large model server specifically include:
[0027] Query the training process information of each large model server using the server identifier, and extract the model training sample data in the training process information;
[0028] Extract several keywords from the model training sample data, and determine the sample training ratio of different processing data domains in the model training sample data according to the ratio of the sum of similarity values between the several keywords and the characteristic phrases corresponding to different processing data domains;
[0029] Sort the values of each processing data domain in the sample training ratio from large to small to generate a sorting table of the processing data domain fitness of each large model server.
[0030] Optionally, the steps of predicting and generating the production data output information of each production data output unit during the target processing period according to the historical production database of each production data output unit specifically include:
[0031] Query the historical production database of each production data output unit, and extract the production data output amount in the unit processing cycle when different types of production tasks are executed recorded in the historical production database;
[0032] Predict the production data output amount of each production data output unit during the target processing period according to the production task types of each production data output unit in different processing cycles during the target processing period;
[0033] Based on the production data output volume, the production data domain type corresponding to the production task type, and the production data transmission restriction set, construct the production data output information of each production data output unit in the target processing period; wherein, the production data transmission restriction set includes data integrity restriction and real-time restriction.
[0034] Optionally, according to the historical network condition information of each production data output unit, predict the data transmission quality steps between each production data output unit and each large model server in each processing cycle of the target processing period, specifically including:
[0035] Query the historical network condition information of each production data output unit, extract the historical data transmission quality parameters of each production data output unit and each large model server in the non-mapped responsibility area recorded in the historical network condition information, and construct a data transmission quality training sample containing several data transmission quality parameter features composed of data transmission timestamps and data transmission quality parameters;
[0036] Use the data transmission quality training sample to train the constructed initial convolutional neural network model. When the number of training times reaches the target number or the model converges, obtain the trained data transmission quality prediction model for each production data output unit and each large model server;
[0037] Use the data transmission quality prediction model to predict the data transmission quality parameters between each production data output unit and each large model server in the non-mapped responsibility area in each processing cycle of the target processing period; wherein, the data transmission quality parameters include data packet loss rate and data transmission delay.
[0038] Optionally, based on the production data output information of each production data output unit in the target processing period, the data processing domain fitness value ranking table of each large model server, and the data processing capacity quantification information in each processing cycle of the target processing period, plan the production data output units allocated to each large model server in each processing cycle of the target processing period, and generate data processing scheduling strategy steps, specifically including:
[0039] Based on the production data output information of each production data output unit in the target processing period, the data processing domain fitness value ranking table of each large model server, and the data processing capacity quantification information in each processing cycle of the target processing period;
[0040] Taking the sum of the production data output amounts of a number of production data output units allocated to each large model server in each processing cycle of the target processing period being simultaneously less than the first data processing capacity quantization value and the second data processing capacity quantization value of the large model server in the corresponding processing cycle as the first constraint condition, and taking the data transmission quality parameter corresponding to the data transmission quality of the production data output units belonging to the non-mapping responsibility area allocated to each large model server in each processing cycle of the target processing period satisfying the data allowable packet loss rate corresponding to the data integrity limitation in the data transmission and feedback process and the data allowable transmission delay corresponding to the real-time limitation as the second constraint condition, and taking the sum of the order values of the production data domain types in the production data output information of the production data output units allocated to all large model servers in each processing cycle of the target processing period in the processing data domain fitness ranking table of the large model server being the smallest as the optimization objective, and optimizing and solving the production data output units allocated to each large model server in each processing cycle of the target processing period;
[0041] Generate a data processing scheduling strategy based on the production data output units allocated to each large model server in each processing cycle of the target processing period.
[0042] In addition, to achieve the above object, the present invention also provides a text content cleaning and annotation system based on a large model, including:
[0043] An acquisition module, configured to acquire the server association information of each large model server in the distributed large model data processing architecture, and extract the server identifier and the data processing mapping responsibility area in the server association information;
[0044] A determination module, configured to use the server identifier to determine the data processing domain fitness ranking table of each large model server and the data processing capacity quantization information in each processing cycle in the target processing period; wherein, the data processing capacity quantization information includes a first data processing capacity quantization value for text content cleaning and a second data processing capacity quantization value for text content annotation;
[0045] A query module, configured to use the data processing mapping responsibility area to access the production unit deployment location database and query a number of production data output units having an initial data processing matching relationship with each large model server;
[0046] A prediction module, configured to predict and generate the production data output information of each production data output unit in the target processing period according to the historical production database of each production data output unit; and predict the data transmission quality between each production data output unit and each large model server in each processing cycle of the target processing period according to the historical network condition information of each production data output unit;
[0047] A generation module, configured to plan the production data output units assigned to each large model server in each processing cycle of the target processing period based on the production data output information of each production data output unit in the target processing period, the data processing domain fitness value ranking table of each large model server, and the data processing capacity quantization information of each processing cycle in the target processing period, and generate a data processing scheduling strategy.
[0048] An execution module, configured to control each large model server to perform the text content cleaning and annotation actions of the production data transmitted by the assigned production data output unit in each processing cycle of the target processing period according to the data processing scheduling strategy.
[0049] The beneficial effects of the present invention are as follows: A text content cleaning and annotation method and system based on a large model are proposed. By using the server identifier to determine the data processing domain fitness value ranking table and the data processing capacity quantization information, using the data processing mapping responsibility area to query the affiliated production data output unit, predicting and generating the production data output information and the data transmission quality, establishing a constraint condition set including the data processing capacity quantization information and the data transmission quality and an optimization target determined by the data processing domain fitness value ranking table, comprehensively considering the influence of factors such as the matching of data processing requirements and capabilities, the domain relevance of large model servers, and the data transmission quality, planning the data processing scheduling strategy in a distributed large model data processing architecture, and then controlling the execution of production data transmission and data cleaning and annotation, while achieving the load balancing of large model servers under the influence of multiple factors in the adaptation scenario, improving the efficiency and quality of data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a flowchart of the text content cleaning and annotation method based on a large model of the present invention;
[0051] Figure 2 It is a structural diagram of the text content cleaning and annotation system based on a large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] An embodiment of the present invention provides a text content cleaning and annotation method based on a large model, referring to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the text content cleaning and annotation method based on a large model of the present invention.
[0054] In this embodiment, a method for cleaning and annotating text content based on a large model, the method includes the following steps:
[0055] S100: Obtain the server association information of each large model server in the distributed large model data processing architecture, and extract the server identifier and the data processing mapping responsibility area in the server association information;
[0056] S200: Use the server identifier to determine the data processing domain adaptation value sorting table of each large model server and the data processing capacity quantization information of each processing cycle in the target processing period; wherein, the data processing capacity quantization information includes a first data processing capacity quantization value for text content cleaning and a second data processing capacity quantization value for text content annotation;
[0057] S300: Use the data processing mapping responsibility area to access the production unit deployment location database and query a number of production data output units that have an initial data processing matching relationship with each large model server;
[0058] S400: Predict and generate the production data output information of each production data output unit in the target processing period according to the historical production database of each production data output unit; predict the data transmission quality between each production data output unit and each large model server in each processing cycle in the target processing period according to the historical network condition information of each production data output unit;
[0059] S500: Based on the production data output information of each production data output unit in the target processing period, the data processing domain fitness value sorting table of each large model server, and the data processing capacity quantization information of each processing cycle in the target processing period, plan the production data output units allocated to each large model server in each processing cycle in the target processing period, and generate a data processing scheduling strategy;
[0060] S600: According to the data processing scheduling strategy, control each large model server to perform the text content cleaning and annotation actions of the production data transmitted by the allocated production data output unit in each processing cycle in the target processing period.
[0061] It should be noted that currently, there are some platforms that can use large models to clean and label the text data provided by users and return the cleaned and labeled data to the users. However, the following limitations still exist in the actual application of this solution: First, by using several large model servers set up distributively at the backend of the platform, the cleaning and labeling requirements of the data within the scope of each large model server can be met. However, due to the different data processing requirements (including the amount of data processed and the type of data processed) of different production data output units within the scope of each region at different times, and the different processing capabilities of each large model server, there may be peak pressure on the data processing load at different times. The traditional method of server load balancing cannot achieve satisfactory results in scenarios with the above multiple influencing factors. Second, since each large model server mainly serves the data processing tasks of several production data output units within its scope, each large model server needs to have stronger professionalism in data processing in a specific field, rather than in all fields (usually, large language models trained in multiple fields have the problem of weak generalization ability. For example, sentiment words in e-commerce "negative reviews" may be misinterpreted according to medical semantics). For such scenarios, the existing server load balancing solutions do not comprehensively consider the domain relevance of large models. Third, since the production task types of different production data output units may be different, their requirements for data processing latency will also be different (for example, the data processing requests sent by artificial intelligence customer service in the e-commerce industry require high real-time performance, while the real-time performance required for analyzing and summarizing business data in the business field is not high). Such requirements will also have a greater impact on the implementation of server load balancing, thereby affecting the efficiency and quality of data cleaning and labeling in the overall region.
[0062] To solve the above problems, in this embodiment, by using the server identifier to determine the data processing domain adaptation value ranking table and the data processing capacity quantification information, querying the production data output unit to which it belongs by using the data processing mapping responsibility area, predicting and generating the production data output information and data transmission quality, establishing a constraint condition set including the data processing capacity quantification information and the data transmission quality and an optimization target determined by the data processing domain adaptation value ranking table, comprehensively considering the influence of factors such as the matching of data processing requirements and capabilities, the domain relevance of large model servers, and the data transmission quality, planning the data processing scheduling strategy under the distributed large model data processing architecture, and then controlling the execution of production data transmission and data cleaning and labeling, while achieving the load balancing of large model servers under the influence of multiple factors in the adaptation scenario, improving the efficiency and quality of data processing.
[0063] In a preferred embodiment, the steps of using the server identifier to determine the data processing domain adaptation value ranking table of each large model server and the data processing capacity quantification information in each processing cycle during the target processing period specifically include:
[0064] S210: Query the task running list and the server historical processing data information library of each large model server using the server identifier, and determine the data processing capacity quantization information of each large model server for each processing cycle in the target processing period;
[0065] S220: Query the training process information of each large model server using the server identifier, extract the model training sample data in the training process information, analyze the model training sample data, and generate a sorting table of the processing data domain adaptation values of each large model server.
[0066] In this embodiment, after obtaining the server identifier of each large model server in the distributed large model data processing architecture, the server identifier can be used to query the data processing capacity quantization information of each large model server for each processing cycle in the target processing period, analyze the model training sample data, and generate a sorting table of the processing data domain adaptation values of each large model server, which can provide a set of constraint conditions for constructing an optimization algorithm when generating a subsequent data processing scheduling strategy.
[0067] On this basis, the step of querying the task running list and the server historical processing data information library of each large model server using the server identifier to determine the data processing capacity quantization information of each large model server for each processing cycle in the target processing period specifically includes:
[0068] S211: Query the task running list and the server historical processing data information library of each large model server using the server identifier;
[0069] S212: Extract the task execution period and the task execution data volume of each to-be-executed task in the task running list, divide the task execution period into each processing cycle in the target processing period, and determine the data volume to be processed by each large model server for each processing cycle in the target processing period;
[0070] S213: Extract the processing data volume of each large model server for each processing cycle in the historical data processing process recorded in the server historical processing data information library, and determine the standard processing data volume of each large model server for each processing cycle;
[0071] S214: Determine the data processing capacity quantization information of each large model server for each processing cycle in the target processing period according to the difference between the data volume to be processed by each large model server for each processing cycle in the target processing period and the standard processing data volume for each processing cycle.
[0072] In practical applications, the tasks to be executed include text content cleaning tasks and text content annotation tasks, the amount of data to be processed includes the amount of text content to be cleaned and the amount of text content to be annotated, and the standard amount of processed data includes the standard amount of text content cleaned and the standard amount of text content annotated.
[0073] Furthermore, the step of determining the quantitative information of the data processing capacity of each large model server in each processing cycle during the target processing period according to the difference between the amount of data to be processed and the standard amount of processed data in each processing cycle of each large model server specifically includes:
[0074] S2141: Calculate the first quantitative value of data processing capacity for text content cleaning according to the difference between the amount of text content to be cleaned and the standard amount of text content cleaned in each processing cycle of each large model server during the target processing period;
[0075] S2142: Calculate the second quantitative value of data processing capacity for text content annotation according to the difference between the amount of text content to be annotated and the standard amount of text content annotated in each processing cycle of each large model server during the target processing period;
[0076] S2143: Use the first quantitative value of data processing capacity for text content cleaning and the second quantitative value of data processing capacity for text content annotation to determine the quantitative information of the data processing capacity of each large model server in each processing cycle during the target processing period.
[0077] In this embodiment, first, the amount of data to be processed in each processing cycle of each large model server is queried using the server identifier, and then according to the standard amount of processed data of each large model server recorded in the historical processing data information library (this standard amount of processed data can be measured by the maximum amount of processed data that does not affect the normal operation of the server recorded in the historical processing data information library for each processing cycle). After that, by considering the text content cleaning task and the text content annotation task, the difference between the two is calculated to obtain the first quantitative value of data processing capacity for text content cleaning and the second quantitative value of data processing capacity for text content annotation, and finally, the quantitative information of the data processing capacity of each large model server in each processing cycle during the target processing period is obtained.
[0078] In a preferred embodiment, the step of using the server identifier to query the training process information of each large model server, extracting the model training sample data in the training process information, and analyzing the model training sample data to generate a sorting table of the adaptation values of the processing data fields of each large model server specifically includes:
[0079] S221: Query the training process information of each large model server using the server identifier, and extract the model training sample data in the training process information;
[0080] S222: Extract several keywords from the model training sample data, and determine the sample training ratio of different data processing fields in the model training sample data according to the ratio of the sum of similarity values between the several keywords and the characteristic phrases corresponding to different data processing fields;
[0081] S223: Sort the values of each data processing field in the sample training ratio from large to small to generate a fitness ranking table of the data processing fields of each large model server.
[0082] In this embodiment, secondly, query the training process information of each large model server using the server identifier, and determine the specific fields for each large model server to learn more deeply by analyzing the ratio of the sample data fields in the model training sample data, so as to serve as the professionalism for different data processing fields, and generate a fitness ranking table of the data processing fields of each large model server, which is used as the optimization objective in the subsequent construction of the optimization algorithm, so that the entire large model server architecture adopts the optimal data processing task allocation method for the field, and improves the overall data cleaning and annotation accuracy.
[0083] In a preferred embodiment, the steps of predicting the production data output information of each production data output unit in the target processing period according to the historical production database of each production data output unit specifically include:
[0084] S410: Query the historical production database of each production data output unit, and extract the production data output volume in the unit processing cycle when different types of production tasks are executed recorded in the historical production database;
[0085] S420: Predict the production data output volume of each production data output unit in the target processing period according to the production task types in different processing cycles of each production data output unit in the target processing period;
[0086] S430: Based on the production data output volume, the production data field type corresponding to the production task type, and the production data transmission restriction set, construct the production data output information of each production data output unit in the target processing period; wherein, the production data transmission restriction set includes data integrity restriction and real-time restriction.
[0087] On this basis, the steps of predicting the data transmission quality between each production data output unit and each large model server in each processing cycle of the target processing period according to the historical network condition information of each production data output unit specifically include:
[0088] S440: Query the historical network status information of each production data output unit, extract the historical data transmission quality parameters of each production data output unit and each large model server in the non-mapped responsibility area recorded in the historical network status information, and construct a data transmission quality training sample containing several data transmission quality parameter features composed of data transmission timestamps and data transmission quality parameters;
[0089] S450: Use the data transmission quality training sample to train the constructed initial convolutional neural network model. When the number of training times reaches the target number or the model converges, obtain a trained data transmission quality prediction model for each production data output unit and each large model server;
[0090] S460: Use the data transmission quality prediction model to predict the data transmission quality parameters of each production data output unit in each processing cycle and each large model server in the non-mapped responsibility area during the target processing period; wherein, the data transmission quality parameters include data packet loss rate and data transmission delay.
[0091] In this embodiment, considering that the data processing requirements of different production data output units within each affiliated area may be different at different times (including the amount of data processed and the type of data processed), and the production task types of different production data output units may be different, the requirements for data processing delay will also be different (for example, the data processing requests sent by artificial intelligence customer service in the e-commerce industry require high real-time performance, while the real-time performance required for analyzing and summarizing business data in the business field is not high). By querying the historical production database of each production data output unit, predict the production data output volume of each production data output unit during the target processing period, and construct production data output information including the production data output volume, the production data domain type corresponding to the production task type, and the production data transmission restriction set. Then, by querying the historical network status information of each production data output unit, use the data transmission quality prediction model to predict the data packet loss rate and data transmission delay of each production data output unit in each processing cycle and each large model server in the non-mapped responsibility area during the target processing period, which can provide a constraint condition set for constructing an optimization algorithm for the subsequent generation of data processing scheduling strategies. Through the construction of multiple constraints, the rationality of the data processing scheduling strategy and the overall data cleaning and annotation accuracy of multiple production data output units within the area can be improved.
[0092] In a preferred embodiment, based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server, and the data processing capacity quantization information of each processing cycle in the target processing period, plan the production data output units allocated to each processing cycle of each large model server in the target processing period, and generate a data processing scheduling strategy step, which specifically includes:
[0093] S510: Based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server, and the data processing capacity quantization information of each processing cycle in the target processing period;
[0094] S520: Take the sum of the production data output amounts of several production data output units allocated to each processing cycle of each large model server in the target processing period being less than the first data processing capacity quantization value and the second data processing capacity quantization value of the large model server in the corresponding processing cycle as the first constraint condition, and take the data transmission quality parameter corresponding to the data transmission quality of the production data output units belonging to the non-mapped responsibility area allocated to each processing cycle of each large model server in the target processing period satisfying the data integrity limit corresponding to the data transmission and feedback process and the data allowable packet loss rate corresponding to the real-time limit and the data allowable transmission delay as the second constraint condition. Take the minimum sum of the order values of the production data field types in the production data output information of the production data output units allocated to each processing cycle of all large model servers in the target processing period in the processing data field fitness ranking table of the large model server as the optimization objective, and optimize and solve the production data output units allocated to each processing cycle of each large model server in the target processing period;
[0095] S530: Generate a data processing scheduling strategy based on the production data output units allocated to each processing cycle of each large model server in the target processing period.
[0096] In this embodiment, by obtaining the server identifier and the data processing mapping responsibility area of each large model server in the distributed large model data processing architecture, using the server identifier to determine the data processing domain adaptation value ranking table and the data processing capacity quantification information, using the data processing mapping responsibility area to query the affiliated production data output unit, predicting and generating the production data output information and the data transmission quality, by establishing a constraint condition set including the data processing capacity quantification information and the data transmission quality and an optimization target determined based on the data processing domain adaptation value ranking table, and adopting an optimization algorithm to plan the production data output unit allocated to each processing cycle of each large model server in the target processing period, and then generating a data processing scheduling strategy and controlling the production data transmission of each production data output unit and the data cleaning and annotation of each large model server. Thus, by comprehensively considering the matching of data processing requirements and capabilities, the domain relevance of large model servers, and the impact of factors such as data transmission quality, planning the data processing scheduling strategy under the distributed large model data processing architecture can improve the quality of data cleaning and annotation services provided by the distributed large model data processing architecture to a large number of production data output units in different regional ranges, and while achieving the load balancing of large model servers under the influence of multiple factors in the adaptation scenario, improve the efficiency and quality of data processing.
[0097] Refer to Figure 2 , Figure 2 which is the structural block diagram of the embodiment of the text content cleaning and annotation system based on the large model of the present invention.
[0098] As Figure 2 shown, the text content cleaning and annotation system based on the large model proposed in the embodiment of the present invention includes:
[0099] An acquisition module 10, configured to acquire the server association information of each large model server in the distributed large model data processing architecture, and extract the server identifier and the data processing mapping responsibility area from the server association information;
[0100] A determination module 20, configured to use the server identifier to determine the data processing domain adaptation value ranking table of each large model server and the data processing capacity quantification information of each processing cycle in the target processing period; wherein, the data processing capacity quantification information includes a first data processing capacity quantification value for text content cleaning and a second data processing capacity quantification value for text content annotation;
[0101] A query module 30, configured to use the data processing mapping responsibility area to access the production unit deployment location database and query a plurality of production data output units having an initial data processing matching relationship with each large model server;
[0102] A prediction module 40, configured to predict and generate production data output information of each production data output unit in a target processing period according to the historical production database of each production data output unit; and predict the data transmission quality between each production data output unit and each large model server in each processing cycle of the target processing period according to the historical network status information of each production data output unit.
[0103] A generation module 50, configured to plan the production data output units assigned to each large model server in each processing cycle of the target processing period based on the production data output information of each production data output unit in the target processing period, the sorting table of data processing domain fitness values of each large model server, and the quantified information of data processing capabilities in each processing cycle of the target processing period, and generate a data processing scheduling strategy.
[0104] An execution module 60, configured to control each large model server to perform text content cleaning and annotation actions on the production data transmitted by the assigned production data output units in each processing cycle of the target processing period according to the data processing scheduling strategy.
[0105] For other embodiments or specific implementation manners of the text content cleaning and annotation system based on a large model of the present invention, reference may be made to the above method embodiments, which will not be elaborated herein.
[0106] It can be understood that in the description of this specification, the descriptions referring to terms such as "one embodiment", "another embodiment", "other embodiments", or "the first embodiment to the Nth embodiment" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0107] It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.
[0108] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A method for cleaning and annotating text content based on a large model, characterized in that, The method includes the following steps: Obtain the server association information of each large model server in the distributed large model data processing architecture, and extract the server identifier and the data processing mapping responsibility area from the server association information; Use the server identifier to determine the data processing domain fitness value sorting table of each large model server and the data processing capacity quantization information for each processing cycle in the target processing period; wherein, the data processing capacity quantization information includes a first data processing capacity quantization value for text content cleaning and a second data processing capacity quantization value for text content annotation; Use the data processing mapping responsibility area to access the production unit deployment location database and query a number of production data output units that have an initial data processing matching relationship with each large model server; Predict and generate the production data output information of each production data output unit in the target processing period according to the historical production database of each production data output unit; predict the data transmission quality between each production data output unit and each large model server for each processing cycle in the target processing period according to the historical network condition information of each production data output unit; Based on the production data output information of each production data output unit in the target processing period, the data processing domain fitness value sorting table of each large model server, and the data processing capacity quantization information for each processing cycle in the target processing period, plan the production data output units assigned to each large model server for each processing cycle in the target processing period, and generate a data processing scheduling strategy; Specifically, it includes: based on the production data output information of each production data output unit in the target processing period, the data processing domain fitness value sorting table of each large model server, and the data processing capacity quantization information for each processing cycle in the target processing period; Taking the sum of the production data output amounts of the several production data output units assigned to each large model server for each processing cycle in the target processing period being less than the first data processing capacity quantization value and the second data processing capacity quantization value of the large model server in the corresponding processing cycle as the first constraint condition, taking the data transmission quality parameter corresponding to the data transmission quality between the production data output units belonging to the non-mapping responsibility area assigned to each large model server for each processing cycle in the target processing period and the large model server satisfying the data allowable packet loss rate corresponding to the data integrity limit in the data transmission and feedback process and the data allowable transmission delay corresponding to the real-time limit as the second constraint condition, and taking the sum of the order values of the production data domain types in the production data output information of the production data output units assigned to all large model servers for each processing cycle in the target processing period in the processing data domain fitness sorting table of the large model server being the smallest as the optimization objective, and optimizing and solving the production data output units assigned to each large model server for each processing cycle in the target processing period; Generate a data processing scheduling strategy based on the production data output units assigned to each large model server for each processing cycle in the target processing period; According to the above data processing scheduling strategy, control each large model server to perform the text content cleaning and annotation actions of the production data transmitted by the assigned production data output unit in each processing cycle during the target processing period.
2. The method for cleaning and annotating text content based on a large model according to claim 1, wherein The steps of using the server identifier to determine the data processing domain adaptation value sorting table of each large model server and the data processing capacity quantization information in each processing cycle during the target processing period specifically include: Using the server identifier, query the task running list and the server historical processing data information library of each large model server to determine the data processing capacity quantization information of each large model server in each processing cycle during the target processing period; Using the server identifier, query the training process information of each large model server, extract the model training sample data in the training process information, analyze the model training sample data, and generate the data processing domain adaptation value sorting table of each large model server.
3. The method for cleaning and annotating text content based on a large model according to claim 2, wherein The steps of using the server identifier to query the task running list and the server historical processing data information library of each large model server to determine the data processing capacity quantization information of each large model server in each processing cycle during the target processing period specifically include: Using the server identifier, query the task running list and the server historical processing data information library of each large model server; Extract the task execution period and the task execution data volume of each to-be-executed task in the task running list, divide the task execution period into each processing cycle during the target processing period, and determine the to-be-processed data volume of each large model server in each processing cycle during the target processing period; Extract the processed data volume of each large model server in each processing cycle during the historical data processing recorded in the server historical processing data information library, and determine the standard processed data volume of each large model server in each processing cycle; According to the difference between the to-be-processed data volume of each large model server in each processing cycle during the target processing period and the standard processed data volume in each processing cycle, determine the data processing capacity quantization information of each large model server in each processing cycle during the target processing period.
4. The method for cleaning and annotating text content based on a large model according to claim 3, wherein, The to-be-executed tasks include text content cleaning tasks and text content annotation tasks, the to-be-processed data volume includes the text content to-be-cleaned data volume and the text content to-be-annotated data volume, and the standard processed data volume includes the text content standard cleaning data volume and the text content standard annotation data volume.
5. The method for cleaning and annotating text content based on a large model according to claim 4, wherein, The steps of determining the data processing capacity quantization information of each large model server in each processing cycle during the target processing period according to the difference between the to-be-processed data volume of each large model server in each processing cycle during the target processing period and the standard processed data volume in each processing cycle specifically include: Calculate the first data processing capacity quantization value for text content cleaning according to the difference between the text content to-be-cleaned data volume of each large model server in each processing cycle during the target processing period and the text content standard cleaning data volume in each processing cycle; Calculate the second data processing capacity quantization value for text content cleaning according to the difference between the amount of data to be labeled of the text content in each processing cycle and the standard labeled data amount of the text content in each processing cycle during the target processing period for each large model server; Use the first data processing capacity quantization value for text content cleaning and the second data processing capacity quantization value for text content cleaning to determine the data processing capacity quantization information of each large model server in each processing cycle during the target processing period.
6. The method for cleaning and annotating text content based on a large model according to claim 2, wherein, The steps of using the server identifier to query the training process information of each large model server, extracting the model training sample data in the training process information, and analyzing the model training sample data to generate the adaptation value ranking table of the processing data fields of each large model server specifically include: Use the server identifier to query the training process information of each large model server and extract the model training sample data in the training process information; Extract several keywords from the model training sample data, and determine the sample training ratio of different processing data fields in the model training sample data according to the ratio of the sum of similarity values between the several keywords and the characteristic phrases corresponding to different processing data fields; Sort the values of each processing data field in the sample training ratio from large to small to generate the adaptation degree ranking table of the processing data fields of each large model server.
7. The method for cleaning and annotating text content based on a large model according to claim 1, wherein The steps of predicting and generating the production data output information of each production data output unit in the target processing period according to the historical production database of each production data output unit specifically include: Query the historical production database of each production data output unit and extract the production data output amount in the unit processing cycle when different types of production tasks are executed recorded in the historical production database; Predict the production data output amount of each production data output unit in the target processing period according to the production task types of each production data output unit in different processing cycles in the target processing period; Based on the production data output amount, the production data field type corresponding to the production task type, and the production data transmission restriction set, construct the production data output information of each production data output unit in the target processing period; wherein, the production data transmission restriction set includes data integrity restriction and real-time restriction.
8. The method for cleaning and annotating text content based on a large model according to claim 1, wherein The steps of predicting the data transmission quality between each production data output unit and each large model server in each processing cycle during the target processing period according to the historical network condition information of each production data output unit specifically include: Query the historical network condition information of each production data output unit, extract the historical data transmission quality parameters of each production data output unit and each large model server in the non-mapped responsibility area recorded in the historical network condition information, and construct a data transmission quality training sample containing several data transmission quality parameter features composed of data transmission timestamps and data transmission quality parameters; Training the constructed initial convolutional neural network model with the data transmission quality training samples, and obtaining a trained data transmission quality prediction model for each production data output unit and each large model server when the number of training times reaches the target number or the model converges; Using the data transmission quality prediction model to predict the data transmission quality parameters of each production data output unit and each large model server in each processing cycle during the target processing period in each non-mapped responsibility area; wherein, the data transmission quality parameters include data packet loss rate and data transmission delay.
9. A text content cleaning and annotation system based on a large model, characterized in that, Including: An acquisition module, configured to acquire the server association information of each large model server in the distributed large model data processing architecture, and extract the server identifier and the data processing mapping responsibility area in the server association information; A determination module, configured to use the server identifier to determine the data processing domain fitness value ranking table of each large model server and the data processing capacity quantization information in each processing cycle during the target processing period; wherein, the data processing capacity quantization information includes a first data processing capacity quantization value for text content cleaning and a second data processing capacity quantization value for text content annotation; A query module, configured to use the data processing mapping responsibility area to access the production unit deployment location database and query a plurality of production data output units that have an initial data processing matching relationship with each large model server; A prediction module, configured to predict and generate the production data output information of each production data output unit during the target processing period according to the historical production database of each production data output unit; and predict the data transmission quality of each production data output unit and each large model server in each processing cycle during the target processing period according to the historical network condition information of each production data output unit; A generation module, configured to plan the production data output units allocated to each large model server in each processing cycle during the target processing period based on the production data output information of each production data output unit during the target processing period, the data processing domain fitness value ranking table of each large model server, and the data processing capacity quantization information in each processing cycle during the target processing period, and generate a data processing scheduling strategy; Specifically including: based on the production data output information of each production data output unit during the target processing period, the data processing domain fitness value ranking table of each large model server, and the data processing capacity quantization information in each processing cycle during the target processing period; Taking the sum of the production data output amounts of several production data output units assigned to each large model server in each processing cycle of the target processing period being simultaneously less than the first data processing capacity quantization value and the second data processing capacity quantization value of the large model server in the corresponding processing cycle as the first constraint condition, taking the data transmission quality parameter corresponding to the data transmission quality of the production data output units belonging to the non-mapped responsibility area assigned to each large model server in each processing cycle of the target processing period satisfying the data allowable packet loss rate corresponding to the data integrity limit in the data transmission and feedback process and the data allowable transmission delay corresponding to the real-time limit as the second constraint condition, and taking the sum of the order values of the production data domain types in the production data output information of the production data output units assigned to all large model servers in each processing cycle of the target processing period in the processing data domain fitness ranking table of the large model server being the smallest as the optimization objective, and optimizing and solving the production data output units assigned to each large model server in each processing cycle of the target processing period; Generating a data processing scheduling strategy based on the production data output units assigned to each large model server in each processing cycle of the target processing period; An execution module, configured to control each large model server to perform text content cleaning and annotation actions on the production data transmitted by the assigned production data output units in each processing cycle of the target processing period according to the data processing scheduling strategy.
Citation Information
Patent Citations
Theme web crawler method and device and medium
CN110069690A
Intelligent data labeling method and system
CN119647479A