Text content cleaning and labeling method and system based on large model

By obtaining and analyzing the server-related information and data processing mapping responsibility areas on the big model server, predicting the production data output information and data transmission quality, comprehensively considering the influence of multiple factors, and planning data processing scheduling strategies, the problems of load peak pressure, domain correlation, and different data processing delay requirements in the distributed big model data processing architecture are solved, and efficient and balanced data cleaning and labeling services are achieved.

CN120104979AActive Publication Date: 2025-06-06GONGYEYUN MFG (SICHUAN) INNOVATION CENT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510593037.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-06
Estimated Expiration
2045-05-09

Smart Images

  • Figure CN120104979A_ABST
    Figure CN120104979A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a text content cleaning and labeling method and system based on a large model, and the method comprises the steps: determining a data processing field adaptive value sorting table and data processing capability quantification information through server identification, querying a belonging production data output unit through a data processing mapping responsibility region, prediction is carried out to generate production data output information and data transmission quality, and a constraint condition set containing data processing capability quantification information and data transmission quality and an optimization target determined by a data processing field adaptive value sorting table are established; influences of factors such as data processing requirements and capability matching, large model server field correlation and data transmission quality are comprehensively considered, a data processing scheduling strategy under a distributed large model data processing architecture is planned, and then production data transmission and execution of data cleaning and labeling are controlled. The data processing efficiency and quality are improved while large model server load balancing adapting to scene multi-factor influence is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a large model-based text content cleaning and annotation method and system. Background Art

[0002] Data cleaning and labeling are the core links of data preprocessing, which play a key role in improving data quality and mining data value. With the rapid development of artificial intelligence technology, the cleaning and labeling of text data has become increasingly important. Traditional text cleaning and labeling methods often rely on manual operations, which are not only inefficient but also prone to errors. With the continuous growth of data volume, manual cleaning and labeling can no longer meet actual needs.

[0003] Currently, some platforms can use large models to clean and annotate text data provided by users, and return the cleaned and annotated data to users. However, this solution still has the following limitations in actual applications; first, using a number of large model servers distributed on the back end of the platform can solve the needs of data cleaning and annotation within the area to which each large model server belongs. However, due to the different data processing requirements of different production data output units in each area at different times (including data processing volume and data processing type), the processing capacity of each large model server is also different. Data processing load peak pressure may occur at different times. Traditional server load balancing methods cannot achieve satisfactory results under the above-mentioned multiple influencing factors; second, since each large model server mainly serves the data processing tasks of several production data output units within the area to which it belongs, each large model server needs to have stronger professionalism in data processing in a specific field, rather than in the entire field (usually large language models trained in multiple fields have weak generalization capabilities, such as mistakenly treating e-commerce "bad reviews" as "bad reviews"). The sentiment words in the video are understood as medical semantics). For such scenarios, the existing server load balancing solutions have not yet comprehensively considered the domain relevance of large models. Third, since the production task types of different production data output units may be different, their requirements for data processing latency will also be different (for example, the data processing requests sent by artificial intelligence customer service in the e-commerce industry require high real-time performance, while the real-time performance required for analyzing and summarizing business data in the commercial field is not high). Such requirements will also have a great impact on the implementation of server load balancing, thereby affecting the efficiency and quality of data cleaning and annotation within the overall area.

[0004] Therefore, how to improve the quality of data cleaning and annotation services provided to a large number of production data output units in different regions using a distributed large-model data processing architecture, and improve the efficiency and quality of data processing while achieving load balancing of large-model servers under the influence of multiple factors in the scene, is a technical problem that needs to be solved urgently. Summary of the invention

[0005] The main purpose of the present invention is to provide a text content cleaning and annotation method and system based on a large model, aiming to solve at least one of the above technical problems.

[0006] To achieve the above object, the present invention provides a text content cleaning and annotation method based on a large model, the method comprising the following steps: Obtaining server association information of each large model server in the distributed large model data processing architecture, and extracting the server identifier and data processing mapping responsibility area in the server association information; Using the server identifier, determine a data processing field fitness value ranking table of each large model server and data processing capacity quantification information of each processing cycle in a target processing period; wherein the data processing capacity quantification information includes a first data processing capacity quantification value for text content cleaning and a second data processing capacity quantification value for text content annotation; Using the data processing mapping responsibility area, accessing the production unit deployment location database, querying a number of production data output units that have an initial data processing matching relationship with each large model server; According to the historical production database of each production data output unit, the production data output information of each production data output unit in the target processing period is predicted and generated; according to the historical network status information of each production data output unit, the data transmission quality between each production data output unit and each large model server in each processing cycle in the target processing period is predicted; Based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server and the quantitative information of the data processing capacity of each processing cycle in the target processing period, the production data output unit allocated to each large model server in each processing cycle of the target processing period is planned to generate a data processing scheduling strategy; According to the data processing scheduling strategy, each large model server is controlled to perform text content cleaning and annotation actions of the production data transmitted by the assigned production data output unit in each processing cycle of the target processing period.

[0007] Optionally, the step of using the server identifier to determine a data processing field fitness value ranking table of each large model server and quantitative information of data processing capacity of each processing cycle in a target processing period specifically includes: Using the server identifier, query the task running list and server historical processing data information base of each large model server to determine the quantitative information of the data processing capacity of each large model server in each processing cycle in the target processing period; The server identifier is used to query the training process information of each large model server, extract the model training sample data in the training process information, analyze the model training sample data, and generate a processing data domain fitness value ranking table for each large model server.

[0008] Optionally, using the server identifier, querying the task running list and the server historical processing data information base of each large model server to determine the quantitative information step of the data processing capacity of each large model server in each processing cycle in the target processing period, specifically includes: Using the server identifier, query the task running list and server historical processing data information base of each large model server; Extracting the task execution period and the task execution data volume of each to-be-executed task in the task running list, dividing the task execution period into each processing cycle in the target processing period, and determining the amount of to-be-processed data of each large model server in each processing cycle in the target processing period; Extracting the amount of data processed by each large model server in each processing cycle during the historical data processing process recorded in the server historical processing data information library, and determining the standard amount of data processed by each large model server in each processing cycle; According to the difference between the amount of data to be processed in each processing cycle of each large model server in the target processing period and the standard amount of data processed in each processing cycle, the quantitative information of the data processing capacity of each large model server in each processing cycle in the target processing period is determined.

[0009] Optionally, the tasks to be executed include text content cleaning tasks and text content annotation tasks, the amount of data to be processed includes the amount of text content to be cleaned and the amount of text content to be annotated, and the standard processing data includes the amount of text content standard cleaned data and the amount of text content standard annotated data.

[0010] Optionally, the step of determining the quantitative information of the data processing capacity of each large model server in each processing cycle in the target processing period according to the difference between the amount of data to be processed in each processing cycle of each large model server in the target processing period and the standard amount of data processed in each processing cycle specifically includes: Calculate a first data processing capacity quantification value for text content cleaning according to the difference between the amount of text content to be cleaned data in each processing cycle of each large model server in the target processing period and the amount of text content standard cleaned data in each processing cycle; Calculate a second data processing capacity quantification value for text content cleaning according to the difference between the amount of text content to-be-annotated data in each processing cycle of each large model server in the target processing period and the amount of text content standard annotated data in each processing cycle; The data processing capacity quantification information of each large model server in each processing cycle in the target processing period is determined by using the first data processing capacity quantification value for text content cleaning and the second data processing capacity quantification value for text content cleaning.

[0011] Optionally, using the server identifier, querying the training process information of each large model server, extracting the model training sample data in the training process information, analyzing the model training sample data, and generating a processing data domain fitness value ranking table for each large model server, specifically includes: Using the server identifier, query the training process information of each large model server, and extract the model training sample data in the training process information; Extracting a number of keywords from the model training sample data, and determining the sample training ratios of different processing data fields in the model training sample data according to the ratio of the cumulative sum of similarity values ​​between the keywords and the feature phrases corresponding to different processing data fields; The values ​​of each processing data field in the sample training ratio are sorted from large to small to generate a processing data field fitness ranking table for each large model server.

[0012] Optionally, the step of predicting and generating production data output information of each production data output unit in a target processing period according to a historical production database of each production data output unit specifically includes: Querying the historical production database of each production data output unit, and extracting the production data output within a unit processing cycle when different types of production tasks are executed, which are recorded in the historical production database; Predicting the output volume of production data of each production data output unit in the target processing period according to the production task types of each production data output unit in different processing cycles in the target processing period; Based on the production data output volume, the production data field type corresponding to the production task type and the production data transmission restriction set, the production data output information of each production data output unit in the target processing period is constructed; wherein the production data transmission restriction set includes data integrity restriction and real-time restriction.

[0013] Optionally, according to the historical network status information of each production data output unit, the step of predicting the data transmission quality between each production data output unit and each large model server in each processing cycle during the target processing period specifically includes: Query the historical network status information of each production data output unit, extract the historical data transmission quality parameters between each production data output unit and each large model server in the non-mapped responsibility area recorded in the historical network status information, and construct a data transmission quality training sample containing a plurality of data transmission quality parameter features consisting of a data transmission timestamp and a data transmission quality parameter; The constructed initial convolutional neural network model is trained using the data transmission quality training samples, and when the number of training times reaches the target number of times or the model converges, a trained data transmission quality prediction model for each production data output unit and each large model server is obtained; The data transmission quality prediction model is used to predict the data transmission quality parameters of each production data output unit and each large model server in each non-mapped responsibility area in each processing cycle of the target processing period; wherein the data transmission quality parameters include data packet loss rate and data transmission delay.

[0014] Optionally, based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server, and the quantitative information of the data processing capacity of each processing cycle in the target processing period, the production data output unit allocated to each large model server in each processing cycle of the target processing period is planned, and the data processing scheduling strategy step is generated, which specifically includes: Based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server and the data processing capacity quantification information of each processing cycle in the target processing period; The first constraint condition is that the sum of the production data output amounts of several production data output units allocated to each large model server in each processing cycle of the target processing period is less than the first data processing capacity quantification value and the second data processing capacity quantification value of the large model server in the corresponding processing cycle at the same time; the second constraint condition is that the data transmission quality parameter corresponding to the data transmission quality of the production data output unit belonging to the non-mapped responsibility area allocated to each large model server in each processing cycle of the target processing period and the large model server satisfies the data transmission quality of the data integrity restriction and the data transmission delay corresponding to the real-time restriction of the data transmission and return process; the optimization target is that the sum of the order values ​​of the production data field type in the production data output information of the production data output units allocated to all large model servers in each processing cycle of the target processing period in the processing data field fitness ranking table of the large model server is minimized, and the production data output unit allocated to each large model server in each processing cycle of the target processing period is optimized; A data processing scheduling strategy is generated based on the production data output units allocated to each large model server in each processing cycle during the target processing period.

[0015] In addition, in order to achieve the above-mentioned purpose, the present invention also provides a text content cleaning and annotation system based on a large model, comprising: An acquisition module, used to acquire server association information of each large model server in the distributed large model data processing architecture, and extract the server identifier and data processing mapping responsibility area in the server association information; A determination module, used to determine the data processing field fitness value ranking table of each large model server and the data processing capacity quantification information of each processing cycle in the target processing period by using the server identifier; wherein the data processing capacity quantification information includes a first data processing capacity quantification value for text content cleaning and a second data processing capacity quantification value for text content annotation; A query module, for accessing the production unit deployment location database using the data processing mapping responsibility area, and querying a number of production data output units having an initial data processing matching relationship with each large model server; A prediction module is used to predict and generate production data output information of each production data output unit in a target processing period according to the historical production database of each production data output unit; and to predict the data transmission quality of each production data output unit and each large model server in each processing cycle of the target processing period according to the historical network status information of each production data output unit; A generation module is used to plan the production data output units allocated to each large model server in each processing cycle of the target processing period based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server, and the quantitative information of the data processing capacity of each processing cycle in the target processing period, and generate a data processing scheduling strategy; The execution module is used to control each large model server to execute the text content cleaning and annotation actions of the production data transmitted by the assigned production data output unit in each processing cycle of the target processing period according to the data processing scheduling strategy.

[0016] The beneficial effects of the present invention are as follows: a text content cleaning and annotation method and system based on a large model are proposed, by using a server identifier to determine a data processing field fitness value ranking table and data processing capacity quantification information, using a data processing mapping responsibility area to query the corresponding production data output unit, predicting and generating production data output information and data transmission quality, establishing a constraint condition set including data processing capacity quantification information and data transmission quality and an optimization target determined by a data processing field fitness value ranking table, comprehensively considering the influence of factors such as data processing demand and capacity matching, large model server field relevance and data transmission quality, planning a data processing scheduling strategy under a distributed large model data processing architecture, and then controlling the execution of production data transmission and data cleaning and annotation, while achieving load balancing of large model servers under the influence of multiple factors of the adaptation scenario, improving the efficiency and quality of data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flow chart of the text content cleaning and annotation method based on the large model of the present invention; Figure 2 This is a structural diagram of the text content cleaning and annotation system based on a large model of the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0019] The embodiment of the present invention provides a text content cleaning and annotation method based on a large model, referring to Figure 1 , Figure 1 It is a flow chart of an embodiment of a method for cleaning and annotating text content based on a large model according to the present invention.

[0020] In this embodiment, a text content cleaning and annotation method based on a large model is provided, and the method comprises the following steps: S100: Obtain server association information of each large model server in the distributed large model data processing architecture, and extract the server identifier and data processing mapping responsibility area in the server association information; S200: using the server identifier, determining a data processing field fitness value ranking table of each large model server and data processing capacity quantification information of each processing cycle in a target processing period; wherein the data processing capacity quantification information includes a first data processing capacity quantification value for text content cleaning and a second data processing capacity quantification value for text content annotation; S300: using the data processing mapping responsibility area, accessing the production unit deployment location database, querying a number of production data output units having an initial data processing matching relationship with each large model server; S400: predicting and generating production data output information of each production data output unit in a target processing period according to the historical production database of each production data output unit; predicting the data transmission quality between each production data output unit and each large model server in each processing cycle in the target processing period according to the historical network status information of each production data output unit; S500: Based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server, and the quantitative information of the data processing capacity of each processing cycle in the target processing period, the production data output unit allocated to each large model server in each processing cycle of the target processing period is planned, and a data processing scheduling strategy is generated; S600: According to the data processing scheduling strategy, each large model server is controlled to perform text content cleaning and annotation actions of the production data transmitted by the assigned production data output unit in each processing cycle of the target processing period.

[0021] It should be noted that there are currently some platforms that can use large models to clean and annotate text data provided by users, and return the cleaned and annotated data to users. However, this solution still has the following limitations in actual applications; first, using a number of large model servers distributed on the back end of the platform can solve the needs of data cleaning and annotation within the area to which each large model server belongs. However, due to the different data processing requirements of different production data output units in each area at different times (including data processing volume and data processing type), the processing capacity of each large model server is also different, and data processing load peak pressure may occur at different times. The traditional server load balancing method cannot achieve satisfactory results in the scenario of the above multiple influencing factors; second, since each large model server mainly serves the data processing tasks of several production data output units within the area to which it belongs, each large model server needs to have stronger professionalism in data processing in a specific field, rather than in the entire field (usually large language models trained in multiple fields have weak generalization capabilities, such as mistakenly treating e-commerce "bad reviews" as "bad reviews"). The sentiment words in the video are understood as medical semantics). For such scenarios, the existing server load balancing solutions have not yet comprehensively considered the domain relevance of large models. Third, since the production task types of different production data output units may be different, their requirements for data processing latency will also be different (for example, the data processing requests sent by artificial intelligence customer service in the e-commerce industry require high real-time performance, while the real-time performance required for analyzing and summarizing business data in the commercial field is not high). Such requirements will also have a great impact on the implementation of server load balancing, thereby affecting the efficiency and quality of data cleaning and annotation within the overall area.

[0022] In order to solve the above problems, this embodiment uses the server identifier to determine the data processing field fitness value ranking table and data processing capacity quantification information, uses the data processing mapping responsibility area to query the production data output unit, predicts and generates production data output information and data transmission quality, establishes a constraint condition set including data processing capacity quantification information and data transmission quality, and an optimization goal determined by the data processing field fitness value ranking table, comprehensively considers the influence of factors such as data processing demand and capacity matching, large model server field relevance, and data transmission quality, plans the data processing scheduling strategy under the distributed large model data processing architecture, and then controls the execution of production data transmission and data cleaning and annotation, while achieving load balancing of large model servers under the influence of multiple factors of the adaptation scenario, and improves the efficiency and quality of data processing.

[0023] In a preferred embodiment, the server identifier is used to determine the data processing field fitness value ranking table of each large model server and the data processing capacity quantitative information step of each processing cycle in the target processing period, specifically including: S210: using the server identifier, querying the task running list and the server historical processing data information base of each large model server, and determining the quantitative information of the data processing capacity of each large model server in each processing cycle in the target processing period; S220: Utilize the server identifier to query the training process information of each large model server, extract the model training sample data in the training process information, analyze the model training sample data, and generate a sorted table of processing data domain fitness values ​​of each large model server.

[0024] In this embodiment, after obtaining the server identifier of each large model server in the distributed large model data processing architecture, the server identifier can be used to query the quantitative information of the data processing capacity of each large model server in each processing cycle in the target processing period, analyze the model training sample data and generate a ranking table of fitness values ​​of the processing data domain of each large model server, which can provide a set of constraints for constructing the optimization algorithm for the subsequent generation of data processing scheduling strategies.

[0025] On this basis, using the server identifier, querying the task running list and server historical processing data information library of each large model server, and determining the quantitative information steps of the data processing capacity of each large model server in each processing cycle in the target processing period, specifically includes: S211: using the server identifier, querying the task running list and server historical processing data information base of each large model server; S212: extracting the task execution period and the task execution data volume of each to-be-executed task in the task running list, dividing the task execution period into each processing cycle in the target processing period, and determining the to-be-processed data volume of each large model server in each processing cycle in the target processing period; S213: extracting the amount of data processed by each large model server in each processing cycle during the historical data processing process recorded in the server historical processing data information library, and determining the standard amount of data processed by each large model server in each processing cycle; S214: Determine the quantitative information of the data processing capacity of each large model server in each processing cycle in the target processing period according to the difference between the amount of data to be processed in each processing cycle of each large model server in the target processing period and the standard amount of data processed in each processing cycle.

[0026] In practical applications, the tasks to be executed include text content cleaning tasks and text content annotation tasks, the amount of data to be processed includes the amount of text content to be cleaned and the amount of text content to be annotated, and the standard processing data includes the amount of text content standard cleaning data and the amount of text content standard annotation data.

[0027] Furthermore, according to the difference between the amount of data to be processed in each processing cycle of each large model server in the target processing period and the standard amount of data processed in each processing cycle, the step of determining the quantitative information of the data processing capacity of each large model server in each processing cycle in the target processing period specifically includes: S2141: Calculate a first data processing capability quantification value for text content cleaning according to the difference between the amount of text content to-be-cleaned data in each processing cycle of each large model server in the target processing period and the amount of text content standard cleaned data in each processing cycle; S2142: Calculate a second data processing capability quantification value for text content cleaning according to the difference between the amount of text content to-be-annotated data in each processing cycle of each large model server in the target processing period and the amount of text content standard annotated data in each processing cycle; S2143: Determine the data processing capacity quantification information of each large model server in each processing cycle in the target processing period by using the first data processing capacity quantification value for text content cleaning and the second data processing capacity quantification value for text content cleaning.

[0028] In this embodiment, the server identifier is first used to query the amount of data to be processed in each processing cycle of each large model server in the target processing time period, and then the standard processing data amount of each large model server in each processing cycle recorded in the historical processing data information library is used (the standard processing data amount can be measured by the maximum processing data amount recorded in the historical processing data information library for each processing cycle that does not affect the normal operation of the server). After that, by considering the text content cleaning task and the text content labeling task, the difference between the two is calculated to obtain a first data processing capacity quantification value for text content cleaning and a second data processing capacity quantification value for text content cleaning, and finally the data processing capacity quantification information of each large model server in each processing cycle in the target processing time period is obtained.

[0029] In a preferred embodiment, the server identifier is used to query the training process information of each large model server, extract the model training sample data in the training process information, analyze the model training sample data, and generate a processing data domain fitness value ranking table for each large model server, which specifically includes: S221: using the server identifier, querying the training process information of each large model server, and extracting the model training sample data in the training process information; S222: extracting a plurality of keywords from the model training sample data, and determining sample training ratios of different processing data fields in the model training sample data according to a ratio of the cumulative sum of similarity values ​​between the plurality of keywords and feature phrases corresponding to different processing data fields; S223: Sort the values ​​of each processing data field in the sample training ratio from large to small, and generate a processing data field fitness ranking table for each large model server.

[0030] In this embodiment, the server identifier is used to query the training process information of each large model server, and the specific field for deeper learning of each large model server is determined by analyzing the proportion of sample data fields in the model training sample data. This is used as the professionalism for different processing data fields, and a fitness ranking table for the processing data field of each large model server is generated, which is used as the optimization target in the subsequent construction of the optimization algorithm, so that the entire large model server architecture adopts the best data processing task allocation method in the field, thereby improving the overall data cleaning and labeling accuracy.

[0031] In a preferred embodiment, the step of predicting and generating the production data output information of each production data output unit in the target processing period according to the historical production database of each production data output unit specifically includes: S410: querying a historical production database of each production data output unit, and extracting the production data output volume within a unit processing cycle when different types of production tasks are executed, which is recorded in the historical production database; S420: predicting the output volume of production data of each production data output unit in the target processing period according to the production task types of each production data output unit in different processing cycles in the target processing period; S430: Based on the production data output volume, the production data field type corresponding to the production task type and the production data transmission restriction set, construct the production data output information of each production data output unit in the target processing period; wherein the production data transmission restriction set includes data integrity restriction and real-time restriction.

[0032] On this basis, according to the historical network status information of each production data output unit, the data transmission quality steps of each production data output unit and each large model server in each processing cycle during the target processing period are predicted, specifically including: S440: querying the historical network status information of each production data output unit, extracting the historical data transmission quality parameters between each production data output unit and each large model server in the non-mapped responsibility area recorded in the historical network status information, and constructing a data transmission quality training sample including a plurality of data transmission quality parameter features consisting of a data transmission timestamp and a data transmission quality parameter; S450: using the data transmission quality training samples to train the constructed initial convolutional neural network model, and when the number of training times reaches the target number of times or the model converges, obtaining a trained data transmission quality prediction model for each production data output unit and each large model server; S460: Utilize the data transmission quality prediction model to predict the data transmission quality parameters of each production data output unit and each large model server in each non-mapped responsibility area in each processing cycle during the target processing period; wherein the data transmission quality parameters include data packet loss rate and data transmission delay.

[0033] In this embodiment, considering that different production data output units in each region have different data processing requirements (including data processing volume and data processing type) in different time periods and that the production task types of different production data output units may be different, their requirements for data processing delay may also be different (for example, the data processing requests sent by artificial intelligence customer service in the e-commerce industry require high real-time performance, while the real-time performance required for analyzing and summarizing business data in the business field is not high), the production data output volume of each production data output unit in the target processing period is predicted by querying the historical production database of each production data output unit, and production data output information including the production data output volume, the production data field type corresponding to the production task type and the production data transmission restriction set is constructed, and then the historical network status information of each production data output unit is queried, and the data packet loss rate and data transmission delay of each production data output unit in each processing cycle of the target processing period and each non-mapped responsibility area are predicted by using the data transmission quality prediction model, so as to provide a set of constraints when constructing the optimization algorithm for the generation of subsequent data processing scheduling strategies, and through the construction of multiple constraints, the rationality of the data processing scheduling strategies and the overall data cleaning and labeling accuracy of multiple production data output units in the region can be improved.

[0034] In a preferred embodiment, based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server, and the quantitative information of the data processing capacity of each processing cycle in the target processing period, the production data output unit allocated to each large model server in each processing cycle of the target processing period is planned, and the data processing scheduling strategy steps are generated, which specifically include: S510: Based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server and the data processing capacity quantification information of each processing cycle in the target processing period; S520: The first constraint condition is that the sum of the production data output amounts of several production data output units allocated to each large model server in each processing cycle of the target processing period is less than the first data processing capacity quantization value and the second data processing capacity quantization value of the large model server in the corresponding processing cycle. The second constraint condition is that the data transmission quality parameter corresponding to the data transmission quality of the production data output unit that belongs to the non-mapped responsibility area allocated to each large model server in each processing cycle of the target processing period and the large model server satisfies the data transmission quality of the data integrity restriction and the data transmission delay corresponding to the real-time restriction of the data transmission and return process. The optimization target is that the sum of the order values ​​of the production data field type in the production data output information of the production data output units allocated to all large model servers in each processing cycle of the target processing period in the processing data field fitness ranking table of the large model server is minimized. The production data output unit allocated to each large model server in each processing cycle of the target processing period is optimized. S530: Generate a data processing scheduling strategy based on the production data output unit allocated to each large model server in each processing cycle during the target processing period.

[0035] In this embodiment, by obtaining the server identification and data processing mapping responsibility area of ​​each large model server in the distributed large model data processing architecture, using the server identification to determine the data processing field fitness value ranking table and data processing capacity quantification information, using the data processing mapping responsibility area, querying the production data output unit to which it belongs, predicting and generating production data output information and data transmission quality, by establishing a constraint condition set including data processing capacity quantification information and data transmission quality and an optimization target determined based on the data processing field fitness value ranking table, an optimization algorithm is used to plan the production data output unit assigned to each large model server in each processing cycle of the target processing period, and then a data processing scheduling strategy is generated and the production data transmission of each production data output unit and the data cleaning and labeling of each large model server are controlled. Therefore, by comprehensively considering the matching of data processing requirements and capabilities, the influence of factors such as the domain relevance of large model servers and the quality of data transmission, the data processing scheduling strategy under the distributed large model data processing architecture is planned, which can improve the quality of data cleaning and labeling services provided to a large number of production data output units in different regions using the distributed large model data processing architecture, and improve the efficiency and quality of data processing while achieving load balancing of large model servers under the influence of multiple factors of the adaptation scenario.

[0036] Reference Figure 2 , Figure 2 It is a structural block diagram of an embodiment of a text content cleaning and annotation system based on a large model of the present invention.

[0037] like Figure 2 As shown, the text content cleaning and annotation system based on a large model proposed in an embodiment of the present invention includes: An acquisition module 10 is used to acquire server association information of each large model server in the distributed large model data processing architecture, and extract the server identifier and data processing mapping responsibility area in the server association information; A determination module 20 is used to determine the data processing field fitness value ranking table of each large model server and the data processing capacity quantification information of each processing cycle in the target processing period by using the server identifier; wherein the data processing capacity quantification information includes a first data processing capacity quantification value for text content cleaning and a second data processing capacity quantification value for text content annotation; A query module 30 is used to access the production unit deployment location database using the data processing mapping responsibility area to query a number of production data output units that have an initial data processing matching relationship with each large model server; The prediction module 40 is used to predict and generate the production data output information of each production data output unit in the target processing period according to the historical production database of each production data output unit; and predict the data transmission quality between each production data output unit and each large model server in each processing cycle of the target processing period according to the historical network status information of each production data output unit; A generating module 50 is used to plan the production data output units allocated to each large model server in each processing cycle of the target processing period based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server, and the data processing capacity quantification information of each processing cycle in the target processing period, and generate a data processing scheduling strategy; The execution module 60 is used to control each large model server to perform text content cleaning and annotation actions on the production data transmitted by the assigned production data output unit in each processing cycle of the target processing period according to the data processing scheduling strategy.

[0038] Other embodiments or specific implementations of the text content cleaning and annotation system based on a large model of the present invention can refer to the above-mentioned method embodiments and will not be described in detail here.

[0039] It is understood that, in the description of this specification, the description with reference to the terms "one embodiment", "another embodiment", "other embodiments", or "first to Nth embodiments" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0040] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.

[0041] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A text content cleaning and annotation method based on a large model, characterized in that: The method comprises the following steps: Obtaining server association information of each large model server in the distributed large model data processing architecture, and extracting the server identifier and data processing mapping responsibility area in the server association information; Using the server identifier, determine a data processing field fitness value ranking table of each large model server and data processing capacity quantification information of each processing cycle in a target processing period; wherein the data processing capacity quantification information includes a first data processing capacity quantification value for text content cleaning and a second data processing capacity quantification value for text content annotation; Using the data processing mapping responsibility area, accessing the production unit deployment location database, querying a number of production data output units that have an initial data processing matching relationship with each large model server; According to the historical production database of each production data output unit, the production data output information of each production data output unit in the target processing period is predicted and generated; according to the historical network status information of each production data output unit, the data transmission quality between each production data output unit and each large model server in each processing cycle in the target processing period is predicted; Based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server and the quantitative information of the data processing capacity of each processing cycle in the target processing period, the production data output unit allocated to each large model server in each processing cycle of the target processing period is planned to generate a data processing scheduling strategy; According to the data processing scheduling strategy, each large model server is controlled to perform text content cleaning and annotation actions of the production data transmitted by the assigned production data output unit in each processing cycle of the target processing period.

2. The text content cleaning and annotation method based on a large model as claimed in claim 1 is characterized in that: The step of using the server identifier to determine the data processing field fitness value ranking table of each large model server and the data processing capacity quantitative information of each processing cycle in the target processing period specifically includes: Using the server identifier, query the task running list and server historical processing data information base of each large model server to determine the quantitative information of the data processing capacity of each large model server in each processing cycle in the target processing period; The server identifier is used to query the training process information of each large model server, extract the model training sample data in the training process information, analyze the model training sample data, and generate a processing data domain fitness value ranking table for each large model server.

3. The text content cleaning and annotation method based on a large model as described in claim 2 is characterized in that: The server identifier is used to query the task running list and the server historical processing data information base of each large model server to determine the quantitative information steps of the data processing capacity of each large model server in each processing cycle in the target processing period, specifically including: Using the server identifier, query the task running list and server historical processing data information base of each large model server; Extracting the task execution period and the task execution data volume of each to-be-executed task in the task running list, dividing the task execution period into each processing cycle in the target processing period, and determining the amount of to-be-processed data of each large model server in each processing cycle in the target processing period; Extracting the amount of data processed by each large model server in each processing cycle during the historical data processing process recorded in the server historical processing data information library, and determining the standard amount of data processed by each large model server in each processing cycle; According to the difference between the amount of data to be processed in each processing cycle of each large model server in the target processing period and the standard amount of data processed in each processing cycle, the quantitative information of the data processing capacity of each large model server in each processing cycle in the target processing period is determined.

4. The text content cleaning and annotation method based on a large model as claimed in claim 3 is characterized in that: The tasks to be executed include text content cleaning tasks and text content annotation tasks, the amount of data to be processed includes the amount of text content to be cleaned and the amount of text content to be annotated, and the amount of standard processed data includes the amount of text content standard cleaned data and the amount of text content standard annotated data.

5. The text content cleaning and annotation method based on a large model as claimed in claim 4 is characterized in that: The step of determining the quantitative information of the data processing capacity of each large model server in each processing cycle in the target processing period according to the difference between the amount of data to be processed in each processing cycle of each large model server in the target processing period and the standard amount of data to be processed in each processing cycle includes: Calculate a first data processing capacity quantification value for text content cleaning according to the difference between the amount of text content to be cleaned data in each processing cycle of each large model server in the target processing period and the amount of text content standard cleaned data in each processing cycle; Calculate a second data processing capacity quantification value for text content cleaning according to the difference between the amount of text content to-be-annotated data in each processing cycle of each large model server in the target processing period and the amount of text content standard annotated data in each processing cycle; The data processing capacity quantification information of each large model server in each processing cycle in the target processing period is determined by using the first data processing capacity quantification value for text content cleaning and the second data processing capacity quantification value for text content cleaning.

6. The text content cleaning and annotation method based on a large model as claimed in claim 2 is characterized in that: The steps of using the server identifier to query the training process information of each large model server, extracting the model training sample data in the training process information, analyzing the model training sample data, and generating a processing data domain fitness value ranking table for each large model server specifically include: Using the server identifier, query the training process information of each large model server, and extract the model training sample data in the training process information; Extracting a number of keywords from the model training sample data, and determining the sample training ratios of different processing data fields in the model training sample data according to the ratio of the cumulative sum of similarity values ​​between the keywords and the feature phrases corresponding to different processing data fields; The values ​​of each processing data field in the sample training ratio are sorted from large to small to generate a processing data field fitness ranking table for each large model server.

7. The text content cleaning and annotation method based on a large model as claimed in claim 1 is characterized in that: The step of predicting and generating production data output information of each production data output unit in a target processing period according to a historical production database of each production data output unit specifically includes: Querying the historical production database of each production data output unit, and extracting the production data output within a unit processing cycle when different types of production tasks are executed, which are recorded in the historical production database; Predicting the output volume of production data of each production data output unit in the target processing period according to the production task types of each production data output unit in different processing cycles in the target processing period; Based on the production data output volume, the production data field type corresponding to the production task type and the production data transmission restriction set, the production data output information of each production data output unit in the target processing period is constructed; wherein the production data transmission restriction set includes data integrity restriction and real-time restriction.

8. The text content cleaning and annotation method based on a large model as claimed in claim 1, characterized in that: According to the historical network status information of each production data output unit, the data transmission quality steps of each production data output unit and each large model server in each processing cycle during the target processing period are predicted, specifically including: Query the historical network status information of each production data output unit, extract the historical data transmission quality parameters between each production data output unit and each large model server in the non-mapped responsibility area recorded in the historical network status information, and construct a data transmission quality training sample containing a plurality of data transmission quality parameter features consisting of a data transmission timestamp and a data transmission quality parameter; The constructed initial convolutional neural network model is trained using the data transmission quality training samples, and when the number of training times reaches the target number of times or the model converges, a trained data transmission quality prediction model for each production data output unit and each large model server is obtained; The data transmission quality prediction model is used to predict the data transmission quality parameters of each production data output unit and each large model server in each non-mapped responsibility area in each processing cycle of the target processing period; wherein the data transmission quality parameters include data packet loss rate and data transmission delay.

9. The text content cleaning and annotation method based on a large model as claimed in claim 1, characterized in that: Based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server, and the quantitative information of the data processing capacity of each processing cycle in the target processing period, the production data output unit allocated to each large model server in each processing cycle of the target processing period is planned, and the data processing scheduling strategy steps are generated, which specifically include: Based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server and the data processing capacity quantification information of each processing cycle in the target processing period; The first constraint condition is that the sum of the production data output amounts of several production data output units allocated to each large model server in each processing cycle of the target processing period is less than the first data processing capacity quantification value and the second data processing capacity quantification value of the large model server in the corresponding processing cycle at the same time; the second constraint condition is that the data transmission quality parameter corresponding to the data transmission quality of the production data output unit belonging to the non-mapped responsibility area allocated to each large model server in each processing cycle of the target processing period and the large model server satisfies the data transmission quality of the data integrity restriction and the data transmission delay corresponding to the real-time restriction of the data transmission and return process; the optimization target is that the sum of the order values ​​of the production data field type in the production data output information of the production data output units allocated to all large model servers in each processing cycle of the target processing period in the processing data field fitness ranking table of the large model server is minimized, and the production data output unit allocated to each large model server in each processing cycle of the target processing period is optimized; A data processing scheduling strategy is generated based on the production data output units allocated to each large model server in each processing cycle during the target processing period.

10. A text content cleaning and annotation system based on a large model, characterized in that: include: An acquisition module, used to acquire server association information of each large model server in the distributed large model data processing architecture, and extract the server identifier and data processing mapping responsibility area in the server association information; A determination module, used to determine the data processing field fitness value ranking table of each large model server and the data processing capacity quantification information of each processing cycle in the target processing period by using the server identifier; wherein the data processing capacity quantification information includes a first data processing capacity quantification value for text content cleaning and a second data processing capacity quantification value for text content annotation; A query module, for accessing the production unit deployment location database using the data processing mapping responsibility area, and querying a number of production data output units having an initial data processing matching relationship with each large model server; A prediction module is used to predict and generate production data output information of each production data output unit in a target processing period according to the historical production database of each production data output unit; and to predict the data transmission quality of each production data output unit and each large model server in each processing cycle of the target processing period according to the historical network status information of each production data output unit; A generation module is used to plan the production data output units allocated to each large model server in each processing cycle of the target processing period based on the production data output information of each production data output unit in the target processing period, the data processing field fitness value ranking table of each large model server, and the quantitative information of the data processing capacity of each processing cycle in the target processing period, and generate a data processing scheduling strategy; The execution module is used to control each large model server to execute the text content cleaning and annotation actions of the production data transmitted by the assigned production data output unit in each processing cycle of the target processing period according to the data processing scheduling strategy.

Citation Information

Patent Citations

  • Theme web crawler method and device and medium

    CN110069690A

  • Digital management system and method based on multi-source data information

    CN118626771A

  • Intelligent data labeling method and system

    CN119647479A

  • Platform for facilitating development of intelligence in an industrial internet of things system

    US20200348662A1