Data co-processing method and system based on large model

Through the data collaborative processing method driven by large-models, unified feature representation is generated and adaptive weight calculation is performed, which solves the problems of insufficient intelligence level and low cross-platform data integration efficiency in the existing technology, and realizes efficient task resource allocation and collaborative performance evaluation.

CN120494431AActive Publication Date: 2025-08-15FUJIAN RONGJI SOFTWARE +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510941365.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-08-15
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

The existing technology has shortcomings in terms of intelligence level, dynamic adaptability and cross-platform data integration efficiency, resulting in insufficient data collaborative processing capabilities.

Method used

A large model is introduced to perform feature extraction and labeling of multi-source heterogeneous data, and a unified feature representation is generated. Combined with data classification, region division and adaptive weight calculation, the task allocation threshold is dynamically adjusted to optimize cross-domain data integration.

Benefits of technology

It improves the intelligence level and dynamic adaptability of data collaborative processing, improves the efficiency of cross-platform data integration, and realizes efficient task resource allocation and collaborative performance evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494431A_ABST
    Figure CN120494431A_ABST
Patent Text Reader

Abstract

The invention discloses a data collaborative processing method and system based on a large model, relates to the technical field of data collaboration, is used for solving the problems of insufficient intelligent level, insufficient dynamic adaptability and low cross-platform data integration efficiency, and is used for generating uniform feature representation by introducing the large model to perform feature extraction and labeling on multi-source heterogeneous data. And a basis is provided for data classification and task allocation. Through data classification marking, high-priority data and low-priority data are pre-divided, the task adjustment range is narrowed, and resource consumption is reduced. For low-priority data, a collaborative weight coefficient is obtained based on region division, environment information collection and adaptive weight calculation, and the collaborative weight accuracy is improved. The task execution efficiency is dynamically detected, the cooperation performance is comprehensively evaluated in combination with the cooperation weight, the task allocation threshold value is adjusted according to the evaluation result, task resource dynamic allocation is achieved, and the cooperation efficiency and the cross-domain integration capacity are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data collaboration technology, and more specifically, to a data collaborative processing method and system based on a large model. Background Art

[0002] Driven by big data and artificial intelligence technologies, collaborative data processing methods and systems are becoming essential tools for multi-domain collaboration and intelligent decision-making support. However, existing technologies still have limitations in terms of intelligence, dynamic adaptability, and cross-platform data integration capabilities, hindering their practical application in complex scenarios.

[0003] The existing technology has the following deficiencies: At present, the data feature representation capability is limited, the dynamic adjustment mechanism of data priority is imperfect, the adaptive calculation of data collaborative weights is insufficient, and the task allocation threshold tuning method is missing. These results lead to insufficient intelligence, lack of dynamic adaptability, and inefficient cross-platform data integration. Therefore, a data collaborative processing method and system based on a large model is proposed.

[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a data collaborative processing method and system based on a large model. By introducing a dynamic adaptation mechanism driven by a large model, combined with multi-source data fusion and intelligent task allocation strategy, it optimizes cross-domain data integration capabilities and solves the problems of insufficient intelligence level, lack of dynamic adaptability and low efficiency of cross-platform data integration mentioned in the background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solutions: a data collaborative processing method and system based on a large model, comprising the following steps: Step S1: Obtain multi-source heterogeneous data and pre-process them, extract and annotate data features through a large model, and generate a unified data feature representation in the form of a standardized feature vector; Step S2: Classify and label the data sources according to the data feature representation, build a data classification index table based on historical collaborative data, and divide the data into high-priority data and low-priority data; Step S3: Divide the low-priority data into regions, collect environmental information of each data region, and use an adaptive weight calculation method to obtain a collaborative weight coefficient for each data region; Step S4: Perform dynamic task allocation detection on the data of each data sub-region, determine the task execution efficiency of each data sub-region, comprehensively consider the collaborative weight coefficient and task execution efficiency, and use the weighted average method to evaluate the collaborative performance of each data sub-region; Step S5: adjusting the preset task allocation threshold according to the collaborative performance evaluation results of each data sub-region, and reallocating task resources for different data sub-regions.

[0007] In a preferred embodiment, in step S1, when acquiring multi-source heterogeneous data, the original data is obtained from different data platforms through distributed data acquisition nodes, and the semantic understanding ability of the large model is used to parse and annotate the original data to generate a standardized feature vector as a unified data feature representation.

[0008] In a preferred embodiment, in step S2, the historical collaborative data includes the task completion time and resource consumption of different data sources in the collaborative processing process. The historical database is accessed to obtain relevant information of all data sources recorded in the past collaborative processing process, which is converted into corresponding collaborative efficiency indicators, and classified and stored according to the data source to form a data classification index table.

[0009] In a preferred embodiment, in step S2, data sources corresponding to collaborative efficiency indicators exceeding a preset classification ratio are classified as high-priority data, otherwise they are classified as low-priority data.

[0010] In a preferred embodiment, in step S3, the distribution information of the data source is obtained through the distributed data collection nodes, and the coverage and data density of the data source are extracted from the data distribution information; The ratio of data density to coverage is used as the data distribution index to divide data sources into regions.

[0011] In a preferred embodiment, in step S3, the environmental information of the data sub-region includes regional data flow and regional data delay. The regional data flow is obtained by detecting the network bandwidth occupancy of each data sub-region at the same time point; Multiply the network bandwidth occupancy rate of the data sub-region by the number of network nodes in the corresponding data sub-region; The product result is used as the regional data flow of the data sub-region; Regional data delay is detected by setting the same time interval to test the network response time of each data sub-region; Calculate the response time difference between two adjacent time points and take the average value as the regional data delay of each data sub-region; The adaptive weight calculation method is used to obtain the collaborative weight coefficient of each data sub-region. The specific collaborative weight calculation coefficient formula is expressed as: ; Where, is the regional discharge coefficient, is the regional delay coefficient, 、 is the value corresponding to the optimal solution, 、 is the value corresponding to the worst solution, is the collaborative weight coefficient.

[0012] In a preferred embodiment, in step S4, when performing dynamic task allocation detection on the data of each data sub-region, the task queue length of each data sub-region is first detected; Calculate the average value as the initial task load of the corresponding data sub-region, and allocate equal task resources to each data sub-region; After the same time interval, the task queue length of each data sub-region is tested again, and the average value is calculated as the final task load of the corresponding data sub-region; The difference between the final task load and the initial task load is used as the task execution efficiency of each data sub-region; The weighted average method is used to evaluate the collaborative performance of each data region by combining the collaborative weight coefficient and task execution efficiency. The specific formula is as follows: ;

[0013] Where, is the collaborative performance evaluation value of the data region, To standardize task execution efficiency, To assign weights, is the collaborative weight coefficient.

[0014] In a preferred embodiment, in step S5, the collaborative performance evaluation values of the data sub-regions are averaged to obtain an average collaborative performance evaluation value of the data sub-regions; The average value of the collaborative performance evaluation of the data sub-region is used as the screening benchmark, and the collaborative performance evaluation value of the data sub-region is compared with the screening benchmark; When the collaborative performance evaluation value of a data sub-region is lower than the screening benchmark, the ratio of the collaborative performance evaluation value of the corresponding data sub-region to the screening benchmark is used as the adjustment factor for the corresponding data sub-region; The preset task allocation threshold is multiplied by the adjustment factor to obtain the adjusted task allocation threshold.

[0015] The data collaborative processing system based on the large model includes a data acquisition module, a data classification module, a region division module, a collaborative performance evaluation module, and a task allocation and adjustment module; The data acquisition module is used to obtain raw data from multi-source data platforms and perform preprocessing to generate a unified data feature representation in the form of standardized feature vectors; The data classification module is used to classify and mark the data sources according to the data feature representation and build a data classification index table; The regional division module is used to divide the data source into regions according to the data distribution information and collect the environmental information of each data region; The collaborative performance evaluation module is used to evaluate the collaborative performance based on the environmental information and task execution efficiency of each data sub-region; The task allocation adjustment module is used to adjust the task allocation threshold according to the collaborative performance evaluation results of each data sub-region and reallocate task resources.

[0016] In the data collaborative processing system based on large models, the data acquisition module obtains raw data from different data platforms through distributed data acquisition nodes, and uses high-speed communication links to connect the distributed data acquisition nodes to ensure the real-time and stability of data transmission.

[0017] Technical effects and advantages of the present invention: The present invention extracts and labels features of multi-source heterogeneous data by introducing a large model, generates a unified data feature representation, and lays the foundation for subsequent data classification and task allocation. By classifying and marking the data source, the data is divided into high priority and low priority in advance, the task allocation adjustment range is narrowed, and resource consumption is reduced. When the data is classified as low priority, the collaborative weight coefficient is obtained through regional division and environmental information collection, combined with the adaptive weight calculation method, to improve the accuracy of the collaborative weight of the data sub-region. By dynamically detecting the task execution efficiency of each data sub-region, and comprehensively evaluating the collaborative performance by combining the collaborative weight coefficient and the task execution efficiency, it can more accurately reflect the actual collaborative ability of each data sub-region, adjust the task allocation threshold according to the collaborative performance evaluation result, and realize the dynamic allocation of task resources, thereby improving the overall collaborative efficiency and cross-domain data integration capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flowchart of the data collaborative processing method based on a large model of the present invention.

[0019] Figure 2 This is a schematic diagram of the module structure of the data collaborative processing system based on the large model of the present invention. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. Example

[0021] The present invention provides a data collaborative processing method and system based on a large model. The core of the method is to optimize cross-domain data integration capabilities by introducing a dynamic adaptation mechanism driven by a large model, combining multi-source data fusion and intelligent task allocation strategies.

[0022] The following will be combined with the attached Figure 1 and attached Figure 2 The specific structure and its labels are described in detail.

[0023] like Figure 1 As shown in the figure, the overall process of the data collaborative processing method based on the large model from data collection to task allocation and adjustment includes several key steps, which are Figure 2 The system modules shown in the figure work together to complete the task.

[0024] The system includes a data acquisition module, a data classification module, a region division module, a collaborative performance evaluation module, and a task allocation and adjustment module. The modules are logically connected to achieve data flow and functional coordination.

[0025] In the specific implementation process, the data acquisition module first executes step S1, that is, obtaining multi-source heterogeneous data and performing preprocessing.

[0026] The data acquisition module obtains raw data from different data platforms through distributed data acquisition nodes. These nodes are distributed in different physical locations and connected to each other through high-speed communication links to ensure the real-time and stability of data transmission.

[0027] After the raw data is transmitted to the data acquisition module, the data is parsed and labeled by calling the semantic understanding capabilities of the large model.

[0028] The large model uses deep learning algorithms to extract semantic information from the data and generate standardized feature vectors as a unified data feature representation.

[0029] The key to this process lies in the semantic parsing capability of the large model, which can generate highly consistent feature representations by understanding complex data, thus providing basic support for subsequent steps.

[0030] Then, the process proceeds to step S2, where the data classification module classifies and labels the data sources according to the data feature representation.

[0031] The data classification module accesses the historical database to obtain relevant information of all data sources recorded in the past collaborative processing process and converts it into corresponding collaborative efficiency indicators.

[0032] These collaborative efficiency indicators are classified and stored according to data sources to form a data classification index table. Based on the preset classification ratio, the data sources corresponding to collaborative efficiency indicators that exceed the classification ratio are classified as high-priority data. Otherwise, it is classified as low-priority data. This classification process relies on the task completion time and resource consumption of historical collaborative data to ensure that the classification results have high accuracy.

[0033] The data classification module passes the classified and labeled data to the region division module for further processing.

[0034] In step S3, the region division module divides the low priority data into regions. The region division module first obtains the distribution information of the data source through the distributed data collection nodes, and extracts the coverage and data density of the data source.

[0035] The ratio of data density to coverage is used as the data distribution index to divide data sources into regions.

[0036] The regional division module is also responsible for collecting environmental information of each data sub-region, including regional data traffic and regional data delay.

[0037] Regional data traffic is determined by detecting the network bandwidth occupancy of each data sub-region at the same time point, multiplying the network bandwidth occupancy rate of the data sub-region by the number of network nodes in the corresponding data sub-region, and using the product as the regional data traffic of the data sub-region.

[0038] Regional data delay is determined by setting the same time interval to test the network response time of each data sub-region separately, calculating the response time difference between two adjacent time points and taking the average value as the regional data delay of each data sub-region.

[0039] The region division module passes this environmental information to the collaborative performance evaluation module to provide input for its subsequent calculations.

[0040] Step S4 is executed by the collaborative performance evaluation module, whose main task is to perform dynamic task allocation detection on the data of each data sub-region and determine the task execution efficiency of each data sub-region.

[0041] The collaborative performance evaluation module first detects the task queue length of each data sub-region and calculates the average value as the initial task load of the corresponding data sub-region.

[0042] By allocating equal amounts of task resources to each data sub-region, the task queue length of each data sub-region is tested twice after the same time interval, and the average value is calculated as the final task load of the corresponding data sub-region.

[0043] The difference between the final task load and the initial task load is used as the task execution efficiency of each data sub-region. The collaborative performance evaluation module uses Min-Max normalization to process the task execution efficiency of each data sub-region and obtain a standardized task execution efficiency value.

[0044] At the same time, the collaborative performance evaluation module comprehensively considers the standardized task execution efficiency value and the collaborative weight coefficient of the corresponding data sub-region, and uses the weighted average method to evaluate the collaborative performance of each data sub-region.

[0045] The calculation of the collaborative weight coefficient depends on the regional traffic coefficient, regional delay coefficient and allocation weight provided by the regional division module. The specific calculation process includes basic value calculation, determination of optimal and worst solutions, calculation of Euclidean distance and final solution of the collaborative weight coefficient.

[0046] The specific collaborative weight calculation coefficient formula is expressed as: ; Where, is the regional discharge coefficient, is the regional delay coefficient, 、 is the value corresponding to the optimal solution, 、 is the value corresponding to the worst solution, is the collaborative weight coefficient; The weighted average method is used to evaluate the collaborative performance of each data region by combining the collaborative weight coefficient and task execution efficiency. The specific formula is as follows: ; Where, is the collaborative performance evaluation value of the data region, To standardize task execution efficiency, To assign weights, is the collaborative weight coefficient.

[0047] The collaborative performance evaluation module passes the evaluation results to the task allocation adjustment module for the next step.

[0048] In step S5, the task allocation adjustment module adjusts the preset task allocation threshold according to the collaborative performance evaluation results of each data sub-region.

[0049] The collaborative performance evaluation values of the data sub-regions are averaged to obtain the average collaborative performance evaluation value of the data sub-regions; The task allocation adjustment module uses the average collaborative performance evaluation value of the data sub-region as a screening benchmark and compares the collaborative performance evaluation value of the data sub-region with the screening benchmark.

[0050] When the collaborative performance evaluation value of the data sub-region is higher than the screening benchmark, the preset task allocation threshold is not adjusted; When the collaborative performance evaluation value of a data sub-region is lower than the screening benchmark, the ratio of the collaborative performance evaluation value of the corresponding data sub-region to the screening benchmark is used as the adjustment factor of the corresponding data sub-region, and the preset task allocation threshold is multiplied by the adjustment factor to obtain the adjusted task allocation threshold.

[0051] The task allocation adjustment module monitors the task queue length of each data sub-region in real time. When the task allocation threshold of the corresponding data sub-region is reached, the task resources are reallocated for the corresponding data sub-region.

[0052] The connection relationship and collaboration between the above modules ensure the efficient operation of the entire system.

[0053] The data acquisition module transmits the original data to the data classification module through the distributed data acquisition nodes; The data classification module generates classification labels and passes the data to the region division module; The regional division module completes the regional division and transmits the environmental information to the collaborative performance evaluation module; The collaborative performance evaluation module evaluates the collaborative performance and transmits the result to the task allocation adjustment module; The task allocation adjustment module dynamically adjusts task resources based on the evaluation results.

[0054] This modular design enables the system to flexibly respond to complex multi-source data collaborative processing needs in different scenarios.

[0055] In practical applications, the present invention can be widely used in cross-domain data integration scenarios, such as industrial Internet of Things, smart cities, and medical information systems.

[0056] For example, in the traffic management scenario of a smart city, the data acquisition module can obtain real-time traffic data from multiple sensors and monitoring devices. The data classification module divides the data into high-priority data (such as traffic accident alarms) and low-priority data (such as general road conditions). The area division module divides the data area according to the traffic flow and delay characteristics of the urban area. The collaborative performance evaluation module evaluates the collaborative performance of each area. The task allocation adjustment module dynamically adjusts the task resource allocation of each area based on the evaluation results, thereby improving the overall traffic management efficiency.

[0057] It can be seen from the above specific implementation methods that the present invention realizes efficient collaborative processing of multi-source heterogeneous data through modular system design and detailed process steps, solving the problems of insufficient intelligence level, lack of dynamic adaptability and low efficiency of cross-platform data integration in the existing technology.

[0058] In order to better enable relevant personnel in this technical field to fully understand and implement the present invention, the specific implementation principle of the present invention is further supplemented below with reference to a specific application scenario.

[0059] In the traffic management scenario of smart cities, the present invention realizes efficient integration and dynamic task allocation of multi-source heterogeneous data through a data collaborative processing method and system based on a large model.

[0060] The system's operation process starts with the data acquisition module, which obtains real-time traffic data from traffic monitoring cameras, sensor networks and vehicle terminals through distributed data acquisition nodes.

[0061] This data includes vehicle flow, speed, road congestion, and traffic accident alarm information. Distributed data collection nodes are located in different areas of the city and interconnected through high-speed communication links to ensure real-time and stable data transmission.

[0062] After the raw data is transmitted to the data acquisition module, the semantic understanding capability of the large model is called upon to parse and annotate the data.

[0063] For example, for traffic accident alarm information, the large model can identify its urgency and generate a standardized feature vector as a unified data feature representation.

[0064] The key to this process lies in the deep learning algorithm of the large model, which can extract semantic information from complex data, such as identifying the type, location and severity of the accident from the text description, thereby providing a highly consistent foundation for subsequent steps.

[0065] Subsequently, the data classification module classifies and labels the data sources according to the data feature representation. The data classification module accesses the historical database to obtain relevant information on all data sources recorded in the past collaborative processing process and converts it into corresponding collaborative efficiency indicators.

[0066] For example, in a traffic management scenario, the collaborative efficiency indicator can reflect the response time and resource consumption of a certain data source in past tasks.

[0067] The data classification module classifies and stores these indicators according to the data source to form a data classification index table.

[0068] According to the preset classification ratio, data sources corresponding to collaborative efficiency indicators exceeding the classification ratio are classified as high-priority data, such as traffic accident alarm information; Otherwise, it is classified as low-priority data, such as general traffic information. This classification mechanism relies on the task completion time and resource consumption of historical collaborative data to ensure high accuracy of the classification results.

[0069] After the classification is completed, the data classification module passes the classified and marked data to the region division module.

[0070] In the regional division module, low-priority data is divided into regions based on its distribution information. The regional division module first obtains the coverage and data density of the data source through distributed data collection nodes, and calculates the ratio of data density to coverage as the data distribution index.

[0071] For example, in a traffic management scenario, the data distribution index can reflect the density of traffic data in a certain area.

[0072] The regional division module is also responsible for collecting environmental information of each data sub-region, including regional data traffic and regional data delay.

[0073] Regional data traffic is determined by detecting the network bandwidth occupancy of each data sub-region at the same time point, multiplying the network bandwidth occupancy rate of the data sub-region by the number of network nodes in the corresponding data sub-region, and using the product as the regional data traffic of the data sub-region.

[0074] For example, during peak hours, the network bandwidth usage in certain areas is high, which will directly affect the data traffic in that area.

[0075] Regional data delay is determined by setting the same time interval to test the network response time of each data sub-region separately, calculating the response time difference between two adjacent time points and taking the average value as the regional data delay of each data sub-region.

[0076] This environmental information is passed to the collaborative performance evaluation module to provide input for its subsequent calculations.

[0077] The main task of the collaborative performance evaluation module is to perform dynamic task allocation detection on the data of each data sub-region and determine the task execution efficiency of each data sub-region.

[0078] In the traffic management scenario, the collaborative performance evaluation module first detects the task queue length of each data sub-area and calculates the average value as the initial task load of the corresponding data sub-area.

[0079] For example, in a certain area, the task queue length may reflect the number of traffic incidents that currently need to be handled in the area.

[0080] By allocating equal amounts of task resources to each data sub-region, the task queue length of each data sub-region is tested twice after the same time interval, and the average value is calculated as the final task load of the corresponding data sub-region.

[0081] The difference between the final task load and the initial task load is used as the task execution efficiency of each data sub-region. The collaborative performance evaluation module uses Min-Max normalization to process the task execution efficiency of each data sub-region and obtain a standardized task execution efficiency value.

[0082] At the same time, the collaborative performance evaluation module comprehensively considers the standardized task execution efficiency value and the collaborative weight coefficient of the corresponding data sub-region, and uses the weighted average method to evaluate the collaborative performance of each data sub-region.

[0083] The calculation of the collaborative weight coefficient depends on the regional flow coefficient, regional delay coefficient and allocation weight provided by the regional division module.

[0084] For example, in a certain area, if the regional flow coefficient is high and the regional delay coefficient is low, the collaborative weight coefficient of this area may be high, indicating that this area occupies a more important position in the collaborative performance evaluation.

[0085] The collaborative performance evaluation module transmits the evaluation results to the task allocation adjustment module.

[0086] In the task allocation adjustment module, the preset task allocation threshold is adjusted according to the collaborative performance evaluation results of each data sub-region.

[0087] The task allocation adjustment module first calculates the average collaborative performance evaluation value of the data sub-region as a screening benchmark, and compares the collaborative performance evaluation value of the data sub-region with the screening benchmark.

[0088] For example, in a certain area, if the collaborative performance evaluation value is higher than the screening benchmark, the preset task allocation threshold is not adjusted; If the collaborative performance evaluation value is lower than the screening benchmark, the ratio of the collaborative performance evaluation value of the corresponding data sub-region to the screening benchmark is used as the adjustment factor of the corresponding data sub-region, and the preset task allocation threshold is multiplied by the adjustment factor to obtain the adjusted task allocation threshold.

[0089] The task allocation adjustment module monitors the task queue length of each data sub-region in real time. When the task allocation threshold of the corresponding data sub-region is reached, the task resources are reallocated for the corresponding data sub-region.

[0090] For example, in a certain area, if the task queue length exceeds the task allocation threshold, the task allocation adjustment module will dynamically increase the task resource allocation in the area to alleviate task pressure.

[0091] It can be seen from the above steps that the present invention realizes efficient collaborative processing of multi-source heterogeneous data in the traffic management scenario of smart cities.

[0092] The data acquisition module transmits the original data to the data classification module through distributed data acquisition nodes. The data classification module generates classification tags and passes the data to the area division module. The area division module completes the area division and passes the environmental information to the collaborative performance evaluation module. The collaborative performance evaluation module evaluates the collaborative performance and passes the results to the task allocation adjustment module. The task allocation adjustment module dynamically adjusts task resources according to the evaluation results.

[0093] This modular design enables the system to flexibly respond to complex multi-source data collaborative processing needs in different scenarios.

[0094] In actual operation, the technical effects of the present invention are fully reflected.

[0095] For example, during rush hour, the system can quickly identify traffic accident alarm information and classify it as high-priority data, prioritizing the allocation of task resources for processing.

[0096] For low-priority data such as general road conditions, the system dynamically adjusts the task resource allocation of each area through regional division and collaborative performance evaluation, thereby improving overall traffic management efficiency.

[0097] In addition, the introduction of large models has significantly improved the intelligence level of the system. By understanding and extracting features from complex data, it has generated highly consistent data feature representations, providing a solid foundation for subsequent steps.

[0098] The application of the adaptive weight calculation method further improves the accuracy of the collaborative weights of data regions, making the collaborative performance evaluation more accurate.

[0099] Ultimately, through dynamic task allocation and resource adjustment, the system optimized cross-domain data integration capabilities, solving the problems of insufficient intelligence, lack of dynamic adaptability, and low efficiency of cross-platform data integration in existing technologies.

[0100] The contents not described in detail in the specification belong to the existing technology known to those skilled in the art, and the model parameters of each electrical appliance are not specifically limited, and conventional equipment can be used.

[0101] In this technical solution, the electrical control components not mentioned belong to the existing technology and are not shown in the figure and will not be described here. Example

[0102] The data collaborative processing system based on a large model includes a data collaborative processing method based on a large model, including the following steps: Step S1: Obtain multi-source heterogeneous data and preprocess them, extract and annotate data features through a large model, and generate a unified data feature representation; Step S2: Classify and label the data sources according to the data feature representation, build a data classification index table based on historical collaborative data, and divide the data into high-priority data and low-priority data; Step S3: Divide the low-priority data into regions, collect environmental information of each data region, and use an adaptive weight calculation method to obtain a collaborative weight coefficient for each data region; Step S4: Perform dynamic task allocation detection on the data of each data sub-region, determine the task execution efficiency of each data sub-region, comprehensively consider the collaborative weight coefficient and task execution efficiency, and use the weighted average method to evaluate the collaborative performance of each data sub-region; Step S5: adjusting the preset task allocation threshold according to the collaborative performance evaluation results of each data sub-region, and reallocating task resources for different data sub-regions.

[0103] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0104] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0105] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0106] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0107] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0108] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0109] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0110] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0111] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0112] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0113] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0114] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A data collaborative processing method based on a large model, characterized by: The following steps are involved: Step S1: Obtain multi-source heterogeneous data and pre-process them, extract and annotate data features through a large model, and generate a unified data feature representation in the form of a standardized feature vector; Step S2: Classify and label the data sources according to the data feature representation, build a data classification index table based on historical collaborative data, and divide the data into high-priority data and low-priority data; Step S3: Divide the low-priority data into regions, collect environmental information of each data region, and use an adaptive weight calculation method to obtain a collaborative weight coefficient for each data region; Step S4: Perform dynamic task allocation detection on the data of each data sub-region, determine the task execution efficiency of each data sub-region, comprehensively consider the collaborative weight coefficient and task execution efficiency, and use the weighted average method to evaluate the collaborative performance of each data sub-region; Step S5: adjusting the preset task allocation threshold according to the collaborative performance evaluation results of each data sub-region, and reallocating task resources for different data sub-regions.

2. The data collaborative processing method based on a large model according to claim 1, characterized in that: In step S1, when acquiring multi-source heterogeneous data, the original data is obtained from different data platforms through distributed data acquisition nodes, and the semantic understanding ability of the large model is used to parse and annotate the original data to generate a standardized feature vector as a unified data feature representation.

3. The data collaborative processing method based on a large model according to claim 2, characterized in that: In step S2, the historical collaborative data includes the task completion time and resource consumption of different data sources in the collaborative processing process. The historical database is accessed to obtain relevant information of all data sources recorded in the past collaborative processing process, which is converted into corresponding collaborative efficiency indicators. The data is classified and stored according to the data source to form a data classification index table.

4. The data collaborative processing method based on a large model according to claim 3 is characterized in that: In step S2, data sources corresponding to collaborative efficiency indicators exceeding a preset classification ratio are classified as high-priority data, otherwise they are classified as low-priority data.

5. The data collaborative processing method based on a large model according to claim 1 is characterized in that: In step S3, the distribution information of the data source is obtained through the distributed data collection nodes, and the coverage and data density of the data source are extracted from the data distribution information; The ratio of data density to coverage is used as the data distribution index to divide data sources into regions.

6. The data collaborative processing method based on a large model according to claim 5, characterized in that: In step S3, the environmental information of the data sub-region includes regional data traffic and regional data delay. The regional data traffic is detected by detecting the network bandwidth occupancy of each data sub-region at the same time point; Multiply the network bandwidth occupancy rate of the data sub-region by the number of network nodes in the corresponding data sub-region; The product result is used as the regional data flow of the data sub-region; Regional data delay is detected by setting the same time interval to test the network response time of each data sub-region; Calculate the response time difference between two adjacent time points and take the average value as the regional data delay of each data sub-region; The adaptive weight calculation method is used to obtain the collaborative weight coefficient of each data sub-region. The specific collaborative weight calculation coefficient formula is expressed as: ; Where, is the regional discharge coefficient, is the regional delay coefficient, 、 is the value corresponding to the optimal solution, 、 is the value corresponding to the worst solution, is the collaborative weight coefficient.

7. The data collaborative processing method based on a large model according to claim 1, characterized in that: In step S4, when performing dynamic task allocation detection on the data of each data sub-region, the task queue length of each data sub-region is first detected; Calculate the average value as the initial task load of the corresponding data sub-region, and allocate equal task resources to each data sub-region; After the same time interval, the task queue length of each data sub-region is tested again, and the average value is calculated as the final task load of the corresponding data sub-region; The difference between the final task load and the initial task load is used as the task execution efficiency of each data sub-region; The weighted average method is used to evaluate the collaborative performance of each data region by combining the collaborative weight coefficient and task execution efficiency. The specific formula is as follows: ; Where, is the collaborative performance evaluation value of the data region, To standardize task execution efficiency, To assign weights, is the collaborative weight coefficient.

8. The data collaborative processing method based on a large model according to claim 1, characterized in that: In step S5, the collaborative performance evaluation values of the data sub-regions are averaged to obtain an average collaborative performance evaluation value of the data sub-regions; The average value of the collaborative performance evaluation of the data sub-region is used as the screening benchmark, and the collaborative performance evaluation value of the data sub-region is compared with the screening benchmark; When the collaborative performance evaluation value of a data sub-region is lower than the screening benchmark, the ratio of the collaborative performance evaluation value of the corresponding data sub-region to the screening benchmark is used as the adjustment factor for the corresponding data sub-region; The preset task allocation threshold is multiplied by the adjustment factor to obtain the adjusted task allocation threshold.

9. A data collaborative processing system based on a large model, for implementing the data collaborative processing method based on a large model according to any one of claims 1 to 8, characterized in that: It includes data collection module, data classification module, area division module, collaborative performance evaluation module and task allocation and adjustment module; The data acquisition module is used to obtain raw data from multi-source data platforms and perform preprocessing to generate a unified data feature representation in the form of standardized feature vectors; The data classification module is used to classify and mark the data sources according to the data feature representation and build a data classification index table; The regional division module is used to divide the data source into regions according to the data distribution information and collect the environmental information of each data region; The collaborative performance evaluation module is used to evaluate the collaborative performance based on the environmental information and task execution efficiency of each data sub-region; The task allocation adjustment module is used to adjust the task allocation threshold according to the collaborative performance evaluation results of each data sub-region and reallocate task resources.

10. The data collaborative processing system based on a large model according to claim 9, characterized in that: The data acquisition module obtains raw data from different data platforms through distributed data acquisition nodes, and uses high-speed communication links to connect distributed data acquisition nodes to ensure the real-time and stability of data transmission.

Citation Information

Patent Citations

  • Forward transmission network resource allocation method and device

    CN112636995A

  • Resource scheduling strategy generation method and device and terminal equipment

    CN113986562A

  • Distributed computing power scheduling management system and method

    CN118132228A

  • News analysis method and system based on multi-modal large model

    CN118535978A

  • Resource scheduling method and system for containerized workloads in edge-cloud collaborative scene

    CN118708351A