Data collaboration processing method and system based on large model

By employing a large-model-driven dynamic adaptation mechanism and a multi-source data fusion strategy, the system addresses the shortcomings in intelligence and cross-platform integration of data collaborative processing in existing technologies. This enables efficient data collaborative processing and optimized resource allocation, thereby enhancing the system's intelligence and adaptability.

CN120494431BActive Publication Date: 2025-11-25FUJIAN RONGJI SOFTWARE +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510941365.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-11-25
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Existing technologies have shortcomings in terms of intelligence level, dynamic adaptability, and cross-platform data integration capabilities, resulting in limited data feature representation capabilities, imperfect dynamic adjustment mechanisms for data priorities, insufficient adaptive calculation of data collaborative weights, and a lack of methods for optimizing task allocation thresholds, which affect the level of intelligence and the efficiency of cross-platform data integration.

Method used

By introducing a dynamic adaptation mechanism driven by a large model, combined with multi-source data fusion and intelligent task allocation strategies, data features are extracted and labeled to generate a unified data feature representation. Data is then classified and divided into regions. The collaborative weight coefficient is obtained using an adaptive weight calculation method, and task resource allocation is optimized through collaborative performance evaluation and task allocation adjustment.

Benefits of technology

It has improved the intelligence and dynamic adaptability of data collaborative processing, enhanced cross-domain data integration capabilities and overall collaborative efficiency, and achieved more accurate regional collaborative performance evaluation and dynamic resource adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494431B_ABST
    Figure CN120494431B_ABST
Patent Text Reader

Abstract

The application discloses a data collaborative processing method and system based on a large model, relates to the technical field of data collaboration, and is used for solving the problems of insufficient intelligent level, lack of dynamic adaptability, and low cross-platform data integration efficiency. The method introduces a large model to perform feature extraction and labeling on multi-source heterogeneous data, generates unified feature representation, and provides a basis for data classification and task allocation. Through data classification marking, high and low priority data is pre-divided, the task adjustment range is reduced, and resource consumption is reduced. For low-priority data, a collaborative weight coefficient is obtained based on regional division, environmental information collection, and adaptive weight calculation, and the accuracy of the collaborative weight is improved. The task execution efficiency is dynamically detected, the collaborative performance is comprehensively evaluated in combination with the collaborative weight, the task allocation threshold is adjusted according to the evaluation result, the dynamic allocation of tasks and resources is realized, and the collaborative efficiency and cross-domain integration capability are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data collaboration, more particularly, the present application relates to a data collaboration processing method and system based on a large model. BACKGROUND

[0002] The data collaboration processing method and system gradually become an important tool for multi-field collaboration and intelligent decision support under the promotion of big data and artificial intelligence technology. However, the existing technology still has certain limitations in the intelligent level, dynamic adaptability and cross-platform data integration capability, which affects its actual application effect in complex scenarios.

[0003] The existing technology has the following deficiencies:

[0004] At present, the data feature representation capability is limited, the data priority dynamic adjustment mechanism is imperfect, the data collaboration weight adaptive calculation is insufficient, and the task allocation threshold optimization method is missing, resulting in insufficient intelligent level, lack of dynamic adaptability and low cross-platform data integration efficiency, therefore, a data collaboration processing method and system based on a large model is proposed.

[0005] The above information disclosed in the background section is only used to strengthen the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0006] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present application provide a data collaboration processing method and system based on a large model, by introducing a large model driven dynamic adaptation mechanism, combining multi-source data fusion and intelligent task allocation strategy, optimizing the cross-domain data integration capability, solving the problems of insufficient intelligent level, lack of dynamic adaptability and low cross-platform data integration efficiency mentioned in the background technology.

[0007] To achieve the above-mentioned purpose, the present application provides the following technical scheme, a data collaboration processing method and system based on a large model, comprising the following steps:

[0008] Step S1: acquiring multi-source heterogeneous data and preprocessing, extracting and labeling data features through a large model, generating standardized feature vector form unified data feature representation;

[0009] Step S2: classifying and marking the data sources according to the data feature representation, combining the historical collaboration data to build a data classification index table, and dividing the data into high priority data and low priority data;

[0010] Step S3: dividing the low priority data into regions, collecting the environmental information of each data sub-region, and using an adaptive weight calculation method to obtain the collaboration weight coefficient of each data sub-region;

[0011] Step S4: Dynamic task allocation detection is performed on the data of each data sub-region, the task execution efficiency of each data sub-region is determined, the coordination weight coefficient and the task execution efficiency are comprehensively coordinated, and the coordination performance of each data sub-region is evaluated by using a weighted average method;

[0012] Step S5: The preset task allocation threshold is adjusted according to the coordination performance evaluation result of each data sub-region, and the task resources are re-allocated for different data sub-regions.

[0013] In a preferred embodiment, in step S1, when acquiring multi-source heterogeneous data, the distributed data acquisition node acquires original data from different data platforms, and the semantic understanding ability of the large model is used to parse and label the original data to generate a standardized feature vector as a unified data feature representation.

[0014] In a preferred embodiment, in step S2, the historical coordination data includes the task completion time and resource consumption of different data sources in the coordination processing process, the historical database is accessed to obtain the relevant information of all data sources recorded in the past coordination processing process, which is converted into corresponding coordination efficiency indicators, and is classified and stored according to the data sources to form a data classification index table.

[0015] In a preferred embodiment, in step S2, the data sources corresponding to the coordination efficiency indicators exceeding the preset classification ratio are classified as high-priority data, otherwise they are classified as low-priority data.

[0016] In a preferred embodiment, in step S3, the distributed data acquisition node is used to acquire the distribution information of the data sources, and the coverage range and data density of the data sources are extracted from the data distribution information.

[0017] The ratio of data density to coverage range is taken as a data distribution index, which is used for regional division of the data sources.

[0018] In a preferred embodiment, in step S3, the environmental information of the data sub-regions includes regional data traffic and regional data delay, and the regional data traffic is detected by detecting the network bandwidth occupation of each data sub-region at the same time point.

[0019] The network bandwidth occupation rate of the data sub-region is multiplied by the number of network nodes corresponding to the data sub-region.

[0020] The product result is taken as the regional data traffic of the data sub-region.

[0021] The regional data delay is detected by detecting the network response time of each data sub-region at the same time interval.

[0022] The response time difference of two adjacent time points is calculated and averaged as the region data delay of each data sub-region;

[0023] The adaptive weight calculation method is used to obtain the synergy weight coefficient of each data sub-region, and the specific synergy weight calculation coefficient formula is expressed as:

[0024] ;

[0025] In the formula, is the region flow coefficient, is the region delay coefficient, , is the optimal solution corresponding value, , is the worst solution corresponding value, is the synergy weight coefficient.

[0026] In a preferred embodiment, in step S4, when detecting the dynamic task allocation of the data of each data sub-region, the task queue length of each data sub-region is first detected;

[0027] The average value is calculated as the initial task load of the corresponding data sub-region, and equal amount of task resources are allocated to each data sub-region;

[0028] After the same time interval, the task queue length of each data sub-region is detected again, and the average value is calculated as the final task load of the corresponding data sub-region;

[0029] The difference between the final task load and the initial task load is taken as the task execution efficiency of each data sub-region;

[0030] The synergy weight coefficient and the task execution efficiency are integrated, and the weighted average method is used to evaluate the synergy performance of each data sub-region, and the specific formula is expressed as follows:

[0031] ;

[0032] In the formula, is the synergy performance evaluation value of the data sub-region, is the standardized task execution efficiency, is the allocation weight, is the synergy weight coefficient.

[0033] In a preferred embodiment, in step S5, the synergy performance evaluation value of the data sub-region is averaged to obtain the synergy performance evaluation average value of the data sub-region;

[0034] The synergy performance evaluation average value of the data sub-region is taken as the screening reference, and the synergy performance evaluation value of the data sub-region is compared with the screening reference;

[0035] When the cooperative performance evaluation value of the data sub-region is lower than the screening reference, the ratio of the cooperative performance evaluation value of the corresponding data sub-region to the screening reference is taken as an adjustment factor of the corresponding data sub-region;

[0036] The preset task allocation threshold is multiplied by the adjustment factor to obtain an adjusted task allocation threshold.

[0037] The data cooperative processing system based on a large model comprises a data acquisition module, a data classification module, a region division module, a cooperative performance evaluation module and a task allocation adjustment module.

[0038] The data acquisition module is used for acquiring original data from a multi-source data platform and performing preprocessing to generate a unified data feature representation in the form of a standardized feature vector.

[0039] The data classification module is used for classifying and marking data sources according to data feature representations, and constructing a data classification index table.

[0040] The region division module is used for dividing data sources into regions according to data distribution information, and collecting environmental information of each data sub-region.

[0041] The cooperative performance evaluation module is used for evaluating cooperative performance according to environmental information and task execution efficiency of each data sub-region.

[0042] The task allocation adjustment module is used for adjusting a task allocation threshold according to the cooperative performance evaluation result of each data sub-region, and reallocating task resources.

[0043] In the data cooperative processing system based on a large model, the data acquisition module acquires original data from different data platforms through distributed data acquisition nodes, and connects the distributed data acquisition nodes by using a high-speed communication link to ensure the real-time performance and stability of data transmission.

[0044] Technical effects and advantages of the present application:

[0045] This invention introduces a large model to extract and label features from multi-source heterogeneous data, generating a unified data feature representation that lays the foundation for subsequent data classification and task allocation. By classifying and labeling data sources, data is pre-divided into high-priority and low-priority categories, narrowing the scope of task allocation adjustments and reducing resource consumption. When data is classified as low-priority, collaborative weight coefficients are obtained through regional division and environmental information collection, combined with an adaptive weight calculation method, improving the accuracy of regional collaborative weights. By dynamically detecting the task execution efficiency of each data region and comprehensively evaluating collaborative performance based on the collaborative weight coefficients and task execution efficiency, the actual collaborative capabilities of each data region can be more accurately reflected. Based on the collaborative performance evaluation results, the task allocation threshold is adjusted to achieve dynamic allocation of task resources, thereby improving overall collaborative efficiency and cross-domain data integration capabilities. Attached Figure Description

[0046] Figure 1 This is a flowchart illustrating the data collaborative processing method based on a large model according to the present invention.

[0047] Figure 2 This is a schematic diagram of the module structure of the data collaborative processing system based on a large model according to the present invention. Detailed Implementation

[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example

[0049] This invention provides a data collaborative processing method and system based on a large model. Its core lies in optimizing cross-domain data integration capabilities by introducing a dynamic adaptation mechanism driven by a large model and combining multi-source data fusion and intelligent task allocation strategies.

[0050] The following will be combined with the appendix Figure 1 and attached Figure 2 The specific structure and its labels are explained in detail.

[0051] like Figure 1 As shown, the overall process of the large-model-based data collaborative processing method, from data acquisition to task allocation and adjustment, includes several key steps, which are... Figure 2 The system modules shown in the image work together to complete the task.

[0052] The system comprises a data acquisition module, a data classification module, a region division module, a collaborative performance evaluation module, and a task allocation adjustment module, and the modules are logically connected to realize data flow and functional cooperation.

[0053] In the implementation process, first, the data acquisition module performs step S1, that is, acquires multi-source heterogeneous data and performs preprocessing.

[0054] The data acquisition module acquires raw data from different data platforms through distributed data acquisition nodes, which are distributed in different physical locations and connected to each other through high-speed communication links to ensure real-time and stability of data transmission.

[0055] After the raw data is transmitted to the data acquisition module, the semantic understanding ability of the large model is called to parse and label the data.

[0056] The large model uses deep learning algorithms to extract semantic information from the data and generates standardized feature vectors as unified data feature representations.

[0057] The key to this process is the semantic parsing ability of the large model, which can generate highly consistent feature representations through understanding of complex data, thereby providing a foundation for subsequent steps.

[0058] Then, step S2 is entered, and the data classification module classifies and labels the data sources according to the data feature representations.

[0059] The data classification module accesses the historical database to obtain all data source-related information recorded in the past collaborative processing process, and converts it into corresponding collaborative efficiency indicators.

[0060] These collaborative efficiency indicators are classified and stored according to data sources to form a data classification index table. Through a preset classification ratio, data sources corresponding to collaborative efficiency indicators exceeding the classification ratio are classified as high-priority data;

[0061] Otherwise, it is classified as low-priority data. This classification process relies on the task completion time and resource consumption of historical collaborative data to ensure high accuracy of the classification results.

[0062] The data classification module passes the data labeled by the classification to the region division module for further processing.

[0063] In step S3, the region division module divides the low-priority data into regions. The region division module first acquires the distribution information of the data sources through the distributed data acquisition nodes, extracts the coverage range and data density of the data sources, and divides the low-priority data into regions according to the distribution information.

[0064] The ratio of data density and coverage is taken as a data distribution index to divide the data sources into regions.

[0065] The region division module is also responsible for collecting environmental information of each data sub-region, including regional data traffic and regional data delay.

[0066] The regional data traffic is detected by detecting the network bandwidth occupation of each data sub-region at the same time point, multiplying the network bandwidth occupation rate of the data sub-region by the number of network nodes corresponding to the data sub-region, and taking the product as the regional data traffic of the data sub-region.

[0067] The regional data delay is detected by setting the same time interval to detect the network response time of each data sub-region, calculating the response time difference between adjacent two time points, and taking the average value as the regional data delay of each data sub-region.

[0068] The region division module transmits these environmental information to the cooperative performance evaluation module to provide input for subsequent calculation.

[0069] Step S4 is executed by the cooperative performance evaluation module, which mainly detects the dynamic task allocation of each data sub-region and determines the task execution efficiency of each data sub-region.

[0070] The cooperative performance evaluation module first detects the task queue length of each data sub-region, calculates the average value as the initial task load of the corresponding data sub-region.

[0071] By allocating equal amount of task resources to each data sub-region, the task queue length of each data sub-region is detected again after the same time interval, and the average value is calculated as the final task load of the corresponding data sub-region.

[0072] The difference between the final task load and the initial task load is taken as the task execution efficiency of each data sub-region. The cooperative performance evaluation module uses Min-Max standardization to process the task execution efficiency of each data sub-region to obtain the standardized task execution efficiency value.

[0073] At the same time, the cooperative performance evaluation module integrates the standardized task execution efficiency value and the cooperative weight coefficient of the corresponding data sub-region, and uses the weighted average method to evaluate the cooperative performance of each data sub-region.

[0074] The calculation of the cooperative weight coefficient depends on the regional traffic coefficient, the regional delay coefficient and the allocation weight provided by the region division module. The specific calculation process includes the calculation of the basic value, the determination of the optimal solution and the worst solution, the calculation of the Euclidean distance and the final solution of the cooperative weight coefficient.

[0075] The specific cooperative weight calculation coefficient formula is expressed as:

[0076] ;

[0077] wherein, is a regional flow coefficient, is a regional delay coefficient, is a corresponding value of the optimal solution, is a corresponding value of the worst solution, is a synergy weight coefficient;

[0078] The synergy performance of each data sub-region is evaluated by using a weighted average method based on the synergy weight coefficient and the task execution efficiency, and the specific formula is as follows:

[0079] ;

[0080] wherein, is the synergy performance evaluation value of the data sub-region, is the standardized task execution efficiency, is the assigned weight, is the synergy weight coefficient.

[0081] The synergy performance evaluation module transmits the evaluation result to the task allocation adjustment module for the next operation.

[0082] In step S5, the task allocation adjustment module adjusts the preset task allocation threshold according to the synergy performance evaluation result of each data sub-region.

[0083] The synergy performance evaluation values of the data sub-regions are averaged to obtain the synergy performance evaluation average value of the data sub-regions;

[0084] The task allocation adjustment module takes the synergy performance evaluation average value of the data sub-regions as a screening reference, and compares the synergy performance evaluation values of the data sub-regions with the screening reference.

[0085] When the synergy performance evaluation value of the data sub-region is higher than the screening reference, the preset task allocation threshold is not adjusted;

[0086] When the synergy performance evaluation value of the data sub-region is lower than the screening reference, the ratio of the synergy performance evaluation value of the corresponding data sub-region to the screening reference is taken as an adjustment factor of the corresponding data sub-region, and the preset task allocation threshold is multiplied by the adjustment factor to obtain an adjusted task allocation threshold.

[0087] The task allocation adjustment module monitors the task queue length of each data sub-region in real time, and when the task allocation threshold of the corresponding data sub-region is reached, the task resource is re-allocated for the corresponding data sub-region. ​​

[0088] The connection relationship and cooperation mode between the above-mentioned modules ensure efficient operation of the entire system.

[0089] The data acquisition module transmits raw data to the data classification module through distributed data acquisition nodes;

[0090] The data classification module transmits data to the region division module after generating a classification mark;

[0091] The region division module completes region division and transmits environmental information to the collaborative performance evaluation module;

[0092] The collaborative performance evaluation module transmits results to the task allocation adjustment module after evaluating collaborative performance;

[0093] The task allocation adjustment module dynamically adjusts task resources according to evaluation results.

[0094] This modular design enables the system to flexibly cope with complex multi-source data collaborative processing requirements in different scenarios.

[0095] In practical applications, the present application can be widely applied to cross-domain data integration scenarios, such as industrial Internet of Things, smart city, and medical information systems.

[0096] For example, in the traffic management scenario of a smart city, the data acquisition module can obtain real-time traffic data from multiple sensors and monitoring devices, the data classification module classifies the data into high-priority data (such as traffic accident alarms) and low-priority data (such as ordinary road condition information), the region division module divides the data region according to the traffic flow and delay characteristics of the city region, the collaborative performance evaluation module evaluates the collaborative performance of each region, and the task allocation adjustment module dynamically adjusts the task resource allocation of each region according to the evaluation results, thereby improving the overall traffic management efficiency.

[0097] As can be seen from the above specific embodiments, the present application realizes efficient collaborative processing of multi-source heterogeneous data through modular system design and detailed process steps, solving the problems of insufficient intelligence level, lack of dynamic adaptability, and low efficiency of cross-platform data integration in the prior art.

[0098] In order to better enable relevant persons in the art to fully understand and implement the present application, the specific implementation principles of the present application are further supplemented below in conjunction with a specific application scenario.

[0099] In the traffic management scenario of a smart city, the present application realizes efficient integration and dynamic task allocation of multi-source heterogeneous data through a data collaborative processing method and system based on a large model.

[0100] The operation process of the system starts from the data acquisition module, which acquires real-time traffic data from traffic monitoring cameras, sensor networks, and vehicle terminals through distributed data acquisition nodes.

[0101] These data include vehicle flow, speed, road congestion status, and traffic accident alarm information. Distributed data acquisition nodes are distributed in different areas of the city and are connected through high-speed communication links to ensure real-time and stability of data transmission.

[0102] After the raw data is transmitted to the data acquisition module, the semantic understanding ability of the large model is called to analyze and label the data.

[0103] For example, for traffic accident alarm information, the large model can identify its urgency and generate a standardized feature vector as a unified data feature representation.

[0104] The key to this process is the deep learning algorithm of the large model, which can extract semantic information from complex data, such as identifying accident type, location, and severity from text descriptions, thereby providing high-consistency support for subsequent steps.

[0105] Subsequently, the data classification module classifies and labels the data sources according to the data feature representation. The data classification module accesses the historical database to obtain all relevant information of the data sources recorded in the past collaborative processing process, and converts it into corresponding collaborative efficiency indicators.

[0106] For example, in the traffic management scenario, the collaborative efficiency indicator can reflect the response time and resource consumption of a certain data source in past tasks.

[0107] The data classification module stores these indicators according to data sources to form a data classification index table.

[0108] Through the preset classification ratio, the data sources corresponding to the collaborative efficiency indicators exceeding the classification ratio are classified as high-priority data, such as traffic accident alarm information;

[0109] Otherwise, it is classified as low-priority data, such as ordinary road condition information. This classification mechanism relies on the task completion time and resource consumption of historical collaborative data to ensure the accuracy of the classification results.

[0110] After classification, the data classification module transfers the classified and labeled data to the regional division module.

[0111] In the region division module, the low-priority data is divided into regions according to its distribution information. The region division module first obtains the coverage range and data density of the data source through the distributed data collection node, and calculates the ratio of the data density to the coverage range as the data distribution index.

[0112] For example, in the traffic management scenario, the data distribution index can reflect the intensity of traffic data in a certain region.

[0113] The region division module is also responsible for collecting the environmental information of each data sub-region, including the regional data flow and the regional data delay.

[0114] The regional data flow is detected by detecting the network bandwidth occupation of each data sub-region at the same time point, multiplying the network bandwidth occupation rate of the data sub-region by the number of network nodes corresponding to the data sub-region, and taking the product as the regional data flow of the data sub-region.

[0115] For example, during the peak period, the network bandwidth occupation rate of some regions is high, which will directly affect the data flow of the region.

[0116] The regional data delay is detected by setting the same time interval to detect the network response time of each data sub-region, calculating the difference between the response times of two adjacent time points, and taking the average value as the regional data delay of each data sub-region.

[0117] These environmental information is transmitted to the cooperative performance evaluation module to provide input for subsequent calculation.

[0118] The main task of the cooperative performance evaluation module is to detect the dynamic task allocation of each data sub-region and determine the task execution efficiency of each data sub-region.

[0119] In the traffic management scenario, the cooperative performance evaluation module first detects the task queue length of each data sub-region, calculates the average value as the initial task load of the corresponding data sub-region.

[0120] For example, in a certain region, the task queue length may reflect the number of traffic events that need to be processed in the region.

[0121] By allocating equal amount of task resources to each data sub-region, the task queue length of each data sub-region is detected again after the same time interval, and the average value is calculated as the final task load of the corresponding data sub-region.

[0122] The difference between the final task load and the initial task load is the task execution efficiency of each data sub-region. The cooperative performance evaluation module uses Min-Max standardization to process the task execution efficiency of each data sub-region to obtain the standardized task execution efficiency value.

[0123] Meanwhile, the cooperative performance evaluation module comprehensively evaluates the cooperative performance of each data sub-region by using a weighted average method on the basis of the standardized task execution efficiency value and the cooperative weight coefficient of the corresponding data sub-region.

[0124] The calculation of the cooperative weight coefficient depends on the regional traffic coefficient, the regional delay coefficient and the distribution weight provided by the region division module.

[0125] For example, if the regional traffic coefficient is high and the regional delay coefficient is low in a certain region, the cooperative weight coefficient of the region may be high, indicating that the region plays a more important role in the cooperative performance evaluation.

[0126] The cooperative performance evaluation module transmits the evaluation result to the task distribution adjustment module.

[0127] In the task distribution adjustment module, the preset task distribution threshold is adjusted according to the cooperative performance evaluation result of each data sub-region.

[0128] The task distribution adjustment module first calculates the average value of the cooperative performance evaluation of the data sub-region as a screening reference, and compares the cooperative performance evaluation value of the data sub-region with the screening reference.

[0129] For example, if the cooperative performance evaluation value is higher than the screening reference in a certain region, the preset task distribution threshold is not adjusted;

[0130] If the cooperative performance evaluation value is lower than the screening reference, the ratio of the cooperative performance evaluation value of the corresponding data sub-region to the screening reference is taken as an adjustment factor of the corresponding data sub-region, and the preset task distribution threshold is multiplied by the adjustment factor to obtain an adjusted task distribution threshold.

[0131] The task distribution adjustment module monitors the task queue length of each data sub-region in real time, and when the task distribution threshold of the corresponding data sub-region is reached, the task resource is redistributed for the corresponding data sub-region.

[0132] For example, if the task queue length exceeds the task distribution threshold in a certain region, the task distribution adjustment module will dynamically increase the task resource distribution of the region to relieve the task pressure.

[0133] As can be seen from the above steps, the present application realizes efficient cooperative processing of multi-source heterogeneous data in the traffic management scene of a smart city.

[0134] The data acquisition module transmits raw data to the data classification module through distributed data acquisition nodes, the data classification module generates classification labels and transmits data to the region division module, the region division module completes region division and transmits environmental information to the collaborative performance evaluation module, the collaborative performance evaluation module evaluates the collaborative performance and transmits the result to the task allocation adjustment module, and the task allocation adjustment module dynamically adjusts the task resources according to the evaluation result.

[0135] This modular design enables the system to flexibly cope with complex multi-source data collaborative processing requirements in different scenarios.

[0136] In actual operation, the technical effects of the present application are fully embodied.

[0137] For example, during peak traffic hours, the system can quickly identify traffic accident alarm information and classify it as high-priority data, and prioritize task resources for processing.

[0138] For low-priority data such as ordinary road condition information, the system dynamically adjusts the allocation of task resources in each region through region division and collaborative performance evaluation, thereby improving overall traffic management efficiency.

[0139] In addition, the introduction of large models significantly improves the intelligence level of the system, and through understanding and feature extraction of complex data, it generates highly consistent data feature representations, providing a solid foundation for subsequent steps.

[0140] The application of the adaptive weight calculation method further improves the accuracy of the data region collaborative weight, making the collaborative performance evaluation more accurate.

[0141] Finally, through dynamic task allocation and resource adjustment, the system optimizes the cross-domain data integration capability, solving the problems of insufficient intelligence level, lack of dynamic adaptability, and low efficiency of cross-platform data integration in existing technologies.

[0142] The contents not described in detail in the specification are all existing technologies known to those skilled in the art, and the model parameters of each appliance are not specifically limited, and conventional equipment can be used.

[0143] In this technical solution, the electric appliance control elements not mentioned belong to existing technologies, so they are not shown in the figure and will not be described here. Embodiment

[0144] The data collaborative processing system based on a large model includes a data collaborative processing method based on a large model, including the following steps:

[0145] Step S1: Obtain multi-source heterogeneous data and perform preprocessing, extract and label data features through a large model, and generate unified data feature representations;

[0146] Step S2: classifying and marking the data sources according to the data feature representation, constructing a data classification index table in combination with historical collaborative data, and dividing the data into high-priority data and low-priority data;

[0147] Step S3: dividing the low-priority data into regions, collecting environmental information of each data sub-region, and using an adaptive weight calculation method to obtain a collaborative weight coefficient of each data sub-region;

[0148] Step S4: performing dynamic task allocation detection on the data of each data sub-region, determining the task execution efficiency of each data sub-region, and comprehensively evaluating the collaborative performance of each data sub-region by using a weighted average method according to the collaborative weight coefficient and the task execution efficiency;

[0149] Step S5: adjusting the preset task allocation threshold according to the collaborative performance evaluation results of each data sub-region, and reallocating task resources for different data sub-regions.

[0150] The above formulas are all dimensionless numerical calculations, and the formulas are obtained by software simulation of a large amount of data to obtain a formula of the nearest real situation, and the preset parameters in the formula are set by a person skilled in the art according to the actual situation.

[0151] The above embodiments can be realized wholly or partially by software, hardware, firmware or any other combination. When realized by software, the above embodiments can be realized wholly or partially in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center and the like containing one or more available medium collections. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD) or a semiconductor medium. The semiconductor medium can be a solid-state disk.

[0152] It should be understood that the term "and / or" in this document is merely used to describe associated relationship, and it can mean three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In addition, the character " / " in this document generally means that the associated objects before and after the " / " are in an "or" relationship, but can also mean an "and / or" relationship, which can be understood according to the context before and after.

[0153] In this application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of the items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.

[0154] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0155] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0156] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0157] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be realized by other ways. For example, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed objects can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0158] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0159] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.

[0160] If the functions are realized in the form of software functional units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.

[0161] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for data collaborative processing based on a large model, characterized in that: The method comprises the following steps: Step S1: acquiring multi-source heterogeneous data and preprocessing, extracting and labeling data features through the semantic understanding ability of the large model, and generating a standardized feature vector form of unified data feature representation; Step S2: classifying and marking the data sources according to the data feature representation, combining the historical collaborative data to construct a data classification index table, and dividing the data into high-priority data and low-priority data; Step S3: dividing the low-priority data into regions, collecting the environmental information of each data sub-region, and using an adaptive weight calculation method to obtain the collaborative weight coefficient of each data sub-region; Step S4: dynamically allocating and detecting the data of each data sub-region, determining the task execution efficiency of each data sub-region, and comprehensively evaluating the collaborative performance of each data sub-region by combining the collaborative weight coefficient and the task execution efficiency using a weighted average method; Step S5: adjusting the preset task allocation threshold according to the collaborative performance evaluation results of each data sub-region, and reallocating task resources for different data sub-regions.

2. The data collaborative processing method based on a large model according to claim 1, wherein in step S1, when acquiring multi-source heterogeneous data, distributed data acquisition nodes are used to acquire original data from different data platforms, and the semantic understanding ability of the large model is used to analyze and label the original data, and a standardized feature vector is generated as a unified data feature representation.

3. The data collaborative processing method based on a large model according to claim 2, wherein in step S2, the historical collaborative data includes the task completion time and resource consumption of different data sources in the collaborative processing process, and the historical database is accessed to obtain the relevant information of all data sources recorded in the past collaborative processing process, which is converted into corresponding collaborative efficiency indicators, and stored according to the data sources to form a data classification index table.

4. The data collaborative processing method based on a large model according to claim 3, wherein in step S2, the data sources corresponding to the collaborative efficiency indicators exceeding the preset classification proportion are classified as high-priority data, and otherwise as low-priority data.

5. The data collaborative processing method based on a large model according to claim 1, wherein in step S3, the distribution information of the data sources is acquired through the distributed data acquisition nodes, and the coverage range and data density of the data sources are extracted from the data distribution information; The ratio of data density to coverage range is taken as a data distribution index for regional division of the data sources.

6. The data collaborative processing method based on a large model according to claim 5, wherein in step S3, the environmental information of the data sub-regions includes regional data traffic and regional data delay, and the regional data traffic is detected by detecting the network bandwidth occupation of each data sub-region at the same time point; The product of the network bandwidth occupation rate of the data sub-region and the number of network nodes of the corresponding data sub-region is calculated; The product result is taken as the regional data traffic of the data sub-region. ​ ​ ​ ​ ​ The regional data delay is detected by setting the same time interval for each data sub-region network response time; The response time difference between two adjacent time points is calculated, and the average value is taken as the regional data delay of each data sub-region; The adaptive weight calculation method is used to obtain the synergy weight coefficient of each data sub-region, and the specific synergy weight calculation coefficient formula is expressed as: ; In the formula, is a regional flow coefficient, is a regional delay coefficient, , is a value corresponding to an optimal solution, , is a value corresponding to a worst solution, is a synergy weight coefficient.

7. The large model-based data collaboration processing method of claim 1, characterized in that: In step S4, when detecting the dynamic task allocation of each data sub-region, the task queue length of each data sub-region is first detected; The average value is calculated as the initial task load of the corresponding data sub-region, and an equal amount of task resources is allocated to each data sub-region; After the same time interval, the task queue length of each data sub-region is detected again, and the average value is calculated as the final task load of the corresponding data sub-region; The difference between the final task load and the initial task load is taken as the task execution efficiency of each data sub-region; The synergy performance of each data sub-region is evaluated by using the weighted average method to combine the synergy weight coefficient and the task execution efficiency, and the specific formula is expressed as: In the formula, is a data sub-region collaborative performance evaluation value, is a normalized task execution efficiency, is an assigned weight, is a collaborative weight coefficient.

8. The large model-based data collaboration processing method of claim 1, characterized in that: In step S5, the synergy performance evaluation value of the data sub-region is averaged to obtain the synergy performance evaluation average value of the data sub-region; The synergy performance evaluation average value of the data sub-region is taken as the screening reference, and the synergy performance evaluation value of the data sub-region is compared with the screening reference; When the synergy performance evaluation value of the data sub-region is lower than the screening reference, the ratio of the synergy performance evaluation value of the corresponding data sub-region to the screening reference is taken as the adjustment factor of the corresponding data sub-region; The preset task allocation threshold is multiplied by the adjustment factor to obtain the adjusted task allocation threshold.

9. A large model-based data collaborative processing system for implementing the large model-based data collaborative processing method of any one of claims 1-8, characterized in that: It includes a data acquisition module, a data classification module, a region division module, a synergy performance evaluation module, and a task allocation adjustment module; The data acquisition module is used to obtain raw data from a multi-source data platform and perform preprocessing, extract and label data features through the semantic understanding ability of a large model, and generate a standardized feature vector form of unified data feature representation; The data classification module is used to classify and mark data sources according to data feature representation, and construct a data classification index table; The region division module is used to divide data sources into regions according to data distribution information, and collect environment information of each data sub-region; The synergy performance evaluation module is used to evaluate the synergy performance according to the environment information and task execution efficiency of each data sub-region; The task allocation adjustment module is used to adjust the task allocation threshold according to the synergy performance evaluation result of each data sub-region, and re-allocate task resources.

10. The large model-based data collaboration processing system of claim 9, characterized in that: The data acquisition module obtains raw data from different data platforms through distributed data acquisition nodes, and connects the distributed data acquisition nodes using a high-speed communication link to ensure the real-time and stability of data transmission.

Citation Information

Patent Citations

  • Strong-adaptation distributed data distribution method supporting dynamic expansion

    CN119960991A