Multi-modal data archiving processing system based on cloud computing
Through the multimodal data archiving and processing system based on cloud computing, the problem of difficult to uniformly package structured and unstructured data is solved, and standardized archiving and process link recording of multimodal data is realized, which improves data migration and verifiability, and improves the stability and processing efficiency of the system.
Patent Information
- Application Number
- CN202510890896.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
In the existing multimodal data management system, it is difficult to package and archive structured data and unstructured documents in a unified manner, resulting in a lack of general description of archived data, limiting the migrationability and repeatability of data packets, and the existing methods cannot form logically complete and semantic self-described archive packages.
A multimodal data archiving processing system based on cloud computing determines the archive type and priority through configuration modules, the resource module records the processing node status, the processing module extracts and maps structured fields and unstructured content, records the module record the operation relationship, synchronizes the module to bind state data, executes the module to generate two-layer archiving components, and archives the archive module to archive, realizing unified processing of structured and unstructured data and recording of process links.
It realizes standardized packaging for multimodal data archiving, supports rapid decoding and access, improves the semantic description ability and environmental robustness of the data, ensures the integrity and traceability of archived data, and improves processing efficiency and system stability.
Smart Images

Figure CN120386766A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a multi-modal data archiving and processing system based on cloud computing. Background Art
[0002] In existing multi-modal data management systems, structured data, unstructured documents, and meta-information are scattered, making it difficult to uniformly encapsulate and archive them. Some archiving processes rely on local tools and manual rule configurations, with complex implementation methods and poor environmental adaptability. Often, due to inconsistent data semantics and un-unified structure models, the archived data lacks a common description, restricting the portability and repeatable verification ability of data packets. Although some systems support data extraction and de-domain processing, in the archiving process, multi-modal content needs to be processed simultaneously, and existing methods may not be able to form an archived package with complete logic and self-describing semantics.
[0003] For example, there is a lack of a unified identification and reference mechanism between PDF files and XML structure forms, and the data content is disconnected from the process information, which may not be able to completely reproduce the original business logic and affect the authenticity of the archived data. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-modal data archiving and processing system based on cloud computing, aiming to solve the problems mentioned in the background art.
[0005] To solve the above technical problems, the technical solution of the present invention is as follows:
[0006] A multi-modal data archiving and processing system based on cloud computing, the system comprising:
[0007] A configuration module, configured to obtain an archiving request and an archiving task, and determine an archiving type, a processing priority, and a task number according to them to obtain task configuration data;
[0008] A resource module, configured to record the memory identifier, execution path, and cache address of a processing node before the archiving task is started to obtain a resource status data set;
[0009] A processing module, configured to extract a structured field group and an unstructured content block according to the task configuration data, generate a structure reference table by mapping the structured field group by fields, and generate a content time stream by splitting the unstructured content block in chronological order;
[0010] A recording module, configured to record each field mapping and time splitting operation to obtain operation event data, and determine its reference relationship with the processing node to obtain process link data;
[0011] A synchronization module, which is used to match the resource status data set with the processing nodes in the process link data, determine the task number, resource status, and data segment of each processing node, and obtain status-anchored data;
[0012] An execution module, which is used to combine the structure reference table and the content time flow with the status-anchored data into a two-layer archival component, and the two-layer archival component includes a data component flow and an operation link set;
[0013] An archival module, which is used to archive the two-layer archival component and generate process index data.
[0014] Furthermore, the configuration module includes:
[0015] A type judgment unit, which is used to receive an archival request, match and compare the type field of the archival task with the archival type list to obtain a matching item. When the matching item is greater than a preset matching threshold, the archival type is determined;
[0016] A path judgment unit, which is used to hierarchically expand the path field of the archival task, divide the path characters into multiple path segments with a hierarchical separator, and judge whether each path segment conforms to the path pattern list. When the result is yes, it marks the valid path segment and combines the valid path segments into an archival path structure;
[0017] A merging unit, which is used to merge the archival type with the archival path structure to obtain an archival source data identification item and assign an archival task number to it;
[0018] A priority unit, which is used to generate a processing priority according to the archival task and its task queue where it is located, and bind it to the archival task number to obtain task configuration data.
[0019] Furthermore, the resource module includes:
[0020] A node screening unit, which is used to call the cloud platform resource management interface according to the archival task number, obtain the current list of available processing nodes, and screen out the processing nodes whose idle degree and computing power meet the minimum execution standard from them to obtain a node candidate set;
[0021] A resource status unit, which is used to request the resource status of each processing node in the node candidate set, and call the node interface to extract the memory identifier, execution path, and cache address of the processing node to obtain resource status characteristics;
[0022] A binding unit, which is used to form a triple structure body with the resource status characteristics and bind it to the corresponding archival task number to obtain a resource status mapping item;
[0023] A number grouping unit, which is used to group resource status mapping items according to the archiving task number and assign a resource snapshot number to obtain a resource status data set.
[0024] Further, the processing module includes:
[0025] A path parsing unit, which is used to extract a structured field group from the archiving task, perform path parsing on each structured field according to the field path, extract each level of path label to obtain a field path stack, and set the value of the last-level field as the field end value to obtain a path value pair;
[0026] A sub-segment mapping unit, which is used to detect the path value pair, judge the sub-segment type of the field value, and call the corresponding field mapping structure according to it to obtain a structure reference table, and the structure reference table includes a path hierarchy structure, a field type label, and a standard field code;
[0027] A time division unit, which is used to extract an unstructured content block from the archiving task, sort it in chronological order and then divide it, cut it at a fixed time length and insert an anchor point at the cut point to obtain a content slice;
[0028] A sorting unit, which is used to arrange each content slice in order to obtain a content time stream, and the content time stream includes the start and end times of the paragraph, the number of words in the paragraph, the anchor field path, and the paragraph index value.
[0029] Further, the recording module includes:
[0030] A field operation unit, which is used to monitor each field mapping operation during the generation of the structure reference table, seal the path label, field type, and sub-segment mapping structure during path parsing, and record the task number where the operation occurs to obtain field operation data;
[0031] A paragraph operation unit, which is used to track each time slice operation during the generation of the content time stream, classify and identify the anchor points of each content slice, and record the paragraph index, the number of words in the paragraph, and the split number to obtain paragraph operation data;
[0032] A combination unit, which is used to combine and sort the field operation data and the paragraph operation data in the order of their recording times, and add an operation number to them to obtain an operation event data set;
[0033] A node mapping unit, which is used to query the processing node number of each operation event in the operation event data set in the resource status data set, determine its mapping relationship, and obtain process link data.
[0034] Further, the synchronization module includes:
[0035] A resource matching unit, configured to extract the processing node numbers of each operation event in the process link data, and query the triple structure with the same processing node number and task number in the resource status dataset to obtain a matching resource group;
[0036] A status binding unit, configured to use the memory identifier, execution path, and cache address of the matching resource group as status parameters, and associate them with the path label or paragraph index in the operation event to obtain a status binding item;
[0037] A marking merging unit, configured to mark the task number and data segment of the status binding item, generate a status anchor sub-item, and merge all status anchor sub-items into status anchor data according to the task number.
[0038] Further, the execution module includes:
[0039] A field anchoring unit, configured to retrieve the processing node number and memory address corresponding to the path label from the status anchor data according to the path label of each field path in the structure reference table, and construct a field anchoring pair;
[0040] A field binding unit, configured to locate the co-occurrence position of the paragraph index and the field path in the field anchoring pair in the content time flow, judge the logical relationship between the paragraph index and the field path, and construct a field binding item;
[0041] A combination unit, configured to combine the processing node number and memory address in the field anchoring pair with the field path and paragraph index in the field binding item into a four-element index block, and attach a task number to it to obtain a data component unit;
[0042] A sorting and grouping unit, configured to sort the data component units according to the hierarchical structure of the field path, and archive and group them according to the task number to obtain a data component stream.
[0043] Further, the execution module further includes:
[0044] An operation node unit, configured to extract the operation number, processing node number, and task number of each operation event from the status anchor data to obtain an operation node set;
[0045] An operation sequence unit, configured to sort the processing nodes in the operation node set in ascending order according to the event number, and group them based on the task number to obtain a task operation sequence;
[0046] An event linked list unit, configured to read the event number and data segment item by item in the task operation sequence, and establish a one-way event connection pointer between two adjacent operation events to obtain event linked list data;
[0047] An identification merging unit is used to add a head anchor identification and a tail end identification to each event linked list, construct task operation link entries, and combine all task operation link entries to obtain an operation link set.
[0048] Furthermore, the archive type list includes multiple archive type items. Each archive type item includes a type number field, a keyword set field, and an archive template identification field. The keyword set field includes several matching keywords related to this archive type. When determining the archive type, the type field in the archive task is extracted for keywords and compared with the keyword set field. When the number of matching keywords exceeds the preset matching threshold, it is determined that this archive type item is the archive type of the archive task.
[0049] The path pattern list includes multiple path template items. Each path template item includes a template number field, a path format field, a maximum level field, and a wildcard allowance field. The path format field is a legal path structure format, including one or more levels of path segments, and the path segments are divided by delimiters. After the path field in the archive task is expanded into multiple path segments, it is judged item by item whether the number and content of the path segments conform to the path format field and the maximum level field. When it conforms to the path template item, the path segment is marked as a valid path segment.
[0050] Furthermore, the data component stream is used to perform content backtracking of the archive task during the archiving process; the operation link set is used to perform operation backtracking of the archive task during the archiving process.
[0051] The above solution of the present invention has at least the following beneficial effects:
[0052] By determining the task type and priority, and classifying, extracting, and processing structured and unstructured data, the present invention uniformly constructs the two types of data into a two-layer archive component. The collaborative processing of this structured field group and unstructured content block no longer relies on traditional local archive tools for manual matching and integration, but completes the integration operation through a standardized data interface, enabling the archived data to have a unified structure, clear path levels, and attribution information, supporting the subsequent rapid decoding and access of different data, improving the packaging standardization degree of multi-modal data archiving, and solving the problems of scattered data and inconsistent packaging.
[0053] In the present invention, a structure reference table and a content time stream are established during the processing, and each processing node and task number are recorded, so that the corresponding relationship between content, structure and process is retained in the final archived component, making the archived data not only a collection of stock information, but also a compressed expression of the business processing path. The structure reference table describes the field path and tags, and the content time stream reflects the time process of unstructured data. These information together constitute the context semantics of the data, enabling the restoration of the semantics of the archived data when viewing or auditing the archived package.
[0054] The present invention realizes the perception and adaptation of the resource state in the current cloud computing environment by recording the memory identifier, execution path and cache address of the processing node. This binding relationship between the resource state and the task number is used for path positioning and execution scheduling in the subsequent archiving process. In scenarios such as node dynamic allocation and task migration that may occur in the cloud platform, it is possible to continue to stably complete the archiving task based on the recorded resource state data, ensuring the recovery ability after task interruption, thereby improving the environmental robustness and stability of the entire system.
[0055] The present invention performs field mapping and time slicing operations on structured field groups and unstructured content blocks respectively, realizes parallel processing logically, and then performs normalization processing and uniformly constructs a two-layer archived component. This processing flow of parallel first and then fusion has higher processing efficiency than the traditional serial data integration method, and at the same time avoids the problem of mutual interference between structured data and text content in the traditional method.
[0056] The present invention generates a two-layer archived component including a data component stream and an operation link set by fusing the structure reference table, the content time stream and the state anchor data, breaking the design limitation of the traditional archiving system that only archives data but not processes, and realizing the integration of content archiving and process archiving. This archived component can be stored as a standard data packet in the cloud and supports being parsed, verified and reconstructed by other archiving management platforms. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is a flowchart of a multi-modal data archiving processing system based on cloud computing provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0059] As Figure 1As shown in the figure, an embodiment of the present invention provides a multi-modal data archiving and processing system based on cloud computing. The system includes:
[0060] A configuration module, which is used to obtain an archiving request and an archiving task, and determine the archiving type, processing priority, and task number according to them to obtain task configuration data;
[0061] A resource module, which is used to record the memory identifier, execution path, and cache address of a processing node before the archiving task starts to obtain a resource status data set;
[0062] A processing module, which is used to extract a structured field group and an unstructured content block according to the task configuration data, generate a structure reference table by mapping the structured field group by fields, and generate a content time stream by splitting the unstructured content block in chronological order;
[0063] A recording module, which is used to record each field mapping and time splitting operation to obtain operation event data, and determine its reference relationship with the processing node to obtain process link data;
[0064] A synchronization module, which is used to match the resource status data set with the processing nodes in the process link data, determine the task number, resource status, and data segment of each processing node to obtain status anchoring data;
[0065] An execution module, which is used to combine the structure reference table and the content time stream with the status anchoring data into a two-layer archiving component. The two-layer archiving component includes a data component stream and an operation link set;
[0066] An archiving module, which is used to archive the two-layer archiving component to generate process index data.
[0067] In an embodiment of the present invention, the configuration module is used to obtain an archiving request and an archiving task, and determine the archiving type, processing priority, and task number according to them to obtain task configuration data, clarify the task execution strategy and priority, avoid resource contention and processing delay, and provide a basis for subsequent task scheduling; the resource module is used to record the memory identifier, execution path, and cache address of a processing node before the archiving task starts to obtain a resource status data set, establish a resource mapping between the archiving task and the running environment, and provide a basis for subsequent data processing operations; the processing module is used to extract a structured field group and an unstructured content block according to the task configuration data, generate a structure reference table by mapping the structured field group by fields, and generate a content time stream by splitting the unstructured content block in chronological order, and through field-level mapping and time series organization, achieve format regularization and clear logic within the data type, and solve the problem of the mixture of structured content and unstructured content and the lack of unified processing in the prior art.
[0068] A recording module is used to record each field mapping and time slicing operation, obtain operation event data, and determine its reference relationship with processing nodes to obtain process link data, realizing the traceability of each operation in the archiving task processing process and enhancing the traceability of data processing; A synchronization module is used to match the resource status data set with the processing nodes in the process link data, determine the task number, resource status, and data segment of each processing node to obtain status-anchored data, ensuring the correspondence between the archived content and the execution status and providing a basis for subsequent reconstruction tasks; An execution module is used to combine the structure reference table and the content time flow with the status-anchored data into a double-layer archiving component, and the double-layer archiving component includes a data component flow and an operation link set, which not only retains the structural content but also retains the generation process, having the dual evidence-preserving capabilities of content and process; An archiving module is used to archive the double-layer archiving component to generate process index data, ensuring that the archived data can be completely restored, verified, and analyzed.
[0069] Among them, the archiving module is used to archive the double-layer archiving component to generate process index data, specifically including:
[0070] First, receive the double-layer archiving component output by the execution module, which includes a data component flow and an operation link set. The system uses this double-layer archiving component as the archiving input unit and binds a corresponding task number, data component type identifier, and generation timestamp to each component. This binding operation is completed through an archiving configuration table, which is initialized based on the archiving type and task number recorded in the task configuration data and is used to manage the meta-information structure of the archiving output. Each data component unit in the data component flow contains a four-tuple content of field path, paragraph index, processing node number, and memory address. The system generates a unique storage reference identifier for each unit through a hash pointer mechanism and writes it into the archiving record index table.
[0071] After completing the index writing of the data component flow, the system continues to process the operation link set. The operation link set consists of a linked list structure sorted by operation number, and each chain represents the operation execution sequence of an archiving task, including fields such as chain head identifier, event number, processing node number, task number, and chain tail identifier. After receiving the operation link set, the archiving module first classifies and archives the linked list according to the task number and assigns a unique sequence identifier based on the linked list structure. Subsequently, the system proofreads the event number in each link with the operation event data table, extracts the associated field values, field paths, and slice anchor information to form an operation tracking node set. In this node set, each node contains the operation source path, node resource location, corresponding data component unit index value, and operation timestamp. The system combines this node set with the original task number to generate a process index entry and writes it into the process index table as the process-level index information of the archive package.
[0072] The process index data is finally output by the archiving module to the archiving metadata repository, forming a complete archiving package together with the double-layer archiving components. The system supports uniformly identifying, compressing, encrypting the archiving package and uploading it to the specified cloud storage target address. At the same time, the archiving index is registered in the metadata management platform, including the task number, archiving type, archiving component summary and generation time.
[0073] In a preferred embodiment of the present invention, the configuration module includes:
[0074] A type judgment unit, configured to receive an archiving request, match and compare the type field of the archiving task with the archiving type list to obtain a matching item. When the matching item is greater than a preset matching threshold, the archiving type is determined;
[0075] A path judgment unit, configured to hierarchically expand the path field of the archiving task, divide the path characters into multiple path segments with a hierarchical separator, and judge whether each path segment conforms to the path mode list. When the result is yes, it marks the valid path segment and combines the valid path segments into an archiving path structure;
[0076] A merging unit, configured to merge the archiving type with the archiving path structure to obtain an archiving source data identification item and assign an archiving task number to it;
[0077] A priority unit, configured to generate a processing priority according to the archiving task and its task queue, and bind it to the archiving task number to obtain task configuration data.
[0078] In an embodiment of the present invention, a type judgment unit is configured to receive an archiving request, match and compare the type field of the archiving task with an archiving type list to obtain a matching item. When the matching item is greater than a preset matching threshold, the archiving type is determined, which avoids the efficiency limitation caused by manual selection and effectively solves the problems of unclear archiving types and inconsistent classification criteria in the traditional archiving method; a path judgment unit is configured to hierarchically expand the path field of the archiving task, divide the path characters into multiple path segments with a hierarchical separator, and determine whether each path segment conforms according to a path pattern list. When the result is yes, it marks the valid path segments and combines the valid path segments into an archiving path structure, which avoids problems such as path chaos and hierarchical errors and provides a clear path reference for resource scheduling and index establishment of subsequent archiving tasks; a merging unit is configured to merge the archiving type with the archiving path structure to obtain an archiving source data identification item and assign an archiving task number to it, constructing a unique identification mechanism at the task level in the archiving system, which is convenient for task tracking, resource binding, and data structure association; a priority unit is configured to generate a processing priority according to the archiving task and its task queue where it is located, and bind it to the archiving task number to obtain task configuration data, realizing the reasonable allocation of resources and effectively avoiding the problems of resource blockage and critical task delay.
[0079] Among them, the priority unit is configured to generate a processing priority according to the archiving task and its task queue where it is located, and bind it to the archiving task number to obtain task configuration data, specifically including:
[0080] Based on the waiting times of all tasks in the archiving queue where the archiving task is located, and then normalizing them through the Sigmoid function to ensure that tasks with longer waiting times are more likely to be scheduled preferentially, obtaining the task waiting score item; according to the proportional relationship between the current system resource idle degree and the minimum resource requirements of the task, when resources are sufficient, it is inclined to process tasks with high resource occupancy in advance, obtaining the resource adaptability score item; constructing a product logarithm function based on the hierarchical depth of the field path and the length of the corresponding paragraph to quantify the difficulty of archiving data processing, ensuring that complex tasks obtain reasonable scheduling priorities, obtaining the data complexity score item; performing a weighted sum of the task waiting score item, the resource adaptability score item, and the data complexity score item to determine the positive drive of the archiving task processing priority, obtaining the value score item; calculating the clarity degree of the path structure of the archiving task according to the proportion of valid path fields, obtaining the path structure confidence penalty item; obtaining the processing priority according to the ratio of the value score item to the path structure confidence penalty item. When the proportion of path segments marked as valid in the archiving task is relatively high, the penalty item approaches 1 and has no obvious weakening effect on the processing priority. When the path segment structure is incomplete or the matching fails, this ratio decreases, resulting in an increase in the denominator, thereby weakening the task priority. To achieve consistent reference of each module in the task processing link, after generating the processing priority, the archiving task number is called, and this priority label is bound to the archiving task number one by one to form a complete task configuration data structure.
[0081] Among them, the calculation formula for the processing priority is: ,
[0082] Among them, is the processing priority of the archiving task, is the waiting time of the archiving task, is the average waiting time of all archiving tasks in the archiving queue, is the minimum waiting time in the archiving queue, is the amount of idle resources of the current cloud platform, is the minimum resource requirement of the archiving task, is the number of path fields in the archiving task, is the index of the path field, is the path field in the archiving task 's level, is the maximum level of the path field in the archiving task, is the path field in the archiving task 's length, is the minimum length of the path field in the archiving task, is the number of valid path segments in the archiving task, is the coefficient.
[0083] Among them, is the weight coefficient, and their sum is 1.
[0084] In the financial data archiving scenario, tasks are mostly high-frequency and high-density information such as transaction records, audit logs, and settlement vouchers. The financial system is highly sensitive to processing delays, and the waiting time of tasks must be an important consideration in scheduling; however, since such systems usually run in a dedicated environment with sufficient resource guarantees, the risk of resource shortage is relatively low, and resource adaptation items have relatively little impact on task scheduling; on the other hand, archiving tasks usually have concise fields and shallow structures, and complexity has little impact on the processing capacity requirements. Therefore, the proportion of this item can be appropriately reduced. Therefore, the task waiting score item should be significantly larger, the resource adaptation score item should be moderate, and the data complexity score item should be smaller, with values of 0.6, 0.3, and 0.1 respectively.
[0085] In the enterprise operation data archiving scenario, task types include system call logs, user behavior records, monitoring status updates, etc. The biggest bottleneck in archiving such data is usually resource scheduling rather than waiting time, because archiving tasks often trigger in batches, and system resource shortage becomes the key factor in scheduling and sorting; on the other hand, the field structure complexity of such tasks varies greatly, including nested objects, dynamic fields, etc., and the processing cost is significantly affected by the content. Therefore, the complexity item also needs to be considered key; and since enterprise systems focus on overall throughput efficiency, although the queuing time is meaningful, its impact on priority sorting is less than that of resources and content. Therefore, the resource adaptation score item is the largest, followed by the data complexity score item, and the task waiting score item, with values of 0.2, 0.5, and 0.3 respectively.
[0086] In the intelligent manufacturing data archiving scenario, archiving tasks come from multi-modal sources such as device operation data, sensor readings, and image acquisition information, and are often accompanied by a production control system with extremely high real-time requirements. In this environment, a slightly longer queuing waiting time may cause synchronization failure, seriously affecting the correctness of task scheduling. Therefore, the waiting time item must be treated as the dominant item; however, at the same time, the manufacturing scenario often uses edge computing nodes or resource-constrained terminals, and the resource status is highly unstable. The importance of resource adaptation and waiting time is almost equivalent; in addition, a large number of fields in the archived content have deep hierarchical structures (such as multi-channel data streams, periodic signals, etc.), and complexity will also significantly increase the processing load. Therefore, the gap between the three weights should not be too large, but generally, the task waiting score item and the resource adaptation score item are the dominant ones, with values of 0.4, 0.45, and 0.15 respectively.
[0087] Among them, is the penalty factor coefficient, and its value is 1.0. This value is set through quantitative analysis of the relationship between path confidence and priority. During the archival task scheduling process, the structural integrity of the path field directly affects whether the task can be accurately identified and processed. If there are many invalid structures in the path segment, it will lead to failures in subsequent field mapping, resource anchoring, and data component generation. Considering that when the path field confidence is insufficient, although the execution opportunity of the task should not be directly deprived, its priority still needs to be moderately suppressed. Therefore, is fixedly set to 1.0, so that when the proportion of valid paths is insufficient, the overall decline in priority is linearly weakened, which is neither overly punitive nor sufficient to distinguish between tasks with clear structures and tasks with fuzzy structures.
[0088] In a preferred embodiment of the present invention, the resource module includes:
[0089] A node screening unit, configured to call the cloud platform resource management interface according to the archival task number, obtain the list of currently available processing nodes, and screen out the processing nodes whose idle degree and computing power meet the minimum execution standard from them to obtain a node candidate set;
[0090] A resource status unit, configured to send a resource status request to each processing node in the node candidate set, and call the node interface to extract the memory identifier, execution path, and cache address of the processing node to obtain resource status characteristics;
[0091] A binding unit, configured to form the resource status characteristics into a triple structure body and bind it to the corresponding archival task number to obtain a resource status mapping item;
[0092] A number grouping unit, configured to group the resource status mapping items according to the archival task number and assign a resource snapshot number to obtain a resource status data set.
[0093] In an embodiment of the present invention, a node screening unit is configured to call a cloud platform resource management interface according to an archiving task number, obtain a list of currently available processing nodes, and screen out processing nodes with a free degree and computing power meeting the minimum execution standard from them to obtain a node candidate set, avoiding delays and failures caused by insufficient resources or scheduling conflicts during task processing; a resource status unit is configured to request the resource status of each processing node in the node candidate set, call a node interface to extract the memory identifier, execution path, and cache address of the processing node to obtain resource status characteristics, enabling the system to clarify in what computing environment the current archiving will run; a binding unit is configured to form a triple structure body with the resource status characteristics and bind it to the corresponding archiving task number to obtain a resource status mapping item, realizing the unique indexing ability of tasks to node resources and providing a basis for subsequent field path location of processing nodes; a number grouping unit is configured to group the resource status mapping items according to the archiving task number and assign a resource snapshot number to obtain a resource status data set, which is not only convenient for subsequent execution modules to call and reproduce, but also convenient for result verification and archiving after the task is completed.
[0094] Among them, the node screening unit is configured to call a cloud platform resource management interface according to an archiving task number, obtain a list of currently available processing nodes, and screen out processing nodes with a free degree and computing power meeting the minimum execution standard from them to obtain a node candidate set, specifically including:
[0095] The system will automatically start the node screening process. The node screening unit sends a query request to the resource control center according to the received archiving task number through a standard cloud platform API interface (such as a resource management API supporting the RESTful protocol), and calls the resource status synchronization module to obtain a list of processing nodes in the "idle", "partially idle", or "schedulable" state in the cloud computing platform at the current moment. In this list of processing nodes, each item contains basic information such as the unique identification code of the node, the number of CPU cores, the current CPU load, the total memory capacity, the used memory, whether GPU resources are bound, the GPU occupancy rate, the network latency score, and the task execution history, and is returned in a structured format for subsequent screening.
[0096] The system will then select processing nodes that meet the following two constraint conditions from the candidate node list according to the processing priority corresponding to the task: First, the current free degree of the node needs to be greater than or equal to a preset scheduling threshold, and the calculation method of the free degree is: (1 - current CPU load / total number of cores) x 100%; Second, the computing power of the node needs to reach the minimum standard for task execution, and this minimum standard is the configuration lower limit of the resource requirements required by the task preset model, including the minimum memory capacity (for example, not less than 16GB), whether the GPU is available, whether the network I / O meets the bandwidth threshold (for example, greater than 500Mbps), etc.
[0097] Among them, the number grouping unit is used to group the resource status mapping items according to the archiving task number, and assign a resource snapshot number to obtain a resource status data set, specifically including:
[0098] First, the system traverses all current resource status mapping items. According to the archiving task number field contained in each mapping item, it performs logical grouping by task number to ensure that the resource status mapping items in each group belong to the same archiving task. The grouping operation is implemented using a HashMap structure, with the task number as the primary key and the resource status mapping item list as the value set.
[0099] After completing the logical grouping of the resource status mapping items, the number grouping unit will generate a unique resource snapshot number for the resource status group under each task number. The resource snapshot number is a sortable and non-repeating timestamp number generated within the system (such as in the format of "RS20250614T154500001"), which is used to identify the resource environment snapshot during task execution. This number not only serves as the unique identifier for the binding status of this batch of resources but can also be used for reverse query and status reuse during task execution to ensure the consistency between task scheduling and resource scheduling.
[0100] In a preferred embodiment of the present invention, the processing module includes:
[0101] The path parsing unit is used to extract the structured field group from the archiving task, perform path parsing on each structured field according to the field path, extract each level of path label to obtain a field path stack, and set the value of the last level field as the field end value to obtain a path value pair;
[0102] The sub-segment mapping unit is used to detect the path value pair, judge the sub-segment type of the field value, and call the corresponding field mapping structure according to it to obtain a structure reference table, where the structure reference table includes a path hierarchy structure, a field type label, and a standard field code;
[0103] The time division unit is used to extract the unstructured content block from the archiving task, sort it in chronological order and then divide it, cut it into fixed time lengths and insert anchors at the cut points to obtain content slices;
[0104] The sorting unit is used to arrange each content slice in order to obtain a content time stream, where the content time stream includes the start and end times of the paragraph, the number of words in the paragraph, the anchor field path, and the paragraph index value.
[0105] In an embodiment of the present invention, a path parsing unit is configured to extract a structured field group from an archiving task, perform path parsing on each structured field according to a field path, extract each level of path label, obtain a field path stack, and set the value of the last level field as the field end value to obtain a path value pair, which can convert a traditional flat data structure into a tree-like field path stack, clearly express the relationship between fields, and lay a structural foundation for subsequent field mapping and comparison; a sub-segment mapping unit is configured to detect the path value pair, determine the sub-segment type of the field value, and call the corresponding field mapping structure according to it to obtain a structure reference table, where the structure reference table includes a path hierarchy structure, a field type label, and a standard field code, unify the expression method of heterogeneous fields, and provide a unified reference interface for subsequent field anchoring and cross-task comparison; a time division unit is configured to extract an unstructured content block from the archiving task, sort it in chronological order and then divide it, cut it into slices of a fixed time length and insert an anchor point at the cut point to obtain content slices, organize the unstructured content in terms of time dimension, solve the problem of disorder and non-traceability of unstructured content in traditional archiving, and avoid resource waste or processing exceptions caused by a single data block being too large or too small; a sorting unit is configured to arrange each content slice in order to obtain a content time stream, where the content time stream includes the start and end times of a paragraph, the number of words in the paragraph, the anchor field path, and the paragraph index value, which not only realizes the timing logic control of data archiving, but also provides a basis for subsequent data binding and event indexing.
[0106] Among them, the path parsing unit is configured to extract a structured field group from an archiving task, perform path parsing on each structured field according to a field path, extract each level of path label, obtain a field path stack, and set the value of the last level field as the field end value to obtain a path value pair, specifically including:
[0107] Read a list of structured field groups from the data structure of the archiving task. For each structured field in the list, extract it in the form of a key-value pair of "path string - value", where the path string is a combination of hierarchical labels separated by " / ", such as " / patient / record / bloodType"; the system splits the path string with " / " as the delimiter, extracts each level of path label step by step, and pushes it into the field path stack in order, and at the same time extracts and sets the final field value corresponding to the field as the field end value; subsequently, the system assembles the path stack and the field end value into a path value pair structure, and the path value pair includes: a list of path labels, the depth of the field hierarchy, the end value, and the unique path index identifier, which are used for the unique positioning and logical binding of subsequent fields.
[0108] Among them, the sub-segment mapping unit is used to detect path-value pairs, determine the sub-segment type of the field value, and call the corresponding field mapping structure according to it to obtain a structure reference table. The structure reference table includes a path hierarchy structure, a field type label, and a standard field code, specifically including:
[0109] The system performs a field value type recognition process on the end value in the path-value pair. According to the preset field value discrimination rule set, it identifies the sub-segment type corresponding to the field value through methods such as regular matching, literal type analysis, and context statistical analysis, including but not limited to: numeric type (integer, float), text type (string, text block), enumeration type (enum), time type (timestamp), etc.; after the sub-segment type recognition is completed, the system retrieves the corresponding field mapping structure in the structured field standard mapping library. This mapping structure maps each type of field value to the corresponding standard path template, field type label (such as "identification field", "numeric field"), and a globally unique standard field code; the system constructs a structure reference table based on this. The structure reference table is a multi-field structure record table, and each record corresponds to a path-value pair. Its fields include: complete path hierarchy information (i.e., the label combination from the root path to the end path), field type label (used to assist semantic indexing and extraction), and standard field code (used as the basic identifier for cross-task and cross-system data comparison). This structure reference table will serve as the key semantic interface for subsequent field binding and structure anchoring.
Claims
1. A multi-modal data archiving and processing system based on cloud computing, characterized in that, The system includes: A configuration module, which is used to obtain an archiving request and an archiving task, and determine the archiving type, processing priority, and task number according to them to obtain task configuration data; A resource module, which is used to record the memory identifier, execution path, and cache address of the processing node before the archiving task is started to obtain a resource status data set; A processing module, which is used to extract a structured field group and an unstructured content block according to the task configuration data, generate a structure reference table by mapping the structured field group by fields, and generate a content time stream by slicing the unstructured content block in chronological order; A recording module, which is used to record each field mapping and time slicing operation to obtain operation event data, and determine its reference relationship with the processing node to obtain process link data; A synchronization module, which is used to match the resource status data set with the processing nodes in the process link data to determine the task number, resource status, and data segment of each processing node to obtain status anchoring data; An execution module, which is used to combine the structure reference table and the content time stream with the status anchoring data into a two-layer archiving component, and the two-layer archiving component includes a data component stream and an operation link set; An archiving module, which is used to archive the two-layer archiving component to generate process index data.
2. The multi-modal data archiving and processing system based on cloud computing according to claim 1, wherein The configuration module includes: A type judgment unit, which is used to receive an archiving request, match and compare the type field of the archiving task with the archiving type list to obtain a matching item. When the matching item is greater than the preset matching threshold, the archiving type is determined; A path judgment unit, which is used to hierarchically expand the path field of the archiving task, divide the path characters into multiple path segments with a hierarchical separator, and judge whether each path segment conforms to the path pattern list. When the result is yes, it marks the valid path segment and combines the valid path segments into an archiving path structure; A merging unit, which is used to merge the archiving type and the archiving path structure to obtain an archiving source data identification item, and assign an archiving task number to it; A priority unit, which is used to generate a processing priority according to the archiving task and its task queue, and bind it to the archiving task number to obtain task configuration data.
3. The multi-modal data archiving and processing system based on cloud computing according to claim 2, characterized in that The resource module includes: A node screening unit, which is used to call the cloud platform resource management interface according to the archiving task number to obtain a list of currently available processing nodes, and screen out the processing nodes whose idle degree and computing power meet the minimum execution standard from them to obtain a node candidate set; A resource status unit, which is used to request the resource status of each processing node in the node candidate set, and call the node interface to extract the memory identifier, execution path, and cache address of the processing node to obtain resource status characteristics; A binding unit, which is used to form the resource status characteristics into a triple structure body and bind it to the corresponding archiving task number to obtain a resource status mapping item; A number grouping unit, which is used to group the resource status mapping items according to the archiving task number and assign a resource snapshot number to obtain a resource status data set.
4. The multi-modal data archiving and processing system based on cloud computing according to claim 3, wherein The processing module includes: A path parsing unit, which is used to extract a structured field group from an archiving task, perform path parsing on each structured field according to a field path, extract each level of path label, obtain a field path stack, and set the value of the last-level field as the field end value to obtain a path-value pair; A sub-segment mapping unit, which is used to detect the path-value pair, judge the sub-segment type of the field value, and call the corresponding field mapping structure according to it to obtain a structure reference table, where the structure reference table includes a path hierarchy structure, a field type label, and a standard field code; A time division unit, which is used to extract unstructured content blocks from an archiving task, sort them in chronological order and then divide them, cut them into fixed time lengths and insert anchors at the cut points to obtain content slices; A sorting unit, which is used to arrange each content slice in order to obtain a content time stream, where the content time stream includes the start and end times of a paragraph, the number of words in the paragraph, the anchor field path, and the paragraph index value.
5. The multi-modal data archiving and processing system based on cloud computing according to claim 4, wherein, The recording module includes: A field operation unit, which is used to listen to each field mapping operation during the generation of the structure reference table, seal the path label, field type, and sub-segment mapping structure during path parsing, and record the task number where the operation occurs to obtain field operation data; A paragraph operation unit, which is used to track each time slice operation during the generation of the content time stream, classify and identify the anchors of each content slice, and record the paragraph index, the number of words in the paragraph, and the segmentation number to obtain paragraph operation data; A combination unit, which is used to combine and sort the field operation data and the paragraph operation data in the order of their recording times, and add an operation number to them to obtain an operation event data set; A node mapping unit, which is used to query the processing node number of each operation event in the operation event data set in the resource status data set, determine its mapping relationship, and obtain process link data.
6. The multi-modal data archiving and processing system based on cloud computing according to claim 5, characterized in that The synchronization module includes: A resource matching unit, which is used to extract the processing node number of each operation event in the process link data, and query a triple structure with the same processing node number and task number in the resource status data set to obtain a matching resource group; A status binding unit, which is used to use the memory identifier, execution path, and cache address of the matching resource group as status parameters, and associate them with the path label or paragraph index in the operation event to obtain a status binding item; A marking merging unit, which is used to mark the task number and data segment of the status binding item, generate a status anchor sub-item, and merge all status anchor sub-items into status anchor data according to the task number.
7. The multi-modal data archiving and processing system based on cloud computing according to claim 6, wherein, The execution module includes: A field anchoring unit, which is used to retrieve the processing node number and memory address corresponding to the path label from the status anchor data according to the path label of each field path in the structure reference table, and construct a field anchoring pair; A field binding unit, which is used to locate the co-occurrence position of the paragraph index and the field path in the field anchoring pair in the content time stream, judge the logical relationship between the paragraph index and the field path, and construct a field binding item; Combination unit, which is used to combine the processing node number and memory address in the field anchor pair with the field path and paragraph index in the field binding item into a four - element index block, and attach a task number to it to obtain a data component unit; Sorting and grouping unit, which is used to sort the data component units according to the hierarchical structure of the field path and archive and group them according to the task number to obtain a data component stream.
8. The multi-modal data archiving and processing system based on cloud computing according to claim 7, wherein, The execution module further includes: Operation node unit, which is used to extract the operation number, processing node number and task number of each operation event from the status - anchored data to obtain an operation node set; Operation sequence unit, which is used to sort the processing nodes in the operation node set in ascending order according to the event number and group them based on the task number to obtain a task operation sequence; Event linked - list unit, which is used to read the event number and data segment one by one in the task operation sequence and establish a one - way event connection pointer between two adjacent operation events to obtain event linked - list data; Identification merging unit, which is used to add a head - anchor identification and a tail - end termination identification to each event linked - list, construct a task operation link entry, and combine all task operation link entries to obtain an operation link set.
9. The multi-modal data archiving and processing system based on cloud computing according to claim 8, characterized in that, The archive type list includes multiple archive type items. Each archive type item includes a type number field, a keyword set field and an archive template identification field. The keyword set field includes several matching keywords related to this archive type. When determining the archive type, the type field in the archive task is extracted for keywords and compared with the keyword set field. When the number of matching keywords exceeds the preset matching threshold, it is determined that this archive type item is the archive type of the archive task; The path pattern list includes multiple path template items. Each path template item includes a template number field, a path format field, a maximum level field and a wildcard permission field. The path format field is a legal path structure format, including one - level or multi - level path segments, and the path segments are divided by delimiters. After the path field in the archive task is expanded into multiple path segments, it is judged item by item whether the number and content of the path segments conform to the path format field and the maximum level field. When it conforms to the path template item, the path segment is marked as a valid path segment.
10. The multi-modal data archiving and processing system based on cloud computing according to claim 9, wherein, The data component stream is used to perform content backtracking of the archive task during the archiving process; the operation link set is used to perform operation backtracking of the archive task during the archiving process.
Citation Information
Patent Citations
Digital archives management method and system based on blockchain technology
CN107947922A
Heterogeneous multi-source system electronic file archiving method
CN114090591A
Method and system for independent application of fused and archived multi-modal data
CN119226236A
Cloud multi-modal data dynamic archiving system for big data
CN120179184A
Automated archival partitioning and synchronization on heterogeneous data systems
US20180225352A1
Cited By
Cloud platform data analyzing and processing system based on front-end segmentation
CN121117088A
Intelligent archive classified storage system
CN121580968A