Open source data cross-domain multi-modal resource extraction and conversion method based on pre-training large model fine tuning

By constructing a multimodal semantic description set and label offset-guided scheduling, the problem of unstable cross-modal alignment was solved, achieving high-precision cross-domain multimodal resource extraction and transformation, and improving the stability and efficiency of data processing.

CN122388044APending Publication Date: 2026-07-14神州数码数云科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610510721.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-17
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies suffer from cross-modal alignment instability in multi-source heterogeneous data processing, leading to decreased semantic matching and a lack of segment-by-segment consistency verification, which affects the integrity and usability of data representation.

Method used

By extracting text semantic direction and image structural contour features, a multimodal semantic description set is constructed. Semantic direction comparison and coordinate mapping are performed, label offset is adjusted to guide scheduling, path segmentation is constructed, modality alignment and path coherence determination are achieved, and cross-domain multimodal resource extraction and transformation results are generated.

Benefits of technology

It achieves high-precision structured extraction and transformation of cross-domain multimodal resources, improves the stability and efficiency of data processing, and reduces unnecessary computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122388044A_ABST
    Figure CN122388044A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data extraction processing, in particular to an open-source data cross-domain multi-modal resource extraction and conversion method based on pre-training of a large model, which comprises the following steps: recognizing text semantic directions and image structure contour features to construct a multi-modal semantic description set; performing semantic coordinate mapping and vector deviation comparison to realize cross-modal dynamic alignment; constructing a label offset guide scheduling result based on label difference intensity and processing module state; forming a consistency constraint judgment result by judging the stability of path segmentation connection features; and combining stable paths and conversion rules to complete storage and writing of resource extraction results and format conversion. The application can realize high-precision structured extraction and conversion of cross-domain multi-modal resources through linkage regulation and control of semantic driving, label scheduling, path constraint and cooperative control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data extraction and processing technology, and in particular to an open-source method for cross-domain multimodal data extraction and transformation based on fine-tuning of pre-trained large models. Background Technology

[0002] The field of data extraction and processing technology includes data acquisition, data parsing, data annotation, data fusion, and data structuring. The core of this technology lies in the unified expression and semantic alignment of multi-source heterogeneous data. Through text parsing, image feature extraction, audio signal parsing, and structured mapping, unstructured or semi-structured data is transformed into a computable structured data form. Its overall technical system covers key aspects such as data collection rule formulation, field-level information recognition, cross-modal semantic association, entity relationship extraction, and data format conversion. Furthermore, it achieves consistent expression and improved usability among different data types through rule definition, model reasoning, and semantic mapping mechanisms.

[0003] Among them, the open-source data cross-domain multimodal resource extraction and transformation method based on pre-trained large model fine-tuning refers to the parameter adjustment of pre-trained language models and vision models on open-source datasets to adapt them to specific domain data features, and to perform unified extraction and format conversion processing for multimodal data such as text, images, and audio. This technical matter covers pre-trained model weight loading, domain data annotation sample construction, word segmentation sequence encoding, image region feature encoding, multimodal feature alignment, cross-domain label mapping, entity attribute extraction, relation triple generation, and structured data format reconstruction. Specifically, it involves constructing a training sample set containing text annotation, image annotation, and audio transcription results, updating model parameters using a sequence-to-sequence learning approach, and using a unified semantic labeling system to identify entities and attributes in different modal data. Then, it converts the extraction results into a unified structured data record format according to preset field rules.

[0004] Existing technologies rely on a unified semantic labeling system and static mapping relationships for cross-modal alignment in the processing of multi-source heterogeneous data. However, in scenarios where data distribution changes or cross-domain differences are significant, semantic correspondence offsets are difficult to correct in a timely manner, leading to a decrease in the stability of matching between text and image information. For example, using the established semantic mapping method when the distribution of image features changes will result in incorrect associations. At the same time, the overall processing flow adopts a balanced processing strategy for information in each dimension, causing significantly different semantic parts to participate in the calculation without distinguishing between them and stable parts. This increases ineffective computational overhead and reduces processing efficiency. Furthermore, the lack of a segment-by-segment consistency verification mechanism in the generation of structured relationships makes it easy for discontinuous node connections or relationship mismatches to occur when constructing complex relationship chains, affecting the integrity and usability of the final data representation. Summary of the Invention

[0005] To address the technical problems existing in the prior art, this invention provides an open-source method for cross-domain multimodal resource extraction and transformation based on fine-tuning of a pre-trained large model. The technical solution is as follows: An open-source method for cross-domain multimodal data extraction and transformation based on fine-tuning of pre-trained large models includes the following steps: S1: Extract text and image data stored in the data center server, identify semantic direction information and structural contour features in each modality of data, organize semantic change trends and structural distribution patterns according to data units, establish the combination order of semantic descriptions, and generate a multimodal semantic description set. S2: Extract text and image features from the multimodal semantic description set, perform semantic direction comparison and coordinate mapping, determine the directional consistency of the two types of modal data when expressing the same object, adjust the semantic reference direction to achieve modal alignment, and generate cross-modal semantic alignment results; S3: Retrieve the label correspondence in the cross-modal semantic alignment result, determine the degree of difference and adjustment direction of cross-domain labels in different processing modules, filter the processing modules that meet the priority update conditions, and construct the label offset guidance scheduling result; S4: Based on the processing sequence in the label offset-guided scheduling result, perform path segmentation construction on the structured resource relationship, count the semantic connection features between adjacent path segments, determine the degree of node consistency in the relationship chain during the construction process, form a correspondence between path segments and semantic stability, and generate path segmentation consistency constraint judgment results.

[0006] As a further aspect of the present invention, the multimodal semantic description set includes semantic main axis direction, local structural trend, and semantic description order; the cross-modal semantic alignment result specifically includes direction vector deviation, distribution similarity, and semantic coordinate mapping matrix; the label offset guided scheduling result includes label difference mapping strength, processing module adjustment direction, and execution priority queue; and the path segment consistency constraint judgment result specifically refers to node connection mismatch rate, path backtracking trigger frequency, and path construction coherence.

[0007] As a further aspect of the present invention, the step of obtaining S1 is as follows: S101: Obtain text data files and image data files stored in the hard disk array, transmit them to the graphics processing unit through the network interface module, perform semantic analysis on the extracted sentence sequence from the text, compare the semantic vector obtained after analysis with the preset semantic axis feature library and semantic change trend library in turn, determine the semantic direction information of the text based on the comparison action, and associate it with the corresponding data identifier to obtain a semantic feature mapping table. S102: Based on the semantic direction information and data identifier in the semantic feature mapping table, extract the direction vector reflecting the structural contour and the range of change of local structure in each image region, and classify the frequency of occurrence of structural features in the overall image sampling points. Through this processing method, confirm the salience of structural features in spatial distribution and obtain the structural feature distribution parameter set. S103: Call the saliency distribution information in the structural feature distribution parameter group and associate it with the text semantic direction in the semantic feature mapping table. By the spatial mapping relationship and temporal position of semantic and structural features in the same resource identifier, determine the overlapping interval of semantic expression between modalities, and deduce the initial association path of multimodal data based on the interval information to obtain the multimodal semantic description set.

[0008] As a further aspect of the present invention, the step of obtaining S2 is as follows: S201: Based on the arranged feature content in the multimodal semantic description set, extract the projection vectors of text and image in a unified semantic coordinate system, deconstruct the direction angle of the projection vectors, match the deflection component of the text vector with the normal direction of the image structure item by item, calculate the deviation of feature items with the same semantic dimension, and obtain the semantic direction offset set. S202: Call the deviation values ​​and modal mapping relationships in the semantic direction offset set, perform dynamic alignment operation on each set of feature sequences in the multimodal semantic description set, and use the cosine value of the vector angle, the Euclidean distance distribution degree and the cross-modal feature overlap as reference factors to identify alignment points in the semantic space that can achieve directional coordination, and obtain the semantic alignment state matrix. S203: Based on the semantic relationship between text and image in the semantic alignment state matrix, the semantic direction is gradually fine-tuned, and feature items whose semantic offset remains within a preset threshold during the alignment process are selected. The feature identifiers are then bound to the index sequence in the unified coordinate system to obtain the cross-modal semantic alignment result.

[0009] As a further aspect of the present invention, the step of obtaining S3 is as follows: S301: Based on the feature identifiers selected from the cross-modal semantic alignment results, obtain the status information of the processing modules configured in the computing unit, and extract the label difference parameters, module adjustment vectors and computing load parameters bound to the identifiers. Call the feature identifiers and parameter groups to perform binding operations to obtain the scheduling decision parameter set. S302: Based on the value status of the tag difference parameter, module adjustment vector and computing load parameter in the scheduling decision parameter set, perform multiple sets of logical judgment operations, perform consistency screening on the value status, determine whether the module adjustment direction is similar to the tag difference direction, whether the computing load is within the idle threshold, and whether the tag difference exceeds the preset mutation boundary, and obtain the processing module status tag group. S303: Call the module identifier and status information in the status tag group of the processing module, perform priority aggregation operation on the module identifiers with high directional convergence, low load status and high difference intensity, and combine the identifier with the corresponding processing order content to obtain the tag offset guidance scheduling result.

[0010] As a further aspect of the present invention, the step of obtaining S4 is as follows: S401: Based on the processing order listed in the label offset guidance scheduling result, obtain the node connection record and structure rollback record generated during the path construction process, and perform state mapping on the two types of records according to the logical topology order, convert the node connection event and the verification failure event into a continuous logical sequence, form a corresponding construction trajectory for each relational path, and obtain a path operation state sequence set. S402: Based on the construction trajectory corresponding to each path identifier in the path operation state sequence set, perform logical judgment on the start and end nodes of adjacent connection segments, identify the connection relationship between semantic description and relation type in the sequence, and perform discrimination processing on connection consistency. Then, distinguish and encode the stable connection segments and unstable connection segments after discrimination to obtain the path coherence state discrimination table. S403: Call the status coding results corresponding to each path identifier in the path continuity status discrimination table, perform integrity judgment on the overall path in the resource extraction process, classify the continuity segment codes and non-continuous segment codes accordingly, form the continuity judgment relationship of resource relationship in the conversion process, and generate path segment consistency constraint judgment results.

[0011] As a further aspect of the present invention, the method further includes: S5: Combining the stable path, its structured type, and conversion rules in the path segmentation consistency constraint judgment result, the extraction result is stored, written, and formatted through the server bus, and the structured record generated in the multimodal resource processing system is output to generate cross-domain multimodal resource extraction and conversion results. The cross-domain multimodal resource extraction and transformation results include structural path index data, resource relationship mapping information, and the final transformed file set.

[0012] As a further aspect of the present invention, the step of obtaining S5 is as follows: S501: Based on the path identifier with a stable degree of construction in the path segmentation consistency constraint judgment result, obtain the corresponding structured conversion rules and resource extraction relationship data, and perform sequential verification of resource data frames according to the path index. Perform consistency judgment on the logical identifier of the data frame and the structured path sequence to obtain the structured conversion delivery sequence. S502: Based on the resource data frames arranged in the structured conversion delivery sequence, obtain the current write status parameters of the storage module, match and judge the resource data frames with the storage parameters, send the data frames that can be converted to the storage array in sequence through the server bus, and at the same time, link the status of the conversion confirmation signal and the bus response signal to obtain the extraction conversion response trajectory. S503: Call the resource identifier and conversion status information in the extraction and conversion response trajectory, confirm the status of the data frames that have been written and the conversion records that have generated responses, associate and collect the corresponding resource content with its structured conversion status, form the content range of the system that has been extracted and structured converted, and generate cross-domain multimodal resource extraction and conversion results.

[0013] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this invention, a multimodal semantic description set is constructed by extracting text semantic direction and image structural contour features. A dynamic alignment relationship between different modal data is established by combining semantic coordinate mapping and vector direction deviation comparison. In conjunction with the priority scheduling of label difference intensity and processing module adjustment direction, and the mismatch rate and backtracking features reflected by segmented connection features in the linkage path construction process, the path continuity status in the resource extraction process is comprehensively judged. Through the linkage and regulation between semantic main axis driving, label offset scheduling, path segmentation constraints and software and hardware collaborative control, the high-precision structured extraction and transformation results of cross-domain multimodal resources are promoted. Attached Figure Description

[0014] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a flowchart illustrating the acquisition process of S1 in this invention; Figure 3 This is a flowchart illustrating the acquisition process of S2 in this invention; Figure 4 This is a flowchart illustrating the acquisition process of S3 in this invention; Figure 5 This is a flowchart illustrating the acquisition process of S4 in this invention; Figure 6 This is a flowchart of the acquisition process for S5 of the present invention. Detailed Implementation

[0015] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0016] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0017] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0018] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0019] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0020] Please see Figure 1 This invention provides a technical solution: an open-source data cross-domain multimodal resource extraction and transformation method based on pre-trained large model fine-tuning, comprising the following steps: S1: Extract text and image data stored in the data center server, identify semantic direction information and structural contour features in each modality of data, organize semantic change trends and structural distribution patterns according to data units, establish the combination order of semantic descriptions, and generate a multimodal semantic description set. S2: Extract text and image features from the multimodal semantic description set, perform semantic direction comparison and coordinate mapping, determine the directional consistency of the two types of modal data when expressing the same object, adjust the semantic reference direction to achieve modal alignment, and generate cross-modal semantic alignment results; S3: Retrieve the label correspondence in the cross-modal semantic alignment results, determine the degree of difference and adjustment direction of cross-domain labels in different processing modules, filter the processing modules that meet the priority update conditions, and construct the label offset guidance scheduling results; S4: Based on the processing sequence in the label offset-guided scheduling result, perform path segmentation construction on the structured resource relationship, count the semantic connection features between adjacent path segments, judge the consistency of nodes in the relationship chain during the construction process, form the correspondence between path segments and semantic stability, and generate path segmentation consistency constraint judgment results; S5: Combining the stable path, its structured type, and conversion rules in the path segmentation consistency constraint judgment result, the extraction result is stored, written, and formatted through the server bus. The structured record generated in the multimodal resource processing system is output, generating cross-domain multimodal resource extraction and conversion results.

[0021] The multimodal semantic description set includes semantic main axis direction, local structural trend, and semantic description order. The cross-modal semantic alignment results specifically include direction vector deviation, distribution similarity, and semantic coordinate mapping matrix. The label offset-guided scheduling results include label difference mapping strength, processing module adjustment direction, and execution priority queue. The path segment consistency constraint judgment results specifically include node connection mismatch rate, path backtracking trigger frequency, and path construction coherence. The cross-domain multimodal resource extraction and transformation results include structural path index data, resource relationship mapping information, and the final transformation file set.

[0022] Please see Figure 2 The steps to obtain S1 are as follows: S101: Obtain text data files and image data files stored in the hard disk array, transmit them to the graphics processing unit through the network interface module, perform semantic analysis on the extracted sentence sequence from the text, compare the semantic vector obtained after analysis with the preset semantic axis feature library and semantic change trend library in turn, determine the semantic direction information of the text based on the comparison action, and associate it with the corresponding data identifier to obtain a semantic feature mapping table. During data acquisition, high-bandwidth SATA or NVMe network interface modules are used to deeply access the distributed hard disk array of the data center server. This process employs an asynchronous, non-blocking I / O scheduling mechanism to retrieve multi-source heterogeneous text and image data files in a high-concurrency manner, ensuring low latency and high throughput during memory mapping of massive files. For the acquired text data, a deep learning word segmenter based on the Transformer architecture is introduced for serialization parsing. Through a self-attention mechanism based on the context, highly expressive high-dimensional semantic feature vectors are accurately extracted. Subsequently, these vectors are loaded in batches into the high-speed video memory of the graphics processing unit (GPU). The semantic principal axis feature library and semantic change trend library, pre-residing in the video memory, are invoked. Under the acceleration of tens of thousands of GPU cores, an optimized cosine similarity algorithm is used for full matrix comparison. When the cosine value of the angle between the semantic vector and the principal axis feature space is greater than a preset adaptive threshold, it can be determined with high confidence that the text fragment contains a specific semantic direction. Finally, a globally unique UUID data identifier is generated using a distributed algorithm, which is then deeply bound to the identified semantic direction and persistently written into a memory-level semantic feature mapping table, thereby constructing a text semantic benchmark with extremely strong structured attributes.

[0023] S102: Based on the semantic direction information and data identifier in the semantic feature mapping table, extract the direction vector reflecting the structural contour and the range of change of local structure in each image region, and classify the frequency of occurrence of structural features in the overall image sampling points. Through this processing method, confirm the salience of structural features in spatial distribution and obtain the structural feature distribution parameter set. Based on the data identifiers in the established semantic feature mapping table, high-precision cross-domain image feature extraction is performed. First, a deep residual network (such as an improved ResNet architecture) is used to perform multi-scale convolution processing on the retrieved associated image files to accurately extract the gradient direction vectors and edge normal features that reflect the physical object's structural contour. Next, fine-grained spatial distribution classification and parameterized quantization are performed on the variation range of these local structures. The specific numerical calculation logic includes two dimensions: in the local dimension, a sliding window mechanism is used to count the number of pixels contained in a specific structural feature (e.g., high-frequency straight edges or closed contours) within a local image region as the numerator, and the total number of pixels in that independent region as the denominator, performing floating-point division to accurately obtain the local structure density index; in the global dimension, the absolute total frequency of this feature in all sampled points of the entire image is comprehensively scanned. Subsequently, a statistical model is introduced to deeply calculate the variance and expected value (mean) of the spatial distribution of these features. Based on the magnitude of variance and the degree of aggregation of density, the significance level of structural features in spatial topology is strictly confirmed (such as extremely high significance, high significance, or low significance), and finally the structural feature distribution parameter set is encapsulated and output.

[0024] Table 1: Calculation Table of Image Structural Feature Distribution Parameters As shown in Table 1, the saliency of different feature terms in spatial distribution was clarified through quantitative analysis of image sampling points, and a set of structural feature distribution parameters was formed.

[0025] S103: Call the saliency distribution information in the structural feature distribution parameter group and associate it with the text semantic direction in the semantic feature mapping table. By the spatial mapping relationship and temporal position of semantic and structural features in the same resource identifier, determine the overlapping interval of semantic expression between modalities, and deduce the initial association path of multimodal data based on the interval information to obtain the multimodal semantic description set. The algorithm retrieves the core information about the high-significance distribution from the structural feature distribution parameter set generated in the previous step and performs deep association and temporal tracking with the text semantic direction residing in the semantic feature mapping table. At this stage, a unified virtual multimodal coordinate system is first established, rigidly aligning the text feature vectors under the same Resource Unique Identifier (UUID) with the image visual feature vectors in the coordinate space. Subsequently, a complex association tracking engine is activated, utilizing a cross-attention calculation mechanism to quantify and evaluate the temporal overlap between regions with significant physical structures in the image and the core semantic keywords extracted from the text in terms of descriptive logic and timestamp dimensions. If the overlap rate of the mapping coordinates of the two modalities in the feature manifold space is detected to be significantly higher than a preset confidence boundary (e.g., a 0.75 threshold), the algorithm determines that these features constitute a strong overlap interval in the semantic expression between the multimodalities. Based on the extracted high-value interval information, a graph inference algorithm is used to deeply deduce the implicit association path between the logical order of text description and the changing trend of the image's underlying structure. These validated paths and features are assembled together to ultimately generate a set of rigorously structured multimodal semantic descriptions, providing high-quality data support for subsequent fine-tuning of large cross-domain models.

[0026] Please see Figure 3 The steps to obtain S2 are as follows: S201: Based on the arranged feature content in the multimodal semantic description set, extract the projection vectors of text and image in a unified semantic coordinate system, deconstruct the direction angle of the projection vector, match the deflection component of the text vector with the normal direction of the image structure item by item, calculate the deviation of feature items with the same semantic dimension, and obtain the semantic direction offset set. Based on the constructed multimodal semantic description set, the sorted feature content in text and image data is uniformly mapped to a high-dimensional semantic metric coordinate system. To accurately capture subtle deviations between modalities, a deep vector deconstruction operation is performed, utilizing Principal Component Analysis (PCA) or Singular Value Decomposition (SVD) techniques to separate the components of the multimodal projection vectors along each orthogonal basis direction. Subsequently, a feature matching engine is activated to rigorously compare the deconstructed text feature deviation components with the normal direction vectors of the image structure in the topological manifold space. During this process, relative deviation calculations are performed in high-dimensional space specifically for feature groups with the same semantic dimension. Specifically, the absolute deviation between the text deviation vector and the image normal vector in angular space, as well as the Euler angle difference, are calculated. Through this series of complex geometric and algebraic operations, a precise quantitative description of semantic misalignments arising when different modalities express the same business object can be achieved, outputting a semantic direction offset set containing detailed deviation values ​​for each cross-modal feature pair, providing precise target coordinates for subsequent fine-tuning and alignment.

[0027] S202: Call the deviation values ​​and modal mapping relationships in the semantic direction offset set, perform dynamic alignment operation on each set of feature sequences in the multimodal semantic description set, and use the cosine value of the vector angle, the Euclidean distance distribution degree and the cross-modal feature overlap as reference factors to identify alignment points in the semantic space that can achieve directional coordination and obtain the semantic alignment state matrix. A dynamic optimization alignment mechanism is initiated on the full feature sequence in the multimodal semantic description set by invoking a semantic direction offset set containing precise deviation values ​​and modal mapping relationships. This core process relies on a multi-dimensional weighted fusion evaluation algorithm to calculate the dynamic alignment score of each set of modal features in real time. Three key reference factors are set: the cosine of the vector angle represents directional consistency, the reciprocal of the Euclidean distance distribution degree represents spatial proximity, and the cross-modal feature overlap rate represents content coverage. The algorithm assigns specific dynamic weights to these three factors (e.g., setting weight ratios of 0.4, 0.3, and 0.3) and performs a summation operation. During the calculation, a heuristic search algorithm traverses the entire high-dimensional semantic space, iteratively adjusting the relative positions of feature vectors to find the optimal coordinate system spatial point that maximizes the weighted evaluation score. These identified extreme points represent the alignment benchmark points where multimodal data can achieve perfect directional coordination at the semantic level. The optimal scores of all features and their corresponding spatial coordinates are arrayed and stored, and the final output is a semantic alignment state matrix containing the global optimal solution state.

[0028] S203: Based on the semantic relationship between text and image in the semantic alignment state matrix, the semantic direction is gradually fine-tuned, and feature items whose semantic offset remains within a preset threshold during the alignment process are selected. The feature identifiers are then bound to the index sequence in the unified coordinate system to obtain the cross-modal semantic alignment results. The semantic alignment state matrix records the semantic scoring relationship between text and images, initiating a parameter-level, stepwise fine-tuning correction process for the overall semantic direction. Based on the global error distribution presented in the matrix, a strict dynamic offset tolerance threshold is set (e.g., a maximum allowable error of less than 0.15). Subsequently, the engine performs high-concurrency item-by-item screening and filtering on the massive feature pairs within the matrix, decisively eliminating features that cannot be effectively integrated due to significant differences in underlying modal logic, retaining only high-quality feature items with extremely high alignment accuracy and minimal error. For these selected core feature items, the underlying layer uses gradient descent fine-tuning technology to perform extremely subtle direction calibration, completely eliminating residual semantic offsets caused by feature extraction. After confirming that the fine-tuning is satisfactory, a hard memory binding operation is performed, physically locking the globally unique identifier of the fine-tuned cross-modal features with the absolute index sequence allocated in a unified multi-dimensional coordinate system. This operation completely breaks down the modal barriers between text and images, ensuring that multimodal data can accurately share the same absolute semantic origin when finally performing structured knowledge extraction, thereby generating the final cross-modal semantic alignment result.

[0029] Please see Figure 4 The steps to obtain S3 are as follows: S301: Based on the feature identifiers selected from the cross-modal semantic alignment results, obtain the status information of the processing modules configured in the computing unit, and extract the label difference parameters, module adjustment vectors and computing load parameters bound to the identifier. Call the feature identifier and parameter group to perform binding operations to obtain the scheduling decision parameter set. Based on the rigorously selected and locked feature identifiers in the cross-modal semantic alignment results, the main control program accesses the global resource manager of the large-scale heterogeneous computing units through the internal high-speed RPC bus. It polls in real time to obtain the instantaneous running status and lifecycle information of all configured underlying processing modules (including feature parsing modules, modality fusion modules, knowledge mapping modules, etc.) in the current architecture. On this basis, it performs deep context parameter binding extraction operations, accurately extracting cross-domain label difference strength parameters closely related to specific feature identifiers, adjustment vector data guiding module logic changes, and CPU and GPU core computing load rates reflecting the current hardware computing power bottlenecks from the cache. These real-time parameters representing the dynamic running environment status are then combined and formatted in a multi-dimensional manner with the feature identifiers representing static data attributes. This process effectively breaks down the information silos between the data processing flow and the underlying hardware resource monitoring, successfully constructing a scheduling decision parameter set that balances the scale of semantic differences with the underlying hardware carrying capacity, providing core data support for effectively addressing the potential surge in computational overhead caused by cross-domain conversion.

[0030] S302: Based on the value status of the tag difference parameter, module adjustment vector and computing load parameter in the scheduling decision parameter set, execute multiple sets of logical judgment operations, perform consistency screening on the value status, determine whether the module adjustment direction is similar to the tag difference direction, whether the computing load is within the idle threshold, and whether the tag difference exceeds the preset mutation boundary, and obtain the processing module status tag group. The system reads key values ​​such as tag difference parameters, module adjustment vectors, and computational load, which are encapsulated in the scheduling decision parameters. It then activates its internal high-performance logical decision rule engine, executing multiple sets of rigorous Boolean logic judgments and filters in parallel. First, it performs a directional trend screening: calculating the spatial inner product of the module adjustment vector and the cross-domain tag difference vector. If the result is positive, it confirms that the adjustment directions of the two are converging. Second, it performs a computational load screening: real-time verification of the computational load parameters to determine if they are securely within the preset 70% safe idle level to prevent scheduling overload. Finally, it performs a mutation anomaly screening: calculating the absolute value of the cross-domain tag difference to determine if it exceeds the preset 0.5 level semantic mutation boundary. The outputs of these three sets of logical judgments are combined and, based on complex truth table matching rules, refined status labeling is applied to different processing modules. For example, modules that meet the criteria of directional convergence, good load, and significant difference mutations are labeled "high priority" and "requires urgent correction," thus comprehensively obtaining a set of processing module status labels reflecting the current business urgency and resource status.

[0031] Table 2: Processing Module Scheduling Logic Verification Table As shown in Table 2, the execution priority of each processing module in the cross-domain conversion process was clarified through multi-dimensional logical comparison.

[0032] S303: Call the module identifier and status information in the status tag group of the processing module, perform priority aggregation operation on the module identifiers with high directional convergence, low load status and high difference intensity, and combine the identifier with the corresponding processing order content to obtain the tag offset guided scheduling result; The system invokes the detailed module identifiers and Boolean state judgment information recorded in the processing module status label group to initiate the underlying dynamic resource scheduler for task rearrangement. For module identifiers that simultaneously possess high directional convergence, are in a low computational load state, and face high-intensity label semantic differences, a preemptive priority aggregation and escalation operation is performed, forcibly moving these urgently needing correction module identifiers to the head of the global execution queue. After establishing the execution order, container orchestration technology is used to strongly bind and sandbox these high-priority module identifiers with the accompanying large-model cross-modal fine-tuning scripts, structured transformation logic rules, and other processing order content. This automated scheduling mechanism ensures that computational resources are preferentially allocated to nodes that have experienced severe semantic shifts due to drastic changes in cross-domain data distribution, achieving millisecond-level dynamic capture and adaptive correction of modal deviations. Through this intelligent, anti-blocking queue scheduling strategy, not only is the resource waste caused by redundant computation significantly reduced, but the final result is a highly effective and instructive label offset-guided scheduling result.

[0033] Please see Figure 5 The steps to obtain S4 are as follows: S401: Based on the processing order listed in the label offset-guided scheduling result, obtain the node connection record and structure rollback record generated during the path construction process, and perform state mapping on the two types of records according to the logical topology order. Convert the node connection event and the verification failure event into a continuous logical sequence, form a corresponding construction trajectory for each relational path, and obtain a path operation state sequence set. During the high-speed operation of the cross-domain multimodal resource relationship extraction engine, the underlying monitoring probes collect interaction trajectory data in real time and non-intrusively throughout the entire complex path construction lifecycle. A log parsing engine is used to deeply mine the memory stack, accurately extracting two core types of data generated during relationship extraction: first, node connection records, containing millisecond-level confirmation logs of successful logical connections between entity nodes and attributes; second, structural rollback records, containing reset and disconnection / reconnection logs triggered by context rule validation failures or semantic conflicts. Subsequently, graph database topology logic is introduced to perform strong state mapping based on timestamps and topology depth on these two types of discrete log records. Through a state machine model, the originally isolated node connection success events and constraint validation failure events are smoothly transformed into continuous logical time-series sequences representing the health of path growth. For each complex entity relationship dependency chain generated in the knowledge extraction graph, an immutable full-lifecycle construction trajectory archive is created, resulting in a well-structured and state-complete set of path operation state sequences.

[0034] S402: Based on the construction trajectory corresponding to each path identifier in the path operation state sequence set, perform logical judgment on the start and end nodes of adjacent connection segments, identify the connection relationship between semantic description and relation type in the sequence, and perform discrimination processing on connection consistency. Then, distinguish and encode the stable connection segments and unstable connection segments after discrimination to obtain the path coherence state discrimination table. The system analyzes the complete construction trajectory corresponding to each independent path identifier in the path operation state sequence set. A sliding window algorithm is used to perform fine-grained logical coherence analysis on the starting and ending nodes of adjacent relationship connection segments. It accurately identifies the tightness of contextual coherence between multimodal semantic description rules and entity relationship types throughout the sequence transformation process, and introduces a quantified mismatch rate index for coherence discrimination. The specific calculation logic is as follows: the total number of semantic verification failures triggered within adjacent segments is used as the numerator, and the total number of handshake attempts to connect nodes within that logical segment is used as the denominator to obtain the connection mismatch rate. A strict threshold scale is set for discretization and coding: if the mismatch rate is strictly less than the tolerance threshold of 0.1, it is safely coded as "S (Stable)"; if the mismatch rate fluctuates between 0.1 and 0.3, it is warned and coded as "M (Medium)"; if high-frequency triggering of backoff causes the mismatch rate to exceed 0.3, it is heavily marked and coded as "U (Unstable)". Through this rigorous quantitative calculation and layered defense, vulnerable nodes hidden in the relationship chain are effectively identified and highlighted, and a detailed path coherence status judgment table is obtained.

[0035] Table 3: Consistency Judgment Table for Segmented Connection of Relationship Path As shown in Table 3, by quantifying the mismatch rate, the abstract relational path is transformed into a monitorable coherent state code.

[0036] S403: Call the status coding results corresponding to each path identifier in the path continuity status discrimination table, perform integrity judgment on the overall path in the resource extraction process, classify the continuity segment codes and non-continuous segment codes, form the continuity judgment relationship of resource relationship in the conversion process, and generate path segment consistency constraint judgment results. The system uses the dynamically calculated state coding results for each path identifier in the path coherence status discrimination table to perform a final state assessment of the overall health and topological integrity of the global logical path of the entire multimodal resource extraction task. An aggregation statistical algorithm is then activated to accurately calculate the absolute percentage of highly stable "S"-level codes in the entire link coding sequence, which serves as the core quantitative index for measuring the path coherence. A strict quality interception gate is set: the algorithm engine will only allow passage when the overall coherence index of a complex graph path is consistently higher than the extremely high threshold of 0.9, determining that the resource extraction relationship path fully meets the strict cross-domain consistency constraints. During this process, coherent and incoherent segments are physically isolated and categorized for storage; incoherent segments are either reverted for reorganization or trigger a manual review alarm mechanism. This judgment model based on segment decay and overall percentage verification effectively intercepts semantic illusions and relationship breaks that may occur during model fine-tuning and data extraction, ultimately generating a high-confidence path segment consistency constraint judgment result.

[0037] Please see Figure 6 The steps to obtain S5 are as follows: S501: Based on the path identifier with a stable degree of construction in the path segmentation consistency constraint judgment result, obtain the corresponding structured transformation rules and resource extraction relationship data, and perform sequential verification of resource data frames according to the path index. Perform consistency judgment on the logical identifier of the data frame and the structured path sequence to obtain the structured transformation delivery sequence. Based on the consistency constraint judgment results of path segmentation, all stable path identifiers with extremely high construction level and full consistency verification are accurately screened. For these high-quality entity relationships, the scheduler automatically pulls structured mapping transformation rules that match its business scenario from the rule base as needed (such as complex JSON schemas or deeply nested XML transformation templates customized for specific businesses), and simultaneously retrieves source modality extraction data cached in memory. At this stage, the core processor initiates a strict data frame pipeline order verification mechanism. The algorithm extracts the globally unique logical ID encapsulated in the header of each resource data frame and performs high-frequency comparison with the topology sequence index presented by the stable path to ensure that the delivery order of each frame of data is completely consistent with the logical sequence of the business graph. Any data frames that are out of order, have mismatched IDs, or contain dirty data will be immediately intercepted at this stage and transferred to the dead-letter queue. Through this rigorous consistency defense, the thoroughly cleaned and sorted data frames are repackaged to obtain a structured transformation delivery sequence with extremely strong time order protection and format legality.

[0038] S502: Based on the resource data frames arranged in the structured conversion delivery sequence, obtain the current write status parameters of the storage module, match and judge the resource data frames with the storage parameters, send the data frames that can be converted to the storage array in sequence through the server bus, and at the same time, link the status of the conversion confirmation signal and the bus response signal to obtain the extraction conversion response trajectory. The system acquires a rigorously ordered structured conversion delivery sequence and wakes up the underlying hardware abstraction layer. Through the management interface, it captures real-time physical parameters of the target storage module's current write status, including the disk array's current available storage capacity, IOPS throughput limit, and cache hit rate. It dynamically matches the estimated volume of pending resource data frames in the queue with the available bandwidth parameters of the current storage components and performs overload assessment. After confirming that the storage load is at a healthy level, it initiates a DMA (Direct Memory Access) mechanism, sequentially sending data frames that can safely undergo format conversion to the underlying storage array via the server's high-speed backplane bus. Simultaneously, a high-precision asynchronous event listener is deployed at the bus control layer to perform bidirectional status monitoring of the disk write conversion acknowledgment signal (ACK) from the storage array and the bus's transmission handshake response signal. Upon detecting packet loss or timeout, a microsecond-level retransmission mechanism is immediately triggered. The entire process, from memory transmission and format parsing / reconstruction to successful physical sector disk write, is recorded, providing a rigorous extraction and conversion response trajectory.

[0039] S503: Call the resource identifier and conversion status information in the extraction and conversion response trajectory, confirm the status of the data frames that have been written and the conversion records that have generated responses, associate and collect the corresponding resource content with its structured conversion status, form the content range that has been extracted and structured converted, and generate cross-domain multimodal resource extraction and conversion results.

[0040] As the final closed loop of the entire cross-domain extraction framework, detailed extraction and transformation response logs are invoked to extract the unique resource identifier (UUID) and the final successful transformation status code fed back by the underlying storage. A high-concurrency asynchronous auditing process is initiated to perform strict one-to-one status confirmation and hash verification between the data frame queue that has been logically written in memory and the transformation disk-writing records that have actually generated physical IO responses in the storage array. After confirming that the multimodal raw data has been losslessly converted to the target structured format, the content blocks of the core resources are strongly correlated and aggregated with their structured index coordinates and dependency mapping graphs in the distributed files. At this point, the scope of content that has successfully completed cleaning, alignment, extraction, and format conversion in this batch processing task is completely locked, and the final structured path index data, resource relationship mapping information, and standardized conversion file set are packaged and generated. This marks the successful overcoming of the cross-modal semantic gap and the output of cross-domain multimodal resource extraction and transformation results with extremely high business direct usability.

[0041] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An open-source method for cross-domain multimodal resource extraction and transformation based on fine-tuning of pre-trained large models, characterized in that... Includes the following steps: S1: Extract text and image data stored in the data center server, identify semantic direction information and structural contour features in each modality of data, organize semantic change trends and structural distribution patterns according to data units, establish the combination order of semantic descriptions, and generate a multimodal semantic description set. S2: Extract text and image features from the multimodal semantic description set, perform semantic direction comparison and coordinate mapping, determine the directional consistency of the two types of modal data when expressing the same object, adjust the semantic reference direction to achieve modal alignment, and generate cross-modal semantic alignment results; S3: Retrieve the label correspondence in the cross-modal semantic alignment result, determine the degree of difference and adjustment direction of cross-domain labels in different processing modules, filter the processing modules that meet the priority update conditions, and construct the label offset guidance scheduling result; S4: Based on the processing sequence in the label offset-guided scheduling result, perform path segmentation construction on the structured resource relationship, count the semantic connection features between adjacent path segments, determine the degree of node consistency in the relationship chain during the construction process, form a correspondence between path segments and semantic stability, and generate path segmentation consistency constraint judgment results.

2. The open-source data cross-domain multimodal resource extraction and transformation method based on pre-trained large model fine-tuning as described in claim 1, characterized in that: The multimodal semantic description set includes semantic main axis direction, local structural trend, and semantic description order. The cross-modal semantic alignment result specifically includes direction vector deviation, distribution similarity, and semantic coordinate mapping matrix. The label offset guided scheduling result includes label difference mapping strength, processing module adjustment direction, and execution priority queue. The path segment consistency constraint judgment result specifically refers to node connection mismatch rate, path backtracking trigger frequency, and path construction coherence.

3. The method for open-source data cross-domain multimodal resource extraction and transformation based on pre-trained large model fine-tuning as described in claim 1, characterized in that, The steps for obtaining S1 are as follows: S101: Obtain text data files and image data files stored in the hard disk array, transmit them to the graphics processing unit through the network interface module, perform semantic analysis on the extracted sentence sequence from the text, compare the semantic vector obtained after analysis with the preset semantic axis feature library and semantic change trend library in turn, determine the semantic direction information of the text based on the comparison action, and associate it with the corresponding data identifier to obtain a semantic feature mapping table. S102: Based on the semantic direction information and data identifier in the semantic feature mapping table, extract the direction vector reflecting the structural contour and the range of change of local structure in each image region, and classify the frequency of occurrence of structural features in the overall image sampling points. Through this processing method, confirm the salience of structural features in spatial distribution and obtain the structural feature distribution parameter set. S103: Call the saliency distribution information in the structural feature distribution parameter group and associate it with the text semantic direction in the semantic feature mapping table. By the spatial mapping relationship and temporal position of semantic and structural features in the same resource identifier, determine the overlapping interval of semantic expression between modalities, and deduce the initial association path of multimodal data based on the interval information to obtain the multimodal semantic description set.

4. The method for open-source data cross-domain multimodal resource extraction and transformation based on pre-trained large model fine-tuning as described in claim 1, characterized in that, The steps for obtaining S2 are as follows: S201: Based on the arranged feature content in the multimodal semantic description set, extract the projection vectors of text and image in a unified semantic coordinate system, deconstruct the direction angle of the projection vectors, match the deflection component of the text vector with the normal direction of the image structure item by item, calculate the deviation of feature items with the same semantic dimension, and obtain the semantic direction offset set. S202: Call the deviation values ​​and modal mapping relationships in the semantic direction offset set, perform dynamic alignment operation on each set of feature sequences in the multimodal semantic description set, and use the cosine value of the vector angle, the Euclidean distance distribution degree and the cross-modal feature overlap as reference factors to identify alignment points in the semantic space that can achieve directional coordination, and obtain the semantic alignment state matrix. S203: Based on the semantic relationship between text and image in the semantic alignment state matrix, the semantic direction is gradually fine-tuned, and feature items whose semantic offset remains within a preset threshold during the alignment process are selected. The feature identifiers are then bound to the index sequence in the unified coordinate system to obtain the cross-modal semantic alignment result.

5. The open-source data cross-domain multimodal resource extraction and transformation method based on pre-trained large model fine-tuning as described in claim 1, characterized in that, The steps for obtaining S3 are as follows: S301: Based on the feature identifiers selected from the cross-modal semantic alignment results, obtain the status information of the processing modules configured in the computing unit, and extract the label difference parameters, module adjustment vectors and computing load parameters bound to the identifiers. Call the feature identifiers and parameter groups to perform binding operations to obtain the scheduling decision parameter set. S302: Based on the value status of the tag difference parameter, module adjustment vector and computing load parameter in the scheduling decision parameter set, perform multiple sets of logical judgment operations, perform consistency screening on the value status, determine whether the module adjustment direction is similar to the tag difference direction, whether the computing load is within the idle threshold, and whether the tag difference exceeds the preset mutation boundary, and obtain the processing module status tag group. S303: Call the module identifier and status information in the status tag group of the processing module, perform priority aggregation operation on the module identifiers with high directional convergence, low load status and high difference intensity, and combine the identifier with the corresponding processing order content to obtain the tag offset guidance scheduling result.

6. The open-source data cross-domain multimodal resource extraction and transformation method based on pre-trained large model fine-tuning as described in claim 1, characterized in that, The steps for obtaining S4 are as follows: S401: Based on the processing order listed in the label offset guidance scheduling result, obtain the node connection record and structure rollback record generated during the path construction process, and perform state mapping on the two types of records according to the logical topology order, convert the node connection event and the verification failure event into a continuous logical sequence, form a corresponding construction trajectory for each relational path, and obtain a path operation state sequence set. S402: Based on the construction trajectory corresponding to each path identifier in the path operation state sequence set, perform logical judgment on the start and end nodes of adjacent connection segments, identify the connection relationship between semantic description and relation type in the sequence, and perform discrimination processing on connection consistency. Then, distinguish and encode the stable connection segments and unstable connection segments after discrimination to obtain the path coherence state discrimination table. S403: Call the status coding results corresponding to each path identifier in the path continuity status discrimination table, perform integrity judgment on the overall path in the resource extraction process, classify the continuity segment codes and non-continuous segment codes accordingly, form the continuity judgment relationship of resource relationship in the conversion process, and generate path segment consistency constraint judgment results.

7. The open-source data cross-domain multimodal resource extraction and transformation method based on pre-trained large model fine-tuning as described in claim 1, characterized in that, The method further includes: S5: Combining the stable path, its structured type, and conversion rules in the path segmentation consistency constraint judgment result, the extraction result is stored, written, and formatted through the server bus, and the structured record generated in the multimodal resource processing system is output to generate cross-domain multimodal resource extraction and conversion results. The cross-domain multimodal resource extraction and transformation results include structural path index data, resource relationship mapping information, and a final set of transformed files.

8. The method for open-source data cross-domain multimodal resource extraction and transformation based on pre-trained large model fine-tuning as described in claim 7, characterized in that, The steps for obtaining S5 are as follows: S501: Based on the path identifier with a stable degree of construction in the path segmentation consistency constraint judgment result, obtain the corresponding structured conversion rules and resource extraction relationship data, and perform sequential verification of resource data frames according to the path index. Perform consistency judgment on the logical identifier of the data frame and the structured path sequence to obtain the structured conversion delivery sequence. S502: Based on the resource data frames arranged in the structured conversion delivery sequence, obtain the current write status parameters of the storage module, match and judge the resource data frames with the storage parameters, send the data frames that can be converted to the storage array in sequence through the server bus, and at the same time, link the status of the conversion confirmation signal and the bus response signal to obtain the extraction conversion response trajectory. S503: Call the resource identifier and conversion status information in the extraction and conversion response trajectory, confirm the status of the data frames that have been written and the conversion records that have generated responses, associate and collect the corresponding resource content with its structured conversion status, form the content range of the system that has been extracted and structured converted, and generate cross-domain multimodal resource extraction and conversion results.