An AI large model-based cross-system data collection method
By adopting a cross-system data acquisition method based on AI large models, multimodal behavioral feature vectors are generated and weights are dynamically adjusted. These vectors are then parsed and mapped to SQL query templates and RAG retrieval paths. This solves the problems of intent alignment deviation and temporal optimization in cross-system acquisition, and achieves efficient data acquisition and knowledge graph construction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIN RONG HUI XIN XI JI SHU YOU XIAN GONG SI
- Filing Date
- 2025-08-15
- Publication Date
- 2026-04-10
AI Technical Summary
In the dynamic fusion of multi-source heterogeneous data, existing technologies rely on static weight allocation to align behavioral trajectories with the intent of natural language commands, resulting in significant deviations in the parsing of cross-system acquisition targets. The priority scheduling of time-series data cleaning lacks quantification of the chaotic characteristics of operation intervals, leading to the loss of high-value information. Furthermore, the generation of knowledge graphs does not consider spatiotemporal dimensions, limiting the accuracy of cross-system event tracing.
By using a cross-system data collection method based on AI large models, multimodal behavioral feature vectors are generated. Combined with a dynamic intent analysis layer, entropy-constrained intent vectors with corrected weights are generated. The integrated target is parsed and mapped to SQL query templates and RAG enhanced retrieval paths. Data cleaning priorities are dynamically scheduled to generate time-series optimized datasets, and finally, a cross-system related knowledge graph is constructed.
It achieves precise alignment of cross-modal intents, reduces the error between SQL query templates and RAG retrieval paths, ensures the real-time performance and accuracy of cross-system data collection targets, and improves the real-time performance and accuracy of information parsing.
Smart Images

Figure CN121144333B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of cross-system integration, and particularly relates to a cross-system data collection method based on an AI large model. BACKGROUND
[0002] In recent years, multi-modal behavior analysis technology has realized quantitative characterization of user operation trajectories, and a dynamic intention analysis model combined with natural language processing (NLP) can generate an enhanced instruction set. At the same time, the development of a cross-system protocol converter supports the compilation of a structured query template into a native database statement, and the advancement of a RAG (retrieval augmented generation) framework improves the semantic retrieval efficiency of unstructured data.
[0003] However, the prior art has deficiencies in the dynamic fusion of multi-source heterogeneous data. The intention alignment of behavior trajectories and natural language instructions relies on static weight allocation, resulting in a large deviation in the analysis of cross-system collection targets. The time series data cleaning priority scheduling lacks quantification of the chaotic characteristics of operation intervals, causing the loss of high-value information. In the knowledge graph generation process, the fusion of structured statistical features and unstructured semantic fragments does not consider the spatiotemporal dimension association, and the cross-system event tracing accuracy has limitations. SUMMARY
[0004] In view of the above-mentioned existing problems, the present application is proposed.
[0005] Therefore, the present application provides a cross-system data collection method based on an AI large model to solve the problems of intention vector distortion caused by the time sequence misalignment of behavior trajectories and language instructions and the insufficient optimization of heterogeneous data time sequences.
[0006] To solve the above technical problems, the present application provides the following technical solutions:
[0007] In a first aspect, the present application provides a cross-system data collection method based on an AI large model, which includes collecting user behavior trajectories and generating multi-modal behavior feature vectors.
[0008] The multi-modal behavior feature vectors are input into a dynamic intention analysis layer, and natural language instructions received in real time are fused to generate an entropy-constrained intention vector after weight correction.
[0009] According to the entropy-constrained intention vector, the collected integrated target is analyzed, the structured data components are mapped to a SQL query template, the unstructured data components are mapped to a RAG augmented retrieval path, and a unified instruction set is output.
[0010] The unified instruction set is executed, the SQL query template is compiled into a native query statement through a protocol converter, and the RAG augmented retrieval path is converted into an unstructured data component scanning operation to obtain an initial data set.
[0011] Utilize the operation time interval sequence in the user behavior trajectory, dynamically schedule the cleaning priority of the initial data set, and generate a timing optimization data set;
[0012] Input the timing optimization data set into an enhanced retrieval engine, fuse the statistical features extracted by the AI large model from the structured data component and the semantic fragments generated by RAG from the unstructured data component, and generate a cross-system associated knowledge graph.
[0013] As a preferred scheme of the cross-system data acquisition method based on the AI large model, wherein: the user behavior trajectory includes a user pen pressure intensity sequence, a screen touch area coordinate sequence and an operation time interval sequence.
[0014] As a preferred scheme of the cross-system data acquisition method based on the AI large model, wherein: the user behavior trajectory is collected, and a multi-modal behavior feature vector is generated, and the specific steps are as follows,
[0015] Perform Daubechies wavelet decomposition on the pen pressure intensity sequence, extract the detail coefficient and output the pressure energy entropy value;
[0016] According to the touch area coordinate sequence, a moving speed vector is obtained, and a pulse code sequence mean value is generated through an emulated retinal pulse firing rate engine;
[0017] The operation time interval sequence is reconstructed in phase space to generate an embedding vector sequence, and an operation rhythm complexity feature value is generated through a permutation entropy algorithm;
[0018] The pressure energy entropy value, the pulse code sequence mean value and the operation rhythm complexity feature value are input into a cortical column competitive encoder for normalization, and a multi-modal behavior feature vector is output.
[0019] As a preferred scheme of the cross-system data acquisition method based on the AI large model, wherein: the entropy-constrained intention vector after the correction weight is generated, and the specific steps are as follows,
[0020] The multi-modal behavior feature vector is modulated into a multi-frequency sinusoidal pulse sequence, and the real-time received natural language instruction is encoded into a Gaussian distribution pulse sequence;
[0021] The phase difference of the multi-frequency sinusoidal pulse sequence and the Gaussian distribution pulse sequence is synchronized through a pre-trained oscillator coupling network, the phase difference change rate is calculated, and a gating signal is generated by performing a gating operation;
[0022] Integrate the gating signal and the phase difference change rate to generate an entropy-constrained decay factor;
[0023] The weight distribution of the multi-modal behavior feature vector is modulated using an entropy constraint attenuation factor to generate an entropy constraint intention vector after weight correction.
[0024] As a preferred scheme of the cross-system data acquisition method based on the AI large model, the cross-system acquisition target includes a structured data component and an unstructured data component generated by entropy constraint intention vector decomposition.
[0025] The structured data component includes user identification weight, time sensitivity weight, and data type preference weight.
[0026] The unstructured data component includes semantic sensitivity weight, urgency weight, and correlation strength weight.
[0027] As a preferred scheme of the cross-system data acquisition method based on the AI large model, the output unified instruction set has the following specific steps,
[0028] The structured data component vector is dynamically activated by a differentiable symbol rule generator to generate a SQL query template.
[0029] The neural concept space projection is performed on the unstructured data component to generate a semantic potential field gradient path by Monte Carlo integration, and an RAG enhanced retrieval path is output.
[0030] The SQL query template and the RAG enhanced retrieval path are fused by a cross-modal semantic encoder to output a low-dimensional instruction vector as a unified instruction set.
[0031] As a preferred scheme of the cross-system data acquisition method based on the AI large model, the SQL query template is a dynamically condition-activated relational database query instruction carrier.
[0032] The RAG enhanced retrieval path is a semantic focus retrieval path generated by neural concept space projection.
[0033] As a preferred scheme of the cross-system data acquisition method based on the AI large model, the initial data set is obtained by the following specific steps,
[0034] The unified instruction set is decoded into a structured query micro-instruction vector and an unstructured scan micro-instruction vector.
[0035] Based on the semantic sensitivity weight, urgency weight, and correlation strength weight in the unstructured data component and the unstructured scan micro-instruction vector, a precision semantic scanning stream is generated by a dynamic precision control mechanism.
[0036] The SQL query template is parameterized and compiled with the user identification weight, the time-sensitive weight and the data type preference weight in the structured data component to generate a native query statement;
[0037] The heterogeneous execution flow is dynamically arranged through a gradient-constrained entropy minimization strategy, the native query statement and the semantic scanning flow are executed, and an initial data set is acquired.
[0038] As a preferred scheme of the AI large model-based cross-system data acquisition method, the time sequence optimization data set is generated, and the specific steps are as follows,
[0039] The operation time interval sequence in the user behavior trajectory is subjected to phase space reconstruction, the maximum Lyapunov index and the statistical distribution feature are calculated, and a chaotic feature vector is generated.
[0040] The dynamic scheduling weight is calculated through a chaotic entropy weight generation mechanism based on the chaotic feature vector;
[0041] The dynamic scheduling weight is combined with the size of each data subset in the initial data set, and is mapped to the cleaning priority of the initial data set to generate a dynamic cleaning queue.
[0042] The dynamic cleaning queue is subjected to chaotic filtering and timestamp calibration, and a time sequence optimization data set is output.
[0043] As a preferred scheme of the AI large model-based cross-system data acquisition method, the cross-system association knowledge graph is generated, and the specific steps are as follows,
[0044] The structured data entity is extracted from the time sequence optimization data set as a node, the unstructured semantic fragment is extracted as a hyperedge, a space-time hypergraph is constructed, and a time window label is bound to the hyperedge;
[0045] According to the bound time window label, the time dimension feature is extracted, the node feature of the space-time hypergraph is extracted through a pre-trained space-time feature fusion network, the time dimension feature is dynamically associated through fusion, and a node feature vector is output.
[0046] The statistical distribution feature is injected into the node feature vector through a cross-modal attention mechanism to generate a fusion feature matrix.
[0047] Based on the fusion feature matrix, the relationship strength between entities is acquired, the entity relationship triplets meeting a preset strength threshold are extracted, and a time window label is bound to generate a cross-system association knowledge graph.
[0048] The application has the beneficial effects that: the phase of the natural language instruction is synchronized with the behavior characteristics of the neural pulse coupling oscillator, cross-modal intention precise alignment is realized, phase difference fluctuation is controlled, the entropy constraint decay factor generated based on the phase difference change rate dynamically suppresses the conflict feature weight, the error of the SQL query template and the RAG retrieval path is reduced, and the real-time performance and accuracy of cross-system target analysis are ensured. BRIEF DESCRIPTION OF DRAWINGS
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0050] Fig. 1 Flowchart of the cross-system data acquisition method based on the AI large model.
[0051] Fig. 2 Flowchart for generating multi-modal behavior feature vectors.
[0052] Fig. 3 Flowchart for generating entropy-constrained intention vectors after weight correction.
[0053] Fig. 4 Flowchart for generating a cross-system associated knowledge graph. DETAILED DESCRIPTION
[0054] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification.
[0055] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application, therefore the present application is not limited to the specific embodiments disclosed below.
[0056] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an independent or alternative embodiment that excludes other embodiments.
[0057] REFERENCE Figs. 1-4 For one embodiment of the present application, the embodiment provides a cross-system data acquisition method based on an AI large model, comprising the following steps:
[0058] S1, collect user behavior trajectory, generate multi-modal behavior feature vector.
[0059] S1.1: the user behavior trajectory includes a user pen pressure intensity sequence, a screen touch area coordinate sequence, and an operation time interval sequence.
[0060] It should be noted that the user behavior trajectory is synchronized with the user operation through the NLU enhanced cross-system focused crawler, and the natural language understanding capability maps the physical operation to the semanticized crawling instruction of the browser interface element. For example, the screen touch area coordinate sequence directly drives the crawler focus to move in the web page, simulating the human visual gaze trajectory.
[0061] The user pen pressure intensity sequence refers to the force value sequence applied to the pen tip recorded by the pressure sensor in time sequence when the user operates the writing tool;
[0062] The screen touch area coordinate sequence refers to the position sequence of the contact point in the screen coordinate system recorded in time sequence when the touch screen is touched;
[0063] The operation time interval sequence refers to the action interval duration sequence recorded in time sequence between the user's continuous touch operations.
[0064] S1.2: Daubechies wavelet decomposition is performed on the pen pressure intensity sequence, and the detail coefficient is extracted to output the pressure energy entropy value;
[0065] Specifically, Daubechies wavelet is used to perform decomposition on the pen pressure intensity sequence to generate a wavelet coefficient set containing approximation coefficients and detail coefficients, the detail coefficients are extracted from the wavelet coefficient set, the energy value of the detail coefficients is obtained, the energy distribution is output through the square and operation, the energy distribution is normalized to form a probability distribution, and the entropy value of the probability distribution is obtained from the probability distribution to generate the pressure energy entropy value.
[0066] S1.3: obtain a moving speed vector according to the touch area coordinate sequence, and generate a pulse code sequence mean value through an emulated retinal pulse firing rate engine;
[0067] Specifically, the position difference between adjacent coordinate points and the time interval are used to derive the instantaneous speed component, which is used as the moving speed vector. The moving speed vector is input into the emulated retinal pulse firing rate engine. The emulated retinal pulse firing rate engine simulates the firing characteristics of retinal neurons based on the moving speed vector. The emulated retinal pulse firing rate engine maps the magnitude of the moving speed vector to the pulse firing frequency. The greater the magnitude of the moving speed vector, the higher the pulse firing frequency. A binary pulse code sequence is output according to the pulse firing frequency within a fixed time window. According to the proportion of high level in the binary pulse code sequence, the pulse code sequence mean value is output.
[0068] S1.4: phase space reconstruction is performed on the operation time interval sequence to generate an embedding vector sequence, and an operation rhythm complexity feature value is generated by an algorithm of permutation entropy;
[0069] Specifically, the delay time is determined according to the autocorrelation of the operation time interval sequence, the embedding dimension is determined according to the pseudo-neighbor characteristic of the operation time interval sequence, and the subsequence with the embedding dimension is cut from the operation time interval sequence in time sequence according to the delay time and the embedding dimension, each subsequence is arranged in time sequence as an embedding vector; all embedding vectors are arranged in time sequence to form an embedding vector sequence;
[0070] The embedding vector sequence is input into the permutation entropy algorithm: the values in each embedding vector are arranged in ascending order to obtain a permutation pattern, the number of occurrences of each type of permutation pattern in all embedding vectors is counted, the occurrence probability of each type of permutation pattern is obtained, a permutation pattern probability distribution is formed, and the entropy value of the permutation pattern probability distribution is obtained as the operation rhythm complexity feature value.
[0071] It should be noted that the autocorrelation of the operation time interval sequence refers to a quantitative index of the degree of linear dependence between the operation time interval value at any time and the operation time interval value after a delay of several time units in the operation time interval sequence;
[0072] The pseudo-neighbor characteristic of the operation time interval sequence refers to the geometric relationship feature that the operation time interval sequence sample points appear as adjacent points in a low-dimensional space, but are actually not adjacent in a higher-dimensional space.
[0073] S1.5: input the pressure energy entropy value, the pulse code sequence mean value and the operation rhythm complexity feature value into the cortical column competitive encoder, normalize, and output a multi-modal behavior feature vector.
[0074] Specifically, the pressure energy entropy value, the pulse code sequence mean value and the operation rhythm complexity feature value are input into the cortical column competitive encoder, and are spliced into an original feature vector in the order of the pressure energy entropy value as the first dimension, the pulse code sequence mean value as the second dimension and the operation rhythm complexity feature value as the third dimension. The cortical column competitive encoder allocates dynamic weights by comparing the relative dimension value size relationship of the pressure energy entropy value, the pulse code sequence mean value and the operation rhythm complexity feature value, the larger the dimension value, the higher the weight allocated, and performs linear normalization on the weighted original feature vector to output a normalized three-dimensional vector as a multi-modal behavior feature vector.
[0075] S2, input the multi-modal behavior feature vector into the dynamic intention analysis layer, fuse the real-time received natural language instructions, and generate an entropy-constrained intention vector after correction weight.
[0076] S2.1: modulate the multi-modal behavior feature vector into a multi-frequency sinusoidal pulse sequence, and encode the real-time received natural language instruction into a Gaussian distribution pulse sequence;
[0077] Specifically, each dimension of the multi-modal behavior feature vector is extracted, and an independent sinusoidal frequency is assigned to each dimension. The larger the dimension index value, the higher the frequency assigned. The dimension is converted into the amplitude value of the corresponding sinusoidal wave, and a pulse sequence triggered by the peak position of the multi-frequency sinusoidal wave on the time axis is generated, forming a multi-frequency sinusoidal pulse sequence.
[0078] Meanwhile, the natural language instruction is segmented to generate a word sequence. A word embedding algorithm is performed on the word sequence, and each word is a real number vector. The mean vector of all real number vectors is calculated, and the components of the mean vector are used as Gaussian distribution parameters to generate a random pulse time sequence conforming to Gaussian distribution, forming a Gaussian distribution pulse sequence.
[0079] S2.2: synchronize the phase difference between the multi-frequency sinusoidal pulse sequence and the Gaussian distribution pulse sequence through the pre-trained oscillator coupling network, calculate the phase difference rate of change, and generate a gating signal by performing a gating operation;
[0080] It should be noted that the pre-training process of the oscillator coupling network is as follows: collect multi-frequency sinusoidal pulse sequence samples and Gaussian distribution pulse sequence samples as input training data, calculate the output phase difference in the forward propagation under the initial random coupling parameter state of the oscillator coupling network, measure the deviation value of the output phase difference from the theoretical ideal phase difference, calculate the gradient direction of the coupling parameter by the back propagation algorithm, update the coupling parameter value by using the gradient descent strategy, and repeat the above process until the average absolute error of the output phase difference and the ideal phase difference is less than the preset phase difference threshold. The preset phase difference threshold is set based on the stability requirement of the gating signal and the real-time constraint, and the example value is 0.05. The coupling parameter is solidified, and the oscillator coupling network is obtained.
[0081] Specifically, the dynamic changing gating reference quantity is automatically generated according to the inherent vibration characteristics of the oscillator coupling network, and the phase difference rate of change is compared with the gating reference quantity in size in real time. When the phase difference rate of change is greater than the gating reference quantity, a positive polarity switching state is triggered, and when the phase difference rate of change is less than the gating reference quantity, a negative polarity switching state is triggered. The positive and negative polarity switching states at continuous time points are directly connected into a discrete pulse sequence, which is the gating signal.
[0082] The phase difference between the multi-frequency sinusoidal pulse sequence and the Gaussian distribution pulse sequence is synchronized through the oscillator coupling network, and the phase difference rate of change is calculated, which is expressed as:
[0083] ;
[0084] In the formula, Indicates the rate of change of phase difference. Indicates the current moment. Indicates instantaneous phase difference, This represents the fundamental frequency of a Gaussian distributed pulse sequence. This represents the fundamental frequency of a multi-frequency sinusoidal pulse sequence. This indicates the oscillator synchronization adjustment parameters. Represents the sine function. This represents the scaling factor for the weights of multimodal behavioral features. This represents summing the total dimensions of the multimodal behavior feature vectors. This represents the dimension index of the multimodal behavior feature vector. The first feature vector representing multimodal behavior Dimension value, Represents the Dirac operator, Indicates the first The moment when the dimensional pulse occurs.
[0085] It should be noted that, The dimension is rad / s. and The dimension is rad / s. Dimensional and dimensionless The combined output has the dimension of rad / s. The units are rad and Dimensions are The product is and weighted dimensionless Ultimately, all quantities are standardized to rad / s.
[0086] The multimodal behavioral feature weight scaling factor is derived from the dynamic range calibration value of the behavioral feature vector output by the cortical column competitive encoder, with an example value of 0.05.
[0087] S2.3: Specifically, the gating signal is integrated with the rate of change of phase difference to generate an entropy-constrained attenuation factor, the expression of which is:
[0088] ;
[0089] In the formula, This represents the entropy constraint decay factor. This represents the natural exponential function. Indicates the attenuation intensity coefficient. Indicates from arrive Perform integration. Indicates a single instruction processing cycle. Indicates the current time The gating signal, denotes a time decay factor, denotes a time differential unit.
[0090] It should be noted that, is dimensionless, is dimensionless, is dimensionless, is dimensionless, , is dimensionless, , is dimensionless, and the final output is dimensionless, and the matching dimensionless , the dimension is kept uniform;
[0091] The decay intensity coefficient is dynamically calibrated based on the chaos degree of the behavior-language intention alignment, and an example value is 0.7 in a high chaos scene and 0.3 in a low chaos scene. The time decay factor is derived from the window width calibration parameter of the operation rhythm speed, and an example value is 0.1 in a regular operation and 0.5 in an emergency operation.
[0092] S2.4: Modulate the weight distribution of the multi-modal behavior feature vector using the entropy constraint decay factor to generate an entropy constraint intention vector after weight correction.
[0093] Specifically, the original weight of each dimension of the multi-modal behavior feature vector is extracted, the entropy constraint decay factor is combined with the original weight of each dimension as a global scaling coefficient, the scaled dimension weight is generated, and the weight sum normalization is performed on the scaled dimension weight to generate the normalized scaling weight.
[0094] Each dimension of the multi-modal behavior feature vector is combined with the corresponding normalized scaling weight to obtain a weighted dimension, and the weighted dimension is reorganized in the original dimension order of the multi-modal behavior feature vector to output the reconstructed feature vector as the entropy constraint intention vector after weight correction.
[0095] Preferably, the dynamic intention analysis layer synchronizes the phase difference between the multi-modal behavior feature vector and the natural language instruction through the oscillator coupling network to generate an entropy constraint decay factor to dynamically modulate the feature weight. Compared with the conventional static weight distribution model, the feature conflict problem caused by the time sequence misalignment between the behavior trajectory and the language instruction is solved. The static weight distribution model ignores the real-time phase relationship between the operation rhythm and the language instruction, resulting in a high deviation rate of the structured query template and the unstructured retrieval path. Through the entropy constraint decay factor, the conflict feature dimension is inhibited in real time, the joint error rate of the SQL query template and the RAG retrieval path is reduced, and the real-time performance and accuracy of the cross-system acquisition target analysis are improved.
[0096] S3. The integrated target is parsed according to the entropy-constrained intention vector, the structured data component is mapped to a SQL query template, and the unstructured data component is mapped to a RAG enhanced search path, and a unified instruction set is output.
[0097] S3.1: The cross-system acquisition target includes a structured data component and an unstructured data component generated by entropy-constrained intention vector decomposition;
[0098] It should be noted that the structured data component includes a user identification weight, a time-sensitive weight, and a data type preference weight;
[0099] The user identification weight refers to a quantitative indicator in the entropy-constrained intention vector that represents the strength of the uniqueness of the user's identity, which is used for identity filtering priority determination in structured queries;
[0100] The time-sensitive weight refers to a quantitative indicator in the entropy-constrained intention vector that represents the strength of the data time validity, which is used for time window dynamic scaling control in structured queries;
[0101] The data type preference weight refers to a quantitative indicator in the entropy-constrained intention vector that represents the user's demand preference for a specific data structure type (such as numerical type / text type), which is used for field selection optimization in structured queries.
[0102] It should be noted that the unstructured data component includes a semantic sensitivity weight, an urgency weight, and an association strength weight.
[0103] The semantic sensitivity weight refers to a quantitative indicator in the entropy-constrained intention vector that represents the strength of the semantic association demand, which is used for semantic matching depth control in unstructured data retrieval;
[0104] The urgency weight refers to a quantitative indicator in the entropy-constrained intention vector that represents the real-time requirement of unstructured data acquisition, which is used for non-structured data scanning task scheduling priority allocation;
[0105] The association strength weight refers to a quantitative indicator in the entropy-constrained intention vector that represents the strength of the cross-data source association demand, which is used for range expansion control in unstructured data retrieval.
[0106] S3.2: The structured data component vector is dynamically activated by a differentiable symbolic rule generator to generate a SQL query template;
[0107] Specifically, the user identification weight in the structured data component vector is compared with a preset user identity threshold, and the user identification filtering condition is activated when the user identification weight is greater than the preset user identity threshold; the time sensitivity weight in the structured data component vector is compared with a preset time sensitivity threshold, and the time window limiting condition is activated when the time sensitivity weight is greater than the preset time sensitivity threshold; the data type preference weight in the structured data component vector is compared with a preset type preference threshold, and the specific data type field selection condition is activated when the data type preference weight is greater than the preset type preference threshold.
[0108] The activated user identification filtering condition is mapped to a WHERE clause user ID placeholder of SQL; the activated time window limiting condition is mapped to a BETWEEN clause time parameter placeholder of SQL; and the activated specific data type field selection condition is mapped to a SELECT clause field name enumeration placeholder of SQL;
[0109] The generated WHERE clause user ID placeholder, BETWEEN clause time parameter placeholder and SELECT clause field name enumeration placeholder are combined according to a standard SQL syntax structure to output an SQL query template.
[0110] It should be noted that the preset user identity threshold is set based on the minimum confidence requirement of the user identity verification scene, and an example value is 0.65; the preset time sensitivity threshold is set based on the upper limit of the chaos feature tolerance of the operation interval sequence, and an example value is 0.55; and the preset type preference threshold is set based on the median of the data type distribution entropy, and an example value is 0.48.
[0111] S3.3: Perform neural concept space projection on the unstructured data component, generate a semantic potential field gradient path through Monte Carlo integration, and output an RAG enhanced retrieval path;
[0112] Specifically, the semantic sensitivity weight, urgency weight and correlation strength weight in the unstructured data component are input into a neural concept space projection engine, and the neural concept space projection engine maps each weight component to a coordinate vector in the neural concept space to form an initial projection point;
[0113] A spherical semantic potential field is generated with the initial projection point as the center and the semantic sensitivity weight value as the radius, the urgency weight value scales the intensity of the semantic potential field, and the correlation strength weight value expands the effective action range of the semantic potential field;
[0114] A plurality of candidate retrieval paths are randomly generated within the effective action range of the semantic potential field, the average semantic potential field intensity of each candidate retrieval path is counted, and the candidate retrieval path with the maximum average semantic potential field intensity is selected as the reference path;
[0115] From the starting point of the reference path, iteratively extend the path node in the direction of the fastest growth of the semantic potential field intensity, connect all path nodes to form a semantic potential field gradient path, and directly output the semantic potential field gradient path as the RAG enhanced retrieval path.
[0116] It should be noted that the RAG enhanced retrieval path enhances the natural language understanding ability of the NLU enhanced cross-system focused crawler, converts the semantic potential field gradient into a dynamic crawling route in the browser interface, controls the semantic focus depth of the crawler on text / image elements (such as parsing deep nested DOM structures when the weight is high), expands the cross-element association range of the crawler (such as crawling the associated fields of adjacent tables), optimizes the rendering waiting strategy (such as bypassing dynamic loading to directly extract static snapshots when the urgency is high).
[0117] S3.4: Fuse the SQL query template and the RAG enhanced retrieval path through the cross-modal semantic encoder, and output a low-dimensional instruction vector as a unified instruction set.
[0118] Specifically, the structured encoding branch of the SQL query template is input into the cross-modal semantic encoder, the syntax structure features of the SQL query template are parsed, and the syntax tree vector is mapped; the unstructured encoding branch of the RAG enhanced retrieval path is input into the cross-modal semantic encoder, the node topology features of the RAG enhanced retrieval path are parsed, and the path node vector is mapped.
[0119] The syntax tree vector and the path node vector are input into the bidirectional attention fusion layer of the cross-modal semantic encoder, the syntax tree vector is taken as the query vector, and the path node vector is taken as the key-value pair vector to obtain the attention weight matrix from structured to unstructured; the path node vector is taken as the query vector, and the syntax tree vector is taken as the key-value pair vector to obtain the attention weight matrix from unstructured to structured; and the two attention weight matrices are combined element by element to generate a fusion attention graph.
[0120] The fusion attention graph is input into the dimension reduction compression layer of the cross-modal semantic encoder, the dimension reduction compression layer extracts the first three principal component directions through principal component analysis, projects the fusion attention graph to the principal component directions, and generates a three-dimensional low-dimensional instruction vector, which is directly output as a unified instruction set.
[0121] It should be noted that the SQL query template is a dynamic condition activated relational database query instruction carrier; and the RAG enhanced retrieval path is a semantic focus retrieval path generated by neural concept space projection.
[0122] More preferably, the structured data component weight generation SQL query template is dynamically activated by the differentiable symbol rule generator, the unstructured data component is converted into RAG by combining the neural concept space projection to enhance the retrieval path, and then compressed into a low-dimensional instruction vector by the cross-modal semantic encoder. Compared with the fixed template mapping method, the analysis deviation is significantly caused by the solidified weight distribution. Through the dynamic component decomposition of the entropy constraint intention vector, the accurate identity filtering is activated according to the user identification weight, and the retrieval range is dynamically adjusted according to the semantic sensitivity weight, which significantly improves the analysis accuracy of the cross-system collected target.
[0123] S4, execute a unified instruction set, compile the SQL query template into a native query statement through a protocol converter, and convert the RAG enhanced retrieval path into an unstructured data component scanning operation to obtain an initial data set.
[0124] S4.1: decode the unified instruction set into a structured query micro-instruction vector and an unstructured scanning micro-instruction vector;
[0125] Specifically, the three-dimensional low-dimensional instruction vector of the unified instruction set is input into an instruction decoder. The instruction decoder splices the first dimension coordinate value and the second dimension coordinate value into a two-dimensional vector, and separately extracts the third dimension coordinate value into a one-dimensional vector;
[0126] The two-dimensional vector is directly output as a structured query micro-instruction vector, wherein the first dimension coordinate value is mapped as the user identification parameter intensity in the SQL template, and the second dimension coordinate value is mapped as the time window scaling factor;
[0127] The one-dimensional vector is directly output as an unstructured scanning micro-instruction vector.
[0128] S4.2: based on the semantic sensitivity weight, the urgency weight and the association strength weight in the unstructured scanning micro-instruction vector and the unstructured data component, generate a precision semantic scanning stream through a dynamic precision control mechanism;
[0129] Specifically, the numerical value of the unstructured scanning micro-instruction vector maps a basic precision level, and the greater the unstructured scanning micro-instruction vector, the higher the basic precision level;
[0130] The semantic sensitivity weight scales the basic precision level. When the semantic sensitivity weight is greater than the preset semantic sensitivity threshold, the basic precision level is improved. When the semantic sensitivity weight is less than the preset semantic sensitivity threshold, the basic precision level is reduced, and a semantic enhanced precision level is generated;
[0131] The association strength weight and the semantic enhanced precision level are combined to determine the semantic association span of the unstructured data scanning. The greater the numerical value of the association strength weight, the wider the semantic association span;
[0132] The emergency weight value is mapped to a scan task scheduling priority, and when the emergency weight is greater than a preset emergency threshold, the highest scheduling priority is allocated, and when the emergency weight is less than the preset emergency threshold, the lowest scheduling priority is allocated;
[0133] The scanning accuracy defined by the integrated semantic enhanced precision level, the scanning range defined by the semantic correlation span, and the execution order defined by the scheduling priority are integrated to generate a precision semantic scan stream.
[0134] It should be noted that the execution of the precision semantic scan stream relies on the natural language understanding (NLU) capability of the NLU enhanced cross-system focused crawler, and the semantic level analysis of visible interface elements: based on the semantic sensitivity weight in the unstructured scan micro-instruction vector, the target element in the browser or application interface is located through the pre-trained visual-language joint model, and the human visual cognitive path is simulated;
[0135] The spatial constraints in the natural language instruction are analyzed in real time to generate a focus area mask, and the scanning engine is controlled to only capture visible data components within the mask to achieve accurate extraction of "naked-eye recognizable data";
[0136] For dynamically rendered content (such as Captcha or pop-up windows), noise is filtered through an attention mechanism to ensure that the collection result is consistent with the semantics of the user instruction.
[0137] It should be noted that the preset semantic sensitivity threshold is set based on the semantic matching recall rate requirement of unstructured data retrieval, and an example value is 0.6; the preset emergency threshold is set based on the upper limit of the cross-system operation response delay, and an example value is 0.7.
[0138] S4.3: Parameterize and compile the SQL query template with the user identification weight, time sensitivity weight, and data type preference weight in the structured data component to generate a native query statement;
[0139] It should be noted that the preset user identity threshold is set based on the minimum confidence requirement of the user identity verification scenario, and an example value is 0.65.
[0140] Specifically, the user ID placeholder in the SQL query template is extracted, the user identification weight is combined with the user identity reference value to generate a user ID specific value; if the user identification weight is greater than the preset user identity threshold, the user ID specific value is used to replace the user ID placeholder; if the user identification weight is less than the preset user identity threshold, the user ID placeholder and the associated operator are deleted;
[0141] The time parameter placeholder in the SQL query template is extracted, and based on the current timestamp, the time sensitivity weight is combined with the time window reference span to generate a dynamic time window range, and the start timestamp and end timestamp of the dynamic time window range are bound to the time parameter placeholder;
[0142] extracting the field name enumeration placeholder in the SQL query template, comparing the data type preference weight with the predefined data type priority list, screening the top three data types in the priority list closest to the data type preference weight, and replacing the field name enumeration placeholder with the screened field name;
[0143] After completing all placeholder replacement, combining clauses into components according to standard SQL syntax structure, and outputting the native query statement.
[0144] It should be noted that the predefined data type priority list refers to a field type sequence list arranged in descending order of query frequency in the structured data dictionary, which is used to map the data type preference weight to the high-frequency field subset.
[0145] S4.4: Dynamically scheduling a heterogeneous execution flow by a gradient-constrained entropy minimization strategy, executing the native query statement and the semantic scan flow, and obtaining the initial data set.
[0146] Specifically, the central processing unit load rate, memory occupancy rate and input / output throughput of the protocol converter execution environment are collected in real time to form an execution state vector, and the probability distribution entropy value of the execution state vector is obtained;
[0147] The complexity value of the native query statement is extracted as the gradient constraint direction, and the urgency weight of the precision semantic scan flow is extracted as the gradient constraint strength. When the probability distribution entropy value of the execution state vector is greater than the preset entropy tolerance threshold, the task scheduling order is adjusted in the direction of decreasing complexity value;
[0148] The native query statement compilation task and the precision semantic scan flow execution task are split into microtasks, a microtask priority sequence is generated according to the gradient constraint direction and the gradient constraint strength, and the microtasks are submitted to the protocol converter execution engine in the order of the priority sequence;
[0149] After the protocol converter execution engine completes all microtasks, the structured query result set and the unstructured semantic fragment set are merged to output the initial data set.
[0150] It should be noted that the preset entropy tolerance threshold is set based on the resource contention tolerance upper limit of the protocol converter, and the example value is 0.8.
[0151] S5, using the operation time interval sequence in the user behavior trajectory to dynamically schedule the cleaning priority of the initial data set, and generating a time-optimized data set.
[0152] S5.1: Specifically, the operation time interval sequence in the user behavior trajectory is reconstructed in phase space, the maximum Lyapunov index and statistical distribution characteristics are calculated, and a chaotic feature vector is generated, the expression is:
[0153] ;
[0154] ;
[0155] wherein, represents a chaotic feature vector, represents the mean of the rate of change of adjacent operation time intervals, represents the ratio of the standard deviation of operation time intervals to the arithmetic mean of the sequence of operation time intervals, represents the standard deviation of operation time intervals, represents the normalized permutation entropy, represents the total length of the sequence of operation time intervals, represents the reciprocal of the total length of the sequence of operation time intervals, represents the traversal summation from to , represents the index of operation time intervals, represents the number of adjacent interval pairs, represents the natural logarithm function, represents the th operation time interval value, represents the th operation time interval value, represents the absolute value of the ratio of adjacent operation time intervals, represents the traversal summation from to , represents the arithmetic mean of the sequence of operation time intervals, represents the square of the deviation of a single interval from the mean, represents the entropy value normalization coefficient, represents the traversal summation from to , represents all permutation pattern species, represents the permutation pattern index, represents the occurrence probability of the th permutation pattern, represents the logarithm function with base 2 acting on .
[0156] It should be noted that is dimensionless, D is dimensionless after the unit of time is eliminated, is dimensionless, and finally is dimensionless.
[0157] S5.2: Specifically, the dynamic scheduling weight is calculated based on the chaotic feature vector through a chaotic entropy weight generation mechanism, and the expression is:
[0158] ;
[0159] wherein, denotes the dynamic scheduling weight, denotes the entropy decay gain coefficient, denotes the maximum value, denotes the mean value of the adjacent operation time interval change rate, denotes the change rate activation threshold, denotes the maximum value of the mean value of the adjacent operation time interval change rate and the change rate activation threshold, denotes the normalized permutation entropy, denotes the ratio of the standard deviation of the operation time interval to the arithmetic mean of the operation time interval sequence, denotes the minimum positive value. It should be noted that all input quantities (,
[0160] , , , and ) in the above formula are dimensionless, the final output is dimensionless, and the dimension is kept uniform;
[0161] The change rate activation threshold is set based on physiological microtremor, and an example value is 0.1 for healthy people, 0.3 for Parkinson's patients, and 0.05 for athletes.
[0162] S5.3: Combine the dynamic scheduling weight with the size of each data subset in the initial data set to map the cleaning priority of the initial data set, and generate a dynamic cleaning queue;
[0163] Specifically, the size value of each data subset in the initial data set is extracted, and the size value of the data subset is combined with the corresponding dynamic scheduling weight value to obtain a priority coefficient. The larger the priority coefficient value is, the higher the cleaning priority is;
[0164] All data subsets are arranged in descending order of priority coefficient values, and data subsets with the same priority coefficient value are arranged in ascending order of data acquisition time stamp;
[0165] The sorted data subset sequence is output as a dynamic cleaning queue, and the first sequence is the highest priority cleaning task.
[0166] S5.4: Perform chaotic filtering and time stamp calibration on the dynamic cleaning queue to output a time series optimized data set.
[0167] Specifically, the maximum Lyapunov index associated with each data subset in the dynamic cleaning queue is extracted, and when the maximum Lyapunov index is greater than zero, it is determined that the data subset is a chaotic feature significant data subset. A low-pass filtering algorithm is applied to the chaotic feature significant data subset to suppress high-frequency chaotic noise components, and the original state of the data subset with a maximum Lyapunov index less than or equal to zero is retained.
[0168] According to the statistical distribution characteristics of the operation time interval sequence, the average interval value of the operation time interval sequence is obtained, and the chaotic filtered data subset timestamp is aligned according to the time window width boundary with the average interval value as the time window width;
[0169] After the timestamp alignment, all data subsets are rearranged in time sequence, and all data subsets are merged to form a time sequence optimized data set.
[0170] S6, input the time sequence optimized data set into the enhanced retrieval engine, fuse the statistical features extracted by the AI large model of the structured data component and the semantic fragments generated by the RAG of the unstructured data component, and generate a cross-system associated knowledge graph.
[0171] S6.1: Extract structured data entities from the time sequence optimized data set as nodes, extract unstructured semantic fragments as hyperedges, construct a spatio-temporal hypergraph, and bind a time window label to the hyperedge;
[0172] Specifically, the structured data part in the time sequence optimized data set is scanned, and data entities with unique identifiers are identified. The key attributes (user identifier, timestamp and data type label) of each data entity are encapsulated as node attributes, and all nodes are output to form a node set;
[0173] Parse the unstructured semantic fragments in the time sequence optimized data set, and each unstructured semantic fragment is associated with a set of semantic core entities. When an unstructured semantic fragment is associated with three or more nodes at the same time, the unstructured semantic fragment is defined as a hyperedge;
[0174] Extract the timestamp of the unstructured semantic fragment in the time sequence optimized data set, obtain the start time and end time of the time window to which the timestamp belongs, and bind the time interval from the start time to the end time to the corresponding hyperedge as a time window label to form a hyperedge set;
[0175] Input the node set and the hyperedge set into the spatio-temporal hypergraph constructor to establish the connection relationship between the hyperedge and the node, arrange all hyperedges in ascending order according to the time sequence of the time window label, and output the spatio-temporal hypergraph structure with the time window label.
[0176] It should be noted that the unstructured semantic fragment is derived from the natural language understanding process of the NLU enhanced cross-system focused crawler: visible interface elements (such as button labels, table titles) are encoded into semantic vectors by CLIP as semantic primitives of hyper-edges; the time window label corresponds to the temporal context of the crawler's crawling action (such as the causal chain of click events and data rendering); the cross-system association strength is guaranteed by the cross-domain parsing ability of the crawler (such as automatically identifying iframe nested content).
[0177] S6.2: According to the bound time window label, extract the time dimension feature, extract the node feature of the spatio-temporal hypergraph through the pre-trained spatio-temporal feature fusion network, dynamically associate the time dimension feature, and output the node feature vector;
[0178] It should be noted that the pre-training process of the spatio-temporal feature fusion network is: according to the synthesized spatio-temporal hypergraph dataset (the node attribute is subject to the real data distribution, and the hyper-edge is generated by the random walk algorithm and bound with the time window label containing the time fluctuation parameter), the supervised training is performed, the spatio-temporal feature fusion network extracts the node feature and the time feature, and the node classification loss function and the time correlation loss function are calculated after the variance weighted fusion, the parameters are iteratively updated through the back propagation and the adaptive matrix estimator optimizer, the network parameters are fixed, and the spatio-temporal feature fusion network is obtained.
[0179] Specifically, the time window label bound in the spatio-temporal hypergraph is extracted, the starting time difference of adjacent time window labels is obtained to generate a time window label interval sequence, the time window label interval sequence is input into a time feature extractor, and a time dimension feature vector is output;
[0180] The spatio-temporal feature fusion network traverses each node of the spatio-temporal hypergraph, aggregates the semantic fragment content of all hyper-edges connected to the current node, and fuses the semantic information of the connected hyper-edges through graph convolution operation to generate an initial node feature vector;
[0181] The time window label interval sequence statistical value of the associated hyper-edge of the current node is extracted, the variance value of the time window label interval sequence is obtained, and the variance value is used as the weighting coefficient of the time dimension feature vector. The weighted sum operation is performed on the initial node feature vector;
[0182] The weighted and fused feature vector is output as the final node feature vector.
[0183] S6.3: Inject the statistical distribution feature into the node feature vector through the cross-modal attention mechanism to generate a fusion feature matrix;
[0184] Specifically, according to the statistical distribution feature of the operation time interval sequence, the skewness, kurtosis and variance values are extracted, and the skewness, kurtosis and variance values are spliced into a statistical feature vector;
[0185] The node feature vector is taken as a query vector, the statistical feature vector is taken as a key-value pair vector, a similarity weight distribution of the query vector and the key-value pair vector is obtained through a cross-modal attention mechanism, and an attention weight matrix is generated;
[0186] A weighted sum operation is performed on the statistical feature vector using the attention weight matrix, a weighted statistical feature vector is output, the weighted statistical feature vector is spliced with the node feature vector along a feature dimension, and a fusion feature matrix is generated.
[0187] S6.4: Based on the fusion feature matrix, the inter-entity relationship strength is obtained, the entity relationship triplets meeting the preset strength threshold are extracted, and the time window label is bound to generate a cross-system associated knowledge graph.
[0188] It should be noted that the preset strength threshold is set based on the upper limit of the false positive rate of cross-system association, and the example value is 0.75.
[0189] Specifically, the feature vectors of all entity pairs in the fusion feature matrix are traversed, the inner product similarity of each pair of entity feature vectors is obtained, and the inner product similarity value is taken as the inter-entity relationship strength value;
[0190] The inter-entity relationship strength value is compared with the preset strength threshold, and the entity pairs with the inter-entity relationship strength value greater than the preset strength threshold are retained, and the entity pairs with the inter-entity relationship strength value less than the preset strength threshold are discarded;
[0191] The retained entity pairs are added with semantic relationship labels (the semantic relationship labels are identified by semantic role labeling of unstructured semantic fragments), and a triplet structure of head entity-relation-tail entity is formed;
[0192] The time window labels of the hyper-edges corresponding to the head entity and the tail entity in the triplet are extracted, when the time window labels of the head entity and the tail entity are different, the intersection of the time window label time intervals is taken as the bound time window label, and when the time window labels of the head entity and the tail entity are the same, the time window label is directly bound;
[0193] All the triplets with the bound time window labels are input into a knowledge graph constructor, the knowledge graph constructor organizes the triplets in the time order of the time window labels, and generates a cross-system associated knowledge graph.
[0194] In summary, the present application synchronizes the phase of the neural pulse coupled oscillator behavior feature and the natural language instruction, realizes accurate alignment of cross-modal intent, controls phase difference fluctuation, dynamically suppresses conflict feature weight based on the entropy constraint decay factor generated based on the phase difference change rate, reduces the error of the SQL query template and the RAG retrieval path, and guarantees the real-time and accuracy of cross-system acquisition target analysis.
[0195] It should be noted that the above examples are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced, without departing from the spirit and scope of the technical solutions of the present application, which should be covered in the scope of the claims of the present application.
Claims
1. An AI large model-based cross-system data collection method, characterized in that: include, Collect user behavior trajectories and generate multimodal behavior feature vectors; The multimodal behavioral feature vector is input into the dynamic intent analysis layer, and then fused with real-time received natural language commands to generate an entropy-constrained intent vector with corrected weights. The specific steps are as follows. The multimodal behavioral feature vectors are modulated into a multi-frequency sinusoidal pulse sequence, while the real-time received natural language commands are encoded into a Gaussian distributed pulse sequence. The phase difference between a multi-frequency sinusoidal pulse sequence and a Gaussian distributed pulse sequence is synchronized through a pre-trained oscillator coupling network. The rate of change of the phase difference is calculated, and a gating operation is performed to generate a gating signal. Integrate the gated signal with the rate of change of phase difference to generate an entropy-constrained attenuation factor; The weight distribution of the multimodal behavioral feature vector is modulated using an entropy-constrained decay factor to generate an entropy-constrained intent vector with corrected weights. Based on the integration target collected by the entropy-constrained intent vector parsing, the structured data components are mapped to SQL query templates, while the unstructured data components are mapped to RAG enhanced retrieval paths, and a unified instruction set is output. Execute a unified instruction set, compile the SQL query template into a native query statement through a protocol converter, and convert the RAG enhanced retrieval path into an unstructured data component scanning operation to obtain the initial dataset; By utilizing the operation time interval sequence in user behavior trajectory, the cleaning priority of the initial dataset is dynamically scheduled to generate a time-series optimized dataset; The time-series optimized dataset is input into the enhanced retrieval engine, and the statistical features extracted from the structured data components by the AI large model and the semantic fragments generated by the unstructured data components by RAG are integrated to generate a cross-system related knowledge graph. 2.The AI large model-based cross-system data collection method of claim 1, wherein: The user behavior trajectory includes a sequence of pen grip pressure intensity, a sequence of screen touch area coordinates, and a sequence of operation time intervals.
3. The cross-system data acquisition method based on a large AI model as described in claim 2, characterized in that: The specific steps for collecting user behavior trajectories and generating multimodal behavior feature vectors are as follows. Perform Daubechies wavelet decomposition on the pen grip pressure intensity sequence to extract detail coefficients and output pressure energy entropy values; The movement speed vector is obtained based on the coordinate sequence of the touch area, and the average value of the pulse code sequence is generated by a simulated retinal pulse firing rate engine. The phase space of the operation time interval sequence is reconstructed to generate an embedded vector sequence, and the operation rhythm complexity feature value is generated by the permutation entropy algorithm. The pressure energy entropy value, pulse code sequence mean, and operation rhythm complexity feature value are input into the cortical column competitive encoder, normalized, and output as a multimodal behavior feature vector.
4. The cross-system data acquisition method based on a large AI model as described in claim 1, characterized in that: The cross-system acquisition targets include structured data components and unstructured data components generated by entropy-constrained intention vector decomposition; The structured data components include user identifier weight, time-sensitive weight, and data type preference weight; The unstructured data components include semantic sensitivity weights, urgency weights, and association strength weights.
5. The cross-system data acquisition method based on a large AI model as described in claim 4, characterized in that: The specific steps for outputting a unified instruction set are as follows. A SQL query template is generated by dynamically activating structured data component vectors through a differentiable symbol rule generator. Perform neural concept space projection on unstructured data components, generate semantic potential gradient paths through Monte Carlo integration, and output RAG-enhanced retrieval paths; By fusing SQL query templates and RAG-enhanced retrieval paths through a cross-modal semantic encoder, a low-dimensional instruction vector is output as a unified instruction set.
6. The cross-system data acquisition method based on a large AI model as described in claim 5, characterized in that: The SQL query template is a dynamically condition-activated relational database query instruction carrier; The RAG-enhanced retrieval path is a semantically focused retrieval path generated by neural concept space projection.
7. The cross-system data acquisition method based on a large AI model as described in claim 1, characterized in that: The specific steps for obtaining the initial dataset are as follows. The unified instruction set is decoded into structured query microinstruction vectors and unstructured scan microinstruction vectors; Based on the semantic sensitivity weight, urgency weight, and correlation strength weight in the unstructured scanning micro-instruction vector and unstructured data components, a precision semantic scanning stream is generated through a dynamic precision control mechanism. The SQL query template is parameterized and compiled with the user identifier weight, time sensitivity weight, and data type preference weight in the structured data components to generate the native query statement; The heterogeneous execution flow is dynamically orchestrated using a gradient-constrained entropy minimization strategy to execute native query statements and semantic scanning flow, thereby obtaining the initial dataset.
8. The cross-system data acquisition method based on a large AI model as described in claim 7, characterized in that: The specific steps for generating the time-series optimization dataset are as follows. Perform phase space reconstruction on the operation time interval sequence in the user behavior trajectory, calculate the maximum Lyapunov exponent and statistical distribution characteristics, and generate chaotic feature vectors; Dynamic scheduling weights are calculated based on chaotic feature vectors using a chaotic entropy weight generation mechanism. By combining the dynamic scheduling weights with the size of each data subset in the initial dataset, a cleaning priority for the initial dataset is generated, thus creating a dynamic cleaning queue. Perform chaotic filtering and timestamp calibration on the dynamic cleaning queue, and output a time-series optimized dataset.
9. The cross-system data acquisition method based on a large AI model as described in claim 8, characterized in that: The specific steps for generating the cross-system related knowledge graph are as follows: Structured data entities are extracted from the time-series optimization dataset as nodes, unstructured semantic fragments are extracted as hyperedges, a spatiotemporal hypergraph is constructed, and time window labels are bound to the hyperedges. Based on the bound time window labels, extract time dimension features, extract node features of the spatiotemporal hypergraph through a pre-trained spatiotemporal feature fusion network, dynamically associate the time dimension features, and output node feature vectors. Statistical distribution features are injected into node feature vectors through a cross-modal attention mechanism to generate a fused feature matrix; The strength of relationships between entities is obtained based on the fusion feature matrix. Entity relationship triples that meet the preset strength threshold are extracted and bound with time window labels to generate a cross-system related knowledge graph.
Citation Information
Patent Citations
Unified construction method for intention recognition system and semantic vector recall system
CN119691577A
Multi-modal time sequence mixed query system and method
CN120086243A