Data query method and system based on large model

By building a three-layer pipeline large model and a hybrid precision calculation framework based on query complexity prediction, combined with the Sentence-BERT model and dynamic cache elimination strategy, the problem that existing data query methods cannot effectively capture user intentions and flexibly adjust calculation accuracy is solved, and high-precision and secure data query effect is achieved.

CN120197219AActive Publication Date: 2025-06-24MAILEFENG (XIAMEN) E-COMMERCE CO LTD

Patent Information

Application Number
CN202510626220.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-24
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

Existing data query methods cannot effectively capture the user's behavioral characteristics and intentions, resulting in reduced accuracy and user experience of query results. When processing complex queries, it is impossible to flexibly adjust the calculation accuracy and resource allocation, resulting in reduced performance.

Method used

A three-layer pipeline model is built using a data query method based on the big model, and a three-layer pipeline model is collected and quantified by user behavior monitoring in real time. A mixed accuracy calculation framework based on query complexity prediction is established, and the calculation accuracy is dynamically adjusted. Query statement vectors are generated through the Sentence-BERT model, a dynamic cache elimination strategy is established, and finally a hierarchical detection strategy is constructed to desensitize sensitive information.

Benefits of technology

It improves the understanding of query intention and the desensitization effect of sensitive information, enhances the accuracy and security of data query, and improves user experience and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197219A_ABST
    Figure CN120197219A_ABST
Patent Text Reader

Abstract

The invention discloses a data query method and system based on a large model, and relates to the technical field of data processing.The method comprises the steps that a three-layer assembly line large model is constructed, and all layers are connected in series through a data bus; user behavior monitoring is carried out, when a user submits a query request, behavior characteristics of the user are collected in real time, the intention intensity of the user is quantified through behavior characteristic fusion, and incremental updating is carried out; constructing a mixed precision calculation framework based on query complexity prediction, and compensating precision loss; generating a query statement vector through the model, and establishing a dynamic cache elimination strategy; and constructing a hierarchical detection strategy, desensitizing the queried content, and meanwhile, setting a rule to authorize the user to restore the original data. According to the method, the understanding of the query intention and the desensitization effect of sensitive information are improved, the response speed and the resource utilization rate of the system are improved, the privacy of user data is guaranteed, the safety and the practicability of the data query method are improved, and the dual goals of comprehensive user experience optimization and data protection are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a data query method and system based on a large model. Background Art

[0002] In the past few decades, with the rapid development of information technology, data query methods have also undergone significant evolution. From the initial simple keyword-based queries to the current intelligent queries using complex algorithms and models, the progress of data query technology is not only reflected in the improvement of functions and efficiency, but also in the continuous enhancement of its intelligence. Especially in recent years, the wide application of deep learning and big data analysis methods has enabled data query systems to more accurately understand and analyze users' query intentions. The core of this technological development lies in how to utilize large models, combine user behavior characteristics and query complexity prediction, optimize the query response time and the relevance of results, so as to better meet user needs.

[0003] Although the current query methods perform well in many aspects, there are still some key deficiencies that affect the accuracy and efficiency of queries. For example, most of the existing technologies fail to fully consider users' behavior characteristics, resulting in the system being unable to achieve personalized responses when dealing with different users' query intentions. In addition, when dealing with complex queries, many systems cannot dynamically adapt to changes in query conditions due to the fixity of algorithms, thus reducing the quality of query results. In terms of intention analysis and query output, existing solutions often lack an effective real-time feedback mechanism and cannot quickly respond to user behavior, thereby affecting the efficiency and accuracy of queries. The data query method based on a large model proposed by our invention can effectively solve these existing deficiencies by establishing a multi-level data processing pipeline, a multi-dimensional user behavior monitoring model, and a dynamic precision calculation framework based on query complexity; among them, the personal information extracted in the present invention complies with national laws and regulations and is carried out with the authorization of the parties concerned. Summary of the Invention

[0004] In view of the above existing problems, the present invention solves the following technical problems: solving the problem that traditional data query methods cannot effectively capture users' behavior characteristics and intentions, resulting in a reduction in the accuracy of query results and user experience. Solving the problem that existing methods often cannot flexibly adjust the calculation precision and resource allocation when dealing with complex queries, resulting in performance degradation. Solving the challenge of how to implement effective retrieval while ensuring data security as data privacy protection is increasingly emphasized; among them, the personal information extracted in the present invention complies with national laws and regulations and is carried out with the authorization of the parties concerned.

[0005] To solve the above technical problems, a data query method based on a large model is proposed, including,

[0006] Build a three - layer pipeline large model, with each layer connected in series through a data bus; conduct user behavior monitoring. When a user submits a query request, real - time collect the user's behavior characteristics, quantify the user's intention intensity through behavior feature fusion, and adopt a dual - channel update strategy for incremental update; build a mixed - precision calculation framework based on query complexity prediction to compensate for precision loss; use the Sentence - BERT model to generate query statement vectors and establish a dynamic cache elimination strategy; build a hierarchical detection strategy to desensitize the queried content and at the same time set rules to authorize users to restore the original data.

[0007] As a preferred solution of a data query method based on a large model according to the present invention, wherein: the three - layer pipeline large model includes data processing through query input, query processing, and query output, and is connected in series through a data bus;

[0008] Through user behavior monitoring in the query output layer, when query data is input, capture the user's behavior characteristics, and input the results to the query processing layer for intention analysis and cache processing, and build a user intention intensity model for update, and perform desensitized output through the query output layer.

[0009] As a preferred solution of a data query method based on a large model according to the present invention, wherein: the user behavior monitoring includes conducting user behavior monitoring, and at the same time starting to understand the user intention intensity and building an intention classification network. When a user submits a query request, real - time collect the user's behavior characteristics, including click - through rate, stay time, and time interval, and perform a normalization operation on the behavior characteristics;

[0010] Consider the user's click - through rate and stay time, quantify the user's intention intensity through behavior feature fusion, adopt a deep neural network architecture. The input layer of the deep neural network architecture receives the normalized behavior feature vector, and after non - linear transformation through three fully - connected layers, outputs a normalized intention credibility score value between 0 and 1 : ;

[0011] wherein, is the click - through rate, is the stay time, is the time interval, is the activation function, is the base of the natural logarithm.

[0012] As a preferred solution of a data query method based on a large model according to the present invention, wherein: the quantization of the user intention intensity includes, after outputting the intention credibility score result, the intention credibility score result is injected into the intention understanding engine in real time, triggering a dual-channel update strategy,

[0013] When > 0.5, it is determined as a high-confidence query scenario, and the current query request is automatically marked as a high-quality training sample, and the supervised learning process is started. The user intention intensity model inputs the query request submitted by the current user into the intention classification network for intention prediction, calculates the cross-entropy loss between the intention prediction result and the actual intention, and uses the stochastic gradient descent algorithm with momentum to update the network weight parameters. The learning rate is dynamically adjusted according to the intention credibility score value, and the higher the score, the larger the learning rate;

[0014] When ≤ 0.5, it is determined as a low-confidence query scenario, and it is automatically switched to the self-supervised learning mode. The current query statement is subjected to two differential data augmentations to generate positive sample pairs and input them into the siamese network, and the feature space similarity loss is calculated. At the same time, the historical model parameter distribution is introduced as a regularization constraint, and the KL divergence is used to measure the difference between the current parameters and the parameters of the previous version to ensure the knowledge continuity in the update process of the user intention intensity model;

[0015] After incremental update, the query content of the user is used as the query statement to be processed and input into the query processing layer for intention analysis.

[0016] As a preferred solution of a data query method based on a large model according to the present invention, wherein: the compensation for precision loss includes that after the system receives the query statement to be processed, it calls the syntax parser for in-depth analysis and dynamically constructs a multi-dimensional feature vector, including the syntax tree depth, the number of key operation nodes, and the context window length;

[0017] Based on the multi-dimensional feature vector, calculate the query complexity score : ;

[0018] Among them, is the syntax tree depth, is the number of key operation nodes, is the context window length, is the vector of the user intention intensity, which is used to represent the clarity of the intention. When is larger, the score is higher;

[0019] The current score result is input into the computing resource scheduler in real time, triggering a differential precision execution strategy:

[0020] When happens, the computing precision adopts FP32, and the hardware unit allocation strategy adopts exclusive occupation of TensorCore units;

[0021] When happens, the computing precision adopts FP16, and the hardware unit allocation strategy adopts shared allocation of CUDA stream processors;

[0022] When happens, the computing precision adopts INT8, and the hardware unit allocation strategy adopts batch packaging and execution on the AI acceleration module;

[0023] When in a highly trusted query scenario, for low-precision computing units including INT8 and FP16, quantization errors are generated. During the backpropagation stage, gradient compensation is implemented. After compensation according to the calculation results, the compensated query statement is cached.

[0024] As a preferred solution of the data query method based on a large model according to the present invention, wherein: the dynamic cache eviction strategy includes using the Sentence-BERT model to generate 768-dimensional query vectors and calculating the value of all cache entries: ;

[0025] Among them, is the value score of cache entry k, is the access count of cache entry k, and 0.7 is used to adjust the influence degree of frequency on the cache value to avoid excessive dominance of high-frequency items; is the cosine similarity of the query vector, is the current timestamp, is the last access time, is the vector representation of the current query, is the vector of cache entry k, is the access frequency of all cache entries;

[0026] Sort by in descending order and retain the maximum value in . Insert the new query result into the cache with . Calculate the total cache value . When and 's ratio is greater than 0.3, trigger the sensitive detection enhancement mode. When the new query result inserted into the cache triggers the sensitive detection enhancement mode, mark the current new query result as a sensitive word.

[0027] As a preferred solution of a data query method based on a large model according to the present invention, wherein: the desensitization includes that when the sensitive detection enhancement mode is triggered, a hierarchical detection strategy is adopted, including entity-level detection, relationship-level detection, and structure-level detection;

[0028] Calculate the sensitive score : ;

[0029] Wherein, is the variable index, is the three-level detection weight, and the three levels are respectively valued at 0.5, 0.3, and 0.2; is the word frequency of the sensitive word j, is the position of the distribution entropy value, and q is the total number of detection levels;

[0030] When it is the high-sensitivity height, and all contents are completely masked with asterisks;

[0031] When it is the medium-sensitivity height, and the part that triggers the sensitive detection enhancement mode is masked with asterisks;

[0032] When it is the low sensitivity, and the original text is output;

[0033] Generate an access token and number for all masked contents using HMAC-SHA256. The key management adopts a threshold signature scheme to divide the user permissions, assign the access token to users with different permissions, and define the user permissions according to the user identity during user registration. When the user needs to query the masked content, input the number of the access token, automatically identify whether the number is correct and whether the corresponding user has the permission, and confirm the user identity through password or face recognition to prevent theft. When the number recognition and user identity are correct and the corresponding user has the viewing permission, the original text is output.

[0034] Another object of the present invention is to provide a data query system based on a large model.

[0035] As a preferred solution of a data query system based on a large model according to the present invention, wherein, it includes a query input module, a query processing module, and a query output module;

[0036] Construct a three-layer pipeline large model, and each layer is connected in series through a data bus;

[0037] The query input module includes a user monitoring unit and an update unit. The user monitoring unit establishes a user behavior monitoring model. When a user submits a query request, it collects the user's behavior characteristics in real time. The update unit fuses and quantifies the user intention intensity based on the behavior characteristics and performs incremental updates using a dual-channel update strategy.

[0038] The query processing module includes a compensation unit and a cache unit. The compensation unit constructs a mixed-precision calculation framework based on query complexity prediction to compensate for precision loss. After compensating according to the calculation results, it caches the compensated query statement. The cache unit generates query statement vectors using Sentence-BERT and establishes a dynamic cache elimination strategy.

[0039] The query output module constructs a hierarchical detection strategy to desensitize the queried content and at the same time sets rules to authorize users to restore the original data.

[0040] Advantages of the present invention: By constructing a three-layer pipeline large model, the present invention respectively realizes a hierarchical structure for query input, processing, and output, ensuring that the user's original query and behavior characteristics can be effectively captured and analyzed, thereby improving the understanding of query intentions and the desensitization effect of sensitive information. This process greatly enhances the accuracy and security of data queries. Then, by establishing a real-time user behavior monitoring model, it timely collects the user's behavior characteristics, quantifies the user intention intensity, and adopts a dynamic update mechanism. This not only greatly improves the accuracy of query intentions but also makes the overall learning of the model more in line with the actual needs of users. Constructing a mixed-precision calculation framework, dynamically adjusting the calculation precision, effectively reduces the consumption of computing resources. At the same time, implementing a gradient compensation mechanism in high-trust queries ensures the stability of the model and the accuracy of query results. Using Sentence-BERT to generate query vectors and implementing a dynamic cache elimination strategy improves the system's response speed and resource utilization rate, ensuring the timeliness and effectiveness of query results. Finally, constructing a multi-granularity sensitive word detection model, protecting the privacy of user data through a hierarchical detection strategy, and combining a flexible permission management mechanism to ensure that only authorized users can access sensitive information, significantly improves the security and practicality of the entire data query method, thus achieving the dual goals of comprehensive user experience optimization and data protection. Brief Description of the Drawings

[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0042] Figure 1The overall flowchart of a data query method based on a large model provided by an embodiment of the present invention.

[0043] Figure 2 The system schematic diagram of a data query system based on a large model provided by an embodiment of the present invention. Detailed implementation manners

[0044] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0046] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is mutually exclusive with other embodiments individually or selectively.

[0047] The present invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of the present invention, for ease of explanation, the cross-sectional views showing the device structure will be enlarged locally out of the general scale, and the schematic diagrams are only examples and should not limit the protection scope of the present invention here. In addition, in actual production, three-dimensional spatial dimensions including length, width, and depth should be included.

[0048] At the same time, in the description of the present invention, it should be noted that the orientation or positional relationships indicated by terms such as "upper, lower, inner, and outer" are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention. In addition, the terms "first, second, or third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0049] Unless otherwise clearly defined and limited in the present invention, the terms "installation, connection, and coupling" shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can also be a mechanical connection, an electrical connection, or a direct connection, and can also be indirectly connected through an intermediate medium, or can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0050] Example 1. Refer to Figure 1 , which is the first embodiment of the present invention. This embodiment provides a data query method based on a large model, including:

[0051] S1: Construct a three-layer pipeline large model, and each layer is connected in series through a data bus.

[0052] Furthermore, the three-layer pipeline large model includes data processing through query input, query processing, and query output, and is connected in series through a data bus; the following process is processed based on the three-layer pipeline large model.

[0053] The query output layer monitors user behavior. When query data is input, it captures the user's behavior characteristics, and inputs the results to the query processing layer for intention analysis, cache processing, and construction and update of the user intention intensity model, and performs desensitized output through the query output layer; among them, the personal information extracted involved in the present invention complies with national laws and regulations and is carried out with the authorization of the parties.

[0054] S2: Conduct user behavior monitoring. When the user submits a query request, the user's behavior characteristics are collected in real time, the user intention intensity is quantified through behavior feature fusion, and an incremental update is performed using a dual-channel update strategy.

[0055] It should also be noted that when conducting user behavior monitoring, at the same time, start to understand the user intention intensity and construct an intention classification network. When the user submits a query request, the user's behavior characteristics are collected in real time, including click-through rate, stay time, and time interval, and normalization operations are performed on the behavior characteristics;

[0056] The click-through rate is the number of query operations counted per unit time, the stay time is the duration of staying on the result page recorded, and the time interval is the time difference calculated for adjacent queries;

[0057] Considering the user click-through rate and stay time, the user intention intensity is quantified through behavior feature fusion. Using a deep neural network architecture, the input layer of the deep neural network architecture receives the normalized behavior feature vector, and after non-linear transformation through three fully connected layers, outputs a standardized intention credibility score value between 0 and 1 : ;

[0058] Among them, is the click-through rate, is the dwell time, is the time interval, is the activation function, is the base of the natural logarithm.

[0059] It should be noted that after the output of the intention credibility score result, the intention credibility score result is injected into the intention understanding engine in real time, triggering a dual-channel update strategy;

[0060] When > 0.5, it is determined as a high-confidence query scenario, and the current query request is automatically marked as a high-quality training sample, starting the supervised learning process. The user intention intensity model inputs the query request submitted by the current user into the intention classification network for intention prediction, calculates the cross-entropy loss between the intention prediction result and the actual intention, and updates the network weight parameters using the stochastic gradient descent algorithm with momentum. The learning rate is dynamically adjusted according to the intention credibility score value, and the higher the score, the larger the learning rate;

[0061] When ≤ 0.5, it is determined as a low-confidence query scenario, and it automatically switches to the self-supervised learning mode. The current query statement is subjected to two differential data augmentations (including operations such as synonym replacement and word order swapping), generating positive sample pairs to input into the twin network, calculating the feature space similarity loss. At the same time, the historical model parameter distribution is introduced as a regularization constraint, and the KL divergence is used to measure the difference between the current parameters and the previous version parameters to ensure the knowledge continuity during the model update process;

[0062] After incremental update, the user query content is used as the query statement to be processed and input into the query processing layer for intention analysis.

[0063] S3: Construct a mixed-precision calculation framework based on query complexity prediction to compensate for precision loss.

[0064] Furthermore, after the system receives the query statement to be processed, it calls the syntax parser for in-depth analysis and dynamically constructs a multi-dimensional feature vector, including the depth of the syntax tree, the number of key operation nodes, and the context window length:

[0065] The depth of the syntax tree is the path length from the root node to the deepest leaf node;

[0066] The number of key operation nodes is to count the number of key operators such as JOIN, WHERE, GROUP BY, etc.;

[0067] The context window length is the length of the query statement measured in characters;

[0068] Calculate the query complexity score based on the multi-dimensional feature vector : ;

[0069] Among them, is the depth of the syntax tree, is the number of key operation nodes, is the length of the context window, is the vector of user intention intensity, which is used to represent the clarity of the intention. When is larger, the score is higher;

[0070] Input the current score result into the computing resource scheduler in real time to trigger a differentiated precision execution strategy:

[0071] When , the computing precision adopts FP32, and the hardware unit allocation strategy adopts exclusive occupation of TensorCore units;

[0072] When , the computing precision adopts FP16, and the hardware unit allocation strategy adopts shared allocation of CUDA stream processors;

[0073] When , the computing precision adopts INT8, and the hardware unit allocation strategy adopts batch packaging and execution on the AI acceleration module;

[0074] Specifically, FP32 is the precision format of 32-bit floating point, FP16 is the precision format of 16-bit floating point, and INT8 is the precision format of 8-bit floating point.

[0075] When in a high-trust query scenario, for low-precision computing units including INT8 and FP16, quantization errors are generated, and gradient compensation is implemented in the backpropagation stage: ;

[0076] Among them, is the gradient compensation amount, is the gradient norm of the loss function in the backpropagation calculation, is the activation function;

[0077] After compensating according to the calculation result, cache the compensated query statement.

[0078] S4: Use the Sentence-BERT model to generate query statement vectors and establish a dynamic cache eviction strategy.

[0079] It should be noted that the Sentence-BERT model is used to generate 768-dimensional query vectors, and calculate all cache items Value: ;

[0080] Wherein, is the value score of cache item k, is the access count of cache item k, and 0.7 is used to adjust the influence degree of frequency on cache value to avoid high-frequency items from overly dominating; is the cosine similarity of the query vector, is the current timestamp, is the last access time, is the vector representation of the current query, is the vector of cache item k, is the access frequency of all cache items;

[0081] Arrange in descending order and retain the maximum value in, insert the new query result into the cache with and calculate the total cache value , when the ratio of is greater than 0.3, trigger the sensitive detection enhancement mode. When the new query result inserted into the cache triggers the sensitive detection enhancement mode, mark the current new query result as a sensitive word.

[0082] S5: Construct a hierarchical detection strategy to desensitize the queried content and set rules to authorize users to restore the original data.

[0083] Specifically, when the sensitive detection enhancement mode is triggered, the hierarchical detection strategy adopted includes entity-level detection, relationship-level detection, and structure-level detection;

[0084] The entity-level detection is to identify sensitive entities such as names and addresses. For example, personal identity information (name, ID number / passport number, mobile phone number / fixed phone number), geographical location information (detailed address accurate to the house number, GPS coordinates with an accuracy of <10 meters), biometric identifiers (fingerprint / iris hash value, voiceprint feature vector); among them, the personal information extracted in the present invention complies with national laws and regulations and is carried out with the authorization of the parties concerned.

[0085] The relationship-level detection is to discover sensitive patterns such as "A transfers to B". For example, fund flow patterns (transfer relationship: "A remits X yuan to B", large transaction: "single transaction > 500,000 yuan, etc.), account association patterns (multiple accounts bound to the same mobile phone number, abnormal IP address association), behavior time series patterns (high-frequency queries > 10 times / minute, access during non-working hours, etc.); among them, the personal information extracted in the present invention complies with national laws and regulations and is carried out with the authorization of the parties concerned.

[0086] The structural-level detection is to analyze sensitive combinations in the table Schema. For example, database structure risks (such as the coexistence of three columns of sensitive fields (e.g., "name + ID card + bank card number"), high-privilege view definitions (SELECT * FROM salary), etc.), metadata features (sensitive words in field naming (including "password" / "secret"), non-normalized designs (storing credit card CVV codes)), interface exposure risks (API returns too many fields (more than 200% of the query requirements), unencrypted JDBC connection configurations); among them, all the personal information extracted in the present invention complies with national laws and regulations and is carried out with the authorization of the parties concerned.

[0087] Calculate the sensitivity score : ;

[0088] Among them, is the variable index, is the weight of the three-level detection, and the three levels take values of 0.5, 0.3, and 0.2 respectively; is the word frequency of the sensitive word j, is the position is the distribution entropy value of, q is the total number of detection levels, q = 3;

[0089] When it is the high-sensitivity level, and all content is completely masked with asterisks. For example, for the ID card number 11111111, the mask is ********; among them, all the personal information extracted in the present invention complies with national laws and regulations and is carried out with the authorization of the parties concerned.

[0090] When it is the medium-sensitivity level, and the part that triggers the sensitive detection enhancement mode is masked with asterisks. For example, for Zhang San, the mask is Zhang *; among them, all the personal information extracted in the present invention complies with national laws and regulations and is carried out with the authorization of the parties concerned.

[0091] When it is the low-sensitivity level, and the original text is output; among them, all the personal information extracted in the present invention complies with national laws and regulations and is carried out with the authorization of the parties concerned.

[0092] Generate access tokens and numbers for the content of all masks using HMAC-SHA256. The key management adopts a threshold signature scheme to divide user permissions and assign access tokens to users with different permissions. User permissions are defined according to the user's identity during user registration. When a user needs to query the mask content, enter the number of the access token, automatically identify whether the number is correct and whether the corresponding user has permission, and confirm the user's identity through password or face recognition to prevent theft. When the number recognition and user identity are correct and the corresponding user has the viewing permission, the original text will be output.

[0093] Example 2, the second example of the present invention, which is different from Example 1 in that:

[0094] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0095] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0096] More specific examples (nonexhaustive list) of computer-readable media include the following: electrical connections (electronic devices) with one or more wirings, portable computer disk cartridges (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0097] It should be understood that the various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0098] Example 3, referring to Figure 2 , is the third embodiment of the present invention. This embodiment provides a large model-based data query system, including a query input module, a query processing module, and a query output module;

[0099] Construct a three-layer pipeline large model, and each layer is connected in series through a data bus;

[0100] The query input module includes a user monitoring unit and an update unit. The user monitoring unit establishes a user behavior monitoring model. When a user submits a query request, it real-time collects the user's behavior characteristics. Through the update unit, the behavior characteristics are fused to quantify the user intention intensity, and an incremental update is performed using a dual-channel update strategy;

[0101] The query processing module includes a compensation unit and a cache unit. The compensation unit constructs a mixed-precision calculation framework based on query complexity prediction to compensate for precision loss. After compensation according to the calculation results, the compensated query statement is cached. Through the cache unit, a query statement vector is generated using Sentence-BERT, and a dynamic cache elimination strategy is established;

[0102] The query output module constructs a hierarchical detection strategy to desensitize the queried content and at the same time sets rules to authorize users to restore the original data.

[0103] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and all of them should be covered by the scope of the claims of the present invention.

Claims

1. A data query method based on a large model, characterized in that: include, Build a three-layer pipeline model, with each layer connected in series through a data bus; Monitor user behavior. When a user submits a query request, collect the user's behavior characteristics in real time, quantify the user's intention intensity through behavior feature fusion, and use a dual-channel update strategy for incremental updates. Build a mixed-precision computing framework based on query complexity prediction to compensate for precision loss; Use the Sentence-BERT model to generate query sentence vectors and establish a dynamic cache elimination strategy; Build a hierarchical detection strategy to desensitize the queried content, and set rules to authorize users to restore the original data.

2. A data query method based on a large model as claimed in claim 1, characterized in that: The three-layer pipeline model includes data processing through query input, query processing, and query output, and connected in series through a data bus; The query output is monitored through user behavior. When the query data is input, the user's behavior characteristics are captured, and the results are input into the query processing, intent analysis and cache processing are performed, and a user intent strength model is constructed and updated, and the query output is desensitized.

3. A data query method based on a large model as claimed in claim 2, characterized in that: The user behavior monitoring includes: monitoring user behavior based on a three-layer pipeline model, and simultaneously starting to understand the user's intention intensity and build an intention classification network. When the user submits a query request, the user's behavior characteristics are collected in real time, including click rate, dwell time, time interval, and the behavior characteristics are normalized; Considering the user click rate and dwell time, the user intention strength is quantified by integrating behavioral features. A deep neural network architecture is adopted. The input layer of the deep neural network architecture receives the normalized behavioral feature vector. After nonlinear transformation of the three fully connected layers, a standardized intention credibility score value between 0 and 1 is output. : ; in, Click-through rate, For the stay time, is the time interval, is the activation function, is the base of natural logarithms.

4. A data query method based on a large model as claimed in claim 3, characterized in that: The quantification of the user intention strength includes: after outputting the intention credibility scoring result, the intention credibility scoring result is injected into the intention understanding engine in real time to trigger the dual-channel update strategy, when When >0.5, it is determined to be a high-credibility query scenario, and the current query request is automatically marked as a high-quality training sample. The supervised learning process is started, and the user intention strength model inputs the query request submitted by the current user into the intention classification network for intent prediction, calculates the cross entropy loss between the intent prediction result and the actual intent, and uses the stochastic gradient descent algorithm with momentum to update the network weight parameters. The learning rate is dynamically adjusted according to the intent credibility score value. The higher the score, the greater the learning rate; when When ≤0.5, it is judged as a low-credibility query scenario and automatically switches to self-supervised learning mode. The current query statement is enhanced twice with differential data, and positive samples are generated for input into the twin network. The feature space similarity loss is calculated. At the same time, the historical model parameter distribution is introduced as a regularization constraint, and the KL divergence is used to measure the difference between the current parameters and the parameters of the previous version to ensure the knowledge continuity during the update process of the user intent strength model. After incremental updating, the user's query content is input into the query processing layer as a pending query statement for intent analysis.

5. A data query method based on a large model as claimed in claim 4, characterized in that: The compensation for precision loss includes, after the system receives the query statement to be processed, calling the syntax parser to perform in-depth analysis and dynamically constructing a multi-dimensional feature vector, including the syntax tree depth, the number of key operation nodes, and the context window length; Calculate query complexity score based on multi-dimensional feature vector : ; in, is the syntax tree depth, is the number of key operation nodes, is the context window length, is the vector of user intention strength, which represents the degree of clarity of intention. The bigger, The higher the score; The current rating result Real-time input computing resource scheduler triggers differentiated precision execution strategies: when When , the calculation accuracy uses FP32, and the hardware unit allocation strategy uses exclusive occupation of Tensor Core units; when When , the calculation accuracy adopts FP16, and the hardware unit allocation strategy adopts shared allocation of CUDA stream processors; when When , the calculation accuracy is INT8, and the hardware unit allocation strategy adopts batch packaging to execute in the AI ​​acceleration module; When in a high-confidence query scenario, quantization errors are generated for low-precision computing units including INT8 and FP16. Gradient compensation is implemented in the back-propagation stage. After compensation is performed based on the calculation results, the compensated query statements are cached.

6. A data query method based on a large model as claimed in claim 5, characterized in that: The dynamic cache elimination strategy includes using the Sentence-BERT model to generate a 768-dimensional query vector and calculating the value: ; in, Score the value of cache item k, is the number of accesses to cache item k. 0.7 is used to adjust the impact of frequency on cache value to avoid excessive dominance of high-frequency items. is the cosine similarity of the query vector, is the current timestamp, is the last access time, is the vector representation of the current query, is the vector of cache items k, is the access frequency of all cache items; according to Keep descending order The new query result is the maximum value of Insert cache and count the total value of cache ,when and When the ratio is greater than 0.3, the sensitive detection enhancement mode is triggered. When the new query result inserted into the cache triggers the sensitive detection enhancement mode, the current new query result is marked as a sensitive word.

7. A data query method based on a large model as claimed in claim 6, characterized in that: The desensitization includes, when the sensitive detection enhancement mode is triggered, adopting a hierarchical detection strategy including entity-level detection, relationship-level detection and structure-level detection; Calculating sensitivity scores : ; in, is the variable index, is the three-level detection weight, and the three levels are 0.5, 0.3, and 0.2 respectively; is the frequency of sensitive word j, For location The distribution entropy value of , q is the total number of detection levels; when When , it is a high-sensitivity altitude, and all contents are completely masked with asterisks; when When , it is a medium-sensitive height, and the use of asterisks will trigger the partial mask of the sensitive detection enhanced mode; when When , it is low sensitivity and the original text is output; Use HMAC-SHA256 to generate access tokens and numbers for all masked contents. The threshold signature scheme is used for key management. Users are divided into permissions and access tokens are assigned to users with different permissions. User permissions are defined according to user identities when users register. When users need to query masked contents, they enter the number of the access token, automatically identify whether the number is correct and whether the corresponding user has permissions, and confirm the user identity through a password or face recognition to prevent theft. When the number identification and user identity are correct and the corresponding user has viewing permissions, the original text is output.

8. A data query system based on a large model, applied to a data query method based on a large model as claimed in any one of claims 1 to 7, characterized in that: It includes a query input module, a query processing module, and a query output module; Build a three-layer pipeline model, with each layer connected in series through a data bus; The query input module includes a user monitoring unit and an updating unit. The user monitoring unit establishes a user behavior monitoring model. When a user submits a query request, the user's behavior characteristics are collected in real time. The behavior characteristics are integrated and quantified by the updating unit to quantify the user's intention strength, and a dual-channel updating strategy is used for incremental updating. The query processing module includes a compensation unit and a cache unit. The compensation unit constructs a mixed precision calculation framework based on query complexity prediction to compensate for precision loss. After compensation according to the calculation result, the compensated query statement is cached. The query statement vector is generated by the cache unit using Sentence-BERT to establish a dynamic cache elimination strategy. The query output module constructs a hierarchical detection strategy, desensitizes the queried content, and sets rules to authorize users to restore the original data.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a data query method based on a large model as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a data query method based on a large model as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Three-dimensional digital system capable of displaying inundation compensation information of downstream beach area of Yellow River

    CN111177300A

  • Internet of Things system

    CN116368355A

  • Dynamic data pipeline construction method based on artificial intelligence and multi-modal data processing

    CN119830200A

  • Computer security based on artificial intelligence

    US20170214701A1

  • Temperature compensation method and system for solar air conditioner for vehicle

    WO2018176598A1

Cited By

  • Data authority management method and system based on four-dimensional authority matrix and dynamic adaptation

    CN121765706A