Data intelligent service method, device and equipment under mixed model and storage medium

By employing a data intelligence service approach based on a hybrid model, user intent is identified and data resources are analyzed, solving the problem of inaccurate data retrieval in existing technologies and achieving efficient and intelligent data asset services and analysis.

CN122633708APending Publication Date: 2026-08-25CHINA MERCHANTS BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610753481.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing data search tools cannot understand users' semantic and logical relationships, resulting in inaccurate and incomplete search results, low data retrieval efficiency, difficulty in meeting complex data asset retrieval needs, and inability to effectively provide data asset analysis and knowledge Q&A.

Method used

By adopting a data intelligence service approach under a hybrid model, we can predict service demand by identifying intent categories in user questions and answers and retrieval information, and provide intelligent data asset services based on intent category analysis of data resources.

Benefits of technology

It significantly improves the intelligence level and retrieval accuracy of data asset services, lowers the threshold for users to understand data, improves the efficiency and accuracy of data asset analysis, and reduces reliance on manual interpretation and the risk of errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122633708A_ABST
    Figure CN122633708A_ABST
Patent Text Reader

Abstract

The application discloses a data intelligent service method and device under a mixed model, equipment and a storage medium, relates to the technical field of artificial intelligence technology, and comprises the following steps: acquiring user question and answer information and user search information; identifying at least one of a corresponding query intention category, an inquiry intention category, an exploration behavior category and a search behavior category based on the user question and answer information or the user search information, and determining at least one of query intention category information, inquiry intention category information, exploration behavior category information and search behavior category information; predicting a corresponding service demand based on at least one of the query intention category information, the inquiry intention category information, the exploration behavior category information and the search behavior category information, and determining service demand prediction information; and analyzing a corresponding data resource based on the service demand prediction information, and determining a data resource analysis result. The application identifies query, inquiry, exploration and search intentions, performs data interpretation, and adapts intelligent service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to data intelligence service methods, apparatus, devices and storage media under hybrid models. Background Technology

[0002] With the continuous development of the digital economy and the deepening of digital transformation, the amount of data is expanding rapidly, accumulating massive data assets. In their daily work, data analysts need to quickly and accurately find the data assets they need and be able to understand complex data knowledge and data asset metadata to support data insight and delivery in high-time-efficiency scenarios such as business opportunity mining and risk prevention.

[0003] Currently, the existing approach involves using data search tools that rely on simple keyword matching and manually set fixed rules to retrieve data assets. When a user enters a query, the system returns search results based on keyword string matching or preset rules, and displays basic metadata information of the data assets in the results list. However, the existing approach cannot understand the user's semantics and logical relationships. Users find it difficult to accurately grasp the keywords that describe their needs, resulting in inaccurate, incomplete, and excessively vague search results. This leads to low data retrieval efficiency, difficulty in handling complex data asset retrieval needs, and an inability to effectively provide users with data asset analysis and knowledge Q&A. Therefore, how to accurately and effectively provide data asset services and analysis has become an urgent problem to be solved. Summary of the Invention

[0004] The main objective of this application is to provide a data intelligence service method, apparatus, device, and storage medium under a hybrid model, aiming to solve the technical problem of how to accurately and effectively perform data asset services and analysis.

[0005] To achieve the above objectives, this application proposes a data intelligence service method under a hybrid model, the method comprising: Obtain user question and answer information and user search information; Based on the user question and answer information or the user search information, at least one of the corresponding query intent category, inquiry intent category, exploration behavior category and search behavior category is identified, and at least one of the query intent category information, inquiry intent information, exploration behavior category and search behavior category information is determined; Based on at least one of the query intent category information, the inquiry intent category information, the exploration behavior category information, and the retrieval behavior category information, the corresponding service demand is predicted to determine the service demand prediction information. Based on the service demand forecast information, the corresponding data resources are analyzed to determine the data resource analysis results.

[0006] In one embodiment, the step of identifying at least one of the corresponding query intent category, inquiry intent category, exploration behavior category, and retrieval behavior category based on the user question-and-answer information or the user retrieval information, and determining at least one of the query intent category information, inquiry intent information, exploration behavior information, and retrieval behavior category information includes: Retrieve user's question history information; Based on the user questions and intent tags corresponding to the user question history information, the training batch sample size, learning rate, iteration rounds, random seed and weight pruning in the hyperparameters of the predefined intent recognition model are adjusted to determine the target intent recognition model; Based on the user question and answer information, the target intent recognition model is input to identify at least one of the query intent category and inquiry intent category under the corresponding model arrangement, and to determine at least one of the query intent category information and inquiry intent information; Based on the filtering conditions and keywords in the user search information, at least one of the corresponding exploration behavior category and search behavior category is identified, and at least one of the exploration behavior category information and search behavior category information is determined.

[0007] In one embodiment, the step of predicting the corresponding service demand based on at least one of the query intent category information, the inquiry intent category information, the exploration behavior category information, and the retrieval behavior category information, and determining the service demand prediction information, includes: Based on the query intent category information or the exploration behavior category information, the predefined recommendation model is input to predict the recommendation requirements in the corresponding service requirements and determine the service recommendation requirement information. Based on the query intent category information, the predefined knowledge question answering model is input to predict the knowledge question answering requirements in the corresponding service requirements and determine the knowledge question answering requirement information. Based on the search behavior category information, the search requirements in the corresponding service needs are predicted to determine the search requirement information; Service demand prediction information is obtained based on at least one of the service recommendation demand information, the knowledge question and answer demand information, and the retrieval demand information.

[0008] In one embodiment, the step of predicting the recommendation demand in the corresponding service demand based on the query intent category information or the exploration behavior category information input into a predefined recommendation model to determine the service recommendation demand information includes: Obtain data service metadata information, which includes basic table information, field information, and upstream and downstream lineage information; Based on the data service metadata information and the query intent category information or the exploration behavior category information, the predefined recommendation model is input to predict the corresponding data asset profile and extended keywords, and to determine the data profile information and extended keyword information. Based on the extended keyword information, determine the recall data service information under the corresponding multi-path recall strategy; Based on the data profile information and the recall service profile corresponding to the recall data service information, the recommended requirements in the corresponding service requirements are analyzed to determine the service recommendation requirement information.

[0009] In one embodiment, the step of determining the recall data service information under the corresponding multi-path recall strategy based on the extended keyword information includes: Based on the extended keyword information, the predefined recommendation model performs vector retrieval on the corresponding converted text vector to determine the vector retrieval result. Based on the vector retrieval results, the corresponding keywords are used to perform full-text retrieval using a predefined inverted index to determine the full-text retrieval results under the multi-path recall strategy. The full-text search results are compared with a predefined knowledge graph using a graph mining algorithm to determine the recall data service information under the corresponding multi-path recall strategy.

[0010] In one embodiment, the step of predicting the knowledge question-answering requirements in the corresponding service requirements based on the input of the query intent category information into a predefined knowledge question-answering model, and determining the knowledge question-answering requirement information, includes: Based on the query intent category information, a predefined knowledge base is input for retrieval, and the knowledge base retrieval results are determined. The knowledge base includes basic information of data tables, field information, code value information, graph information, and script information. Based on the retrieval results of the knowledge base, the documents in the knowledge base are sliced ​​to determine the sliced ​​text block information; A hybrid retrieval method combining predefined keyword retrieval and vector semantic retrieval is used to recall the fragments corresponding to the sliced ​​text block information from the knowledge base, and the recalled fragment information is determined. Based on the recalled fragment information, a predefined knowledge question-answering model is input to predict the knowledge question-answering requirements in the corresponding service requirements, thereby obtaining knowledge question-answering requirement information.

[0011] In one embodiment, the step of analyzing the corresponding data resources based on the service demand prediction information and determining the data resource analysis results includes: Retrieve user metadata information and associated statement script information; Based on the user metadata information and the associated statement script information, the predefined analysis model is input to identify the corresponding data table names and determine the data table list information; Based on the document name, basic description, and field code value in the data list information, the data resources corresponding to the service demand prediction information are analyzed to obtain data resource analysis results, which include service scenario description information and application scenario information.

[0012] Furthermore, to achieve the above objectives, this application also proposes a data intelligence service device under a hybrid model, the data intelligence service device under the hybrid model comprising: The acquisition module is used to acquire user question and answer information and user search information; The processing module is used to identify at least one of the corresponding query intent category, inquiry intent category, exploration behavior category and retrieval behavior category based on the user question and answer information or the user retrieval information, and to determine at least one of the query intent category information, inquiry intent information, exploration behavior information and retrieval behavior category information; The processing module is further configured to predict the corresponding service demand based on at least one of the query intent category information, the inquiry intent category information, the exploration behavior category information, and the retrieval behavior category information, and determine the service demand prediction information; The execution module is used to analyze the corresponding data resources based on the service demand prediction information and determine the data resource analysis results.

[0013] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the data intelligence service method under the hybrid model described above.

[0014] One or more technical solutions proposed in this application have at least the following technical effects: This embodiment proposes a data intelligence service method under a hybrid model, which involves acquiring user question-and-answer information and user search information; identifying at least one of the corresponding query intent category, inquiry intent category, exploration behavior category, and search behavior category based on the user question-and-answer information or the user search information, and determining at least one of the query intent category information, inquiry intent category information, exploration behavior information, and search behavior category information; predicting the corresponding service demand based on at least one of the query intent category information, inquiry intent information, exploration behavior information, and search behavior category information, and determining service demand prediction information; and analyzing the corresponding data resources based on the service demand prediction information, and determining the data resource analysis results. This application identifies the intent categories of queries, inquiries, explorations, and searches through user question-and-answer information or user search information, and predicts knowledge-based question-and-answer, service recommendations, and search needs. It then automatically identifies data table names to generate a list of data usage tables. Based on the text names, basic descriptions, and field code values ​​in the list of data usage tables, it performs semantic analysis on data resources to obtain data resource analysis results. This transforms complex, technical scripting languages ​​and data asset metadata into natural language descriptions that are easy for business personnel to understand, significantly reducing the understanding threshold of data assets, improving the efficiency and accuracy of data usage scenario analysis, while reducing reliance on manual interpretation and the risk of errors. This significantly enhances the intelligence level, search accuracy, and user understanding efficiency of data asset services, lowering the barrier for users to use data. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the data intelligence service method architecture under the hybrid model of this application; Figure 2 This is a flowchart illustrating an embodiment of the data intelligence service method under the hybrid model of this application. Figure 3 This is a flowchart illustrating Embodiment 2 of the data intelligence service method under the hybrid model of this application. Figure 4 This is a flowchart illustrating Embodiment 3 of the data intelligence service method under the hybrid model of this application. Figure 5A simplified flowchart illustrating the data intelligence service method under the hybrid model provided in this application embodiment; Figure 6 This is a schematic diagram of the module structure of the data intelligence service device under the hybrid model of this application embodiment; Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the data intelligence service method under the hybrid model in the embodiments of this application.

[0018] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0020] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0021] The main solution of this application embodiment is as follows: Obtain user question-and-answer information and user search information; based on the user question-and-answer information or the user search information, identify at least one of the corresponding query intent category, inquiry intent category, exploration behavior category, and search behavior category, and determine at least one of the query intent category information, inquiry intent category information, exploration behavior information, and search behavior category information; based on at least one of the query intent category information, inquiry intent information, exploration behavior information, and search behavior category information, predict the corresponding service demand, and determine service demand prediction information; based on the service demand prediction information, analyze the corresponding data resources, and determine the data resource analysis results.

[0022] In this embodiment, for ease of description, the following description will focus on the data intelligence service device under the hybrid model as the execution subject.

[0023] Because existing technologies cannot understand users' semantic and logical relationships, users find it difficult to accurately grasp the keywords that describe their own needs, resulting in inaccurate, incomplete, and too many vague search results. This leads to low data retrieval efficiency, difficulty in handling complex service data retrieval needs, and an inability to effectively provide users with data asset analysis and knowledge Q&A.

[0024] This application provides a solution, such as Figure 1 As shown, Figure 1This is a schematic diagram of the data intelligence service method architecture under the hybrid model of this application. The data intelligence service system under the hybrid model is divided into a user layer, a front-end application layer, a basic capability layer, and an AI capability layer. The user layer includes different roles such as data analysts, data managers, and data developers, providing input data for user interaction. The front-end application layer provides two main functional modules: intelligent question answering and intelligent analysis. The intelligent question answering module implements "finding data" and "asking for data" services in a dialogue format. The left side displays historical dialogue records, and the right side displays the current dialogue content. Based on intent recognition capabilities, it determines the user's intent and calls recommendation models or knowledge-based question answering capabilities to respond. The response content is output in a streaming manner and displays information such as asset name, basic description, response time, and references. It also provides related questions such as "You can continue asking" to guide further dialogue. The intelligent analysis module, based on data usage scenario analysis capabilities, uses NL2SQL technology to parse tables and analyze commonly used SQL scripts to analyze usage scenarios. It combines existing logical descriptions and rule-based judgments to determine the asset recommendation index (displayed as a number of stars, calculated based on asset level, timeliness, and eligibility). It uses a "Help me analyze [asset name]" approach to analyze the [logic] of the asset. The system uses a fixed question format of "[Analysis][Usage Scenarios][Recommendation Index]" to stream the logical description, usage scenarios, recommendation index, and key metadata information, and provides guidance on related questions. The front-end application layer also covers service process integration (such as calling asset recommendation capabilities on traditional search pages for quick answers, displaying recommendation results in rectangular cards and supporting jumps; and displaying analysis results in pop-up windows on asset details pages) and user feedback collection (providing functions such as changing / re-answering, error feedback, knowledge sharing, and satisfaction evaluation). The basic capability layer is responsible for supporting backend management, including log management of question and operation records (saving question information, reply information, user information, etc.), A / B testing (supporting different model versions to serve different user groups, checking the configuration during execution and matching the model or using the default model), and data analysis (supporting statistical analysis of question hit rate, asset hit rate, user satisfaction, and model version comparison, displaying training trends in line graphs). The AI ​​capability layer integrates core intelligent capabilities such as recommendation models, intent recognition (including rule engines, semantic matching, and deep learning model orchestration, based on BERT-like small model training, dynamically adjusting hyperparameters), knowledge question answering (based on knowledge base construction and processing, RAG configuration and Prompt training, using a hybrid retrieval method), data usage scenario analysis (including SQL table name recognition and SQL analysis capabilities, guiding large models to output table lists, service scenario descriptions, and application scenario information through prompt word engineering), and performance optimization strategies (large model call compression, thought process pruning, and asynchronous streaming return). Through layered collaboration, it enables intelligent analysis and interactive services for data assets.

[0025] As can be seen from the above embodiments, this application identifies the intent categories of queries, inquiries, explorations, and searches through user question-and-answer information or user search information, and predicts knowledge-based question-and-answer, service recommendations, and search needs. This automatically identifies data table names to generate a list of data usage tables. Based on the text names, basic descriptions, and field code values ​​in the list of data usage tables, semantic analysis is performed on the data resources to obtain data resource analysis results. This transforms complex, technical scripting languages ​​and data asset metadata into natural language descriptions that are easy for business personnel to understand, significantly reducing the understanding threshold of data assets, improving the efficiency and accuracy of data usage scenario analysis, while reducing reliance on manual interpretation and the risk of errors. This significantly improves the intelligence level, search accuracy, and user understanding efficiency of data asset services, and lowers the threshold for users to use data.

[0026] Based on this, embodiments of this application provide a data intelligence service method under a hybrid model, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the data intelligence service method under the hybrid model of this application.

[0027] In this embodiment, the data intelligence service method under the hybrid model includes steps S10 to S40: Step S10: Obtain user question and answer information and user search information; It should be noted that the user Q&A information refers to the interactive content entered by the user in natural language on the system's front-end interface to express their intention to search for data or inquire about knowledge. It may include one or more service keywords, data asset names, specific question descriptions, etc., such as "I want to find the transaction data of retail customers for the past month" or "Please explain the service meaning of the customer information table", allowing users to express complex or vague data needs in a conversational manner.

[0028] It is understood that the user retrieval information refers to the instruction information generated by the user in the search box of a traditional data asset retrieval page, such as the bank's unified data service portal, by entering keywords and selecting filter conditions, in order to precisely limit the search scope. For example, if a user enters the keyword "retail customers" and selects "high timeliness" and "premium assets" as filter conditions, this combination constitutes the user retrieval information, which follows a specific input format or interaction specification and represents the user's clear need to accurately search for data assets.

[0029] In a specific embodiment, as an optional implementation, in response to user interaction at the front-end application layer, the system can receive and record the user's question-and-answer information or the user's search information. For example, when a user submits a piece of text in the dialogue input box of the intelligent question-and-answer module, the system captures the text content through front-end listening events and stores and processes it as user question-and-answer information. For traditional search scenarios, when a user enters keywords in the search box on the search page and clicks the "search" button, or selects or modifies filter options, the system front-end collects the text keywords in the search box and all selected filter conditions (such as "asset grade = premium" and "update cycle = daily update"), and integrates and encapsulates this data into user search information.

[0030] In a specific embodiment, as another optional implementation, the system can also automatically analyze user behavior to generate corresponding information. For example, when a user does not input any filtering conditions but only enters a complete natural language sentence in the search box, such as "I want to analyze the risk factors of credit card delinquency", the system determines that the behavior belongs to "natural language input without filtering conditions". At this time, the input content can be passed to the intent recognition module as user question and answer information, and can also be used as part of the user's search information to enhance the search results under complex semantics.

[0031] In a specific embodiment, the user question and answer information and the user search information will be temporarily stored in the session context, carrying user identification (such as employee ID, affiliated organization) and timestamp, for intent recognition, logging and A / B test routing. If the user does not provide any input information (such as submitting a blank dialogue or searching directly without entering keywords), the system will not obtain valid information and may return corresponding prompt information to guide the user to re-enter.

[0032] Step S20: Based on the user question and answer information or the user search information, identify at least one of the corresponding query intent category, inquiry intent category, exploration behavior category and search behavior category, and determine at least one of the query intent category information, inquiry intent information, exploration behavior information and search behavior category information; It should be noted that the query intent category information refers to the intent tags identified by the system when users actively search for specific data assets through question-and-answer formats. This indicates that the user hopes the system will find data resources that match the description. For example, if a user asks, "Please help me find the customer deposit details table for the past three months," the system determines their intent as "finding data," and the corresponding query intent category information is "finding data." The inquiry intent category information refers to the intent tags identified by the system when users ask knowledge-based questions about the metadata content of existing data assets. This indicates that the user hopes the system will explain or provide knowledge related to the data assets. For example, if a user asks, "What does the cust_status field in the retail customer information table mean?", the system determines their intent as "asking about data," and the corresponding inquiry intent category information is "asking about data." The exploration behavior category information refers to the intent tags identified by the system when users explore traditional data resources... The product search page includes category tags for non-keyword search behaviors such as filtering, browsing, and sorting the asset list through selection criteria. These tags represent users narrowing down the scope of data assets through fuzzy exploration or conditional limitation. For example, if a user selects "Data Subject Domain = Customer Information", "Asset Grade = Premium", and "Update Cycle = Daily Update", the system determines that their behavior is exploratory, and the corresponding exploratory behavior category is "Filtering Exploration". The search behavior category information is the category tag identified by the system for users entering keywords and performing precise matching searches on the traditional data asset search page. This represents users using explicit keywords for string matching or simple semantic matching. For example, if a user enters "Retail Customer Basic Information Table" in the search box and clicks search, the system determines that their behavior is a search behavior, and the corresponding search behavior category is "Keyword Search".

[0033] In a specific embodiment, user question history information is obtained, that is, the system obtains user question history information from the log database. This record contains the original question text of the user's past questions and manually labeled intent tags. These records are periodically classified and organized into three categories of question intents: data finding, data asking, and functional module. A dataset containing user questions and intent tags is constructed and divided into training set, validation set, and test set in a ratio of 80%:10%:10% (each text is classified with a single label, for example, it can only belong to the category of data finding or data asking).

[0034] Based on the user question and intent label corresponding to the user question history information, the hyperparameters of the predefined intent recognition model, including the number of training batch samples, learning rate, number of iterations, random seed, and weight pruning, are adjusted to determine the target intent recognition model. That is, a BERT-based deep learning model is created based on the user question history information. By tracking the intent label recognition accuracy in real time, the hyperparameters are dynamically adjusted, including the number of training batch samples (the number of samples processed in each update), learning rate (parameter update step size), maximum number of iterations (the number of times the dataset is traversed), random seed (to ensure repeatability), and weight pruning (to penalize large weights to prevent overfitting), thereby training the target intent recognition model.

[0035] Based on the user's question and answer information, the target intent recognition model is input to identify at least one of the query intent category and inquiry intent category under the corresponding model orchestration. This determines at least one of the query intent category information and inquiry intent category information. Specifically, after obtaining user question and answer information, it is input into the target intent recognition model and identified according to a predefined model orchestration process. A rule engine is used to match based on regular expressions and keywords, such as matching "find" and "find data" for the "find data" category, "open" and "enter" for the "function module" category, and "lineage" and "code value" for the "jump" category. If the match is successful, the intent category is directly output; if it fails, semantic matching is performed. Semantic analysis is conducted on the user's question and the dataset based on text similarity technology, with a hit threshold of 0.95. If a result greater than the threshold exists, the recognition is successful; if it still fails, a trained deep learning model is called for recognition, with a threshold of 0.5, thereby determining the query intent category (find data category) and / or the inquiry intent category (ask data category), and outputting the corresponding query intent category information and inquiry intent category information.

[0036] Based on the filtering conditions and keywords in the user's search information, at least one of the corresponding exploration behavior category and search behavior category is identified. Specifically, after obtaining the user's search information, the filtering conditions and keywords are parsed. If it is detected that the user has selected at least one filtering condition (such as asset level, update cycle, topic domain, etc.) through a dropdown list, checkbox, or other control and has not entered any keywords in the search box, then the current behavior is determined to belong to the exploration behavior category. The exploration behavior category information for "Filter Exploration Category" is generated, and all selected filters are recorded. For conditional key-value pairs, if it is detected that the user has entered non-empty keyword text in the search box and triggered a search action (such as clicking the search button or pressing the Enter key), the current behavior is determined to belong to the retrieval behavior category, and retrieval behavior category information of "keyword retrieval category" is generated. The original keyword in the search box is extracted as the retrieval condition. If the user uses both filter conditions and keyword search at the same time, the exploration behavior category information and retrieval behavior category information are identified in parallel and combined as input for demand prediction. If the user neither enters keywords nor selects any filter conditions, no behavior category information is generated and the default asset list or prompt information is returned.

[0037] In one feasible implementation, step S20 may include steps A11 to A14: Step A11: Obtain the user's problem history information; It should be noted that the user question history information is user historical interaction data collected from the system log database or session storage. Each record includes at least the original question text asked by the user and question intent tags marked manually or by the system.

[0038] Step A12: Based on the user questions and intent tags corresponding to the user question history information, adjust the training batch sample size, learning rate, iteration rounds, random seed and weight pruning in the hyperparameters of the predefined intent recognition model to determine the target intent recognition model; It should be noted that the target intent recognition model is a deep learning model trained on a BERT-like architecture that can accurately distinguish user intent categories. The BERT-like architecture is a pre-trained language model architecture based on Transformer encoders, represented by BERT (Bidirectional Encoder Representations from Transformers). It adopts an Encoder-Only architecture, completely abandoning the decoder part and focusing on text understanding rather than text generation tasks.

[0039] Understandably, hyperparameters can include the number of training batch samples, learning rate, number of iterations, random seed, weight clipping, user question, and intent label. The number of training batch samples is the number of samples processed simultaneously during each model parameter update. The learning rate controls the step size of parameter updates. The number of iterations represents the number of times the model completely traverses the training dataset. The random seed is used to ensure the repeatability of random operations during model initialization and training. Weight clipping prevents overfitting by penalizing excessively large weight values. The user question and intent label correspond to the text input in the history and its corresponding category label, respectively.

[0040] Additionally, it should be noted that the intent recognition model is an initial untrained BERT-like model structure. Through dynamic adjustment of hyperparameters and supervised learning of training data, it can converge into a deployable target intent recognition model. The BERT-like architecture is composed of multiple layers of Transformer encoders stacked together. Each encoder layer contains a multi-head self-attention mechanism and a feedforward neural network, supplemented by residual connections and layer normalization, thereby achieving true bidirectional context modeling.

[0041] Step A13: Based on the user question and answer information, input the target intent recognition model to identify at least one of the query intent category and inquiry intent category under the corresponding model arrangement, and determine at least one of the query intent category information and inquiry intent category information; Understandably, model orchestration is a logical process that combines rule engines, semantic matching, and deep learning models according to priority. For example, it can perform fast rule matching based on regular expressions and keywords. If the match is successful, it can be output directly. If it fails, semantic matching can be performed. If it still fails, the deep learning model can be called. Through multi-level orchestration, the user's query intent category (find data category) and / or inquiry intent category (ask data category) can be output. The corresponding category information is the identification result label and confidence level.

[0042] Step A14: Based on the filtering conditions and keywords in the user search information, identify at least one of the corresponding exploration behavior category and search behavior category, and determine at least one of the exploration behavior category information and search behavior category information.

[0043] Understandably, the filtering criteria information refers to the asset attribute constraints selected by the user on the traditional search page through controls such as drop-down boxes, checkboxes, and radio buttons, such as asset grade = premium, update cycle = daily updates, subject field = customer information, etc., while the keyword information is the text string that the user enters in the search box for precise matching.

[0044] Step S30: Based on at least one of the query intent category information, the inquiry intent category information, the exploration behavior category information, and the retrieval behavior category information, predict the corresponding service demand to determine the service demand prediction information; Understandably, service needs refer to the goals users expect to achieve or the problems they need to solve when using data asset services. These needs can encompass multiple dimensions such as data retrieval, data understanding, and data analysis, including analytical needs, recommendation needs, knowledge-based Q&A needs, and retrieval needs, all aimed at interpreting assets. Analytical and recommendation needs refer to users' desire for the system to proactively push data assets relevant to their service scenario and potentially aligned with their underlying intentions. For example, when a user states, "I want to analyze the purchasing behavior of retail customers," the implicit recommendation need is for the system to recommend data tables or datasets related to retail customers and their purchasing behavior. Knowledge-based Q&A needs refer to users' desire for the system to explain and answer questions regarding the concepts, field meanings, and service rules of specific data assets. For example, a user might ask, "Does the update_time field in this table represent the data update time or the service occurrence time?" Retrieval needs refer to users' desire for the system to return a list of data assets that precisely or partially match their desired results using specific keywords or combinations of conditions. For example, a user might search for "customer information table" and expect to obtain all assets whose names contain that keyword.

[0045] In a specific embodiment, based on the query intent category information or the exploration behavior category information, a predefined recommendation model is input to predict the recommendation demand in the corresponding service requirements, thereby determining the service recommendation demand information. That is, the system can select the corresponding prediction path according to the intent category information, which significantly improves the accuracy of recommendation demand prediction in complex semantic scenarios. For example, if the query intent category information is "data search" or the exploration behavior category information is "exploration", the system determines that the user currently has a recommendation demand. The query intent category information or exploration behavior category information is input into the predefined recommendation model. The recommendation model outputs service recommendation demand information as service demand prediction information based on the user's question content, data asset profile, and multi-path recall strategy. For example, if the user's query intent is "data search" or the exploration behavior is "exploration", and the question is "I want to find credit card transaction records", the system predicts the recommendation demand as "recommend data assets related to credit card transaction records" and generates recommendation demand information containing a list of candidate assets.

[0046] Based on the query intent category information, the system inputs a predefined knowledge-based question-and-answer model to predict the corresponding knowledge-based question-and-answer requirements in the service needs, thus determining the knowledge-based question-and-answer requirement information. In other words, the system can select the corresponding prediction path based on the query intent category information, significantly improving the accuracy of knowledge-based question-and-answer requirement prediction in complex semantic scenarios. For example, if the query intent category information is "question-number type," the system determines that the user currently has a knowledge-based question-and-answer requirement. It inputs the query intent category information into the predefined knowledge-based question-and-answer model. The knowledge-based question-and-answer model, based on RAG (Retrieval Augmented Generation) technology, retrieves relevant fragments from the knowledge base and generates answers through a large model, thereby outputting knowledge-based question-and-answer requirement information as service demand prediction information. For example, if a user asks, "What is the code value of the gender field in the customer information table?", the system predicts that their knowledge-based question-and-answer requirement is "It is necessary to explain the code value mapping relationship of the gender field," and generates knowledge-based question-and-answer requirement information containing the answer content.

[0047] Based on the retrieval behavior category information, the system predicts the retrieval needs in the corresponding service requirements and determines the retrieval need information. That is, the system can select the corresponding prediction path according to the retrieval behavior category information, which significantly improves the accuracy of retrieval need prediction in complex semantic scenarios. For example, if the retrieval behavior category information is "keyword retrieval" and no query intent or inquiry intent category information is output at the same time, the system determines that the user currently has a retrieval need. At this time, the system does not call the intelligent model, but directly executes the traditional keyword matching retrieval logic based on the keywords and filtering conditions entered by the user, and encapsulates the retrieval need information as service need prediction information. For example, if the user enters "customer basic information table" in the traditional search box and clicks search, the system outputs retrieval need information, including a list of data assets that hit the keyword.

[0048] Service demand prediction information is obtained based on at least one of the service recommendation demand information, the knowledge question and answer demand information, and the retrieval demand information. This enables comprehensive coverage of various data service scenarios such as user data search, data inquiry, and precise retrieval, achieving intelligent fusion and accurate mapping from diverse intentions to structured demands, and significantly improving the completeness of demand prediction and service adaptability.

[0049] Step S40: Analyze the corresponding data resources based on the service demand prediction information and determine the data resource analysis results.

[0050] It should be noted that the data resource analysis results are actionable service analysis results presented to users after gradually deepening the understanding of user intent through multiple rounds of dialogue, real-time feedback, and interactive guidance. This helps users quickly understand the business positioning, core value, applicable scenarios, and key characteristics of data assets, and allows users to further explore or confirm the analysis results through follow-up questions, selections, and feedback, thereby reducing the cost for users to understand data assets. The results may include asset business positioning information, usage scenario information, asset recommendation index, core field descriptions, data timeliness, and application status, and support streaming output, related question guidance, and clickable jumps.

[0051] It is understandable that data resources can be data assets, including database tables (such as user information tables and transaction log tables), views, datasets, data files, and associated metadata information (such as table structure, field definitions, data lineage, service scope, etc.). In other words, these are the core data resources that users focus on, search for, or request to analyze when using a data asset service system. For example, if a user requests to analyze a "snapshot table of basic information of retail customers," this table and its contained fields, code values, upstream and downstream dependencies all fall under the category of data assets.

[0052] In a specific embodiment, user metadata information and associated statement script information are obtained. That is, user metadata information and original SQL scripts can be obtained from the database user logs. The input SQL is standardized and cleaned by removing comments, unifying capitalization, standardizing spaces and newlines, and generating standardized SQL text, i.e., associated statement script information.

[0053] Based on the user metadata information and the associated statement script information, the predefined analysis model identifies the corresponding data table names and determines the user table list information. This involves inputting standardized SQL text into the predefined analysis model, which then guides the model to output the user table list information (including the English and Chinese names of the tables and service descriptions) through specific prompts. The specific prompts are designed as follows: Task: As a database expert, identify all tables from the SQL statements in the ${sql_type} database and output the table names.

[0054] SQL statement: ${input}.

[0055] Require: 1. Tables generally appear near the from and join keywords.

[0056] 2. Remove the table name prefix.

[0057] 3. Do not output the recognition process.

[0058] 4. Table names should be deduplicated.

[0059] 5. Separate table names with newline characters.

[0060] Here, the ${sql_type} variable refers to the database type, such as GAUSS, MySQL, etc.; the ${input} variable refers to the content of the standardized SQL statement to be processed.

[0061] Based on the text name, basic description, and field code values ​​in the data list information, the data resources corresponding to the service demand prediction information are analyzed to obtain data resource analysis results. The data resource analysis results include service scenario description information and application scenario information. That is, the original SQL script and related data asset metadata knowledge (such as Chinese text name, basic description, field code values, etc. in the table) are used as input to construct a multi-dimensional input information set. The large model is guided by structured prompt word templates to analyze and summarize the SQL to obtain data resource analysis results containing service scenario description information and application scenario information.

[0062] The multi-dimensional input information set includes: 1. Original SQL statement: a structured query statement to be interpreted; 2. Data asset metadata knowledge base: containing table-level information (such as Chinese table name, service description, subject domain, and data source) and field-level information (such as Chinese field name, data type, service meaning, code value mapping table, value range, and whether it is a primary key / foreign key); 3. Contextual scenario information: SQL interpretation examples (such as this SQL performs a left join between a snapshot of household registration information on a specific date and a snapshot of the core institution to extract customer data and its affiliated branch information), application scenario examples (such as this temporary table can be used to further analyze the customer's credit limit usage, customer identity information, etc.), and output language style (formal / colloquial / concise / detailed).

[0063] The structured prompt template includes: 1. Task Definition: Assign the role of a database expert and specify the task of interpreting SQL and analyzing application scenarios, requiring concise results. The variable `${sql_type}` refers to the database type, supporting mainstream databases such as GAUSS and MySQL. 2. Input Information: Set variables to pass in the required input information, including the SQL statement and table structure (table name, field name, field code value, and Chinese explanation). The variable `${input}` refers to the SQL statement, `${table_chn}` refers to the table name, `${table_column}` refers to the field name, and `${column_code_values}` refers to the field code value and its Chinese explanation. 3. Requirements and Output Specifications: Specify the output format (e.g., output in JSON format: `{"explanation": service scenario description, "next": application scenario}`), required examples, and handling of dates / constants.

[0064] It should be understood that the structured prompt word template can be represented as: Task: As a database expert, your task is to provide a Chinese explanation of the SQL statement input by the user for the ${sql_type} database. The explanation should be as accurate, concise, and coherent as possible. Please think through the process step by step and provide your thought process; finally, summarize the explanation, extracting the core content, and keep the explanation concise; finally, try to analyze and output the application scenario of this SQL statement.

[0065] SQL statement: ${input}.

[0066] Table structure: ${table_chn}; ${table_column}.

[0067] The code values ​​for the table fields are: ${column_code_values}.

[0068] SQL explanation example: 1. This SQL statement counts the number of customers under the age of 18 and their transaction amounts in different financial product categories (such as wealth management, insurance, funds, etc.), including sub-categories such as private / public wealth management and structured deposits.

[0069] 2. This SQL statement performs a left join between a snapshot of household registration information on a specific date and a snapshot of the core institution to extract customer data and their affiliated branch information.

[0070] Application scenario examples: 1. This temporary table can be used for further analysis of customer credit limit usage, customer identity information, etc.

[0071] 2. This table can be used to further analyze customer purchasing preferences or service usage habits.

[0072] 3. The service department can use this table to analyze the distribution of customers at different time periods and evaluate the effectiveness of marketing activities or customer acquisition strategies.

[0073] Require: 1. Output the explanation content in JSON format, with the format: {"explanation": service scenario description, "next": application scenario}.

[0074] 2. If there are Chinese names for table fields, do not use the English names of the fields in the explanation. Alternatively, you can analyze the Chinese names of the table columns from the context of the filter conditions.

[0075] 3. To imitate the SQL interpretation example, express the complete meaning in one sentence, and also reflect the key information in the Chinese names of the fields.

[0076] 4. Thinking process: When you encounter screening or related conditions, please describe their specific meaning or function.

[0077] 5. When encountering a constant, try to identify whether it is a code value.

[0078] 6. Application scenarios: You can refer to the format of the application scenario examples.

[0079] 7. When encountering dates, please use "specific date" instead.

[0080] By automating SQL preprocessing, identifying large model table names, and integrating metadata with rich examples for semantic analysis, complex technical SQL statements can be quickly transformed into natural language descriptions that service personnel can understand. This significantly reduces the threshold for understanding data assets and improves the efficiency and accuracy of data usage scenario analysis.

[0081] It should be understood that, for user-initiated data asset analysis requests, the system can use standardized question templates to guide users to clarify the scope of analysis. For example, in the intelligent interpretation module, a user can input or select "Help me analyze the [Logical Description] [Usage Scenarios] [Recommendation Index] of the [Retail Customer Basic Information Table]". After parsing the request, the system extracts the logical description information of the table from the metadata database, obtains typical usage scenarios based on historical SQL parsing from the usage scenario analysis module, and obtains the star-rating recommendation index generated based on rules such as asset level, timeliness, and application status from the asset recommendation index calculation module. These contents are then integrated into data resource analysis results and output in streaming format, sequentially displaying the logical description, usage scenarios, recommendation index, and other key metadata information (such as whether it is a high-timeliness table or whether it is available for application).

[0082] In a specific embodiment, as an application example, a fixed question format of "Help me interpret [Please select asset name] [Logical description] [Usage scenario] [Recommendation index]" can be used to guide users to ask questions about interpreting data assets in a simple, fast and efficient way. Users can directly input the data asset name and then select the content to be interpreted to quickly get the answer. The reply content adopts a streaming content output, which sequentially displays the data asset logical description information, usage scenario information, asset recommendation index and other key metadata information, such as whether it is a high timeliness table and whether it can be applied for. It also provides highly relevant questions such as "You can continue to ask" to guide users to further dialogue. Among them, the asset recommendation index is displayed with the effect of star count. The calculation rule logic is as follows: (1) Judging from the asset level, premium / high-efficiency assets = 1 star, other low-level assets = 0 stars (2) Judging from the timeliness, high timeliness = 1 star, medium timeliness = 0.5 stars, low timeliness = 0 stars (3) Judging from whether it can be applied for, recommended application = 1 star, not applicable = 0 stars.

[0083] In specific implementations, as functionality becomes more complex, the overall process time increases, impacting user experience. Several performance optimization strategies can be implemented: First, large model call compression, simplifying multi-round model interactions into a single call. For example, the two tasks of SQL natural language interpretation and application scenario summarization can be merged into a single call through deep integration and optimization of the Prompt project, leveraging the large model's long contextual understanding capabilities to autonomously complete multi-task reasoning. Second, thought process pruning, automatically shutting down non-critical reasoning stages such as syntactic analysis when the input text does not contain complex grammatical structures, resulting in only a loss of approximately 1% accuracy but a 200% speed increase, directly breaking down keyword searches. Third, asynchronous streaming return, prioritizing the return of basic results (such as simple query results, summaries, or default data), while the background continues to execute reasoning for complex tasks and incrementally updates the interface content, thereby meeting users' expectations for immediate interaction and significantly improving system response speed and user experience.

[0084] This embodiment proposes a data intelligence service method under a hybrid model, which involves acquiring user question-and-answer information and user search information; identifying at least one of the corresponding query intent category, inquiry intent category, exploration behavior category, and search behavior category based on the user question-and-answer information or the user search information, and determining at least one of the query intent category information, inquiry intent category information, exploration behavior information, and search behavior category information; predicting the corresponding service demand based on at least one of the query intent category information, inquiry intent information, exploration behavior information, and search behavior category information, and determining service demand prediction information; and analyzing the corresponding data assets based on the service demand prediction information to determine the data resource analysis results. This application addresses the technical challenge of accurately and effectively serving and analyzing data assets. Compared to existing technologies, it identifies the intent categories of queries, inquiries, explorations, and searches using user question-and-answer information or user search information, and predicts knowledge-based Q&A, service recommendations, and search needs. This automatically identifies data table names to generate a list of data usage tables. Based on the text names, basic descriptions, and field code values ​​in the list, semantic analysis is performed on the data resources to obtain data resource analysis results. This transforms complex, technical scripting languages ​​and data asset metadata into natural language descriptions easily understood by business personnel, significantly lowering the barrier to understanding data assets, improving the efficiency and accuracy of data usage scenario analysis, while reducing reliance on and error risks associated with manual interpretation. This significantly enhances the intelligence level, search accuracy, and user understanding efficiency of data asset services, and lowers the barrier to data usage for users.

[0085] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the first embodiment described above can be referred to the above description, and will not be repeated hereafter.

[0086] In this embodiment, refer to Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the data intelligence service method under the hybrid model of this application. Step S30 specifically includes steps S31 to S34: Step S31: Based on the query intent category information or the exploration behavior category information, input the predefined recommendation model to predict the recommendation requirements in the corresponding service requirements and determine the service recommendation requirement information; It should be noted that the service recommendation demand information is a predicted list of candidate data assets that users may need, the relevance score or ranking of each candidate asset, or the reason for recommendation. For example, the service recommendation demand information can be represented as: Recommended Asset List = [{Asset Name: Retail Customer Basic Information Table, Score: 0.95, Reason: This table contains customer identity and contact information}, {Asset Name: Retail Customer Transaction Record Table, Score: 0.87, Reason: This table records customer transaction behavior}].

[0087] Understandably, the recommendation model is a combination of algorithmic models deployed in the AI ​​capability layer to achieve intelligent retrieval and recommendation of data assets. It is a composite model system composed of multiple sub-modules, such as data collection and processing, keyword extraction and association, multi-path recall (vector retrieval, full-text retrieval, graph retrieval), and large-scale model ranking. It is used to convert users' fuzzy, colloquial, and multi-keyword complex natural language queries into accurate data asset matching results and actively push them to users. It supports two input scenarios: one is to receive query intent category information from the intent recognition module (data search category), and the other is to receive user exploration behavior category information captured by the front-end application layer (such as natural language search without filtering conditions).

[0088] In a specific embodiment, data service metadata information is acquired, including basic table information, field information, and upstream / downstream lineage information. Based on the data service metadata information and the query intent category information or the exploration behavior category information, a predefined recommendation model is input to predict the corresponding data asset profile and extended keywords, determining the data profile information and extended keyword information. That is, the system collects and organizes the metadata information of existing data assets, including basic table information (such as English table name, Chinese table name, data directory, logical description, table form, slice data volume, asset level such as premium / high-efficiency / medium-efficiency, etc.), field information (English field name, Chinese field name, field type, code value, etc.), and upstream / downstream lineage information (upstream and downstream related data assets and their metadata information). This data service metadata information, along with the identified query intent category information (data search type) or exploration behavior category information (filtering exploration type), is input into the predefined recommendation model, utilizing the document organization capabilities of the recommendation model. The system leverages its capabilities to highly summarize metadata and uncover hidden connections and insights, generating data asset profiles that include service positioning, core value, and upstream and downstream relationships (specifically, a profile like "Table [Table Name], Chinese Name is [Chinese Name], contains [Field Information], comes from [Upstream System], is used in [Downstream Scenario], daily data volume X, has X downstream users"). Simultaneously, the model extracts keywords from user queries (such as service names, product names, table names, field names, etc.) using a large-scale model and supplements user colloquial expressions with the model's associative capabilities and common sense (e.g., adding the keyword "debit card" to "savings card"), thus obtaining expanded keyword information. By integrating multi-dimensional metadata and user intent, the generated asset profiles more comprehensively represent the service semantics and lineage of data assets, significantly improving the interpretability of recommendation results. Expanded keywords effectively avoid misjudgments caused by vague user descriptions or mismatched technical terms, thereby greatly improving the accuracy and recall of data asset recommendations in complex semantic scenarios.

[0089] Based on the extended keyword information, a predefined recommendation model is used to perform vector retrieval on the corresponding converted text vectors to determine the vector retrieval results. Based on the vector retrieval results, the corresponding keywords are used for full-text retrieval using a predefined inverted index to determine the full-text retrieval results under the multi-path recall strategy. A graph mining algorithm is then used to perform graph retrieval between the full-text retrieval results and a predefined knowledge graph to determine the recall data service information under the corresponding multi-path recall strategy. This allows the multi-path recall strategy to be initiated based on the extended keyword information. The extended keywords are converted into text vectors, and vector retrieval is performed in the asset profile database by calculating vector similarity (e.g., Euclidean distance) to capture semantic matching of synonyms, near-synonyms, and implicit intents, resulting in vector retrieval results. Based on the keywords in the vector retrieval results, a full-text retrieval is performed using a pre-built inverted index to quickly and accurately locate structured metadata documents (such as table names, field descriptions, etc.) containing the specified keywords. The full-text search results are obtained. Based on the knowledge graph constructed from the lineage of the data tables within the row, the service importance of the tables in the full-text search results is quantified using graph mining algorithms such as personalized PageRank. Data assets that are frequently used or located in the core lineage path are prioritized for recall. This determines the final recalled data service information under the multi-path recall strategy. It should be understood that vector retrieval makes up for the lack of semantic understanding in traditional full-text retrieval, full-text retrieval ensures the speed and accuracy of precise matching of metadata, and graph retrieval uses lineage relationships to mine core assets and improve the relevance of recommendations to service scenarios. Therefore, the multi-path recall strategy essentially ensures both retrieval speed and accuracy, and achieves term alignment and semantic mapping. Even if the key fields of the question are not in the table name fields, relevant data assets can still be found. At the same time, user feedback is collected during use and the comprehensive scoring strategy is dynamically optimized, making asset recommendations increasingly accurate and significantly improving the accuracy in complex semantic scenarios.

[0090] Based on the data profile information and the recall service profile corresponding to the recall data service information, the recommendation requirements in the corresponding service needs are analyzed to determine the service recommendation requirement information. That is, taking the asset profiles of each data table in the data profile information and the recall data service information as input, the user's natural language intent and the profile information such as the service positioning, core value, field meaning, lineage, and usage scenarios of each asset are input into the big model. The big model, based on deep semantic analysis, cross-domain association matching, and dynamic context awareness capabilities, performs unified scoring and fine ranking of the multi-way recall results, and selects the multiple data tables that best match the user's question as the recommendation results. Clear and explainable recommendation reasons are generated for each table. Through the big model's accurate matching of the user's deep needs and the semantics of asset services, the relevance and hit rate of the recommendation results can be significantly improved.

[0091] In one feasible implementation, step S31 may include steps B11 to B14: Step B11: Obtain data service metadata information, where the data service metadata information includes table basic information, field information, and upstream and downstream lineage relationship information; It should be noted that the data service metadata information is a data set representing the attributes of data assets themselves and their service semantics, including static definitions of data tables, field-level details, and dependencies between data tables.

[0092] It can be understood that the table basic information is metadata for the overall attributes and service positioning of data assets, including the English name of the table, Chinese name, data directory (subject domain to which it belongs), logical description (service meaning), table form (such as entity table, view, temporary table), sliced data volume (such as daily increment or full data volume level), and asset level (such as high-quality, efficient, medium-efficient, low-efficient), etc., which is used to quickly understand the identity, service background, and quality level of the data table. The field information is a detailed attribute description of each column (field) in the data table, including the English name of the field, Chinese name (service meaning), field type (such as string, integer, date), code value (enumeration value and its corresponding Chinese explanation, such as "0 - male, 1 - female"), and value range, whether it is a primary key / foreign key, etc., which is used to understand the specific data content and service semantics inside the data table. The upstream and downstream lineage relationship information is the dependency and flow relationship between data assets, that is, the data source of the current table (upstream source table, source system) and which downstream tables or applications use the data of the current table (downstream dependent tables, usage scenarios). Through this lineage chain, the generation, processing, and consumption paths of data can be traced, and it can be used to evaluate the core degree and influence range of data assets.

[0093] Step B12: Based on the data service metadata information and the query intent category information or the exploration behavior category information, input them into a predefined recommendation model to predict the corresponding data asset portraits and extended keywords, and determine the data portrait information and extended keyword information; It should be noted that the data portrait information is a service semantic description generated by the large model after highly summarizing the metadata, including service positioning, core value, upstream and downstream associations, etc. For example, "Table customer_info, Chinese name is customer information table, contains customer ID, name, and level fields, is used in the precise marketing scenario, daily data volume is 100,000, and there are 3 downstream tables". The extended keyword information is a set of keywords extracted and supplemented by the large model from the user input. For example, supplementing "储蓄卡" to "借记卡". It can be understood that by integrating metadata and user intent, the recommendation model can automatically generate an interpretable asset portrait, expand the user's fuzzy or colloquial expressions, and provide richer retrieval entries.

[0094] Step B13: Determine the recalled data service information under the corresponding multi-way recall strategy based on the extended keyword information; It should be noted that the recalled data service information is a list of candidate data assets and their service profiles that are initially screened from the data asset library and are related to user needs through a multi-path recall strategy.

[0095] Understandably, the multi-path recall strategy is a hybrid retrieval mechanism that combines vector retrieval, full-text retrieval, and graph retrieval. Vector retrieval uses an embedding model to convert text into vectors and captures semantically similar unstructured asset profiles through similarity calculation. Full-text retrieval uses an inverted index to quickly match keywords in structured metadata. Graph retrieval uses a knowledge graph built from data table lineages and uses graph mining algorithms to quantify the coreness of the table, prioritizing the recall of assets that are in key lineage paths or are frequently used.

[0096] In one feasible implementation, step B13 may include steps C11-C13: Step C11: Based on the extended keyword information, input the predefined recommendation model to perform vector retrieval on the corresponding converted text vector, and determine the vector retrieval result; It should be noted that the vector retrieval result is a list of candidate segments or assets that are semantically similar to the user's intent, recalled from the asset profile. For example, the extended keyword information is mapped to a vector representation, an approximate nearest neighbor search is performed in the pre-generated asset profile vector library, the cosine similarity or Euclidean distance is calculated, and the result with similarity exceeding a set threshold is output as the vector retrieval result. This can capture synonyms, near-synonyms and implicit intent, and make up for the semantic blind spots of keyword matching.

[0097] Step C12: Based on the vector retrieval results, perform full-text retrieval of the corresponding keywords using a predefined inverted index to determine the full-text retrieval results under the multi-path recall strategy; It should be noted that the full-text search results are a collection of documents returned after precise keyword matching. For example, the system extracts key feature words from the vector search results, combines them with the original extended keywords, and quickly locates the metadata records containing these words in the pre-built inverted index. The full-text search results are then sorted according to the matching degree, thereby ensuring the speed and accuracy of the search.

[0098] Step C13: The full-text search results are compared with a predefined knowledge graph using a graph mining algorithm to determine the recall data service information under the corresponding multi-path recall strategy.

[0099] Understandably, graph mining algorithms are used to analyze the importance of nodes, the strength of associations, and path characteristics in graph-structured data. For example, a personalized PageRank algorithm can be used, starting from seed nodes related to users and iteratively propagating weights in the graph to quantify the importance of each data asset node relative to user needs. Knowledge graphs are graph databases built on the upstream and downstream lineage relationships between data tables, where nodes represent data assets (tables) and edges represent lineage dependencies (such as "table A -> table B" indicating that A is the upstream of B).

[0100] Step B14: Based on the data profile information and the recall service profile corresponding to the recall data service information, analyze the recommended requirements in the corresponding service requirements to determine the service recommendation requirement information.

[0101] It is understandable that the service profile is a set of service semantic descriptions corresponding to each candidate data asset obtained through a multi-path recall strategy. This includes the service positioning, core value, upstream and downstream dependencies, meaning of field code values, and scene tags mined from the lineage graph of the asset. It can deeply match user needs from the service semantic level and significantly improve the accuracy and interpretability of recommendation results.

[0102] Step S32: Based on the query intent category information, input a predefined knowledge question answering model to predict the knowledge question answering requirements in the corresponding service requirements and determine the knowledge question answering requirement information; It should be noted that the knowledge Q&A requirement information is used to answer users' questions about data asset metadata. This includes direct answers to user questions, knowledge source fragments cited in the answers, and the confidence level or related explanations of the answers. It is used to provide users with accurate and easy-to-understand explanations of data asset-related concepts (such as table meaning, field definition, code value mapping, service scope, lineage, etc.), thereby lowering the threshold for users to understand data assets.

[0103] In a specific embodiment, a retrieval is performed based on the input of the query intent category information into a predefined knowledge base to determine the knowledge base retrieval results. The knowledge base includes basic data table information, field information, code value information, graph information, and script information. That is, the system can build a knowledge base based on the user's query intent category information (question type) and provide data asset knowledge question answering capabilities based on the RAG framework. It stores basic data table information (such as the English name, Chinese name, update cycle, asset level, etc., split by row in Excel format), field information, code value information, graph information, and ETL / DDL script information.

[0104] Based on the knowledge base retrieval results, the documents in the knowledge base are sliced ​​to determine the sliced ​​text block information. That is, Word, PDF and other documents in the knowledge base retrieval results can be sliced ​​into a maximum length of 500 and punctuation marks can be processed (such as converting consecutive line breaks into single ones and replacing line breaks ending with question marks with question marks). For HTML documents, the tags are parsed first and then segmented.

[0105] A hybrid retrieval method combining predefined keyword retrieval and vector semantic retrieval is used to recall the segments corresponding to the sliced ​​text block information from the knowledge base. The recalled segment information is determined by RAG configuration. Vector representations are generated for the sliced ​​text blocks through an embedding model (such as bge-base-zh). A hybrid retrieval method combining keyword retrieval (inverted index) and vector semantic retrieval is used, with a maximum number of recalled paragraphs set to 4, a total number of recalled segments set to 5, and a threshold of 0.3, thereby recalling the most relevant sliced ​​text block information from the knowledge base.

[0106] Based on the recalled fragment information, a predefined knowledge-based question-answering model is input to predict the knowledge-based question-answering requirements in the corresponding service needs, thus obtaining the knowledge-based question-answering requirement information. This information can then be used to train a large model to generate answers using a prompt. The prompt words are: Background information: ${knowledge}.

[0107] Please answer user questions based on the background information and history.

[0108] When a user asks: ${userQuestion}, do not return phrases like "based on background information" or summary text, and do not display empty information such as "nan".

[0109] If the background information is [], or there is no background information, or the user's question is unrelated or only minimally related to the returned background information, then the following text content will be returned: I'm sorry, I cannot answer your question at the moment. I suggest you join our support group for assistance.

[0110] At this point, after cross-validation and A / B testing optimization, the retrieved fragments are integrated into concise, easy-to-understand, and professional answers to obtain knowledge-based Q&A requirements (such as field meanings, code value translations, lineage paths, etc.).

[0111] By constructing a multi-source knowledge base, employing hybrid retrieval, and fine-tuning the prompt, the system balances precise matching with semantic generalization capabilities. Even if the terminology used in the query does not match the descriptions in the database, it can accurately retrieve relevant knowledge, significantly improving the accuracy of Q&A for data asset metadata and enhancing user experience, while lowering the threshold for data understanding.

[0112] In one feasible implementation, step S32 may include steps D11 to D14: Step D11: Based on the query intent category information, input a predefined knowledge base for retrieval and determine the knowledge base retrieval results. The knowledge base includes basic data table information, field information, code value information, graph information, and script information. It should be noted that the knowledge base retrieval result is the original set of all documents or data entries that may be related to the user's question, returned by the system after preliminary matching in the pre-built knowledge base according to the user's query intent. This result has not yet been processed by slicing and reordering.

[0113] Understandably, basic information in a data table can include the English and Chinese names of assets, update cycles, asset grades, and technical personnel. Field information can include the English and Chinese names of fields and their data types. Code value information includes field enumeration values ​​and their corresponding Chinese explanations. Graph information can be the lineage relationships (upstream and downstream dependencies) between data tables. Script information can be SQL script content such as ETL and DDL.

[0114] Step D12: Based on the knowledge base retrieval results, slice the documents in the knowledge base to determine the slice text block information; It should be noted that the sliced ​​text block information refers to multiple semantically reasonable and appropriately sized text segments obtained by segmenting the original document retrieved from the knowledge base according to certain rules (such as file type, maximum length, and semantic boundaries). Each segment independently contains an understandable context.

[0115] Understandably, slicing is an intelligent segmentation operation for documents of different formats: for Excel spreadsheets, it is split by row to preserve the hierarchical structure; for Word, PPT, PDF, HTML, Markdown, and other files, it is split according to a maximum length of 500 characters, while punctuation is standardized (e.g., consecutive line breaks are converted to single line breaks, and line breaks after a question mark are replaced with a question mark); for HTML documents, tags are parsed first and then segmented. The purpose of slicing is to break down long documents into smaller units that are easier to vectorize and retrieve, thereby improving the relevance and efficiency of the retrieved data.

[0116] Step D13: Using a hybrid retrieval method combining predefined keyword retrieval and vector semantic retrieval, retrieve the fragments corresponding to the sliced ​​text block information from the knowledge base, and determine the retrieved fragment information; It should be noted that the recalled fragment information is a set of several text fragments most relevant to the user's question, selected from all sliced ​​text blocks through a hybrid retrieval strategy, sorted by relevance score, and a maximum recall number and similarity threshold are set, such as a total of 5 recalled fragments and a threshold of 0.3.

[0117] Understandably, keyword retrieval, based on an inverted index, quickly locates text slices containing key terms from the user's question, ensuring accurate matching of structured terms. Vector semantic retrieval, on the other hand, transforms the user's question and text slices into vectors through an embedding model, calculates cosine similarity or Euclidean distance, and captures synonyms, near-synonyms, and implicit semantic relationships. Even if the original words in the database do not appear in the question, relevant fragments can still be recalled. After the two methods are executed in parallel, the results are deduplicated, merged, and reordered to output the recalled fragment information, thus balancing the accuracy and generalization ability of the retrieval.

[0118] Step D14: Based on the recalled fragment information, input a predefined knowledge question-answering model to predict the knowledge question-answering requirements in the corresponding service requirements, and obtain knowledge question-answering requirement information.

[0119] Understandably, a knowledge-based question-answering model is a composite question-answering model deployed at the AI ​​capability layer, consisting of multiple sub-modules such as knowledge base construction and processing, RAG configuration, prompt word training, and a large language model. After receiving the user's query intent category information, it can retrieve text fragments related to the user's question from the pre-built knowledge base, and then use the retrieved fragments as background information, combined with carefully designed prompt words, to guide the large language model to generate an accurate, concise, and easy-to-understand answer.

[0120] Step S33: Based on the search behavior category information, predict the search requirements in the corresponding service requirements to determine the search requirement information; It should be noted that the search requirement information refers to the query parameters and expected results description used to directly obtain the data asset list. It may include the original keywords entered by the user, the key-value pairs of the filtering conditions selected by the user, and the search type that the system expects to perform based on these conditions.

[0121] It is understood that the retrieval demand information is different from recommendation demand information and knowledge question answering demand information. It does not rely on intelligent models for semantic expansion or intent reasoning, but directly represents the retrieval constraints actively specified by the user, and can quickly and accurately return data assets that meet the explicit conditions.

[0122] In a specific embodiment, after obtaining the retrieval behavior category information, the system can directly use the corresponding keywords as retrieval conditions and call the traditional data asset retrieval engine (full-text retrieval based on inverted index) to perform precise matching and sorting in metadata (such as table names, field names, logical descriptions, etc.), and return a list of matched data assets as retrieval requirement information (including asset name, basic description, asset level, etc.). Therefore, for scenarios where users explicitly use keywords for precise searching, the system can avoid complex intent recognition and model inference overhead, and return retrieval results with extremely low latency, meeting the efficiency requirements of users to quickly locate known assets.

[0123] Step S34: Obtain service demand prediction information based on at least one of the service recommendation demand information, the knowledge question answering demand information, and the retrieval demand information.

[0124] It should be noted that the service demand prediction information is the result generated by the system after inferring and predicting the user's potential service demand based on the identified intent category information by calling the corresponding intelligent model or rule logic. It may include the predicted demand type (such as "recommendation demand", "knowledge Q&A demand", "retrieval demand") and the specific parameters under the demand (such as the recommended asset list, the Q&A reply content, and the retrieval results).

[0125] In specific embodiments, if a query intent or exploratory behavior is identified, service recommendation demand information is used as the main body of service demand prediction information; if an inquiry intent is identified, knowledge question-and-answer demand information is used as the main body of service demand prediction information; and if a retrieval behavior is identified, retrieval demand information is used as the main body of service demand prediction information. Therefore, when multiple intents coexist (such as when a user simultaneously inputs natural language and filtering conditions), the system can integrate service recommendation demand information with retrieval demand information, generating comprehensive service demand prediction information with recommendation results as the priority and retrieval results as a supplement. This information is then transmitted to the analysis module in a unified format, enabling flexible adaptation to diverse data service scenarios. This allows for intelligent diversion and integration of scenarios such as data retrieval, data inquiry, precise querying, data understanding, and accurate interpretation. It ensures accurate response under a single intent while balancing recommendation generalization ability and retrieval certainty under complex mixed intents, thereby improving the comprehensiveness of demand prediction and service adaptability.

[0126] This embodiment proposes a data intelligence service method under a hybrid model. Based on the query intent category information or the exploration behavior category information, a predefined recommendation model is input to predict the recommended service requirements, thus determining service recommendation requirement information. Based on the query intent category information, a predefined knowledge question-and-answer model is input to predict the knowledge question-and-answer requirements, thus determining knowledge question-and-answer requirement information. Based on the retrieval behavior category information, the retrieval requirements, thus determining retrieval requirement information. Service requirement prediction information is obtained based on at least one of the service recommendation requirement information, the knowledge question-and-answer information, and the retrieval requirement information. This application addresses the technical challenge of accurately and effectively providing data asset services and analysis. Compared to existing technologies, it identifies different intent categories and invokes corresponding models. Query or exploration intents trigger recommendation models to generate service recommendation requirements, inquiry intents trigger knowledge question-answering models to generate knowledge question-answering requirements, and retrieval intents directly execute keyword full-text retrieval to generate retrieval requirements. Furthermore, it can use at least one of the service recommendation requirements, knowledge question-answering requirements, and retrieval requirements as service requirement prediction information. This enables precise traffic distribution and intelligent integration across scenarios such as data retrieval, data inquiry, accurate querying, data understanding, and precise interpretation, while balancing semantic generalization capabilities with low-latency determinism, comprehensively covering users' diverse data service needs.

[0127] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the first embodiment described above can be referred to the above description, and will not be repeated hereafter.

[0128] In this embodiment, refer to Figure 4 , Figure 4 This is a flowchart illustrating the data intelligence service method embodiment three under the hybrid model of this application. Step S40 specifically includes steps S41 to S43: Step S41: Obtain user metadata information and associated statement script information; It should be noted that the user metadata information is metadata that describes the attributes of the data itself and the semantics of the service, including the table structure of the data assets, the meaning of the fields, the data type, the service description, the subject domain to which it belongs, the data source, the field code value mapping table, the value range, the primary and foreign key relationships, and the upstream and downstream lineage relationships, etc. The related statement script information is the raw SQL script collected from the database user logs, which records the data query, processing and handling logic executed by the user in the actual service analysis.

[0129] Step S42: Based on the user metadata information and the associated statement script information, input a predefined analysis model to identify the corresponding data table name and determine the data table list information; It should be noted that the data table list information is a collection of all referenced data tables identified by the system from the SQL script. Each record includes the English name of the table, the Chinese name of the table (literal name), and a basic service description of the table.

[0130] Understandably, the analysis model is an intelligent analysis module built on a large language model. It parses the standardized and cleaned SQL scripts through a specific prompt word project (clearly defining the role as a database expert, requiring identification of table names near the FROM / JOIN keywords, removal of table name prefixes, deduplication, and separation by newline characters). The data table name is the name of the physical table or view that appears in the SQL statement, located after keywords such as FROM, JOIN, and INTO. This allows it to automatically extract the table dependencies in complex SQL into a structured list, providing a clear scope of data assets for analysis.

[0131] Step S43: Analyze the data resources corresponding to the service demand prediction information based on the text name, basic description and field code value in the data list information to obtain data resource analysis results. The data resource analysis results include service scenario description information and application scenario information.

[0132] It is understood that the service scenario description information is a natural language explanation that can be understood by humans, which transforms the SQL query logic. For example, "This SQL performs a left join between the snapshot of household registration information on a specific date and the snapshot of the core institution to extract customer data and its affiliated branch information," which summarizes the data asset processing process implemented by the SQL statement. The application scenario information is an explanation of what kind of service analysis the data results can be used for, such as "This table can be used to further analyze customer purchase preferences or service usage habits" or "The service department can use this table to count the distribution of customers in different time periods and evaluate the effectiveness of marketing activities," thereby transforming technical SQL statements into analysis results that service personnel can directly understand.

[0133] This embodiment proposes a data intelligence service method under a hybrid model, which acquires user metadata information and related statement script information; based on the user metadata information and the related statement script information, a predefined analysis model is input to identify the corresponding data table names and determine the data usage table list information; based on the text name, basic description, and field code values ​​in the data usage table list information, the data resources corresponding to the service demand prediction information are analyzed to obtain data resource analysis results, which include service scenario description information and application scenario information. This solves the technical problem of how to accurately and effectively perform data asset services and analysis. Compared with existing technologies, this application automatically identifies data table names and generates a data usage table list by using user metadata information and related statement script information. Then, based on the text name, basic description, and field code values ​​in the list, semantic analysis is performed on the data resources, and data resource analysis results are output. This method can transform complex and technical scripting languages ​​and data asset metadata into natural language descriptions that are easy for business personnel to understand, significantly reducing the understanding threshold of data assets, improving the efficiency and accuracy of data usage scenario analysis, and reducing reliance on manual interpretation and the risk of errors.

[0134] For example, to help understand the implementation process of the data intelligence service method under the hybrid model obtained by combining this embodiment with the above embodiment one, please refer to... Figure 5 , Figure 5 A simplified flowchart of a data intelligence service method under a hybrid model is provided, specifically: Referring to Example 1, user question-and-answer information and user search information are obtained; based on the user question-and-answer information or the user search information, at least one of the corresponding query intent category, inquiry intent category, exploration behavior category, and search behavior category is identified to determine at least one of the query intent category information, inquiry intent category information, exploration behavior category information, and search behavior category information; based on at least one of the query intent category information, inquiry intent category information, exploration behavior information, and search behavior category information, the corresponding service demand is predicted to determine service demand prediction information; based on the service demand prediction information, the corresponding data assets are analyzed to determine the data resource analysis results. Referring to Example 2, based on the query intent category information or the exploration behavior category information, a predefined recommendation model is input to predict the recommended needs in the corresponding service requirements, thus determining service recommendation need information; based on the inquiry intent category information, a predefined knowledge question-and-answer model is input to predict the knowledge question-and-answer needs in the corresponding service requirements, thus determining knowledge question-and-answer need information; based on the retrieval behavior category information, the retrieval needs in the corresponding service requirements are predicted, thus determining retrieval need information; and based on at least one of the service recommendation need information, the knowledge question-and-answer need information, and the retrieval need information, service need prediction information is obtained. The system identifies at least one of the four types of intents—query, inquiry, exploration, and retrieval—by acquiring user question-and-answer information or user retrieval information. Based on the identified intent categories, the system invokes recommendation models (for data search or exploration intents, generating service recommendation requirements through keyword association, multi-path recall, and large-scale model ranking), knowledge question-answering models (for data inquiry intents, retrieving fragments from the knowledge base based on RAG and generating answers), or traditional search engines (for precise keyword retrieval) to obtain corresponding service recommendation requirements, knowledge question-answering requirements, or retrieval requirements. At least one of these requirements can be used as service requirement prediction information. Then, by inputting user metadata information and related statement script information into a predefined analysis model, the corresponding data table names are identified, and the data usage table list information is determined. Based on the text names, basic descriptions, and field code values ​​in the data usage table list information, the data resources corresponding to the service requirement prediction information are analyzed to obtain data resource analysis results. This achieves precise triage and intelligent integration of scenarios such as data search, data inquiry, precise querying, data understanding, and accurate interpretation, balancing semantic generalization capabilities with low-latency determinism, comprehensively covering diverse user data service demands, and significantly improving the accuracy, comprehensiveness, and efficiency of data asset services in complex semantic scenarios.

[0135] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the data intelligence service method under the hybrid model of this application. Any simple transformations based on this technical concept are within the protection scope of this application.

[0136] This application also provides a data intelligence service device under a hybrid model; please refer to... Figure 6 The data intelligence service device under the hybrid model includes: Module 10 is used to acquire user question and answer information and user search information; Processing module 20 is used to identify at least one of the corresponding query intent category, inquiry intent category, exploration behavior category and retrieval behavior category based on the user question and answer information or the user retrieval information, and determine at least one of the query intent category information, inquiry intent information, exploration behavior information and retrieval behavior category information; The processing module 20 is further configured to predict the corresponding service demand based on at least one of the query intent category information, the inquiry intent category information, the exploration behavior category information, and the retrieval behavior category information, and determine the service demand prediction information; The execution module 30 is used to analyze the corresponding data assets based on the service demand prediction information and determine the data resource analysis results.

[0137] The processing module 20 is also used to obtain user problem history information; Based on the user questions and intent tags corresponding to the user question history information, the training batch sample size, learning rate, iteration rounds, random seed and weight pruning in the hyperparameters of the predefined intent recognition model are adjusted to determine the target intent recognition model; Based on the user question and answer information, the target intent recognition model is input to identify at least one of the query intent category and inquiry intent category under the corresponding model arrangement, and to determine at least one of the query intent category information and inquiry intent information; Based on the filtering conditions and keywords in the user search information, at least one of the corresponding exploration behavior category and search behavior category is identified, and at least one of the exploration behavior category information and search behavior category information is determined.

[0138] The processing module 20 is further configured to predict the recommended requirements in the corresponding service requirements based on the query intent category information or the exploration behavior category information input into a predefined recommendation model, and determine the service recommendation requirement information; Based on the query intent category information, the predefined knowledge question answering model is input to predict the knowledge question answering requirements in the corresponding service requirements and determine the knowledge question answering requirement information. Based on the search behavior category information, the search requirements in the corresponding service needs are predicted to determine the search requirement information; Service demand prediction information is obtained based on at least one of the service recommendation demand information, the knowledge question and answer demand information, and the retrieval demand information.

[0139] The processing module 20 is also used to obtain data service metadata information, which includes basic table information, field information and upstream and downstream lineage information; Based on the data service metadata information and the query intent category information or the exploration behavior category information, the predefined recommendation model is input to predict the corresponding data asset profile and extended keywords, and to determine the data profile information and extended keyword information. Based on the extended keyword information, determine the recall data service information under the corresponding multi-path recall strategy; Based on the data profile information and the recall service profile corresponding to the recall data service information, the recommended requirements in the corresponding service requirements are analyzed to determine the service recommendation requirement information.

[0140] The processing module 20 is also used to perform vector retrieval on the corresponding converted text vector based on the extended keyword information input into a predefined recommendation model, and determine the vector retrieval result; Based on the vector retrieval results, the corresponding keywords are used to perform full-text retrieval using a predefined inverted index to determine the full-text retrieval results under the multi-path recall strategy. The full-text search results are compared with a predefined knowledge graph using a graph mining algorithm to determine the recall data service information under the corresponding multi-path recall strategy.

[0141] The processing module 20 is also used to perform retrieval based on the query intent category information input into a predefined knowledge base, and determine the knowledge base retrieval results. The knowledge base includes basic information of data tables, field information, code value information, graph information and script information. Based on the retrieval results of the knowledge base, the documents in the knowledge base are sliced ​​to determine the sliced ​​text block information; A hybrid retrieval method combining predefined keyword retrieval and vector semantic retrieval is used to recall the fragments corresponding to the sliced ​​text block information from the knowledge base, and the recalled fragment information is determined. Based on the recalled fragment information, a predefined knowledge question-answering model is input to predict the knowledge question-answering requirements in the corresponding service requirements, thereby obtaining knowledge question-answering requirement information.

[0142] The execution module 30 is also used to obtain user metadata information and associated statement script information; Based on the user metadata information and the associated statement script information, the predefined analysis model is input to identify the corresponding data table names and determine the data table list information; Based on the document name, basic description, and field code value in the data list information, the data resources corresponding to the service demand prediction information are analyzed to obtain data resource analysis results, which include service scenario description information and application scenario information.

[0143] The data intelligence service device under the hybrid model provided in this application, employing the data intelligence service method under the hybrid model in the above embodiments, can solve the technical problem of how to accurately and effectively perform data asset services and analysis. Compared with the prior art, the beneficial effects of the data intelligence service device under the hybrid model provided in this application are the same as those of the data intelligence service method under the hybrid model provided in the above embodiments, and other technical features in the data intelligence service device under the hybrid model are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0144] This application provides a data intelligence service device under a hybrid model, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the data intelligence service method under the hybrid model in the above embodiment 1.

[0145] The following is for reference. Figure 7 This document illustrates a structural diagram of a data intelligence service device suitable for implementing the hybrid model of the embodiments of this application. The data intelligence service device in the hybrid model of the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The data intelligence service device shown in the hybrid model is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0146] like Figure 7As shown, the data intelligence service device in the hybrid model may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to programs stored in ROM (Read Only Memory) 1002 or programs loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the data intelligence service device in the hybrid model. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the data intelligence service device in the hybrid model to exchange data via wireless or wired communication with other devices. Although a data intelligence service device in a hybrid model with various systems is shown in the figure, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented or possessed alternatively.

[0147] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0148] The data intelligence service device under the hybrid model provided in this application, employing the data intelligence service method under the hybrid model in the above embodiments, can solve the technical problem of how to accurately and effectively perform data asset services and analysis. Compared with the prior art, the beneficial effects of the data intelligence service device under the hybrid model provided in this application are the same as the beneficial effects of the data intelligence service method under the hybrid model provided in the above embodiments, and other technical features in the data intelligence service device under the hybrid model are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0149] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0150] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0151] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the data intelligence service method under the hybrid model in the above embodiments.

[0152] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0153] The aforementioned computer-readable storage medium may be included in the data intelligence service device under the hybrid model; or it may exist independently and not be assembled into the data intelligence service device under the hybrid model.

[0154] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by a data intelligence service device under a hybrid model, the data intelligence service device under the hybrid model enables the following actions: acquiring user question-and-answer information and user retrieval information; identifying at least one of the corresponding query intent category, inquiry intent category, exploration behavior category, and retrieval behavior category based on the user question-and-answer information or the user retrieval information, and determining at least one of the query intent category information, inquiry intent category information, exploration behavior category information, and retrieval behavior category information; predicting the corresponding service demand based on at least one of the query intent category information, inquiry intent information, exploration behavior information, and retrieval behavior category information, and determining service demand prediction information; and analyzing the corresponding data assets based on the service demand prediction information, and determining the data resource analysis results.

[0155] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0157] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0158] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the data intelligence service method under the above-described hybrid model, and can solve the technical problem of how to accurately and effectively perform data asset services and analysis. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the data intelligence service method under the hybrid model provided in the above embodiments, and will not be repeated here.

[0159] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A data intelligence service method under a hybrid model, characterized in that, The method includes: Obtain user question and answer information and user search information; Based on the user question and answer information or the user search information, at least one of the corresponding query intent category, inquiry intent category, exploration behavior category and search behavior category is identified, and at least one of the query intent category information, inquiry intent information, exploration behavior category and search behavior category information is determined; Based on at least one of the query intent category information, the inquiry intent category information, the exploration behavior category information, and the retrieval behavior category information, the corresponding service demand is predicted to determine the service demand prediction information. Based on the service demand forecast information, the corresponding data resources are analyzed to determine the data resource analysis results.

2. The method as described in claim 1, characterized in that, The step of identifying at least one of the corresponding query intent category, inquiry intent category, exploration behavior category, and retrieval behavior category based on the user question-and-answer information or the user retrieval information, and determining at least one of the query intent category information, inquiry intent information, exploration behavior information, and retrieval behavior information includes: Retrieve user's question history information; Based on the user questions and intent tags corresponding to the user question history information, the training batch sample size, learning rate, iteration rounds, random seed and weight pruning in the hyperparameters of the predefined intent recognition model are adjusted to determine the target intent recognition model; Based on the user question and answer information, the target intent recognition model is input to identify at least one of the query intent category and inquiry intent category under the corresponding model arrangement, and to determine at least one of the query intent category information and inquiry intent information; Based on the filtering conditions and keywords in the user search information, at least one of the corresponding exploration behavior category and search behavior category is identified, and at least one of the exploration behavior category information and search behavior category information is determined.

3. The method as described in claim 1, characterized in that, The step of predicting service demand based on at least one of the query intent category information, the inquiry intent category information, the exploration behavior category information, and the retrieval behavior category information, and determining the service demand prediction information, includes: Based on the query intent category information or the exploration behavior category information, the predefined recommendation model is input to predict the recommendation requirements in the corresponding service requirements and determine the service recommendation requirement information. Based on the query intent category information, the predefined knowledge question answering model is input to predict the knowledge question answering requirements in the corresponding service requirements and determine the knowledge question answering requirement information. Based on the search behavior category information, the search requirements in the corresponding service needs are predicted to determine the search requirement information; Service demand prediction information is obtained based on at least one of the service recommendation demand information, the knowledge question and answer demand information, and the retrieval demand information.

4. The method as described in claim 3, characterized in that, The step of predicting the service recommendation demand information by inputting the query intent category information or the exploration behavior category information into a predefined recommendation model includes: Obtain data service metadata information, which includes basic table information, field information, and upstream and downstream lineage information; Based on the data service metadata information and the query intent category information or the exploration behavior category information, the predefined recommendation model is input to predict the corresponding data asset profile and extended keywords, and to determine the data profile information and extended keyword information. Based on the extended keyword information, determine the recall data service information under the corresponding multi-path recall strategy; Based on the data profile information and the recall service profile corresponding to the recall data service information, the recommended requirements in the corresponding service requirements are analyzed to determine the service recommendation requirement information.

5. The method as described in claim 4, characterized in that, The step of determining the recall data service information under the corresponding multi-path recall strategy based on the extended keyword information includes: Based on the extended keyword information, the predefined recommendation model performs vector retrieval on the corresponding converted text vector to determine the vector retrieval result. Based on the vector retrieval results, the corresponding keywords are used to perform full-text retrieval using a predefined inverted index to determine the full-text retrieval results under the multi-path recall strategy. The full-text search results are compared with a predefined knowledge graph using a graph mining algorithm to determine the recall data service information under the corresponding multi-path recall strategy.

6. The method as described in claim 3, characterized in that, The step of predicting knowledge question-answering requirements in the corresponding service requirements and determining knowledge question-answering requirement information based on the input of the query intent category information into a predefined knowledge question-answering model includes: Based on the query intent category information, a predefined knowledge base is input for retrieval, and the knowledge base retrieval results are determined. The knowledge base includes basic information of data tables, field information, code value information, graph information, and script information. Based on the retrieval results of the knowledge base, the documents in the knowledge base are sliced ​​to determine the sliced ​​text block information; A hybrid retrieval method combining predefined keyword retrieval and vector semantic retrieval is used to recall the fragments corresponding to the sliced ​​text block information from the knowledge base, and to determine the recalled fragment information; Based on the recalled fragment information, a predefined knowledge question-answering model is input to predict the knowledge question-answering requirements in the corresponding service requirements, thereby obtaining knowledge question-answering requirement information.

7. The method as described in claim 1, characterized in that, The step of analyzing the corresponding data resources based on the service demand forecast information and determining the data resource analysis results includes: Retrieve user metadata information and associated statement script information; Based on the user metadata information and the associated statement script information, the predefined analysis model is input to identify the corresponding data table names and determine the data table list information; Based on the document name, basic description, and field code value in the data list information, the data resources corresponding to the service demand prediction information are analyzed to obtain data resource analysis results, which include service scenario description information and application scenario information.

8. A data intelligence service device based on a hybrid model, characterized in that, The device includes: The acquisition module is used to acquire user question and answer information and user search information; The processing module is used to identify at least one of the corresponding query intent category, inquiry intent category, exploration behavior category and retrieval behavior category based on the user question and answer information or the user retrieval information, and to determine at least one of the query intent category information, inquiry intent information, exploration behavior information and retrieval behavior category information; The processing module is further configured to predict the corresponding service demand based on at least one of the query intent category information, the inquiry intent category information, the exploration behavior category information, and the retrieval behavior category information, and determine the service demand prediction information; The execution module is used to analyze the corresponding data resources based on the service demand prediction information and determine the data resource analysis results.

9. A data intelligence service device based on a hybrid model, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the data intelligence service method under the hybrid model as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the data intelligence service method under the hybrid model as described in any one of claims 1 to 7.