Searching method, system and device, program product and storage medium

By using the intent identification model to analyze user query intent and combining personalized, sorted and deep learning models to generate search results, the shortcomings of traditional search engines in handling complex queries are solved, and more accurate and personalized search results are achieved, which significantly improves the user experience.

CN120011642AInactive Publication Date: 2025-05-16BEIJING SHANGYIN MICRO CORE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510110115.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional search engines have shortcomings in handling complex and fuzzy query statements, making it difficult to understand the semantic intentions of users, especially when facing queries involving specific domain knowledge or complex logic.

Method used

By obtaining search statements, using pre-trained intent identification models to analyze the user's query intent and needs, select the appropriate search method, and generate search results in combination with personalized models, sorting models and deep learning models to ensure that the results match the user's real needs.

Benefits of technology

It significantly improves the user experience, provides search results that are highly matched with users' real needs, and can accurately understand financial terms and complex logic, and adapt to the specific needs of different banking systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011642A_ABST
    Figure CN120011642A_ABST
Patent Text Reader

Abstract

The invention discloses a search method, system and device, a program product and a storage medium. The method comprises the steps of obtaining a search statement; analyzing the search statement based on an intention recognition model to obtain a search intention; based on the search statement, a retrieval mode suitable for the search statement is selected, and the intention recognition model is used for recognizing the intention and demand of an initiator of the search statement according to the search statement; determining preparation information corresponding to the search statement based on the retrieval mode and the search intention; on the basis of at least one of a personalized model, a sorting model and a deep learning model, a search result corresponding to the search statement is generated, and the personalized model is used for screening the preparation information according to the social network relation of an initiator of the search statement and / or preference information. Therefore, the search result better meets the personalized requirements and behavior modes of the user, and the user experience is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a search method, system, device, program product and storage medium. Background Art

[0002] With the rapid development of banking business and the deepening of digital transformation, banking systems have accumulated a huge amount of data, including but not limited to customer information, transaction records, market intelligence, risk assessment reports, etc. These data are not only the basis for bank operations, but also an important resource for banks to conduct decision-making analysis, product innovation, and customer service optimization. However, faced with such a huge amount of data, how to efficiently retrieve, analyze and utilize this information has become a major challenge facing the banking system.

[0003] Traditional search engines, such as Elasticsearch, perform well in processing text searches, but have obvious deficiencies in understanding and parsing complex and ambiguous query statements. These systems mainly rely on keyword matching and simple grammatical analysis, and have a shallow understanding of semantics, making it difficult to capture the true intention behind user queries. In particular, when faced with queries involving specific domain knowledge or complex logic, the performance of traditional search engines is particularly limited. Summary of the invention

[0004] Based on the above problems, the present application provides a search method, system, device, program product and storage medium.

[0005] The embodiments of the present application disclose the following technical solutions:

[0006] A first aspect of an embodiment of the present application provides a search method, including:

[0007] Get the search statement;

[0008] The search statement is parsed based on an intention recognition model to obtain a search intention; based on the search statement, a retrieval method suitable for the search statement is selected, wherein the intention recognition model is used to identify the intention and needs of the initiator of the search statement according to the search statement;

[0009] Determining preliminary information corresponding to the search statement based on the retrieval method and the search intent;

[0010] Based on at least one of a personalization model, a ranking model, and a deep learning model, a search result corresponding to the search statement is generated, wherein the personalization model is used to filter the preliminary information according to the social network relationship and / or preference information of the initiator of the search statement; the ranking model is used to sort the preliminary information according to the historical behavior information of the initiator of the search statement; and the deep learning model is used to filter the preliminary information according to the semantic information of the search statement and the preliminary information.

[0011] In a possible implementation, the preliminary information is information selected from a database based on an inverted index and corresponding to the retrieval method and the search intent; the database is constructed in the following manner:

[0012] Get initial data information;

[0013] Segmenting the initial data information based on the segmentation model to obtain a segmentation result;

[0014] Filter out the segmented words that meet the preset field requirements in the segmented words results as index fields;

[0015] For each selected index field, an index structure including an inverted index is constructed and stored in the database, where the inverted index is used to represent the position of the index field in the database.

[0016] In a possible implementation manner, after obtaining the initial data information, the method further includes:

[0017] When the initial data information includes unstructured data information, the unstructured data information is analyzed using NLP technology to obtain an analysis result;

[0018] Filter out the word segmentations that meet the preset field requirements in the analysis results to construct a knowledge graph;

[0019] Based on the data information in the knowledge graph, an index structure including an inverted index is generated and stored in a database.

[0020] In a possible implementation, the method further includes:

[0021] Obtaining a selected filtering option, wherein the filtering option is used to adjust the display content of the search results according to the actual needs of the initiator of the search statement;

[0022] Filtering the search results based on the filtering conditions corresponding to each filtering option;

[0023] The screening results are displayed on the front-end interface.

[0024] In a possible implementation, the method further includes:

[0025] Based on the number of search statements obtained within a preset time period and a cache elimination algorithm, the size of the available space of the database is adjusted.

[0026] In a possible implementation, the method further includes:

[0027] When the search result includes non-text data information, calling a processing method corresponding to the non-text data information to process the sub-text data information;

[0028] The information in the processing results that meets the preset validity conditions is displayed on the front-end interface.

[0029] A first aspect of the embodiments of the present application provides a search system.

[0030] An acquisition unit, used for acquiring a search statement;

[0031] A parsing unit, configured to parse the search statement based on an intention recognition model to obtain a search intention; based on the search statement, select a retrieval method suitable for the search statement, wherein the intention recognition model is configured to recognize the intention and needs of the initiator of the search statement according to the search statement;

[0032] A determination unit, configured to determine preliminary information corresponding to the search statement based on the search method and the search intention;

[0033] A generation unit is used to generate search results corresponding to a search statement based on at least one of a personalization model, a ranking model, and a deep learning model, wherein the personalization model is used to filter the preliminary information according to the social network relationships and / or preference information of the initiator of the search statement; the ranking model is used to sort the preliminary information according to the historical behavior information of the initiator of the search statement; and the deep learning model is used to filter the preliminary information according to the semantic information of the search statement and the preliminary information.

[0034] A third aspect of an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the search method as described in the first aspect above is implemented.

[0035] A fourth aspect of the embodiments of the present application provides a computer program product. When the computer program product runs on a computer, the computer executes the search method as described in the first aspect above.

[0036] A fifth aspect of an embodiment of the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a terminal device, the terminal device executes the search method as described in the first aspect above.

[0037] Compared with the prior art, this application has the following beneficial effects:

[0038] Obtain search statements, use the pre-trained intent recognition model to parse the user's search statements, and identify the user's query intent and needs. The model can accurately understand the financial terms, complex logic, and specific business rules in the user's query, providing a basis for subsequent retrieval and screening. According to the parsed search intent, select the most suitable retrieval algorithm and database indexing strategy. Based on the selected retrieval method and the parsed search intent, extract relevant information fragments from the database or other data sources as preliminary information. This information is strictly screened to ensure that it is highly relevant to the user's query. Use a personalized model to consider the user's social network relationships and personal preferences to conduct a preliminary screening of the preliminary information. This step ensures that the content presented to the user not only meets their interests, but also reflects their connections with others. Apply a sorting model to sort the preliminary information based on factors such as the user's historical behavior (such as click records, dwell time), so that the content that the user is most likely to be interested in can be ranked first, improving the user experience. With the help of a deep learning model, further analysis of the semantic relevance between the search statement and its corresponding preliminary information helps to remove irrelevant or low-quality results and ensure the accuracy of the answers provided. The combined use of personalized models, ranking models and deep learning models makes search results more in line with users' personalized needs and behavior patterns, thereby significantly improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0040] Figure 1 A flowchart of a search method provided in an embodiment of the present application;

[0041] Figure 2 A structural diagram of a search system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0043] To facilitate understanding of the technical solutions provided by the embodiments of the present application, the professional terms involved in the embodiments of the present application will be explained below.

[0044] The Chinese name of Elasticsearch is often referred to as "Elastic Search" or "Elastic Search". In different contexts, different translations may be chosen based on habits and preferences.

[0045] To facilitate understanding of the technical solutions provided by the embodiments of the present application, the background technology involved in the embodiments of the present application will be described below.

[0046] As mentioned above, although traditional search engines, such as Elasticsearch, provide powerful full-text search and data analysis capabilities, they have limitations in processing complex queries, understanding user semantic intent, and implementing personalized recommendations. Especially in banking systems, user queries often involve multiple data sources and multiple data types, and the accuracy and real-time requirements of query results are extremely high. Therefore, traditional search engines are difficult to meet the banking system's needs for efficient, accurate, and intelligent information retrieval.

[0047] In banking systems, users’ queries often contain a wealth of financial terms, complex business logic, and implicit demand background. For example, users may ask about the returns, risk assessment, historical performance, etc. of a certain investment product, or ask questions involving multiple combinations of conditions. Such queries not only require search engines to have the ability to accurately identify professional terms, but also to be able to understand the logical relationships and contextual information in the query.

[0048] In order to solve the above problems, in the embodiment of the present application, the user enters a query statement through the search interface of the banking system. The pre-trained intent recognition model is used to parse the user's search statement to identify the user's query intent and needs. The model can accurately understand the financial terms, complex logic and specific business rules in the user's query, and provide a basis for subsequent retrieval and screening. According to the search intent obtained by the analysis, the most suitable retrieval algorithm and database index strategy are selected. For example, for queries involving multiple account types, cross-library retrieval can be selected; for queries containing time ranges, time indexes can be used for efficient retrieval. Based on the selected retrieval method and search intent, the preliminary information related to the query is retrieved from the database. The personalized model is used to filter the preliminary information according to the user's social network relationship and preference information. For example, if the user often pays attention to a certain financial product, the personalized model may put the information related to the product in front of the search results. The sorting model is used to sort the preliminary information according to the user's historical behavior information (such as clicks, purchases, browsing records, etc.). The sorting model can learn the user's preferences and behavior patterns, so as to more accurately predict the content that the user may be interested in. The semantic understanding ability of deep learning is used to match and filter the semantic information of the search statement and the preliminary information. Deep learning models can capture more complex semantic relationships, further improving the accuracy and relevance of search results.

[0049] Through the solution of this application, the search engine of the banking system can more accurately understand the user's query intention and provide search results that are highly matched with the user's real needs. At the same time, the comprehensive use of personalized models, sorting models and deep learning models makes the search results more in line with the user's personalized needs and behavior patterns, thereby significantly improving the user experience. In addition, the solution is also highly scalable and flexible, and can adapt to the specific needs of different banking systems, providing strong support for the digital transformation of banking business.

[0050] It should be noted that the search method, system, device and medium provided in the present application can be applied to the field of computer technology. The above is only an example and does not limit the application field of the search method, system, device and medium provided in the present application. In addition, the embodiments of the present application may not limit the execution subject of the search. For example, the search method of the embodiment of the present application can be applied to data processing devices such as terminal devices or servers. Among them, the terminal device can be an electronic device such as a computer, a personal digital assistant (PDA). The server can be a stand-alone server, a cloud server, or a cluster server composed of multiple servers.

[0051] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0052] The following describes a search method provided by the present application through an embodiment. Figure 1 ,Should Figure 1 A flowchart of a search method provided in an embodiment of the present application, the method comprising:

[0053] S101: Obtain a search statement.

[0054] The user enters a search statement through some input method (such as keyboard input, voice input, etc.), and the search statement expresses the user's query requirements.

[0055] S102. Analyze the search statement based on the intent recognition model to obtain the search intent; based on the search statement, select a retrieval method suitable for the search statement, and the intent recognition model is used to identify the intent and needs of the initiator of the search statement based on the search statement.

[0056] Use the intent recognition model to parse the search statement to understand the user's query intent and needs. At the same time, select the most appropriate retrieval method based on the content and characteristics of the search statement. The intent recognition model is a model trained by machine learning or deep learning technology, which can recognize keywords, phrases and context in the search statement to infer the user's intent. According to the type of search statement, select the corresponding retrieval algorithm and database.

[0057] The following is an explanation of the intent recognition model provided in the embodiment of the present application:

[0058] The intention recognition model based on multi-task learning provided in the embodiment of the present application can simultaneously recognize the user's search intention and potential needs, and provide more intelligent search suggestions. This model not only focuses on the needs directly expressed by the user, but also attempts to understand the implicit intentions behind them, thereby providing users with more personalized and accurate services.

[0059] Using MMoE or Multi-Gated Expert Network (PLE) as the infrastructure allows the model to specialize on different tasks (such as search intent identification and potential demand prediction) while sharing the underlying features. Obtain a large number of real user search samples from the transaction records and customer service interactions of the banking system as the basic dataset for model training. Collect direct feedback from users, including but not limited to satisfaction scores, comment text, etc., to optimize the model's understanding of users' potential needs. Through preprocessing and feature extraction of raw data, construct feature vectors that help characterize user behavior patterns, such as using bag-of-words models, TF-IDF, or BERT to represent search text. Create a neural network structure with a shared layer and multiple task-specific layers, where the shared layer is responsible for capturing common features, while the task-specific layers focus on solving their respective task goals. Based on the annotated search logs and user feedback datasets, adjust the model parameters to minimize the multi-task loss function, thereby improving the performance of the model on all tasks. Integrate the trained model into the existing search system to analyze user input in real time and provide more personalized search suggestions and services.

[0060] S103: Determine preliminary information corresponding to the search statement based on the retrieval method and the search intention.

[0061] Using the selected search method and search intent, preliminary information related to the search statement is retrieved from the database.

[0062] During the search process, in one possible implementation, the embodiment of the present application utilizes libraries such as FAISS (Facebook AI Similarity Search) and combines GPU acceleration technology to achieve efficient indexing and retrieval of vectors. That is, the vector search algorithm provided by the embodiment of the present application can support rapid retrieval of large-scale data sets while ensuring the accuracy and real-time performance of search results.

[0063] FAISS is a library developed by Facebook AI Research for efficient similarity search and dense vector clustering. It is capable of searching in vector sets of arbitrary size and includes support code for evaluation and parameter tuning. The advantage of FAISS is that it improves the retrieval speed of vector similarity and reduces memory usage with a small loss of accuracy. GPU acceleration is an important means of multi-vector retrieval based on graph indexing, which can greatly improve the efficiency of vector retrieval. The implementation of GPU acceleration requires the use of GPU programming languages, such as CUDA, to write vector retrieval programs. For example, RAFT is a composable building block library for accelerating machine learning algorithms on GPUs, including nearest neighbors and ANN algorithms used in vector search. These algorithms can greatly benefit from GPU acceleration.

[0064] In the search process, traditional single retrieval methods (such as keyword retrieval or vector retrieval) are usually difficult to cope with complex and diverse search requirements, especially when facing large-scale, heterogeneous data. For example, when the search involves unstructured text, images, or other multimedia content, relying solely on keyword matching may not provide ideal retrieval results; and retrieval based purely on vector similarity may also be inefficient or inaccurate when dealing with certain specific types of searches. Therefore, a method that can dynamically adjust the retrieval strategy based on the search characteristics and data distribution is needed to optimize the retrieval effect.

[0065] In a possible implementation, the embodiment of the present application dynamically selects the most appropriate search method according to the complexity of the search and the characteristics of the data to achieve the best search effect. This strategy can significantly improve the accuracy and efficiency of the search and meet the search needs in different scenarios.

[0066] Define an initial set of retrieval strategies and their corresponding rule sets to guide the first retrieval operation. After receiving the user's search request, parse the search conditions and evaluate their complexity; at the same time, analyze the distribution characteristics of the data to be retrieved. According to the search conditions and data distribution characteristics, apply the strategy pattern or rule engine to select the most appropriate one or more from the predefined retrieval strategies. When necessary, allow manual intervention to fine-tune the strategy selection process. Perform the retrieval task according to the selected strategy, which can be a single retrieval method or a combination of multiple retrieval methods. Collect the results from different retrieval methods, merge and sort them according to certain rules, and form the final retrieval result list. Record the results of each search and user feedback as a basis for subsequent strategy adjustments, and gradually improve the intelligence level of the retrieval system.

[0067] S104. Generate search results corresponding to the search statement based on at least one of a personalization model, a ranking model, and a deep learning model, wherein the personalization model is used to filter the preliminary information according to the social network relationships and / or preference information of the initiator of the search statement; the ranking model is used to sort the preliminary information according to the historical behavior information of the initiator of the search statement; and the deep learning model is used to filter the preliminary information according to the semantic information of the search statement and the preliminary information.

[0068] The preliminary information is further screened and sorted to generate the final search results. This process may use one or more of the personalization model, sorting model, and deep learning model.

[0069] Personalization model, that is, filtering the preliminary information based on the user's social network relationships and preference information. For example, if the user often mentions a brand or topic on the social network, the personalization model may put this information at the front of the search results. Sorting model, that is, sorting the preliminary information based on the user's historical behavior information (such as clicks, purchases, browsing history, etc.). The sorting model can learn the user's preferences and behavior patterns, so as to more accurately predict the content that the user may be interested in. Deep learning model, that is, using the semantic understanding ability of deep learning to match and filter the semantic information of search statements and preliminary information. The deep learning model can capture more complex semantic relationships, thereby providing more accurate search results.

[0070] In practical applications, these three models can be used simultaneously or in combination to provide more comprehensive, accurate and personalized search results. For example, the personalization model can be used to preliminarily screen the preliminary information, then the sorting model can be used to sort the screened results, and finally the deep learning model can be used to fine-tune the sorted results.

[0071] The following describes the functions and construction methods of the three models:

[0072] In this application, gradient boosting tree algorithms such as XGBoost and LightGBM are used in advance, combined with user feedback data for model training to achieve personalized sorting of search results. A deep learning model for financial search is built, which can deeply understand the semantics and context of bank text data, and can also take into account professional terms in the financial field and their special contexts.

[0073] The process of building a deep learning model can include: collecting and organizing bank text data, including but not limited to announcements, reports, press releases, etc., and marking relevant tags for supervised learning. Extracting text features such as word frequency, TF-IDF value, sentence embedding, etc., and using financial domain expertise to build additional features. Using gradient boosting tree algorithms such as XGBoost or LightGBM as the base model, combined with user feedback data for training. Fine-tuning pre-trained NLP models (such as BERT or GPT) to adapt to specific financial text data. Evaluating model performance through methods such as cross-validation, and continuously optimizing model parameters until satisfactory accuracy and recall are achieved. Deploying the trained model to the actual environment to provide users with more accurate and personalized financial search services.

[0074] The pre-built ranking model in this application provides a ranking algorithm that integrates multiple data sources (such as user behavior, click-through rate, document relevance, etc.), which can dynamically adjust the search result ranking, thereby significantly improving the user experience. This multi-source data fusion method allows the system to understand the user's intention more comprehensively and provide more accurate result ranking.

[0075] Integrate user behavior data (such as browsing history, dwell time, click-through rate, etc.), document relevance data, and possible external data sources (such as social media popularity, expert reviews, etc.) to form a comprehensive data support system. Through machine learning algorithms, extract key features from each data source and build a multi-feature fusion sorting model. Dynamically adjust the weight of each data source in the sorting model based on the specific needs of the search and contextual information (such as time, location, user preferences, etc.). Realize real-time or quasi-real-time updates of sorting results to reflect the latest user needs and changes in the external environment. For example, if a user frequently visits a website or application within a specific time period, then the resources related to it during this period may receive a higher score.

[0076] Unlike traditional static sorting methods, the sorting model provided in the embodiments of the present application can respond quickly to new information collected in real time and update the sorting results in a timely manner. This means that when the user's behavior pattern changes, the system can immediately capture this change and adjust the recommendation list accordingly to ensure that the content provided always matches the user's current needs. For example, if the user has recently increased his attention to a certain type of product, then this type of product will receive a higher weight in future search results.

[0077] The personalized model pre-built in this application is based on graph neural network frameworks such as DGL (Deep Graph Library), combined with user data and social network data of the banking system, and is obtained through model training and inference. It aims to capture the social network relationships between users and provide more accurate personalized search results based on the user's historical behavior and preferences. This method not only goes beyond the traditional method that relies only on individual characteristics, but also attempts to simulate the real social interaction process to better reflect the real needs of users.

[0078] Considering that people's behavior is often influenced by the people around them, the model introduces the concept of social networks, which not only focuses on the preferences of individual users, but also considers the behavior patterns of other members in the social circle to which they belong. For example, in the financial field, if a user's friends or colleagues frequently use a service or product, then the user may also be interested in the service or product. The introduction of this "group wisdom" makes the recommendation results closer to the actual needs of users.

[0079] In one possible implementation, a personalized sorting model is constructed based on the user's search history, behavior patterns and other personal information to intelligently sort the search results and prioritize the content that the user is most likely to be interested in. This personalized sorting is intended to improve the user experience and make the search results more in line with the user's actual needs and preferences.

[0080] Collect user interaction data, such as clickstreams, browsing history, and purchase behavior. Clean and denoise the data, and extract useful features as the basis for training personalized sorting models. Build a rich feature library, including but not limited to users' basic attributes (age, gender), behavioral characteristics (last visit time, average stay time), and content characteristics (category, label). Use machine learning or deep learning frameworks (such as TensorFlow, PyTorch) combined with commonly used algorithms in recommendation systems (such as matrix decomposition and neural network recommendation models) to train personalized sorting models. Introduce a reinforcement learning mechanism so that the model can self-optimize in the feedback loop to further improve the recommendation effect.

[0081] In one possible implementation, the system can create a detailed personal profile for each user by integrating the user's historical behavior data (such as browsing history, purchase history, etc.) and preference information (such as favorite product categories, brand preferences, etc.). These profiles contain not only static personal information, but also dynamic behavior trajectories, thereby helping the system to more accurately predict the user's future behavior.

[0082] The following is an explanation of how the database is constructed:

[0083] Get the preliminary data. This is the starting point for database construction. Raw data can be collected from various sources (such as internal systems, external data sources, user input, etc.). The preliminary data can include various types of information such as text, numbers, dates, etc.

[0084] In one possible implementation, the preprocessing process of the prepared data may include data cleaning of the prepared data, including but not limited to rule-based duplicate data detection, outlier identification (such as abnormal transaction amounts, timestamp errors), and machine learning-based anomaly detection models to improve the accuracy and efficiency of data cleaning.

[0085] To ensure consistency and accuracy of the data set, duplicate records must be removed. Redundant data entries can be identified and removed by defining clear rules. For example, in bank transaction records, if two records have the same customer ID, transaction time, and amount, they may be duplicates.

[0086] Outliers are data points that deviate significantly from other observed values, which may be caused by measurement errors or other reasons. For problems such as abnormal transaction amounts or timestamp errors, statistical methods (such as Z-Score, box plots, etc.) or outlier detection algorithms can be used to identify them, and appropriate measures can be taken to deal with these outliers.

[0087] In one possible implementation, the preprocessing process of the prepared data may include format conversion of data in unstructured data formats in the prepared data. Unstructured data refers to data that does not have a fixed format or pattern, such as text files, images, audio and video, etc. Automatically identify and convert multiple unstructured data formats into a unified structured data format while retaining key business logic and semantic information. In the actual implementation process, OCR (optical character recognition) technology, NLP entity recognition and relationship extraction technology, and customized data mapping rules can be combined to achieve automatic data conversion.

[0088] Optical character recognition (OCR) technology can extract editable text information from text content in the form of scanned documents, pictures, etc. This is very useful for processing paper materials such as PDF reports and handwritten notes. Through OCR technology, these unstructured visual information can be automatically converted into text formats that can be directly processed by computers. Natural language processing (NLP) technology can help parse entities (such as names, places, organizations, etc.) in natural language texts and their relationships. This helps to extract valuable structured information from free-form text descriptions, thereby supporting deeper data analysis. For specific fields or application scenarios, special data mapping rules can be formulated to ensure that the converted data can maintain the original semantics and meet the new structural requirements. For example, in financial statement analysis, the correspondence between fields can be set according to accounting standards to achieve accurate data conversion.

[0089] In one possible implementation, natural language processing (NLP) technology is used to perform deep mining on unstructured text data, extract key information (such as entities, relationships, events, etc.), and build a knowledge graph. This process can not only extract valuable information from massive text data, but also organize this information by building a knowledge graph to form a searchable knowledge network.

[0090] For search results containing non-text data such as pictures and videos, image recognition, speech recognition and other technologies are used to parse and process them, extract useful information and display it to users. Multimedia processing aims to break the text limitations, so that unstructured data can also be effectively used, providing users with a richer content experience.

[0091] In one possible implementation, in order to further optimize data management and search performance, it is also necessary to extract key information from the prepared data after cleaning and conversion, and to establish an efficient index structure, such as customer name, account number, transaction type, etc.

[0092] In the actual implementation process, advanced pre-trained language models such as BERT and RoBERTa can be used and fine-tuned in combination with specific business needs to better adapt to the text characteristics of specific fields. This can not only increase the speed of feature extraction, but also enhance the accuracy of the results. Considering the differences in vocabulary usage in different industries or application contexts, corresponding professional terminology lists can be introduced as auxiliary tools to help the model understand more complex contexts. In addition, the knowledge of domain experts can be used to improve and expand these word libraries to ensure that the final output information is both comprehensive and accurate.

[0093] For example, an NLP model that has been specially trained in the financial field can be used for word segmentation; it can identify and segment words and phrases with practical meaning to improve the accuracy and efficiency of the index; and it can be combined with contextual semantic analysis to ensure that the word segmentation results meet the needs of actual application scenarios.

[0094] In one possible implementation, the embodiment of the present application provides a wealth of filtering options, such as time range, amount range, business type, etc., so that users can accurately filter search results according to their own needs. The multi-dimensional filtering function allows users to customize search results more flexibly, so as to quickly find the required information.

[0095] Determine which fields can be used as filter conditions, and set a reasonable value range or option list for each condition. Create appropriate indexes for different filter conditions to reduce search response time. Use cache technology to store popular search results to avoid repeated calculations. Develop a concise and clear front-end interface so that users can intuitively set and adjust filter conditions. When users modify filter conditions through the front-end interface, the number of results or samples affected are immediately displayed.

[0096] By combining personalized sorting with multi-dimensional filtering, we can not only significantly improve the user's search experience, but also help users find the information they really care about more quickly and accurately. This approach not only improves user satisfaction, but also brings higher participation and stickiness to the platform.

[0097] In one possible implementation, the present application provides a distributed architecture with strong load balancing and fault tolerance, which can ensure the stability and reliability of the system under high concurrency and failure conditions. In the implementation process, container orchestration tools such as Kubernetes are used in combination with the distributed characteristics of Elasticsearch to achieve automatic expansion and load balancing of indexes and data.

[0098] In a possible implementation, the application can automatically adjust the cache size and strategy according to the search frequency and the timeliness of the data to reduce the access pressure on the backend database, aiming to maximize the cache hit rate while minimizing the storage cost. In the implementation process, in combination with high-performance cache systems such as Redis, cache elimination algorithms such as LRU (least recently used) are used to achieve intelligent cache management.

[0099] In a possible implementation, the present application can analyze the search pattern and optimize the search statement and search logic to reduce unnecessary calculations and I / O operations. The goal of the optimizer is to improve search efficiency and reduce system burden through intelligent means.

[0100] In actual application scenarios, you can capture every search request and its execution time and result set size. For long-running or frequently occurring searches, mark them as search requests of particular concern. Use the EXPLAIN command or similar tools to parse SQL statements and identify potential bottlenecks. Apply slow search log analysis tools, such as pt-query-digest, to help locate the root cause of the problem. Make preliminary optimization suggestions based on the rule base, such as adding indexes, restructuring the search structure, etc.

[0101] The above are some specific implementations of the search method provided in the embodiment of the present application. Based on this, the present application also provides a corresponding search system. The system provided in the embodiment of the present application will be introduced from the perspective of functional modularization. Figure 2 A structural diagram of a search system provided in an embodiment of the present application.

[0102] The system comprises:

[0103] An acquisition unit 110, configured to acquire a search statement;

[0104] The parsing unit 111 is used to parse the search statement based on the intention recognition model to obtain the search intention; based on the search statement, select a retrieval method suitable for the search statement, and the intention recognition model is used to identify the intention and needs of the initiator of the search statement according to the search statement;

[0105] A determination unit 112, configured to determine preliminary information corresponding to the search statement based on the search method and the search intent;

[0106] The generation unit 113 is used to generate search results corresponding to the search statement based on at least one of a personalization model, a sorting model, and a deep learning model, wherein the personalization model is used to filter the preliminary information according to the social network relationships and / or preference information of the initiator of the search statement; the sorting model is used to sort the preliminary information according to the historical behavior information of the initiator of the search statement; and the deep learning model is used to filter the preliminary information according to the semantic information of the search statement and the preliminary information.

[0107] The embodiments of the present application also provide corresponding devices and computer storage media for implementing the solutions provided by the embodiments of the present application.

[0108] The device includes a memory and a processor, the memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes the method described in any embodiment of the present application.

[0109] The computer storage medium stores codes, and when the codes are executed, a device executing the codes implements the method described in any embodiment of the present application.

[0110] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system or device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.

[0111] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0112] It should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.

[0113] It should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0114] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0115] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0116] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A search method, characterized in that: include: Get the search statement; Parsing the search statement based on the intent recognition model to obtain the search intent; Based on the search statement, a retrieval method suitable for the search statement is selected, and the intention recognition model is used to recognize the intention and needs of the initiator of the search statement according to the search statement; Determining preliminary information corresponding to the search statement based on the retrieval method and the search intent; Generate search results corresponding to the search statement based on at least one of a personalized model, a ranking model, and a deep learning model, wherein the personalized model is used to filter the preliminary information according to the social network relationship and / or preference information of the initiator of the search statement; The ranking model is used to sort the preliminary information according to the historical behavior information of the initiator of the search statement; The deep learning model is used to filter the preliminary information according to the search statement and semantic information of the preliminary information.

2. The method according to claim 1, characterized in that The preliminary information is information selected from the database based on the inverted index and corresponding to the retrieval method and the search intention; The database is constructed in the following manner: Get initial data information; Segmenting the initial data information based on the segmentation model to obtain a segmentation result; Filter out the segmented words that meet the preset field requirements in the segmented words results as index fields; For each selected index field, an index structure including an inverted index is constructed and stored in the database, where the inverted index is used to represent the position of the index field in the database.

3. The method according to claim 2, characterized in that After obtaining the initial data information, the method further includes: When the initial data information includes unstructured data information, the unstructured data information is analyzed using NLP technology to obtain an analysis result; Filter out the word segmentations that meet the preset field requirements in the analysis results to construct a knowledge graph; Based on the data information in the knowledge graph, an index structure including an inverted index is generated and stored in a database.

4. The method according to claim 1, characterized in that The method further comprises: Obtaining a selected filtering option, wherein the filtering option is used to adjust the display content of the search results according to the actual needs of the initiator of the search statement; Filtering the search results based on the filtering conditions corresponding to each filtering option; The screening results are displayed on the front-end interface.

5. The method according to claim 1, characterized in that The method further comprises: Based on the number of search statements obtained within a preset time period and a cache elimination algorithm, the size of the available space of the database is adjusted.

6. The method according to claim 1, characterized in that The method further comprises: When the search result includes non-text data information, calling a processing method corresponding to the non-text data information to process the sub-text data information; The information in the processing results that meets the preset validity conditions is displayed on the front-end interface.

7. A search system, characterized in that: include: An acquisition unit, used for acquiring a search statement; A parsing unit, used to parse the search sentence based on an intent recognition model to obtain a search intent; Based on the search statement, a retrieval method suitable for the search statement is selected, and the intention recognition model is used to recognize the intention and needs of the initiator of the search statement according to the search statement; A determination unit, configured to determine preliminary information corresponding to the search statement based on the search method and the search intention; a generating unit, configured to generate search results corresponding to the search statement based on at least one of a personalized model, a ranking model, and a deep learning model, wherein the personalized model is configured to filter the preliminary information according to a social network relationship and / or preference information of an initiator of the search statement; The ranking model is used to sort the preliminary information according to the historical behavior information of the initiator of the search statement; The deep learning model is used to filter the preliminary information according to the search statement and semantic information of the preliminary information.

8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the search method according to any one of claims 1 to 6 is implemented.

9. A computer program product, characterized in that When the computer program product runs on a computer, the computer executes the search method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the search method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Program channel fuzzy search method and device based on large model, and terminal

    CN121056696A