Collaborative retrieval method and device, electronic equipment and storage medium
By employing a three-tiered hybrid retrieval architecture and data fusion technology, the problems of a single retrieval architecture and insufficient knowledge integration in insurance intelligent services have been solved, enabling efficient and compliant data retrieval and response, and improving user experience and service efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-03-10
AI Technical Summary
Traditional insurance intelligent service solutions suffer from a single retrieval architecture and insufficient knowledge integration, resulting in poor real-time performance, difficulty in cost control, and a lack of caching optimization, which negatively impacts user experience and service efficiency.
A three-level hybrid retrieval architecture is adopted, including a cache database, a relational database, and a vector database. The Least Recently Used algorithm is used to maintain high-frequency query results, multi-dimensional vector representations are generated by field sharding, and structured and unstructured data are aligned and weighted in a unified semantic space for compliance verification. Finally, the final answer is generated and the cache is updated.
It improved the real-time performance of insurance business data retrieval, enhanced the ability to integrate structured and unstructured knowledge, optimized retrieval cost control, ensured service compliance and response efficiency, and improved the user experience in business scenarios such as intelligent customer service.
Smart Images

Figure CN121636561A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of data processing, and particularly relates to a collaborative retrieval method and device, an electronic device and a storage medium. BACKGROUND
[0002] As the core driving force of the digital transformation of the insurance industry, insurance technology is widely used in intelligent customer service, claim processing and risk assessment, and other key business scenarios. In related technologies, through model fine-tuning, collaborative work of retrieval enhancement generation and knowledge distillation, a technical system of insurance intelligent service is constructed. Specifically, the system covers the whole process from data collection to decision control. With the exponential growth of insurance business data, traditional technical solutions face significant challenges in real-time performance, knowledge fusion and cost control, and breakthroughs in existing technical bottlenecks are urgently needed. SUMMARY
[0003] The present disclosure provides a collaborative retrieval method, device, electronic device and storage medium.
[0004] According to a first aspect of the present disclosure, a collaborative retrieval method is provided, comprising: constructing a three-level hybrid retrieval architecture comprising a cache database, a relational database and a vector database; wherein the cache database maintains high-frequency query results, the relational database stores structured insurance business data and sets up a joint index, and the vector database divides data shards according to a preset field and generates a multi-dimensional vector representation; performing retrieval in the three-level hybrid retrieval architecture based on a query request input by a user; aligning structured data retrieval results and unstructured data retrieval results in a unified semantic space, weighting and fusing the aligned data according to a preset result fusion weight, and generating preliminary response data; performing compliance verification on the preliminary response data, and if the verification is passed, generating a final answer and updating the cache database.
[0005] Optionally, the construction of the three-level hybrid retrieval architecture comprising the cache database, the relational database and the vector database comprises: maintaining high-frequency query results in the cache database based on a least recently used algorithm; dividing data shards in the vector database based on a product type field.
[0006] Optionally, the alignment of the structured data retrieval results and the unstructured data retrieval results in the unified semantic space comprises: extracting policy text features using a bidirectional encoder model, and extracting medical report optical character recognition image features based on a residual network.
[0007] Optionally, the performing compliance verification on the preliminary response data, generating a final answer if the verification is passed, and updating the cache database comprise: The regulatory rules are loaded in the relational database, and regular expressions are used for risk matching; The response data that fails the verification triggers an artificial review process, and a violation log is recorded in an audit database.
[0008] Optionally, the method further comprises: According to the compliance verification result, the preloading strategy of the prediction model is dynamically adjusted; wherein the response data that passes the verification updates the cache database according to a weight.
[0009] According to a second aspect of the present disclosure, a cooperative retrieval device is provided, comprising: A construction unit is configured to construct a three-level hybrid retrieval architecture comprising a cache database, a relational database, and a vector database; wherein the cache database maintains high-frequency query results, the relational database stores structured insurance business data and sets up a joint index, and the vector database divides data shards according to a preset field and generates a multi-dimensional vector representation; A retrieval unit is configured to perform retrieval in the three-level hybrid retrieval architecture based on a query request input by a user; A generation unit is configured to perform multi-modal alignment of structured data retrieval results and unstructured data retrieval results in a unified semantic space, and to perform weighted fusion of the aligned data according to a preset result fusion weight to generate preliminary response data; A verification unit is configured to perform compliance verification on the preliminary response data, and generate a final answer if the verification is passed and update the cache database.
[0010] Optionally, the construction unit is further configured to comprise: The cache database maintains high-frequency query results based on a least recently used algorithm; The vector database divides data shards based on a product type field.
[0011] Optionally, the generation unit is further configured to: A bidirectional encoder model is used to extract policy text features, and a residual network is used to extract medical report optical character recognition image features.
[0012] Optionally, the verification unit is further configured to: The regulatory rules are loaded in the relational database, and regular expressions are used for risk matching; The response data that fails the verification triggers an artificial review process, and a violation log is recorded in an audit database.
[0013] Optionally, the device further comprises: An adjusting unit is configured to dynamically adjust a preloading strategy of the prediction model according to the compliance verification result; wherein the response data passing the verification is used to update the cache database according to a weight.
[0014] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect.
[0015] According to a fourth aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method of the first aspect.
[0016] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of the first aspect.
[0017] The cooperative retrieval method, device, electronic device and storage medium provided by the present disclosure can solve the technical problems of poor real-time performance, weak knowledge fusion capability and difficult cost control of traditional insurance intelligent service solutions due to single retrieval architecture, insufficient knowledge fusion and lack of cache optimization, and achieve the technical effects of improving the real-time performance of insurance business data retrieval, enhancing the structured and unstructured knowledge fusion capability, optimizing the retrieval cost control, ensuring the service response compliance, and improving the user experience and service efficiency of business scenarios such as intelligent customer service.
[0018] It should be understood that the contents described in this part are not intended to identify the key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0019] The accompanying drawings are used to better understand the present application and do not limit the present disclosure. Among them: Figure 1 A flowchart of a cooperative retrieval method provided by an embodiment of the present disclosure; Figure 2 A structural diagram of a cooperative retrieval device provided by an embodiment of the present disclosure; Figure 3 A structural diagram of another cooperative retrieval device provided by an embodiment of the present disclosure; Figure 4 A schematic block diagram of an example electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, which should be considered in a descriptive sense only. Thus, it will be apparent to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.
[0021] The cooperative retrieval method, device, electronic device and storage medium of the embodiments of the present disclosure are described below with reference to the accompanying drawings.
[0022] Figure 1 A flowchart of a cooperative retrieval method provided by an embodiment of the present disclosure.
[0023] As shown in Figure 1 , the method comprises the following steps: Step 101, constructing a three-level hybrid retrieval architecture comprising a cache database, a relational database and a vector database; wherein the cache database maintains high-frequency query results, the relational database stores structured insurance business data and sets up a joint index, and the vector database divides data shards according to a preset field and generates a multi-dimensional vector representation; The core function of the cache database is to maintain high-frequency query results, which are usually derived from the daily high-frequency inquiries of insurance users. By storing these results, the need for repeated calculations for similar queries can be avoided, and the overall response efficiency can be improved by directly calling the stored results. The relational database is mainly used to store structured insurance business data, including policy clause-related data, customer portrait-related data, and other structured information. To further optimize search speed, a joint index needs to be set up in the database. Through the joint index, structured data under specific conditions can be quickly located, reducing the scanning range during data query and ensuring the efficiency of structured data retrieval. The vector database needs to divide data shards according to pre-set fields. The pre-set fields can be determined according to insurance business needs, such as dividing by insurance product type. Unstructured data such as insurance case medical reports corresponding to different product types are classified into different shards to reduce the capacity pressure of a single data shard and avoid retrieval delays caused by excessive data concentration. The vector database also needs to generate multi-dimensional vector representations of the divided unstructured data. By adapting to insurance field embedding models, unstructured data can be converted into fixed-dimensional vectors, allowing unstructured data to be quickly searched through vector similarity matching, providing support for subsequent mixed search of unstructured data.
[0024] Step 102, based on the user input query request, performing retrieval in the three-level mixed retrieval architecture; The user input query request covers various insurance-related scenarios, such as querying the scope of protection of a specific insurance product, understanding the effective date of the policy, and understanding the materials required for claim settlement. When performing retrieval, first interact with the cache database. The cache database stores and maintains high-frequency query results. These results are answers to common needs based on past user query behavior. By prioritizing the cache database for retrieval, it can quickly determine whether there is a matching high-frequency result for the current query request. If there is a matching result, it can provide a basis for subsequent response to reduce overall time consumption.
[0025] If the cache database does not retrieve a matching result, interaction with the relational database is continued, the relational database stores structured insurance business data, including the specific content of the policy terms, customer risk score information, etc., and the database has set a joint index, which can quickly locate the corresponding structured data according to the key information in the query request, avoiding scanning the full amount of data to improve retrieval efficiency. At the same time, interaction with the vector database is also required, the vector database has divided data shards according to the preset fields, and has converted unstructured data into multi-dimensional vector representation, the query request can be converted into the corresponding vector form, and in the relevant data shards of the vector database, the unstructured data related to the query request, such as insurance cases and medical reports, is filtered out through vector similarity calculation, so as to comprehensively obtain the data required to support the query response.
[0026] In step 103, the structured data retrieval result and the unstructured data retrieval result are aligned in a unified semantic space, and the aligned data is weighted and fused according to a preset result fusion weight to generate preliminary response data. The structured data retrieval result and the unstructured data retrieval result are aligned in a unified semantic space, and the unified semantic space refers to a space in which different formats of data can be compared in the same dimension semantic layer. Since the structured data retrieval result comes from the relational database and contains information with a clear format such as policy terms and customer portraits, the unstructured data retrieval result comes from the vector database and covers content without a fixed format such as insurance cases and medical reports, the formats of the two are quite different, and they need to be converted by adapting to the processing mode in the insurance field, so that the structured data is converted into the same dimension representation form as the multi-dimensional vector of the unstructured data, thereby realizing the alignment of the two types of data in the semantic layer and ensuring that the information can be effectively associated during subsequent fusion.
[0027] After the multi-modal alignment is completed, the aligned data is weighted and fused according to a preset result fusion weight, and the preset result fusion weight is set according to the importance proportion of the structured data and the unstructured data in the insurance business scenario. Through this weight, the query-related information contained in the two types of data is calculated comprehensively, which not only retains the accuracy and standardization of the structured data, but also incorporates the reference and richness of the unstructured data, and finally generates preliminary response data that can cover the core needs of the query.
[0028] In step 104, compliance verification is performed on the preliminary response data, and if the verification is passed, the final answer is generated and the cache database is updated.
[0029] In the verification process, the content of the preliminary response data is comprehensively checked according to preset risk rules, and the risk rules are set for possible violations in the insurance field, such as expressions that do not meet regulatory requirements, such as guaranteeing absolute safety of returns, or information that conflicts with insurance business specifications. If the preliminary response data passes the compliance verification, a final answer is generated based on the preliminary response data, and the final answer accurately covers the core needs of the user's query request and provides clear and compliant information feedback to the user. At the same time, the cache database needs to be updated, and the query request of the user this time and the corresponding final answer are stored in the cache database, so that the same or similar query request can be directly retrieved from the cache database in the future, further improving the response speed of the search, and ensuring the effective implementation of the function of continuously maintaining high-frequency query results in the cache database.
[0030] In some embodiments, the construction of a three-level hybrid retrieval architecture including a cache database, a relational database, and a vector database includes: Maintaining high-frequency query results in the cache database based on a least recently used algorithm; Dividing data shards in the vector database based on a product type field.
[0031] For the cache database, the least recently used algorithm is used to maintain high-frequency query results, and the core logic of the least recently used algorithm is to manage the query results in the cache according to access frequency and time, and to preferentially retain high-frequency query results that have been accessed recently, while periodically eliminating low-frequency results that have not been accessed for a long time. This can not only avoid the occupation of cache space by invalid low-frequency data, but also ensure that cache resources are concentrated to serve high-frequency needs, and can also allow subsequent high-frequency query requests to quickly hit the cache, reducing the time consumption caused by repeated calculations.
[0032] For the vector database, data shards are divided based on the product type field, and the product type field can include common insurance categories such as health insurance, car insurance, and property insurance. Unstructured data corresponding to different product types, such as insurance case medical reports and claims records of each category, are classified into corresponding data shards. Through this division method, the vector database does not need to traverse all unstructured data when processing retrieval requests for specific product types, but only needs to perform retrieval operations within the corresponding data shards, effectively reducing the retrieval range and reducing the time consumption of data retrieval. At the same time, it is also convenient to independently update and maintain data of different product types, ensuring the efficiency and pertinence of the vector database for unstructured data retrieval.
[0033] In some embodiments, the multi-modal alignment of structured data retrieval results and unstructured data retrieval results in a unified semantic space includes: The policy text features are extracted by using a bidirectional encoder model, and the medical report optical character recognition image features are extracted based on a residual network.
[0034] For policy text data, a bidirectional encoder model is used to extract features. The bidirectional encoder model has the ability to capture semantic information from both the front and back of the text, can deeply understand the professional content such as the scope of coverage, exemption conditions, and effective date covered in the policy text, accurately mine the semantic association inside the text, avoid information bias caused by single direction semantic understanding, and thus generate feature representations that accurately reflect the core meaning of the policy text.
[0035] For medical report optical character recognition image data, features are extracted based on a residual network. The residual network effectively solves the gradient vanishing problem in deep neural network training by introducing a residual connection structure, can fully capture the text layout details and content information in the medical report optical character recognition image, accurately extract key medical data such as diagnosis results, examination indicators, and treatment suggestions contained in the image, and convert image information into feature vectors that can be used for semantic matching. Through the above two feature extraction methods, policy text and medical report optical character recognition image data of two different modalities are converted into features that meet the dimensional requirements of a unified semantic space, ensuring that the two types of data can be effectively aligned at the same semantic level, providing a consistent feature basis for subsequent weight-based result fusion.
[0036] In some embodiments, the compliance verification is performed on the preliminary response data, and if the verification is passed, a final answer is generated and the cache database is updated, comprising: The relationship database loads the regulatory rules, and risk matching is performed using regular expressions; The response data that does not pass the verification triggers an artificial review process and records the violation log to the audit database.
[0037] Subsequently, risk matching is performed using regular expressions, and according to the specific violation content features in the regulatory rules, corresponding regular expression matching patterns are constructed to scan and match the text content of the preliminary response data sentence by sentence, accurately identify potential violations, and ensure that each potential compliance risk can be effectively detected. If the preliminary response data does not pass the compliance verification, i.e., it is detected that there is a violation, an artificial review process is immediately triggered, and a professional compliance reviewer further checks and determines the violation content, and records the violation log to the audit database. The violation log contains detailed information such as the specific text risk level of the violation content and its location in the preliminary response data, providing complete evidence for subsequent compliance tracing and problem rectification.
[0038] If the preliminary response data passes the compliance check, a final answer is generated based on the preliminary response data, the final answer needs to fully cover the user query requirements and meet all compliance requirements, and the cache database is updated, the query request of the user this time and the corresponding final answer are stored in association to ensure that the same or similar query request can be directly retrieved from the cache database in the future, continuously improving the retrieval response efficiency.
[0039] In some embodiments, the method further comprises: According to the compliance check result, the preloading strategy of the prediction model is dynamically adjusted, and the response data passing the check is updated in the cache database according to a weight.
[0040] The compliance check result is not only used to judge whether the preliminary response data meets the insurance industry regulatory rules and business specifications, but also provides a basis for adjusting the preloading strategy of the prediction model. The preloading strategy of the prediction model is mainly to load the query-related data that may appear in advance to the system to reduce the data loading time in subsequent queries. Based on the compliance check result, the preloading strategy is dynamically adjusted, the related data corresponding to the query type that passes the check and appears frequently in the near future is preferentially included in the preloading range, and the preloading of the related data of the query type that frequently fails the check is reduced or excluded, avoiding invalid preloading from occupying system resources, ensuring that the preloading operation is more in line with the actual compliance query requirements, and improving the effectiveness and pertinence of preloading.
[0041] In addition, for the response data passing the compliance check, the response data needs to be updated in the cache database according to a weight. The weight can be set according to the access frequency of the query request corresponding to the response data, the number of times the response data is used, and other factors. The higher the access frequency and the more the number of times the response data is used, the higher the weight corresponding to the response data. When updating the cache database, the response data with a high weight will have a higher storage priority, which can ensure that high-value compliance response data is always retained in the cache database, and can avoid low-frequency compliance response data from occupying too much cache space, further optimizing the maintenance effect of the cache database on high-frequency compliance query results, and providing more efficient cache support for subsequent retrieval operations.
[0042] Corresponding to the above-mentioned collaborative retrieval method, the present application also proposes a collaborative retrieval device. Since the device embodiments of the present application correspond to the above-mentioned method embodiments, for details not disclosed in the device embodiments, reference can be made to the above-mentioned method embodiments, which will not be described in detail in the present application.
[0043] Figure 2 A structural schematic diagram of a collaborative retrieval device provided by the embodiments of the present disclosure is shown in Figure 2 As shown in the figure, it comprises: The constructing unit 21 is configured to construct a three-level hybrid retrieval architecture including a cache database, a relational database and a vector database; the cache database maintains high-frequency query results, the relational database stores structured insurance business data and sets up a joint index, and the vector database divides data shards according to preset fields and generates multi-dimensional vector representations; The retrieving unit 22 is configured to perform retrieval in the three-level hybrid retrieval architecture based on a query request input by a user. The generating unit 23 is configured to perform multi-modal alignment of structured data retrieval results and unstructured data retrieval results in a unified semantic space, and perform weighted fusion on the aligned data according to a preset result fusion weight to generate preliminary response data. The checking unit 24 is configured to perform compliance checking on the preliminary response data, and if the checking is passed, generate a final answer and update the cache database.
[0044] Further, in a possible implementation of the embodiment of the present disclosure, the constructing unit 21 is further configured to include: maintain high-frequency query results in the cache database based on a least recently used algorithm; divide data shards in the vector database based on a product type field.
[0045] Further, in a possible implementation of the embodiment of the present disclosure, the generating unit 23 is further configured to: extract policy text features using a bidirectional encoder model and extract medical report optical character recognition image features based on a residual network.
[0046] Further, in a possible implementation of the embodiment of the present disclosure, the checking unit 24 is further configured to: load supervision rules in the relational database and perform risk matching using a regular expression; trigger an artificial review process for response data that fails to pass the checking, and record violation logs to an audit database.
[0047] Further, in a possible implementation of the embodiment of the present disclosure, as shown in Figure 3 the apparatus further includes: The adjusting unit 25 is configured to dynamically adjust a preloading strategy of the prediction model according to the compliance checking result; wherein, the response data that passes the checking updates the cache database according to a weight.
[0048] It should be noted that the foregoing explanation and description of the method embodiments are also applicable to the apparatus of the embodiment of the present disclosure, and the principles are the same, which are not limited in the embodiment of the present disclosure.
[0049] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0050] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0051] As shown in Figure 4 The device 400 includes a computing unit 401 that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM (Read-Only Memory) 402 or a computer program loaded into a RAM (Random Access Memory) 403 from a storage unit 408. Various programs and data required for the operation of the device 400 can also be stored in the RAM 403. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An I / O (Input / Output) interface 405 is also connected to the bus 404.
[0052] Various components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0053] The computing unit 401 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs various methods and processes described above, such as the collaborative retrieval method. For example, in some embodiments, the collaborative retrieval method can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded onto the RAM 403 and executed by the computing unit 401, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the aforementioned collaborative retrieval method by any other appropriate means, such as by means of firmware.
[0054] Various implementations of the systems and techniques described above herein can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application-Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0055] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0056] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage medium can include but are not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include one or more lines of electrical conductors, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0057] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0058] The systems and techniques described here can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.
[0059] The computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server is one of communication and distribution, with the server receiving requests from the client and transmitting responses via the communication network. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0060] It should be noted that artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors of people (such as learning, reasoning, thinking, planning, etc.), both hardware and software technologies. Artificial intelligence hardware technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc. several major directions.
[0061] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.
[0062] The above detailed description does not limit the scope of the disclosure. Various modifications, combinations, sub-combinations and alternatives can be made to the detailed description. Any modification, equivalent replacement and improvement etc. made within the spirit and principle of the disclosure shall be included in the scope of the disclosure.
Claims
1. A method of cooperative search, characterized by, The method comprises: constructing a three-level hybrid retrieval architecture comprising a cache database, a relational database, and a vector database; wherein the cache database maintains high-frequency query results, the relational database stores structured insurance business data and sets up a joint index, and the vector database divides data shards according to preset fields and generates multi-dimensional vector representations; performing retrieval in the three-level hybrid retrieval architecture based on a query request input by a user; aligning structured data retrieval results and unstructured data retrieval results in a unified semantic space, weighting and fusing the aligned data according to preset result fusion weights, and generating preliminary response data; performing compliance verification on the preliminary response data, and generating a final answer and updating the cache database if the verification is passed.
2. The method of claim 1, wherein, The construction of the three-level hybrid retrieval architecture comprising a cache database, a relational database, and a vector database comprises: maintaining high-frequency query results in the cache database based on a least recently used algorithm; dividing data shards in the vector database based on a product type field.
3. The method of claim 1, wherein, The multi-modal alignment of structured data retrieval results and unstructured data retrieval results in a unified semantic space comprises: extracting policy text features using a bidirectional encoder model and extracting medical report optical character recognition image features based on a residual network.
4. The method of claim 1, wherein, The compliance verification of the preliminary response data, and the generation of a final answer and the updating of the cache database if the verification is passed, comprises: loading supervision rules in the relational database and performing risk matching using regular expressions; response data that fails to pass the verification triggers an artificial review process and records violation logs in an audit database.
5. The method of claim 1, wherein, The method further comprises: dynamically adjusting the preloading strategy of the prediction model according to the compliance verification result; wherein response data that passes the verification updates the cache database according to a weight.
6. A cooperative search apparatus characterized by comprising: The method comprises: a construction unit configured to construct a three-level hybrid retrieval architecture comprising a cache database, a relational database, and a vector database; wherein the cache database maintains high-frequency query results, the relational database stores structured insurance business data and sets up a joint index, and the vector database divides data shards according to preset fields and generates multi-dimensional vector representations; a retrieval unit configured to perform retrieval in the three-level hybrid retrieval architecture based on a query request input by a user; a generation unit configured to align structured data retrieval results and unstructured data retrieval results in a unified semantic space, weight and fuse the aligned data according to preset result fusion weights, and generate preliminary response data; a verification unit configured to perform compliance verification on the preliminary response data, and generate a final answer and update the cache database if the verification is passed.
7. The apparatus of claim 6, wherein, The construction unit is further configured to comprise: maintaining high-frequency query results in the cache database based on a least recently used algorithm; dividing data shards in the vector database based on a product type field.
8. An electronic device, comprising: The method comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.
9. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are for causing the computer to perform the method of any one of claims 1-5.
10. A computer program product, characterised in that, A computer program comprising instructions which, when executed by a processor, implement the method of any one of claims 1-5.