Index data generation method, information retrieval method, device and computer system

By processing the failed target data in the distributed full-text search server unit and redetermining and writing data using the message queue retry module, the problem of low data retrieval efficiency is solved, and more efficient real-time data writing and retrieval is achieved.

CN115080514BActive Publication Date: 2025-06-20INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210531570.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-16
Publication Date
2025-06-20
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

Data retrieval efficiency in the prior art is not high, especially in the era of big data, it is difficult to meet the real-time needs of intelligent searches with millions of data.

Method used

In response to receiving the target data determined according to the configuration file, it is written to the distributed full-text search server unit and record the write result; if the write fails, the corresponding configuration file is written to the message queue retry module, the target data is re-determined and the writing is retryed, and the index data is finally determined based on the successfully written data.

Benefits of technology

It effectively improves the efficiency of real-time data writing, ensures the integrity of the retrieved data, and thus improves the retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115080514B_ABST
    Figure CN115080514B_ABST
Patent Text Reader

Abstract

The present disclosure provides an index data generation method, an information retrieval method, an apparatus, a computer system, a computer-readable storage medium, and a computer program product, which can be used in the fields of big data, information security technology, or other fields. Among them, the index data generation method includes: in response to receiving first target data determined according to a configuration file, writing the first target data into a distributed full-text search server unit and recording the writing result; in response to detecting a target writing result indicating that the writing process fails, writing the target configuration file corresponding to the target writing result into a message queue retry module; in response to determining that the message queue retry module has received the target configuration file, determining second target data according to the target configuration file; writing the second target data into the distributed full-text search server unit; and determining index data according to the first target data and the second target data written into the distributed full-text search server unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of big data and information security, and more particularly, to an index data generation method, an information retrieval method, an apparatus, a computer system, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the development of big data technology, data retrieval has been increasingly applied to many fields such as industrial and agricultural production, construction, logistics, and daily life. Data retrieval is a process or technology of storing selected, sorted, and evaluated data in a certain carrier and retrieving accurate data that can answer questions from a certain data set according to user needs.

[0003] An index is a decentralized storage structure created to accelerate the retrieval of data rows in a table. An index is established for a table and consists of index pages outside the data pages. Each row in the index page contains a logical pointer to accelerate the retrieval of physical data.

[0004] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following problems in the related art: the data retrieval efficiency is not high. Summary of the Invention

[0005] In view of this, the present disclosure provides an index data generation method, an information retrieval method, an apparatus, a computer system, a computer-readable storage medium, and a computer program product.

[0006] One aspect of the present disclosure provides an index data generation method, including: in response to receiving first target data determined according to a configuration file, writing the first target data into a distributed full-text search server unit and recording the writing result; in response to detecting a target writing result indicating that the process of writing the first target data into the distributed full-text search server unit fails, writing a target configuration file corresponding to the target writing result into a message queue retry module; in response to determining that the message queue retry module receives the target configuration file, determining second target data according to the target configuration file; writing the second target data into the distributed full-text search server unit; and determining index data according to the first target data and the second target data written into the distributed full-text search server unit.

[0007] One aspect of the present disclosure provides an information retrieval method, including: obtaining a target retrieval term; and retrieving the target retrieval term based on index data to obtain a retrieval result; wherein the index data is determined according to the index data generation method of the present disclosure.

[0008] Another aspect of the present disclosure provides an index data generation device, comprising: a first writing module, configured to write the first target data into a distributed full-text search server unit in response to receiving the first target data determined according to a configuration file, and record the writing result; a second writing module, configured to write the target configuration file corresponding to the target writing result into a message queue retry module in response to detecting a target writing result indicating that the process of writing the first target data into the distributed full-text search server unit fails; a first determination module, configured to determine second target data according to the target configuration file in response to determining that the message queue retry module receives the target configuration file; a third writing module, configured to write the second target data into the distributed full-text search server unit; and a second determination module, configured to determine index data according to the first target data and the second target data written into the distributed full-text search server unit.

[0009] One aspect of the present disclosure provides an information retrieval device, comprising: an information retrieval device, comprising: an acquisition module, configured to acquire a target retrieval term; and a retrieval module, configured to perform a retrieval on the target retrieval term based on index data to obtain a retrieval result; wherein the index data is determined according to the index data generation device described in the present disclosure.

[0010] Another aspect of the present disclosure provides a computer system, comprising: one or more processors; a memory, configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the index data generation method and the information retrieval method described in the present disclosure.

[0011] Another aspect of the present disclosure provides a computer-readable storage medium, having stored thereon computer-executable instructions, which are used to implement the index data generation method and the information retrieval method described in the present disclosure when executed.

[0012] Another aspect of the present disclosure provides a computer program product, the computer program product comprising computer-executable instructions, which are used to implement the index data generation method and the information retrieval method described in the present disclosure when executed.

[0013] According to an embodiment of the present disclosure, by adopting the technical means of writing the first target data into the distributed full-text search server unit in response to receiving the first target data determined according to the configuration file, recording the writing result; in response to detecting a target writing result indicating a failed writing process, writing the target configuration file corresponding to the target writing result into the message queue retry module; in response to determining that the message queue retry module has received the target configuration file, determining the second target data according to the target configuration file; writing the second target data into the distributed full-text search server unit; and determining index data based on the first target data and the second target data written into the distributed full-text search server unit, since the first target data with a failed write can be processed in a timely manner, the efficiency of real-time data writing can be effectively improved, so at least partially overcomes the technical problem of low retrieval efficiency, and thus achieves the technical effect of improving the retrieval efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above and other objects, features and advantages of the present disclosure will become clearer. In the drawings:

[0015] Figure 1 Schematically shows an exemplary system architecture to which the index data generation method and the information retrieval method according to the embodiments of the present disclosure can be applied;

[0016] Figure 2 Schematically shows a flowchart of the index data generation method according to the embodiments of the present disclosure;

[0017] Figure 3 Schematically shows a flowchart of generating retrieval data based on an integrated intelligent search engine system according to the embodiments of the present disclosure;

[0018] Figure 4 Schematically shows a flowchart of the information retrieval method according to the embodiments of the present disclosure;

[0019] Figure 5 Schematically shows a structural diagram of an integrated intelligent search engine system with the functions of generating index data and information retrieval according to the embodiments of the present disclosure;

[0020] Figure 6 Schematically shows a block diagram of an index data generation device according to the embodiments of the present disclosure;

[0021] Figure 7 Schematically shows a block diagram of an information retrieval device according to the embodiments of the present disclosure; and

[0022] Figure 8 Schematically shows a block diagram of a computer system suitable for implementing the methods described above according to the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.

[0024] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. as used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0026] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0027] In the process of implementing the concept of the present disclosure, the inventors found that in the context of the big data era with a huge explosion of data resources, the trend of using real-time search technology to complete data aggregation and integration has grown exponentially. In addition, for data volumes of hundreds of thousands or even millions, it is difficult to meet the real-time requirements of intelligent search.

[0028] Embodiments of the present disclosure provide an index data generation method, an information retrieval method, an apparatus, a computer system, a computer-readable storage medium, and a computer program product. The index data generation method includes: in response to receiving first target data determined according to a configuration file, writing the first target data into a distributed full-text search server unit and recording the writing result; in response to detecting a target writing result indicating that the process of writing the first target data into the distributed full-text search server unit fails, writing the target configuration file corresponding to the target writing result into a message queue retry module; in response to determining that the message queue retry module has received the target configuration file, determining second target data according to the target configuration file; writing the second target data into the distributed full-text search server unit; and determining index data according to the first target data and the second target data written into the distributed full-text search server unit. The information retrieval method includes: obtaining a target retrieval term; and performing a retrieval on the target retrieval term based on the index data to obtain a retrieval result; wherein the index data is determined according to the index data generation method of the embodiments of the present disclosure.

[0029] Figure 1 FIG. 1 schematically shows an exemplary system architecture 100 to which the index data generation method and the information retrieval method according to embodiments of the present disclosure can be applied. It should be noted that, Figure 1 The shown is only an example of the system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0030] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0031] Users may use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).

[0032] The terminal devices 101, 102, 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0033] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0034] It should be noted that the index data generation method and the information retrieval method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the index data generation device and the information retrieval device provided by the embodiments of the present disclosure can generally be set in the server 105. The index data generation method and the information retrieval method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the index data generation device and the information retrieval device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Alternatively, the index data generation method and the information retrieval method provided by the embodiments of the present disclosure can also be executed by the terminal devices 101, 102, or 103, or can also be executed by other terminal devices different from the terminal devices 101, 102, or 103. Correspondingly, the index data generation device and the information retrieval device provided by the embodiments of the present disclosure can also be set in the terminal devices 101, 102, or 103, or set in other terminal devices different from the terminal devices 101, 102, or 103.

[0035] For example, the to-be-configured file may originally be stored in any one of the terminal devices 101, 102, or 103 (for example, the terminal device 101, but not limited thereto), or stored on an external storage device and can be imported into the terminal device 101. Then, the terminal device 101 may execute the index data generation method provided by the embodiments of the present disclosure locally, or send the to-be-configured file to other terminal devices, servers, or server clusters, and the other terminal devices, servers, or server clusters that receive the to-be-configured file execute the index data generation method provided by the embodiments of the present disclosure.

[0036] For example, the target search term may originally be stored in any one of the terminal devices 101, 102, or 103 (e.g., terminal device 101, but not limited thereto), or stored on an external storage device and can be imported into terminal device 101. Then, terminal device 101 may execute the information retrieval method provided by the embodiments of the present disclosure locally, or send the target search term to other terminal devices, servers, or server clusters, and the other terminal devices, servers, or server clusters that receive the target search term execute the information retrieval method provided by the embodiments of the present disclosure.

[0037] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0038] It should be noted that the index data generation method, information retrieval method, device, computer system, computer-readable storage medium, and computer program product of the present disclosure can be used in the fields of big data and information security technologies, and can also be used in any field other than the fields of big data and information security technologies. The application fields of the index data generation method, information retrieval method, device, computer system, computer-readable storage medium, and computer program product of the present disclosure are not limited.

[0039] Figure 2 Schematically shows a flowchart of the index data generation method according to an embodiment of the present disclosure.

[0040] As Figure 2 shown, the method includes operations S201 to S205.

[0041] In operation S201, in response to receiving the first target data determined according to the configuration file, write the first target data into the distributed full-text search server unit and record the write result.

[0042] According to an embodiment of the present disclosure, the configuration file may configure at least one of attribute information, access address information, file information, retrieval statement information, and user-related information related to obtaining the first target data. The first target data can be determined according to the configuration information in the configuration file. The configuration file may include a file written and configured by a data source provider according to predefined rules.

[0043] For example, the configuration file may be configured with at least one file name information, field information corresponding to each file name information, and information on the file content corresponding to the file name represented by each file name information, etc. The first target data can be determined from the file content corresponding to the field information according to the field information.

[0044] For example, the configuration file may also be configured with at least one of retrieval statement information, access address information, access port information, instance name information, authorized username and password information, etc. By using the authorized username and password, accessing the set address and port, and retrieving data information related to the instance name based on the retrieval statement, the first target data can be obtained.

[0045] According to an embodiment of the present disclosure, the distributed full-text search server unit can be used to store the first target data for retrieval. Distributed computing is a new computing method. Distributed computing can decompose the search engine into many small modules and then allocate them to multiple servers for processing to save the overall computing time and computing resources and improve computing efficiency. Full-text search is a retrieval method implemented based on the indexing program or derivative program of Lucene (a full-text search engine). Each word segment in the data can be scanned, and an index can be established for each word segment to indicate the number of times and positions where the corresponding word segment appears in the article. When the user retrieves, the retrieval program can search according to the pre-established index and feedback the search results to the user.

[0046] According to an embodiment of the present disclosure, the write result may include at least one of write time period information, identification information of the written data, write success or write failure, and other relevant information, and is not limited thereto.

[0047] In operation S202, in response to detecting a target write result indicating that the process of writing the first target data to the distributed full-text search server unit fails, the target configuration file corresponding to the target write result is written to the message queue retry module.

[0048] According to an embodiment of the present disclosure, when it is detected according to the write result that there is first target data with a write failure, the message queue retry module can store and reprocess the target configuration file corresponding to the first target data with a write failure in the form of a queue.

[0049] In operation S203, in response to determining that the message queue retry module receives the target configuration file, the second target data is determined according to the target configuration file.

[0050] According to an embodiment of the present disclosure, after the message queue retry module receives the target configuration file, the second target data can be obtained according to the configuration information in the target configuration file in combination with the foregoing method for obtaining the first target data.

[0051] In operation S204, the second target data is written to the distributed full-text search server unit.

[0052] According to an embodiment of the present disclosure, as the data for retrieval, the second target data can also be stored in the distributed full-text search server unit.

[0053] In operation S205, index data is determined based on the first target data and the second target data written to the distributed full-text search server unit.

[0054] According to an embodiment of the present disclosure, after receiving the first target data and the second target data for retrieval, the distributed full-text search server unit may establish an index based on the information of the first target data and the second target data, and obtain the index data required during the retrieval process.

[0055] Through the above embodiments of the present disclosure, in the case where the first target data fails to be successfully written to the distributed full-text search server unit, the message queue retry module can be combined to perform a retry operation, and the second target data obtained by the retry is written to the distributed full-text search server unit. Since the first target data with a write failure can be processed in a timely manner, the efficiency of real-time data writing can be effectively improved, the integrity of all constructed data can be ensured, and thus the retrieval efficiency can be improved.

[0056] The following further describes the Figure 2 method shown with specific embodiments.

[0057] According to an embodiment of the present disclosure, in response to receiving the first target data determined according to the configuration file, writing the first target data to the distributed full-text search server unit, and recording the write result may include: in response to receiving the configuration file, determining the first target data through the multi-threaded data processing unit. Writing the first target data to the distributed full-text search server unit. In response to detecting that the first target data is successfully written to the distributed full-text search server unit within a preset time period, recording, through the database unit, a write result indicating that the process of writing the first target data to the distributed full-text search server unit is successful. In response to detecting that the first target data is not successfully written to the distributed full-text search server unit within a preset time period, recording, through the database unit, a write result indicating that the process of writing the first target data to the distributed full-text search server unit fails.

[0058] According to an embodiment of the present disclosure, the multi-threaded data processing unit may provide multiple information processing threads. The multiple information processing threads may perform parallel processing on the information in the received configuration file to obtain the first target data. The multi-threaded processing method may refer to the foregoing method for obtaining the first target data, and will not be elaborated herein.

[0059] According to an embodiment of the present disclosure, after the first target data is processed, the first target data may be written to the distributed full-text search server unit for data retrieval in a subsequent retrieval process.

[0060] According to an embodiment of the present disclosure, in the database unit, a database table for storing the write result can be predefined according to the writing information of the write result. Information related to the write result can be recorded in this database table of the database unit. When detecting whether the write result is successful or not, the recorded information in the database table can be directly and effectively detected.

[0061] Through the above embodiments of the present disclosure, combined with the high concurrency of multi-threading, by determining the first target data through the multi-threading processing unit, the data processing efficiency can be effectively improved, and the rate of writing the first target data into the distributed full-text search server unit can be accelerated. In addition, by recording the write result through the database unit, the management and detection of the write result can be facilitated, further improving the data processing efficiency.

[0062] According to an embodiment of the present disclosure, the target configuration file includes a configuration sub-file configured with a predetermined file name and predetermined field information, and a data sub-file related to the predetermined file name. The message queue retry module includes a message processing system unit and a consumer application unit. In response to determining that the message queue retry module receives the target configuration file, determining the second target data according to the target configuration file may include: through the message processing system unit, writing the configuration sub-file and the data sub-file related to the configuration sub-file as a task into the consumer application unit. According to the configuration sub-file, through the consumer application unit, obtaining the second target data from the data sub-file related to the configuration sub-file.

[0063] According to an embodiment of the present disclosure, the predetermined field information configured in the configuration sub-file can be determined according to the field information included in the file content corresponding to the predetermined file name. For example, if the file content corresponding to a certain predetermined file name app1_YYYYMMDD.csd records information such as ID (identification information), date, address, etc., the field information corresponding to the file app1_YYYYMMDD.csd can be configured as: ID|date|address, etc. At this time, the data sub-file can include all the file content with the file name app1_YYYYMMDD.csd.

[0064] According to an embodiment of the present disclosure, the message processing system unit can write the information of the target configuration file corresponding to the first target data that fails to be written into the distributed full-text search server unit into the consumer application unit. This writing process may include: in the case where the number of target configuration files includes multiple or the configuration sub-files configured in the target configuration file include multiple, writing each configuration sub-file and the data sub-file related to the configuration sub-file as a task into the consumer application unit in the form of a queue.

[0065] According to an embodiment of the present disclosure, the consumer application unit may be a unit implemented based on the Kafka (a high-throughput distributed publish-subscribe messaging system) mechanism. The consumer application unit may be configured to obtain second target data from a data sub-file related to the predetermined field information according to the predetermined field information.

[0066] Through the above embodiments of the present disclosure, in the case where the writing of the first target data to the distributed full-text search server unit fails, the message processing system unit and the consumer application unit may be combined to process the target configuration information related to the first target data, so as to obtain second target data corresponding to the first target data that should have been written to the distributed full-text search server unit. The second target data may be written to the distributed full-text search server unit in a timely manner to ensure the integrity of the target data for retrieval and improve the data retrieval efficiency.

[0067] According to an embodiment of the present disclosure, the target configuration file includes at least one of access port information, access user information, and retrieval statement information related to the target database to be accessed. The message queue retry module includes a message processing system unit and a consumer application unit. In response to determining that the message queue retry module receives the target configuration file, according to the target configuration file, determining the second target data may include: writing at least one of the access port information, access user information, and retrieval statement information as a task to the consumer application unit through the message processing system unit. Obtaining the second target data from the target database through the consumer application unit according to at least one of the access port information, access user information, and retrieval statement information.

[0068] According to an embodiment of the present disclosure, the access interface information may include at least one of IP address information to be accessed, port information, and other information related to the access interface. The access user information may include information such as an authorized user name and password. The retrieval statement information may be used to retrieve the required data file.

[0069] According to an embodiment of the present disclosure, the message processing system unit may write the access port information, access user information, and retrieval statement information corresponding to the first target data that fails to be written to the distributed full-text search server unit to the consumer application unit. The writing process may include: when the access port information, access user information, and retrieval statement information are configured, taking the associated access port information, access user information, and retrieval statement information as a task and writing it to the consumer application unit. The writing process may further include: when only the access port information is configured, different access port information may be taken as individual tasks and written to the consumer application unit in the form of a queue. The writing process may further include: when the access port information and at least one retrieval statement information are configured, the access port information and each retrieval statement information related to the access port information may be taken as a task in sequence and written to the consumer application unit in the form of a queue.

[0070] According to an embodiment of the present disclosure, when the access user information is not configured, the access user information may not be considered when writing a task each time. When the access user information is configured, the corresponding access user information may be written to the consumer application unit as part of the task when writing a task each time.

[0071] According to an embodiment of the present disclosure, the consumer application unit may obtain the second target data from the target database according to the configuration information in each task. For example, when only the access port information is included in a task, all the information in the target database may be determined as the second target data. When the access port information and the retrieval statement information are included in a task, the data information found from the target database according to the retrieval statement represented by the retrieval statement information may be determined as the second target data.

[0072] Through the above embodiments of the present disclosure, the second target data may be written to the distributed full-text search server unit in a timely manner, ensuring the integrity of the target data for retrieval and improving the data retrieval efficiency.

[0073] According to an embodiment of the present disclosure, in response to determining that the message queue retry module receives the target configuration file, determining the second target data according to the target configuration file may include: when it is determined that the message queue retry module receives the target configuration file, in response to receiving a modification operation on the target configuration file written in the message queue retry module, determining the second target data according to the modified target configuration file.

[0074] According to an embodiment of the present disclosure, the reasons for the failure of the process of writing the first target data to the distributed full-text search server unit may include at least one of the following: the multi-threaded data processing unit is overloaded when processing the configuration file in parallel, resulting in the interruption or stagnation of the processing process, and the information configured in the configuration file itself is incorrect, etc.

[0075] According to an embodiment of the present disclosure, when it is determined that the message queue retry module receives the target configuration file, the configuration information in the target configuration file can be checked and modified to ensure the accuracy of the configuration information and achieve the acquisition of valid second target data.

[0076] According to an embodiment of the present disclosure, the method for determining the second target data according to the modified target configuration file can refer to the method for determining the second target data according to the target configuration file described above, which will not be elaborated here.

[0077] Through the above embodiments of the present disclosure, writing the target configuration information to the message queue retry module for modification and retry operations can timely eliminate errors, obtain valid second target data, improve the accuracy of obtaining the target data for retrieval, ensure the integrity of the target data for retrieval, and effectively improve the data retrieval efficiency.

[0078] According to an embodiment of the present disclosure, based on the determined configuration file, in combination with the aforementioned multi-threaded data processing unit, database unit, distributed full-text search server unit, message processing system unit, and consumer application unit, a comprehensive intelligent search engine system with a distributed full-text search server cluster architecture can be constructed.

[0079] Figure 3 Schematically shows a flowchart of generating retrieval data based on the comprehensive intelligent search engine system according to an embodiment of the present disclosure.

[0080] As Figure 3 described, the method includes operations S301 to S315.

[0081] In operation S301, upload the configuration sub-file.

[0082] According to an embodiment of the present disclosure, a configuration file unit and a file system unit may also be provided in the comprehensive intelligent search engine system. The configuration file unit can write the configuration sub-file obtained by the data source provider filling in the file name and file field information according to the rules into the file system unit.

[0083] In operation S302, upload the data sub-file.

[0084] According to an embodiment of the present disclosure, a data file unit may also be provided in the comprehensive intelligent search engine system. The data file unit may write source files related to the file names configured in the configuration sub-file into the file system unit.

[0085] In operation S303, the configuration file is read.

[0086] According to an embodiment of the present disclosure, the file system unit may scan file names and move the source files corresponding to the file names after scanning file names in the corresponding file format.

[0087] In operation S304, the file list is partitioned.

[0088] According to an embodiment of the present disclosure, the source files scanned when scanning the same configuration file by the file system unit may be moved and imported under the same path in the order of date and about to be imported. For the source files scanned in different configuration files, they may be moved and imported under different paths in the same responsible manner.

[0089] In operation S305, the upstream database access information is uploaded.

[0090] According to an embodiment of the present disclosure, a first database unit for storing access information provided by a data source provider may also be provided in the comprehensive intelligent search engine system. The first database unit may build a database connection pool, and write relevant access information into the multi-threaded data processing unit by connecting to databases including MySQL, Oracle, SQL Server, etc. provided by the data source provider. The access information may include at least one of database IP address information of the production environment, access port information, instance name information, username and password information, and relevant SQL statement information, etc.

[0091] It should be noted that operations S301 to S304 may implement a method of importing data for retrieval in the form of file collection, and operation S306 may implement another method of importing data for retrieval in the form of directly accessing a database.

[0092] In operation S306, data is read and processed.

[0093] According to an embodiment of the present disclosure, both the files partitioned in operation S301 and the access information obtained in operation S305 may be written into the multi-threaded data processing unit, and the multi-threaded data processing unit processes the partitioned files and the received access information.

[0094] In operation S307, the multi-threaded processing information is recorded.

[0095] According to an embodiment of the present disclosure, a second database unit for recording processing information may also be provided in the integrated intelligent search engine system. The divided result of the file to be processed and the relevant information during the processing of the divided file by the multi-threaded data processing unit can be recorded in the second database unit, which is beneficial to information tracking.

[0096] In operation S308, write to the target cluster.

[0097] According to an embodiment of the present disclosure, the multi-threaded data processing unit may import the divided files into the target cluster of the distributed full-text search server unit according to the configuration field format of the configuration file.

[0098] According to an embodiment of the present disclosure, the multi-threaded data processing unit may also read access information, obtain an access result, and write it to the target cluster of the distributed full-text search server unit.

[0099] In operation S309, determine whether the data is successfully written? If not, perform operations S310 - S315. If so, perform operations S311 - S315.

[0100] In operation S310, re-consume the data that fails to be written.

[0101] According to an embodiment of the present disclosure, based on the Kafka mechanism, through the message processing system unit and the consumer application unit, the abnormal data can be re-queued for processing.

[0102] In operation S311, write to the backup cluster.

[0103] According to an embodiment of the present disclosure, the data obtained by re-consumption can be written to the backup cluster of the distributed full-text search server unit again.

[0104] In operation S312, record the write result.

[0105] According to an embodiment of the present disclosure, the second database unit may also record the information of successful and failed writes when the data is written to the distributed full-text search server unit, which is beneficial to tracking detection and timely diagnosis of errors.

[0106] In operation S313, build a shard index.

[0107] According to an embodiment of the present disclosure, after the distributed full-text search server unit receives the data to be retrieved, it may allocate a source file list according to an algorithm and build index data in combination with the sharding technology. The process of building the index may adopt the Bulk mode, and the Bulk mode is a high-performance way of loading data. The number of shards can be allocated according to the amount of data for building the index.

[0108] In operation S314, move the processed file to a temporary directory.

[0109] According to an embodiment of the present disclosure, the processed file may include a file with an index built. After moving the processed file into the temporary directory, the original file may be deleted, or the processed file may be stored on a backup distributed full-text search server.

[0110] In operation S315, record the write result.

[0111] According to an embodiment of the present disclosure, the data processing statistical result of this crawl may be written into the second database unit, which is convenient for comparing the data obtained from each crawl.

[0112] Through the above embodiments of the present disclosure, since the first target data with a write failure can be processed in a timely manner, the efficiency of real-time data writing can be effectively improved, so at least partially overcoming the technical problem of low retrieval efficiency, and thus achieving the technical effect of improving the retrieval efficiency.

[0113] Figure 4 Schematically shows a flowchart of an information retrieval method according to an embodiment of the present disclosure.

[0114] As Figure 4 shown, the method includes operations S401 to S402.

[0115] In operation S401, obtain a target retrieval term.

[0116] In operation S402, based on the index data, retrieve the target retrieval term to obtain a retrieval result.

[0117] According to an embodiment of the present disclosure, the target retrieval term may include one or more related terms determined according to the retrieval term in the input retrieval box. The related terms may include at least one of related terms with a similarity greater than a first threshold to the literal content of the retrieval term in the input retrieval box, related terms with a similarity greater than a second threshold to the semantic information of the retrieval term in the input retrieval box, and related terms determined by other means.

[0118] According to an embodiment of the present disclosure, the index data is all data determined according to the foregoing index data generation method. The retrieval result may include results related to all related terms. The retrieval result may be displayed in the order from the highest to the lowest relevance to the retrieval term in the input retrieval box.

[0119] Through the above embodiments of the present disclosure, by writing the data to be retrieved into the full-text search server unit in real time and performing distributed deployment, the practicability of the retrieval function can be effectively improved, the retrieval efficiency and retrieval integrity can be improved, and the user experience can be enhanced.

[0120] According to an embodiment of the present disclosure, obtaining a target search term may include: obtaining an initial search term. Using a preset tokenizer to process the initial search term to obtain at least one target search term whose similarity to the initial search term is greater than a preset threshold.

[0121] According to an embodiment of the present disclosure, the initial search term may be determined according to the search term input into the search box. The preset tokenizer may include a tokenizer that can perform tokenization based on multiple tokenization methods and has a vocabulary expansion function. When performing vocabulary expansion, the expandable vocabulary may be related to actual business requirements.

[0122] According to an embodiment of the present disclosure, after processing the initial search term based on the foregoing preset tokenizer, one or more target search terms may be obtained for retrieving search results related to the initial search term.

[0123] Through the above embodiments of the present disclosure, querying the initial search term by tokenizing the initial search term through a multifunctional preset tokenizer can consider the semantic meaning of the words in the initial search term and situations such as less or more input by the user, and can effectively improve the retrieval accuracy and retrieval efficiency.

[0124] Figure 5 Schematically shows a structural diagram of a comprehensive intelligent search engine system having an index data generation and information retrieval function according to an embodiment of the present disclosure.

[0125] As Figure 5 shown, the comprehensive intelligent search engine system 500 includes a configuration file unit 501, a data file unit 502, a file system unit 503, a first database unit 504, a multi-threaded data processing unit 505, a second database unit 506, a full-text search server unit 507, an online service interface unit 508, a message processing system unit 509, and a consumer application unit 510.

[0126] According to an embodiment of the present disclosure, the search engine may include a search engine implemented based on the ElasticSearch framework. On the one hand, through the configuration file unit 501, the data file unit 502, and the file system unit 503, the data or files to be imported can be imported into the file system unit in a specified format with the assistance of configuration sub-files. After importing the data sub-files and configuration sub-files into the file system unit 503, they can be processed by the multi-threaded processing unit 505. On the other hand, after the information input of the first database unit 504 that stores access information is completed, the multi-threaded data processing unit 505 can directly crawl the data from the database into the full-text search server unit 507 through SQL statements, and the crawled result information can be recorded in the second database unit 506. The full-text search server unit 507 can establish an index based on the received data to implement information query.

[0127] According to an embodiment of the present disclosure, in the case where data writing to the full-text search server 507 fails, the multi-threaded data processing unit 505 can import the input message into the processing system unit 509, and then, through the retry mechanism consumer application unit 510 provided by kafka, rewrite the data to the full-text search server unit 507.

[0128] According to an embodiment of the present disclosure, when performing information query, the online service interface unit 508 can be called through the API interface layer. The online service interface unit 508 can provide an input box for inputting a search term. After receiving the search term through the online service interface unit 508, in combination with the processing of the message processing system unit 109, the search result can be obtained by querying from the full-text search server unit 507.

[0129] According to an embodiment of the present disclosure, a predetermined tokenizer can also be combined during the query process to improve the accuracy and integrity of the query result.

[0130] According to an embodiment of the present disclosure, the comprehensive intelligent search engine that completes full-text search through a distributed system may not belong to multiple regions with different geographical locations.

[0131] Through the above embodiments of the present disclosure, by combining the ElasticSearch and Kafka frameworks, the obtained full-text search engine can perform high-similarity matching based on a reasonable Chinese tokenizer through reasonable distributed deployment, real-time reading and writing of data, etc., and can match the most appropriate result from millions of data within milliseconds after inputting a keyword, improving the real-time performance, accuracy, and computing efficiency of the search engine, and improving the user experience. In addition, the comprehensive intelligent search engine can achieve a high-performance, highly available, and load-balanced cluster architecture through F5 load balancing.

[0132] Figure 6 A block diagram of an index data generation device according to an embodiment of the present disclosure is schematically shown.

[0133] As Figure 6 shown, the index data generation device 600 includes a first writing module 610, a second writing module 620, a first determining module 630, a third writing module 640, and a second determining module 650.

[0134] The first writing module 610 is configured to, in response to receiving first target data determined according to a configuration file, write the first target data to the distributed full-text search server unit and record the writing result.

[0135] A second writing module 620, configured to, in response to detecting a target writing result indicating that the process of writing first target data into a distributed full-text search server unit fails, write a target configuration file corresponding to the target writing result into a message queue retry module.

[0136] A first determination module 630, configured to, in response to determining that the message queue retry module receives a target configuration file, determine second target data according to the target configuration file.

[0137] A third writing module 640, configured to write the second target data into the distributed full-text search server unit.

[0138] A second determination module 650, configured to determine index data according to the first target data and the second target data written into the distributed full-text search server unit.

[0139] According to an embodiment of the present disclosure, the target configuration file includes a configuration sub-file configured with a predetermined file name and predetermined field information, and a data sub-file related to the predetermined file name. The message queue retry module includes a message processing system unit and a consumer application unit. The first determination module includes a first writing unit and a first obtaining unit.

[0140] The first writing unit is configured to write, through the message processing system unit, the configuration sub-file and the data sub-file related to the configuration sub-file as a task into the consumer application unit.

[0141] The first obtaining unit is configured to obtain, according to the configuration sub-file, through the consumer application unit, the second target data from the data sub-file related to the configuration sub-file.

[0142] According to an embodiment of the present disclosure, the target configuration file includes at least one of access interface information related to a target database to be accessed, access user information, and retrieval statement information. The message queue retry module includes a message processing system unit and a consumer application unit. The first determination module includes a second writing unit and a second obtaining unit.

[0143] The second writing unit is configured to write, through the message processing system unit, at least one of the access interface information, the access user information, and the retrieval statement information as a task into the consumer application unit.

[0144] The second obtaining unit is configured to obtain, according to at least one of the access interface information, the access user information, and the retrieval statement information, through the consumer application unit, the second target data from the target database.

[0145] According to an embodiment of the present disclosure, the first determination module includes a first determination unit.

[0146] The first determination unit is configured to, when determining that the message queue retry module receives the target configuration file, in response to receiving a modification operation on the target configuration file written in the message queue retry module, determine second target data according to the modified target configuration file.

[0147] According to an embodiment of the present disclosure, the first writing module includes a second determination unit, a third writing unit, a first recording unit, and a second recording unit.

[0148] The second determination unit is configured to, in response to receiving a configuration file, determine first target data through the multi-thread data processing unit.

[0149] The third writing unit is configured to write the first target data into the distributed full-text search server unit.

[0150] The first recording unit is configured to, in response to detecting that the first target data is successfully written into the distributed full-text search server unit within a preset time period, record, through the database unit, a writing result indicating that the process of writing the first target data into the distributed full-text search server unit is successful.

[0151] The second recording unit is configured to, in response to detecting that the first target data is not successfully written into the distributed full-text search server unit within a preset time period, record, through the database unit, a writing result indicating that the process of writing the first target data into the distributed full-text search server unit fails.

[0152] Figure 7 A block diagram of an information retrieval device according to an embodiment of the present disclosure is schematically shown.

[0153] As Figure 7 shown, the information retrieval device 700 includes an acquisition module 710 and a retrieval module 720.

[0154] The acquisition module 710 is configured to acquire a target retrieval term.

[0155] The retrieval module 720 is configured to perform a retrieval on the target retrieval term based on index data to obtain a retrieval result. The index data is determined according to a retrieval data generation method implemented according to the present disclosure.

[0156] According to an embodiment of the present disclosure, the acquisition module includes an acquisition unit and a processing unit.

[0157] The acquisition unit is configured to acquire an initial retrieval term.

[0158] The processing unit is configured to process the initial retrieval term by using a preset word segmenter to obtain at least one target retrieval term whose similarity to the initial retrieval term is greater than a preset threshold.

[0159] Any of a plurality of modules and units according to embodiments of the present disclosure, or at least part of the functions of any of them, may be implemented in one module. Any one or more of the modules and units according to embodiments of the present disclosure may be split into multiple modules for implementation. Any one or more of the modules and units according to embodiments of the present disclosure may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable way of integrating or packaging circuits, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any of them. Alternatively, one or more of the modules and units according to embodiments of the present disclosure may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.

[0160] For example, any of the first writing module 610, the second writing module 620, the first determining module 630, the third writing module 640, and the second determining module 650, or any of the acquisition module 710 and the retrieval module 720 may be combined and implemented in one module / unit, or any one of the modules / units may be split into multiple modules / units. Alternatively, at least part of the functions of one or more of these modules / units may be combined with at least part of the functions of other modules / units and implemented in one module / unit. According to embodiments of the present disclosure, at least one of the first writing module 610, the second writing module 620, the first determining module 630, the third writing module 640, and the second determining module 650, or the acquisition module 710 and the retrieval module 720 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable way of integrating or packaging circuits, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any of them. Alternatively, at least one of the first writing module 610, the second writing module 620, the first determining module 630, the third writing module 640, and the second determining module 650, or the acquisition module 710 and the retrieval module 720 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.

[0161] It should be noted that the part of the index data generation device in the embodiments of the present disclosure corresponds to the part of the index data generation method in the embodiments of the present disclosure. For the description of the part of the index data generation device, please refer to the part of the index data generation method for details, and will not be repeated here.

[0162] It should be noted that in the embodiments of the present disclosure, the information retrieval device part corresponds to the information retrieval method part in the embodiments of the present disclosure. For the description of the information retrieval device part, please specifically refer to the information retrieval method part, which will not be elaborated herein.

[0163] Figure 8 A block diagram of a computer system suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Figure 8 The computer system shown is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.

[0164] As Figure 8 shown, the computer system 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), and so on. The processor 801 may also include on-board memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0165] In the RAM 803, various programs and data required for the operation of the system 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The processor 801 performs various operations of the method flow according to an embodiment of the present disclosure by executing the programs in the ROM 802 and / or the RAM 803. It should be noted that the program may also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 may also perform various operations of the method flow according to an embodiment of the present disclosure by executing the programs stored in the one or more memories.

[0166] According to an embodiment of the present disclosure, the system 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The system 800 may further include one or more of the following components connected to the I / O interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. The drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage portion 808 as needed.

[0167] According to an embodiment of the present disclosure, the method flow according to the embodiment of the present disclosure may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication portion 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. may be implemented by computer program modules.

[0168] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiment; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.

[0169] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, device, or apparatus.

[0170] For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include one or more memories other than the ROM 802 and / or RAM 803 and / or ROM 802 and RAM 803 described above.

[0171] Embodiments of the present disclosure also include a computer program product, which includes a computer program that contains program code for executing the methods provided by the embodiments of the present disclosure. When the computer program product runs on an electronic device, the program code is used to cause the electronic device to implement the index data generation method and information retrieval method provided by the embodiments of the present disclosure.

[0172] When the computer program is executed by the processor 801, the above functions defined in the system / apparatus of the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the systems, apparatuses, modules, units, etc. described above can be implemented by computer program modules.

[0173] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program can also be transmitted and distributed in the form of a signal on a network medium, and downloaded and installed through the communication part 809, and / or installed from the removable medium 811. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0174] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0175] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions. Those skilled in the art will appreciate that the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0176] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in the respective embodiments cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. An index data generation method, comprising: In response to receiving first target data determined according to a configuration file, write the first target data to a distributed full-text search server unit and record the write result; In response to detecting a target write result indicating that the process of writing the first target data to the distributed full-text search server unit fails, write a target configuration file corresponding to the target write result to a message queue retry module, where the target configuration file includes a configuration sub-file configured with a predetermined file name and predetermined field information and a data sub-file related to the predetermined file name, and the message queue retry module includes a message processing system unit and a consumer application unit, and the consumer application unit is a unit implemented based on a distributed publish-subscribe message mechanism; In response to determining that the message queue retry module receives the target configuration file, determine second target data according to the target configuration file, including: through the message processing system unit, write the configuration sub-file and the data sub-file related to the configuration sub-file as a task to the consumer application unit; And according to the configuration sub-file, obtain the second target data from the data sub-file related to the configuration sub-file through the consumer application unit; Write the second target data to the distributed full-text search server unit; and Determine index data according to the first target data and the second target data written to the distributed full-text search server unit.

2. The method according to claim 1, wherein, The target configuration file includes at least one of access interface information, access user information, and retrieval statement information related to a target database to be accessed, and the message queue retry module includes a message processing system unit and a consumer application unit; The step of, in response to determining that the message queue retry module receives the target configuration file, determining second target data according to the target configuration file includes: Through the message processing system unit, write at least one of the access interface information, access user information, and retrieval statement information as a task to the consumer application unit; and According to at least one of the access interface information, access user information, and retrieval statement information, obtain the second target data from the target database through the consumer application unit.

3. The method according to claim 1, wherein, The step of, in response to determining that the message queue retry module receives the target configuration file, determining second target data according to the target configuration file includes: In the case of determining that the message queue retry module receives the target configuration file, in response to receiving a modification operation on the target configuration file written in the message queue retry module, determine second target data according to the modified target configuration file.

4. The method according to claim 1, wherein, The step of, in response to receiving first target data determined according to a configuration file, writing the first target data to a distributed full-text search server unit and recording the write result includes: In response to receiving the configuration file, determine the first target data through a multi-threaded data processing unit; Write the first target data to the distributed full-text search server unit; In response to detecting that the first target data is successfully written into the distributed full-text search server unit within a preset time period, record, through the database unit, a write result indicating that the process of writing the first target data into the distributed full-text search server unit is successful; and In response to detecting that the first target data is not successfully written into the distributed full-text search server unit within the preset time period, record, through the database unit, a write result indicating that the process of writing the first target data into the distributed full-text search server unit fails.

5. An information retrieval method, comprising: Obtain a target search term; And Based on the index data, perform a search on the target search term to obtain a search result; Wherein, the index data is determined according to the method described in any one of claims 1-4.

6. The method according to claim 5, wherein, The obtaining of the target search term includes: Obtain an initial search term; and Use a preset word segmenter to process the initial search term to obtain at least one target search term whose similarity to the initial search term is greater than a preset threshold.

7. An index data generation apparatus, comprising: A first writing module, configured to, in response to receiving first target data determined according to a configuration file, write the first target data into a distributed full-text search server unit and record the write result; A second writing module, configured to, in response to detecting a target write result indicating that the process of writing the first target data into the distributed full-text search server unit fails, write a target configuration file corresponding to the target write result into a message queue retry module, where the target configuration file includes a configuration sub-file configured with a predetermined file name and predetermined field information and a data sub-file related to the predetermined file name, and the message queue retry module includes a message processing system unit and a consumer application unit, and the consumer application unit is a unit implemented based on a distributed publish-subscribe message mechanism; A first determination module, configured to, in response to determining that the message queue retry module receives the target configuration file, determine second target data according to the target configuration file, where the first determination module includes: a first writing unit, configured to write, through the message processing system unit, the configuration sub-file and the data sub-file related to the configuration sub-file as a task into the consumer application unit; and a first obtaining unit, configured to obtain, according to the configuration sub-file, through the consumer application unit, the second target data from the data sub-file related to the configuration sub-file; A third writing module, configured to write the second target data into the distributed full-text search server unit; and A second determination module, configured to determine index data according to the first target data and the second target data written into the distributed full-text search server unit.

8. An information retrieval apparatus, comprising: An obtaining module, configured to obtain a target search term; And A retrieval module, configured to perform a search on the target search term based on the index data to obtain a search result; Wherein, the index data is determined according to the device described in claim 7.

9. A computer system, comprising: One or more processors; A memory, configured to store one or more programs, Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any one of claims 1 to 4 or 5 to 6.

10. A computer-readable storage medium having executable instructions stored thereon, which when executed by a processor cause the processor to implement the method according to any one of claims 1 to 4 or 5 to 6.

11. A computer program product, comprising computer-executable instructions, which when executed are used to implement the method according to any one of claims 1 to 4 or 5 to 6.

Citation Information

Patent Citations

  • Data access method and device

    CN110019080A

  • Method and device for constructing distributed OLAP data analysis based on MPP and full-text index

    CN114490735A