Chinese statement searching method and system based on MongoDB database and medium
By configuring multiple database tables in the MongoDB database and using different word segmenters to process Chinese text, the performance problems of MongoDB database in fuzzy search and Chinese search are solved, and accurate search and efficient search efficiency of keywords in different business fields are achieved.
Patent Information
- Application Number
- CN202411800973.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-16
AI Technical Summary
MongoDB database has performance problems when performing fuzzy searches, and does not support full-text indexing in Chinese, so it is impossible to achieve accurate search of Chinese statements.
By configuring multiple database tables in the MongoDB database, each table is configured with a different Chinese word segmenter, performing the first word segmentation on the same Chinese text, generating fragment text, and creating a full text index. The Chinese statement to be query uses different word segmentation devices to process the second word segmentation, obtains keywords, and uses the full text index to search for fragment text in the MongoDB database.
It realizes the support of MongoDB database for Chinese text and word segmentation effects in different business fields, meets the accurate search requirements of keywords in different business fields, improves search efficiency and reduces system complexity and operation and maintenance costs.
Smart Images

Figure CN120011527A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a Chinese sentence search method, system and medium based on a MongoDB database. Background Art
[0002] MongoDB is an open source non-relational database that has been widely used in many industries and application scenarios for its flexibility, scalability and high performance. MongoDB has become one of the preferred databases for many companies and developers to build modern applications due to its ease of use, powerful data processing capabilities and good ecosystem.
[0003] In the prior art, when using the MongoDB database, there are mainly the following technical problems:
[0004] (1) When using the MongoDB database for fuzzy search, it is found that it has obvious performance issues. When the amount of data is large, the response time may reach more than ten seconds or even several minutes.
[0005] (2) The MongoDB database itself provides a full-text index for implementing functions similar to keyword retrieval. However, the full-text index of the MongoDB database is only valid for English text searches, and the MongoDB database does not support Chinese.
[0006] (3) In the prior art, for MongoDB database, there are also some Chinese word segmenters that attempt to support Chinese word segmentation by integrating third-party plug-ins into the database. However, a fixed single Chinese word segmenter cannot achieve the expected word segmentation effect for different business fields, such as professional terms, names, place names, etc., and ultimately cannot meet the effect of accurate keyword search in different business fields.
[0007] (4) Although other middleware that excels in search can be introduced in terms of search, this will undoubtedly increase the complexity of the system and the cost of operation and maintenance.
[0008] In view of the above technical problems, common solutions in the prior art mainly include:
[0009] (1) Create indexes and query optimization:
[0010] In the case of a large amount of data in a MongoDB database, in order to improve query efficiency, we usually use the method of creating indexes and optimizing query statements. For the content fields to be queried, we can create single-field indexes and composite indexes, etc. Through index query, we can avoid full table scans and significantly improve query efficiency. After creating the index, we can optimize the query statement. Common ones include: ensuring that the query conditions can match the existing indexes, especially for composite indexes, the order of the index fields must be consistent with the order of the fields in the query conditions. Try to make the query be completed completely through the index, that is, all the fields of the query are included in the index, and there is no need to access the document itself, which is the so-called "index coverage query" to avoid additional query overhead. Avoid using index-unfriendly operations: such as using regular expressions in queries, inexact matching of array fields, or calculations on index fields, which may cause the index to not be effectively used.
[0011] The disadvantages of this solution: The common indexes (single-field indexes and composite indexes) created in this solution are effective for exact match queries, but they do not perform index queries for fuzzy queries. The query efficiency is still very low in the case of large amounts of data. MongoDB supports the creation of full-text index queries, but does not support Chinese word segmentation, and cannot perform keyword queries on Chinese sentences after word segmentation.
[0012] (2) Use other search middleware.
[0013] For fuzzy query requirements in large data volumes, you can usually consider introducing and using text-oriented search engine middleware (such as Elasticsearch, Apache Lucene, etc.). When the search engine stores data, it creates an inverted index after word segmentation of the search text, which can meet better keyword search efficiency.
[0014] The shortcomings of this solution are: when introducing other search engine middleware while using MongoDB database as the storage database, it is necessary to consider the data synchronization issues of MongoDB and search middleware at the same time. The software system needs to increase the adaptation of the newly introduced middleware, increase the server and operation and maintenance costs of deploying the middleware, and increase the complexity and operation and maintenance costs of the system. Summary of the invention
[0015] The present invention proposes a Chinese sentence search method, system and medium based on MongoDB database to solve the technical problem that the Chinese word segmenter of MongoDB database integrated with a third-party plug-in cannot achieve the expected word segmentation effect for different business fields and cannot meet the effect of accurate keyword search in different business fields.
[0016] One aspect of the present invention is to provide a Chinese sentence search method based on a MongoDB database, the Chinese sentence search method comprising the following method steps:
[0017] S1. Configure multiple database tables in a MongoDB database; wherein each database table includes a text field symbol, and the same Chinese text is stored under the text field symbols of different database tables;
[0018] S2. Configure a Chinese word segmenter for each database table in the MongoDB database. Different database tables are configured with different Chinese word segmenters.
[0019] S3, using the corresponding Chinese word segmenter to perform first word segmentation processing on the same Chinese text stored under the text field symbols of different database tables, and concatenating the Chinese text after the first word segmentation processing using spaces to obtain segmented Chinese texts corresponding to different database tables;
[0020] S4. For each database table in the MongoDB database, create a fragment text field symbol; store the fragment Chinese texts corresponding to different database tables in the fragment text field symbol of the corresponding database table respectively;
[0021] S5. Create a full-text index for each fragment text field in each database table;
[0022] S6, obtaining the Chinese sentence to be queried;
[0023] Using different Chinese word segmenters configured with different database tables, the second word segmentation process is performed on the Chinese sentence to be queried to obtain keywords corresponding to different database tables;
[0024] S7. Use full-text indexing to search for Chinese text fragments corresponding to different database tables in the MongoDB database, and complete the Chinese sentence search.
[0025] In a preferred embodiment, in step S1, three database tables, namely, an event table, a policy table and a service table, are configured in the MongoDB database.
[0026] In a preferred embodiment, the matter table is configured with a matrix constraint Chinese word segmenter, the policy table is configured with a grammatical analysis Chinese word segmenter, and the service table is configured with a neural network Chinese word segmenter.
[0027] In a preferred embodiment, in step S5, an inverted index is used to create a full-text index for each fragment text field symbol of each database table.
[0028] Another aspect of the present invention is to provide a Chinese sentence search system based on a MongoDB database, the Chinese sentence search system comprising:
[0029] The database table configuration module is used to configure multiple database tables in the MongoDB database; each database table includes a text field symbol, and the same Chinese text is stored under the text field symbols of different database tables;
[0030] The Chinese word segmenter configuration module is used to configure a Chinese word segmenter for each database table in the MongoDB database. Different database tables are configured with different Chinese word segmenters.
[0031] A first word segmentation processing module is used to perform first word segmentation processing on the same Chinese text stored in the text field symbols of different database tables using the corresponding Chinese word segmenter, and to concatenate the Chinese text after the first word segmentation processing using spaces to obtain segmented Chinese texts corresponding to different database tables;
[0032] The fragment text field symbol creation module is used to create a fragment text field symbol for each database table in the MongoDB database; the fragment Chinese texts corresponding to different database tables are stored in the fragment text field symbols of the corresponding database tables respectively;
[0033] The index creation module is used to create a full-text index for each fragment text field in each database table;
[0034] The second word segmentation processing module is used to obtain the Chinese sentence to be queried; different Chinese word segmenters configured with different database tables are used to perform the second word segmentation processing on the Chinese sentence to be queried to obtain keywords corresponding to different database tables;
[0035] The search module is used to use full-text indexing to search for Chinese text fragments corresponding to different database tables in the MongoDB database, thereby completing Chinese sentence searches.
[0036] In a preferred embodiment, in the library table configuration module, three database tables, namely, a configuration item table, a policy table and a service table, are provided in the MongoDB database.
[0037] In a preferred embodiment, the matter table is configured with a matrix constraint Chinese word segmenter, the policy table is configured with a grammatical analysis Chinese word segmenter, and the service table is configured with a neural network Chinese word segmenter.
[0038] In a preferred embodiment, the index creation module uses an inverted index to create a full-text index for each fragment text field symbol of each database table.
[0039] Another aspect of the present invention is to provide a computer storage medium, wherein the computer storage medium is used to store computer execution instructions, and the computer execution instructions are used to execute a Chinese sentence search method based on MongoDB database provided by the present invention.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] The present invention proposes a Chinese sentence search method, system and medium based on MongoDB database. Multiple database tables are configured in MongoDB database. When Chinese text is stored in MongoDB database, the same Chinese text is stored in multiple database tables. By configuring different Chinese word segmenters for different database tables, the same Chinese text stored in different database tables is subjected to first word segmentation processing, so that MongoDB database supports Chinese word segmentation, and the same Chinese text stored in multiple database tables is segmented according to different business fields, so that the Chinese text achieves the expected word segmentation effect in different business fields, and meets the effect of accurate keyword search in different business fields.
[0042] The present invention proposes a Chinese sentence search method, system and medium based on MongoDB database, which creates a fragment text field symbol for each database table in MongoDB database. After Chinese text is segmented according to different business fields, the fragment Chinese texts corresponding to different database tables are respectively stored under the fragment text field symbols of the corresponding database tables, and a full-text index is created for the fragment text field symbol of each database table, without the need to deploy middleware servers and operation and maintenance costs, thereby reducing the complexity and operation and maintenance costs of the system.
[0043] The present invention proposes a Chinese sentence search method, system and medium based on MongoDB database. For the Chinese sentence to be queried, different Chinese word segmenters configured with different database tables are used to perform a second word segmentation process to obtain keywords corresponding to different database tables. Full-text indexing is used to search for Chinese text fragments corresponding to different database tables in the MongoDB database, thereby achieving the effect of accurate search of keywords in different business fields.
[0044] The present invention proposes a Chinese sentence search method, system and medium based on MongoDB database, which significantly improves the search efficiency in the case of large data volume, and the search effect is similar to fuzzy search. The bottom layer relies on the inverted index created by MongoDB, and can quickly retrieve relevant data through keywords. There is no need to introduce other text search engine middleware or database plug-ins, and only based on MongoDB database can achieve better search results, without considering the data synchronization problem between the database and the middleware, and without additional server hardware and operation and maintenance costs.
[0045] The present invention proposes a Chinese sentence search method, system and medium based on MongoDB database, which can customize the Chinese word segmentation algorithm, adapt to the word segmentation retrieval of professional terms in different business fields, and can also customize the filtering of auxiliary words, conjunctions, etc. that affect the search efficiency, and can also customize the filtering of sensitive words, junk words, etc. to avoid being searched.
[0046] The present invention proposes a Chinese sentence search method, system and medium based on MongoDB database. The response time for searching keywords is within 1 second when the data volume is in the tens of millions. Compared with fuzzy search, it has obvious performance improvement. In addition, the present invention can flexibly configure different word segmentation algorithms for business data in different database tables, so as to meet the professional words and sentences in different industry fields and still have good search effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the specific implementation of the present invention or the technical solution in the prior art, the following briefly introduces the drawings required for use in the specific implementation or the prior art description. Obviously, the drawings described below are some implementations of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0048] Figure 1 It is a flow chart of a Chinese sentence searching method based on MongoDB database of the present invention.
[0049] Figure 2 It is a structural block diagram of a Chinese sentence search system based on MongoDB database of the present invention. DETAILED DESCRIPTION
[0050] In order to make the above and other features and advantages of the present invention more clear, the present invention is further described below in conjunction with the accompanying drawings. It should be understood that the specific embodiments given herein are for the purpose of explaining to those skilled in the art and are only exemplary and not restrictive.
[0051] Combination Figure 1 According to an embodiment of the present invention, a Chinese sentence search method based on a MongoDB database is provided, comprising the following steps:
[0052] Step S1, configuring multiple database tables in a MongoDB database; wherein each database table includes a text field symbol, and the same Chinese text is stored under the text field symbols of different database tables.
[0053] In this embodiment, three database tables are configured in the MongoDB database, namely, a matter table (matter), a policy table (policy) and a service table (service) are configured in the MongoDB database.
[0054] Each database table includes a text field symbol (content), and the same Chinese text is stored under the text field symbols (content) of different database tables.
[0055] For example, in this embodiment, the library table structure of the matter table is as follows:
[0056] matter table:
[0057] {
[0058] "_id":ObjectId("5dde1371740cd66298f929fb"),
[0059] "content":"Special awards for designers of software and integrated circuit enterprises",
[0060] }.
[0061] The library table structure of the policy table is as follows:
[0062] Policy table:
[0063] {
[0064] "_id":ObjectId("5dde1371740cd66298f929fb"),
[0065] "content":"Special awards for designers of software and integrated circuit enterprises",
[0066] }.
[0067] The service table structure is as follows:
[0068] service table:
[0069] {
[0070] "_id":ObjectId("5dde1371740cd66298f929fb"),
[0071] "content":"Special awards for designers of software and integrated circuit enterprises",
[0072] }.
[0073] In this embodiment, the same Chinese text "Special rewards for software and integrated circuit enterprise designers" is stored under the text field symbol (content) of the three database tables, namely, the matter table (matter), the policy table (policy) and the service table (service).
[0074] It should be understood that the same Chinese text stored under the text field symbol (content) of different database tables of the present invention is not limited to the above-mentioned Chinese text. In actual applications, there are tens of millions of Chinese texts under the text field symbol (content) of each database table. In order to make the present invention more concise, this embodiment takes the same Chinese text "Special Rewards for Software and Integrated Circuit Enterprise Designers" stored under the text field symbol (content) of the three database tables of the matter table (matter), policy table (policy) and service table (service) as an example for explanation.
[0075] Step S2: configure a Chinese word segmenter for each database table in the MongoDB database, and different database tables are configured with different Chinese word segmenters.
[0076] In this embodiment, three database tables, namely, the matter table, the policy table and the service table, are respectively configured with a Chinese word segmenter, and different database tables are configured with different Chinese word segmenters.
[0077] Furthermore, in this embodiment, the matter table (matter) is configured with a matrix constraint method Chinese word segmenter (A), the policy table (policy) is configured with a grammatical analysis method Chinese word segmenter (B), and the service table (service) is configured with a neural network method Chinese word segmenter (C).
[0078] The Chinese word segmenters corresponding to the three database tables, matter, policy, and service, are saved according to the following corresponding relationship:
[0079] Matter table (matter) - matrix constraint method Chinese word segmenter (A);
[0080] Policy table (policy) - grammar analysis method Chinese word segmenter (B);
[0081] Service table (service) - Neural network Chinese word segmenter (C).
[0082] Different Chinese word segmenters configured for different database tables can be saved in a configuration file or directly saved in a MongoDB database. The present invention does not specifically limit the storage location of the specific Chinese word segmenter, and those skilled in the art can select it according to actual needs.
[0083] The matrix constraint method Chinese word segmenter (A), the grammatical analysis method Chinese word segmenter (B) and the neural network method Chinese word segmenter (C) of the present invention adopt the existing Chinese word segmenters. The Chinese word segmenters corresponding to the three database tables of the matter table (matter), the policy table (policy) and the service table (service) are reasonably selected according to the needs of the actual business field. The commonly used Chinese word segmenters are not limited to the above-mentioned matrix constraint method Chinese word segmenter (A), the grammatical analysis method Chinese word segmenter (B) and the neural network method Chinese word segmenter (C).
[0084] Step S3: Perform a first word segmentation process on the same Chinese text stored under the text field symbols of different database tables using the corresponding Chinese word segmenter, and concatenate the Chinese text after the first word segmentation process using spaces to obtain Chinese text fragments corresponding to different database tables.
[0085] In this embodiment, the Chinese text "Special rewards for designers of software and integrated circuit enterprises" stored in the matter table (matter) is first segmented using the corresponding matrix constraint method Chinese word segmenter (A), and the Chinese text after the first segmentation is concatenated with spaces to obtain the corresponding Chinese text fragment of the matter table (matter):
[0086] "Special awards for designers of software and integrated circuit companies".
[0087] The Chinese text "Special rewards for designers of software and integrated circuit enterprises" stored in the policy table (policy) is first segmented using the corresponding grammar analysis method Chinese word segmenter (B), and the Chinese text after the first segmentation is concatenated with spaces to obtain the corresponding Chinese text fragment of the policy table (policy):
[0088] "Special awards for designers of software and integrated circuits in integrated circuit companies".
[0089] The Chinese text "Special rewards for designers of software and integrated circuit enterprises" stored in the service table (service) is processed by the corresponding neural network Chinese word segmenter (C) for the first word segmentation, and the Chinese text after the first word segmentation is concatenated with spaces to obtain the corresponding Chinese text fragment of the service table (service):
[0090] "Software Integrated Circuit Enterprise Designer Award".
[0091] Step S4: For each database table in the MongoDB database, a fragment text field symbol is created; and the fragment Chinese texts corresponding to different database tables are respectively stored in the fragment text field symbols of the corresponding database tables.
[0092] In this embodiment, a segment text field symbol (segmentContent) is created for each of the three database tables, namely, the matter table (matter), the policy table (policy), and the service table (service) configured in the MongoDB database.
[0093] The Chinese text segments of the matter table (matter), policy table (policy) and service table (service) are stored in the segment text field symbol (segmentContent) of the corresponding database table.
[0094] In this embodiment, after the Chinese text segments corresponding to different database tables are stored in the segment text field symbol (segmentContent) of the corresponding database table, the obtained database table structures of the matter table (matter), policy table (policy) and service table (service) are as follows:
[0095] The database table structure of the matter table is as follows:
[0096] matter table:
[0097] {
[0098] "_id":ObjectId("5dde1371740cd66298f929fb"),
[0099] "content":"Special awards for designers of software and integrated circuit enterprises",
[0100] "segmentContent":"Special awards for designers in software and integrated circuit companies",}.
[0101] The library table structure of the policy table is as follows:
[0102] Policy table:
[0103] {
[0104] "_id":ObjectId("5dde1371740cd66298f929fb"),
[0105] "content":"Special awards for designers of software and integrated circuit enterprises",
[0106] "segmentContent":"Special rewards for designers of software and integrated circuit companies",
[0107] }.
[0108] The service table structure is as follows:
[0109] service table:
[0110] {
[0111] "_id":ObjectId("5dde1371740cd66298f929fb"),
[0112] "content":"Special awards for designers of software and integrated circuit enterprises",
[0113] "segmentContent": "Software integrated circuit enterprise designer rewards",
[0114] }.
[0115] Step S5: Create a full-text index for each fragment text field in the database table.
[0116] In this embodiment, full-text indexes are created for the segment text field symbol (segmentContent) in the matter table (matter), policy table (policy) and service table (service) respectively.
[0117] Preferably, an inverted index is used for the full-text index created for the segment text field symbol (segmentContent) of each database table.
[0118] In this embodiment, the full-text index created for the segment text field symbol (segmentContent) in the matter table (matter), policy table (policy) and service table (service) is as follows:
[0119] db.matter.createIndex({"segmentContent":"text"});
[0120] db.policy.createIndex({"segmentContent":"text"});
[0121] db.service.createIndex({"segmentContent":"text"}).
[0122] Step S6: obtaining the Chinese sentence to be queried, using different Chinese word segmenters configured with different database tables, performing a second word segmentation process on the Chinese sentence to be queried, and obtaining keywords corresponding to different database tables.
[0123] In this embodiment, the Chinese sentence to be queried "software and integrated circuit" is taken as an example, and the matrix constraint method Chinese word segmenter (A) configured by the matter table (matter) is used to perform the second word segmentation processing on the Chinese sentence to be queried "software and integrated circuit" to obtain the keywords corresponding to the matter table (matter): ["software", "and", "integrated circuit"].
[0124] The Chinese word segmenter (B) using the grammatical analysis method configured by the policy table (policy) performs the second word segmentation processing on the Chinese sentence "software and integrated circuit" to be queried, and obtains the keywords corresponding to the policy table (policy): ["software", "and", "integrated circuit", "integrated", "circuit"].
[0125] The neural network Chinese word segmenter (C) configured using the service table (service) performs the second word segmentation processing on the Chinese sentence "software and integrated circuit" to be queried, and obtains the keywords corresponding to the service table (service): ["software", "integrated circuit"].
[0126] In some preferred embodiments, after the second word segmentation processing is performed on the Chinese sentence to be queried, auxiliary words, conjunctions, modal particles, etc. in the word segmentation are filtered out according to the search requirements to improve the search accuracy.
[0127] Step S7: Using full-text indexing, the keywords corresponding to different database tables are searched in the MongoDB database for Chinese text fragments corresponding to different database tables to complete the Chinese sentence search.
[0128] In this embodiment, the keywords ["software", "and", "integrated circuit"] of the corresponding matter table (matter), the keywords ["software", "and", "integrated circuit", "integration", "circuit"] of the corresponding policy table (policy) and the keywords ["software", "integrated circuit"] of the corresponding service table (service) are searched in the MongoDB database using full-text indexing for the fragmented Chinese text stored under the fragment text field symbol (segmentContent) of the corresponding matter table (matter), policy table (policy) and service table (service).
[0129] Finally, the searched Chinese text fragments are assembled to complete the Chinese sentence search. In this embodiment, the result after assembling the Chinese text fragments is as follows:
[0130] db.matter.find({"$text":{"$search":"Software and Integrated Circuits"}});
[0131] db.policy.find({"$text":{"$search":"Software and Integrated Circuits Integrated Circuits"}});
[0132] db.service.find({“$text”:{“$search”:”Software Integrated Circuit”}}).
[0133] The present invention configures different Chinese word segmenters for different database tables, performs the first word segmentation processing on the same Chinese text stored in different database tables, realizes that the MongoDB database supports Chinese word segmentation, and enables the same Chinese text stored in multiple database tables to be segmented according to different business fields, so that the Chinese text can achieve the expected word segmentation effect in different business fields. For the Chinese sentence to be queried, different Chinese word segmenters configured in different database tables are used to perform the second word segmentation processing, obtain the keywords corresponding to the different database tables, use the full-text index, search for the fragmented Chinese text corresponding to the different database tables in the MongoDB database, achieve the effect of accurate search of keywords in different business fields, and can obtain a better keyword matching effect.
[0134] The present invention creates a fragment text field symbol for each database table in the MongoDB database. After Chinese text is segmented according to different business fields, the fragment Chinese texts corresponding to different database tables are respectively stored under the fragment text field symbols of the corresponding database tables, and a full-text index is created for the fragment text field symbol of each database table. There is no need to deploy a middleware server and operation and maintenance costs, which reduces the complexity and operation and maintenance costs of the system, and can shorten the response time of keyword matching while obtaining a better keyword matching effect.
[0135] Combination Figure 2 According to an embodiment of the present invention, a Chinese sentence search system based on MongoDB database is provided, comprising:
[0136] The database table configuration module 100 is used to configure multiple database tables in a MongoDB database; wherein each database table includes a text field symbol, and the same Chinese text is stored under the text field symbols of different database tables.
[0137] Preferably, in the library table configuration module 100, three database tables, namely, a matter table (matter), a policy table (policy) and a service table (service), are configured in the MongoDB database.
[0138] The Chinese word segmenter configuration module 200 is used to configure a Chinese word segmenter for each database table in the MongoDB database, and different database tables are configured with different Chinese word segmenters.
[0139] Preferably, the matter table (matter) is configured with a matrix constraint method Chinese word segmenter (A), the policy table (policy) is configured with a grammatical analysis method Chinese word segmenter (B), and the service table (service) is configured with a neural network method Chinese word segmenter (C).
[0140] The first word segmentation processing module 300 is used to perform first word segmentation processing on the same Chinese text stored under text field symbols of different database tables using corresponding Chinese word segmenters, and concatenate the Chinese text after the first word segmentation processing using spaces to obtain Chinese text fragments corresponding to different database tables.
[0141] The fragment text field symbol creation module 400 is used to create a fragment text field symbol for each database table in the MongoDB database; and store the fragment Chinese texts corresponding to different database tables in the fragment text field symbols of the corresponding database tables respectively.
[0142] The index creation module 500 is used to create a full-text index for each fragment text field symbol of each database table.
[0143] Preferably, the index creation module 500 creates a full-text index for each fragment text field symbol of each database table using an inverted index.
[0144] The second word segmentation processing module 600 is used to obtain the Chinese sentence to be queried; use different Chinese word segmenters configured with different database tables to perform second word segmentation processing on the Chinese sentence to be queried, and obtain keywords corresponding to different database tables.
[0145] The search module 700 is used to use full-text indexing to search for Chinese text fragments corresponding to different database tables in the MongoDB database, thereby completing Chinese sentence search.
[0146] According to an embodiment of the present invention, a computer storage medium is provided for storing computer-executable instructions. The computer-executable instructions are used to execute a Chinese sentence search method based on a MongoDB database provided by the present invention.
[0147] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A Chinese sentence search method based on MongoDB database, characterized in that: The Chinese sentence search method comprises the following steps: S1. Configure multiple database tables in a MongoDB database; wherein each database table includes a text field symbol, and the same Chinese text is stored under the text field symbols of different database tables; S2. Configure a Chinese word segmenter for each database table in the MongoDB database. Different database tables are configured with different Chinese word segmenters. S3, using the corresponding Chinese word segmenter to perform first word segmentation processing on the same Chinese text stored under the text field symbols of different database tables, and concatenating the Chinese text after the first word segmentation processing using spaces to obtain segmented Chinese texts corresponding to different database tables; S4. For each database table in the MongoDB database, create a fragment text field symbol; store the fragment Chinese texts corresponding to different database tables in the fragment text field symbol of the corresponding database table respectively; S5. Create a full-text index for each fragment text field in each database table; S6, obtaining the Chinese sentence to be queried; Using different Chinese word segmenters configured with different database tables, the second word segmentation process is performed on the Chinese sentence to be queried to obtain keywords corresponding to different database tables; S7. Use full-text indexing to search for Chinese text fragments corresponding to different database tables in the MongoDB database, and complete the Chinese sentence search.
2. The Chinese sentence search method according to claim 1, characterized in that: In step S1, three database tables, namely, an event table, a policy table and a service table, are configured in the MongoDB database.
3. The Chinese sentence search method according to claim 2, characterized in that: The matters table is configured with a matrix constraint Chinese word segmenter, the policy table is configured with a syntax analysis Chinese word segmenter, and the services table is configured with a neural network Chinese word segmenter.
4. The Chinese sentence search method according to claim 1, characterized in that: In step S5, a full-text index is created for each fragment text field symbol of each database table, using an inverted index.
5. A Chinese sentence search system based on MongoDB database, characterized in that: The Chinese sentence search system comprises: The database table configuration module is used to configure multiple database tables in the MongoDB database; each database table includes a text field symbol, and the same Chinese text is stored under the text field symbols of different database tables; The Chinese word segmenter configuration module is used to configure a Chinese word segmenter for each database table in the MongoDB database. Different database tables are configured with different Chinese word segmenters. A first word segmentation processing module is used to perform first word segmentation processing on the same Chinese text stored in the text field symbols of different database tables using the corresponding Chinese word segmenter, and to concatenate the Chinese text after the first word segmentation processing using spaces to obtain segmented Chinese texts corresponding to different database tables; The fragment text field symbol creation module is used to create a fragment text field symbol for each database table in the MongoDB database; the fragment Chinese texts corresponding to different database tables are stored in the fragment text field symbols of the corresponding database tables respectively; The index creation module is used to create a full-text index for each fragment text field in each database table; The second word segmentation processing module is used to obtain the Chinese sentence to be queried; different Chinese word segmenters configured with different database tables are used to perform the second word segmentation processing on the Chinese sentence to be queried to obtain keywords corresponding to different database tables; The search module is used to use full-text indexing to search for Chinese text fragments corresponding to different database tables in the MongoDB database, thereby completing Chinese sentence searches.
6. The Chinese sentence search system according to claim 5, characterized in that: In the library table configuration module, there are three database tables in the MongoDB database: configuration item table, policy table, and service table.
7. The Chinese sentence search system according to claim 6, characterized in that: The matters table is configured with a matrix constraint Chinese word segmenter, the policy table is configured with a syntax analysis Chinese word segmenter, and the services table is configured with a neural network Chinese word segmenter.
8. The Chinese sentence search system according to claim 5, characterized in that: The index creation module creates a full-text index for the fragment text field of each database table, using an inverted index.
9. A computer storage medium, characterized in that: The computer storage medium is used to store computer executable instructions, and the computer executable instructions are used to execute the Chinese sentence search method described in any one of claims 1 to 4.