Document descriptor table construction method and device, electronic equipment, and storage medium

By acquiring literature data in real time to construct thesaurus data tuples, constructing institutional tuples according to business needs, and using a standard library for matching and processing, the problems of high cost, low efficiency and fixed dimensions in thesaurus construction are solved, and efficient and personalized literature thesaurus generation is achieved.

CN118153560BActive Publication Date: 2025-12-09AGRI INFORMATION INST OF CHINESE ACAD OF AGRI SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410069911.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-12-09
Estimated Expiration
2044-01-17

AI Technical Summary

Technical Problem

In existing technologies, the thesaurus construction methods cannot logically match the business needs of the research field, resulting in high construction costs, low efficiency and high error rates. Furthermore, the fixed dimensions of the thesaurus cannot meet the personalized needs of users.

Method used

By acquiring literature input data in real time, the thesaurus data tuples are constructed, and institutional tuples are constructed according to business needs. The institutional tuple standard library is used for matching and processing to generate literature thesauruses corresponding to multiple research fields, realizing automatic classification and dynamic dimension adjustment of metadata.

Benefits of technology

It reduces the cost and error rate of thesaurus construction, improves efficiency, meets users' personalized needs, and achieves logical matching based on research fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118153560B_ABST
    Figure CN118153560B_ABST
Patent Text Reader

Abstract

The application provides a literature narrative thesaurus construction method, device, electronic equipment and storage medium, the literature narrative thesaurus construction method includes real-time acquisition literature input data;According to the literature input data of real-time acquisition constructs narrative thesaurus data tuple;According to narrative thesaurus data tuple constructs and business demand dimension corresponding organization multiple tuple;The organization multiple tuple is matched with organization tuple standard library, obtains the matching tuple corresponding to multiple research fields;The matching tuple corresponding to multiple research fields is processed, and each processed matching tuple is used as a row of data in the corresponding field literature narrative thesaurus, and multiple research field corresponding literature narrative thesaurus is generated according to the multiple matching tuples after processing.The application can be logically matched according to the business needs of research field, reduce the construction cost and error rate of literature narrative thesaurus, improve the construction efficiency, and the dimension of literature narrative thesaurus is dynamically adjusted according to the business needs of tuple dimension, meet the personalized needs of users.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a literature narrative term table construction method and device, electronic equipment and storage medium. BACKGROUND

[0002] The narrative term table is a standardized dynamic vocabulary table that displays narrative terms and semantic relationships between narrative terms. It contains many vocabulary related in semantics and hierarchical relationships in a specific licensed field. From a functional aspect, the narrative term table is a bridge between the thoughts of literature indexing personnel and retrieval personnel, a term control tool for converting between natural language (language used in literature) and system language (standardized language of the retrieval system), and a medium for communication between people and systems. Scientific and technological literature resources contain a large amount of information and knowledge, and are an important knowledge base. Therefore, constructing a narrative term table for scientific and technological literature is conducive to the dissemination of information and knowledge. Using a standardized controlled literature narrative term table in the indexing and retrieval process can effectively improve the accuracy of literature retrieval. Since the narrative term table corresponding to each research field may differ, for example, research institutions in the agricultural research field and the computer research field may differ greatly, in related technologies, the same narrative term table is used for literature in different research fields, which cannot be logically matched according to the business requirements of the research field. Since the forms of presentation of the author's institution in scientific and technological literature in different research fields differ, for example, the English capitalization of the Chinese name of the institution differs, it is necessary to manually determine the classification of the metadata of the institution of the author of each literature, and construct a narrative term table according to the classification results. This method of constructing a literature narrative term table has high cost, low efficiency, and a high error rate. Moreover, the dimensions of the existing narrative term table are fixed and cannot be dynamically adjusted according to the business requirements of the tuple dimensions, which cannot meet the individual needs of users. SUMMARY

[0003] The present application provides a literature narrative term table construction method, device, electronic equipment and storage medium to solve the defects that the traditional literature narrative term table construction method cannot be logically matched according to the business requirements of the research field, the method of constructing a literature narrative term table has high cost, low efficiency and a high error rate, and the dimensions of the narrative term table are fixed and cannot meet the individual needs of users.

[0004] The present application provides a literature narrative term table construction method, comprising:

[0005] real-time acquisition of literature input data;

[0006] constructing a narrative term table data tuple according to the real-time acquired literature input data;

[0007] constructing an institution multi-tuple corresponding to the business requirement dimensions according to the narrative term table data tuple;

[0008] Match the institution tuple with the institution tuple standard library to obtain a plurality of matched tuples corresponding to research fields, wherein the institution tuple standard library is constructed according to historical literature data in the field to which the literature belongs.

[0009] Process the plurality of matched tuples corresponding to the research fields, take each processed matched tuple as a row of data in a literature subject heading table corresponding to the field, and generate a plurality of literature subject heading tables corresponding to the research fields according to the plurality of processed matched tuples.

[0010] According to the literature subject heading table construction method provided by the application, the construction of the subject heading table data tuple according to the real-time obtained literature input data comprises:

[0011] A plurality of fields are extracted from the real-time obtained literature, and the fields comprise an institution, an institute, a department, a city, a zip code and a country.

[0012] The subject heading table data tuple is constructed according to the plurality of fields.

[0013] According to the literature subject heading table construction method provided by the application, the construction of the institution tuple corresponding to the business requirement dimension according to the subject heading table data tuple comprises:

[0014] The institution field is cleaned to remove the institution prefix in the institution field.

[0015] The country code prefix in the country field is identified.

[0016] The institution tuple is constructed according to the cleaned institution field and the country code prefix.

[0017] According to the literature subject heading table construction method provided by the application, the institution tuple comprises an institution and a country, and the institution comprises at least one level of unit.

[0018] The number of unit levels included in the institution tuple is determined according to the business requirement dimension.

[0019] According to the literature subject heading table construction method provided by the application, the matching of the institution tuple with the institution tuple standard library to obtain a plurality of matched tuples corresponding to research fields comprises:

[0020] The first key value in the institution tuple is matched with the first key value of the institution-country binary tuple in the institution-country tuple standard library of the corresponding research field.

[0021] If the matching is successful, the second key value in the institution tuple is matched with the second key value of the institution-country binary tuple in the institution-country tuple standard library of the corresponding research field.

[0022] If the matching is successful, the first key value and the second key value in the agency tuple are constructed into a two-dimensional matching tuple corresponding to the research field.

[0023] According to the document narrative table construction method provided by the application, the processing of the matching tuples corresponding to the plurality of research fields comprises:

[0024] A one-dimensional standard symbol is added at the first key value of the first dimension of the two-dimensional matching tuple.

[0025] A two-dimensional standard symbol is added at the first key value of the second dimension of the two-dimensional matching tuple.

[0026] According to the document narrative table construction method provided by the application, the agency tuple standard library is constructed according to historical document data in the field to which the document belongs.

[0027] The historical document data is classified according to research fields.

[0028] The agency and country information is screened from the historical document data in each research field, and an agency-country binary tuple is constructed.

[0029] The agency-country binary tuple is matched with the historical document data in the corresponding research field.

[0030] The agency and country information in each historical document data is attributed to the corresponding agency-country binary tuple, and an agency tuple standard library is obtained.

[0031] The application further provides a document narrative table construction device, comprising:

[0032] The acquisition module is configured to acquire document input data in real time.

[0033] The first construction module is configured to construct a narrative table data tuple according to the document input data acquired in real time.

[0034] The second construction module is configured to construct an agency multi-tuple corresponding to a business requirement dimension according to the narrative table data tuple.

[0035] The matching module is configured to match the agency multi-tuple with the agency tuple standard library to obtain matching tuples corresponding to a plurality of research fields, wherein the agency tuple standard library is constructed according to historical document data in the field to which the document belongs.

[0036] The processing module is configured to process the matching tuples corresponding to the plurality of research fields, and each processed matching tuple is taken as a row of data in a corresponding field document narrative table, and a plurality of research field document narrative tables corresponding to the processed matching tuples are generated.

[0037] The application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the document narrative table construction method according to any one of the above when executing the program.

[0038] The application further provides a non-transitory computer readable storage medium, which stores a computer program, wherein the computer program is executable by a processor to implement the document narrative table construction method according to any one of the above.

[0039] The application provides a document narrative table construction method and device, an electronic device, and a storage medium. The document input data is acquired in real time. The narrative table data tuple is constructed according to the document input data acquired in real time. The organization multi-tuple corresponding to the business demand dimension is constructed according to the narrative table data tuple. The organization multi-tuple is matched with the organization tuple standard library to obtain the matching tuples corresponding to multiple research fields, wherein the organization tuple standard library is constructed according to the historical document data in the field to which the document belongs. The matching tuples corresponding to the multiple research fields are processed, each processed matching tuple is taken as a row of data in the field document narrative table, and the multiple research field document narrative tables corresponding to the multiple processed matching tuples are generated. The logical matching can be performed according to the business demand of the research field, the metadata automatic classification is realized, the construction cost and error rate of the document narrative table are reduced, the construction efficiency is improved, the dimension of the document narrative table can be dynamically adjusted according to the business demand of the tuple dimension, and the personalized demand of the user is met. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0041] Figure 1 is one of the flowcharts of the document narrative table construction method provided by the present application;

[0042] Figure 2 is the second flowchart of the document narrative table construction method provided by the present application;

[0043] Figure 3 is the functional structure diagram of the document narrative table construction device provided by the present application;

[0044] Figure 4 is the structure diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0045] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0046] Figure 1 The flowchart of the literature narrative word table construction method provided by the embodiments of the present application is shown in FIG. 1, and the literature narrative word table construction method provided by the embodiments of the present application comprises the following steps. Figure 1

[0047] Step 101, acquiring literature input data in real time;

[0048] Step 102, constructing a narrative word table data tuple according to the literature input data acquired in real time;

[0049] Step 103, constructing an organization multi-tuple corresponding to a business demand dimension according to the narrative word table data tuple;

[0050] Step 104, matching the organization multi-tuple with an organization tuple standard library to obtain a plurality of matching tuples corresponding to research fields, wherein the organization tuple standard library is constructed according to historical literature data in the field to which the literature belongs;

[0051] Step 105, processing the plurality of matching tuples corresponding to the research fields, taking each processed matching tuple as a row of data in a literature narrative word table in a corresponding field, and generating a plurality of literature narrative word tables corresponding to the research fields according to the plurality of processed matching tuples.

[0052] In the traditional literature narrative word table construction method, different research field literatures share the same narrative word table, and logical matching cannot be performed according to the business demand of the research field. Since the display forms of the publishing organizations in the scientific literatures of different research fields are different, for example, the English capitalization corresponding to the Chinese organization name is different, it is necessary to manually determine the classification of the metadata of the organization of the author of each literature, and construct the narrative word table according to the classification result. This method of constructing the literature narrative word table has high cost, low efficiency and high error rate. Moreover, the dimensions of the existing narrative word table are fixed, and cannot be dynamically adjusted according to the business demand of the tuple dimensions, and cannot meet the individual needs of users.

[0053] ​The application provides a document narrative table construction method, which comprises the following steps: acquiring document input data in real time; constructing a narrative table data tuple according to the document input data acquired in real time; constructing an organization multi-tuple corresponding to a business demand dimension according to the narrative table data tuple; and matching the organization multi-tuple with an organization tuple standard library to obtain a plurality of matching tuples corresponding to research fields, wherein the organization tuple standard library is constructed according to historical document data in a field to which the document belongs. The matching tuples corresponding to the research fields are processed, each processed matching tuple is taken as a row of data in a document narrative table in a corresponding field, and a plurality of document narrative tables corresponding to the research fields are generated according to the processed matching tuples, so that logical matching can be performed according to business demands of the research fields, metadata automatic classification is realized, the cost and error rate of constructing the document narrative table are reduced, and the construction efficiency is improved. In addition, the dimension of the document narrative table can be dynamically adjusted according to business demands of the tuple dimension, so as to meet personalized demands of users.

[0054] Based on any of the above embodiments, in order to construct an organization and fund cleaning narrative table for a research field, a scientific researcher needs to collect narrative table metadata including organizations and funds as document input data, and construct a narrative table data tuple according to the document input data acquired in real time, which comprises the following steps:

[0055] Step 201: extracting a plurality of fields from the document acquired in real time, wherein the fields include organizations, research institutes, departments, cities, zip codes and countries.

[0056] Step 202: constructing a narrative table data tuple according to the plurality of fields.

[0057] In the embodiment of the application, the collected metadata fields include organizations (universities), research institutes (colleges), departments (key laboratories), cities and zip codes, and countries. A classification with organizations (universities) and countries as cores is established according to the collected metadata, and the entries of the collected metadata are matched into the respective corresponding classifications according to the corresponding relationship between the organizations (universities) and the countries.

[0058] In the embodiment of the application, the data tuple supports functions such as data addition, deletion, search and sorting, and can be managed according to data business demands, and can extract and export different types of input data and output data.

[0059] In the embodiment of the application, since the document data is acquired in real time, the output document narrative table also changes with the document input data, and the document narrative table can be dynamically constructed.

[0060] Based on any of the above embodiments, constructing an organization multi-tuple corresponding to a business demand dimension according to the narrative table data tuple comprises the following steps:

[0061] Step 301, cleaning the organization field to remove the organization prefix in the organization field;

[0062] Step 302, identifying the country code prefix in the country field;

[0063] Step 303, constructing the organization tuple according to the cleaned organization field and the country code prefix.

[0064] In the embodiment of the application, the organization tuple includes an organization and a country, and the organization includes at least one level of unit: the number of unit levels included in the organization tuple is determined according to the business demand dimension. For example, the organization includes a first-level unit, a second-level unit, and a third-level unit; the business demand dimension is 2, and the number of unit levels in the organization tuple is 2, for example, the organization tuple is (first-level unit, second-level unit, country).

[0065] Based on any of the above embodiments, the organization tuple is matched with the organization tuple standard library to obtain a plurality of matching tuples corresponding to the research fields, including:

[0066] Step 401, matching the first key value in the organization tuple with the first key value of the organization-country binary tuple in the organization-country tuple standard library corresponding to the research field;

[0067] Step 402, if the matching is successful, matching the second key value in the organization tuple with the second key value of the organization-country binary tuple in the organization-country tuple standard library corresponding to the research field;

[0068] In the embodiment of the application, the first key value in the organization tuple is matched with the first key value of the organization-country binary tuple in the organization-country tuple standard library corresponding to the research field, if the matching fails, the metadata tuple in the organization tuple is returned for the next matching; if the matching is successful, the metadata tuple of the organization tuple and the second key value of the organization-country binary tuple are matched, if the matching is successful, the second key value in the organization tuple is matched with the second key value of the organization-country binary tuple in the organization-country tuple standard library corresponding to the research field; if the matching fails, the metadata tuple in the organization tuple is returned for the next matching, and after all the key values of each organization-country binary tuple are matched, the matching tuple is constructed, and the literature keyword table is generated after accurate processing.

[0069] In the embodiment of the application, the organization tuple standard library is constructed according to historical literature data in the field to which the literature belongs, including:

[0070] Step 4021, classifying the historical literature data according to the research field;

[0071] Step 4022, screening the organization and country information from the historical literature data in each research field to construct the organization-country binary tuple;

[0072] Step 4023, matching the institution-country binary tuple with the historical literature data of the corresponding research field;

[0073] Step 4024, attributing the institution and country information in each historical literature data to the corresponding institution-country binary tuple to obtain an institution tuple standard library.

[0074] In the embodiment of the application, the institution-country binary tuple is constructed by iterating the metadata and screening the institution-country appearing in the source metadata, such as {Univ Massachusetts, USA}, and then the institution-country binary tuple is matched with the metadata to iterate the metadata entries including the institution-country, which are matched to the corresponding institution-country binary tuple. After iterating all the metadata, the matched tuple standard library is obtained. For example, the entry of Nanchang Univ, State Key Lab Food Sci&Technol, Nanchang 330047, Jiangxi, The People’s Republic of China contains the institution of Nanchang Univ and the country of The People’s Republic of China. The entry is matched to {Nanchang Univ, The People’s Republic of China} to obtain the standard library of [Nanchang Univ, The People’s Republic of China] {(Nanchang Univ, State Key Lab Food Sci&Technol, Nanchang 330047, Jiangxi, The People’s Republic of China)}.

[0075] Step 403, if the matching is successful, the first key value and the second key value in the institution multi-tuple are constructed to form a two-dimensional matching tuple of the corresponding research field.

[0076] In the embodiment of the application, the English of the institution in the literature is constructed into a thesaurus, and the two-dimensional matching tuple is, for example:

[0077] [Univ Massachusetts, USA]

[0078] {

[0079] (Univ Massachusetts, Dept Food Sci, Amherst, MA 01003 USA),

[0080] (Univ Massachusetts, Stockbridge Sch Agr, Amherst, MA 01003 USA),

[0081]

[0082] (Univ Massachusetts, Dept Food Sci, Biopolymers&Colloids Lab,Amherst, MA 01003 USA)

[0083] }

[0084] or

[0085] [Nanchang Univ, The People’s Republic of China]

[0086] {

[0087] (Nanchang Univ, State Key Lab Food Sci&Technol, Nanchang 330047,Jiangxi, The People’s Republic of China),

[0088] (Nanchang Univ, State Key Lab Food Sci&Technol, 235 Nanjing East Rd,Nanchang 330047, Jiangxi, The People’s Republic of China),

[0089]

[0090] (Nanchang Univ, State Key Lab Food Sci&Technol, Nanchang 330047, ThePeople’s Republic of China)

[0091] }

[0092] Based on any of the above embodiments, the matching tuples corresponding to a plurality of research fields are processed, including:

[0093] A one-dimensional standard symbol is added at the first key value of the first dimension of the two-dimensional matching tuple;

[0094] A two-dimensional standard symbol is added at the first key value of the second dimension of the two-dimensional matching tuple.

[0095] Based on any of the above embodiments, as shown in the document narrative table construction method includes: Figure 2

[0096] Initialize metadata tuples: construct metadata tuples according to original data of the document, and the tuple fields are constructed as institutions (universities), research institutes (colleges), departments (key laboratories), cities and zip codes, and countries.

[0097] Clean metadata tuples: clean the institution field and the country field, identify the prefixes of the institutions and the prefixes of the country codes, and clean them.

[0098] Constructing institution-country tuples: dynamically constructing institution-country two-tuples according to the cleaned metadata tuples.

[0099] Constructing matching tuples: matching according to the first key value of the metadata tuples and the institution-country two-tuples. If the matching fails, the metadata tuples are returned for the next matching; if the matching succeeds, the metadata tuples are matched according to the second key value of the metadata tuples and the institution-country two-tuples, and if the matching succeeds, the metadata tuples are added to the corresponding key position of the two-tuples to construct matching tuples; if the matching fails, the metadata tuples are returned for the next matching. After all the key values of each institution-country two-tuple are matched, the matching tuples are constructed.

[0100] Processing matching tuples: generating a standard library of matching tuples, adding a one-dimensional standard symbol to the first key value of the first dimension of the two-dimensional matching tuples, and adding a two-dimensional standard symbol to the first key value of the first key value of the second dimension of the two-dimensional matching tuples, to complete the processing of the matching tuples.

[0101] In the embodiments of the present application, the one-dimensional standard symbol is, for example, **, and the two-dimensional standard symbol is, for example, 0 1 ^ and $; the one-dimensional standard symbol and the one-dimensional standard symbol are used to clean the narrative table, and the processed matching tuples are stored in a.txt file, and the.txt file data is, for example:

[0102] **Univ Massachusetts, USA

[0103] 0 1 ^Univ Massachusetts, Dept Food Sci, Amherst, MA 01003 USA$

[0104] 0 1 ^Univ Massachusetts, Dept Food Sci, Amherst, MA 01003 USA.$

[0105] ​01 Univ Massachusetts, Stockbridge Sch Agr, Amherst, MA 01003 USA

[0106] 01 Univ Massachusetts, Dept Food Sci, Biopolymers & Colloids Lab, Amherst, MA 01003 USA

[0107] Nanchang Univ, The People's Republic of China

[0108] 01 Nanchang Univ, State Key Lab Food Sci & Technol, Nanchang 330047, Jiangxi, The People's Republic of China

[0109] 01 Nanchang Univ, State Key Lab Food Sci & Technol, 235 Nanjing East Rd, Nanchang 330047, Jiangxi, The People's Republic of China

[0110] 01 Nanchang Univ, State Key Lab Food Sci & Technol, Nanchang 330047, The People's Republic of China

[0111] 01 Nanchang Univ, State Key Lab Food Sci & Technol, Nanchang, Jiangxi, The People's Republic of China

[0112] 01 Nanchang Univ, Sch Life Sci, Nanchang 330031, Jiangxi, The People's Republic of China

[0113] 0 1 ^Nanchang Univ, Coll Food Sci, State Key Lab Food Sci&Technol,Nanchang 330047, Jiangxi, The People’s Republic of China$

[0114] The literature narrative word table construction method provided by the embodiment of the application can realize automatic and dynamic establishment of a logical corresponding matching relationship between a metadata tuple and a business demand tuple, data matching and data processing of an institution country array and a metadata tuple, and full-process manual intervention is not required, so that the cost is reduced, the efficiency is improved, and the error rate is greatly reduced.

[0115] The literature narrative word table construction device provided by the application is described below, and the literature narrative word table construction device described below can be correspondingly referred to the literature narrative word table construction method described above.

[0116] Figure 3 A schematic diagram of the literature narrative word table construction device provided by the embodiment of the application is shown in Figure 3 The literature narrative word table construction device provided by the embodiment of the application comprises:

[0117] The acquisition module 301 is configured to acquire literature input data in real time.

[0118] The first construction module 302 is configured to construct a narrative word table data tuple according to the literature input data acquired in real time.

[0119] The second construction module 303 is configured to construct an institution multi-tuple corresponding to a business demand dimension according to the narrative word table data tuple.

[0120] The matching module 304 is configured to match the institution multi-tuple with an institution tuple standard library to obtain a plurality of matching tuples corresponding to research fields, wherein the institution tuple standard library is constructed according to historical literature data in a field to which the literature belongs.

[0121] The processing module 305 is configured to process the plurality of matching tuples corresponding to the research fields, take each processed matching tuple as a row of data in a corresponding field literature narrative word table, and generate a plurality of literature narrative word tables corresponding to the research fields according to the plurality of processed matching tuples.

[0122] This invention provides a thesaurus construction device that acquires document input data in real time; constructs thesaurus data tuples based on the acquired data; constructs institutional tuples corresponding to business requirements based on the thesaurus data tuples; matches the institutional tuples with an institutional tuple standard library to obtain matching tuples corresponding to multiple research fields, wherein the institutional tuple standard library is constructed based on historical document data in the field to which the document belongs; processes the matching tuples corresponding to multiple research fields, using each processed matching tuple as a row in the thesaurus of the corresponding field; and generates multiple thesauruses corresponding to multiple research fields based on the processed matching tuples. This device allows for logical matching based on the business requirements of the research fields, enabling automatic metadata classification, reducing the cost and error rate of thesaurus construction, improving construction efficiency, and dynamically adjusting the dimensions of the thesaurus according to the business requirements of the tuple dimensions to meet personalized user needs.

[0123] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440. The processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a thesaurus construction method. This method includes: acquiring document input data in real time; constructing thesaurus data tuples based on the acquired document input data; constructing institutional tuples corresponding to business requirements dimensions based on the thesaurus data tuples; matching the institutional tuples with an institutional tuple standard library to obtain matching tuples corresponding to multiple research fields, wherein the institutional tuple standard library is constructed based on historical document data in the field to which the document belongs; processing the matching tuples corresponding to multiple research fields, using each processed matching tuple as a row in the thesaurus of the corresponding field, and generating multiple thesauruses corresponding to multiple research fields based on the processed matching tuples.

[0124] In addition, the logic instructions in the memory 430 described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0125] In another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the literature keyword table construction method provided by the above method, the method comprising: acquiring literature input data in real time; constructing a keyword table data tuple according to the literature input data acquired in real time; constructing an organization multi-tuple corresponding to the business demand dimension according to the keyword table data tuple; matching the organization multi-tuple with an organization tuple standard library to obtain a plurality of matching tuples corresponding to research fields, wherein the organization tuple standard library is constructed according to historical literature data in the field to which the literature belongs; processing the plurality of matching tuples corresponding to the research fields, taking each processed matching tuple as a row of data in a literature keyword table of a corresponding field, and generating a plurality of literature keyword tables corresponding to the research fields according to the plurality of processed matching tuples.

[0126] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement without creative labor.

[0127] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of the various embodiments or some parts of the embodiments.

[0128] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method of constructing a document thesaurus, characterized by, The method comprises the following steps: real-time acquisition of literature input data; constructing a thesaurus data tuple according to the real-time acquired literature input data; constructing an institution multi-tuple corresponding to a business demand dimension according to the thesaurus data tuple; matching the institution multi-tuple with an institution tuple standard library to obtain a plurality of matching tuples corresponding to research fields, wherein the institution tuple standard library is constructed according to historical literature data in a field to which the literature belongs; and the matching of the institution multi-tuple with the institution tuple standard library to obtain the plurality of matching tuples corresponding to the research fields comprises: matching a first key value in the institution multi-tuple with a first key value of an institution-country binary tuple in an institution-country tuple standard library corresponding to a research field; if the matching is successful, matching a second key value in the institution multi-tuple with a second key value of the institution-country binary tuple in the institution-country tuple standard library corresponding to the research field; and if the matching is successful, constructing a two-dimensional matching tuple corresponding to the research field from the first key value and the second key value in the institution multi-tuple; processing the plurality of matching tuples corresponding to the research fields, taking each processed matching tuple as a row of data in a literature thesaurus in a corresponding field, and generating a plurality of literature thesauruses corresponding to the research fields according to the plurality of processed matching tuples; the processing of the plurality of matching tuples corresponding to the research fields comprises: adding a one-dimensional standard symbol at a first key value in a first dimension of a two-dimensional matching tuple; and adding a two-dimensional standard symbol at a second key value in a second dimension of the two-dimensional matching tuple.

2. The method of claim 1, wherein, the constructing of the thesaurus data tuple according to the real-time acquired literature input data comprises: extracting a plurality of fields from the real-time acquired literature, wherein the fields comprise an institution, an institute, a department, a city, a zip code and a country; constructing a thesaurus data tuple according to the plurality of fields.

3. The method of claim 2, wherein, the constructing of the institution multi-tuple corresponding to the business demand dimension according to the thesaurus data tuple comprises: cleaning the institution field to remove an institution prefix in the institution field; identifying a country code prefix in the country field; constructing an institution multi-tuple according to the cleaned institution field and the country code prefix.

4. The method of claim 1 or 3, wherein, the institution multi-tuple comprises an institution and a country, and the institution comprises at least one level of unit. the number of unit levels included in the institution multi-tuple is determined according to a business demand dimension.

5. The method of claim 1, wherein, the institution tuple standard library is constructed according to historical literature data in a field to which the literature belongs, and the constructing of the institution tuple standard library comprises: classifying the historical literature data according to research fields; screening institution and country information from the historical literature data in each research field to construct institution-country binary tuples; matching the institution-country binary tuples with historical literature data corresponding to the research fields; and attributing institution and country information in each piece of historical literature data to a corresponding institution-country binary tuple to obtain an institution tuple standard library.

6. A document thesaurus construction apparatus characterized by comprising: The method comprises the following steps: an acquisition module, configured to acquire literature input data in real time; a first construction module, configured to construct a thesaurus data tuple according to the real-time acquired literature input data; a second construction module, configured to construct an institution multi-tuple corresponding to a business demand dimension according to the thesaurus data tuple; and The matching module is configured to match the institution multi-tuple with an institution tuple standard library to obtain a plurality of matching tuples corresponding to research fields, wherein the institution tuple standard library is constructed according to historical literature data in the field to which the literature belongs; and the matching of the institution multi-tuple with the institution tuple standard library to obtain the plurality of matching tuples corresponding to the research fields comprises: matching a first key value in the institution multi-tuple with a first key value of an institution country two-tuple in an institution country tuple standard library corresponding to the research field; if the matching is successful, matching a second key value in the institution multi-tuple with a second key value of the institution country two-tuple in the institution country tuple standard library corresponding to the research field; and if the matching is successful, constructing a two-dimensional matching tuple corresponding to the research field from the first key value and the second key value in the institution multi-tuple; The processing module is configured to process the plurality of matching tuples corresponding to the research fields, take each processed matching tuple as a row of data in a literature subject heading table corresponding to the field, and generate a plurality of literature subject heading tables corresponding to the research fields according to the plurality of processed matching tuples; and the processing of the plurality of matching tuples corresponding to the research fields comprises: adding a one-dimensional standard symbol to a first key value in a first dimension of a two-dimensional matching tuple; and adding a two-dimensional standard symbol to a second key value in a second dimension of the two-dimensional matching tuple.

7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the literature subject heading table construction method according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the literature subject heading table construction method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for extracting author affiliation information of English literature published by Chinese authors

    CN104881398A

  • Intelligent indexing-oriented scientific and technical literature keyword indexing method

    CN115204160A