A method and system for constructing a domain thesaurus for information systems
By constructing a domain-specific thesaurus suitable for information system development, and combining it with enterprise characteristics and development processes, the problem that general-purpose thesaurus cannot support enterprise information system development has been solved, and efficient document analysis and knowledge management have been achieved.
Patent Information
- Application Number
- CN202411784256.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Existing general domain thesaurus cannot effectively support document library knowledge mining and asset retrieval in the field of enterprise information system R&D. It lacks understanding of enterprise-specific terms and contexts, resulting in low retrieval efficiency and accuracy, which affects knowledge management and innovation capabilities.
We construct a domain-specific thesaurus suitable for information system development. This involves extracting vocabulary from business domains and information system management systems, combining it with development entities and processes to establish a relationship network, and using word segmentation tools to segment and classify documents, and calculating word similarity.
It improved the practicality and accuracy of the domain thesaurus, optimized asset retrieval, and provided a solid data foundation for building an enterprise-level knowledge system.
Smart Images

Figure CN119808767B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of information system development and knowledge system construction. Specifically, it proposes a method and system for constructing a domain terminology database for information system development. By combining the enterprise's internal R&D processes, role-based document outputs, and the enterprise's own characteristics, a domain terminology database suitable for enterprise information system development is constructed, providing a data foundation for knowledge system construction and asset retrieval optimization. Background Technology
[0002] With the rapid development of information technology, industry-specific thesauruses are becoming increasingly rich, and their openness and sharing are also increasing. Simultaneously, technologies such as natural language processing and Chinese word segmentation are maturing, providing strong support for the construction and application of these thesauruses. However, most current open-source industry-specific thesauruses are general-purpose industry thesauruses, lacking consideration for specific enterprise-specific terminology and contexts. In the field of enterprise information system development, general-purpose thesauruses often cannot effectively support internal enterprise R&D document library knowledge mining, R&D asset retrieval, and asset data analysis.
[0003] Specifically, general-purpose thesaurus may not include enterprise-specific business terms, technical vocabulary, and organizational structure-related terms. This can lead to difficulties in accurately matching internal technical terms during document retrieval and knowledge mining, thus affecting retrieval efficiency and accuracy. Furthermore, the lack of understanding of the specific context of an enterprise in a general-purpose thesaurus may result in information omissions or misunderstandings during the knowledge system construction process, thereby impacting the enterprise's knowledge management and innovation capabilities.
[0004] To address the aforementioned issues, this invention proposes a novel method for constructing a domain-specific thesaurus. This method is specifically designed for the information system development field, combining the unique characteristics and practical needs of enterprises to build a more accurate and practical domain-specific thesaurus. This invention extracts and integrates vocabulary from multiple dimensions, including business domains, organizational structure, and system construction, to form a domain-specific thesaurus suitable for enterprise information system development, thereby providing a solid data foundation for subsequent knowledge system construction and asset retrieval optimization. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention proposes a method and system for constructing a domain thesaurus for information systems.
[0006] To achieve the above objectives, the technical solution adopted by the present invention includes:
[0007] A method for constructing a domain thesaurus for information systems, characterized by comprising the following steps:
[0008] S1. Construct a first thesaurus, which contains first words extracted from the business domain;
[0009] S2. Construct a second thesaurus, which contains a second vocabulary extracted from the information system management system;
[0010] S3. Expand the first lexicon based on the expanded vocabulary, wherein the expanded vocabulary is extracted based on the R&D entities in information system R&D, and the R&D entities are R&D personnel, R&D deliverables, and R&D organizations;
[0011] S4. Based on the information system development process, set corresponding tags for the development entities, and associate the first vocabulary and expanded vocabulary in the first thesaurus with the tags to establish a relationship network;
[0012] S5. Based on the relationship network, a word segmentation tool is used to segment the documents of the R&D entity to form a third lexicon. The word segmentation tool uses the first lexicon and the second lexicon as the basic dictionary.
[0013] S6. Calculate the similarity between the third vocabulary in the third vocabulary library and the first vocabulary and the expanded vocabulary, and classify the third vocabulary accordingly.
[0014] Furthermore, step S1 of constructing the first lexicon further includes:
[0015] S11. Extract keywords related to business processes from business domain documents;
[0016] S12. Categorize the extracted keywords according to the business process;
[0017] S13. Store the categorized keywords in the first thesaurus.
[0018] Furthermore, step S2 of constructing the second lexicon further includes:
[0019] S21. Extract keywords related to the management system from the information system management system documents;
[0020] S22. Categorize the extracted keywords according to the hierarchy and function of the management system;
[0021] S23. Store the categorized keywords in the second thesaurus.
[0022] Furthermore, step S3, which involves expanding the first lexicon based on the expanded vocabulary, further includes:
[0023] S31. Determine the relevant technical terms based on the role and responsibilities of the R&D entity;
[0024] S32. Extract terms related to technical terms from the outputs of the R&D entity;
[0025] S33. Compare the extracted words with the words in the first vocabulary database, remove duplicates, and expand the first vocabulary database.
[0026] Furthermore, step S4, which involves setting corresponding tags for the R&D entity based on the information system R&D process, further includes:
[0027] S41. Define the various stages in the R&D process;
[0028] S42. Identify the corresponding R&D entity for each stage;
[0029] S43. Assign a corresponding label to each R&D entity. The label indicates the role or type of output of the R&D entity at that stage.
[0030] Furthermore, step S5, which involves segmenting the R&D entity document using a word segmentation tool based on a relational network, further includes:
[0031] S51. Use natural language processing technology to preprocess the document;
[0032] S52. Use the first and second dictionaries as the basic dictionary to segment the document into words;
[0033] S53. Match the word segmentation results with the words in the relational network to form a third lexicon.
[0034] Furthermore, step S6, which calculates similarity and classifies based on third vocabulary in a third lexicon, further includes:
[0035] S61. Use similarity calculation algorithms, such as cosine similarity or Jaccard similarity, to calculate the similarity between the third word and the first word and the expanded words.
[0036] S62. Based on the calculated similarity, classify the third word into the category with the highest similarity to the first word or the expanded word.
[0037] Furthermore, the present invention also relates to a domain thesaurus construction system for information systems, characterized in that it comprises:
[0038] The first thesaurus construction module is used to extract the first vocabulary from the business domain and build the first thesaurus.
[0039] The second thesaurus construction module is used to extract second vocabulary from the information system management system and construct the second thesaurus.
[0040] An expanded vocabulary extraction module is used to extract expanded vocabulary based on R&D entities in information system R&D and expand the first vocabulary database. The R&D entities include R&D personnel, R&D deliverables, and R&D organizations.
[0041] The tag setting module is used to set corresponding tags for the R&D entity and associate the first vocabulary and expanded vocabulary in the first thesaurus with the tags to establish a relationship network.
[0042] The third lexicon construction module is used to call the word segmentation tool based on the relationship network to segment the documents of the R&D entity and form the third lexicon. The word segmentation tool uses the first lexicon and the second lexicon as the basic dictionary.
[0043] The vocabulary classification module is used to calculate the similarity between the third vocabulary in the third vocabulary corpus and the first and expanded vocabulary, and to classify the third vocabulary according to the similarity.
[0044] Furthermore, the present invention also relates to an electronic device, characterized in that it includes a processor and a memory;
[0045] The memory is used to store operation instructions;
[0046] The processor is configured to execute the above-described method by invoking the operation instructions.
[0047] Furthermore, the present invention also relates to a computer-readable storage medium, characterized in that the storage medium stores a computer program, which, when executed by a processor, implements the above-described method.
[0048] This invention discloses a method and system for constructing a domain lexicon for information systems, aiming to build a domain lexicon suitable for information system development by combining enterprise characteristics and R&D processes. The method includes the following steps: extracting first terms from the business domain to construct a first lexicon, and extracting second terms from the information system management system to construct a second lexicon; extracting expanded terms to expand the first lexicon based on R&D entities in information system development, such as R&D personnel, R&D deliverables, and R&D organizations; setting corresponding tags for R&D entities and associating the first terms, expanded terms, and tags to establish a relationship network; using a word segmentation tool to segment the documents of the R&D entities to form a third lexicon, with the word segmentation tool using the first and second lexicons as its base dictionary; and classifying the terms in the third lexicon based on their similarity to the first and expanded terms. Therefore, this technical solution can improve the practicality and accuracy of the domain lexicon, optimize asset retrieval, and provide a data foundation for the construction of an enterprise-level knowledge system. Attached Figure Description
[0049] Figure 1 A flowchart illustrating a method for constructing a domain thesaurus for an information system, provided as an embodiment of this application;
[0050] Figure 2 A schematic diagram of the structure of a domain thesaurus construction system for information systems provided in this application embodiment;
[0051] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0052] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting the invention.
[0053] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. The terms “first,” “second,” etc., are merely for clarification of the subject matter and do not limit the subject matter itself. Of course, the subjects defined by “first” and “second” may be the same terminal, device, and user, or the same type of terminal, device, and user. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The term “and / or” as used herein includes all or any unit and all combinations of one or more associated listed items.
[0054] The technical solutions of this application and how the technical solutions of this application solve the above-mentioned technical problems are described in detail below with specific embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0055] like Figure 1 As shown, a method for constructing a domain thesaurus for information systems is characterized by comprising:
[0056] S1. Construct a first thesaurus, which contains first words extracted from the business domain;
[0057] Specifically, keywords related to business processes are extracted from banking business documents. For example, for the deposit business process, keywords such as "deposit," "account," and "interest rate" are extracted. Next, these keywords are categorized according to the business process; for example, "deposit" is categorized under the deposit business process. Finally, the categorized keywords are stored in a primary thesaurus, forming the basic vocabulary for the business domain.
[0058] We extract vocabulary based on business domains to ensure the professionalism and relevance of the vocabulary, providing high-quality foundational data for subsequent knowledge mining and asset retrieval.
[0059] S2. Construct a second thesaurus, which contains a second vocabulary extracted from the information system management system;
[0060] Specifically, keywords related to the management system are extracted from the bank's information system management documents. For example, keywords related to IT service management such as "ITIL," "service desk," and "event management" are extracted. These keywords are then categorized according to the hierarchy and function of the management system; for example, "ITIL" is categorized under the IT service management framework. Finally, the categorized keywords are stored in a second thesaurus, forming the basic vocabulary of the management system.
[0061] Building a second thesaurus can enhance the professionalism of information system management, enabling relevant knowledge of the management system to be effectively retrieved and utilized, thereby improving the management efficiency of the information system.
[0062] S3. Expand the first lexicon based on the expanded vocabulary, wherein the expanded vocabulary is extracted based on the R&D entities in information system R&D, and the R&D entities are R&D personnel, R&D deliverables, and R&D organizations;
[0063] Specifically, based on the roles and responsibilities of the bank's R&D team, relevant technical terms are identified. For example, for R&D personnel, technical terms such as "API," "database," and "front-end development" are identified. Then, words related to these technical terms are extracted from the R&D personnel's outputs, such as "Java" and "SQL" from the codebase. Finally, these extracted words are compared with the words in the first thesaurus, and duplicates are removed before being added to the first thesaurus.
[0064] Expanding the primary thesaurus makes the vocabulary more comprehensive and richer, covering enterprise-specific technical terms and industry jargon, thereby improving the accuracy of document processing and knowledge retrieval.
[0065] S4. Based on the information system development process, set corresponding tags for the development entities, and associate the first vocabulary and expanded vocabulary in the first thesaurus with the tags to establish a relationship network;
[0066] Specifically, define the various stages in the development process of a bank information system, such as requirements analysis, system design, coding implementation, and testing and verification.
[0067] For each stage, a corresponding R&D entity is identified. For example, in the requirements analysis stage, the R&D entities include business analysts, project managers, and requirements engineers; in the system design stage, the R&D entities include system architects and design engineers; in the coding implementation stage, the R&D entities include developers and code reviewers; and in the testing and verification stage, the R&D entities include test engineers and quality assurance specialists.
[0068] Each R&D entity is assigned a corresponding label, which indicates the entity's role or deliverable type at that stage. For example, a business analyst is assigned the label "Requirements Analysis: Business Analyst"; a system architect is assigned the label "System Design: Architect"; a developer is assigned the label "Coding Implementation: Developer"; and a test engineer is assigned the label "Test Verification: Test Engineer".
[0069] Establish a relationship network, associating primary terms and expanded terms in the primary thesaurus through these tags. For example, if "deposit" is a primary term and "online deposit" is an expanded term, these two terms can be associated with the business analyst's tag "demand analysis: business analyst," because these terms are both related to the deposit business process, which the business analyst is responsible for defining during the requirements analysis phase.
[0070] By labeling R&D entities, the responsibilities and contributions of each role in the R&D process can be clearly identified. The relationship network reflects the association between words and the collaboration and process dependencies between R&D entities. The relationship network provides rich contextual information for subsequent word segmentation and vocabulary classification, thereby improving the accuracy and practicality of the domain thesaurus.
[0071] S5. Based on the relationship network, a word segmentation tool is used to segment the documents of the R&D entity to form a third lexicon. The word segmentation tool uses the first lexicon and the second lexicon as the basic dictionary.
[0072] Specifically, natural language processing technology is used to preprocess the documents of the bank's R&D team, such as removing stop words and performing part-of-speech tagging.
[0073] Using the first and second thesaurus as basic dictionaries, the document is segmented into words. The segmented words are then matched with words in the relational network to form a third thesaurus. For example, if the document mentions "deposit," it is matched with the word "deposit" in the first thesaurus, and the relevant information is stored in the third thesaurus.
[0074] By using word segmentation tools, the efficiency and accuracy of document analysis can be improved, enabling the rapid extraction and utilization of key content in information system development documents.
[0075] S6. Calculate the similarity between the third vocabulary in the third vocabulary library and the first vocabulary and the expanded vocabulary, and classify the third vocabulary accordingly.
[0076] Specifically, similarity calculation algorithms, such as cosine similarity or Jaccard similarity, are used to calculate the similarity between words in the third vocabulary and the first and expanded vocabulary. For example, the similarity between "online deposit" and "deposit" is calculated, and based on the calculated similarity, "online deposit" is categorized into the category with the highest similarity to "deposit".
[0077] The solution provided in this application extracts first terms from the business domain to construct a first thesaurus, and extracts second terms from the information system management system to construct a second thesaurus. Based on R&D entities in information system development, such as R&D personnel, R&D deliverables, and R&D organizations, expanded terms are extracted to expand the first thesaurus. Corresponding tags are set for R&D entities, and the first and expanded terms are associated with the tags to establish a relationship network. A word segmentation tool is used to segment the documents of the R&D entities, forming a third thesaurus, with the first and second thesaurus as the base dictionary. Classification is performed based on the similarity between the third terms in the third thesaurus and the first and expanded terms. Therefore, this technical solution can improve the practicality and accuracy of the domain thesaurus, optimize asset retrieval, and provide a data foundation for the construction of an enterprise-level knowledge system.
[0078] Furthermore, another aspect involves a domain thesaurus construction system for information systems, the structure of which is as follows: Figure 2 As shown, it includes:
[0079] The first thesaurus construction module 201 is used to extract first words from the business domain and construct the first thesaurus.
[0080] Specifically, extract keywords related to business processes from business domain documents;
[0081] The extracted keywords are categorized according to business processes;
[0082] The categorized keywords are stored in the first thesaurus.
[0083] The second thesaurus construction module 202 is used to extract second vocabulary from the information system management system and construct the second thesaurus.
[0084] Specifically, extract keywords related to the management system from the information system management system documents;
[0085] The extracted keywords are categorized according to the hierarchy and function of the management system;
[0086] The categorized keywords are stored in a second thesaurus.
[0087] The expanded vocabulary extraction module 203 is used to extract expanded vocabulary based on R&D entities in information system R&D and expand the first vocabulary library. The R&D entities include R&D personnel, R&D deliverables, and R&D organizations.
[0088] Specifically, based on the role and responsibilities of the R&D entity, determine the relevant technical terms;
[0089] Extract terms related to technical jargon from the outputs of R&D entities;
[0090] The extracted words are compared with the words in the first vocabulary database, and after deduplication, they are expanded into the first vocabulary database.
[0091] The tag setting module 204 is used to set corresponding tags for the R&D entity and associate the first vocabulary and expanded vocabulary in the first thesaurus with the tags to establish a relationship network.
[0092] Specifically, define each stage in the R&D process;
[0093] Determine the corresponding R&D entity for each stage;
[0094] Assign a corresponding label to each R&D entity. The label indicates the role or type of output of the R&D entity at that stage.
[0095] The third lexicon construction module 205 is used to call a word segmentation tool based on a relational network to segment the documents of the R&D entity and form a third lexicon. The word segmentation tool uses the first lexicon and the second lexicon as its basic dictionary.
[0096] Specifically, natural language processing techniques are used to preprocess the documents;
[0097] The document is segmented using the first and second dictionaries as the basic dictionary;
[0098] The word segmentation results are matched with words in the relational network to form a third lexicon.
[0099] The vocabulary classification module 206 is used to calculate the similarity between the third vocabulary in the third vocabulary library and the first vocabulary and the expanded vocabulary, and to classify the third vocabulary according to the similarity.
[0100] Specifically, similarity calculation algorithms, such as cosine similarity or Jaccard similarity, are used to calculate the similarity between the third word and the first word and the expanded words.
[0101] Based on the calculated similarity, the third word is categorized into the category with the highest similarity to the first word or the expanded word.
[0102] By using this system, the above methods can be executed and the corresponding technical effects can be achieved.
[0103] Embodiments of the present invention also provide an electronic device for performing the above-described method, which, as an implementation apparatus for the method, includes a processor and a memory;
[0104] Memory, used to store operation instructions;
[0105] The processor is configured to execute, by invoking operation instructions, a method for constructing a domain thesaurus for an information system provided in any embodiment of this application.
[0106] As an example, Figure 3 This diagram illustrates the structure of an electronic device to which this application applies. The electronic device 300 includes a processor 301 and a memory 303. The processor 301 and the memory 303 are connected, for example, via a bus 302. Optionally, the electronic device 300 may also include a transceiver 304. It should be noted that in practical applications, the transceiver 304 is not limited to one. It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the specific structure of the electronic device 300. In other embodiments of this application, the electronic device 300 may include more or fewer components than illustrated, or combine some components, or split some components, or arrange different components. The illustrated components may be implemented as hardware, software, or a combination of software and hardware. Optionally, the electronic device may also include a display screen 305 for displaying images or receiving user operation commands when needed.
[0107] In this embodiment, processor 301 is used to implement the method shown in the above method embodiment. Transceiver 304 may include a receiver and a transmitter. Transceiver 304 is used in this embodiment to enable the electronic device of this embodiment to communicate with other devices during execution.
[0108] Processor 301 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 301 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0109] Processor 301 may also include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units can be independent devices or integrated into one or more processors. The controller can be the central nervous system and command center of the electronic device 300. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. Processor 301 may also include memory for storing instructions and data. In some embodiments, the memory in processor 301 is a cache memory. This memory can store instructions or data that the processor 301 has recently used or is recurring.
[0110] The processor 301 can run the domain thesaurus construction method for information systems provided in the embodiments of this application. The processor 301 may include different devices, such as when integrating a CPU and a GPU, the CPU and GPU can cooperate to execute the domain thesaurus construction method for information systems provided in the embodiments of this application. Some algorithms are executed by the CPU and other algorithms are executed by the GPU to obtain faster processing efficiency.
[0111] Bus 302 may include a pathway for transmitting information between the aforementioned components. Bus 302 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 302 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0112] The memory 303 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory), or high-speed random access memory. It may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), or other optical disc storage, optical disk storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0113] Optionally, the memory 303 is used to store application code that executes the scheme of this application, and the execution is controlled by the processor 301. The processor 301 is used to execute the application code stored in the memory 303 to implement the domain thesaurus construction method for information systems provided in any embodiment of this application.
[0114] The memory 303 can be used to store computer executable program code, which includes instructions. The processor 301 executes various functional applications and data processing of the electronic device 300 by running the instructions stored in the memory 303. The memory 303 may include a program storage area and a data storage area. The program storage area can store the operating system, application code, etc. The data storage area can store data created during the use of the electronic device 300 (such as images and videos captured by a camera application).
[0115] The memory 303 may also store one or more computer programs corresponding to the domain terminology construction method for information systems provided in the embodiments of this application. These one or more computer programs are stored in the memory 303 and configured to be executed by the one or more processors 301. The one or more computer programs include instructions that can be used to perform the various steps in the corresponding embodiments described above.
[0116] Of course, the code for the domain thesaurus construction method for information systems provided in this application embodiment can also be stored in external storage.
[0117] The display screen 305 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a minimized LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 300 may include one or N displays 305, where N is a positive integer greater than 1. The display screen 305 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces (GUIs). For example, the display screen 305 can display photos, videos, web pages, or documents.
[0118] The electronic device provided in this application is applicable to any of the above-described methods. Therefore, the beneficial effects it can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0119] Embodiments of the present invention also provide a computer-readable storage medium capable of implementing all the steps of the methods in the above embodiments, wherein the computer-readable storage medium stores a computer program that, when executed by a processor, implements all the steps of the methods in the above embodiments.
[0120] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0121] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A system that specifies functions in one or more boxes.
[0122] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction set implemented in a process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0123] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes. Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.
[0124] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for constructing a domain thesaurus for information systems, characterized in that, include: S1. Construct a first thesaurus, which contains first words extracted from the business domain; S2. Construct a second thesaurus, which contains a second vocabulary extracted from the information system management system; S3. Expand the first lexicon based on the expanded vocabulary, wherein the expanded vocabulary is extracted based on the R&D entities in information system R&D, and the R&D entities are R&D personnel, R&D deliverables, and R&D organizations; S4. Based on the information system development process, set corresponding tags for the development entities, and associate the first vocabulary and expanded vocabulary in the first thesaurus with the tags to establish a relationship network; S5. Based on the relationship network, a word segmentation tool is used to segment the documents of the R&D entity to form a third lexicon. The word segmentation tool uses the first lexicon and the second lexicon as the basic dictionary. S6. Calculate the similarity between the third vocabulary in the third vocabulary library and the first vocabulary and the expanded vocabulary, and classify the third vocabulary accordingly; Wherein, step S3, which expands the first lexicon based on the expanded vocabulary, further includes: S31. Determine the relevant technical terms based on the role and responsibilities of the R&D entity; S32. Extract terms related to technical terms from the outputs of the R&D entity; S33. Compare the extracted words with the words in the first vocabulary database, remove duplicates, and expand the first vocabulary database; Step S4, which involves setting corresponding tags for the R&D entity based on the information system R&D process, further includes: S41. Define the various stages in the R&D process; S42. Identify the corresponding R&D entity for each stage; S43. Assign a corresponding label to each R&D entity, the label indicating the role or type of output of the R&D entity at this stage; Step S5, which uses a word segmentation tool based on a relational network to segment R&D entity documents, further includes: S51. Use natural language processing technology to preprocess the document; S52. Use the first and second dictionaries as the basic dictionary to segment the document into words; S53. Match the word segmentation results with the words in the relational network to form a third lexicon.
2. The method as described in claim 1, characterized in that, Step S1 of constructing the first lexicon further includes: S11. Extract keywords related to business processes from business domain documents; S12. Categorize the extracted keywords according to the business process; S13. Store the categorized keywords in the first thesaurus.
3. The method as described in claim 1, characterized in that, Step S2, which involves constructing the second lexicon, further includes: S21. Extract keywords related to the management system from the information system management system documents; S22. Categorize the extracted keywords according to the hierarchy and function of the management system; S23. Store the categorized keywords in the second thesaurus.
4. The method as described in claim 1, characterized in that, Step S6, which calculates similarity and classifies based on third vocabulary in a third lexicon, further includes: S61. Use similarity calculation algorithms, such as cosine similarity or Jaccard similarity, to calculate the similarity between the third word and the first word and the expanded words. S62. Based on the calculated similarity, classify the third word into the category with the highest similarity to the first word or the expanded word.
5. A domain thesaurus construction system for information systems, characterized in that, include: The first thesaurus construction module is used to extract the first vocabulary from the business domain and build the first thesaurus. The second thesaurus construction module is used to extract second vocabulary from the information system management system and construct the second thesaurus. An expanded vocabulary extraction module is used to extract expanded vocabulary based on R&D entities in information system R&D and expand the first vocabulary database. The R&D entities include R&D personnel, R&D deliverables, and R&D organizations. The tag setting module is used to set corresponding tags for the R&D entity and associate the first vocabulary and expanded vocabulary in the first thesaurus with the tags to establish a relationship network. The third lexicon construction module is used to call the word segmentation tool based on the relationship network to segment the documents of the R&D entity and form the third lexicon. The word segmentation tool uses the first lexicon and the second lexicon as the basic dictionary. The vocabulary classification module is used to calculate the similarity between the third vocabulary in the third vocabulary database and the first vocabulary and the expanded vocabulary, and to classify the third vocabulary according to the similarity. The expanded vocabulary extraction module is specifically used to extract vocabulary related to professional terms from the outputs of R&D entities; Extract terms related to technical jargon from the outputs of R&D entities; The extracted words are compared with the words in the first vocabulary database, and after deduplication, they are expanded into the first vocabulary database. The tag settings module is specifically used to define the various stages in the R&D process; Determine the corresponding R&D entity for each stage; Assign a corresponding label to each R&D entity, and the label indicates the role or type of output of the R&D entity at that stage; The third lexicon building module is specifically used to preprocess documents using natural language processing technology; The document is segmented using the first and second dictionaries as the basic dictionary; The word segmentation results are matched with words in the relational network to form a third lexicon.
6. An electronic device, characterized in that, Including processor and memory; The memory is used to store operation instructions; The processor is configured to execute the method of any one of claims 1-4 by invoking the operation instructions.
7. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method of any one of claims 1-4.
Citation Information
Patent Citations
Vertical industrial field entity dictionary construction method and device, equipment and storage medium
CN118070784A
Technical specification information dynamic configuration method and system based on scene recognition
CN118446183A