Method and device for importing enterprise data into ERP system and storage medium

By obtaining the application of search terms and semantic recognition models, the efficiency and accuracy of the ERP system in the initialization stage of enterprise information is solved, and the automatic import of enterprise data is realized, and the information management process is simplified.

CN120123409APending Publication Date: 2025-06-10INSPUR GENERSOFT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335779.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing ERP system has time consumption, high error rate and omissions in the initialization stage of enterprise information. Natural language processing technology faces problems such as large amount of data, high computing resources and interference from parallel entities when entering auxiliary information, and cannot achieve full automatic import.

Method used

By obtaining the search terms, the root enterprise is determined, and the relevant enterprise description information collection is extracted from massive enterprise information based on the root enterprise description information, the enterprise entity and relationship are extracted using the semantic recognition model, and the association relationship and permission relationship between the data list are constructed to realize automatic import.

Benefits of technology

It reduces the data processing volume, improves the accuracy of relationship recognition, realizes automatic data import of ERP system, and simplifies the enterprise information management process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123409A_ABST
    Figure CN120123409A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of enterprise resource planning systems, and particularly provides a method and device for importing enterprise data into an ERP system and a storage medium, and the method comprises the steps: obtaining a search word, and obtaining root enterprise description information according to the search word; key words are extracted from the root enterprise description information, an enterprise description information set is obtained according to the key words, and any one piece of enterprise description information in the enterprise description information set comprises the key words; utilizing a semantic recognition model to extract enterprise entities and relationships from the enterprise description information set; and storing the enterprise description information corresponding to the enterprise entities to the data lists corresponding to the enterprise entities, and constructing an association relationship and an authority relationship between the corresponding data lists based on a relationship between the enterprise entities. According to the method and the system, efficient analysis of enterprise data is realized, interference of redundant data is reduced, and key information can be automatically imported into an ERP (Enterprise Resource Planning) system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of enterprise resource planning systems, and particularly relates to a method, device, and storage medium for importing enterprise data into an ERP system. Background Art

[0002] The ERP (Enterprise Resource Planning) system, that is, the enterprise resource planning system, is a product that combines information technology with advanced management ideas relying on information technology. Based on a systematic management idea, the ERP system provides decision-making means for enterprise employees and the decision-making level, and is an important management platform.

[0003] In the existing ERP systems, there are obvious deficiencies in the enterprise information initialization stage. Currently, it is usually necessary to manually collect data and then enter the information by means of Excel import or manual entry. This method not only consumes a large amount of time, but also is prone to errors and omissions, which will have a negative impact on subsequent enterprise information management and analysis work.

[0004] Natural Language Processing (NLP) technology, as an interdisciplinary field of computer science, linguistics, statistics, etc., aims to enable computers to understand, process, and generate human natural language. In enterprise information processing, using natural language processing technology to extract key information that needs to be entered into the ERP system from enterprise information can provide a certain degree of assistance for related work.

[0005] However, when natural language processing technology is applied to assist in ERP system information entry, it also faces many challenges. On the one hand, the enterprise information data volume is huge, which poses high requirements for hardware computing resources; on the other hand, there are many entities involved in enterprise information, and some of the parallel entities do not have a hierarchical relationship, but the model needs to traverse when identifying entity relationships, which not only interferes with the actual relationship identification work, but also leads to a sharp increase in the amount of calculation.

[0006] In addition, even if the key information is successfully extracted with the help of natural language processing technology, it still needs to be manually sorted out logically before the operation of importing into the ERP system can be carried out, and the full-automatic import of the ERP system cannot be achieved. Summary of the Invention

[0007] In view of the above deficiencies of the prior art, the present invention provides a method, device, and storage medium for importing enterprise data into an ERP system to solve the above technical problems.

[0008] In a first aspect, the present invention provides a method for importing enterprise data into an ERP system, including: Obtain a search term, and obtain root enterprise description information according to the search term; Extract keywords from the root enterprise description information, and obtain a set of enterprise description information according to the keywords, where any enterprise description information in the set of enterprise description information contains the keywords; Use a semantic recognition model to extract enterprise entities and relationships from the set of enterprise description information; Save the enterprise description information corresponding to the enterprise entity to a data list corresponding to the enterprise entity, and construct an association relationship and a permission relationship between the corresponding data lists based on the relationships between the enterprise entities.

[0009] In an optional implementation manner, the method further includes: Crawl enterprise description information from a target web page, where the enterprise description information includes enterprise basic information, branch information, and investment information; Name the crawled enterprise description information with the enterprise name and organization code of the affiliated enterprise, and save it to the information library in a document format; Perform duplicate removal processing and regular update on the enterprise description information in the information library.

[0010] In an optional implementation manner, obtaining a search term and obtaining root enterprise description information according to the search term includes: Obtain a search term input by a user, where the search term includes the enterprise code or organization code of the root enterprise; Query the root enterprise description information corresponding to the search term from the information library.

[0011] In an optional implementation manner, extracting keywords from the root enterprise description information and obtaining a set of enterprise description information according to the keywords includes: Use a regular expression to extract target keywords from the root enterprise description information; Use a pre-constructed enterprise name dictionary to extract enterprise names from the root enterprise description information, and output the extracted enterprise names as associated enterprise keywords; Query enterprise description information containing any target keyword or any enterprise keyword from the information library, and save the queried enterprise description information to the set of enterprise description information.

[0012] In an optional implementation manner, using a semantic recognition model to extract enterprise entities and relationships from the set of enterprise description information includes: Randomly select target enterprise description information from the set of enterprise description information; Use a pre-trained semantic recognition model to identify the parallel relationship between enterprise name entities from the target enterprise description information based on the enterprise name dictionary; Create corresponding entity groups for multiple entities with a parallel relationship and create corresponding entity groups for isolated entities without a parallel relationship; Use a pre-trained classification model to identify the relationships between entity groups; Traverse the set of enterprise description information to obtain all enterprise entities and the relationships between enterprise entities.

[0013] In an alternative embodiment, the pre-trained semantic recognition model includes: An input layer, where the input data of the input layer is the text sequence of enterprise description information, and each text sequence is composed of a series of words; the input dimension of the input layer matches the maximum sequence length of the text sequence; An embedding layer for converting the input word index into a corresponding word vector representation; A convolutional layer for extracting local features of the text through convolutional operations; the number of convolutional kernels is 128, the size of the convolutional kernels is 3, and the ReLU activation function is used as the activation function to use the max pooling layer to reduce the dimension of the feature map; An attention mechanism layer for weighting the feature map extracted by the convolutional layer so that the LSTM layer pays more attention to the features of the context-related word vectors of the enterprise name entity and reduces the attention to the enterprise name entity; An LSTM layer for processing the sequence information output by the convolutional layer and capturing long-range dependencies; the number of units is 128; the return sequence is set to True so that the output of each time step can be passed to the next layer; A fully connected layer for mapping the output of the LSTM layer to the category space; the number of neurons in the fully connected layer is 64, and the ReLU activation function is used as the activation function; An output layer for classification using the softmax function to judge the parallel relationship between enterprise name entities; the number of neurons in the output layer is 2, indicating parallel or non-parallel.

[0014] In an alternative embodiment, using the pre-trained semantic recognition model to identify the parallel relationship between enterprise name entities from the target enterprise description information includes: Pre-build an enterprise name dictionary using the Tokenizer class tool in Keras; Convert each enterprise name in the target enterprise description information into the corresponding index in the enterprise name dictionary and generate an entity identifier for the index; Convert the content other than the enterprise name in the target enterprise description information into word vectors; Arrange the word vectors and the indexes corresponding to the enterprise names in the context order into a text sequence; Input the text sequence into the pre-trained semantic recognition model, set the weight value corresponding to the feature of the entity identifier to the minimum value for the weights output by the attention mechanism layer of the pre-trained semantic recognition model, and input the new feature weights into the LSTM layer; Obtain the parallel relationship between the indexes output by the pre-trained semantic recognition model.

[0015] In an alternative embodiment, identifying the relationship between entity groups using a pre-trained classification model includes: Assign a unique corresponding identity vector to each entity group, replace the indexes in the text sequence with the identity vectors of the entity groups to which they belong, and remove duplicate adjacent identity vectors to obtain a new text sequence; Use a recurrent neural network to identify the relationship between entity groups in the new text sequence.

[0016] In a second aspect, there is provided a device, including: A memory for storing a program for importing enterprise data into an ERP system; A processor for implementing the steps of the method for importing enterprise data into an ERP system as provided in the first aspect when executing the program for importing enterprise data into an ERP system.

[0017] In a third aspect, there is provided a computer-readable storage medium having stored thereon a program for importing enterprise data into an ERP system, and when the program for importing enterprise data into an ERP system is executed by a processor, the steps of the method for importing enterprise data into an ERP system as provided in the first aspect are implemented.

[0018] The beneficial effects of the present invention are as follows. The method, device, and storage medium for importing enterprise data into an ERP system provided by the present invention first determine the root enterprise according to the search term, and based on the root enterprise description information, determine the set of enterprise description information related to the current import task from the massive enterprise information. Then, using the set of enterprise description information as sample data, enterprise entities and relationships are extracted. Compared with directly processing the original enterprise information, the data processing volume is greatly reduced. At the same time, according to the extracted entities, the enterprise description information corresponding to the enterprise entities is stored, and equivalent association relationships and permission relationships are established between data tables according to the relationships between the entities, realizing automatic data import. In addition, by first identifying the parallel relationship between entities and excluding the interference of parallel entities in the relationship recognition stage, not only the calculation amount is reduced, but also the accuracy of relationship recognition is improved.

[0019] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect. Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.

[0022] Figure 2 It is a schematic structural diagram of a device provided by an embodiment of the present invention. Detailed implementation manners

[0023] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0025] The following explains the key terms that appear in the present invention.

[0026] LSTM, namely Long Short - Term Memory, is a special type of recurrent neural network (RNN), which can effectively solve the problem of gradient vanishing or gradient explosion in traditional RNNs when processing long sequences, thereby better capturing long - term dependencies in the sequence.

[0027] CNN, namely Convolutional Neural Network, is a type of deep learning model specifically designed for processing data with grid structures (such as image, audio, etc., sequence data), and has achieved great success in the fields of computer vision, natural language processing, etc.

[0028] In deep learning models for natural language processing (NLP), especially when using convolutional neural networks (CNNs) to process text sequences, the feature map is an important concept. A feature map (Feature Map) refers to the matrix of feature representations extracted from the input through specific operations (such as convolutional operations) during the model's processing of the input text sequence. It can be understood as a numerical representation of different feature patterns in the text, and each feature map focuses on capturing a specific type of feature in the input text.

[0029] The ReLU (Rectified Linear Unit) activation function, also known as the rectified linear unit, is a commonly used activation function in the field of deep learning. In a neural network, the role of the activation function is to introduce non-linearity into the network, enabling the neural network to learn complex non-linear relationships.

[0030] The Softmax function is used to convert a real-valued vector into a probability distribution, where each element in the distribution is between 0 and 1, and the sum of all elements is 1. It is commonly used in the output layer of multi-classification problems to convert the model's raw output (also known as logits) into the probability of each class, so that the probability can be used to determine the likelihood of a sample belonging to each class.

[0031] The Tokenizer class in Keras is a utility tool for text preprocessing, mainly used to convert text data into numerical sequences suitable for processing by deep learning models.

[0032] The method for importing enterprise data into an ERP system provided by an embodiment of the present invention is executed by a computer device. Correspondingly, the system for importing enterprise data into an ERP system runs on the computer device.

[0033] Figure 1 It is a schematic flowchart of the method of an embodiment of the present invention. Among them, Figure 1 The execution subject can be a system for importing enterprise data into an ERP system. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0034] As Figure 1 shown, the method includes: S1. Obtain a search term, and obtain root enterprise description information according to the search term.

[0035] The user inputs a search term through the interface. The search term can be an enterprise name, an industry keyword, a product name, etc. The system preprocesses the input search term, such as removing special characters, converting to a unified case format, etc. In addition, stemming or lemmatization techniques in natural language processing can also be used to reduce the search term to its basic form to improve the accuracy of the search.

[0036] According to the preprocessed search terms, the system performs an exact match or fuzzy match query in the existing enterprise information database. If an exact match is used, only the enterprise description information corresponding to the enterprise name that is exactly the same as the search term will be obtained; if a fuzzy match is adopted, the system will use a certain similarity algorithm (such as the edit distance algorithm) to find the enterprise names similar to the search term and obtain their corresponding enterprise description information. At the same time, relevant enterprise description information can also be obtained from external data sources (such as enterprise official websites, industry information platforms, etc.) through web crawler technology, but it is necessary to comply with the crawler protocols and laws and regulations of relevant websites.

[0037] S2. Extract keywords from the root enterprise description information, and obtain a set of enterprise description information according to the keywords. Any enterprise description information in the set of enterprise description information contains the keywords.

[0038] Perform natural language processing on the obtained root enterprise description information. First, perform word segmentation to split the text into individual words or phrases. Open-source word segmentation tools such as Jieba (for Chinese), NLTK (for English and other languages), etc. can be used. Then, through part-of-speech tagging technology, mark the part of speech of each word (such as noun, verb, adjective, etc.). On this basis, adopt keyword extraction algorithms such as TF-IDF (Term Frequency - Inverse Document Frequency) algorithm, TextRank algorithm, etc. to extract representative keywords from the text. The TF-IDF algorithm determines keywords based on the frequency of a word in a document and its rarity in the entire document collection. The TextRank algorithm is based on a graph model and calculates the importance of keywords through the connection relationships and weights between nodes.

[0039] According to the extracted keywords, query again in the enterprise information database and external data sources. The system will traverse each piece of enterprise description information in the database and determine whether it contains at least one keyword. If it does, add this enterprise description information to the set of enterprise description information. For external data sources, also use web crawler technology to obtain relevant web page content and perform keyword matching and screening.

[0040] S3. Use a semantic recognition model to extract enterprise entities and relationships from the set of enterprise description information.

[0041] Select a suitable semantic recognition model, such as a deep learning-based model (e.g., models combined with neural networks such as BERT, LSTM, GRU, etc.) or traditional machine learning models (e.g., support vector machines, conditional random fields, etc.). If a deep learning model is used, a large amount of labeled enterprise entity and relationship data needs to be prepared for training. During the training process, optimize the parameters of the model with the goal of minimizing the loss function (such as the cross-entropy loss function) to improve the accuracy and recall rate of the model. For traditional machine learning models, appropriate features (such as lexical features, syntactic features, semantic features, etc.) need to be extracted and used to train the model.

[0042] Input each piece of text in the enterprise description information set into the trained semantic recognition model. The model will analyze the text and identify the enterprise entities in it (such as enterprise names, product names, branch names, etc.) and the relationships between the entities (such as subordination relationships, cooperation relationships, competition relationships, etc.). For the identified enterprise entities, the model will give their position information in the text (such as start position and end position) for subsequent processing and verification.

[0043] S4. Save the enterprise description information corresponding to the enterprise entity to the data list corresponding to the enterprise entity, and build the association relationship and permission relationship between the corresponding data lists based on the relationships between the enterprise entities.

[0044] Create a corresponding list for each enterprise entity to store the enterprise description information related to the enterprise entity. When an enterprise entity is identified from the enterprise description information set, add the enterprise description information containing the enterprise entity to its corresponding list. At the same time, the data in the list can be further processed, such as deduplication, formatting, etc., to ensure the consistency and accuracy of the data.

[0045] Establish an association relationship between the corresponding enterprise entity data lists according to the relationships between the enterprise entities extracted by the semantic recognition model. For example, if there is a subordination relationship between two enterprise entities, then a pointing relationship can be set between the data lists representing the two enterprise entities to clarify their subordination levels. A graph database (such as Neo4j) can be used to store these enterprise entities and the relationships between them. The graph database can efficiently process complex relationship data and facilitate querying and analysis.

[0046] Set corresponding permission relationships for different data lists according to the relationships between enterprise entities and business requirements. For example, for enterprise entities with a subordination relationship, the data list of the superior enterprise entity has higher permissions and can access and modify some or all of the information in the data list of the subordinate enterprise entity; while for enterprise entities with a cooperation relationship, the permissions between the data lists are independent of each other, and only the publicly available part of the information of the other party can be accessed. The setting of permission relationships can be achieved through a permission management system, which can verify and authorize users' access requests to ensure the security and privacy of data.

[0047] In an embodiment of the present invention, based on step S1, the following will give an embodiment to non-restrictively elaborate on its specific implementation scheme.

[0048] S101. Obtain a search term input by a user, where the search term includes the enterprise code or institutional code of the root enterprise. S102. Query the root enterprise description information corresponding to the search term from the information library.

[0049] Among them, the construction method of the information library includes: (1) Crawl enterprise description information from the target web pages, where the enterprise description information includes enterprise basic information, branch information, and investment information.

[0050] Rank the determined target web pages according to factors such as authority, information richness, and update frequency. For example, the official website of an enterprise usually has the highest priority because the information it provides is the most accurate and comprehensive; industry information platforms, regulatory websites, etc. also have relatively high reference value.

[0051] Common programming languages include Python, and frameworks that can be paired with it such as Scrapy, BeautifulSoup, etc. Scrapy is a powerful crawler framework that can efficiently handle web page requests, data parsing, and storage; BeautifulSoup is a convenient HTML / XML parsing library that can be used to extract specific information from web pages. To avoid being recognized as a crawler and blocked by the target website, reasonable request headers need to be set. The request headers should include information such as User - Agent (simulating browser information), Referer (indicating the request source), etc. For example, set the User - Agent to the identifier of a common browser (such as Chrome, Firefox). Some websites will adopt anti - crawler mechanisms such as captchas, IP bans, and frequency limits. For captchas, OCR (Optical Character Recognition) technology can be used for recognition, or manual assistance can be adopted to solve them; for IP bans, a proxy IP pool can be used to regularly change the IP address of the request; for frequency limits, the request frequency of the crawler needs to be reasonably controlled to avoid overly frequent requests.

[0052] Use an HTML / XML parser to parse the web page content crawled. For web pages with regular structures, XPath or CSS selectors can be used to locate and extract enterprise description information. For example, using the XPath expression / / div[@class='company - info'] can extract all elements on the web page with the class name company - info Information within the label. Classify and organize the extracted information according to the basic information of the enterprise (such as enterprise name, legal representative, establishment date, business scope, etc.), branch information (such as branch name, address, person in charge, etc.) and investment information (such as investment amount, investment time, investor, etc.).

[0053] (2)Name the crawled enterprise description information with the enterprise name and institutional code of the affiliated enterprise, and save it in the information library in document format.

[0054] Clean the extracted enterprise description information to remove noise data such as HTML tags, special characters, and extra spaces. For example, use regular expressions to remove , HTML tags such as . Standardize key information such as enterprise names and institutional codes. For example, unify the case format of enterprise names, and perform format verification and supplementation on institutional codes (such as unifying them into 18-digit social credit codes).

[0055] Name the files in the way of combining enterprise names and institutional codes, such as [Enterprise Name]_[Institutional Code].txt. To avoid problems caused by special characters in file names, the enterprise name can be further processed to only retain legal characters. Before saving the file, check whether the file name already exists in the information repository. If it exists, the file name needs to be appropriately modified, such as adding a serial number after the file name to ensure the uniqueness of the file name.

[0056] According to the requirements and usage scenarios of the information repository, select appropriate document formats, such as text files (.txt), JSON files (.json), CSV files (.csv), etc. Text files are simple and easy to read, suitable for storing pure text information; JSON files have good structuring and readability, suitable for storing complex nested data; CSV files are convenient for data import and export, suitable for data analysis. Determine the save path of the information repository and set the corresponding read and write permissions. Ensure that only authorized users can access and modify the files in the information repository to guarantee data security.

[0057] (3) Remove duplicates and update the enterprise description information in the information repository regularly.

[0058] Select key features in the enterprise description information (such as enterprise name, institutional code, establishment date, etc.) as the basis for deduplication. Calculate the feature values of each record in the information repository and perform hash processing on them to obtain unique hash values. Find duplicate records by comparing the hash values.

[0059] Determine the update cycle of the information repository according to the update frequency of enterprise information and business requirements. For example, for some enterprises with rapidly changing information, it can be set to update once a week; for enterprises with relatively stable information, it can be set to update once a month or once a quarter.

[0060] Develop an automated script to regularly trigger the crawler program to re-crawl the target web pages. Compare the newly crawled enterprise description information with the existing information in the information repository. For updated information, replace the original record; for new information, add it to the information repository. At the same time, record the time and content of each update for subsequent auditing and traceability.

[0061] In an embodiment of the present invention, based on step S2, the following will give an embodiment to non-restrictively elaborate on its specific implementation scheme.

[0062] S201. Extract target keywords from the root enterprise description information using regular expressions.

[0063] First, deeply analyze the text features of the root enterprise description information to clarify the possible patterns of target keywords. For example, if the target keyword is the product name of an enterprise, it may contain specific industry terms, brand names, etc.; if it is the business area of an enterprise, there may be some common descriptive words, such as "software development", "financial services", etc.

[0064] According to the characteristics of the keywords, construct corresponding regular expressions. Taking the extraction of product names as an example, if the product name usually starts with a brand and is followed by a specific model or series name, a regular expression like [brand name][\w\s-]+ can be constructed, where [\w\s-] means to match letters, numbers, spaces, and hyphens, and + means to match the previous character class one or more times. For Chinese keywords, considering the Chinese encoding and character range, use [\u4e00-\u9fa5] to match Chinese characters.

[0065] Use part of the root enterprise description information to test the constructed regular expression to see if it can accurately extract the target keywords. If the extraction result is inaccurate, adjust and optimize the regular expression, and it may be necessary to add or modify elements such as character classes and quantifiers.

[0066] Before applying the regular expression, preprocess the root enterprise description information. Remove special characters, HTML tags (if the information is from a web page), etc. in the text, and convert the text to a unified case format to improve the accuracy of regular expression matching.

[0067] Use the regular expression library (such as the re module) in a programming language (such as Python) to perform matching operations on the preprocessed root enterprise description information. Traverse each possible matching position in the text and extract the matching target keywords.

[0068] Screen the extracted target keywords to remove some meaningless or non-compliant words. For example, remove some overly general words (such as "company", "enterprise", etc.). At the same time, perform deduplication on the keywords to prevent duplicate keywords from entering the subsequent processing process.

[0069] S202. Extract the enterprise name from the root enterprise description information using a pre-constructed enterprise name dictionary, and output the extracted enterprise name as an associated enterprise keyword.

[0070] (1) Construct an enterprise name dictionary Collect enterprise name data from various sources, including enterprise official websites, industrial and commercial registration information, industry directories, etc. Ensure the accuracy and integrity of the data, and conduct preliminary cleaning and sorting on the collected enterprise names to remove duplicate and incorrect names.

[0071] Use the Tokenizer class tool in Keras to build an enterprise name dictionary. The Tokenizer class can build a word index dictionary from text data. The specific steps are as follows: Create a Tokenizer object and set num_words to None, indicating that all different enterprise names are to be retained; Fit the enterprise name data to let Tokenizer learn the vocabulary of enterprise names; Obtain the enterprise name dictionary (word index); Print all enterprise names and their corresponding indexes; Use the json module in Python to save the dictionary as a JSON file.

[0072] Regularly update the enterprise name dictionary, add newly established enterprise names, and delete enterprise names that have been cancelled or merged. At the same time, verify and correct the names in the dictionary to ensure their accuracy.

[0073] (2)Query associated enterprise keywords Scan the root enterprise description information word by word, and compare each consecutive string with the enterprise name dictionary. The sliding window method can be adopted, starting from shorter strings and gradually increasing the length to check if there is a matching enterprise name in the dictionary.

[0074] Give priority to exact matching, that is, check whether the scanned string is exactly the same as the enterprise name in the dictionary. If the exact matching fails, fuzzy matching can be considered, such as using an edit distance algorithm (such as Levenshtein distance) to calculate the similarity between the scanned string and the name in the dictionary. When the similarity exceeds a certain threshold, it is considered a successful match.

[0075] For the extracted enterprise names, verify them in combination with the context information to ensure that they actually refer to an enterprise. For example, check whether there are enterprise-related descriptive words before and after the name, such as "company", "group", etc.

[0076] Convert the extracted enterprise names into a unified format, such as removing extra spaces, punctuation marks, etc., to ensure the standardization of keywords. Store the processed enterprise names as associated enterprise keywords in a suitable data structure, such as a list or a set, for subsequent query and use.

[0077] S203. Query the enterprise description information containing any target keyword or any enterprise keyword from the information repository, and save the queried enterprise description information to the enterprise description information set.

[0078] (1) Information repository index construction According to the scale of the information repository and the query requirements, select an appropriate index type. For a text information repository, full-text indexing such as an inverted index can be used. The inverted index records in which documents each keyword appears and can quickly locate the documents containing a specific keyword.

[0079] Use the index creation function provided by a database management system (such as MySQL, Elasticsearch, etc.) to index the enterprise description information in the information repository. When the information repository is updated, update the index in a timely manner to ensure the accuracy and timeliness of the index.

[0080] (2) Query operation execution Construct a query statement based on the target keyword and the enterprise keyword. In the case of using full-text indexing, the query statement usually uses a keyword matching method, such as using the match query in Elasticsearch. For example, construct a query statement to query the enterprise description information containing any one of the target keywords or any one of the enterprise keywords.

[0081] Execute the query statement to obtain the matching enterprise description information from the information repository. Filter the query results to remove duplicate records and information that does not meet the requirements, such as description information with empty or incomplete content.

[0082] (3) Saving the enterprise description information set Select an appropriate data structure to save the queried enterprise description information, such as a list, dictionary, or database table. If further processing and analysis of the information are required, lists or dictionaries in Python can be used; if long-term storage and management are needed, the information can be saved to a database table.

[0083] Save the filtered enterprise description information to the enterprise description information set. During the saving process, verify the information to ensure its integrity and accuracy. At the same time, record the saving time and relevant query conditions for subsequent traceability and auditing.

[0084] In an embodiment of the present invention, based on step S3, the following will give an embodiment to non-restrictively elaborate on its specific implementation scheme.

[0085] S301. Randomly select target enterprise description information from the enterprise description information set.

[0086] Implement random selection using a random number generator in a programming language (such as Python). In Python, the choice function of the random module can be used to randomly select a target enterprise description information from the enterprise description information set.

[0087] S303. Use a pre-trained semantic recognition model to identify the coordination relationship between enterprise name entities based on the enterprise name dictionary from the target enterprise description information.

[0088] (1) Convert each enterprise name in the target enterprise description information into the corresponding index in the enterprise name dictionary, and generate an entity identifier for the index.

[0089] Traverse the target enterprise description information, and use the enterprise name dictionary to convert each enterprise name into the corresponding index. If the enterprise name is not in the dictionary, its index can be set to a special value (such as 0) to indicate unknown. Generate an entity identifier for each index, such as 1.

[0090] (2) Convert the content other than the enterprise name in the target enterprise description information into word vectors.

[0091] Use a pre-trained word vector model (such as Word2Vec, GloVe, or FastText) to convert the content other than the enterprise name in the target enterprise description information into word vectors. For each non-enterprise name word, find its vector representation in the pre-trained word vector model. If the word is not in the vocabulary of the model, a zero vector or a random vector can be used to represent it.

[0092] (3) Arrange the word vectors and the indexes corresponding to the enterprise names in the context order into a text sequence.

[0093] Arrange the word vectors and the indexes corresponding to the enterprise names in the context order into a text sequence. A list can be used to store this sequence, where the indexes of the enterprise names are represented by integers, and the non-enterprise name words are represented by their corresponding word vectors.

[0094] (4) Input the text sequence into the pre-trained semantic recognition model, set the weight value of the feature corresponding to the entity identifier to the minimum value for the weight output by the attention mechanism layer of the pre-trained semantic recognition model, and input the new feature weight into the LSTM layer.

[0095] The pre-trained semantic recognition model includes: An input layer, the input data of the input layer is the text sequence of the enterprise description information, and each text sequence is composed of a series of words; the input dimension of the input layer matches the maximum sequence length of the text sequence; An embedding layer, which is used to convert the input word index into the corresponding word vector representation; The convolutional layer is used to extract local features of the text through convolutional operations; the number of convolutional kernels is 128, the size of the convolutional kernels is 3, and the ReLU activation function is used as the activation function to reduce the dimension of the feature map using the max pooling layer; The attention mechanism layer is used to weight the feature map extracted by the convolutional layer, so that the LSTM layer pays more attention to the features of the context-related word vectors of the enterprise name entity and reduces the attention to the enterprise name entity; The LSTM layer is used to process the sequence information output by the convolutional layer and capture long-range dependencies; the number of units is 128; the return sequence is set to True so that the output of each time step can be passed to the next layer; The fully connected layer is used to map the output of the LSTM layer to the category space; the number of neurons in the fully connected layer is 64, and the ReLU activation function is used as the activation function; The output layer is used to perform classification using the softmax function to judge the parallel relationship between enterprise name entities; the number of neurons in the output layer is 2, indicating parallel or non-parallel.

[0096] Since the input layer of the pre-trained semantic recognition model has a fixed input dimension, the text sequence needs to be padded or truncated to the maximum sequence length. The pad_sequences function of Keras can be used for padding.

[0097] The model is trained using the pre-collected and labeled enterprise description information.

[0098] Among the weights output by the attention mechanism layer of the pre-trained semantic recognition model, the weight value corresponding to the feature of the entity identifier is set to the minimum value. By traversing the attention weight matrix, the position corresponding to the entity identifier can be found and its weight can be set to an extremely small value (such as 1e - 9 ). The new feature weights are input into the LSTM layer for processing. In Keras, this can be achieved by calling the corresponding layer of the model. The output of the LSTM layer passes through the fully connected layer and the output layer, and the softmax function is used for classification to obtain the parallel relationship between the indices. By comparing the output probability values, the category with the highest probability can be selected as the prediction result.

[0099] By modifying the attention weight matrix, the position corresponding to the entity identifier is found and its weight is set to an extremely small value, reducing the attention of the LSTM layer to the entity, focusing on the features of the context word vectors, reducing the interference of the entity content on the recognition of the parallel relationship, and thus improving the efficiency and accuracy of the parallel relationship recognition.

[0100] The inference process of the pre-trained semantic model includes: Adjust the weights of the attention mechanism layer. After the weights are output from the attention mechanism layer of the pre-trained semantic recognition model, set the weight values of the features corresponding to the entity identifiers to the minimum value. Specifically, traverse the attention weight matrix, find the positions corresponding to the entity identifiers, and then set the weights at these positions to an extremely small value, such as 1e -9 .

[0101] Input to the LSTM layer. Input the features with adjusted weights into the LSTM layer of the model for processing. In a model built using Keras, this input operation can be achieved by calling the corresponding layer.

[0102] Obtain the coordination relationship. After the LSTM layer finishes processing, the output result will pass through the fully connected layer and the output layer. Use the softmax function for classification in the output layer to obtain the coordination relationship between the indices. By comparing the output probability values, select the category with the highest probability as the prediction result to determine whether there is a coordination relationship between enterprise name entities.

[0103] (5) Obtain the coordination relationship between the indices output by the pre-trained semantic recognition model.

[0104] S303. Create corresponding entity groups for multiple entities with a coordination relationship, and create corresponding entity groups for isolated entities without a coordination relationship.

[0105] According to the coordination relationship output by the pre-trained semantic recognition model, divide the enterprise name entities with a coordination relationship into one entity group, and each enterprise name entity without a coordination relationship becomes an entity group.

[0106] S304. Use the pre-trained classification model to identify the relationships between entity groups.

[0107] Assign a unique corresponding identity vector to each entity group. For example, use the hash value of the index combination of the members within an entity group as the identity vector. Replace the indices in the text sequence with the identity vectors of the entity groups to which they belong to obtain a new text sequence. By traversing the new text sequence, compare whether adjacent vectors are the same. If they are the same, only keep one. Finally, obtain a deduplicated text sequence.

[0108] After such processing, replace multiple entities in the text sequence with the corresponding entity groups, and only identify the relationships between entity groups in the subsequent recognition process, greatly reducing the data processing volume and the interference of coordinated entities on relationship recognition.

[0109] The final text sequence is input into a pre-trained recurrent neural network (such as LSTM or GRU) for relationship recognition. In Keras, a simple recurrent neural network model is constructed for prediction. For example, a simple LSTM recurrent neural network model is constructed using Keras to perform relationship recognition prediction on the final text sequence: Data preparation: Random final text sequence data and relationship labels are generated, and the relationship labels are converted into one-hot encoding to facilitate multi-classification training of the model. The train_test_split function is used to divide the data into a training set and a test set.

[0110] Model construction: A Sequential model is used to construct an LSTM model. First, an LSTM layer is added to process the sequence data, then a fully connected layer is added for feature mapping, and finally an output layer is added, using the softmax activation function for multi-classification.

[0111] Model compilation: The model is compiled using the adam optimizer and the categorical_crossentropy loss function, and accuracy is selected as the evaluation metric.

[0112] Model training: The fit method is used to train the model, setting the number of training epochs and the batch size, and using the test set for validation.

[0113] Prediction: The predict method is used to predict the test set, and the prediction results are converted into class labels.

[0114] S305. Traverse the set of enterprise description information to obtain all enterprise entities and the relationships between enterprise entities.

[0115] The steps of S301 - S304 are repeatedly executed for each description information in the set of enterprise description information, and the enterprise entities and the relationships between entities obtained each time are summarized to finally obtain all the relationships between enterprise entities and enterprise entities.

[0116] In a scenario of importing enterprise information into an ERP system, the following processes are included: (1)Identification and information initialization of the root enterprise in the target enterprise information database. To quickly initialize enterprise information, it is necessary to identify the "root enterprise". For a group, the "root enterprise" refers to the highest-level enterprise, and there can be multiple "root enterprises". The enterprise information database will be constructed based on these "root enterprises". The "root enterprise" is more like the clue source for constructing the enterprise information database. Through this clue source, information such as the enterprises invested by these enterprises, branches, etc. is searched to construct the enterprise information database. For the information of the root enterprise, only its social credit code or the complete enterprise name needs to be provided, and there is no need to prepare the complete enterprise information data. (2)Prepare to access external data tools. Start accessing external data and obtain the required information by calling the API of external data or other means. This includes at least steps such as API call, data reception, and preliminary processing of return values.

[0117] (3)Prepare data cleaning and preprocessing tools. The obtained information is processed, de-duplicated, error-checked, and data-formatted through a complete data processing tool, so as to form the correct enterprise information or the information to be utilized.

[0118] (4)Set the maximum hierarchical limit for searching enterprise information of invested enterprises and branches. Search for the invested enterprises and branches of the "root enterprise", and then use the invested enterprises and branches to further obtain the information of the enterprises and branches invested by these invested enterprises and branches. This means that the program may not stop, and the process of obtaining enterprise information may not stop. However, enterprises are not concerned about overly deep investment relationships or branch relationships, so it is necessary to set a maximum search hierarchical limit.

[0119] (5)Obtain the basic enterprise information of the "root" enterprise. Through the social credit code or the complete enterprise name of the root enterprise provided by the user, the basic information of the enterprise can be obtained, including the registered place, main business, shareholders, contact information, and many other information. After cleaning and preprocessing the obtained external data, these information can be solidified into the system.

[0120] (6)Obtain the information of the invested enterprises of the root enterprise. Through the social credit code or the complete enterprise name of the root enterprise provided by the user, the enterprises invested by the enterprise can be obtained, and according to the method in step 5, the information of these enterprises can also be further solidified into the system. Among them, if the user is not interested in the information of the invested enterprises and only focuses on the branches of the enterprise, then this step can be skipped by adjusting the parameters.

[0121] (7) Recursively obtain investment enterprise information. The investment enterprise information obtained in step 6 can continue to search for its investment enterprise and branch information until the target enterprise's investment enterprise information and branch information can no longer be searched or exceeds the hierarchical limit set by the system. The acquisition and processing methods of branches will be described in steps 8 and 9. Note that the solidified enterprise will not be solidified again, and the existing enterprise information will be updated.

[0122] In this step, if the enterprise is only concerned with the investment enterprise relationship and does not care about the branches, then this step will only search for the investment enterprise without searching for the branches. The function can be adjusted through simple parameter adjustments.

[0123] If no investment enterprise information is obtained in step 6, this step will be skipped.

[0124] (8) Obtaining branch information of the root enterprise Through the social credit code or full name of the root enterprise provided by the user, the enterprises that are branches of the root enterprise can be obtained, and the information of these enterprises can be further solidified into the system according to the method in step 5.

[0125] If the user is not interested in branch information and only cares about the enterprise's investment enterprises, this step can be skipped by adjusting parameters.

[0126] (9) Recursively obtain branch information The information of the invested enterprise obtained in step 8 can be used to continue searching for its invested enterprise and branch information until the target enterprise can no longer search for its invested enterprise information and branch information or exceeds the hierarchical limit set by the system. Note that the solidified enterprise will not be solidified again, and the existing enterprise information will be updated.

[0127] In this step, if the enterprise is only concerned with the relationship of branches and does not care about the invested enterprises, then this step will only search for branches without searching for invested enterprises. The function can be adjusted through simple parameter adjustments.

[0128] If no branches are obtained in step 8, this step will be skipped.

[0129] (10) Information solidification The enterprise information obtained in the above steps is solidified into the system. At this point, the enterprise's information database is automatically initialized.

[0130] Figure 2 The method for importing enterprise data into an ERP system provided by the embodiments of this application can be applied to devices. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the device includes but is not limited to laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of this application described herein and / or claimed.

[0131] Among them, the device 200 may include: a processor 210, a memory 220, and a communication unit 230. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It can be a bus structure, a star structure, or may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0132] Among them, the memory 220 can be used to store the execution instructions of the processor 210. The memory 220 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks. When the execution instructions in the memory 220 are executed by the processor 210, the device 200 can execute some or all of the steps in the above method embodiments.

[0133] The processor 210 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 220, and by invoking data stored in the memory, it performs various functions of the electronic device and / or processes data. The processor may be composed of an integrated circuit (IC), for example, it may be composed of a single packaged IC, or it may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 210 may include only a central processing unit (CPU). In the embodiments of the present invention, the CPU may be a single computing core or may include multiple computing cores.

[0134] The communication unit 230 is used to establish a communication channel so that the storage device can communicate with other devices. It receives user data sent by other devices or sends user data to other devices.

[0135] The present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the embodiments provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0136] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes, and includes several instructions to enable a computer device (which may be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0137] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.

[0138] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the systems or modules can be in electrical, mechanical, or other forms.

[0139] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0140] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0141] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, and they should all be covered within the protection scope of the present invention.

Claims

1. A method for importing enterprise data into an ERP system, characterized in that: include: Obtaining a search term, and obtaining root enterprise description information according to the search term; Extracting keywords from the root enterprise description information, and acquiring an enterprise description information set according to the keywords, wherein any enterprise description information in the enterprise description information set contains the keywords; Extracting enterprise entities and relationships from the enterprise description information set using a semantic recognition model; The enterprise description information corresponding to the enterprise entity is saved in the data list corresponding to the enterprise entity, and the association relationship and authority relationship between the corresponding data lists are constructed based on the relationship between the enterprise entities.

2. The method according to claim 1, characterized in that: The method further comprises: crawling enterprise description information from the target web page, wherein the enterprise description information includes basic enterprise information, branch information, and capital injection information; The crawled enterprise description information is named with the enterprise name and organization code of the enterprise to which it belongs, and saved in the information database in a document format; De-duplicate and regularly update the enterprise description information in the database.

3. The method according to claim 1, characterized in that Obtaining a search term, and obtaining root enterprise description information according to the search term, including: Acquire a search term input by a user, wherein the search term includes an enterprise code or an organization code of a root enterprise; The root enterprise description information corresponding to the search term is searched from the information base.

4. The method according to claim 1, characterized in that: Extracting keywords from the root enterprise description information and acquiring an enterprise description information set according to the keywords includes: Extracting target keywords from the root enterprise description information using regular expressions; Extracting enterprise names from the root enterprise description information using a pre-built enterprise name dictionary, and outputting the extracted enterprise names as associated enterprise keywords; The enterprise description information containing any target keyword or any enterprise keyword is searched from the information database, and the searched enterprise description information is saved in the enterprise description information set.

5. The method according to claim 1, characterized in that Extracting enterprise entities and relationships from the enterprise description information set using a semantic recognition model includes: Randomly select target enterprise description information from the enterprise description information set; Using the pre-trained semantic recognition model based on the enterprise name dictionary to identify the parallel relationship between enterprise name entities from the target enterprise description information; Create corresponding entity groups for multiple entities with parallel relationships, and create corresponding entity groups for isolated entities without parallel relationships; Identify relationships between groups of entities using pre-trained classification models; The enterprise description information set is traversed to obtain all enterprise entities and the relationships between enterprise entities.

6. The method according to claim 5, characterized in that The pre-trained semantic recognition model includes: An input layer, wherein the input data of the input layer is a text sequence of enterprise description information, each text sequence is composed of a series of words; the input dimension of the input layer matches the maximum sequence length of the text sequence; Embedding layer, which is used to convert the input word index into the corresponding word vector representation; Convolution layer, used to extract local features of text through convolution operation; the number of convolution kernels is 128, the convolution kernel size is 3, and the activation function uses the ReLU activation function to use the maximum pooling layer to reduce the dimension of the feature map; The attention mechanism layer is used to weight the feature graph extracted by the convolution layer so that the LSTM layer pays more attention to the features of the context-related word vectors of the company name entity and pays less attention to the company name entity; LSTM layer, used to process the sequence information output by the convolutional layer and capture long-distance dependencies; the number of units is 128; return sequence is set to True so that the output of each time step is passed to the next layer; The fully connected layer is used to map the output of the LSTM layer to the category space. The number of neurons in the fully connected layer is 64, and the activation function uses the ReLU activation function. The output layer is used to perform classification using the softmax function to determine the parallel relationship between enterprise name entities; the number of neurons in the output layer is 2, indicating parallel or non-parallel.

7. The method according to claim 6, characterized in that The pre-trained semantic recognition model is used to identify the parallel relationship between enterprise name entities from the target enterprise description information based on the enterprise name dictionary, including: Pre-build a dictionary of company names using Keras' Tokenizer class tool; Convert each enterprise name in the target enterprise description information into a corresponding index in the enterprise name dictionary, and generate an entity identifier for the index; Convert the contents other than the company name in the target company description information into word vectors; Arrange the indexes corresponding to the word vectors and company names into a text sequence in context order; Input the text sequence into the pre-trained semantic recognition model, set the weight value of the feature corresponding to the entity identifier output by the attention mechanism layer of the pre-trained semantic recognition model to the minimum value, and input the new feature weight into the LSTM layer; Obtain the parallel relationship between the indexes output by the pre-trained semantic recognition model.

8. The method according to claim 7, characterized in that Leverage pre-trained classification models to identify relationships between groups of entities, including: Assign a unique identity vector to each entity group, replace the index in the text sequence with the identity vector of the entity group to which it belongs, and remove duplicates from adjacent duplicate identity vectors to obtain a new text sequence; Recurrent neural networks are used to identify relationships between groups of entities in new text sequences.

9. A device, characterized in that: include: Memory, used to store enterprise data and import it into the ERP system program; A processor is used to implement the steps of the method for importing enterprise data into an ERP system as described in any one of claims 1 to 8 when executing the program for importing enterprise data into an ERP system.

10. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a program for importing enterprise data into an ERP system, and when the program for importing enterprise data into an ERP system is executed by a processor, the steps of the method for importing enterprise data into an ERP system as described in any one of claims 1 to 8 are implemented.