Association relation identification method and device, electronic equipment and storage medium
Through pre-training language models and relationship extraction models, unstructured data are processed, and the cosine similarity measurement and triple loss function training model is used to solve the problem of inaccurate identification of enterprise relationships in the existing technology, and automated and accurate relationship recognition is achieved, which improves the customer mining effect of supply chain finance.
Patent Information
- Application Number
- CN202510691020.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
AI Technical Summary
The existing technology is difficult to efficiently and accurately identify the relationship between enterprises from unstructured data, resulting in poor potential customer mining results in supply chain finance, especially in the scenario of complex enterprise relationships.
The pre-trained language model is used to perform data preprocessing and relation extraction, the unstructured data is processed through the relation extraction model, the target correlation relationship is determined using the cosine similarity metric, and the model is trained in combination with metric learning and triple-tuple loss function, so as to improve the model's ability to distinguish different categories.
It realizes automatic identification of unstructured data, accurately determines the target relationship between enterprises, improves the efficiency and accuracy of potential customers in supply chain finance, and reduces human and material costs.
Smart Images

Figure CN120598656A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to an association relationship identification method, device, electronic device and storage medium. Background Art
[0002] Supply chain finance is a specialized area of commercial bank lending and a financing channel for businesses, particularly small and medium-sized enterprises. In recent years, supply chain finance has grown rapidly, enabling small and medium-sized businesses upstream and downstream of core enterprises to access more convenient financial services and expand development opportunities. Expanding the supply chain to benefit more businesses is a key issue that must be addressed in this development.
[0003] To encourage more businesses to participate in supply chain financing, timely and accurate data on inter-enterprise relationships is crucial for identifying potential clients for supply chain finance. However, with the continued development of corporate group operations and the financial industry, inter-enterprise relationships are becoming increasingly complex. In addition to existing structured data, loan officers are seeking to obtain more real-time data from online resources to supplement their understanding of enterprise relationships in supply chain finance. However, this data often exists in unstructured form, making it costly to process.
[0004] Therefore, finding an automated method to extract enterprise relationships from unstructured data is an urgent need in current supply chain finance. Summary of the Invention
[0005] The present invention provides a method, device, electronic device and storage medium for automatically determining target association relationships between objects from unstructured data.
[0006] According to one aspect of the present invention, a method for identifying an association relationship is provided, comprising:
[0007] Acquire data to be identified and a set of relationship types, wherein the data to be identified includes unstructured data for identifying association relationships between objects, and the set of relationship types includes multiple candidate association relationships;
[0008] Preprocessing the data to be identified to be adapted to the relation extraction model to obtain a first input sequence;
[0009] The first input sequence and the relationship type set are processed separately by the relationship extraction model to obtain a target relationship among the candidate relationships corresponding to the data to be identified.
[0010] According to another aspect of the present invention, there is provided an association relationship identification device, comprising:
[0011] An acquisition module, configured to acquire data to be identified and a set of relationship types, wherein the data to be identified includes unstructured data for identifying association relationships between objects, and the set of relationship types includes a plurality of candidate association relationships;
[0012] A preprocessing module, configured to preprocess the data to be identified so as to be adapted to the relation extraction model, thereby obtaining a first input sequence;
[0013] A processing module is used to process the first input sequence and the relationship type set respectively through the relationship extraction model to obtain a target relationship among the candidate relationships corresponding to the data to be identified.
[0014] According to another aspect of the present invention, an electronic device is provided, comprising:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the method according to any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method described in any embodiment of the present invention when executed.
[0019] The technical solution of an embodiment of the present invention obtains data to be identified, including unstructured data, and a set of relationship types including multiple candidate relationships. The data to be identified is preprocessed and then processed together with the set of relationship types using a relationship extraction model to determine target relationships. This achieves automated identification of unstructured data to determine target relationships between objects represented by the unstructured data.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 This is a flowchart of a method for identifying association relationships provided according to the first embodiment of the present invention;
[0023] Figure 2 This is a schematic diagram of an overall process for determining an association relationship provided by an embodiment of the present invention;
[0024] Figure 3 is a schematic diagram of a preprocessed input sequence provided by an embodiment of the present invention;
[0025] Figure 4 This is a coding diagram provided in an embodiment of the present application;
[0026] Figure 5 This is a schematic diagram of a training process of a model to be trained provided in an embodiment of the present application;
[0027] Figure 6 This is a flow chart of a relationship extraction model for determining target association relationships provided by an embodiment of the present application;
[0028] Figure 7 This is a structural diagram of an association relationship identification device provided according to the second embodiment of the present invention;
[0029] Figure 8 It is a structural diagram of an electronic device for implementing the association relationship identification method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] The acquisition, storage, use, and processing of data in the technical solution of the present invention comply with the relevant provisions of relevant laws and regulations.
[0033] Supply chain finance (SCF) is a financing method in which banks connect core enterprises with upstream and downstream businesses to provide flexible financial products and services. The core purpose of SCF is to optimize capital allocation and improve capital flow efficiency for small and medium-sized enterprises (SMEs) upstream and downstream of the supply chain. To attract more businesses to participate in SCF, timely and accurate information about inter-enterprise relationships is essential. This is crucial for identifying potential customers and mitigating risks in SCF. With the continuous development of enterprise group operations and the financial industry, inter-enterprise relationships are becoming increasingly complex and hidden. In addition to existing structured data, loan officers attempt to obtain more real-time data from online resources to supplement SCF's enterprise relationships. However, this data often exists in unstructured form, consuming significant manpower and resources. Therefore, finding an automated method to extract enterprise relationships from unstructured data is a pressing need in SCF.
[0034] In the traditional model, businesses in the supply chain are mostly discovered manually by account managers or recommended by core enterprises. Account managers have limited energy and are unable to fully tap into the marketing value of existing customers. This model makes it difficult to identify small and medium-sized enterprises and self-employed individuals at the end of the supply chain, often facing capital shortages and limited financing channels. Therefore, accurately and quickly identifying these businesses and providing them with convenient financial services is key to expanding the scale of supply chain financing. Related technologies use classification and clustering algorithms to group businesses into similar groups. Then, using association rule algorithms, they discover potential relationships between these customers, thereby discovering more customers. This approach works well for customer discovery in scenarios with simple customer organizational relationships (i.e., in supply chain finance, identifying potential customers upstream and downstream of the supply chain). However, the results are less accurate in supply chain scenarios with complex and unstructured enterprise relationships.
[0035] In order to solve the above problems, the present invention provides a method for identifying an association relationship. The method for identifying an association relationship provided by the present invention is described below.
[0036] Example 1
[0037] Figure 1 This is a flowchart of a method for identifying association relationships provided according to the first embodiment of the present invention. This embodiment is applicable to identifying association relationships between objects in unstructured data. The method can be executed by an association relationship identification device. The association relationship identification device can be implemented in the form of hardware and / or software. The association relationship identification device can be configured in an electronic device. The electronic device can be a mobile phone, computer, personal digital assistant, or other device capable of data processing. Figure 1 As shown, the method includes:
[0038] S110 . Acquire data to be identified and a set of relationship types, wherein the data to be identified includes unstructured data used to identify association relationships between objects, and the set of relationship types includes multiple candidate association relationships.
[0039] The data to be identified can be considered data for which relationships between objects are to be identified. For example, the data to be identified is natural language text used to identify relationships between objects. For example, the data to be identified is "Xiao Ming Company has supplied Xiao Wang Company for many years and has a good cooperative relationship." The objects can be considered the entities for which relationships are to be identified, such as companies.
[0040] There may be multiple data to be identified, and each data to be identified is processed separately, with the same identification means. This embodiment is described by taking one data to be identified as an example.
[0041] Unstructured data refers to data without a predefined format or structure, typically including text, images, audio, video, etc. These data types are diverse and difficult to store and manage using traditional relational databases.
[0042] An association relationship can be considered information indicating the connection between objects, such as the relationship between objects in a supply chain. This disclosure does not limit the specific content of an association relationship, and may include: upstream supplier-enterprise, downstream distributor-enterprise, core enterprise-enterprise, issuer-enterprise, recipient-enterprise, acceptor-enterprise, pledgee-enterprise, creditor-enterprise, confirmer-enterprise, government agency-enterprise, and NA. NA indicates no association relationship between the two enterprises.
[0043] The relationship type set can be considered as a set formed by candidate association relationships. In the process of determining a target association relationship for the data to be identified, a candidate association relationship can be selected from the relationship type set as the target association relationship.
[0044] The candidate association relationship can be considered as a candidate association relationship, and the target association relationship can be considered as an association relationship represented by the data to be identified that is selected from the candidate association relationships.
[0045] In this embodiment, the data to be identified can be obtained through a human-computer interaction interface, from other electronic devices, or from a network via a data interface. The method of obtaining the data to be identified is not limited herein. The relationship type set can be a pre-set set, and the candidate association relationships in the relationship type set can be flexibly added, deleted, or modified.
[0046] S120 , preprocessing the data to be identified to be adapted to the relationship extraction model to obtain a first input sequence.
[0047] Relation extraction can refer to identifying the association relationships between entities from unstructured text for further application downstream.
[0048] The relation extraction model can be considered as a model that performs relation extraction to determine the target association relationship between objects. The relation extraction model can be a pre-trained language model, which can refer to a model that is pre-trained on a large-scale (such as a data volume greater than a set threshold) data set, such as a neural network model. The pre-trained language model can use a large-scale data set to initialize the model parameters, and then adapt the model to a specific task through fine-tuning or transfer learning. The pre-trained language model can be a model that uses the Transformer encoder architecture to learn the bidirectional representation of text. With the support of a large amount of unlabeled data, the rich semantic information in the text is fully learned through multiple deep Transformers and their huge parameter volume, thereby obtaining a high-quality general representation, and only simple fine-tuning in downstream tasks is required to obtain excellent results. The relation extraction model in the present invention can be implemented using a bidirectional Transformer encoder, and a high-quality deep context coding representation is generated for the input sequence by stacking multiple layers of Transformer encoders.
[0049] For the relation extraction task of a relation extraction model, explicitly providing entity information to the relation extraction model enables the model to better recognize entities and improves its ability to learn entities during the model's training phase. Therefore, before encoding using the relation extraction model, the data to be recognized must first be preprocessed. This preprocessing is tailored to the relation extraction model, allowing it to better recognize entities.
[0050] The first input sequence can be considered to be the input sequence formed after preprocessing the data to be identified. The preprocessing of the identification data in this operation can include adding a set tag to the data to be identified. The set tag can be considered to be a pre-set tag that serves as an identifier. The set tag can be added to the data to be identified at the location where the tag is required. In this embodiment, the added set tag can serve as a form of entity information used to implement tagging.
[0051] In one embodiment, preprocessing the data to be identified to be adapted to the relation extraction model to obtain a first input sequence includes:
[0052] Adding a set mark to the data to be identified to obtain a first input sequence;
[0053] The setting flags include one or more of the following:
[0054] a start mark located at the start position of the data to be identified;
[0055] an end marker located at the end position of the data to be identified;
[0056] Entity markers located on both sides of the object.
[0057] A start tag can be thought of as marking the beginning of the data to be identified. It's added to the beginning of the data to be identified. An end tag can be thought of as marking the end of the data to be identified. The start and end positions are the beginning and end positions of the data to be identified, respectively. An entity tag can be thought of as a tag used to identify an entity. Different entities correspond to different entity tags. Entity tags are added to both sides of the marked object.
[0058] S130: Process the first input sequence and the relationship type set respectively through the relationship extraction model to obtain a target relationship among the candidate relationship relationships corresponding to the data to be identified.
[0059] In this operation, the first input sequence and the relationship type set can be input into the relationship extraction model respectively. The relationship extraction model processes the first input sequence and the relationship type set respectively to determine a target relationship for the first input sequence from a plurality of candidate relationships included in the relationship type set.
[0060] The output of the relationship extraction model is not limited here. For example, the model can directly output the target relationship. Alternatively, the model can output the target relationship and its confidence score. Alternatively, the model can output multiple target relationships arranged in order of importance. Alternatively, the model can output multiple target relationships arranged in order of importance and the corresponding confidence score for each relationship. Furthermore, the model can output objects in the data to be identified.
[0061] For example, if the data to be identified is "Company Xiaoming has supplied Company Xiaowang for many years and has a good cooperative relationship," the relationship extraction model could output the target relationship between upstream supplier and company. Alternatively, the target relationship between Company Xiaoming and Company Xiaowang could be upstream supplier and company. This includes both the object and target relationships.
[0062] The relationship identification method provided by the present invention obtains data to be identified, including unstructured data, and a set of relationship types comprising multiple candidate relationships. The data to be identified is preprocessed and then, together with the set of relationship types, is run through a relationship extraction model to determine target relationships. This method achieves automated identification of unstructured data to determine target relationships between objects represented by the unstructured data.
[0063] Based on the above embodiment, a modified embodiment of the above embodiment is proposed. It should be noted that, in order to simplify the description, only the differences from the above embodiment are described in the modified embodiment.
[0064] In one embodiment, the step of processing the first input sequence and the set of relationship types by the relationship extraction model to obtain a target relationship among the candidate relationship relationships corresponding to the data to be identified includes:
[0065] The following operations are performed through the relationship extraction model:
[0066] processing the first input sequence to obtain a first relation representation vector of the first input sequence;
[0067] For each candidate association relationship in the relationship type set, determining a first semantic vector corresponding to the candidate association relationship;
[0068] For each first semantic vector, determining a cosine similarity measure between the first semantic vector and the first relationship representation vector;
[0069] The candidate association relationship corresponding to the largest cosine similarity measure among the cosine similarity measures is determined as the target association relationship of the data to be identified.
[0070] The first relation representation vector can be considered the relation representation vector of the first input sequence. This relation representation vector may be a combination of the global sentence vector representation of the first input sequence and the entity representations of the entities in the first input sequence. The sentence vector representation can be considered a vector that can represent the first input sequence. The entity identifier can be considered a vector that can represent the entity.
[0071] The first semantic information may be considered as semantic information of the candidate association relationship. The semantic information may be information representing the semantics of the candidate association relationship. The cosine similarity metric may be considered as a distance metric performed using cosine similarity.
[0072] Metric learning aims to model the distance between sample features, or similarity. Through training and learning, the distance between similar samples (or samples of the same type) decreases, while the distance between dissimilar samples (or samples of different types) increases, thereby enabling the model to better distinguish between different samples. In deep metric learning, deep neural networks extract sample feature vectors, and then use distance metrics to measure the similarity between samples. Cosine similarity is an extension of the inner product of vectors. It normalizes the length of the vectors based on the inner product, so it does not consider the length of the vectors. Cosine similarity focuses on the directional differences between vectors in multidimensional space. The cosine similarity function has a range of [-1, 1]. Larger values indicate greater similarity between two vectors, and vice versa.
[0073] In this embodiment, the first input sequence is encoded through a relation extraction model, and a global sentence vector representation and an entity representation of the entity in the first input sequence are obtained from the encoded output vector. After processing the two, a first relation representation vector is obtained to facilitate cosine similarity measurement.
[0074] Each candidate association relationship in the relationship type set is traversed, and for each candidate association relationship, the candidate association relationship is encoded, and the first semantic information of the candidate association relationship is extracted from the vector output by the encoding.
[0075] For each candidate association relationship, a cosine similarity measurement is performed between the first semantics corresponding to the candidate association relationship and the first relationship representation vector to determine the distance between the two.
[0076] The cosine similarity measures corresponding to all candidate association relationships in the relationship type set are sorted by size, such as eliminating them from large to small, selecting the largest cosine similarity measure, and determining the candidate association relationship corresponding to the largest cosine similarity measure as the target association relationship.
[0077] In one embodiment, processing the first input sequence to obtain a first relation representation vector of the first input sequence includes:
[0078] Encoding the first input sequence to obtain a sequence-encoded vector;
[0079] Obtaining a first latent vector corresponding to the start marker in the sequence encoded vector;
[0080] Determining the first latent vector as the sentence vector representation of the data to be recognized;
[0081] For each object in the to-be-identified data, determine a second latent vector for a target tag corresponding to the object, where the number of the second latent vectors is equal to the number of tags for the corresponding object, and the target tag represents the following information in a marked form: the object corresponding to the target tag in the first input sequence and the entity tag associated with the object;
[0082] For each object, determining the average value of all second latent vectors corresponding to the object;
[0083] For each of the average values, add the average value to the sentence vector representation to obtain a third latent vector;
[0084] Combining the third latent vectors to obtain a fourth latent vector;
[0085] Determine a first relationship representation vector corresponding to the fourth latent vector.
[0086] The sequence encoding vector can be thought of as the vector formed by encoding the first input sequence. The latent vector can be the output vector after the input text sequence (such as words) is processed by the encoder in each Transformer layer of the model. The latent vector captures the contextual information of the input text and can be used for downstream tasks.
[0087] The first latent vector may be the latent vector of the start tag. The second latent vector may be the latent vector of the target tag. The target tag may be the tag corresponding to the object and entity tags in the first input sequence when input into the relation extraction model.
[0088] The input to a relation extraction model is a sequence of text, typically represented as tokens (i.e., target tags). A token is the basic unit of text data, typically a word, subword, or character. The text is segmented into a series of tokens. This approach allows the model to handle unseen words.
[0089] The number of the second latent vector of an object is equal to the number of each tag of the object, that is, the number of tokens corresponding to the object and the entity tag corresponding to the object is equal.
[0090] In this embodiment, the first latent vector corresponding to the start marker in the sequence encoding vector is used as the sentence vector. During the encoding process of the first input sequence, the global features of the first input sequence are encoded into the start marker, so the first latent vector corresponding to the start marker can be used as the sentence vector.
[0091] Traverse each object in the data to be identified. Determine the second latent vector of the object and the corresponding entity token. Calculate the average of all the second latent vectors of the object.
[0092] For the average value corresponding to each object, the average value is added to the sentence vector representation to obtain a composite feature that combines the sentence vector representation and the entity representation, namely the third latent vector.
[0093] The fourth latent vector can be formed by combining the third latent vector. The method of combining is not limited here, and the combination can be achieved through vector concatenation operation.
[0094] After determining the fourth latent vector, the fourth latent vector can be substituted into a formula to determine the first relationship representation vector. This formula can be a formula related to the fourth latent vector and model parameters of the relationship extraction model, and is not limited here.
[0095] In one embodiment, determining, for each candidate association relationship in the relationship type set, a first semantic vector corresponding to the candidate association relationship includes:
[0096] For each candidate association relationship in the relationship type set, encode the candidate association relationship to obtain an encoded relationship vector, wherein the starting position of the candidate association relationship includes a start marker;
[0097] The fifth latent vector corresponding to the start tag in the relationship encoding vector is used as the first semantic vector corresponding to the candidate association relationship.
[0098] The encoded relation vector can be considered the vector obtained by encoding the candidate relation. A start marker can be added to the starting position of the candidate relation. The start marker added to the candidate relation and the start marker added to the data to be identified can be the same marker or different markers, without limitation here, such as the [CLS] marker.
[0099] In this embodiment, for each candidate relationship, the relationship extraction model is used to encode the candidate relationship to obtain a relationship encoding vector. The fifth latent vector corresponding to the start tag is then extracted from the relationship encoding vector to serve as the first semantic vector of the candidate relationship. For example, the latent state corresponding to the [CLS] tag is used as the vector representation of the candidate relationship, i.e., the first semantic vector.
[0100] In one embodiment, the training operation of the relationship extraction model includes:
[0101] Acquire an unstructured data set and a training relationship set, wherein the unstructured data set includes unstructured data for identifying association relationships between training objects, and the training relationship set includes multiple training association relationships;
[0102] marking the positive correlation between the training object and the training object in the unstructured data;
[0103] Selecting negative association relationships between the training objects from the training relationship set;
[0104] Preprocessing the unstructured data to be adapted to the model to be trained to obtain a second input sequence;
[0105] The model to be trained is trained using the second input sequence, the positive correlation relationship, and the negative correlation relationship to obtain a relationship extraction model.
[0106] An unstructured dataset can be considered the set of training samples from the model training phase. Training samples can be unstructured data used to identify relationships between training objects. Training objects can be objects involved in the model training phase. A training relationship set can be considered the set of labels assigned to the unstructured dataset. Training relationships can be considered the labels used during the training phase. Training relationships can be understood as the relationships used during the model training phase.
[0107] In order to adapt to the characteristics of corporate relationships in the field of supply chain finance, the present invention uses existing daily operational data (which may be operational data that is determined to be associated with the relationship between enterprises) and Internet data packets as data sources to construct an unstructured data set. Through the data interface package (the data interface package includes multiple interfaces for obtaining data), news reports, market data, etc. of the enterprise can be obtained, including sufficient natural language text that can be used to identify the relationship between enterprises, which can be used to construct an unstructured data set for relationship extraction dedicated to the field of supply chain finance. The acquisition, storage, use, processing, etc. of data in the present invention comply with the provisions of relevant laws and regulations, and are authorized by the party to which the data belongs.
[0108] The present invention sets a variety of tags for unstructured data sets, such as 11 tags, to describe the relationship between enterprises, including: upstream supplier-enterprise, downstream distributor-enterprise, core enterprise-enterprise, issuer-enterprise, recipient-enterprise, acceptor-enterprise, pledgee-enterprise, affixer-enterprise, confirmer-enterprise, government agency-enterprise, and NA. The set of the above 11 tags is called the training relationship set of the unstructured data set, and the training relationship set can also be used as a relationship type set in the model application stage. In addition to the tags in the training relationship set, the data in the relationship type set can also add new tags. Subsequently, the data is annotated according to certain rules, and the annotation content includes entities, such as enterprise 1, enterprise 2, and the relationship between the two. That is, the training objects (i.e., enterprise 1, enterprise 2) in the unstructured data and the positive relationship between the training objects are annotated. For example, the natural language text "Xiao Ming Company has supplied goods to Xiao Wang Company for many years and the cooperation is good", Table 1 is a schematic table for labeling unstructured data provided in an embodiment of the present application. The labeling content is shown in Table 1. Xiao Ming Company can be an expanded enterprise.
[0109] Table 1
[0110]
[0111] Based on the characteristics of the supply chain finance sector and its inter-enterprise relationships, this paper constructs a fine-grained, unstructured dataset for relationship extraction in the supply chain finance sector, drawing on business data and internet data. This dataset, combined with a relationship extraction model, better matches the customer mining needs of the supply chain finance industry. The relationship extraction model implements relationship extraction and metric learning during use, improving the accuracy of determining target associations and enhancing the model's ability to perceive and identify different categories, resulting in better classification boundaries between categories and greater adaptability to the fine-grained dataset proposed in this paper.
[0112] Positive relationships can be considered positive examples of training objects. That is, relationships between training objects in unstructured data. Taking Table 1 as an example, if the unstructured data is the text in the first column of Table 1, the positive relationships are the relationships shown in Table 1.
[0113] Negative relationships can be relationships other than positive relationships selected from the training relationship set and can be used as negative samples to train the model, such as downstream distributors and enterprises.
[0114] In this embodiment, the unstructured data is preprocessed to adapt it to the model to be trained. The preprocessing methods are the same as those used in the application phase and are not described in detail here. For example, a set flag is added to obtain a second input sequence. The second input sequence can be an input sequence formed after preprocessing the unstructured data. The model to be trained can be considered to be a model to be trained to obtain a relationship extraction model, such as a neural network model.
[0115] In this embodiment, a second input sequence, a positive correlation, and a negative correlation can be input into the model to be trained. The model to be trained processes the second input sequence, the positive correlation, and the negative correlation, respectively, to determine a loss function to adjust the model to be trained, thereby obtaining a trained relationship extraction model. During the training process, the model can bring the second input sequence closer to the positive correlation, i.e., reduce the distance, and move the second input sequence further away from the negative correlation, i.e., increase the distance.
[0116] The processing operation of the model to be trained on the second input sequence is the same as the processing operation of the relationship extraction model on the first input sequence. The processing operation of the model to be trained on positive and negative associations is the same as the operation of the relationship extraction model on candidate associations, which is not limited here.
[0117] In one embodiment, the training of the to-be-trained model using the second input sequence, the positive association relationship, and the negative association relationship to obtain a relationship extraction model includes:
[0118] The following operations are performed using the model to be trained:
[0119] processing the second input sequence to obtain a second relation representation vector corresponding to the second input sequence;
[0120] Determining a second semantic vector corresponding to the positive association relationship;
[0121] Determining a third semantic vector corresponding to the negative association relationship;
[0122] respectively determining a first similarity measure between the second semantic vector and the second relationship representation vector and a second similarity measure between the third semantic vector and the second relationship representation vector;
[0123] Determining a triplet loss function corresponding to the first similarity metric and the second similarity metric;
[0124] The model parameters of the model to be trained are adjusted based on the triplet loss function until an end condition is met.
[0125] The second relationship representation vector can be considered as the relationship representation vector of the second input sequence. This relationship representation vector can be a combination of the global sentence vector representation of the second input sequence and the entity representation of the entity in the second input sequence. The determination method refers to the determination of the first relationship representation vector and is not repeated here. The second semantic vector can be a vector corresponding to the semantics of a positive association relationship. The third semantic vector can be a vector corresponding to the semantics of a negative association relationship. The determination method of the second and third semantic vectors refers to the first semantic vector and is not repeated here.
[0126] The first similarity measure may be a measure of similarity between the second semantic vector and the second relationship representation vector. The second similarity measure may be a measure of similarity between the third semantic vector and the second relationship representation vector. In this embodiment, the first similarity measure and the second similarity measure may be determined by cosine similarity.
[0127] The triplet loss function can be considered a loss function for metric learning. Based on the contrastive loss, it defines anchor samples (anchors), sets positive and negative samples around the anchor samples, and further considers the distance relationship between the positive and negative sample pairs and the anchor samples. In the triplet loss function, each triplet consists of three samples: an anchor sample, a positive sample, and a negative sample. The anchor sample and the positive sample belong to the same category, while the anchor sample and the negative sample belong to different categories. The goal of the triplet loss function is to minimize the distance between the anchor sample and the positive sample while maximizing the distance between the anchor sample and the negative sample.
[0128] In this embodiment, after determining the first and second similarity metrics, they can be substituted into the calculation function of the triple loss function to obtain the triple loss function. The model parameters of the to-be-trained model are then adjusted based on the triple loss function until a termination condition is met. The termination condition can be the number of iterations or the convergence of the loss value corresponding to the triple loss function. The termination condition is not limited here.
[0129] The present invention is described below by way of example. The association relationship identification method provided by the present invention can be considered as a supply chain customer mining method based on relationship extraction. A data set, i.e., an unstructured data set, is constructed based on the characteristics of the association relationships between enterprises in the field of supply chain finance. A pre-trained language model based on the encoder architecture of Transformer to learn the bidirectional representation of text and a relationship extraction method of metric learning are designed and trained and optimized on the data set. The relationship extraction model is used to mine the association relationships between enterprises, and the confidence level, i.e., the cosine similarity metric and the corresponding association relationship, is outputted to facilitate the mining of potential customers in the supply chain based on the association relationship.
[0130] Among them, metric learning aims to model the distance between sample features, that is, similarity. Through training and learning, the distance between similar samples (or samples of the same type) is reduced, while the distance between dissimilar samples (or samples of different types) is increased, so that the model has a stronger ability to distinguish samples of different categories.
[0131] The supply chain finance customer mining method proposed in this invention, based on a pre-trained relationship extraction model, extracts relationships from unstructured data between enterprises and finds the associations between supply chain enterprises, so as to enable more small and medium-sized enterprises to participate in supply chain financing.
[0132] Existing technologies rely on a large amount of characteristic data from enterprises, but in real-world scenarios, conducting sufficient and effective research and modeling of enterprises in batches requires a lot of manpower and material resources, often accompanied by a long cycle. The present invention can obtain unstructured data sets through a data interface, reducing manpower and material resources. At the same time, customer mining methods based on traditional machine learning models are limited by their low accuracy and model performance, making it difficult to adapt to and meet the complex and changing needs of today's actual business scenarios. The relationship extraction model provided by the present invention has a high degree of accuracy in determining target association relationships.
[0133] Existing knowledge graph-based technologies, however, are limited to relatively broad relationships such as "cooperation" and "competition." This leads to a certain mismatch between practical applications and business needs, resulting in poor adaptability in complex scenarios. Furthermore, they require complex graph algorithms designed to mine potential customers, which is highly dependent on the business scenario, resulting in poor portability. The relationship type set provided by this invention allows for flexible data updates. As long as the relationship extraction model captures the semantic information in the relationship type set, the target relationship can be determined.
[0134] First, the core of the relationship extraction task lies in enabling the model to correctly understand the semantics of the text and, consequently, to correctly classify it. Therefore, the quality of the text encoding directly impacts the model's performance. Second, as a multi-classification task, the performance of the relationship extraction model is positively correlated with the classifier's ability to distinguish between different categories. In relationship extraction tasks, there is often a strong commonality between training samples and labels. For example, the semantics of the sentence "Xiao Ming Company has supplied Xiao Wang Company for many years, and the cooperation is excellent" is highly similar to its label "upstream supplier - enterprise." Therefore, capturing label semantics enhances the model's ability to infer relationship types.
[0135] Based on the above two points, the technical solution of the present invention is summarized as follows: First, the present invention adopts a relation extraction model to encode the text and its positive and negative label text; secondly, after using the relation extraction model to encode the global features of the input sentence (i.e., the sentence vector representation) and the local features corresponding to the two entities (i.e., the average value of all second latent vectors corresponding to the objects), a classification method based on metric learning is used to design a model around the triple loss function in deep metric learning. The pre-trained language model is used to encode the context sentence as the anchor sample representation, and the corresponding positive and negative label texts are encoded as positive and negative sample representations to capture the label semantics. The cosine similarity is used as the distance metric, and the triple loss function is used to train the model so that the positive label representation is closer to the context sentence than the negative label representation, thereby improving the model's ability to distinguish different categories and improving the generalization performance.
[0136] Figure 2 This is a schematic diagram of the overall process of determining an association relationship provided by an embodiment of the present invention, see Figure 2 , the process of determining the association relationship is as follows:
[0137] (1) Constructing a dataset, namely an unstructured dataset and a training relation set:
[0138] Based on the characteristics of enterprise relationships in the supply chain finance field, we predefine relationship types, form a training relationship set, and build a dedicated unstructured dataset. We use enterprise operational data as the data source, and also use open source data packages as additional data supplements.
[0139] (2) Input preprocessing: Add start tags, end tags, and entity tags to the text. During the model training phase, the unstructured data is preprocessed to adapt it to the model to be trained, thereby obtaining a second input sequence. During the model application phase, the data to be identified is preprocessed to adapt it to the relation extraction model, thereby obtaining a first input sequence.
[0140] (3) Obtain the relationship representation vector:
[0141] Use a pre-trained language model to encode the text, outputting sentence vectors and entity vectors, which are combined into a relation representation vector. During the application phase, the pre-trained language model can be a relation extraction model, with the sentence vector corresponding to the sentence vector representation, the entity vector corresponding to the average value, and the relation representation vector corresponding to the first relation representation vector. During the application phase, the pre-trained model can be the model to be trained, with the relation representation vector corresponding to the second relation representation vector.
[0142] (4) Obtaining label semantic vectors:
[0143] Use the pre-trained language model to encode the corresponding positive and negative label texts, and output the semantic vectors of the positive and negative labels, namely the second semantic vector and the third semantic vector; in the application stage, the candidate association relationship is encoded through the relationship extraction model to obtain the corresponding first semantic vector.
[0144] (5) Calculate similarity measure:
[0145] During the model training phase, similarity metrics between the second relationship representation vector and the second semantic vector corresponding to the positive label, and between the second relationship representation vector and the third semantic vector corresponding to the negative label, namely, the first similarity metric and the second similarity metric, are calculated respectively.
[0146] In the model application stage, a similarity measure between the first relationship representation vector and the first semantic vector, ie, a cosine similarity measure, is determined so as to select an association relationship corresponding to the maximum cosine similarity measure.
[0147] (6) Model training and prediction:
[0148] In the model training phase, the model is trained based on a triplet loss function determined by the first similarity metric and the second similarity metric.
[0149] During the model application phase, the maximum cosine similarity metric and the corresponding association relationship are determined.
[0150] The following is an exemplary description of the pre-processing operation in the above process. Figure 3 is a schematic diagram of a preprocessed input sequence provided by an embodiment of the present invention. The input sequence in the application phase is the first input sequence. The input sequence in the training phase is the second input sequence. Figure 3 , additional special tags are introduced, namely the start tag, end tag and entity tag, and the entity tags are explicitly added to both sides of the entity. For example, the entity tags on both sides of Enterprise 1 and Enterprise 2 are [E1], [ / E1] and [E2], [ / E2] respectively. At the same time, the [CLS] tag (i.e., start tag) and the [SEP] tag (i.e., end tag) are added at the beginning and end of the sentence (such as the data to be identified or unstructured data) as the start and end of the sequence, respectively, to obtain the final input sequence T. Such as the first input sequence or the second input sequence.
[0151] The following describes the operation of obtaining the relation representation vector:
[0152] Figure 4 This is a coding diagram provided in the embodiment of the present application, see Figure 4In the present invention, the model can encode text. In the application stage, the relation extraction model encodes the first input sequence corresponding to the data to be recognized. In the training stage, the model to be trained encodes the second input sequence corresponding to the unstructured data. The first input sequence and the second input sequence are denoted as input sequences. Assume that the length of the input sequence T is k, and it is input into the model for encoding to obtain the output sequence, i.e., the sequence encoded vector: Encoder(T) = [H [CLS] , H1, ..., H k-1 ].
[0153] In the output vector, let H [CLS] is the hidden vector corresponding to the [CLS] tag (start tag), such as the first hidden vector, H i To H j is the latent vector of each tag (i.e., target tag, e.g., from Toki to Tokj, corresponding to [E1]Enterprise 1[ / E1]), such as the second latent vector, H m To H n is the second latent vector corresponding to each tag of enterprise 2. [CLS] 、H i 、H j 、H m 、H n ∈R h , that is, the h-dimensional real vector space, where h is the hidden state dimension of the Transformer output. [CLS] , directly record it as the global sentence feature of the input text, that is, the sentence vector representation H0: H0 = H [CLS] .
[0154] Take the average value of the latent vectors corresponding to all tags of the enterprise (for example, for each object, determine the average value of all second latent vectors corresponding to the object) as the local entity feature H e1 and H e2 , that is, the characteristics of enterprise 1 and enterprise 2:
[0155]
[0156] For the sentence vector representation H0, and the two entities represent H e1 , H e2 , the information is combined by using the vector addition operation method, such as for each of the average values, the average value is added to the sentence vector representation to obtain the third latent vector:
[0157] H1=H e1 +H0;
[0158] H2=H e2 +H0.
[0159] Among them, H1, H2∈R h , which fuses sentence vector representation and entity representation and can be regarded as a composite feature. Subsequently, the two composite features are combined, that is, the third latent vectors are combined to obtain the fourth latent vector: R = concat(H1, H2).
[0160] Among them, concat(·) represents the vector concatenation operation. Finally, a fully connected layer is used to perform a linear transformation on R to obtain the final relational representation X, such as the first relational representation vector: X = tanh(W×R+b).
[0161] Where W∈R h×2h is the linear transformation matrix, b∈R h is the bias term. W and b are the model parameters of the model. The values of W and b are adjusted during the model training phase to obtain the relationship extraction model.
[0162] The operations of obtaining the first relationship representation vector in the model application phase and obtaining the second relationship representation vector in the model training phase may be the above-mentioned process of obtaining X.
[0163] The following describes the operation of obtaining the semantic vector:
[0164] Let Y be the set of relation types in the application phase or the set of relations used for training.
[0165] In the model training phase, given an input sequence T and its corresponding label y∈Y, y is used as a positive label, that is, a positive association relationship, and a negative label y'∈Y\y is randomly selected for it, that is, a label other than the positive label is selected from the training relationship set as a negative label. Let the texts corresponding to the positive and negative sample labels be T respectively. p , T n , in T p , T n The [CLS] tag is added to the starting position of the positive association relationship and the negative association relationship, respectively. Suppose the text lengths of the two are k p and k n In order to capture the rich semantics contained in the sample labels, the model to be trained is also used to encode the positive and negative label texts, that is, to encode the positive or negative association relationship, and obtain the encoded vector for training:
[0166]
[0167] Use the hidden state corresponding to the [CLS] tag in the positive and negative label text as its corresponding vector representation X p , X n , that is, the second semantic vector or the third semantic vector:
[0168]
[0169] where X p , X n ∈R h Therefore, the positive and negative label text T p , T n The text corresponding to the input sentence, that is, the input vector T, is mapped into the same embedding space, and the distance and similarity between the three can be measured using distance metrics.
[0170] The following describes the determination of the similarity metric, corresponding to the operations of determining the cosine similarity metric in the application phase and determining the first similarity metric and the second similarity metric in the training phase.
[0171] The present invention uses cosine similarity as a distance metric to quantitatively characterize the distance and similarity between the anchor sample representation X and the positive and negative label text representations Xp and Xn, and uses the following formula for calculation:
[0172]
[0173] Among them, X l is the vector representation of the positive or negative label text, that is, X p or X n , X i Represents the components of each dimension of the anchor sample vector X (i.e., the second relationship representation vector), X l,i Represents X p or X n The first similarity measure Dcosine(X,X p ) and the second similarity measure Dcosine(X,X n ) respectively characterizes the similarity between the anchor sample representation and the positive and negative sample representations. The model parameters can be optimized through the loss function so that Dcosine(X,X p ) increases and Dcosine(X,X n ) decreases.
[0174] In the model application stage, X l It can be the first semantic information, and the corresponding anchor sample vector X is the first relationship representation vector, Dcosine(X,X l ) is the cosine similarity metric.
[0175] The following describes the model training and prediction:
[0176] The triplet loss function is calculated as follows:
[0177] L triplet =max(0, D cosine(X, X p ) 2 -D cosine (X, X n ) 2 +α.
[0178] α is a hyperparameter that indicates how much smaller the distance between the anchor sample and the positive sample should be than the distance between the anchor sample and the negative sample. cosine (X, X p ) 2 It can be considered as Dcosine(X,X p ) squared. D cosine (X, X n ) 2 Similarly, for Dcosine(X,X n ) squared.
[0179] In the prediction process of the model, the relationship type is predicted by ranking, that is, the association relationship is determined. Specifically, let the sample T, that is, the relationship representation obtained by the first input sequence through the relationship extraction model is X, that is, the first relationship representation vector, for each candidate label in the relationship type set, that is, the candidate association relationship y i ∈Y, also use the relation extraction model to get y i The context representation of the corresponding text X i, That is, the first semantic vector, calculate X and X i The cosine similarity Dcosine(X,X i ), that is, determining the cosine similarity measure between the first semantic vector and the first relationship representation vector. And ranking them in descending order of similarity, when predicting, output the largest Dcosine(X,X i ) corresponding to the label, that is, the candidate association relationship corresponding to the largest cosine similarity measurement among the cosine similarity measurements is determined as the target association relationship of the data to be identified.
[0180] Because the above method captures the semantics of the label and ranks it based on the similarity between the label semantics and the semantics of the context sentence, this means that the model is no longer limited to the predefined label set, that is, the training relationship set. For a specific relationship type in the prediction process (this relationship type is included in the relationship type set), regardless of whether its sample size is sufficient or even whether it has training samples, as long as the model captures the text semantic representation corresponding to the relationship type, that is, the first semantic vector, the model can rank it accordingly and make a prediction. Based on the above characteristics, this method is very suitable for handling relationship extraction tasks with relatively small sample sizes.
[0181] The following describes the training process of the training model. Figure 5This is a schematic diagram of the training process of a model to be trained provided in an embodiment of the present application. Figure 5 , the second input sequence, the positive label (i.e., positive correlation) and the negative label (i.e., negative correlation) are input into the model to be trained. In the present invention, the model to be trained is one, Figure 5 The second input sequence is processed separately, and the positive label and negative label dimensions are processed twice, showing the model to be trained. The two models to be trained are the same model to be trained. The second input sequence is encoded by the model to be trained to obtain a second relationship representation vector, that is, the anchor representation. The positive sample is processed by the model to be trained to obtain a second semantic vector, that is, the positive label representation. The negative sample is processed by the model to be trained to obtain a third semantic vector, that is, the negative label representation. The triple loss function is determined by the positive label representation, the negative label representation and the anchor representation, that is, the first similarity measure between the second semantic vector and the second relationship representation vector and the second similarity measure between the third semantic vector and the second relationship representation vector are determined respectively; the triple loss function corresponding to the first similarity measure and the second similarity measure is determined to train the model to be trained. Figure 5 In the above example, upstream suppliers-enterprises are taken as positive labels, and downstream distributors-enterprises are taken as negative labels.
[0182] Figure 6 This is a flow chart of a relationship extraction model provided in an embodiment of the present application to determine target association relationships, see Figure 6 The first input sequence is input into the relation extraction model, and the relation extraction model encodes the first input sequence to obtain a first relation representation vector. For each candidate association relationship in the relation type set, a first semantic vector corresponding to the candidate association relationship is determined. Figure 6 The relationship extraction model is shown twice in the figure from the dimensions of processing the first input sequence and processing the candidate association relationship. The relationship extraction model of the present invention is a model. The cosine similarity measure is determined for each first semantic vector and the first relationship representation vector. That is, the candidate association relationship in the figure: the cosine similarity measure determined by the first semantic vector corresponding to the upstream supplier-enterprise and the first relationship representation vector is 0.83. The candidate association relationship: the cosine similarity measure determined by the first semantic vector corresponding to the acceptor-enterprise and the first relationship representation vector is 0.39. The candidate association relationship: the cosine similarity measure determined by the first semantic vector corresponding to the downstream distributor-enterprise and the first relationship representation vector is 0.07. The candidate association relationship corresponding to the largest cosine similarity measure is the target association relationship, and the target association relationship, that is, the upstream supplier-enterprise, is finally output.
[0183] This paper extracts relationships during customer discovery and recommendation in the supply chain finance sector. Based on the characteristics of enterprise relationships within this sector, it constructs a specialized dataset—an unstructured dataset, a set of relationship types, and a training set of relationships—that aligns with business scenarios. Using relationship extraction technology, it efficiently extracts relationships between enterprises from unstructured data, saving significant labor costs. Business personnel can use the output of the relationship extraction model to determine inter-enterprise partnerships and identify enterprises with financing needs.
[0184] The present invention uses a pre-trained language model and metric learning as components of the relationship extraction model. While capturing the contextual information of the input sentence, it also utilizes its label text information, making the model more perceptive of the various relationship types in the data set. Using triple loss as the training target of the model, it makes samples of the same category closer together and samples of different categories farther apart, and also creates better classification boundaries between categories, thereby greatly improving the recognition effect of relationship extraction and assisting business personnel in customer mining. At the same time, metric learning makes the model no longer strongly dependent on a large number of training samples, and it can still maintain good recognition capabilities for categories with few samples, reducing the manual labeling costs brought by a large amount of training data.
[0185] Relationship extraction is an upstream task of the knowledge graph. The supply chain relationship extraction method proposed in this invention provides a basis for the subsequent establishment of supply chain business knowledge graphs and risk prediction applications.
[0186] Example 2
[0187] Figure 7 1 is a schematic diagram of the structure of an association relationship identification device provided according to the second embodiment of the present invention. The device can be integrated into an electronic device. Figure 7 As shown, the device includes:
[0188] An acquisition module 710 is configured to acquire data to be identified and a set of relationship types, wherein the data to be identified includes unstructured data used to identify relationships between objects, and the set of relationship types includes a plurality of candidate relationships;
[0189] A preprocessing module 720 is configured to preprocess the data to be identified so as to be adapted to the relationship extraction model to obtain a first input sequence;
[0190] The processing module 730 is configured to process the first input sequence and the relationship type set respectively through the relationship extraction model to obtain a target relationship among the candidate relationships corresponding to the data to be identified.
[0191] In one embodiment, the pre-processing module 720 is specifically configured to:
[0192] Adding a set mark to the data to be identified to obtain a first input sequence;
[0193] The setting flags include one or more of the following:
[0194] a start mark located at the start position of the data to be identified;
[0195] an end marker located at the end position of the data to be identified;
[0196] Entity markers located on both sides of the object.
[0197] In one embodiment, the processing module 730 includes:
[0198] The following operations are performed through the relationship extraction model:
[0199] a processing unit, configured to process the first input sequence to obtain a first relation representation vector of the first input sequence;
[0200] A first determining unit, configured to determine, for each candidate association relationship in the relationship type set, a first semantic vector corresponding to the candidate association relationship;
[0201] a second determining unit, configured to determine, for each first semantic vector, a cosine similarity measure between the first semantic vector and the first relationship representation vector;
[0202] The third determining unit is configured to determine the candidate association relationship corresponding to the largest cosine similarity measure among the cosine similarity measures as the target association relationship of the data to be identified.
[0203] In one embodiment, the processing unit is specifically configured to:
[0204] Encoding the first input sequence to obtain a sequence-encoded vector;
[0205] Obtaining a first latent vector corresponding to the start marker in the sequence encoded vector;
[0206] Determining the first latent vector as the sentence vector representation of the data to be recognized;
[0207] For each object in the to-be-identified data, determine a second latent vector for a target tag corresponding to the object, where the number of the second latent vectors is equal to the number of tags for the corresponding object, and the target tag represents the following information in a marked form: the object corresponding to the target tag in the first input sequence and the entity tag associated with the object;
[0208] For each object, determining the average value of all second latent vectors corresponding to the object;
[0209] For each of the average values, add the average value to the sentence vector representation to obtain a third latent vector;
[0210] Combining the third latent vectors to obtain a fourth latent vector;
[0211] Determine a first relationship representation vector corresponding to the fourth latent vector.
[0212] In one embodiment, the first determining unit is specifically configured to:
[0213] For each candidate association relationship in the relationship type set, encode the candidate association relationship to obtain an encoded relationship vector, wherein the starting position of the candidate association relationship includes a start marker;
[0214] The fifth latent vector corresponding to the start tag in the relationship encoding vector is used as the first semantic vector corresponding to the candidate association relationship.
[0215] In one embodiment, the relationship extraction model is obtained by training through a training module, and the training operations performed by the training model include:
[0216] Acquire an unstructured data set and a training relationship set, wherein the unstructured data set includes unstructured data for identifying association relationships between training objects, and the training relationship set includes multiple training association relationships;
[0217] marking the positive correlation between the training object and the training object in the unstructured data;
[0218] Selecting negative association relationships between the training objects from the training relationship set;
[0219] Preprocessing the unstructured data to be adapted to the model to be trained to obtain a second input sequence;
[0220] The model to be trained is trained using the second input sequence, the positive correlation relationship, and the negative correlation relationship to obtain a relationship extraction model.
[0221] In one embodiment, the training model is specifically used to:
[0222] The following operations are performed using the model to be trained:
[0223] processing the second input sequence to obtain a second relation representation vector corresponding to the second input sequence;
[0224] Determining a second semantic vector corresponding to the positive association relationship;
[0225] Determining a third semantic vector corresponding to the negative association relationship;
[0226] respectively determining a first similarity measure between the second semantic vector and the second relationship representation vector and a second similarity measure between the third semantic vector and the second relationship representation vector;
[0227] Determining a triplet loss function corresponding to the first similarity metric and the second similarity metric;
[0228] The model parameters of the model to be trained are adjusted based on the triplet loss function until an end condition is met.
[0229] The association relationship identification device provided in the embodiment of the present invention can execute the association relationship identification method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0230] Example 3
[0231] Figure 8 Schematic diagram of the structure of an electronic device that implements the association relationship identification method of an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0232] like Figure 8 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by the at least one processor 11, and the computer program is executed by the at least one processor 11 so that the at least one processor 11 can execute the method provided by the present invention.
[0233] The processor 11 can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 12 or a computer program loaded from a storage unit 18 into a random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0234] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0235] The processor 11 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the association relationship identification method.
[0236] In some embodiments, the association relationship identification method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the association relationship identification method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the association relationship identification method in any other appropriate manner (for example, by means of firmware).
[0237] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0238] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0239] In the context of the present invention, a computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the association relationship identification method provided by the present invention when executed.
[0240] The present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the computer program implements the method provided according to the embodiment of the present invention.
[0241] Computer-readable storage medium can be a tangible medium that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage medium can include but is not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device or any suitable combination of the foregoing.
[0242] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD monitor)) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0243] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0244] A computing system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship arises through computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and virtual private server (VPS) services.
[0245] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0246] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for identifying association relationships, characterized in that: include: Acquire data to be identified and a set of relationship types, wherein the data to be identified includes unstructured data for identifying association relationships between objects, and the set of relationship types includes multiple candidate association relationships; Preprocessing the data to be identified to be adapted to the relation extraction model to obtain a first input sequence; The first input sequence and the relationship type set are processed separately by the relationship extraction model to obtain a target relationship among the candidate relationships corresponding to the data to be identified.
2. The method according to claim 1, characterized in that The preprocessing of the data to be identified to be adapted to the relation extraction model to obtain a first input sequence includes: Adding a set mark to the data to be identified to obtain a first input sequence; The setting flags include one or more of the following: a start mark located at the start position of the data to be identified; an end marker located at the end position of the data to be identified; Entity markers located on both sides of the object.
3. The method according to claim 1, characterized in that The step of processing the first input sequence and the relationship type set respectively by the relationship extraction model to obtain a target relationship among the candidate relationships corresponding to the data to be identified includes: The following operations are performed through the relationship extraction model: processing the first input sequence to obtain a first relation representation vector of the first input sequence; For each candidate association relationship in the relationship type set, determining a first semantic vector corresponding to the candidate association relationship; For each first semantic vector, determining a cosine similarity measure between the first semantic vector and the first relationship representation vector; The candidate association relationship corresponding to the largest cosine similarity measure among the cosine similarity measures is determined as the target association relationship of the data to be identified.
4. The method according to claim 3, characterized in that The processing of the first input sequence to obtain a first relation representation vector of the first input sequence includes: Encoding the first input sequence to obtain a sequence-encoded vector; Obtaining a first latent vector corresponding to the start marker in the sequence encoded vector; Determining the first latent vector as the sentence vector representation of the data to be recognized; For each object in the to-be-identified data, determine a second latent vector for a target tag corresponding to the object, where the number of the second latent vectors is equal to the number of tags for the corresponding object, and the target tag represents the following information in a marked form: the object corresponding to the target tag in the first input sequence and the entity tag associated with the object; For each object, determining the average value of all second latent vectors corresponding to the object; For each of the average values, add the average value to the sentence vector representation to obtain a third latent vector; Combining the third latent vectors to obtain a fourth latent vector; Determine a first relationship representation vector corresponding to the fourth latent vector.
5. The method according to claim 3, characterized in that The determining, for each candidate association relationship in the relationship type set, a first semantic vector corresponding to the candidate association relationship includes: For each candidate association relationship in the relationship type set, encode the candidate association relationship to obtain an encoded relationship vector, wherein the starting position of the candidate association relationship includes a start marker; The fifth latent vector corresponding to the start tag in the relationship encoding vector is used as the first semantic vector corresponding to the candidate association relationship.
6. The method according to claim 1, characterized in that The training operation of the relation extraction model includes: Acquire an unstructured data set and a training relationship set, wherein the unstructured data set includes unstructured data for identifying association relationships between training objects, and the training relationship set includes multiple training association relationships; marking the positive correlation between the training object and the training object in the unstructured data; Selecting negative association relationships between the training objects from the training relationship set; Preprocessing the unstructured data to be adapted to the model to be trained to obtain a second input sequence; The model to be trained is trained using the second input sequence, the positive correlation relationship, and the negative correlation relationship to obtain a relationship extraction model.
7. The method according to claim 6, characterized in that The step of training the model to be trained by using the second input sequence, the positive correlation relationship, and the negative correlation relationship to obtain a relationship extraction model includes: The following operations are performed using the model to be trained: processing the second input sequence to obtain a second relation representation vector corresponding to the second input sequence; Determining a second semantic vector corresponding to the positive association relationship; Determining a third semantic vector corresponding to the negative association relationship; respectively determining a first similarity measure between the second semantic vector and the second relationship representation vector and a second similarity measure between the third semantic vector and the second relationship representation vector; Determining a triplet loss function corresponding to the first similarity metric and the second similarity metric; The model parameters of the model to be trained are adjusted based on the triplet loss function until an end condition is met.
8. A device for identifying association relationships, characterized in that: include: An acquisition module, configured to acquire data to be identified and a set of relationship types, wherein the data to be identified includes unstructured data for identifying association relationships between objects, and the set of relationship types includes a plurality of candidate association relationships; A preprocessing module, configured to preprocess the data to be identified so as to be adapted to the relation extraction model, thereby obtaining a first input sequence; A processing module is used to process the first input sequence and the relationship type set respectively through the relationship extraction model to obtain a target relationship among the candidate relationships corresponding to the data to be identified.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method according to any one of claims 1 to 7 when executed.