Object information extraction method, device, storage medium and electronic device
By extracting and clustering the object data, class labels and relationship information are automatically generated, and the dependence problem of manual annotation in neural network recommendations is solved, and the accuracy and efficiency of the recommendation system are improved.
Patent Information
- Application Number
- CN202210070390.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-01-21
AI Technical Summary
In the prior art, neural network recommendations rely on manual annotation training data, resulting in high information acquisition costs and difficult to guarantee accuracy, affecting the recommendation accuracy.
By performing feature extraction and clustering on the initial data of the object, the object's category label and relationship information are automatically generated, and the target feature information is generated for unsupervised clustering and recommendation.
It realizes the generation of accurate recommendation effects without manual annotation, improves the accuracy and efficiency of the recommendation system, and reduces labor costs.
Smart Images

Figure CN116521884B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to object information extraction methods, devices, storage media, and electronic devices. Background Art
[0002] With the widespread adoption of computer technology, intelligent applications focused on intelligently meeting user needs have seen significant development. For example, automatic item and media recommendations are being developed. As the demand for recommendations grows, intelligent recommendations based on neural networks have become a growing trend. However, neural network recommendations rely on the richness of training data. Currently, obtaining training data relies heavily on manual annotation, which is difficult to achieve with sufficient precision and automation, and is also costly, making it difficult to fully meet the requirements of neural network training, thus affecting the accuracy of recommendations. Summary of the Invention
[0003] In order to solve at least one of the above technical problems, embodiments of the present application provide an object information extraction method, device, storage medium, and electronic device.
[0004] In one aspect, an embodiment of the present application provides a method for extracting object information, the method comprising:
[0005] Acquire object initial data, where the object initial data includes object data corresponding to a plurality of objects, and each object data includes an object identifier and object description information;
[0006] Feature extraction is performed on the object identification and object description information of each object to obtain object identification features and object description features respectively;
[0007] Based on the object description features of each object, clustering the objects to determine the category label corresponding to each object;
[0008] Determining, based on the object identifier and category label of each object, relationship information of each object, wherein the relationship information represents a subordinate relationship between the object identifier of each object and a category involved in a clustering result;
[0009] Target feature information of the object is generated according to the relationship information and the object identification feature.
[0010] On the other hand, an embodiment of the present application provides an object information extraction device, the device comprising:
[0011] An initial data acquisition module is used to acquire object initial data, wherein the object initial data includes object data corresponding to a plurality of objects, and each object data includes an object identifier and object description information;
[0012] A feature extraction module is used to extract features from the object identification and object description information of each object, and obtain object identification features and object description features respectively;
[0013] A clustering module, configured to perform clustering processing on each of the objects based on the object description features of each object, and determine a category label corresponding to each object;
[0014] a relationship extraction module, configured to determine relationship information of each object based on the object identifier and category label of each object, wherein the relationship information represents a subordinate relationship between the object identifier of each object and the category involved in the clustering result;
[0015] A feature generation module is used to generate target feature information of the object according to the relationship information and the object identification feature.
[0016] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the above-mentioned object information extraction method.
[0017] On the other hand, an embodiment of the present application provides an electronic device comprising at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements the above-mentioned object information extraction method by executing the instructions stored in the memory.
[0018] On the other hand, an embodiment of the present application provides a computer program product, including a computer program or instructions, which implements the above-mentioned object information extraction method when executed by a processor.
[0019] The object information extraction method provided by this application automatically performs unsupervised clustering on multiple objects by clustering them to obtain label information corresponding to the objects. The object information corresponding to the object is obtained based on the object and its corresponding label information. This relationship information is reflected in the object information to obtain target object information containing the relationship information. Recommendations based on this object information can achieve more accurate recommendation results. This application can automatically generate category labels for items without relying on manual labeling. Furthermore, the information effectively incorporated into the category labels can be input as numerical features into the recommendation algorithm model, helping the recommendation system to make more accurate recommendations and having strong generalizability. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 This is a schematic diagram of a feasible implementation framework of the object information extraction method provided in the embodiments of this specification;
[0022] Figure 2 This is a flow chart of an object information extraction method provided by an embodiment of this specification;
[0023] Figure 3 Schematic diagram of the skip-gram model provided in the embodiment of the present application;
[0024] Figure 4 This is a schematic diagram of the graph embedding feature generation process provided by an embodiment of the present application;
[0025] Figure 5 Schematic diagram of the object information extraction method provided in an embodiment of the present application;
[0026] Figure 6 is a block diagram of an object information extraction device provided in an embodiment of the present application;
[0027] Figure 7 This is a schematic diagram of the hardware structure of a device provided in an embodiment of the present application for implementing the method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of them. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the embodiments of the present application.
[0029] It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0030] In order to make the purpose, technical solutions and advantages disclosed in the embodiments of the present application more clearly understood, the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present application and are not intended to limit the embodiments of the present application.
[0031] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "multiple" means two or more. In order to facilitate understanding of the above-mentioned technical solutions and the technical effects produced by the embodiments of this application, the embodiments of this application first explain the relevant professional terms:
[0032] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0033] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0034] Convolutional Neural Networks (CNNs) are a type of feedforward neural network with a deep structure that incorporates convolutional computations. They are a representative algorithm for deep learning. Convolutional neural networks possess the ability to learn representations and perform translation-invariant classification of input information based on their hierarchical structure, hence the name "translation-invariant artificial neural network."
[0035] One-hot code: Also known as a one-hot code in English literature, it intuitively refers to a code system with as many bits as there are states, with only one bit set to 1 and all others set to 0. Typically, in communication network protocol stacks, one-hot codes with eight or sixteen states are used, with the system occupying one state code and the remaining bits available for user use.
[0036] With the widespread adoption of computer technology, intelligent applications designed to meet user needs have seen significant growth, with recommendation apps showing the most significant growth. These apps are now permeating every aspect of our lives, offering recommendations for products, media resources, quick information, and expert advice.
[0037] Taking advertising recommendations as an example, recommendation systems are a crucial component of internet advertising recommendation services. Item and ad features serve as inputs to recommendation algorithms, directly determining their effectiveness. Recommended items have various relationships based on their attributes, which can be viewed as a graph. Graph embedding algorithms leverage the topological structure formed by these relationships, integrating them into feature vectors and generating dense features describing the items. Typical methods include DeepWalk, Line, GCN, TransE, and TransD.
[0038] The DeepWalk algorithm draws on the ideas of the word2vec algorithm, a commonly used word embedding method in natural language processing. Word2vec describes the relationships between words through sentence sequences in a corpus, thereby learning vector representations of words. Similar to word2vec, the DeepWalk algorithm uses the relationships between nodes in a graph to learn vector representations of nodes. In DeepWalk, nodes are sampled in the graph using a random walk to simulate the corpus data, and then the node relationships are learned using the word2vec method.
[0039] The Line algorithm defines the first-order and second-order similarities of nodes and designs an objective function that preserves both local and global information about the network. To address the limitations of the stochastic gradient descent algorithm in large-scale network embedding, a new edge sampling algorithm is proposed to improve optimization efficiency.
[0040] GCN (Graph Convolutional Networks) aggregates the feature information of the nodes themselves and the node topology information by performing convolution operations on the graph, and has achieved good results in solving semi-supervised learning problems.
[0041] The goal of the TransE algorithm is to represent all entities and relationships in the knowledge base as a low-dimensional vector. In a knowledge base, entities can be considered head nodes or tail nodes, and relationships exist between nodes. For example, tigers belong to animals, and the triple represented by (head, relation, tail) is (tiger, belongs to, animal). Tigers and animals belong to entities, and the relationship is belongs to. Assuming h, r, and t represent the head vector, relation vector, and tail vector, respectively, TransE believes that the correct relationship is h+r=t, and incorrect triples do not satisfy this relationship. The potential energy of the triple is represented by the bi-norm of (h+r->t), and the optimization goal is to minimize the potential energy of the entire sample set. The TransE algorithm can effectively embed entity relationships into features, but it cannot solve the problems of one-to-many and many-to-one relationships, and the calculation must rely on the entity-label relationship.
[0042] The TransD algorithm projects the head and tail nodes so that the triples satisfy the equation head (after projection) + relation = tail (after projection). This effectively solves the graph embedding problem for many-to-one and one-to-many relationships and improves computational efficiency. However, the TransD algorithm also relies on the entity-to-label relationship; without the label relationship, effective embedding cannot be achieved.
[0043] In summary, related knowledge graph-based graph embedding methods rely on attribute labels for items. If these labels are missing, the knowledge graph cannot be generated, and thus the embedding algorithm cannot be used. Attribute labeling relies on attributes with physical meaning, requiring extensive manual annotation and resulting in low efficiency. Pre-defined categories require manual intervention. For example, interest profile attributes include sports, music, and movies. These three categories are manually defined, and each of these categories further subdivides into more detailed attributes. Relying solely on manual definition is not only labor-intensive but also fails to accurately define and cover all possible attributes. Furthermore, with a large number of interactions, manual labeling is impractical. Labeling methods rely on pre-defined categories, and the level of granularity in the category definition directly impacts the expressiveness of features generated from these labels. For example, for sports products and advertisements, the advertisements are first categorized into different sports categories, such as basketball, football, and table tennis. Then, the corresponding advertised items are identified and labeled. However, in reality, it is impossible to pre-categorize all advertisements and products. This results in many advertisements and products being unable to be categorized, hindering item labeling.
[0044] In order to realize automatic labeling of attribute labels, extract object information containing relationship information, and achieve the purpose of improving recommendation accuracy based on the relationship information in the object information, an embodiment of the present application provides an object information extraction method, which automatically performs unsupervised clustering on objects by clustering multiple objects to obtain label information corresponding to the objects, obtains relationship information corresponding to the objects based on the objects and the label information corresponding to the objects, and reflects the relationship information in the object information to obtain target object information containing relationship information. Recommendations based on this object information can achieve more accurate recommendation effects.
[0045] The embodiments of the present application may involve cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or a local area network to realize the calculation, storage, processing, and sharing of data. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the application of cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the rapid development and application of the Internet industry, each item may have its own identification mark in the future, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately. All kinds of industry data require strong system backing support, which can only be achieved through cloud computing.
[0046] The method provided in the embodiment of the present application may also involve a blockchain, that is, the method provided in the embodiment of the present application may be implemented based on a blockchain, or the data involved in the method provided in the embodiment of the present application may be stored based on a blockchain, or the execution subject of the method provided in the embodiment of the present application may be located in a blockchain. Blockchain is a new application model of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanism, and encryption algorithm. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. The blockchain may include a blockchain underlying platform, a platform product service layer, and an application service layer.
[0047] The underlying blockchain platform can include processing modules such as user management, basic services, smart contracts, and operation management. Among them, the user management module is responsible for the identity information management of all blockchain participants, including maintaining public and private key generation (account management), key management, and maintaining the correspondence between the user's real identity and the blockchain address (authority management), etc., and under authorization, it supervises and audits the transactions of certain real identities and provides risk control rule configuration (risk control audit); the basic service module is deployed on all blockchain node devices to verify the validity of business requests, and records valid requests to storage after consensus is reached. For a new business request, the basic service first adapts the interface for parsing and authentication (interface adaptation), and then encrypts the business information through the consensus algorithm (consensus management). The smart contract module is responsible for the registration, issuance, triggering and execution of contracts. Developers can define the contract logic in a programming language and publish it to the blockchain (contract registration). According to the logic of the contract terms, the contract logic is triggered by calling keys or other events to trigger execution. The contract logic is completed, and the contract upgrade and cancellation functions are also provided. The operation management module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation and real-time status visualization output of the product during the product release process, such as alarms, network management, and node device health status.
[0048] The platform's product service layer provides the basic capabilities and implementation framework for typical applications. Developers can build on these basic capabilities, overlay business features, and complete the blockchain implementation of business logic. The application service layer provides application services based on blockchain solutions for business participants to use.
[0049] See also Figure 1 , Figure 1 This is a schematic diagram of a feasible implementation framework of the object information extraction method provided in the embodiment of this specification. Figure 1As shown, the implementation framework may include at least a terminal device 01 and a data processing server 02. The terminal device 01 may be a device located on the Internet, and may provide users with various optional Internet-based recommendation services, which are provided by the client in the embodiment of the present application. The terminal device 01 includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, and in-vehicle terminals. The client in the embodiment of the present application may provide applications in various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, video media, smart communities, and instant messaging.
[0050] The data processing server 02 provides services to users by interacting with the terminal device 01. Specifically, the terminal device 01 can issue a recommendation request to the data processing server 02, and determine the recommended object by extracting information about related objects. Specifically, the data processing server 02 can obtain the initial data of the object, and the initial data of the object includes object data corresponding to multiple objects, and each object data includes object identification and object description information. Feature extraction is performed on the object identification and object description information in each object respectively, and object identification features and object description features are obtained accordingly. Based on the object description features of each object, clustering processing is performed on each of the above objects to determine the category label corresponding to each of the above objects. Based on the object identification and category label of each object, the relationship information of each object is determined, and the relationship information represents the subordinate relationship between the object identification of each object and the category involved in the clustering result. Based on the relationship information and the object identification features, the target feature information of the object is generated. Based on the target feature information of the object, the recommendation degree corresponding to the object is determined. The object is recommended based on the recommendation degree.
[0051] The following introduces an object information extraction method according to an embodiment of the present application. The information involved in the present disclosure may be information authorized by the user or fully authorized by all parties.
[0052] Figure 2 A flow chart of an object information extraction method provided by an embodiment of the present application is shown. The embodiment of the present application provides the method operation steps as described above in the embodiment or flow chart, but more or fewer operation steps may be included based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the order of many steps and does not represent the only execution order. When the actual system, terminal device or server product is executed, it can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment) according to the method shown in the embodiment or the accompanying drawings. The above method may include:
[0053] S101. Obtain the initial object data. The above-mentioned initial object data includes object data corresponding to multiple objects, and each object data includes an object identifier and object description information.
[0054] The embodiments of the present application do not limit the object. The object is a general term for various contents that can be recommended, such as advertisements, items, resources, information, etc. The object identifier includes the identifier information of the object or other objects related to the object, and the object description information is the information related to the object. Taking an advertisement as an example, the identifier corresponding to the item associated with the advertisement is the object identifier; and the information in the advertisement, such as text information or multimedia information, is the object description information. The initial object data can be the advertisement data to be pushed or the historically pushed advertisement data. From this advertisement data, the data corresponding to multiple objects can be determined, and each object data includes an object identifier and object description information.
[0055] S102. Respectively perform feature extraction on the object identifier and object description information in each object data, and correspondingly obtain object identifier features and object description features.
[0056] The embodiments of the present application do not limit the method for performing feature extraction on the object identifier and object description information. For example, various neural networks such as convolutional neural networks, deep convolutional neural networks, recurrent convolutional neural networks, and encoder-decoder networks can be used for feature extraction.
[0057] Taking the text information in the advertisement as the object description information as an example, before performing feature extraction on the object description information, preprocessing operations on the object description information can also be performed. The preprocessing operations can include at least one of the following:
[0058] (1) Remove punctuation marks, letters, numbers, and special symbols, for example:
[0059] [a-zA-Z0-9’!"#$%&\'()*+,-. / :;<=>?@,。?★、…【】《》?“”‘’![\\]^_`{|}~]+
[0060] (2) Remove specified stop words, such as: "de", "o", "ya"...
[0061] (3) Provide a custom word segmentation dictionary, such as the name of a certain commodity, common words in the advertisement.
[0062] In one embodiment, after preprocessing, the following processing can be performed: the object description information in each object data is segmented to obtain the segmentation sequence corresponding to each object; the segmentation sequence corresponding to each object is embedded with feature extraction to obtain the object description feature. Of course, if a custom segmentation dictionary is obtained, segmentation processing can be performed based on the segmentation dictionary to ensure that the words in the segmentation dictionary are not cut into two different words. After processing the text information in the advertisement, the word sequence corresponding to each text information can be obtained, as shown in Table 1:
[0063] Table 1
[0064]
[0065] In advertising scenarios, object descriptions are textual. Advertisement identifiers can be used as their corresponding identifiers, corresponding to the item identifiers that the advertisements target. After generating the word segmentation sequences, duplicates must be removed to ensure that each sequence is unique.
[0066] The embodiment of the present application does not limit the method used for embedding feature extraction. For example, a skip-gram model can be trained to perform embedding feature extraction. Figure 3 As shown in the figure, it shows a schematic diagram of the skip-gram model. In general, the skip-gram model is a model that predicts the context through the current word, that is, the input is the word vector of a specific word, and the output is the context word vector corresponding to the specific word. Figure 3 For example, a specific word in the text sequence is w(t), and the context size is 4. The specific word w(t) is the input of the skip-gram model, and the four context words w(t-1), w(t-2), w(t+1), and w(t+2) are the outputs of the skip-gram model, where w represents the text sequence and t represents the sequence number of the specific word.
[0067] Each word input to the skip-gram model is encoded in a one-hot format. The input layer is of size 1*N, where N is the number of output prediction vectors. The hidden layer is represented by a weight matrix W of size V*N, where each row represents the embedding of a prediction vector, where V is the length of the prediction vector and N is the number of neurons in the hidden layer, which is also the dimension of the embedding vector. The weight matrix from the hidden layer to the output layer is represented by O and has a size of N*V.
[0068] The output from input to hidden is: Where h is the output of hidden and X is the input of input. The output from hidden to output is: u = O T h, where u is the output of output. Use softmax to normalize the output to [0,1], and get the probability of word j appearing under the condition that a certain word t appears. This probability is expressed by formula (1):
[0069] Formula (1):
[0070] Among them, u j Represents the predicted vector corresponding to word j, u k represents the prediction vector corresponding to word k, and exp represents the exponential function with the natural constant e as the base. w(t) represents the text in which word t appears, w(j) represents the text in which word j appears, and P(w(j)|w(t)) represents the probability of word j appearing in the text given word t.
[0071] The skip-gram model is a single-input, multiple-output model, so it predicts multiple words in the context of word t. Thus, we can get formula (2) and formula (3), and according to formula (2) and (3), we can get the loss function (4). Formula (2-4) is specifically:
[0072] Formula (2):
[0073] Among them, context represents context, w context represents the text of the context in which it appears, P(w context |w(t)) represents the probability of the context appearing when word t appears in the text.
[0074] Formula (3):
[0075] Formula (4): loss=log∑ k∈v exp(u k )-u j,j∈context .
[0076] By minimizing the loss function and updating the parameters of the weight matrix W, we obtain a trained skip-gram model. By inputting the text sequence into the skip-gram model, we can obtain a unique feature representation. In other words, by inputting the word segmentation sequence corresponding to each object into the skip-gram model, we can obtain the object description features described above. For each word, we can obtain its corresponding vector representation, as shown in Table 2.
[0077] Table 2
[0078] Word Number Word Text Vector representation 00000001 Special Offers [0.1,0.3,-0.2,0.9,0.67] 00000002 Buy [0.6,-0.12,-0.3,-0.86,0.56]
[0079] The object description feature of a text sequence can be expressed as the weighted sum of the vector representations corresponding to all words in the text sequence. Of course, the mean can also be taken. Taking advertisements as an example, every item corresponds to at least one advertisement, so an item may have multiple object description features.
[0080] Of course, other models can also be used to extract object description features in this application, such as FastText and Item2Vec. FastText is an open-source word embedding and text classification tool developed by Facebook. Item2Vec is a method for vectorizing content.
[0081] S103 . Based on the object description features of each of the objects, cluster the objects to determine a category label corresponding to each of the objects.
[0082] The embodiments of the present application do not limit the method used to cluster objects. For example, clustering can be performed using partitioning methods, hierarchical methods, density-based methods, grid-based methods, model-based methods, etc. The partitioning method constructs clusters by splitting. The hierarchical method performs a hierarchical decomposition of a given data set until certain conditions are met. A fundamental difference between the density-based method and other methods is that it is not based on various distances, but on density. This can overcome the shortcoming that distance-based algorithms can only find "quasi-circular" clusters. The first step in solving the graph theory clustering method is to establish a graph that is suitable for the problem. The nodes of the graph correspond to the smallest units of the data being analyzed, and the edges (or arcs) of the graph correspond to the similarity measures between the smallest processing unit data. The grid-based method first divides the data space into a grid structure with a finite number of units, and all processing is based on a single unit. The model-based method assumes a model for each cluster, and then looks for a data set that can well meet this model.
[0083] In one embodiment, clustering can be performed by the following operations: determining a preset number of object description features and using the preset number of object description features as cluster centers; adjusting the cluster centers based on the object description features of each object to obtain a clustering result, wherein the clustering result includes the preset number of clusters, each cluster including at least one object description feature. Specifically, the clustering process can be completed by the following operations:
[0084] (1) Select the dataset D formed by the description features of each object = {x0, x1, x2, ..., x m}, the total number of object description features is m, and the maximum number of iterations is N. Of course, N can be set as needed.
[0085] (2) Randomly select K object description features from the dataset D as the initial K cluster centers: {μ0,μ1,μ2,…μ k}
[0086] (3) Initialize K sets {C0, C1, C2, ... C k}, these collections can be empty during the initialization phase.
[0087] (4) For each object description feature in the dataset D, calculate the distance between each object description feature and the K cluster centers. This embodiment uses Euclidean distance, but other distances may also be used. For each object description feature, the object description feature is placed in the corresponding set with the cluster center having the smallest distance to it.
[0088] Assuming x and y are two n-dimensional vectors, the Euclidean distance between x and y is defined as:
[0089] (5) After all the object description features in the dataset D have been calculated and their distances from the cluster centers have been completed and the clusters have been divided, the cluster centers of each cluster are recalculated to generate new K cluster centers. Based on the newly generated cluster centers, {μ0,μ1,μ2,…μ k}.
[0090] (6) If the distance between the newly calculated cluster center and the original cluster center is less than a certain set threshold (indicating that the position of the recalculated centroid has not changed much and tends to be stable, or converged), it can be considered that the clustering has achieved the desired result and the algorithm terminates. Otherwise, repeat steps (4)-(6) until the number of iterations equals N.
[0091] Taking advertisements as an example, in step S103, the object description features of each advertisement text can be obtained. By clustering the object description features, advertisements can be divided into different categories. The advertisement categories in a cluster are similar. For example, related advertisements such as sports shoes and sports equipment will be divided into one cluster, while milk powder, diapers, etc. will be divided into another cluster. It can be considered that each cluster represents a type, and the number of types K is the number of cluster centers set in advance. It can be found that as the number of cluster centers increases, the number of clusters generated by clustering will also increase, which means that the classification will become more and more refined. Therefore, the K value can be defined in advance to adjust the precision of the label to meet the requirements of different scenarios. After the clustering is completed, each advertisement or item will correspond to a label category, as shown in Table 3.
[0092] Table 3
[0093] Item Identification Tag 1 Tag 2 Tag 3 …… Tag K 000001 1 0 0 …… 0 000002 0 0 1 …… 0 000003 0 0 0 …… 1
[0094] The labels generated by this process can not only help depict business-related portraits, but also be directly input into the recommendation model as numerical features for calculation.
[0095] S104. Determine the relationship information of each object based on the object identifier and category label of each object, where the relationship information represents the subordinate relationship between the object identifier of each object and the category involved in the clustering result.
[0096] The above-mentioned relationship information is represented by at least one triple, and the above-mentioned triple includes a first element, a second element and a third element. The above-mentioned first element corresponds to the object identifier, and can be specifically represented by the object identifier feature. The above-mentioned third element corresponds to the category label, and can be specifically represented by the label feature corresponding to the category label. The embodiment of the present application does not limit the method for obtaining the label feature. For example, the label feature can be manually set for each category label, or the cluster center of the cluster corresponding to the category label in the clustering result can be determined as the label feature. The above-mentioned second element represents the subordinate relationship between the above-mentioned first element and the above-mentioned third element. Taking the advertising scene as an example, the first element can be the identifier corresponding to the item pointed to by the advertisement, and the generated triple is shown in Table 4. Of course, there may be one-to-many or many-to-one situations between the item identifier and the category label, which all fall within the scope of implementation of the embodiment of the present application.
[0097] Table 4
[0098]
[0099] S105. Generate target feature information of the object based on the relationship information and the object identification features.
[0100] In one embodiment, the target feature information includes object identification information and object genus information. Generating the target feature information of the object based on the relationship information and the object identification feature includes: obtaining the object identification information and the object genus information based on the first element, the second element, the third element, and the object identification feature.
[0101] The object identification information and the object category information can be obtained based on the knowledge graph method. The knowledge graph method is an algorithm proposed for relational data. Its core idea is to vectorize the relations and entities in the knowledge graph, and make the head + relation equal to the tail as much as possible by mapping the head and tail to a space constructed by the entities and relations in a set of triples, thereby performing a low-dimensional dense representation of each piece of knowledge in the knowledge graph. In the recommendation scenario, items and labels can be represented as head nodes and tail nodes in the knowledge graph respectively. There is a corresponding relationship between items and labels, such as (basketball, belongs to, sports equipment), "basketball" represents the head, "sports equipment" represents the tail, and "belongs to" represents the relationship.
[0102] Specifically, multiple positive samples can be constructed based on the relationship information of each object. Multiple negative samples are generated based on the above multiple positive samples. Based on the preset mapping, the sample identification prediction information and sample category prediction information corresponding to each sample are obtained, and the above sample is a positive sample or a negative sample. Based on the sample identification prediction information and sample category prediction information corresponding to each sample, the mapping loss is calculated. The above preset mapping is optimized based on the above mapping loss. According to the above first element, the above second element, the above third element, the above object identification feature and the above preset mapping, the above object identification information and the above object category information are obtained. Specifically, with the above first element, the above second element, the above third element and the above object identification feature as the input of the above preset mapping, the above object identification information and the above object category information can be obtained.
[0103] In one embodiment, the TransD algorithm can be used to calculate the object identification information and the object category information. Its goal is to make head projection (object identification information corresponding to the object identification) + relation = tail projection (object category information pointed to by the category label). This algorithm is more effective in handling one-to-many and many-to-one relationships.
[0104] like Figure 4 As shown in the figure, for a given triple (h, r, t), h represents the head of the triple, r represents the relation of the triple, and tail represents the tail of the triple. The meaning of the triple can be referred to in Table 4 above. Its feature vector is expressed as: {h, r, t, h p ,r p ,t p}, where h,h p ,t,t p ∈R n ,r,r p ∈R m, h p ,t p are the projection vectors of h and t, r p is the projection vector of r. The relationship between r and the projection matrix of head and tail is defined as:
[0105] Where I is the identity matrix, m and n are the dimensions of the vector space;
[0106] M rh and M rt All belong to the preset mappings mentioned above.
[0107] After mapping, we can get:
[0108] h ⊥ (sample identification prediction information) = M rh h
[0109] t ⊥ (sample category prediction information) = M rt t.
[0110] The triplets generated in the previous article are regarded as positive samples, and negative samples can be randomly generated by randomly replacing the tail node of the triplet of the positive sample so that the replaced triplet does not exist in the positive sample, that is, the newly obtained (h, r, t) relationship does not hold, and the newly generated triplet is regarded as a negative sample.
[0111] The loss generated for each sample can be expressed as Among them, h||2≤1,||t||2≤1,||r||2≤1,||h ⊥ ||2≤1,||t ⊥ ||2≤1. Assume S pos and S neg Denote positive sample triplets and negative sample triplets respectively, the resulting total mapping loss can be expressed as formula (5):
[0112] Formula (5): Among them, h and t represent the first and third elements respectively, γ is the interval parameter, Δ pos is the set of positive sample triples, Δ neg is the set of negative sample triplets.
[0113] Gradient descent is used to optimize this preset mapping to obtain the head and tail projections corresponding to each head and tail, namely the object identification information and object genus information. Since the previous method used knowledge graphs to generate object identification information and object genus information containing relationship information, these object identification information and object genus information can also be considered as graph embedding features.
[0114] The embodiment of the present application provides an object information extraction method that can be applied to extract information of various objects as long as the object has an identification and description. Figure 5 As shown, taking advertisements as an example, the advertisement push content data can be first organized, cleaned, and invalid data removed. Then, after segmenting, deduplicating, and serializing the advertisement text information, the word2vec method is used to perform feature extraction operations to obtain object identification features and object description features. Clustering is performed based on the generated object description features to obtain the category label corresponding to each object in the advertisement. The item labels, category labels, and relationships of the advertisement are organized into knowledge graph triples. Using the knowledge graph triples, graph embedding features are generated. The advertisement push data of the advertising platform has a strong physical meaning and can reflect the attributes of the advertised items themselves. However, due to the diversity of items and advertisements, many items and advertisements cannot be directly labeled, and manual labeling is inefficient and costly. Based on the method in the embodiment of the present application, the relationship between items and category labels can be integrated into the graph embedding features. The embodiment of the present application can effectively improve the expressive power of the graph embedding features, and has the characteristics of simplicity, efficiency, and easy scalability. In addition, since the advertisement push data contains the recommended item information and advertisement text information, using historical advertisement push data can help the recommendation system more accurately judge user needs and improve user experience. For example, the recommendation degree corresponding to the above-mentioned object is determined based on the target feature information of the object obtained in the previous text; the recommendation accuracy of the above-mentioned object is significantly improved by recommending the above-mentioned object based on the above-mentioned recommendation degree. Of course, the present application does not limit the method for determining the recommendation degree. For example, the target feature information can be used as the input of the recommendation model, and the output of the recommendation model can be used as the recommendation degree. The recommendation model can refer to related technologies and will not be described in detail here. Moreover, according to the foregoing, the present application can automatically generate category labels for items without relying on manual labeling. Moreover, the generated graph embedding features effectively incorporate the information of the category label, which can be input as numerical features into the recommendation algorithm model to help the recommendation system complete more accurate recommendations and have strong generalizability. In other embodiments, the intermediate results generated during the implementation of this embodiment can also be used as a user's recent interest portrait to assist other businesses.
[0115] The embodiment of the present application provides an object information extraction method, which performs unsupervised category division for objects by clustering, and obtains relationship information based on the category division results, so that the target feature information corresponding to the object can be extracted based on the relationship information combined with the knowledge graph method. In this way, computing efficiency and label coverage can be improved, and labor costs can be saved. In the machine learning model, the computer does not know and does not need to know the physical meaning of each dimension of the feature. The physical meaning of the feature has no practical significance for the machine learning process. Therefore, the clustering method can be used to automatically divide items and advertisements into multiple categories. Each category does not have a directly defined physical meaning, but similar advertisements and items can be divided into the same category. As the number of categories increases, the labels will become more refined, and the generation process does not require manual intervention. Using the knowledge graph method, the characteristics of the category are integrated into the graph embedding feature, so that items with similar labels are closer in numerical feature expression, which can effectively improve the expression ability of the target feature information.
[0116] Taking advertisements as an example, the textual information pushed by advertisements can be used to convert textual information with strong physical meaning into computable low-dimensional vectors. This can be used as input for clustering algorithms or as numerical features for recommendation systems, improving recommendation effectiveness and enriching feature dimensions. Automatically generating category attribute labels eliminates the need to manually predefine the item categories and labels corresponding to advertisements, reducing manpower investment and improving construction efficiency. The level of sophistication of item labels can be set through parameters, and their coverage of categories far exceeds that of manual labeling. These labels can be adjusted at any time, offering high coverage and flexibility. By adopting the TransD knowledge graph algorithm, item label information can be effectively integrated into graph embedding features, better handling one-to-many and many-to-one relationships, and improving the expressive power of features.
[0117] Please refer to Figure 6 , which shows a block diagram of an object information extraction device in this embodiment, the device includes:
[0118] The initial data acquisition module 101 is used to acquire object initial data, where the object initial data includes object data corresponding to a plurality of objects, and each object data includes an object identifier and object description information;
[0119] The feature extraction module 102 is used to extract features from the object identification and object description information of each object, and obtain object identification features and object description features respectively;
[0120] A clustering module 103 is configured to perform clustering processing on each of the objects based on the object description features of each object, and determine a category label corresponding to each of the objects;
[0121] The relationship extraction module 104 is used to determine the relationship information of each object according to the object identifier and category label of each object, wherein the relationship information represents the subordinate relationship between the object identifier of each object and the category involved in the clustering result;
[0122] The feature generation module 105 is configured to generate target feature information of the object according to the relationship information and the object identification feature.
[0123] In one embodiment, the relationship information is represented by at least one triple, the triple including a first element, a second element, and a third element, the first element corresponds to an object identifier, the third element corresponds to a category label, and the second element represents a subordinate relationship between the first element and the third element;
[0124] The target feature information includes object identification information and object category information;
[0125] The feature generation module 105 is configured to perform the following operations:
[0126] The object identification information and the object attribute information are obtained according to the first element, the second element, the third element and the object identification feature.
[0127] In one embodiment, the feature generation module 105 is configured to perform the following operations:
[0128] Construct multiple positive samples based on the relationship information of each object;
[0129] Generate multiple negative samples based on the above multiple positive samples;
[0130] Obtaining sample identification prediction information and sample category prediction information corresponding to each sample based on a preset mapping, where the sample is a positive sample or a negative sample;
[0131] Calculate the mapping loss based on the sample identification prediction information and sample category prediction information corresponding to each sample;
[0132] Optimizing the preset mapping based on the mapping loss;
[0133] The object identification information and the object attribute information are obtained according to the first element, the second element, the third element, the object identification feature and the preset mapping.
[0134] In one embodiment, the apparatus further comprises a recommendation module configured to determine a recommendation degree corresponding to the object based on target feature information of the object; and recommend the object based on the recommendation degree.
[0135] In one embodiment, the feature extraction module is used to perform word segmentation processing on the object description information in each of the objects to obtain a word segmentation sequence corresponding to each of the objects; and perform embedded feature extraction processing on the word segmentation sequence corresponding to each of the objects to obtain the object description features.
[0136] In one embodiment, the above-mentioned clustering module is used to determine a preset number of object description features and use the above-mentioned preset number of object description features as cluster centers; the above-mentioned cluster centers are adjusted based on the object description features of each object to obtain clustering results, and the above-mentioned clustering results include a preset number of clusters, and each cluster includes at least one object description feature.
[0137] The device embodiment and method embodiment of the present application are based on the same inventive concept and will not be described in detail here.
[0138] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the object information extraction method described above.
[0139] The embodiment of the present application further provides a computer-readable storage medium, which can store a plurality of instructions. The instructions can be suitable for a processor to load and execute the object information extraction method described above in the embodiment of the present application.
[0140] In one embodiment, the object information extraction method includes:
[0141] Obtaining object initial data, the object initial data including object data corresponding to a plurality of objects, each object data including an object identifier and object description information;
[0142] Feature extraction is performed on the object identification and object description information of each object to obtain object identification features and object description features respectively;
[0143] Based on the object description features of each of the above objects, clustering is performed on each of the above objects to determine the category label corresponding to each of the above objects;
[0144] Determining, based on the object identifier and category label of each object, relationship information of each object, wherein the relationship information represents a subordinate relationship between the object identifier of each object and a category involved in the clustering result;
[0145] Target feature information of the object is generated based on the relationship information and the object identification feature.
[0146] In one embodiment, the relationship information is represented by at least one triple, the triple including a first element, a second element, and a third element, the first element corresponds to an object identifier, the third element corresponds to a category label, and the second element represents a subordinate relationship between the first element and the third element;
[0147] The target feature information includes object identification information and object category information;
[0148] The above-mentioned generating target feature information of the above-mentioned object based on the above-mentioned relationship information and the above-mentioned object identification feature includes:
[0149] The object identification information and the object attribute information are obtained according to the first element, the second element, the third element and the object identification feature.
[0150] In one embodiment, the method further includes:
[0151] Construct multiple positive samples based on the relationship information of each object;
[0152] Generate multiple negative samples based on the above multiple positive samples;
[0153] Obtaining sample identification prediction information and sample category prediction information corresponding to each sample based on a preset mapping, where the sample is a positive sample or a negative sample;
[0154] Calculate the mapping loss based on the sample identification prediction information and sample category prediction information corresponding to each sample;
[0155] Optimizing the preset mapping based on the mapping loss;
[0156] The object identification information and the object category information are obtained based on the first element, the second element, the third element, and the object identification feature, including:
[0157] The object identification information and the object attribute information are obtained according to the first element, the second element, the third element, the object identification feature and the preset mapping.
[0158] In one embodiment, the method further includes:
[0159] Determining a recommendation degree corresponding to the object based on the target feature information of the object;
[0160] The above-mentioned objects are recommended based on the above-mentioned recommendation degrees.
[0161] In one embodiment, the object description information in each object is text information, and the feature extraction of the object identification and object description information in each object is performed to obtain the object identification features and object description features, respectively, including:
[0162] Perform word segmentation processing on the object description information in each of the above objects to obtain a word segmentation sequence corresponding to each of the above objects;
[0163] The word segmentation sequence corresponding to each of the above objects is subjected to embedding feature extraction processing to obtain the above object description features.
[0164] In one embodiment, clustering the objects based on the object description features of each object includes:
[0165] Determine a preset number of object description features, and use the preset number of object description features as cluster centers;
[0166] The cluster centers are adjusted based on the object description features of each object to obtain a clustering result. The clustering result includes a preset number of clusters, and each cluster includes at least one object description feature.
[0167] Furthermore, Figure 7 A schematic diagram of the hardware structure of a device for implementing the method provided in the embodiment of the present application is shown. The above-mentioned device may participate in constituting or include the apparatus or system provided in the embodiment of the present application. Figure 7 As shown, the device 10 may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 7 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 7 More or fewer components than shown, or with Figure 7 Different configurations shown.
[0168] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the device 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0169] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the above-mentioned method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, to implement the above-mentioned object information extraction method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the device 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0170] The transmission device 106 is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communications provider of the device 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In one embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0171] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of device 10 (or mobile device).
[0172] It should be noted that the above-mentioned order of the embodiments of the present application is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above-mentioned embodiments of the present application are described in terms of specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0173] Each embodiment of the present application is described in a progressive manner. Similar portions between the embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device and server embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0174] Those skilled in the art will understand that all or part of the steps for implementing the above embodiments may be accomplished by hardware, or may be accomplished by instructing the relevant hardware through a program, and the above program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0175] The above is only a preferred embodiment of the embodiment of the present application and is not intended to limit the embodiment of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiment of the present application should be included in the scope of protection of the embodiment of the present application.
Claims
1. A method for extracting object information, characterized in that: The method comprises: Acquire object initial data, where the object initial data includes object data corresponding to a plurality of objects, and each object data includes an object identifier and object description information; Feature extraction is performed on the object identification and object description information of each object respectively, and the object identification feature and object description feature are obtained accordingly. The object refers to the recommended content; Based on the object description features of each object, performing unsupervised clustering processing on each object to determine a category label corresponding to each object, wherein the category label indicates the category of the cluster and has no directly defined physical meaning; Determine, based on the object identifier and category label of each object, relationship information of each object, wherein the relationship information represents a subordinate relationship between the object identifier of each object and a category involved in the clustering result, and the relationship information is represented by at least one knowledge graph triple, wherein the knowledge graph triple includes a first element corresponding to the object identifier, a second element, and a third element corresponding to a label feature of the category label, wherein the second element represents the subordinate relationship between the object identifier and the category label, and the label feature is the cluster center of the cluster corresponding to the category label in the clustering result; In combination with the knowledge graph method, the first element, the second element, the third element, and the object identification feature are used as inputs of a preset mapping to obtain target feature information. The target feature information includes a graph embedding feature formed by object identification information containing relationship information and object genus information. The graph embedding feature can be used as a numerical feature input of a recommendation algorithm model. The method for obtaining the preset mapping includes: Construct multiple positive samples based on the relationship information of each object; generating a plurality of negative samples according to the plurality of positive samples; Obtaining sample identification prediction information and sample category prediction information corresponding to each sample based on a preset mapping, wherein the sample is a positive sample or a negative sample; Calculating a mapping loss based on the sample identification prediction information and the sample category prediction information corresponding to each sample; The preset mapping is optimized based on the mapping loss.
2. The method according to claim 1, characterized in that The method further comprises: determining a recommendation degree corresponding to the object based on target feature information of the object; The object is recommended based on the recommendation degree.
3. The method according to claim 2, characterized in that The object description information of each object is text information, and the feature extraction is performed on the object identification and object description information of each object respectively, and the object identification feature and the object description feature are obtained accordingly, including: Performing word segmentation processing on the object description information of each object to obtain a word segmentation sequence corresponding to each object; Embedding feature extraction is performed on the word segmentation sequence corresponding to each object to obtain the object description feature.
4. The method according to claim 2, characterized in that The clustering of the objects based on the object description feature of each object includes: Determining a preset number of object description features, and using the preset number of object description features as cluster centers; The cluster center is adjusted based on the object description feature of each object to obtain a clustering result, where the clustering result includes a preset number of clusters, and each cluster includes at least one object description feature.
5. An object information extraction device, characterized in that: The device comprises: An initial data acquisition module is used to acquire object initial data, wherein the object initial data includes object data corresponding to a plurality of objects, and each object data includes an object identifier and object description information; A feature extraction module is used to extract features from the object identification and object description information of each object, and obtain object identification features and object description features respectively. The object refers to the recommended content; a clustering module, configured to perform unsupervised clustering processing on each of the objects based on the object description features of each object, and determine a category label corresponding to each of the objects, wherein the category label indicates the category of the cluster and has no directly defined physical meaning; a relationship extraction module, configured to determine, based on the object identifier and category label of each object, the relationship information of each object, wherein the relationship information represents the subordinate relationship between the object identifier of each object and the category involved in the clustering result; the relationship information is represented by at least one knowledge graph triple, wherein the knowledge graph triple includes a first element corresponding to the object identifier, a second element, and a third element corresponding to a label feature of the category label, wherein the second element represents the subordinate relationship between the object identifier and the category label, and the label feature is the cluster center of the cluster corresponding to the category label in the clustering result; In combination with the knowledge graph method, the first element, the second element, the third element, and the object identification feature are used as inputs of a preset mapping to obtain target feature information. The target feature information includes a graph embedding feature formed by object identification information containing relationship information and object genus information. The graph embedding feature can be used as a numerical feature input of a recommendation algorithm model. The method for obtaining the preset mapping includes: Construct multiple positive samples based on the relationship information of each object; generating a plurality of negative samples according to the plurality of positive samples; Obtaining sample identification prediction information and sample category prediction information corresponding to each sample based on a preset mapping, wherein the sample is a positive sample or a negative sample; Calculating a mapping loss based on the sample identification prediction information and the sample category prediction information corresponding to each sample; The preset mapping is optimized based on the mapping loss.
6. The device according to claim 5, characterized in that The device further includes a recommendation module, configured to: determining a recommendation degree corresponding to the object based on target feature information of the object; The object is recommended based on the recommendation degree.
7. The device according to claim 6, characterized in that The feature extraction module is used for: Performing word segmentation processing on the object description information of each object to obtain a word segmentation sequence corresponding to each object; Embedding feature extraction is performed on the word segmentation sequence corresponding to each object to obtain the object description feature.
8. The device according to claim 7, characterized in that The clustering module is used to: Determining a preset number of object description features, and using the preset number of object description features as cluster centers; The cluster center is adjusted based on the object description feature of each object to obtain a clustering result, where the clustering result includes a preset number of clusters, and each cluster includes at least one object description feature.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement an object information extraction method according to any one of claims 1 to 4.
10. An electronic device, characterized in that: The invention comprises at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements an object information extraction method as described in any one of claims 1 to 4 by executing the instructions stored in the memory.
11. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, an object information extraction method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Method, device and equipment for generating network security knowledge graph, and storage medium
CN109347798A
Data processing method and device, storage medium and equipment
CN111258995A
Recommendation information generation method and device, storage medium and electronic equipment
CN113065911A