A Social Relationship Analysis Method, System, and Storage Medium Based on Multimodal Data

By extracting social text and image information, combining transformer model and Si-SCAN algorithm, the problem of unused image information in the prior art is solved, and a more comprehensive and accurate social relationship analysis is achieved, and potential social relationships are discovered.

CN115293920BActive Publication Date: 2025-07-18XIAMEN MEIYA PICO INFORMATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210971424.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-12
Publication Date
2025-07-18
Estimated Expiration
2042-08-12

AI Technical Summary

Technical Problem

The existing social relationship analysis methods fail to make full use of image information and lack multi-dimensional information portrayal, resulting in the incomplete construction of social relationships.

Method used

Social text and image information are extracted, converted into text features and image features, and personnel social network diagrams are constructed. Multimodal fusion model based on transformer and Si-SCAN graph clustering algorithm are used to conduct in-depth relationship analysis based on personnel intimacy and fusion features.

Benefits of technology

Through multimodal information fusion, the accuracy of social relationship clustering is improved, potential social associations can be discovered, and the clustering results of social networks are optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115293920B_ABST
    Figure CN115293920B_ABST
Patent Text Reader

Abstract

The present invention proposes a social relationship analysis method based on multimodal data, including: S1, extracting the social text and social image information of a person, respectively converting them into text features and image features, and statistically calculating the intimacy of the person, and constructing a person social network graph based on the intimacy of the person; S2, inputting the text features and image features into a multimodal fusion model based on transformer to obtain fusion features; S3, using the Si-SCAN graph clustering algorithm to analyze the person social network graph to obtain a social relationship clustering result, wherein the Si-SCAN graph clustering algorithm is constructed by introducing the intimacy of the person and the fusion feature information on the basis of the SCAN algorithm. The present invention deeply analyzes the social relationship based on the information of two modalities, text and image, designs a multimodal information fusion model, learns the interaction relationship between cross-modalities, and generates a graph node embedding representation of multimodal fusion. Through graph clustering analysis, the deep relationship analysis of the social network is realized, and potential social associations can be effectively discovered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of social relationship analysis, and specifically relates to a social relationship analysis method, system and storage medium based on multi-modal data. Background Art

[0002] With the rapid development of Internet technology, online social applications and media have also spread rapidly, and people's daily lives have become increasingly inseparable from various social information. A social network is also a relationship network between people. By analyzing various types of communication information in the social network, the social relationships between users can be discovered. The prediction methods for social relationship links are mainly divided into two categories, one is the method based on similarity measurement, and the other is the method based on probabilistic graphical models.

[0003] The social relationship mining method based on similarity measurement predicts relationship links by calculating the similarity between nodes. The higher the similarity between nodes, the greater the likelihood of generating a link. Liben-Nowell et al. adopted an unsupervised learning method [Liben-Nowell D, Kleinberg J. The link-prediction problem for social networks[J]. Journal of the American society for information science and technology, 2007, 58(7): 1019-1031], utilized the similarity of the content and structure published by users, and calculated the similarity between node pairs based on the principle of homophily of network nodes for social relationship link prediction. Lichtenwalter et al. [Lichtenwalter R N, Chawla N V. Vertex collocation profiles: subgraph counting for link analysis and prediction[C] / / Proceedings of the 21st international conference on World Wide Web. 2012: 1019-1028] considered the local structural information around node pairs based on the concept of vertex collocation profile (VCP) for topological link analysis and prediction. De et al. [De A, Ganguly N, Chakraba rti S. Discriminative link prediction using local links, node features and community structu re[C] / / 2013 IEEE 13th International Conference on Data Mining. IEEE, 2013: 1009-1018] combined global attributes, local attributes, and the connection density of the community middle layer to predict social relationships from multiple dimensions.

[0004] The social relationship mining method based on probabilistic graphical models uses Bayesian graphical models to model the joint probability between nodes. However, it is difficult to directly use Bayesian graphical models to discover the complex relationships hidden in social networks. Wang et al. proposed a local probabilistic graphical model [Wang C, Satuluri V, Parthasarathy S. Local probabilistic models for link prediction [C] / / Seventh IEEE international conference on data mining (ICDM 2007). IEEE, 2007: 322-331], which discovers hidden social relationships by estimating the co-occurrence probability of network node pairs. Zhu et al. [Zhu Y, Yan X, Getoor L, et al. Scalable text and link analysis with mixed-topic link models [C] / / Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining. 2013: 473-481] proposed a mixed-topic link model based on the idea of topic models and combined with the mixed-membership block model, realizing unsupervised learning of topic classification and association prediction of social relationships. Zhang et al. [Zhang J, Wang C, Yu P S, et al. Learning latent friendship propagation networks with interest awareness for link prediction [C] / / Proceedings of the 36th international ACM SIGIR conference on Research and development in information retrieval. 2013: 63-72] proposed a latent friendship propagation network (LFPN) to describe the network evolution centered around individuals, and used the latent friendship propagation network to model the social behavior of individuals, taking link information as the result of the combination of making friends and interests.

[0005] The above existing social relationship analysis methods usually utilize the text information in the social network combined with the network structure for analysis, and then measure the relationship between the text and the structural features to discover the social relationships in the network. However, social information such as image information is not fully utilized, and there is a lack of multi-dimensional information description in the relationship construction of social figures. Summary of the Invention

[0006] In view of the above problems, the present application proposes a social relationship analysis method based on multi-modal data, including the following steps:

[0007] S1, extract the social text and social image information of the person, convert them into text features and image features respectively, and count the person intimacy. Based on the person intimacy, construct a person social network graph;

[0008] S2, input the text features and image features into a multi-modal fusion model based on transformer to obtain fusion features;

[0009] S3, adopt the Si-SCAN graph clustering algorithm to analyze the person social network graph to obtain a social relationship clustering result, where the Si-SCAN graph clustering algorithm is constructed by introducing person intimacy and fusion feature information on the basis of the SCAN algorithm.

[0010] The above solution deeply analyzes the social relationship based on the information of two modalities, text and image. Through the design of the multi-modal information fusion model, it learns the interaction relationship between different modalities and generates a graph node embedding representation of multi-modal fusion. Through graph clustering analysis, it realizes the in-depth relationship analysis of the social network and can effectively discover potential social associations.

[0011] Preferably, S2 includes: concatenate the text features and image features and input them into the encoder of the multi-modal fusion model. Implementing information interaction between different modalities in the form of a single-stream model can reduce the computational complexity.

[0012] Preferably, S2 further includes:

[0013] Construct a text-image feature pair Z0 = [CLS, E, SEP, Q], where E is the text feature of person i, Q is the image feature of person i, [CLS] is an identifier, and [SEP] is a separator;

[0014] Construct text position encoding and image position encoding;

[0015] Add the text features and text position encoding, and the image features and image position encoding respectively, and input them into the encoder;

[0016] Select the output vector at the identifier position of the last layer of the encoder as the fusion feature vector z.

[0017] The above solution combines image and text modality information, designs a multi-modal fusion model based on transformers, deeply interacts and fuses the two modality data, and generates the embedding representation of nodes in subsequent graph clustering analysis. By means of complementary learning of multi-modal information, the accuracy of subsequent graph clustering analysis results can be improved.

[0018] Further, constructing the personnel social network graph in S1 includes using personnel as nodes of the personnel social network graph, and connecting the nodes corresponding to two personnel who appear in the same social image with an undirected edge.

[0019] Preferably, the node similarity in the Si-SCAN graph clustering algorithm in S3 includes structural similarity, personnel intimacy, and fusion feature similarity.

[0020] Further, the structural similarity is the ratio of the number of common neighbors of two nodes to the geometric mean, and is expressed by the formula:

[0021]

[0022] where σ1(v, w) is the structural similarity between node v and node w, and Γ(v) and Γ(w) are the sets of neighbor nodes of node v and node w respectively;

[0023] The personnel intimacy is expressed by the formula:

[0024] σ2(v, w) = α · p(v, w)

[0025] where σ2(v, w) is the personnel intimacy between node v and node w, p(v, w) is the co-occurrence times of the two nodes in the social image, and α is the adjustment coefficient;

[0026] The fusion feature similarity is expressed by the formula:

[0027]

[0028] where σ3(v, w) is the fusion feature similarity between node v and node w, z v 、z w are the fusion feature vectors of node v and node w respectively;

[0029] The node similarity σ(v, w) between node v and node w is expressed by the formula:

[0030] σ(v, w) = σ1(v, w) + σ2(v, w) + σ3(v, w).

[0031] The above solution designs a Si-SCAN clustering method, which introduces the measurement information of personnel intimacy and multi-modal features on the basis of SCAN [Xu X, Yuruk N, Feng Z, et al. Scan: a structural clustering algorithm for networks [C] / / Proceedings of the 13th ACM SIGKDD international conference on Knowledge discovery and data mining. 2007: 824-833]. It not only considers the structural data of network nodes, but also introduces multi-modal fusion information and image association information, and designs a node similarity measurement method from multiple dimensions, which can better measure the relationship between node characteristics and local structures in the graph data structure, thereby optimizing the clustering results.

[0032] Preferably, in S1, a BiLSTM-CRF model is used to extract text label information from social texts; a word vector model is used to convert the text label information into text features.

[0033] Preferably, in S1, a FACENet model is selected to extract image features.

[0034] In a second aspect, the present application proposes a social relationship analysis system based on multi-modal data, including:

[0035] An information extraction and feature conversion module, configured to extract social text and social image information of personnel, convert them into text features and image features respectively, and count personnel intimacy, and construct a personnel social network graph based on personnel intimacy;

[0036] A multi-modal data fusion module, configured to input the text features and image features into a multi-modal fusion model based on transformer to obtain fusion features;

[0037] A social relationship clustering module, configured to analyze the personnel social network graph by using the Si-SCAN graph clustering algorithm to obtain a social relationship clustering result, and the Si-SCAN graph clustering algorithm is constructed by introducing personnel intimacy and fusion feature information on the basis of the SCAN algorithm.

[0038] In a third aspect, the present application proposes a computer-readable storage medium for social relationship analysis based on multi-modal data, on which one or more computer programs are stored, and when the one or more computer programs are executed by a computer processor, the above-mentioned method is implemented.

[0039] The present invention proposes a social relationship mining technology framework by combining multiple technical methods such as named entity recognition, face recognition, multimodal fusion, and graph clustering. The methods adopted by each module in this framework are extensible and replaceable, and can be flexibly applied to other relationship mining scenarios. Specifically, by designing a multimodal fusion model based on transformers, the social relationships between people are mined using different modal information, and the complementary fusion of different modal features is carried out to make up for the information loss of a single modality and reduce information redundancy, so that the representation of multimodal fusion features can be effectively learned; by designing the Si-SCAN graph clustering method, on the basis of the SCAN algorithm, combined with structural similarity, multimodal feature similarity, and personnel intimacy, the similarity measurement method is optimized by making full use of multimodal and multi-dimensional information, thereby improving the clustering results of the model and helping to mine potential social associations in the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The drawings help to further understand the present application. The elements of the drawings are not necessarily in proportion to each other. For the sake of description, only the parts related to the present invention are shown in the drawings.

[0041] Figure 1 It is a schematic flowchart of a social relationship analysis method based on multimodal data in an embodiment;

[0042] Figure 2 It is a framework diagram of a social relationship analysis based on multimodality in another embodiment;

[0043] Figure 3 It is the model structure of the named entity recognition model BiLSTM plus conditional random field (CRF) in another embodiment;

[0044] Figure 4 It is a schematic diagram of the face feature extraction model structure in another embodiment;

[0045] Figure 5 It is a schematic diagram of the transformer multimodal fusion model structure in another embodiment;

[0046] Figure 6 It is a schematic diagram of the social relationship analysis system structure based on multimodal data in another embodiment;

[0047] Figure 7 It is a schematic diagram of the computer system structure of an electronic device suitable for implementing the embodiments of the present application in another embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The present application will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention.

[0049] Figure 1 It is a schematic flowchart of a social relationship analysis method based on multi-modal data in an embodiment, specifically including:

[0050] S1, Extract the social text and social image information of personnel, convert them into text features and image features respectively, and count the personnel intimacy. Based on the personnel intimacy, construct a personnel social network diagram;

[0051] S2, Input the text features and image features into a multi-modal fusion model based on transformer to obtain fusion features;

[0052] S3, Use the Si-SCAN graph clustering algorithm to analyze the personnel social network diagram to obtain the social relationship clustering result. The Si-SCAN graph clustering algorithm introduces personnel intimacy and fusion feature information on the basis of the SCAN algorithm.

[0053] Figure 2 It is a multi-modal based social relationship analysis framework diagram in an embodiment. This solution is mainly divided into three parts:

[0054] 1. Social information collection. With the help of named entity recognition and face recognition technologies, automatically extract text and image information in the social network, and use word vector models and face recognition algorithms to convert the information of the two modalities into feature vectors.

[0055] 2. Embedding representation of graph nodes. By designing a multi-modal fusion model based on transformer, interactively learn the information between different modalities and output multi-modal fusion features as the embedding representation of graph nodes.

[0056] 3. Graph clustering algorithm analysis. Combine network structure, image information and multi-modal graph node features, design a similarity measurement method Si-SCAN to divide the social network, and mine the social relationships between personnel.

[0057] In a specific embodiment, the analysis process of the social relationship analysis method based on multi-modal data includes:

[0058] 1. Text data labeling. Collect various social information related to personnel in the social network. For the convenience of representation, number all the personnel in the collected data from 1,..., N, and each number corresponds to one person. First, extract all the text content related to the user, and set to extract key text information from four aspects: educational background, living area, career direction, and hobbies. Specifically, eight kinds of information, namely education level, school, place of work, place of birth, work industry, job position, skill label, and interest label, are used to depict the user characteristics.

[0059] 2. Extract entity tag information. Use named entity recognition technology (NER) to automatically identify various entity information such as person names, place names, organization names, times, dates, etc. in a large amount of text. Figure 3 This is a schematic diagram of the model structure of the classic named entity recognition model BiLSTM plus conditional random field (CRF) selected in this embodiment. Each word in the text is converted into a word feature c. Using random initialization or a pre-trained model, the predicted label of each word is output through the NER model. BIO sequence labeling is selected in this embodiment. The extracted entity information is corresponding to the set text data labels. The entity information corresponding to each type of label does not exceed 8. If it is insufficient, it is filled with the character [UNK] to obtain 64 entity tag information.

[0060] 3. Text feature representation. Use the word2vec [Mikolov T, Chen K, Corrado G, et al. Efficient estimation of word representations in vector space [J]. arXiv preprint arXiv: 1301.3781, 2013] word vector model to convert the entity tag information into text features, denoted as where E i = e1,..., e k , e represents the word vector corresponding to the entity tag, k is the maximum number of entity tags with a value of 64, and the feature dimension de is 128 dimensions.

[0061] 4. Representation of personnel intimacy. Collect the group photo images of personnel through the social network platform. In order to represent the degree of personnel association in the group photo pictures, use statistical methods to count the number of times each pair of personnel appears in all group photos as the representation of personnel intimacy.

[0062] 5. Extract face features. Convert the group photo picture information into image features. After removing duplicates from all pictures, use face recognition technology to extract the face features of each person in the group photo. Obtain their face features in different group photos through the face recognition model where Q i = [q1,..., q m, m <= 8, where q represents the face feature vector of the i-th person in a group photo, and m represents the maximum number of group photos of a person. If the number of group photos of this person is less than 8, random vectors are used for feature completion; if the number of group photos is greater than 8, they are sorted in descending order according to the number of people in the group photos, and the top 8 photos with the largest number of people are selected to extract face features. In this embodiment, the FACENet [Schroff F, Kalenichenko D, Philbin J. Facenet: A unified embedding for face recognition and clustering [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2015: 815-823] model is selected to extract face features, and the feature dimension d q is 128 dimensions. Figure 4 is the schematic structural diagram of the face feature extraction model selected in this embodiment.

[0063] 6. Construct a social network graph. Each person represents a node, and for those who appear in the same group photo, an undirected edge is connected between the corresponding nodes to construct an undirected social network graph.

[0064] 7. Embedding representation of graph nodes. In this embodiment, combining image and text modality information, a multi-modal fusion model based on transformer is designed to deeply interact and fuse the two modality data to generate the embedding representation of the nodes. To reduce the computational complexity, the text and image features are concatenated and then input into the transformer encoder to achieve information interaction between different modalities in the form of a single-stream model.

[0065] Figure 5 is the schematic structural diagram of the transformer multi-modal fusion model in this embodiment. For any person i, construct a text-image feature pair

[0066] Z0 = [CLS, E, SEP, Q]

[0067] where E is the text feature of person i, Q is the image feature of person i, [CLS] is an identifier, and [SEP] is a separator.

[0068] After that, construct the position encodings of text and image. The text position encoding represents the embedding of the position information of the input 8 entity label sequences, and the image position encoding represents the embedding of the number of each image in the image dataset. By randomly initializing vectors at different positions as the position encodings, the dimension of the position encodings is the same as that of the text feature and the image feature.

[0069] Then, the text features and image features are added to the text position code and image position code respectively, and input into the transformer encoder for interactive learning in both text and image modes. Through continuous iterative updates of multiple layers of transformers, the correlation information between facial features and text features in different scenarios is learned. The output vector of the dth layer of the encoder is

[0070] Z d = Transformer(Z d-1 )+Z d-1 , d = 1, ..., 6

[0071] The output vector of the position of the encoder identifier [CLS] of the last layer is selected as the multimodal fusion feature representation z. In the preferred embodiment, the hidden layer size is set to 768, the number of multi-head attention heads is 12, and the number of transformer layers is 6. The training of the multimodal feature fusion model adopts the image-text alignment task (Image-Text Matching, ITM). The input image-text pair is randomly replaced, and then the model is used to predict whether there is a corresponding relationship between the input image and text, and the model is continuously optimized.

[0072] 8. Graph clustering analysis. A graph clustering method is used to divide the social network graph and mine the social relationships between people. This embodiment designs a Si-SCAN clustering method, in which the node similarity includes structural similarity, personal intimacy and fusion feature similarity.

[0073] The structural similarity is the ratio of the number of common neighbors of two nodes to the geometric mean, which can be expressed as:

[0074]

[0075] Among them, σ1(v, w) is the structural similarity between node v and node w, Γ(v) and Γ(w) are the sets of neighbor nodes of node v and w respectively;

[0076] The formula for personnel intimacy is:

[0077] σ2(v,w)=α·p(v,w)

[0078] Wherein, σ2(v, w) is the personnel intimacy between node v and node w, p(v, w) is the number of co-occurrences of the two nodes in the social image, and α is the adjustment coefficient, which is set to 0.1 in this embodiment;

[0079] The fusion feature similarity is expressed as:

[0080]

[0081] Among them, σ3(v, w) is the fusion feature similarity between node v and node w, and z v and z w are the fusion feature vectors of nodes v and w respectively;

[0082] Finally, the node similarity σ(v, w) between node v and node w is expressed by the formula:

[0083] σ(v, w) = σ1(v, w) + σ2(v, w) + σ3(v, w)

[0084] Define in the Si-SCAN algorithm, the neighbor nodes of node v

[0085] N(v) = {w ∈ Γ(v) | σ(v, w) ≥ ε}

[0086] The core node is

[0087]

[0088] Si-SCAN is the same as the SCAN algorithm. By calculating the neighbor nodes and core nodes in the graph, all nodes in the social network graph are clustered and partitioned.

[0089] In a specific embodiment, the Si-SCAN algorithm executes the following steps:

[0090] 1) At initialization, mark all nodes in the graph as unassigned nodes;

[0091] 2) For each unassigned node v, determine whether v belongs to a core node according to the similarity definition. If it is not a core node, mark it as a non-member. If it is a core node, expand a new clustering cluster, and assign the unassigned nodes and non-member nodes among its directly reachable nodes to this cluster, and repeat until each node is traversed;

[0092] 3) For the nodes marked as non-members, if the node is connected to two different clusters, it is determined as a bridge node, otherwise it is an outlier.

[0093] 9. Output of social relationship analysis results. Through Si-SCAN, the clustering analysis results of the social network graph are obtained. People with social associations will be classified into the same category, thereby being able to discover potential social relationships in the social network. The output of bridge nodes and outliers is also helpful for subsequent in-depth analysis of social network relationships.

[0094] Figure 6 It is a schematic diagram of the structure 600 of the social relationship analysis system based on multi-modal data in another embodiment of the present application, including:

[0095] The information extraction and feature transformation module 601 is configured to extract the social text and social image information of a person, transform them into text features and image features respectively, and calculate the intimacy of the person, and construct a person social network graph based on the intimacy of the person;

[0096] The multimodal data fusion module 602 is configured to input the text features and image features into a multimodal fusion model based on transformer to obtain fusion features;

[0097] The social relationship clustering module 603 is configured to analyze the person social network graph by using the Si-SCAN graph clustering algorithm to obtain the social relationship clustering result. The Si-SCAN graph clustering algorithm introduces the intimacy of the person and the fusion feature information on the basis of the SCAN algorithm.

[0098] Figure 7 FIG. shows a schematic structural diagram of a computer system 700 of an electronic device suitable for implementing the embodiments of the present application. The electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0099] As Figure 7 shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0100] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. The drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed.

[0101] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 709 and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above functions defined in the method of the present application are performed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0102] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).

[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0104] On the other hand, this application also provides a computer-readable storage medium, which can be included in the electronic device described in the above embodiments; or it can exist separately and not be assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: extract the social text and social image information of a person, convert them into text features and image features respectively, and calculate the intimacy of the person, and construct a person social network graph based on the intimacy of the person; input the text features and image features into a multi-modal fusion model based on transformer to obtain fusion features; use the Si-SCAN graph clustering algorithm to analyze the person social network graph to obtain a social relationship clustering result, and the Si-SCAN graph clustering algorithm introduces the intimacy of the person and the fusion feature information on the basis of the SCAN algorithm.

[0105] In the above embodiments, a deep social relationship analysis method and system based on multi-modal data fusion are designed, which fully utilize the text and image information in the social network, more comprehensively and accurately depict the potential social relationships between different personnel, evaluate the relationship strength, and can effectively mine the potential social associations in the network.

[0106] Although the content of the present application is specifically shown and introduced in combination with the preferred embodiments, those skilled in the art should understand that without departing from the spirit and scope of the present application defined by the appended claims, various changes made to the present application in form and details without creative efforts are within the protection scope of the present application.

Claims

1. A social relationship analysis method based on multi-modal data, characterized in that, Including the following steps: S1. Extract the social text and social image information of the person, convert them into text features and image features respectively, and count the person intimacy. Based on the person intimacy, construct a person social network graph; S2. Input the text features and image features into a multi-modal fusion model based on transformer to obtain fusion features; S3. Use the Si-SCAN graph clustering algorithm to analyze the person social network graph to obtain a social relationship clustering result. The Si-SCAN graph clustering algorithm is constructed by introducing the person intimacy and the fusion feature information on the basis of the SCAN algorithm. In the Si-SCAN graph clustering algorithm, the node similarity includes structural similarity, person intimacy and fusion feature similarity. The structural similarity is the ratio of the number of common neighbors of two nodes to the geometric mean, which is expressed by the formula: where σ1(v,w) is the structural similarity between node v and node w, and Γ(v) and Γ(w) are the sets of neighbor nodes of nodes v and w respectively; The person intimacy is expressed by the formula: σ2(v,w) = α·p(v,w) where σ2(v,w) is the person intimacy between node v and node w, p(v,w) is the co-occurrence times of the two nodes in the social image, and α is an adjustment coefficient; The fusion feature similarity is expressed by the formula: Among them, σ3(v,w) is the fusion feature similarity between node v and node w, z v 、z w are the fused feature vectors of nodes v and w respectively; The node similarity σ(v,w) between node v and node w is expressed by the formula: σ(v,w) = σ1(v,w)+σ2(v,w)+σ3(v,w).

2. The social relationship analysis method based on multimodal data according to claim 1, wherein S2 includes: After splicing the text features and image features, input them into the encoder of the multi-modal fusion model.

3. A method for analyzing social relationships based on multimodal data according to claim 2, characterized in that, S2 also includes: Construct a text-image feature pair Z0 = [CLS,E,SEP,Q], where E is the text feature of person i, Q is the image feature of person i, [CLS] is an identifier, and [SEP] is a separator; Construct text position encoding and image position encoding; Add the text features and text position encoding, and the image features and image position encoding respectively, and input them into the encoder; Select the output vector at the identifier position of the last layer of the encoder as the fusion feature vector z.

4. A social relationship analysis method based on multimodal data according to claim 1, characterized in that The construction of the person social network graph in S1 includes using the person as the node of the person social network graph, and connecting the nodes corresponding to two persons who appear in the same social image with an undirected edge.

5. A social relationship analysis method based on multi-modal data according to claim 1, characterized in that, S1 specifically includes: Use the BiLSTM-CRF model to extract text label information from social text; Use a word vector model to convert the text label information into text features.

6. A social relationship analysis method based on multi-modal data according to claim 1, characterized in that, S1 specifically includes: Select the FACENet model to extract image features.

7. A social relationship analysis system based on multimodal data, characterized in that, Including: An information extraction and feature conversion module configured to extract the social text and social image information of the person, convert them into text features and image features respectively, count the person intimacy, and construct a person social network graph based on the person intimacy; A multi-modal data fusion module configured to input the text features and image features into a multi-modal fusion model based on transformer to obtain fusion features; The social relationship clustering module is configured to analyze the personnel social network graph by using the Si-SCAN graph clustering algorithm to obtain a social relationship clustering result. The Si-SCAN graph clustering algorithm introduces the personnel intimacy and the fusion feature information on the basis of the SCAN algorithm. In the Si-SCAN graph clustering algorithm, the node similarity includes structural similarity, personnel intimacy, and fusion feature similarity. The structural similarity is the ratio of the number of common neighbors of two nodes to the geometric mean, which is expressed by the formula: where σ1(v,w) is the structural similarity between node v and node w, and Γ(v) and Γ(w) are the sets of neighbor nodes of node v and node w respectively; The personnel intimacy is expressed by the formula: σ2(v,w) = α·p(v,w) where σ2(v,w) is the personnel intimacy between node v and node w, p(v,w) is the co-occurrence times of the two nodes in the social image, and α is an adjustment coefficient; The fusion feature similarity is expressed by the formula: Among them, σ3(v, w) is the fusion feature similarity between node v and node w, and z v , z w are the fusion feature vectors of nodes v and w, respectively; The node similarity σ(v,w) between node v and node w is expressed by the formula: σ(v,w) = σ1(v,w) + σ2(v,w) + σ3(v,w).

8. A computer-readable storage medium for social relationship analysis based on multimodal data, on which one or more computer programs are stored, characterized in that, When the one or more computer programs are executed by a computer processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • User closeness-based mixed recommending system and method

    CN102880691A

  • A microblog social circle mining method and system based on an artificial immune network

    CN109597924A