Interdisciplinary team identification method based on scientific research social network

By using the BERT multi-label classification model and label propagation algorithm, the interdisciplinary feature representation of scientific researchers and the identification of interdisciplinary teams is solved, and the problem of difficulty in accurately identifying interdisciplinary teams in the existing technology is solved, achieving more efficient and accurate team recognition effects.

CN120123911APending Publication Date: 2025-06-10SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510277501.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing technology is difficult to accurately identify interdisciplinary teams, which ignores the interdisciplinary attributes of scientific researchers, resulting in poor identification results.

Method used

By pre-training and classifying the results data of scientific researchers using the BERT multi-label classification model, the interdisciplinary feature representation of scientific researchers is constructed, and the interdisciplinary team leaders and members are identified based on scientific research social networks and tag dissemination algorithms.

Benefits of technology

It has achieved more comprehensive and accurate interdisciplinary team recognition, improved identification efficiency and accuracy, and can better reflect the influence and collaboration value of scientific researchers in the team.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123911A_ABST
    Figure CN120123911A_ABST
Patent Text Reader

Abstract

The invention provides an interdisciplinary team identification method based on a scientific research social network. According to the method, a multi-level interdisciplinary feature extraction mechanism and a scientific research social network are combined, and interdisciplinary semantic features, cooperation history and centrality influence are creatively integrated to identify interdisciplinary teams. Firstly, a pre-training BERT model is used for scientific research achievements, and a multi-label classification result and interdisciplinary features of researchers are obtained. Secondly, constructing a scientific research social network, and calculating cooperation weight and interdisciplinary semantic similarity in combination with time period division and contribution degree analysis; and finally, identifying team core members and leaders by adopting a label propagation algorithm, and identifying a strategy optimization result through stability analysis and combination with the leaders. According to the method, the team stability and collaboration are remarkably enhanced, the generalization ability of the model is improved, and the method is suitable for interdisciplinary team identification in a scientific research social network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for identifying interdisciplinary teams based on a scientific research social network, belonging to the fields of social network and data mining. Background Art

[0002] With the continuous innovation and development of technology, multi-modal interdisciplinary cooperation has played an important role in increasing scientific research productivity. Interdisciplinary teams are more likely to achieve scientific breakthroughs compared to single-discipline research. However, with the popularization of interdisciplinary cooperation, the breadth and complexity of scientific research cooperation have been increasing continuously, leading to an expansion of the scale of scientific research teams, making the cooperation relationships among team members and between different teams more complex. Scientific researchers may participate in multiple teams in different fields simultaneously, resulting in a large increase in the number of interdisciplinary cooperation teams. Therefore, accurately identifying interdisciplinary teams is of great significance for promoting scientific research cooperation and the transformation of research results.

[0003] Current methods for identifying scientific research teams mainly rely on hierarchical clustering algorithms or algorithms based on association rules. Although these methods can reveal the structure of teams and the relationships between members to a certain extent, they usually ignore the interdisciplinary attributes of scientific researchers. Therefore, how to accurately identify interdisciplinary teams has become a technical problem to be solved urgently. Summary of the Invention

[0004] The main object of the present invention is to address the problem of identifying interdisciplinary teams, and a method for identifying interdisciplinary teams based on a scientific research social network is proposed.

[0005] The present invention is implemented by the following measures:

[0006] A method for identifying interdisciplinary teams based on a scientific research social network, comprising the following steps:

[0007] S1. Use the public paper dataset D open to pre-train the BERT multi-label classification model, and then use the pre-trained BERT multi-label classification model to classify the achievement dataset D res of scientific researchers, and construct the interdisciplinary feature representation V a of scientific researchers according to the classification results of scientific research achievements;

[0008] S2. Construct a scientific research social network G according to the cooperation records of scientific researchers (such as co-authored papers, participated projects, shared patents, etc.), and construct overlapping-time-span scientific research social networks by taking five years as a time window and sliding every three years where is the scientific research social network with the time window length, and T N is the number of constructed overlapping-time-span scientific research social networks. In each scientific research social network Construct the cooperation weight W of researchers considering the contribution degree and cooperation times of researchers to scientific research achievements, and represent V according to the interdisciplinary characteristics of researchers in step S1 a Calculate the interdisciplinary semantic similarity Sim of researchers;

[0009] S3. Based on the scientific research social network G constructed in step S2 T , cooperation weight W and interdisciplinary semantic similarity Sim, use the label propagation algorithm LabelAlg to identify interdisciplinary team leaders and team members, and obtain the interdisciplinary team distribution C of researchers T ;

[0010] S4. For the scientific research social network G constructed in step S2 T and the interdisciplinary team distribution C obtained in step S3 T Based on the stability criterion and the joint leader team merging strategy, perform stability member identification and joint leader team merging to obtain the final interdisciplinary team distribution C.

[0011] Furthermore, step S1 specifically includes:

[0012] S1.1: Construct a classification system D = {d 1 , d 2 , …, d m} containing m disciplinary fields, where d j represents a disciplinary field, j ∈ {1, 2, …, m}, such as computer science, biology, physics, etc.;

[0013] S1.2: Obtain the public paper dataset D open through academic open resources, and the acquisition methods include but are not limited to:

[0014] (1) Open access journal databases (such as PubMedCentral, arXiv, PLOS ONE)

[0015] (2) Academic institution knowledge bases (such as MIT OpenCourseWare, CERN Document Server)

[0016] (3) Academic graph data (such as Microsoft Academic Graph, AMiner open dataset)

[0017] (4) Proceedings of past years of authoritative conferences / journals in the field (obtained through the publisher API interface)

[0018] Use the public paper dataset D open to pre-train the BERT model using the method M bertPerform pre-training to obtain a pre-trained BERT multi-label classification model BERT-PretrainD, where the classification labels follow the classification system D in step S1.1, and M bert The steps are as follows:

[0019] (1) For each paper in D open use the input representation method of the BERT model to obtain the input sequence X open :

[0020] (2) Use the BERT model to perform a multi-label classification task on the input sequence X open to obtain the multi-label classification probability distribution:

[0021] (3) For the multi-label classification task, use Focal Loss (FL) to replace the traditional cross-entropy loss as the loss function. FL can handle the class imbalance problem in multi-label classification by adjusting the losses of easy and hard samples;

[0022] (4) Use the above method for all papers in D open to obtain the pre-trained BERT multi-label classification model BERT-PretrainD.

[0023] S1.3: Construct the achievement dataset D of researchers res , and the construction methods include but are not limited to:

[0024] (1) The knowledge management system within the research institution (such as the university IR institutional repository)

[0025] (2) Public data on academic social platforms (such as the API interfaces of ResearchGate and AcademicTree)

[0026] (3) The achievement list on the personal homepage of scholars (which needs to be crawled following the robots protocol)

[0027] (4) The achievement list in the team's declaration materials (which needs to be de-identified)

[0028] For each research achievement in the achievement dataset D of researchers res use the input representation method of the BERT model (defined the same as in step S1.2) to obtain the input sequence X res . For the input sequence X res use the pre-trained BERT multi-label classification model BERT-PretrainD in step S1.2 to perform a multi-label classification task to obtain the classification probability distribution z of each research achievement res, and use the output of the last layer of BERT-PretrainD as the feature representation h of the research results res .

[0029] S1.4: Use the classification probability distribution z in step S1.3 res and the feature representation h of the research results res to obtain the interdisciplinary feature representation V of the researchers a .

[0030] Furthermore, step S2 specifically includes:

[0031] S2.1: Construct a set of researchers A = {a 1 , a 2 , …, a N}, where a k and a l represent the k-th and l-th researchers respectively, k, l ∈ {1, 2, …, N} represents the number, and N represents the number of researchers. Construct a research social network G = (V, E) based on the cooperation records of the researchers A (such as co-authored papers, participated projects, shared patents, etc.), where the node set V represents the researcher nodes and the edge set E represents the cooperation relationships of the researchers. And construct a research social network with overlapping time periods by taking five years as the time window and sliding every three years where is the research social network with the time window length, T i ∈ {T 1 , T 2 , … T N}, T N is the number of research social networks with overlapping time periods constructed. Consider the contribution degree and cooperation times of the researchers to the research results in each research social network to construct the cooperation weight W of the researchers. For researchers a k and a l , the edge weight W kl is calculated as follows:

[0032]

[0033] where:

[0034] β k and β l represent the author positions of researchers a k and a l in the cooperative research results respectively;

[0035] and represent researchers a k and al The author contribution weight in the collaborative research results;

[0036] W kl represents the closeness of cooperation between researchers a k and a l ;

[0037] S2.2: Calculate the interdisciplinary semantic similarity of researchers.

[0038] According to the interdisciplinary feature representation defined in step S1.4, use the cosine similarity formula to calculate the cross-semantic similarity value Sim for each pair of researchers with a cooperative relationship.

[0039] Furthermore, step S3 specifically includes:

[0040] S3.1: Identify the leader node. Team leaders usually have strong academic influence, rich interdisciplinary cooperation experience, and high interdisciplinary similarity. This step is based on the centrality influence index C center to quantify the importance of researchers in the social network, and identify the team leader through the core position of the node in the network. For the researcher node a k , its centrality influence index C center is defined by the following formula:

[0041] C center (a k ) = ρ * C degree (a k ) + δ * C between (a k )

[0042] where:

[0043] C degree (a k ) represents the degree centrality of node a k , and its calculation formula is: C degree (a k ) = ∑ l W kl , where W kl is the cooperation weight between researchers a k and a l ;

[0044] C between (a k ) represents the betweenness centrality of node a k , and its calculation formula is where σ st represents the number of shortest paths between node s and node t, and σ st (a k)Indicates the number of shortest paths passing through node a k ;

[0045] ρ and δ are weight coefficients, satisfying ρ + δ = 1.

[0046] Centrality influence index C center By comprehensively considering the cooperation breadth of researchers and their role as a bridge in information transmission in the cooperation network, the importance of researchers in the scientific research social network can be evaluated more comprehensively, so as to identify potential interdisciplinary team leaders.

[0047] Sorted from high to low, select the top 5% of the nodes with the centrality influence index C in the scientific research social network C T constructed in step S2 as the leader node set L. center

[0048] S3.2: Based on the scientific research social network G T constructed in step S2, the cooperation weight W, and the interdisciplinary semantic similarity Sim, combined with the leader node set L identified in step S3.1, use the label propagation algorithm LabelAlg to identify interdisciplinary team members. The steps of the label propagation algorithm LabelAlg are as follows:

[0049] Input parameters: number of iterations T, label count threshold r, interdisciplinary attribute ratio ε;

[0050] Output parameter: distribution G of researchers in interdisciplinary teams T ;

[0051] In the initialization stage, for each node a in the leader node set L L , calculate its initial label set l L , and set the initial label sets of the remaining nodes in G T to be empty;

[0052] In the iterative propagation stage, repeat the following steps until the maximum number of iterations T is reached:

[0053] (4) For each researcher node, calculate the label selection weight w in the label sets of its neighbor nodes;

[0054] (5) For the selected researcher node a k , add the label with the maximum label selection weight w to the label set l k of the selected researcher node a k .

[0055] Label screening and community output stage. For each researcher node a k 's label set l k ​, screening and eliminating tags whose tag ratio does not reach the tag frequency threshold r to obtain each researcher node a k The final team distribution C(a k ), thereby obtaining the interdisciplinary team distribution C corresponding to the scientific research social network G T . T .

[0056] Furthermore, step S4 specifically includes:

[0057] S4.1: Identifying stable members in the team.

[0058] Among the interdisciplinary teams C corresponding to the scientific research social network G obtained in step S3 T , there will be some overlap. Therefore, in this step, based on the stability criterion S of researchers, the interdisciplinary team C T is identified for stable members, thereby obtaining the stable interdisciplinary team C T , where the stability criterion S refers to the number of times a researcher continuously participates in the same team or cooperates with the same researchers in multiple time periods. Researchers with higher stability usually maintain a core role in the teams of multiple time periods, indicating that their cooperation potential in interdisciplinary teams is stronger. The definition of stable members is researchers who continuously appear in the same team in multiple time periods. R

[0059] S4.2: Merging co-leader teams.

[0060] For the stable interdisciplinary team C obtained in step S4.1 R , the co-leader team merging strategy is used to merge the co-leader teams to obtain the final interdisciplinary team distribution C, where the steps of the co-leader team merging strategy are as follows:

[0061] For two teams C 1 and C 2 , calculate their member overlap degree O(C 1 , C 2 ). If the member overlap degree O(C 1 , C 2 ) ≥ 0.8, and both teams have clear leaders L 1 and L 2 , then the two teams are merged to obtain the co-leader team C joint , where C joint = C 1 ∪ C 2 , and the leaders L 1 and L 2 are used as co-leaders.

[0062] ​Compared with the prior art, the beneficial effects of the present invention are as follows:

[0063] (1) The present invention makes full use of the interdisciplinary feature representation of scientific researchers. By pre-training the BERT multi-label classification model with a public paper dataset, and then classifying the scientific research achievements by field, the interdisciplinary feature representation of scientific researchers can be calculated, and a more comprehensive and accurate interdisciplinary feature of scientific researchers can be obtained.

[0064] (2) The present invention uses the centrality influence index to identify the leaders of scientific research teams. Through the label propagation algorithm, the rapid propagation of leader labels is realized, which ensures the accuracy of interdisciplinary team identification while improving the identification efficiency of interdisciplinary teams.

[0065] (3) By comprehensively considering the frequency and contribution degree of scientific researchers' scientific research cooperation and introducing the calculation of interdisciplinary semantic similarity, the present invention can more accurately construct a scientific research social network and effectively reflect the influence and collaboration value of scientific researchers in the team.

[0066] (4) The present invention divides the scientific research social network based on time periods, introduces a stability criterion to identify the stable members of the team in the scientific research social network in different time periods, and introduces an overlap degree to merge the joint leader teams, so as to form a more synergistic interdisciplinary team. Description of the Drawings

[0067] Figure 1 It is a flowchart of the steps of the interdisciplinary identification method based on social network in the specific implementation scheme of the present invention. Specific Embodiments

[0068] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. As Figure 1 shown, the present invention provides an interdisciplinary identification method based on social network, which specifically includes the following steps:

[0069] S1: Construct the interdisciplinary feature representation of scientific researchers;

[0070] S1.1: According to domain knowledge, construct a classification system D = {d 1 , d 2 , …, d m} that contains m subject fields, such as computer science, biology, physics, etc.

[0071] S1.2: Use the public paper dataset D open to pre-train the BERT model using the pre-training method M bertPerform pre-training to obtain a pre-trained BERT multi-label classification model BERT-PretrainD, where the classification labels follow the classification system D in step S1.1, and M bert The steps are as follows:

[0072] (1) For each paper i in D open Use the input representation method of the BERT model to obtain the input sequence X i :

[0073]

[0074] Where:

[0075] title token Is the lexical representation of the title part;

[0076] keywords token Is the lexical representation of the keyword part;

[0077] abstract token Is the lexical representation of the abstract part;

[0078] [CLS] and [SEP] are special symbols in the BERT model, representing the sequence start and separator respectively;

[0079] Is the p-th word representation of paper i;

[0080] n token Is the lexical length of paper i.

[0081] (2) Perform a multi-label classification task on the input sequence X i Using the BERT model to obtain the multi-label classification probability distribution z i :

[0082]

[0083] Where Represents the classification probability of paper i in the subject field d j Of.

[0084] (3) Use Focal Loss (FL) as the loss function for the multi-label classification task, and its formula is:

[0085] FL(p t ) = -α t (1 - p t ) γ log(p t )

[0086] Where:

[0087] p t is the predicted probability, representing the probability that paper i belongs to subject field d j of the probability

[0088] σ is the sigmod function, used to map z ij to the probability value in [0, 1];

[0089] α t is the adjustment factor for category t, used to control class imbalance;

[0090] γ is the adjustment factor, used to control the degree of attention to difficult-to-classify samples.

[0091] (4) Use the output of the last layer of the BERT model as the feature representation h of paper i i : h i = BERT(X i ) [CLS] .

[0092] (5) Apply the above method to all papers in D open to obtain the pre-trained BERT multi-label classification model BERT-PretrainD.

[0093] S1.3: Use the pre-trained BERT multi-label classification model BERT-PretrainD to perform multi-label classification tasks on the achievement dataset D of researchers res to obtain the classification results of each research achievement and the feature representation of the research achievement. The specific steps are as follows:

[0094] (1) For each research achievement in D res , use the input representation method of BERT to obtain the input sequence X res (defined in the same way as in step S1.2).

[0095] (2) Use the pre-trained BERT multi-label classification model BERT-PretrainD to perform multi-label classification tasks on the input sequence X res to obtain the multi-label classification probability distribution z res (defined in the same way as in step S1.2).

[0096] (3) Use the output of the last layer of the BERT model as the feature representation h of the research achievement dataset D res : h res = BERT(X res ) res ) [CLS] .

[0097] (4) For D resFor each scientific research achievement, use the above method to obtain the probability distribution z of each scientific research achievement in each subject field res and the feature representation h res .

[0098] S1.4: Use the classification probability distribution z in step S1.3 res and the feature representation h of the scientific research achievement res to obtain the interdisciplinary feature representation V of the researcher a , and the specific steps are as follows:

[0099] For the scientific research achievement i of researcher a, assume it belongs to h subject fields d 1 , d 2 …, d h , then for each subject field d j use the weighted probability formula to obtain the weighted probability as where represents the weighted probability of the scientific research achievement i of researcher a in the subject field d j , and h is the number of subject fields to which the scientific research achievement belongs.

[0100] For all the scientific research achievements of researcher a in the subject field d j use the weighted representation formula to obtain the weighted representation

[0101]

[0102] where n a represents the total number of scientific research achievements of researcher a, β i represents the author position of researcher a in the scientific research achievement i represents the author contribution weight of researcher a in the scientific research achievement i.

[0103] For the weighted representations of researcher a in all subject fields d 1 , d 2 ,…, d m use the composite representation to obtain the interdisciplinary feature representation V of the researcher , where a S2: Construct a scientific research social network, and calculate the interdisciplinary semantic similarity Sim of researchers according to the interdisciplinary feature representation V of researchers in S1

[0104] S2.1: Construct a scientific research social network a Calculate the interdisciplinary semantic similarity Sim of researchers

[0105] S2.1: Construct a scientific research social network

[0106] Construct a set of researchers A = {a 1 , a2 ,…, a N}, where a k represents the k-th researcher, and N represents the number of researchers. A research social network G=(V, E) is constructed based on the cooperation records of researchers (such as co-authored papers, participated projects, shared patents, etc.), where the node set V represents researchers and the edge set E represents the cooperation relationships between researchers. And a research social network with overlapping time periods is constructed with a five-year time window and a three-year sliding interval where is the research social network with the length of the time window, and T N is the number of research social networks with overlapping time periods constructed. In each research social network , the cooperation weight W of researchers is constructed considering the contribution degree and cooperation times of researchers to research results. For researchers a k and a l , the edge weight W kl is calculated as follows:

[0107]

[0108] where:

[0109] β k and β l respectively represent the author positions of researchers a k and a l in the cooperative research results;

[0110] and respectively represent the author contribution weights of researchers a k and a l in the cooperative research results;

[0111] W kl represents the cooperation tightness between researchers a k and a l .

[0112] S2.2: Calculate the interdisciplinary semantic similarity Sim of researchers;

[0113] For each pair of researchers a k and a l , the interdisciplinary semantic similarity value Sim k and V l of the interdisciplinary feature representations V kl defined in step S1.4 is calculated using cosine similarity as follows:

[0114]

[0115] Wherein:

[0116] V k and V l respectively represent the interdisciplinary feature representations of researcher a k and a l ;

[0117] Sim kl represents the interdisciplinary semantic similarity between researcher a k and a l . The larger the value, the higher the similarity.

[0118] S3: Based on the scientific research social network G T , cooperation weight W, and interdisciplinary semantic similarity Sim constructed in step S2, use the label propagation algorithm LabelAlg to identify the interdisciplinary team leaders and team members, and obtain the interdisciplinary team distribution G T of researchers;

[0119] S3.1: Identify the leader nodes. Team leaders usually have strong academic influence, rich interdisciplinary cooperation experience, and high interdisciplinary similarity. This step quantifies the importance of researchers in the social network based on the centrality influence index C center to identify team leaders through the core position of nodes in the network. For researcher node a k , the calculation formula of its centrality influence index C center is defined as:

[0120] C center (a k ) = ρ * C degree (a k ) + δ * C between (a k )

[0121] Wherein:

[0122] C degree (a k ) represents the degree centrality of node a k , and the calculation formula is: C degree (a k ) = ∑ l W kl , where W kl is the cooperation weight between researcher a k and a l ;

[0123] C between (a k ) represents a kThe betweenness centrality of a node, the calculation formula is where σ st represents the number of shortest paths between node s and node t, σ st (a k ) represents the number of shortest paths passing through node a k ;

[0124] ρ and δ are weight coefficients, satisfying ρ + δ = 1.

[0125] The centrality influence index C center Comprehensively considering the cooperation breadth of researchers and their role as bridges in information transmission in the cooperation network can more comprehensively evaluate the importance of researchers in the scientific research social network, so as to identify potential interdisciplinary team leaders.

[0126] Sorted from high to low, select the nodes with the top 5% of the centrality influence index C T in the scientific research social network G center constructed in step S2 as the leader node set L, where the calculation method of L is:

[0127] L = {a k |C center (a k ) ≥ threshold}

[0128] where:

[0129] C center (a k ) represents the centrality influence index of researcher a k ;

[0130] threshold is the threshold determined by sorting the top 5% of C center (a k ).

[0131] S3.2: Based on the scientific research social network G T constructed in step S2, the cooperation weight W and the interdisciplinary semantic similarity Sim, combined with the leader node set L identified in step S3.1, use the label propagation algorithm LabelAlg to identify interdisciplinary team members, where the steps of the label propagation algorithm LabelAlg are as follows:

[0132] Input parameters: number of iterations T, label count threshold r, interdisciplinary attribute ratio ε;

[0133] Output parameter: interdisciplinary team distribution C T ;

[0134] In the initialization stage, for each node a in the leader node set L L, calculate its initial label set \(l\) L : \(l\) L = \(ID(a\) L ), where \(ID(a\) L ) represents the unique identifier of researcher \(a\) L , and the initial label sets of the remaining nodes in \(G\) T are set to be empty;

[0135] Iterative propagation stage. For each iteration step \(t\in\{1,2,\ldots,T\}\), perform the following steps:

[0136] (1) For each researcher node \(a\) k , calculate the label selection weight \(w\) in the label sets of its neighbor nodes. Assume the neighbor node is \(a\) l , then the calculation method of the label selection weight is:

[0137] \(w\) kl = \(\varepsilon\times Sim\) kl +(1 - \(\varepsilon)\times W\) kl

[0138] Where:

[0139] \(Sim\) kl represents the interdisciplinary semantic similarity between researcher \(a\) k and \(a\) l ;

[0140] \(W\) kl represents the cooperation weight between researcher \(a\) k and \(a\) l .

[0141] (2) For the selected researcher node \(a\) k , add the label with the maximum label selection weight \(w\) to the label set \(l\) k of the selected researcher node \(a\) k .

[0142] Label screening and community output stage. For the label set \(l\) k of each researcher node \(a\) k , screen and remove the labels whose label ratio does not reach the label frequency threshold \(r\), that is, for each label \(l\) in the label set of \(a\) k , calculate its ratio \(p\) l , if \(p\) l ≥\(r\), then retain the label \(l\) as the label of \(a\) k , and obtain the final team distribution \(C(a\) k ) of each researcher node \(a\) k ), thus obtaining the interdisciplinary team distribution \(G\) T corresponding to the scientific research social network \(G\) T .

[0143] S4: The scientific research social network G constructed in step S2 T and the interdisciplinary team distribution C obtained in step S3 T Based on the stability standard and the joint leader team merging strategy, stable members are identified and joint leader teams are merged to obtain the final interdisciplinary team distribution C.

[0144] S4.1: Identify key members of the team.

[0145] The corresponding scientific research social network G obtained in step S3 T The interdisciplinary team T There will be some overlap, so this step is based on the stability criterion S of the researchers and the interdisciplinary team G T Identify stability members to obtain a stability interdisciplinary team C R , where the stability standard S refers to the number of times a researcher has continued to participate in the same team or collaborate with the same researchers over multiple time periods. Researchers with higher stability usually maintain core roles in teams over multiple time periods, indicating that they have greater potential for collaboration in interdisciplinary teams. Stable members are defined as researchers who continue to appear in the same team over multiple time periods. Researchers who frequently change between different teams are considered mobile members. Calculation method of stability standard S and stability of interdisciplinary teams C R The merging criteria are as follows:

[0146] Assume there are T time periods, and the number of teams identified in each time period is T 1 ,T 2 ,…,T T , each team contains several researchers. For each researcher a k , defined in time period T t The participating teams are T k,t , then the stability criterion S(a k ) is calculated as:

[0147]

[0148] in:

[0149] T k,t Indicates that researchers k In time period T t The teams involved;

[0150] T k,t-1 Indicates that researchers k In time period T t-1 The teams involved;

[0151] |T k,t ∩T k,t-1 represents the number of teams that researcher a k participated in jointly during time period T t and T t-1 ;

[0152] |T k,t-1 | represents the number of teams that researcher a k participated in during time period T t-1 .

[0153] For researcher a k , using the stability threshold θ, divide them into stable members M stable or mobile members M mobile : If S(a k )≥θ, then a k ∈M stable ; If S(a k )<θ, then a k ∈M mobile . Add the stable members M stable to the core team as members of the merged team.

[0154] S4.2: For the C R obtained in step S4.1, use the joint leader team merging strategy to merge the joint leader teams, and obtain the final interdisciplinary team distribution C, where the steps of the joint leader team merging strategy are as follows:

[0155] For two teams C 1 and C 2 , calculate their member overlap O(C 1 ,C 2 ):

[0156]

[0157] Where: |C 1 ∩C 2 | represents the number of common members of the two teams, and min(|C 1 |,|C 2 |) represents the smaller number of members in the two teams.

[0158] If for two teams C 1 and C 2 , if the member overlap O(C 1 ,C 2 )≥0.8, and both teams have clear leaders L 1 and L 2 , then merge the two teams to obtain the joint leader team Cjoint , where C joint = C 1 ∪ C 2 , and the leaders L 1 and L 2 are taken as co-leaders.

Claims

1. A method for identifying interdisciplinary teams based on scientific research social networks, characterized in that: The following steps are involved: S1: Constructing interdisciplinary characteristics representation of researchers; S2: Construct a scientific research social network and the cooperation weights of researchers, and calculate the interdisciplinary semantic similarity of researchers based on the interdisciplinary characteristics of researchers in S1; S3: Based on the scientific research social network, cooperation weight and interdisciplinary semantic similarity constructed in step S2, the label propagation algorithm LabelAlg is used to identify the leaders and members of the interdisciplinary team, and the interdisciplinary team distribution of scientific researchers is obtained; S4: Based on the stability standard and the joint leader team merging strategy, the scientific research social network constructed in step S2 and the interdisciplinary team distribution obtained in step S3 are subjected to stable member identification and joint leader team merging to obtain the final interdisciplinary team distribution.

2. The method for identifying an interdisciplinary team based on a scientific research social network according to claim 1, characterized in that: The step S1 specifically includes: S1.1: Construct a classification system D = {d1, d2, …, d m }, where d j represents the subject field, j∈{1,2,…,m}; S1.2: Using the public paper dataset D open Use pre-training method M for the BERT model bert Pre-training is performed to obtain a pre-trained BERT multi-label classification model BERT-PretrainD, where the classification labels follow the classification system D in step S1.1, and M bert The steps are: (1) For D open Each paper in uses the BERT model’s input representation method to obtain the input sequence X open : (2) For the input sequence X open Use the BERT model for multi-label classification tasks and obtain the multi-label classification probability distribution: (3) Focal Loss is used to replace the traditional cross entropy loss as the loss function for multi-label classification tasks. Focal Loss handles the class imbalance problem in multi-label classification by adjusting the loss of difficult and easy samples. (4) For D open All papers in the paper use the above steps (1)-(3) to obtain the pre-trained BERT multi-label classification model BERT-PretrainD; S1.3: Dataset D for researchers’ achievements res Each scientific research result in the BERT model uses the input representation method to obtain the input sequence X res ; For the input sequence X res Use the BERT multi-label classification model BERT-PretrainD pre-trained in step S1.2 to perform multi-label classification tasks and obtain the classification probability distribution z of each scientific research result res , and use the last layer output of BERT-PretrainD as the feature representation of scientific research results h res ; S1.4: Use the classification probability distribution z from step S1.3 res and the characteristic representation of scientific research results res , we get the interdisciplinary feature representation V of researcher a a .

3. The method for identifying an interdisciplinary team based on a scientific research social network according to claim 2, characterized in that: The step S2 specifically includes: S2.1: Construct a set of researchers A = {a1, a2, …, a N }, where a k and a l Denote the kth and lth researchers respectively, k, l∈{1,2,…,N}, N represents the number of researchers; construct a scientific research social network G=(V,E) based on the cooperation records of researcher A, where the node set V represents the researcher nodes and the edge set E represents the cooperation relationship of the researchers; and construct a scientific research social network with overlapping time periods by sliding every three years with a five-year time window. in is the scientific research social network of the time window length, T i ∈{T1,T2,....T N }, T N is the number of scientific research social networks constructed in overlapping time periods; in each scientific research social network Considering the contribution of researchers to scientific research results and the number of collaborations, the collaboration weight W of researchers is constructed. k and a l , edge weight W kl The calculation formula is: Where: β k and β l Represents researchers a k and a l Authorship in collaborative research results; and Represents researchers a k and a l The weight of the author's contribution in the collaborative research results; W kl Indicates that researchers k and a l the closeness of cooperation between them; S2.2: Calculate the interdisciplinary semantic similarity of researchers; based on the interdisciplinary feature representation defined in step S1.4, use the cosine similarity formula to calculate the cross-semantic similarity value Sim for each pair of researchers with a cooperative relationship.

4. The method for identifying an interdisciplinary team based on a scientific research social network according to claim 3, characterized in that: The step S3 specifically includes: S3.1: Identify leader nodes; based on the centrality influence indicator C center To quantify the importance of researchers in social networks, we can identify team leaders by the core position of nodes in the network. k , its central influence index C center The calculation formula is defined as: C center (a k )=ρ*C degree (a k )+δ*C between (a k ) Where: C drgree (a k ) is represented by a k The degree centrality of a node is calculated as: degree (a k )=∑ l W kl , where W kl For researchers k and a l The cooperation weight between between (a k ) is represented by a k The betweenness centrality of a node is calculated as where σ st represents the number of shortest paths between node s and node t, σ st (a k ) means passing through node a k The number of shortest paths; ρ and δ are weight coefficients, satisfying ρ+δ=1; According to the order from high to low, the scientific research social network G constructed in step S2 T Select the central influence indicator C center The top 5% of nodes serve as the leader node set L; S3.2: Based on the scientific research social network G constructed in step S2 T , cooperation weight W and interdisciplinary semantic similarity Sim, combined with the leader node set L identified in step S3.1, use the label propagation algorithm LabelAlg to identify interdisciplinary team members.

5. The method for identifying an interdisciplinary team based on a scientific research social network according to claim 4, characterized in that: The steps of the label propagation algorithm LabelAlg in step S3.1 are as follows: Input parameters: number of iterations T, label number threshold r, cross-disciplinary attribute ratio ε; Output parameter: Distribution of interdisciplinary teams of researchers C T ; Initialization phase; For each node a in the leader node set L L , calculate its initial label set l L , G T The initial label sets of the remaining nodes in are set to empty; Iterative propagation phase; Repeat the following steps until the maximum number of iterations T is reached: (1) For each researcher node, calculate the label selection weight w between it and its neighbor node label set; (2) For the selected researcher node a k , add the label with the maximum label selection weight w to the selected researcher node a k The label set l k middle; (3) Label screening and community output stage: For each researcher node a k The label set l k , filter out labels whose label ratio does not reach the label number threshold r, and get each researcher node a k Final team distribution C(a k ), thus obtaining the corresponding scientific research social network G T Interdisciplinary team distribution C T .

6. The method for identifying an interdisciplinary team based on a scientific research social network according to claim 5, characterized in that: The step S4 specifically includes: S4.1: Identify stable members in the team; Based on the stability standard S of researchers, identify stable members in the interdisciplinary team C T Identify stability members to obtain a stability interdisciplinary team C R , where the stability standard S refers to the number of times a researcher has continued to participate in the same team or collaborate with the same researchers in multiple time periods; researchers with high stability have maintained core roles in teams in multiple time periods, indicating that they have stronger potential for collaboration in interdisciplinary teams; stable members are defined as researchers who continue to appear in the same team in multiple time periods; S4.2: Merge the joint leadership team; the stability interdisciplinary team C obtained in step S4.1 R The joint leader team merging strategy is used to merge the joint leader teams to obtain the final interdisciplinary team distribution C. The steps of the joint leader team merging strategy are as follows: for two teams C1 and C2, calculate their member overlap O(C1, C2). If the member overlap O(C1, C2) ≥ 0.8 and both teams have clear leaders L1 and L2, then merge the two teams to obtain the joint leader team C. joint , where C joint =C1∪C2, and make leaders L1 and L2 as joint leaders.