A generalized representation and collaborative fusion method for multi-source data

By generalizing and collaboratively fusing multi-source data, the problem of insufficient generalization in multi-source data fusion is solved, and accurate and reliable data fusion effects are achieved.

CN114330520BActive Publication Date: 2025-09-19HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111577599.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-09-19
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

The existing technology lacks generalization when fusing multi-source data, resulting in a lack of accuracy and reliability in the fusion results, especially when the data scale is small.

Method used

A generalized representation and collaborative fusion method of multi-source data is adopted. By dividing the data units into abnormal intervals on the left, normal intervals and abnormal intervals on the right, and defining generalization functions and collaborative distances, a generalized array structure is constructed to achieve collaborative fusion between data units.

Benefits of technology

Accurate decision-making based on multi-source data is achieved, precise, reliable and generalized fusion results are obtained, and the problem of insufficient data coverage is overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114330520B_ABST
    Figure CN114330520B_ABST
Patent Text Reader

Abstract

The present invention discloses a generalized representation and collaborative fusion method for multi-source data. From the perspective of the decision-making impact of the data, the present invention provides a generalized representation of the combined data and the knowledge behind the data for multi-source data, thereby defining a unified data structure that can be used for horizontal precise comparison. For the generalized representation of multi-source data, the present invention provides a fusion method based on the degree of collaboration between entity nodes from the perspective of links, realizing the collaborative division and fusion of entity nodes in a generalized array structure. The multi-source data collaborative fusion method based on generalized representation can obtain accurate, reliable and generalized multi-source data fusion results for precise decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-source data representation and fusion, and in particular to a generalized representation and collaborative fusion method for multi-source data. Background Art

[0002] How to obtain more generalizable fusion results from source data is a key issue in data fusion. This issue is particularly prominent when the data size is small, and this also applies to multi-source data fusion. Generally, improving the generalizability of fusion results is to increase the size and source of the data. Even so, it is impossible to cover all types of data units. This inherent flaw leads to insufficient generalization in data fusion results. For example, because temperatures of 36.2 and 36.5 are both normal temperature data, not fever data, there is no clear distinction between them. If a temperature dataset covers temperatures of 36.2 but not 36.5, the fusion results for this temperature dataset may lack generalizability. In fact, the absence of a temperature of 36.5 does not affect the diagnosis of fever based on the fusion results of the temperature data. Therefore, if a generalized representation can be used to mitigate the negative impact of missing non-critical data, the fusion results of a given dataset will have better generalizability.

[0003] Traditionally, the representation of a set of data units is confined to a population interval, with different data units distinguished based on their characteristic values ​​within that population interval. Therefore, while some data units may have large differences in their characteristic values, the differences in their decision impacts may be relatively small. On the other hand, while some data units may have small differences in their characteristic values, the differences in their decision impacts may be relatively large. For example, because a temperature of 37.3 is a sign of a fever, the difference in decision impact between temperatures of 37.2 and 37.3 is significant. However, because temperatures of 37.2 and 36.2 are both normal, they have nearly equal decision impacts.

[0004] From the perspective of decision impact, traditional data representation cannot accurately reflect the true differences between data units within an overall interval. Therefore, it is necessary to divide an overall interval into several sub-intervals, such as the abnormal interval on the left, the normal interval, and the abnormal interval on the right. Data units within the same sub-interval have the same interval index and different interval values. Data representation that covers decision impact appears to be more accurate and can handle situations where a sub-interval does not cover some data units. In addition, when the data representation covers decision impact, it can effectively distinguish data units with the same characteristic value but different interval index, and can also directly compare data units with the same interval index but different characteristic values. These characteristics reflect the generalization of data representation based on impact intervals.

[0005] A generalized data representation not only contains the value information of the source data but also the knowledge behind it, providing a more accurate representation of the source data. In other words, a generalized data representation is a combined representation of data and knowledge. Even so, a reasonable fusion strategy plays a crucial role in achieving accurate and reliable fusion results. Based on an array structure constructed from multi-source generalized data units, the partitioning and fusion of generalized data units is a collaborative process guided by the horizontal and vertical relationships between the data units.

[0006] In an array structure space composed of multiple link networks, a data unit located in one link network is called a branch node, and an entity node is composed of data units generated by the same entity and located in different link networks. The horizontal relationship between branch nodes depends on the horizontal collaborative distance between them. Using the generalized representation of the data elements in each branch node, the horizontal collaborative distance between two branch nodes is defined as the cumulative collaborative distance of the corresponding data elements in them. The vertical relationship between entity nodes is actually a collaborative relationship in the array structure. Here, collaborative fusion based on generalized representation is committed to obtaining accurate, reliable and generalized fusion results for precise decision-making, and overcoming the defect of insufficient data element coverage. Summary of the Invention

[0007] In order to solve the above technical problems existing in the prior art, the present invention proposes a generalized representation and collaborative fusion method for multi-source data. The specific technical solution is as follows:

[0008] A generalized representation and collaborative fusion method for multi-source data, characterized in that the method comprises the following steps: step 1) generalized representation of multi-source data;

[0009] Step 2) Collaborative fusion based on generalized representation.

[0010] Furthermore, the step 1) is specifically as follows:

[0011] m source data sets DS1,…,DS k ,…,DS m There are n entities (nodes) E1,…,E i ,…,E n Thus, a data set DS containing n branch nodes (or data units) from different entities is generated. k With expression DS k ={s k,1 ,…,s k,i ,…,s k,n}, entity node E composed of branch nodes distributed in different data sets i With expression E i =[s 1,i ,…,s k,i ,…,sm,i ] T ; If the data unit s k,i The data elements in come from L different data vectors (Data vector, DV), then we can define s k,i for

[0012] Given a normal range Data elements Its normalized representation is defined as

[0013]

[0014] here, and Indicates the lower and upper limits of the normal range; if Located in the abnormal area on the left Or the abnormal interval on the right Then its normalized expression is defined as or in, and Represents the minimum element of the non-normal interval on the left and the maximum element of the non-normal interval on the right;

[0015] Generalization function depending on In the interval index Position in the indicated interval

[0016]

[0017] Furthermore, the step 2) is specifically as follows:

[0018] Given a data unit and The distance D(s) between them k,i1 ,s k,i2 ) is the cumulative synergy distance between corresponding data elements therein;

[0019]

[0020] The synergy distance between two data elements is related to the synergy index of the data vectors where the two data elements are located;

[0021]

[0022] Given a generalized data vector and Ascending form of and Data vector and The rate of change of the data elements in Data vector and The degree of coordination is defined as follows:

[0023]

[0024] The synergy index of a data vector is the sum of the synergy between this data vector and other data vectors in the same dataset:

[0025]

[0026] Given a dataset DS k Middle branch node s k,i n-1 ascending distances D(s between the node and other branch nodes) k,i ,s′ k,1 ),…,D(s k,i ,s′ k,i′ ),…,D(s k,i ,s′ k,n-1 ), s k,i The distance change rate R d (s k,i ) is defined as the adjacent distance difference (D(s k,i ,s′ k,2 )-D(s k,i ,s′ k,1 )) / D(s k,i ,s′ k,1 ),…,(D(s k,i ,s′ k,i′+1 )-D(s k,i ,s′ k,i′ )) / D(s k,i ,s′ k,i′ ),…,(D(s k,i ,s′ k,n-1 )-D(s k,i ,s′ k,n-2 )) / D(s k,i ,s′ k,n-2 )(i≠i′); the average value of the branch node s k,i Upper distance D u (s k,i ) and lower distance D l (s k,i ) According to s k,i The average distance D a (s k,i ) and distance change rate CRd (s k,i ) is calculated to obtain:

[0027] D u (s k,i )=D a (s k,i )(1+CR d (s k,i )) (7)

[0028] D l (s k,i )=D a (s k,i )(1-CR d (s k,i )) (8)

[0029] For dataset DS k Middle branch node s k,i1 and s k,i2 , the degree of association between them depends on the association level and relative distance between them:

[0030] Cr d (s k,i1 ,s k,i2 )=Cr l (s k,i1 ,s k,i2 )-D r (s k,i1 ,s k,i2 ) (9)

[0031] The relative distance between two branch nodes with adjacency set and subset association levels depends on the maximum and minimum distances of their intervals:

[0032]

[0033] Two entity nodes E i1 =[s 1,i1 ,…,s k,i1 ,…,s m,i1 ] T and E i2 =[s 1,i2 ,…,s k,i2 ,…,s m,i2 ] T The correlation between them is the sum of the correlation between the corresponding branch nodes:

[0034]

[0035] A single generalized representation of a branch node is determined by the decision impact of the data elements in it:

[0036]

[0037] Entity node E i The generalized representation of can be defined as

[0038] R g (E i )=[R g (s 1,i ),…,R g (s k,i ),…,R g (s m,i )] T ;

[0039] Two entity nodes E i1 and E i2 The synergy expression and synergy degree between them are defined as follows:

[0040]

[0041] C d (E i1 ,E i2 )=Cr d (E i1 ,E i2 )Co d (E i1 ,E i2 ) (14)

[0042] The upper cooperativity of an entity node will be used as the cooperativity threshold of this entity node:

[0043] C u (E i )=C a (E i )(1+CR c (E i )) (15)

[0044] Here, C a (E i ) and CR c (E i ) is the entity node E i The average synergy and synergy change rate;

[0045] Assume that entity subset ss h Middle entity node E h,1 ,…,E h,t ,…,E h,T The number of adjacent entity nodes is N(E h,1 ),…,N(E h,t ),…,N(E h,T), an entity node (such as E h,t ) is equal to the number of its adjacent entity nodes and the entity subset ss h The ratio of the sum of all adjacent entity nodes in:

[0046]

[0047] Corresponding entity subset ss h The target entity node E h By ss h The weights and data representation of the entity nodes in the .

[0048] Beneficial effects of the present invention:

[0049] 1) This invention provides a generalized representation for multi-source data, combining the data and underlying knowledge, and forms a generalized array structure based on this representation. This generalized representation defines a unified data structure for accurate horizontal comparison of multi-source data from the perspective of its decision impact. This generalized representation involves partitioning data vectors into intervals, normalizing the representation, defining a decision impact index, and defining a generalization function.

[0050] 2) The present invention provides a fusion method based on the degree of synergy between entity nodes for the generalized representation of multi-source data. This collaborative fusion method realizes the collaborative division and fusion of entity nodes in the generalized array structure from the perspective of linking. It mainly includes the steps of defining the collaborative distance between branch nodes, forming a link network, constructing an array structure, calculating the correlation between entity nodes, calculating the degree of synergy between entity nodes, and dividing and fusion based on the degree of synergy. The multi-source data collaborative fusion method based on generalized representation can obtain accurate, reliable and generalized multi-source data fusion results for precise decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1A Schematic diagram of the generalized representation of data vectors with decision knowledge of different normal interval sizes. Figure 1B Schematic diagram of the generalized representation of data vectors with decision knowledge of uniform normal interval size.

[0052] Figure 2A is the distribution diagram of uncoordinated data elements, Figure 2B To compare the data element distribution diagram of the collaboration, Figure 2C is the distribution diagram of the change rate of uncoordinated data elements, Figure 2D The data element change rate distribution diagram for comparison and collaboration.

[0053] Figure 3A is the distance distribution between branch nodes with significant differences, Figure 3B The distance distribution between branch nodes with slight differences.

[0054] Figure 4A Schematic diagram of the formation of the link network. Figure 4B This is a schematic diagram of the construction of a generalized array structure. Figure 4C This is a simplified representation diagram of a branch node. Figure 4D A schematic diagram of the relationship between consistently adjacent entity nodes. Figure 4E Schematic diagram of the relationship between entity nodes of the consistent subset, Figure 4F Schematic diagram of the relationship between entity nodes in an inconsistent network. Figure 4G This is a schematic diagram of collaborative partitioning of entity nodes. Figure 4H Schematic diagram of collaborative fusion of entity nodes. DETAILED DESCRIPTION

[0055] The specific embodiments of the present invention are further described below with reference to the accompanying drawings:

[0056] Step 1) Generalized representation of multi-source data

[0057] Given n entities (nodes) E1,…,E i ,…,E n The generated m source data sets DS1,…,DS k ,…,DS m , dataset DS k Contains n branch nodes (or data units) from different entity nodes and has the expression DS k ={s k,1 ,…,s k,i ,…,s k,n}, entity node E i It consists of branch nodes distributed in different data sets and has the expression E i =[s 1,i ,…,s k,i ,…,s m,i ] T If the data unit s k,i The data elements in come from L different data vectors (Data vector, DV), then we can define s k,i for Typically, data elements in different data vectors have different value ranges and characterize different characteristic expressions of an entity. Here, each data vector interval is divided into the left abnormal range (LAR), the normal range (NR), and the right abnormal range (RAR). A normal range refers to a normal characteristic expression range of an entity. Figure 1AAs shown in Figure 2, different data vectors have different normal interval sizes. The left and right abnormal intervals usually refer to the lower and higher intervals of the normal interval.

[0058] In order to directly match the data elements in different data vectors, the inconsistent normal intervals are normalized to a unified normal interval (such as [0,1]). Given a normal interval Data elements Its normalized representation is defined as

[0059]

[0060] here, and Indicates the lower and upper limits of the normal range. Located in the abnormal area on the left Or the abnormal interval on the right Then its normalized expression is defined as or in, and Indicates the minimum element of the non-normal interval on the left and the maximum element of the non-normal interval on the right. Figure 1B As shown in Figure 2, after the data elements are normalized, different data vectors have a unified normal interval. However, the non-normal intervals on the left or right of different data vectors may still be non-uniform.

[0061] If the data elements in an entity node are in the normal interval, the entity is in a normal state. In this case, the decision impact index of these normal data elements is set to 0. If these elements are in the abnormal interval on the left or right, the entity is in an abnormal state. The decision impact of an abnormal state on an entity may be positive or negative. For data elements with positive or negative decision impact, their decision impact index is set to 1 or -1. Sometimes, an abnormal state may have a super positive or negative decision impact. For data elements with super positive or negative decision impact, their decision impact index is set to 2 or -2. So far, the generalized representation of a data element is defined as a combination of the normalized representation and the decision impact index of this data element. Usually, the generalization function depending on In the interval index The position in the interval pointed to. Generally speaking, the interval indices of the left abnormal interval, the normal interval, and the right abnormal interval can be set to -1, 0, and 1.

[0062]

[0063] Step 2) Collaborative fusion based on generalized representation

[0064] After generalization, the distances between data units in a dataset are used to classify these data units and are calculated based on the co-distances between corresponding data elements in the data units. and The distance D(s) between them k,i1 ,s k,i2 ) is the accumulation of collaborative distances between corresponding data elements.

[0065]

[0066] The collaborative distance between two data elements is related to the collaborative index of the data vectors containing the two data elements.

[0067]

[0068] like Figure 2A , Figure 2B As shown in , the degree of synergy between two data vectors depends on the distribution of data elements in them. Obviously, Figure 2B The two data vectors are compared Figure 2A The two data vectors in are more cooperative because Figure 2B The two data vectors have similar data element distribution trends. Figure 2C , Figure 2D Display and Figure 2A , Figure 2B The corresponding data element change rate distribution also reflects the same collaborative situation. Given a generalized data vector and Ascending form of and Data vector and The rate of change of the data elements in Data vector and The degree of coordination between them is defined as follows.

[0069]

[0070] The synergy index of a data vector is the sum of the synergies between this data vector and other data vectors in the same dataset.

[0071]

[0072] Therefore, the synergy index of a data vector reflects the overall synergy between this data vector and other data vectors in the same dataset. The larger the synergy index of a data vector, the shorter the synergy distance between the data elements in this data vector and the closer the relationship.

[0073] The partitioning of a data set is based on the closest link relationship between the branch nodes (or data units). The closest link branch of a branch node is defined by the distance between this branch node and other branch nodes in the same data set and the distance threshold of this branch node. The common distance threshold of a branch node is the average value of this branch node and other branch nodes. Here, the definition of the distance threshold of a branch node includes factors such as average distance and distance change rate. Given a data set DS k Middle branch node s k,i n-1 ascending distances D(s between the node and other branch nodes) k,i ,s′ k,1 ),…,D(s k,i ,s′ k,i′ ),…,D(s k,i ,s′ k,n-1 ), s k,i The distance change rate R d (s k,i ) is defined as the adjacent distance difference (D(s k,i ,s′ k,2 )-D(s k,i ,s′ k,1 )) / D(s k,i ,s′ k,1 ),…,(D(s k,i ,s′ k,i′+1 )-D(s k,i ,s′ k,i′ )) / D(s k,i ,s′ k,i′ ),…,(D(s k,i ,s′ k,n-1 )-D(s k,i ,s′ k,n-2 )) / D(s k,i ,s′ k,n-2 )(i≠i′). Figure 3A , Figure 3B As shown, the branch node s k,i Upper distance D u (s k,i ) and lower distance D l (s k,i ) According to s k,i The average distance D a (s k,i ) and distance change rate CR d (sk,i ) calculated.

[0074] D u (s k,i )=D a (s k,i )(1+CR d (s k,i )) (7)

[0075] D l (s k,i )=D a (s k,i )(1-CR d (s k,i )) (8)

[0076] exist Figure 3A In , there are obvious differences between the upper, average, and lower distances because the related branch nodes have a larger distance change rate. However, because the related branch nodes have a smaller distance change rate, Figure 3B The differences between the upper middle, average and lower distances are relatively small. Figure 3A , Figure 3B The upper distance threshold for a branch node shown includes some edge distances. Therefore, the upper distance is selected as the distance threshold for a branch node to include some edge branch nodes in its nearest connected branch nodes. The additional information about a branch node's edge branch nodes can enrich the branch node's nearest connected information and benefit relevant decisions based on the fusion results.

[0077] A branch node and its nearest linked branch node form an adjacent branch node set. All adjacent branch node sets in a data set form a link network with several link components. Each link component represents a branch structure subset. Figure 4A As shown in , the link network of all data sets forms an array structure. For an entity node in the array structure, its constituent branch nodes distributed in different link networks are located in the consistent or inconsistent branch node subsets. Figure 4B As shown, the entity node E i2 The constituent branch nodes of E are distributed in the same branch node subset, but the entity node E i1The constituent branch nodes are distributed in inconsistent branch node subsets. Obviously, the different distribution of branch nodes in the entity node will affect the degree of coordination between the entity nodes. In a link network, the distribution of branch nodes depends on the association relationship between them. Here, the association relationship between branch nodes is divided into three levels: adjacent set, subset, and network. The adjacent set level indicates that there are the closest links between related branch nodes, the subset level indicates that related branch nodes are in the same subset or link component, and the network level indicates that related branch nodes are only in the same link network. For the dataset DS k Middle branch node s k,i1 and s k,i2 ,The degree of association between them depends on the level of association and the relative distance between them.

[0078] Cr d (s k,i1 ,s k,i2 )=Cr l (s k,i1 ,s k,i2 )-D r (s k,i1 ,s k,i2 ) (9)

[0079] Here, the association level indices of the neighbor set, subset, and network are set to 2, 1, and 0. The relative distance between two branch nodes with the network association level is set to 0, while the relative distance between two branch nodes with the neighbor set and subset association levels depends on the maximum and minimum distances of their intervals.

[0080]

[0081] Two entity nodes E i1 =[s 1,i1 ,…,s k,i1 ,…,s m,i1 ] T and E i2 =[s 1,i2 ,…,s k,i2 ,…,s m,i2 ] T The correlation degree between them is the sum of the correlation degrees between the corresponding branch nodes.

[0082]

[0083] The degree of synergy between entity nodes is not only related to the degree of association between them, but also to the degree of synergy between their generalized representations. The generalized representation of an entity node is composed of the single generalized representation of its branch nodes. Figure 4CAs shown in , a single generalized representation of a branch node is calculated based on the normalized representation of the data elements and the decision impact. In other words, a single generalized representation of a branch node is determined by the decision impact of the data elements.

[0084]

[0085] Therefore, the entity node E i The generalized representation of can be defined as R g (E i )=[R g (s 1,i ),…,R g (s k,i ),…,R g (s m,i )] T Two entity nodes E i1 and E i2 The synergy expression degree and synergy degree between them are defined as follows.

[0086]

[0087] C d (E i1 ,E i2 )=Cr d (E i1 ,E i2 )Co d (E i1 ,E i2 ) (14)

[0088] Similar to the formation of a branch node link network of a dataset, the entity node link network (actually an array network) is formed based on the cooperativity between all entity nodes. Therefore, the upper cooperativity of an entity node will serve as the cooperativity threshold of this entity node.

[0089] C u (E i )=C a (E i )(1+CR c (E i )) (15)

[0090] Here, C a (E i ) and CR c (E i ) is the entity node E i The average synergy and synergy change rate of the nodes are calculated according to the average distance and distance change rate of a branch node. Figure 4D , Figure 4E , Figure 4F As shown in , in the array network formed by the adjacency sets of all entity nodes, there are three levels of relationships between entity nodes: consistent adjacency, consistent subset, and inconsistent network. Given two entity nodes with consistent adjacency or consistent subset relationships, all corresponding branch nodes between them have direct adjacent (or nearest) links or are in the same subset. If some corresponding branch nodes between two entity nodes do not have direct links or are in the same subset, then the two entity nodes are in an inconsistent network relationship. Figure 4G As shown, all entity nodes form different entity subsets or components.

[0091] Figure 4H The branch nodes that show an entity subset distributed in each data set will be uniformly fused into a target branch node. In an entity subset, the number of adjacent entity nodes of an entity node determines the fusion weight of this entity node. Assume that the entity subset ss h Middle entity node E h,1 ,…,E h,t ,…,E h,T The number of adjacent entity nodes is N(E h,1 ),…,N(E h,t ),…,N(E h,T ), an entity node (such as E h,t ) is equal to the number of its adjacent entity nodes and the entity subset ss h The ratio of the sum of the number of adjacent entity nodes in .

[0092]

[0093] Corresponding entity subset ss h The target entity node E h By ss h The weights and data representation of the entity nodes in the .

[0094] Example:

[0095] 1) Input a given multi-source dataset (e.g., “n patient entities E1,…,E i ,…,E n The generated multi-source medical history, physical examination, lung examination, auxiliary examination, blood test, stool test, urine test, PCT test, emergency biochemical test data set") DS1,…,DS k ,…,DS mAll data elements in the dataset. First, we divide the intervals of each data vector in each dataset (abnormal interval on the left, normal interval, and abnormal interval on the right) and define the interval index and decision influence index of each interval. On this basis, we normalize all data elements according to formula (1) and define the generalized representation of normalized data elements according to formula (2).

[0096] 2) After the data elements are generalized, the degree of synergy between the generalized data vectors composed of the generalized data elements is calculated according to formula (5), and the synergy index of each generalized data vector is calculated according to formula (6). On this basis, the synergy distance between the data elements is calculated according to formula (4), and the distance between branch nodes (or data units) is further defined according to formula (3). According to the distance between each branch node and other branch nodes in the same data set, the upper distance calculated according to formula (7) is used as the distance threshold of the branch node, and the adjacent branch node set of the branch node is further obtained. The adjacent branch node set of all branch nodes forms a link network with several link components, and several link networks constitute a generalized array structure. In the array structure, the correlation between branch nodes in the same link network is calculated according to formula (10) and formula (9), and the correlation between entity nodes is further calculated according to formula (11). Based on the generalized representation of the data elements and the decision influence index, a single generalized representation of each branch node is calculated according to formula (12). Based on the single generalized representation of the branch node, the collaborative representation degree between entity nodes is calculated according to formula (13), and the collaborative degree between entity nodes is further calculated according to formula (14). Based on the collaborative degree between entity nodes, the collaborative degree threshold of each entity node is calculated according to formula (15). On this basis, all entity nodes are divided into several entity subsets according to the collaborative degree between entity nodes and the collaborative degree threshold of each entity node, and the fusion weight of each entity node is calculated according to formula (16), so that the entity nodes in each entity subset are collaboratively fused into a single target entity node.

Claims

1. A generalized representation and collaborative fusion method for multi-source data, characterized by The method comprises the following steps: Step 1) Generalized representation of multi-source data; Step 2) Collaborative fusion based on generalized representation; The step 1) is specifically as follows: Multi-source datasets DS1,…,DS k ,…,DS m There are n patient entities E1,…,E i ,…,E n The generated multi-source medical history, physical examination, lung examination, routine blood test, routine stool examination, routine urine test, PCT test, and emergency biochemical test data; Therefore, the dataset DS contains n branch nodes from different entities k With expression DS k ={s k,1 ,…,s k,i ,…,s k,n +, entity node E composed of branch nodes distributed in different data sets i With expression E i =,s 1,i ,…,s k,i ,…,s m,i - T ; If the data unit s k,i The data elements in come from L different data vectors, then we can define s k,i for Given a normal range Data elements Its normalized representation is defined as here, and Indicates the lower and upper limits of the normal range; if Located in the abnormal area on the left Or the abnormal interval on the right Then its normalized expression is defined as or in, and Represents the minimum element of the non-normal interval on the left and the maximum element of the non-normal interval on the right; Generalization function depending on In the interval index Position in the indicated interval The step 2) is specifically as follows: Given a data unit and The distance D(s) between them k,i1 ,s k,i2 ) is the cumulative synergy distance between corresponding data elements therein; The synergy distance between two data elements is related to the synergy index of the data vectors where the two data elements are located; Given a generalized data vector and Ascending form of and Data vector and The rate of change of the data elements in Data vector and The degree of coordination is defined as follows: The synergy index of a data vector is the sum of the synergy between this data vector and other data vectors in the same dataset: Given a dataset DS k Middle branch node s k,i n-1 ascending distances D(s between the node and other branch nodes) k,i ,s′ k,1 ),…,D(s k,i ,s′ k,i′ ),…,D(s k,i ,s′ k,n-1 ), s k,i The distance change rate R d (s k,i ) is defined as the adjacent distance difference (D(s k,i ,s′ k,2 )-D(s k,i ,s′ k,1 )) / D(s k,i ,s′ k,1 ),…,(D(s k,i ,s′ k,i′+1 )-D(s k,i ,s′ k,i′ )) / D(s k,i ,s′ k,i′ ),…,(D(s k,i ,s′ k,n-1 )-D(s k,i ,s′ k,n-2 )) / D(s k,i ,s′ k,n-2 )(i≠i′); the average value of the branch node s k,i Upper distance D u (s k,i ) and lower distance D l (s k,i ) According to s k,i The average distance D a (s k,i ) and distance change rate CR d (s k,i ) is calculated to obtain: D u (s k,i )=D a (s k,i )(1+CR d (s k,i )) (7) D l (s k,i )=D a (s k,i )(1-CR d (s k,i )) (8) For dataset DS k Middle branch node s k,i1 and s k,i2 , the degree of association between them depends on the association level and relative distance between them: Cr d (s k,i1 ,s k,i2 )=Cr l (s k,i1 ,s k,i2 )-D r (s k,i1 ,s k,i2 )(9) The relative distance between two branch nodes with adjacency set and subset association levels depends on the maximum and minimum distances of their intervals: Two entity nodes E i1 =,s 1,i1 ,…,s k,i1 ,…,s m,i1 - T and E i2 =,s 1,i2 ,…,s k,i2 ,…,s m,i2 - T The correlation between them is the sum of the correlation between the corresponding branch nodes: A single generalized representation of a branch node is determined by the decision impact of the data elements in it: Entity node E i The generalized representation of can be defined as R g (E i )=,R g (s 1,i ),…,R g (s k,i ),…,R g (s m,i )- T ; Two entity nodes E i1 and E i2 The synergy expression and synergy degree between them are defined as follows: C d (HAVE BEEN i1 ,HAVE BEEN i2 )=Cr d (HAVE BEEN i1 ,HAVE BEEN i2 )Co d (HAVE BEEN i1 ,HAVE BEEN i2 )(14) The upper cooperativity of an entity node will be used as the cooperativity threshold of this entity node: C u (AND i )=C a (AND i )(1+CR c (AND i ))(15) Here, C a (E i ) and CR c (E i ) is the entity node E i The average synergy and synergy change rate; Assume that entity subset ss h Middle entity node E h,1 ,…,E h,t ,…,E h,T The number of adjacent entity nodes is N(E h,1 ),…,N(E h,t ),…,N(E h,T ), an entity node E h,t The fusion weight is equal to the number of its adjacent entity nodes and the entity subset ss h The ratio of the sum of all adjacent entity nodes in: Corresponding entity subset ss h The target entity node E h By ss h The weights and data representation of the entity nodes in the .

Citation Information

Patent Citations

  • Rectal cancer pathology image classification method based on multi-channel collaborative capsule network

    CN111191660A

  • Unified dot matrix plane representation method for multi-source medical examination data

    CN113192157A