Intelligent statistical method and system for innovation and entrepreneurship data
By constructing a temporal semantic hypergraph and utilizing semantic dynamics computation and entropy weight measurement modules, the topological structure is dynamically reconstructed, solving the problem of distorted statistical results of innovation and entrepreneurship data in existing technologies, and realizing accurate identification and visualization support for innovation activities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN SANY IND VOCATIONAL & TECH COLLEGE
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-10
AI Technical Summary
Existing statistical methods for innovation and entrepreneurship data cannot effectively distinguish between routine operational noise and innovative development, and lack the dynamic evolution logic of innovation elements across time cycles, resulting in distorted statistical results and a lack of interpretability and forward-looking support.
By constructing a temporal semantic hypergraph, utilizing semantic dynamics computation and entropy weight measurement modules, the topological structure is dynamically reconstructed, the curvature of innovative trajectories is identified, and the value of events is quantified through interactive entropy weights. High-frequency noise nodes are stripped away, and high-order hyperedges are encapsulated to achieve nonlinear gain weights.
It improves the purity of innovation index calculation, truly reflects the innovation activity trend, enhances the retrieval efficiency of large-scale time series data, and provides support for visualized innovation evolution paths.
Smart Images

Figure CN121833801A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of big data processing and artificial intelligence technology, specifically to an intelligent statistical method and system for innovation and entrepreneurship data. Background Technology
[0002] With the deepening implementation of the digital economy and innovation-driven development strategy, accurate statistics and assessment of regional or industry innovation and entrepreneurship vitality have become crucial for government decision-making and capital market investment. Current innovation and entrepreneurship data statistics primarily rely on the collection and analysis of large-scale, multi-source, heterogeneous data, covering multiple dimensions such as business registration information, patent application documents, investment and financing news reports, and recruitment data. Through cleaning, integrating, and calculating indicators from this massive amount of data, the aim is to reconstruct the operating conditions and development trends of innovation entities.
[0003] However, existing data statistics and analysis technologies typically employ structured storage based on relational databases or simple knowledge graph construction methods. This paradigm reveals significant limitations when dealing with high-dimensional, dynamic, and semantically complex innovative data. Existing statistical methods generally use linear summation logic, generating statistical reports through simple keyword matching or tag counting. While this approach can quickly process structured fields, it often overlooks the semantic value differences and temporal evolution characteristics behind the data. For example, in traditional statistical models, routine administrative changes for a company (such as license renewal) and substantial technological transformations (such as changes in business scope accompanied by new patent applications) are often counted with equal weight. This results in statistical results mixed with a large amount of low-value routine operational noise, failing to effectively distinguish between survival-oriented operations and innovative development, thus distorting statistical indices.
[0004] Furthermore, existing technologies struggle to effectively capture the dynamic evolutionary logic of innovation elements across time cycles. Innovation is often a complex, non-linear process involving orthogonal or co-evolution across multiple dimensions, including technological breakthroughs, capital injection, and talent mobility. Existing time-series data analysis primarily relies on static comparisons of discrete time slices, lacking continuous modeling of drift trajectories in the semantic vector space of entities. This means the system cannot calculate whether the evolutionary direction of the innovation subject has undergone a drastic shift (i.e., the curvature of the innovation trajectory), nor can it identify seemingly isolated event sequences that are actually deeply logically connected. Due to the lack of this deep correlation capability based on semantic dynamics and topological structure, existing statistical systems can only provide lagging results and cannot reconstruct a clear innovation evolution chain, resulting in data output lacking interpretability and forward-looking support. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides an intelligent statistical method and system for innovation and entrepreneurship data. It solves the problems that existing innovation and entrepreneurship statistical methods, which rely on linear accumulation, cannot effectively eliminate conventional business noise, are difficult to capture complex cross-dimensional innovation evolution chains, and result in distorted statistical results and a lack of logical support.
[0006] To achieve the above objectives, the present invention provides an intelligent statistical method for innovation and entrepreneurship data.
[0007] The method primarily achieves accurate quantification of innovation momentum by constructing and dynamically reconstructing a temporal semantic hypergraph. First, it utilizes a multimodal data access module to acquire multi-source, heterogeneous innovation and entrepreneurship data, performing data cleaning and entity alignment processing. Unlike traditional relational database storage, this invention maps the cleaned discrete events as hyperedges containing timestamps, constructing an initial-state temporal semantic hypergraph. In this hypergraph, the node set represents innovation subjects and elements, the hyperedge set represents event associations, and the storage index is divided according to the time dimension, thus preserving the temporal evolution characteristics of the data.
[0008] After the basic graph is constructed, this invention introduces a semantic dynamics computation mechanism. The semantic evolution computation module is used to vectorize entities in the temporal semantic hypergraph. By calculating the vector changes of entities in adjacent time windows, a semantic drift vector representing the entity's development direction is obtained. Furthermore, the innovation trajectory curvature of nodes is determined based on the rate of change of the semantic drift vector's direction. This curvature index can identify whether an innovation entity has experienced a sudden change in its technological path or a transformation in its business model, distinguishing it from linear, stable development.
[0009] To assess the information value of innovative events, this invention utilizes an entropy weighting module to calculate the interaction entropy weight. This step statistically analyzes the co-occurrence probability of node combinations within a hyperedge within a historical time window and introduces a time decay factor. Historically rare but recently sudden combinations are assigned higher interaction entropy weights, thereby quantifying the information scarcity of innovative events; conversely, high-frequency, common combinations are assigned lower weights.
[0010] Based on the above calculation results, this invention utilizes a topology dynamic reconstruction module to optimize the graph structure at the physical level, which is the core processing step of this invention. This step performs hyperedge splitting and hyperedge fusion operations based on the interaction entropy gradient. Specifically, for hyperedges with low interaction entropy weights, the system identifies high-frequency noise nodes (such as general equipment, routine administrative changes, etc.), removes them from the connection relationships, and retains only the core nodes to construct dimensionality-reduced sub-hyperedges, achieving automatic noise reduction. For hyperedges with high interaction entropy weights in continuous time series, the system calculates the inner product of the semantic drift vectors of the core entities. When the inner product approaches zero, thus satisfying the orthogonal complement feature, it indicates that the entities have undergone complementary evolution in different dimensions (such as from the technology research and development dimension to the capital operation dimension). At this time, the system generates a virtual higher-order hyperedge, encapsulates the original discrete hyperedges into a whole, and assigns nonlinear gain weights.
[0011] Finally, the statistical analysis service module is used to respond to statistical query requests. Effective hyperedges for the target region or industry are retrieved from the reconstructed temporal semantic hypergraph. The final innovation momentum index is obtained through a composite calculation of interaction entropy weights, nonlinear gain weights, and innovation trajectory curvature. This process not only outputs numerical results but also generates an evidence subgraph containing virtual higher-order hyperedges, visually demonstrating the source path of innovation momentum.
[0012] A second aspect of this invention provides an intelligent statistical system for innovation and entrepreneurship data.
[0013] The system comprises a multimodal data access module, a semantic evolution calculation module, an entropy weight measurement module, a topology dynamic reconstruction module, and a statistical analysis service module. The multimodal data access module is used to standardize data access and construct an initial graph; the semantic evolution calculation module is responsible for performing dynamic feature extraction based on vector space, outputting semantic drift vectors and innovation trajectory curvature; the entropy weight measurement module quantifies the scarcity value of events based on information theory principles, outputting interactive entropy weights; the topology dynamic reconstruction module, as the core control unit, dynamically adjusts the topological structure of the graph based on the input curvature and weight indicators, performing hyperedge splitting denoising and fusion gain; and the statistical analysis service module performs nonlinear weighted index aggregation calculations based on the reconstructed graph structure and provides visualized evidence tracing services.
[0014] This invention provides an intelligent statistical method and system for innovation and entrepreneurship data. It has the following beneficial effects:
[0015] 1. This invention calculates the interaction entropy weight of hyperedges and uses a topology dynamic reconstruction module to perform a global degree-based splitting operation on low-entropy hyperedges. This automatically identifies and removes high-frequency background nodes generated by innovation entities in their routine business activities. This technical approach severs the strong correlation of non-substantive innovation elements at the physical connection level, solves the technical problem that traditional statistical methods rely solely on quantity accumulation and cannot distinguish pseudo-innovation data, and improves the purity of innovation index calculation.
[0016] 2. This invention utilizes a semantic evolution computation module to measure the curvature of the innovation trajectory of nodes and identifies the orthogonal complementary features between continuous high-entropy hyperedges. Discrete events in different semantic dimensions are encapsulated as virtual high-order hyperedges. This mechanism enables the statistical system to no longer be limited to the discrete evaluation of single-point events, but to logically connect the complete path of enterprise transformation and amplify the statistical contribution of such substantive evolutionary behavior through nonlinear gain weights, thereby more realistically reflecting the innovation activity of a region or industry.
[0017] 3. This invention reconstructs the topology of a temporal semantic hypergraph based on value density. By encapsulating high-value sequences and reducing the dimensionality of low-value connections, it reduces redundant index paths and improves the retrieval efficiency of large-scale time-series data. At the same time, the statistical analysis service module can generate evidence subgraphs based on the reconstructed virtual higher-order hyperedges, restoring abstract statistical values to visualize the evolution path of key events, providing logically supported traceability evidence for government or investment institutions' decision-making. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the method flow of the present invention;
[0019] Figure 2 This is a schematic diagram of the system modules of the present invention. Detailed Implementation
[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] See attached document Figure 1This invention provides an intelligent statistical method for innovation and entrepreneurship data, which can be executed by an intelligent statistical system for innovation and entrepreneurship data. The system may include: a multimodal data access module, a semantic evolution calculation module, an entropy weight measurement module, a topology dynamic reconstruction module, and a statistical analysis service module. The multimodal data access module is used to collect and process multi-source heterogeneous data to construct an initial temporal semantic hypergraph. The semantic evolution calculation module is used to calculate the semantic drift vector of entity nodes in the time dimension and the innovation trajectory curvature. The entropy weight measurement module is used to calculate the interaction entropy weight of node combinations within hyperedges. The topology dynamic reconstruction module is used to perform topological splitting or fusion operations on the hypergraph based on the interaction entropy weights and innovation trajectory curvature. The statistical analysis service module is used to perform weighted statistics based on the reconstructed hypergraph and output the results.
[0022] The method includes the following steps:
[0023] Step S100: Utilize the multimodal data access module to acquire multi-source heterogeneous innovation and entrepreneurship data. Clean and align the data to entities, and map the processed discrete events as hyperedges containing timestamps, constructing an initial temporal semantic hypergraph. The temporal semantic hypergraph includes a node set, a hyperedge set, and a time dimension index.
[0024] Step S200: The semantic evolution calculation module is used to vectorize the entities in the node set, and the semantic drift vector of the node is calculated based on the vector change of adjacent time windows. Then, the innovation trajectory curvature of the node is determined according to the direction change rate of the drift vector.
[0025] Step S300: The entropy weight measurement module is used to count the historical co-occurrence probability of node combinations within the hyperedge, and the interaction entropy weight of each hyperedge is calculated in combination with the time decay factor. The interaction entropy weight represents the information scarcity of the innovation event.
[0026] Step S400: Perform hypergraph structure optimization using the topology dynamic reconstruction module. The structure optimization includes hyperedge splitting operation and hyperedge fusion operation based on interaction entropy gradient. The hyperedge splitting operation decomposes hyperedges below a first preset threshold into a core set and a background set; the hyperedge fusion operation encapsulates a continuous hyperedge sequence above a second preset threshold and possessing orthogonal complement characteristics into a virtual higher-order hyperedge.
[0027] In step S500, the statistical analysis service module responds to the statistical query request, retrieves the effective hyperedges of the target region or target industry in the temporal semantic hypergraph reconstructed in step S400, and calculates the innovation momentum index based on the interaction entropy weight and innovation trajectory curvature.
[0028] The following section details the specific implementation of the temporal semantic hypergraph construction in step S100. The multimodal data access module first accesses multimodal raw data, including enterprise registration information, intellectual property patent texts, investment and financing news reports, and policy documents. The multimodal data access module performs preprocessing on the raw data, including removing format noise, unifying character encoding, and extracting entities and concepts using named entity recognition technology. Entities include enterprise names, legal representatives, investment institution names, and names of universities and research institutes; concepts include technical keywords, industry classification tags, and policy thesaurus terms. After entity extraction, the system establishes a globally unique entity identifier index, aligning and merging descriptions referring to the same object from different data sources.
[0029] After data preprocessing, the system constructs a temporal semantic hypergraph model. Unlike traditional binary graph structures, the basic storage unit of a temporal semantic hypergraph is a hyperedge. The system maps each independent innovation event to a hyperedge. A hyperedge connects two or more heterogeneous nodes, representing the concurrent relationships between multiple entities and elements within that innovation event. For example, the hyperedge corresponding to a financing event simultaneously connects the financing company node, multiple investment institution nodes, financing round label nodes, and related technology field nodes.
[0030] The system assigns a generation timestamp to each hyperedge, corresponding to the actual time the innovation event occurred. To support temporal evolution analysis, the system establishes an inverted index structure based on time windows at the storage level. The system divides the continuous timeline into multiple discrete time windows, each corresponding to a data slice. Hyperedges are stored in the corresponding time window slice according to their generation timestamps. For long-term events or states that persist across multiple time windows, the system generates snapshot copies of the hyperedge within the corresponding series of time windows and establishes temporal pointer connections between the copies. In this way, the temporal semantic hypergraph fully preserves the dynamic topological structure of each element in the innovation and entrepreneurship ecosystem over time, providing a data foundation with temporal attributes for subsequent semantic drift calculations and entropy measurement.
[0031] In this embodiment, the construction of the initial state temporal semantic hypergraph in step S100 specifically includes three sub-processes: data cleaning and entity alignment, hyperedge mapping construction, and temporal inverted index storage.
[0032] First, regarding the data cleaning and entity alignment process, the multimodal data access module performs format standardization processing on the incoming raw data stream. Since innovation and entrepreneurship data comes from a wide range of sources, encompassing unstructured text (such as news reports and patent specifications) and semi-structured tables (such as business registration data), the system first removes HTML tags, garbled characters, and non-printable characters from the data using regular expressions and character encoding conversion technology. Subsequently, the system performs Named Entity Recognition (NER) operations, using a pre-built domain dictionary and sequence labeling model to identify entities with independent semantics from the text. Entities are divided into different type sets, denoted as... This includes, but is not limited to, sets of corporate entities, sets of person entities (including legal persons, inventors, and investors), sets of technical keywords, sets of geographical locations, and sets of capital elements. To resolve ambiguity in multi-source data, the system maintains a global mapping table, assigning a globally unique identifier (GUID) to each identified independent entity. If Company A and A Technology Development Co., Ltd. appear in different data sources, the system determines that they are the same entity by matching string similarity and associating them with business registration numbers, and maps them to the same GUID, thus completing entity alignment.
[0033] Secondly, the hyperedge mapping construction process is executed. This is the core data organization method of this invention. Unlike the binary relationship between nodes and edges used in traditional graph databases, the hypergraph structure defined in this invention... From the set of nodes and superedge set Composition, that is Each of these hyperedges It is a non-empty subset of nodes, denoted as This system is used to fully describe a complex innovation event. It iterates through cleaned data entries, encapsulating all elements involved in a single event into a hyperedge. For example, for a data record of a technology company and a university jointly applying for a patent in the field of artificial intelligence, the system constructs a hyperedge that simultaneously connects the technology company ID, university ID, inventor ID, artificial intelligence technology tag ID, and patent classification number ID. This mapping method preserves the high-dimensional correlation characteristics of multi-party collaboration in innovation activities, avoiding information loss and semantic fragmentation caused by decomposing complex events into multiple binary edges.
[0034] Finally, the time-series inverted index storage procedure is executed. To support subsequent calculations of the evolution of innovation trends, the system manages the hypergraph by slicing it along a time dimension. The system sets a fixed time window length (e.g., in months or quarters) and discretizes the time axis into sequences. Each hyperedge Based on the time of their corresponding event occurrence, each entry is categorized into a specific time window slice. At the physical storage level, the system constructs a versioned inverted index based on entities. This index structure uses the entity's GUID as the key and the list of superedges containing that entity as the value. Unlike inverted indexes used in conventional document retrieval, the inverted index entries in this embodiment store triples. ,in This is the identifier for the superedge. For time window identification, The initial weights are used. When an entity participates in different innovation events at different times, these events are recorded in the inverted index in an append-only manner, strictly ordered by timestamp.
[0035] Furthermore, to accelerate the querying of graph structures within a specific time period, the system maintains a local adjacency list within each time window slice. This list records the set of all active hyperedges within that time window. For persistent states spanning long periods (such as the validity period of a high-tech enterprise qualification), the system employs a state continuation mechanism, inserting a hyperedge reference corresponding to that state into the index of each consecutive time window in which the qualification is valid, or marking the effective lifespan of the hyperedge in the metadata. Through the above storage strategy, the system achieves structured storage of massive dynamic innovation and entrepreneurship data, ensuring the integrity of data at a single moment and providing efficient time-slice retrieval capabilities for calculating vector drift and interaction entropy in subsequent steps.
[0036] In this embodiment, step S200, which involves semantic drift and curvature calculation based on vector space, specifically includes three closely related processing stages: entity temporal embedding representation, semantic drift vector construction, and innovative trajectory curvature measurement.
[0037] First, the entity temporal embedding representation stage is performed. To quantify the semantic information in unstructured data, the system utilizes a semantic evolution computation module to map high-dimensional sparse symbolic data into low-dimensional dense real-number vectors. For any entity node... The system first identifies it within a specific time window. The set of all superedges that participate in the inner circle is denoted as . For each hyperedge in the set containing textual description information (e.g., patent abstracts, descriptions of business scope changes, and investment and financing news articles), the system calls a pre-trained deep language model (e.g., a model based on the Transformer architecture) to encode and extract the feature vector sequence within that time window. To obtain the comprehensive semantic state of the entity at the current moment, the system employs an attention-weighted aggregation mechanism... The feature vectors of all hyperedges in the middle are fused to generate the entity. In the time window The state vector, denoted as The state vector resides in a multi-dimensional semantic space, where different dimensions represent different technological fields, business models, or industry attributes. If no new events occur within a certain time window, the system employs a historical vector smoothing strategy with a time decay coefficient to maintain state continuity.
[0038] Secondly, the semantic drift vector construction phase is performed. This phase aims to capture the dynamic changes of entities over time. The system calculates the entity's position within the current time window. The state vector and the previous time window The difference between the state vectors. This difference is defined as the semantic drift vector, denoted as . The calculation logic is as follows:
[0039] ;
[0040] in, and All are normalized state vectors. Semantic drift vector. Length of the module This represents the degree of drastic change in innovation (e.g., whether it's a minor product iteration or a major business transformation), while The direction represents the specific evolutionary path of innovation (for example, shifting from web development to artificial intelligence).
[0041] Finally, the innovation trajectory curvature measurement stage is performed. This is a crucial step in distinguishing path-dependent development from breakthrough transformation. The system does not consider the drift at a single moment, but rather examines the geometric relationship between the drift vectors of two consecutive time steps. (Innovation trajectory curvature) Defined as the current drift vector The drift vector from the previous time step The complement of the cosine distance between the two points. Its specific calculation is expressed as follows:
[0042] ;
[0043] in, The Euclidean norm (magnitude) of a vector. For a very small positive number (e.g.) This is used to prevent calculation errors when the denominator is zero.
[0044] Geometrically speaking, when When the value approaches 0, it indicates and The directions are basically consistent; entities move along straight lines in the semantic space, which corresponds to cumulative innovative behaviors such as technological advancement or business expansion. When the value approaches 1 (i.e., the two vectors are orthogonal) or even larger (i.e., the two vectors are in opposite directions), it indicates that the evolutionary direction of the entity has deviated. This corresponds to abrupt innovative behaviors such as cross-border transformation, opening up new tracks, or restructuring of business models. The system will calculate the... The values are stored in the node's dynamic attributes and serve as dynamic factors that influence the statistical weights in subsequent steps.
[0045] In this embodiment, the information theory-based hyperedge interaction entropy weight measurement in step S300 is specifically executed through the entropy weight measurement module. Its core lies in utilizing the concept of self-information in information theory, combined with the time dimension, to quantify the scarcity of the innovative events represented by each hyperedge. This process mainly includes three technical steps: node combination co-occurrence probability estimation, introduction of a time decay factor, and interaction entropy weight calculation.
[0046] First, the co-occurrence probability of node combinations is estimated. This step aims to determine the prevalence of a particular combination of innovative elements from a statistical perspective. For any hyperedge to be calculated in the temporal semantic hypergraph... The system first extracts the set of nodes contained in the hyperedge. The entropy weighting module utilizes the time-series inverted index constructed in step S100, within the historical time window (i.e., the generation time). Initiate a combined query within the range of all previous time slices. The system counts the frequency of this specific node combination appearing together in historical data, and records it as... At the same time, obtain the total number of events within the historical time window, denoted as . .
[0047] Based on the above statistical values, the system calculates the prior co-occurrence probability of the hyperedge at the current time. To avoid statistical bias caused by sample sparsity (i.e., some extremely rare combinations have a frequency of 0, resulting in a probability of 0), the system employs Laplace smoothing or a similar smoothing technique. The calculation logic is as follows:
[0048] ;
[0049] in, and This is a preset smoothing parameter. This probability value... This directly reflects the degree of homogenization of innovation models: if A higher value indicates that this combination (e.g., e-commerce + website development) has occurred frequently historically and is a typical business model; if... The extremely low value indicates that this combination (such as brain-computer interface + art therapy) is historically unprecedented and possesses potential for exploration and scarcity.
[0050] Secondly, a time decay factor is introduced. Because the value of innovation and entrepreneurship data exhibits a non-linear decay characteristic over time (i.e., recent innovation events have greater reference value for current trends than historical events in the distant past), the system defines a time decay coefficient based on an exponential function. Let... This is the current point in time for conducting statistical analysis. For super-edge The point in time of generation, For time span. Time decay coefficient. Set as ,in ( This is a hyperparameter for controlling the decay rate. This mechanism ensures that even with the same innovative combinations, the closer they occur, the higher their retention weight.
[0051] Finally, the interaction entropy weights are calculated. The system combines the co-occurrence probabilities mentioned above with the time decay coefficient, and defines the hyperedge according to Shannon's information theory. Inter-entropy weight This weight represents the amount of information or surprise contained in the event. The specific calculation expression is:
[0052] ;
[0053] Among them, the logarithmic function This implements the function of reversing probability into information content: the smaller the probability, the larger its negative logarithm, giving scarce events a higher basic weight. The multiplication term incorporates the time dimension into the weighting system. After calculation, the system will... Write it as a dynamic attribute value to the hyperedge In the metadata fields, This is an exponential decay function of the influence factor. It is a decay constant that determines how uncertainty decays over time.
[0054] Through the above steps, the system completes the value stratification of massive hyperedges: a large amount of repetitive, follow-the-leader low-value data is assigned extremely low entropy values, while events with unique element combinations and recent occurrences are assigned high entropy values. This quantitative result does not rely on manual labeling or black-box model predictions, but is calculated entirely based on the data's own distribution characteristics, providing a unique quantitative decision basis for the subsequent dynamic reconstruction of the hypergraph topology in step S400 (i.e., determining which edges should be retained and which should be merged).
[0055] In this embodiment, the entropy gradient-based hypergraph topology dynamic reconstruction in step S400 is performed by the topology dynamic reconstruction module. This step differs from traditional database maintenance (which typically only involves adding, deleting, and modifying data). Instead, it is a process of physically or logically reshaping the storage structure based on data value (i.e., the interaction entropy calculated in step S300) and evolutionary characteristics (i.e., the drift vector calculated in step S200). This process mainly includes two inverse topology operators: an entropy reduction splitting operation (Fission) for low-value data and an entropy increase fusion operation (Fusion) for high-value data.
[0056] First, an entropy reduction splitting operation is performed. This operation aims to address the supernode phenomenon in the hypergraph caused by high-frequency common terms (Stop-nodes), where certain concepts (such as Internet, sales, and general business items) are connected by a massive number of edges, leading to combinatorial explosion during graph traversal and diluting effective associations. The system sets a first preset threshold, denoted as the low entropy threshold. When the system detects a superedge Inter-entropy weight Below At that time, the splitting procedure is initiated. The specific logic of the split is as follows:
[0057] Node frequency analysis: System scanning of superedges For all nodes within the given time window, query their global degree (Degree) within the current time window.
[0058] Set partitioning: The system divides nodes into a core set (CoreSet) and a background set (BackgroundSet). The background set contains high-frequency nodes whose global degree exceeds a preset noise threshold; the core set contains the remaining low-frequency entity nodes.
[0059] Physical stripping: The system removes the original hyperedge at the storage layer. Deconstruction. For nodes in the background set, the system removes their pointers from the inverted index. The pointer, or marked as a weak connection, is ignored by default in regular queries. For nodes in the core set, the system constructs a new low-dimensional sub-hyperedge. This replaces the original hyperedges. Through this operation, the system effectively reduces noise and sparsifies the graph, ensuring that subsequent statistical calculations can focus on relationships with substantial discernibility.
[0060] Secondly, an entropy-increasing fusion operation is performed. This operation aims to discover and solidify implicit innovation chains. In innovation activities, a major technological breakthrough often consists of a series of independent events that are sequential in time but complementary in semantic dimension. The system sets a second preset threshold, denoted as the high-entropy threshold. .
[0061] The system searches for a sequence of continuous hyperedges that satisfy the following two strict conditions by scanning the temporal hypergraph through a sliding window.
[0062] Condition 1: The interaction entropy weight of each hyperedge in the sequence is higher than... That is, each single event has a high amount of information.
[0063] Condition 2: The semantic drift vectors of the core entities contained in adjacent hyperedges in the sequence satisfy the orthogonal complement feature. Specifically, the system calculates the drift vectors of adjacent time steps. and The inner product of the vectors. If the inner product is close to zero (i.e. the vectors are orthogonal), it indicates that the innovative entity has expanded on different semantic dimensions in consecutive time steps (for example, first innovating on the algorithm architecture dimension, and then on the hardware adaptation dimension), rather than simply repeating on the same dimension.
[0064] Once a sequence that meets the above conditions is identified, the system triggers the fusion procedure:
[0065] Virtual node instantiation: The system creates a new object type in the graph database: virtual higher-order hyperedge, denoted as... This virtual hyperedge does not directly connect to the original entity, but rather acts as a container to encapsulate the sequence. All original hyperedge IDs in.
[0066] Weighted Nonlinear Gain: To highlight the value of this systematic innovation in statistics, the system assigns a new weight to the virtual higher-order hyperedge that has been nonlinearly amplified. The calculation formula is as follows:
[0067] ;
[0068] in, Let be the interaction entropy of each original hyperedge in the sequence. Represents a sequence of sets The summation of the entropy of all events in the process. It is a decay constant that determines how uncertainty decays over time. This formula ensures that the statistical weight of chain innovation is much greater than the simple linear sum of the weights of its individual components.
[0069] Index redirection: The system creates a new entry in the inverted index. This occurs when a retrieval request hits the sequence. When the index points to any core node in the virtual higher-order hyperedge, it should first point to that virtual higher-order hyperedge. This allows users to directly access the complete innovation evolution path, rather than discrete fragments of events.
[0070] Through the above-mentioned splitting and fusion mechanism, the system described in this embodiment reconstructs the original flat data structure into a three-dimensional structure with hierarchy and value gradient, laying the topological foundation for the accurate statistics in step S500.
[0071] In this embodiment, the innovation momentum statistics based on the reconstructed graph described in step S500 are executed by the statistical analysis service module. This step is a key step in transforming the dynamic data structure constructed in the preceding steps into specific business indicators. Its core technical feature is that the statistical traversal is no longer based on the physical number of records in the original data, but on the effective set of hyperedges after the topology reconstruction in step S400, combined with the node dynamic curvature calculated in step S200 for composite weighting.
[0072] First, statistical query parsing and subgraph extraction are performed. When the system receives a statistical request from a user terminal or upper-layer application (e.g., a request to calculate the innovation momentum of the biopharmaceutical industry in a specific administrative region within a specific time window), the statistical analysis service module first performs semantic parsing on the query conditions, converting them into a set of node identifiers in the graph (e.g., a set of codes for a specific region, a set of technical keywords for a specific industry). Subsequently, the system extracts the subgraph from the reconstructed temporal semantic hypergraph. The subgraph extraction operation is performed. During this process, since the low-entropy hyperedges have already been split, a large number of low-value associations that were originally indexed by general high-frequency words are no longer detected by the retrieval mechanism, and the system only extracts the set of valid hyperedges. The set It includes both unmerged high-entropy independent hyperedges and virtual high-order hyperedges encapsulated by multiple original hyperedges.
[0073] Secondly, perform aggregate calculation of the innovation momentum index. Systematically traverse the set. For each hyperedge in the model, its contribution to the overall innovation momentum is calculated. This calculation model differs from traditional linear summation; instead, it employs a composite metric model that multiplies scarcity by evolutionary strength. An innovation momentum index is defined for a region or industry. The calculation logic is as follows:
[0074] ;
[0075] in, Represents a set or state Relevant indicators or basic weights for comprehensive evaluation Representation and State A collection of related events;
[0076] Depends on the hyperedge The type. If If it is a regular independent hyperedge, then Equal to the interaction entropy weight calculated in step S300 ;like If it is a virtual higher-order hyperedge, then Equal to the gain weight after fusion This ensures that chain-like innovation across time cycles dominates the statistics.
[0077] Mean trajectory curvature This represents the average innovation trajectory curvature of all core entity nodes within the hyperedge at the current moment. The system reads the curvature value from the node's dynamic attributes. The average value is calculated. This factor serves as a dynamic multiplier to reward innovative entities undergoing dramatic transformations or technological upheavals.
[0078] Adjustment coefficient is a non-negative hyperparameter used to adjust the weight of the influence of the intensity of the main evolution on the final result.
[0079] Finally, the system outputs the structured statistical results. The system not only outputs... In addition to scalar numerical values, the system also simultaneously generates interpretive evidence subgraphs. The system serializes and encapsulates the most contributing virtual higher-order hyperedges and their contained original event sequences into a JSON or XML data packet. This data packet details the source path of innovation momentum (e.g., demonstrating the complete evolutionary chain of a company from algorithm development to chip design). In this way, the system achieves penetrating queries from macro-statistical indicators to micro-evolutionary evidence, objectively reflecting the quality and momentum of innovation activities within the region or industry.
[0080] This embodiment illustrates in detail the operational logic of the above-mentioned intelligent statistical method in actual business data flow through a specific application scenario.
[0081] In this embodiment, the scenario is to monitor the innovation vitality within a certain optoelectronic industrial park. A typical innovation entity within this area, Company A Optoelectronic Technology Co., Ltd. (hereinafter referred to as Entity A), is selected. The system treats routine operational activities and their activities along the timeline as processing objects. This scenario demonstrates how the system distinguishes between routine operational activities and substantive innovative activities, and achieves accurate statistics through topology reconstruction.
[0082] Data access and initial mapping phase:
[0083] The system first in the first time window (For example, in the first quarter of 2022) Access to information about entities A business registration change record.
[0084] Original data: Company A Optoelectronics Technology has expanded its business scope to include general equipment leasing.
[0085] Hyperedge mapping: The system generates hyperedges .
[0086] Processing result: The system queried the historical database and found the probability of combinations of manufacturing enterprises and equipment leasing. Extremely high (e.g., 0.45), resulting in a high cross-entropy. It is at a low level.
[0087] Subsequently, the system in the second time window (For example, in the first quarter of 2023) Access a message about an entity News and patent data related to joint research and development.
[0088] Original data: Company A Optoelectronics Technology Co., Ltd. and University B Photonics Laboratory jointly released a holographic display module based on metasurface.
[0089] Hyperedge mapping: The system generates hyperedges ={Entity General equipment leasing, business registration changes, etc.
[0090] Processing results: System statistics revealed that the combination of traditional optoelectronic companies, top university laboratories, and metasurface technology is extremely rare in historical data, with a co-occurrence probability of... Extremely low (e.g., 10) -5 Therefore, the calculated interaction entropy Extremely high.
[0091] Semantic drift and curvature calculation stage:
[0092] System retrieves entity exist semantic vector at time step This vector is mainly distributed in the semantic space of traditional manufacturing, optical cold processing and other dimensions.
[0093] exist At that moment, due to the super-edge High-weight input, entity semantic vectors There has been a significant shift towards nanophotonics and micro / nano fabrication.
[0094] Drift vector calculation: System calculation .
[0095] Innovative Trajectory Curvature Measurement: System Comparison The drift vector at time and The historical drift trend prior to the specified time (assuming the historical trend mainly remained in the direction of capacity expansion). Calculations show that... The direction is almost perpendicular to the historical direction, resulting in a curvature of the innovation trajectory. The value approaches 1. This is marked by the system as a technological shift.
[0096] Topology dynamic reconfiguration execution phase:
[0097] Based on the above calculation results, the system performs a physical reconstruction of the local hypergraph structure containing entity A:
[0098] For hyper-edge Execute entropy reduction split (Fission):
[0099] because The interaction entropy is lower than the set first preset threshold. Furthermore, the node general equipment rental included therein is a high-frequency node (high GlobalDegree) in the entire database.
[0100] The system will super-edge The general equipment leasing node was broken down and moved into the background noise index library, severing the direct strong connection between entity A and the high-frequency node. In subsequent calculations of the optoelectronic industry innovation index, this event will no longer be counted as a valid innovation contribution, thus achieving automatic noise reduction.
[0101] Perform entropy-increasing fusion (Fusion) on the hyperedge sequence:
[0102] The system detected that Immediately following the moment At a certain point, entity A underwent a Series A financing event, generating a super-edge. Entity A, Bing Hard Technology Fund, Series A financing ,and It also has high interaction entropy.
[0103] Further analysis of the system revealed that the hyperedge The drift vector of (technological breakthrough) is mainly in the technological semantic dimension, while the hyperedge The drift vector of (capital injection) is mainly in the financial semantic dimension. The two are orthogonal complements of each other in the vector space, that is, they respectively supplement the innovation attributes of entity A in different dimensions.
[0104] Once the above conditions are met, the system creates a virtual higher-order hyperedge. ,Will and It is encapsulated as a whole. The physical storage of this virtual hyperedge is no longer discrete events, but an aggregated object called Entity A-Metasurface Technology Industrialization Project. Its weight... no and It is not a simple summation of weights, but a value amplified by a non-linear exponent (e.g., to the power of 1.5).
[0105] During the statistical results output phase, when a user initiates a query request: "Assess the innovation momentum of the optoelectronic industrial park in 2023," the system traverses the reconstructed graph and retrieves the relevant data. .
[0106] When calculating the final score, the system executes the following logic:
[0107] ;
[0108] Among them, due to After being merged and enlarged, and the entity curvature Extremely high (indicating a dramatic transformation), the project's statistical contribution value was significantly increased.
[0109] Ultimately, the system outputs not only a statistical figure, but also an evidence subgraph, directly demonstrating the results of the analysis. (technology) and (Capital) integration This serves as core evidence supporting the rise in the region's innovation index. This example demonstrates that the method can effectively filter out elements from daily business operations (such as...). It detects and amplifies the pseudo-innovation data noise generated by the data, and keenly captures and amplifies the substantive innovation chain across dimensions. This objectively reflects the true evolution of innovation in the region.
[0110] See attached document Figure 2 This embodiment elaborates on the structural composition of the hardware logic entity or software functional unit that performs the above method.
[0111] The intelligent statistical system for innovation and entrepreneurship data provided in this embodiment runs on a server cluster or cloud computing environment containing computing and storage units at the physical level. At the logical functional level, the system includes: a multimodal data access module, a semantic evolution calculation module, an entropy weight measurement module, a topology dynamic reconstruction module, and a statistical analysis service module. These modules interact with each other via an internal data bus or API interface, sharing the underlying temporal semantic hypergraph repository.
[0112] The multimodal data access module serves as the system's input interface and basic data construction unit. Internally, it further includes a data cleaning unit, an entity alignment unit, and a hyperedge mapping unit.
[0113] The data cleaning unit receives raw data streams from different data sources, performs regular expression matching to remove HTML tags and non-text characters, and uses a pre-built domain dictionary to perform word segmentation on unstructured text.
[0114] The entity alignment unit maintains a global entity mapping table. This unit identifies entities referring to the same physical object in different data sources by calculating the edit distance of strings and the Jaccard similarity coefficient, and assigns them globally unique entity identifiers (GUIDs).
[0115] The hyperedge mapping unit is responsible for transforming the cleaned discrete events into a hypergraph structure. This unit parses the related elements in each data record, constructs a hyperedge object containing a set of nodes and a generation timestamp, and writes this object into an inverted index database partitioned by time windows. The inverted index database uses time slices as physical storage buckets, supporting efficient batch reads by time range.
[0116] The semantic evolution computation module is used to quantify the dynamic evolution features of entities. This module connects to a vector database and contains embedding encoding units and drift curvature computation units.
[0117] The embedding encoding unit integrates a pre-trained deep language model (such as a Transformer-based model). This unit reads the text attributes in the hyperedge and transforms them into a high-dimensional dense vector. For multiple records of the same entity within the same time window, this unit performs an attention-weighted aggregation operation to generate the temporal state vector of the entity.
[0118] The drift curvature calculation unit retrieves the entity's state vectors for adjacent time windows from the vector database. This unit first performs vector subtraction to obtain the semantic drift vector, then calculates the cosine of the angle between two consecutive drift vectors, and finally uses complement operations to obtain the innovation trajectory curvature. The calculation result is used as a dynamic attribute field and written back to the entity's index entry for the corresponding time window.
[0119] The entropy weighting module is responsible for calculating the scarcity weights of innovative events. This module includes a co-occurrence statistics unit and an entropy calculation unit.
[0120] The co-occurrence statistics unit maintains a historical combination frequency table. For a newly generated hyperedge, the unit extracts the node combinations it contains and queries the historical frequency table for the cumulative occurrence count and total number of historical events for that combination, thereby calculating the prior co-occurrence probability. The unit also integrates a Laplace smoothing algorithm to handle data sparsity issues.
[0121] The entropy calculation unit stores a time decay function model. This unit receives the current statistical time point and the hyperedge generation time point, calculates the time decay coefficient, and combines it with the prior co-occurrence probability to calculate the interaction entropy weight. This weight value is directly associated with the hyperedge data structure, serving as the basis for subsequent topology reconstruction decisions.
[0122] The topology dynamic reconstruction module is the core control unit of this system, used to dynamically adjust the physical storage structure of the map based on the value characteristics of the data. This module specifically includes an entropy reduction splitting unit and an entropy increase fusion unit.
[0123] The entropy reduction splitting unit is used to periodically scan low-entropy hyperedges in the graph. When the interaction entropy weight of a hyperedge is detected to be lower than a first preset threshold, the unit performs a decoupling operation: it identifies the node with the highest global degree within the hyperedge, removes it from the main index path, and retains only the connections between core nodes, thereby generating a dimensionality-reduced sub-hyperedge.
[0124] The entropy-increasing fusion unit is used to identify high-entropy hyperedge sequences with orthogonal complement features. This unit includes a sliding window detector to verify whether the entropy values of consecutive hyperedges are all higher than a second preset threshold, and whether the drift vectors of their core nodes satisfy orthogonality. When the conditions are met, the unit instantiates a virtual high-order hyperedge object in the database, encapsulates the original hyperedge sequence, and redirects the relevant index pointers, enabling subsequent queries to directly hit this aggregated object.
[0125] The statistical analysis service module provides users with final data services. This module includes a query parser and a momentum aggregation engine.
[0126] The query parser receives natural language or structured query requests from users and converts them into query instructions for the node set of the reconstructed graph.
[0127] The momentum aggregation engine traverses effective hyperedges in the reconstructed graph. It reads the final weights (interaction entropy or virtual weights) of the hyperedges and the innovation trajectory curvature of the nodes, performs a non-linear weighted summation operation, and generates an innovation momentum index for a region or industry. Simultaneously, the engine also features subgraph serialization, which can restore the most contributing virtual higher-order hyperedges into visualized evolutionary path data, outputting it as supporting evidence for the statistical results.
[0128] During runtime, each of the above modules strictly adheres to a unidirectional dependency relationship in the data flow: data is accessed first for semantic computation, then for entropy weight measurement, followed by topology reconstruction, and finally statistical analysis. Through this modular design, the system decouples the complex unstructured data processing flow into independently maintainable functional units, ensuring stability and scalability when processing large-scale time-series data.
Claims
1. An intelligent statistical method for innovation and entrepreneurship data, characterized in that, Includes the following steps: Step S100: Use the multimodal data access module to acquire multi-source heterogeneous innovation and entrepreneurship data, clean and align the innovation and entrepreneurship data, and map the processed discrete events as hyperedges containing timestamps to construct an initial state temporal semantic hypergraph; the temporal semantic hypergraph includes a set of nodes, a set of hyperedges, and a storage index divided by the time dimension. Step S200: The entities in the temporal semantic hypergraph are vectorized using the semantic evolution calculation module, and the semantic drift vector of the node is obtained by calculating the vector change of adjacent time windows. Then, the innovation trajectory curvature of the node is determined according to the direction change rate of the semantic drift vector. Step S300: The entropy weight measurement module is used to calculate the historical co-occurrence probability of the combination of nodes inside the hyperedge in the temporal semantic hypergraph, and the interaction entropy weight of each hyperedge is obtained by combining the time decay factor. The interaction entropy weight is used to quantify the information scarcity of innovative events. Step S400: Using the topology dynamic reconstruction module, structural optimization is performed on the temporal semantic hypergraph based on the innovative trajectory curvature obtained in step S200 and the interaction entropy weight obtained in step S300. The structural optimization includes hyperedge splitting operation and hyperedge fusion operation based on interaction entropy gradient, to generate the reconstructed temporal semantic hypergraph. Step S500: In response to the statistical query request, the statistical analysis service module retrieves the effective hyperedges of the target region or target industry in the reconstructed temporal semantic hypergraph, and obtains the innovation momentum index by combining the interaction entropy weight and the innovation trajectory curvature.
2. The intelligent statistical method for innovation and entrepreneurship data according to claim 1, characterized in that, In step S100, the specific method for constructing the temporal semantic hypergraph of the initial state includes: Establish an inverted index structure based on time windows to divide the continuous time axis into multiple discrete time windows; Each independent innovation event is mapped to a hyperedge, which connects two or more heterogeneous nodes; Based on the generation timestamp of the superedge, the superedge is stored in the inverted index entry of the corresponding time window; The inverted index uses the entity's globally unique identifier as the index key and stores a triple containing the hyperedge identifier, time window identifier, and initial weight.
3. The intelligent statistical method for innovation and entrepreneurship data according to claim 1, characterized in that, In step S200, the specific calculation logic for determining the curvature of the innovation trajectory of a node is as follows: Obtain the entity's comprehensive semantic state vector in the current time window and the comprehensive semantic state vector in the previous time window. Calculate the difference between the comprehensive semantic state vectors of the current time window and the previous time window to obtain the semantic drift vector at the current moment. Obtain the semantic drift vector of the entity at the previous time step, and obtain the direction angle value by calculating the cosine distance between the semantic drift vector at the current time step and the semantic drift vector at the previous time step; The innovation trajectory curvature is determined by subtracting the cosine distance of the included angle from the calculated value. The larger the value of the innovation trajectory curvature, the more severe the deviation of the entity's evolution direction.
4. The intelligent statistical method for innovation and entrepreneurship data according to claim 1, characterized in that, In step S300, the specific calculation logic for the interaction entropy weight of each hyperedge is as follows: The co-occurrence frequency of the node combinations within the hyperedge is counted within a historical time window, and the prior co-occurrence probability of the hyperedge is calculated using a smoothing algorithm; Based on the time span between the current statistical time point and the hyperedge generation time point, the time decay coefficient is calculated using an exponential function. The interaction entropy weight is obtained by multiplying the negative logarithm of the prior co-occurrence probability with the time decay coefficient.
5. The intelligent statistical method for innovation and entrepreneurship data according to claim 1, characterized in that, In step S400, the step of the hyperedge splitting operation for handling low-value associations includes: Determine whether the interaction entropy weight of the target hyperedge is lower than a first preset threshold; If the value is lower than the first preset threshold, scan the global degree of all nodes within the target hyperedge; Nodes with a global degree higher than a preset noise threshold are classified as the background set, and the remaining nodes are classified as the core set. In the storage structure of the temporal semantic hypergraph, the connection between the background set nodes and the target hyperedge is removed, and only the core set is retained to construct the dimensionality-reduced sub-hyperedge.
6. The intelligent statistical method for innovation and entrepreneurship data according to claim 1, characterized in that, In step S400, the step of the hyperedge fusion operation for aggregating the innovation chain includes: In the temporal semantic hypergraph, a continuous sequence of time windows is identified, and it is determined whether the interaction entropy weight of each hyperedge in the sequence is higher than a second preset threshold. The orthogonality index is obtained by calculating the inner product between the semantic drift vectors of the core entities contained in adjacent hyperedges in the sequence; If the inner product approaches zero and thus satisfies the orthogonal complement feature, a virtual higher-order hyperedge is generated, and all the original hyperedges in the sequence are encapsulated in the virtual higher-order hyperedge. An index entry is established pointing to the virtual higher-order hyperedge, and a nonlinear gain weight is assigned to the virtual higher-order hyperedge.
7. The intelligent statistical method for innovation and entrepreneurship data according to claim 6, characterized in that, The orthogonal complement feature represents the evolution of the core entity in different semantic dimensions in consecutive time steps; the nonlinear gain weight is obtained by summing the interaction entropy weights of each original hyperedge in the sequence and then performing a power amplification calculation.
8. The intelligent statistical method for innovation and entrepreneurship data according to claim 6, characterized in that, In step S500, the specific logic for calculating the innovation momentum index is as follows: Traverse the set of valid hyperedges retrieved in the reconstructed temporal semantic hypergraph. For each hyperedge in the set, determine the basic weight. If the hyperedge is the virtual higher-order hyperedge, the basic weight is the nonlinear gain weight; otherwise, it is the interaction entropy weight. The average curvature is obtained by calculating the average of the innovation trajectory curvature of all core entities within the hyperedge at the current moment; The one-sided contribution value is obtained by calculating the product of the basic weight and the dynamic adjustment factor. The dynamic adjustment factor is composed of a value plus the product of a preset adjustment coefficient and the average value. The innovation momentum index is obtained by summing the unilateral contribution values of all hyperedges in the set.
9. The intelligent statistical method for innovation and entrepreneurship data according to claim 1, characterized in that, The intelligent statistical method also includes: The statistical analysis service module is used to generate evidence subgraphs; The evidence subgraph is a data object that provides a structured representation of the virtual higher-order hyperedge with the highest contribution in the set of effective hyperedges and the original event sequence encapsulated within it. The evidence subgraph shows the source path of the innovation momentum index.
10. An intelligent statistical system for innovation and entrepreneurship data, characterized in that, An intelligent statistical method for innovation and entrepreneurship data as described in any one of claims 1-9, comprising: The multimodal data access module is used to acquire multi-source heterogeneous innovation and entrepreneurship data, clean and align the innovation and entrepreneurship data, and map the processed discrete events into hyperedges containing timestamps to construct a temporal semantic hypergraph of the initial state. The semantic evolution calculation module is used to vectorize the entities in the temporal semantic hypergraph, and obtain the semantic drift vector of the node by calculating the vector change of adjacent time windows, and then determine the innovative trajectory curvature of the node according to the directional change rate of the semantic drift vector. The entropy weight measurement module is used to calculate the historical co-occurrence probability of node combinations within a hyperedge and obtain the interaction entropy weight of each hyperedge by combining the time decay factor. The topology dynamic reconstruction module, which is connected to the semantic evolution calculation module and the entropy weight measurement module, is used to receive the innovation trajectory curvature and the interaction entropy weight, and accordingly perform hyperedge splitting and hyperedge fusion operations based on the interaction entropy gradient on the initial temporal semantic hypergraph to generate the reconstructed temporal semantic hypergraph. The statistical analysis service module, which is connected to the topology dynamic reconstruction module, is used to respond to statistical query requests, retrieve valid hyperedges in the reconstructed temporal semantic hypergraph, and obtain the innovation momentum index through the composite calculation of the interaction entropy weight and the innovation trajectory curvature.
Citation Information
Cited By
Enterprise business stability ai intelligent monitoring method based on multi-modal data fusion
CN122387806A