Enterprise digital transformation index construction system based on knowledge graph and NLP
By introducing contextual source annotation, semantic segmentation extraction, and redirection modules into the enterprise digital transformation index construction system, the semantic drift problem was solved, the accuracy of semantic understanding and the stability of results were achieved, and the reliability of enterprise digital level assessment was ensured.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG UNIV OF FINANCE & ECONOMICS
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-15
AI Technical Summary
In the process of constructing the enterprise digital transformation index, due to differences in writing context, expression habits and industry semantics, internal enterprise documents and external reports are prone to implicit semantic drift in the expression of the same concepts, which leads to the destruction of cross-domain feature boundaries, disordered semantic weight distribution, and affects the stability and credibility of the assessment.
By establishing a context source annotation module, a semantic segmentation and extraction module, a semantic redirection module, and a semantic weight balancing module, the system addresses the semantic differences between internal corporate documents and external reports by annotating context sources, extracting and redirecting semantic segments, limiting the scope of contextual consistency, and dynamically balancing semantic weights to ensure the accuracy of semantic understanding and the stability of results.
It effectively avoids the erroneous merging of concepts in different contexts, improves the stability and consistency of semantic structure, prevents abnormal fluctuations in the index calculation process, ensures the long-term continuity and credibility of the enterprise digital transformation index, and provides a reliable decision-making basis for assessing the digital level of enterprises.
Smart Images

Figure CN122047249A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of enterprise digital transformation technology, specifically to an enterprise digital transformation index construction system based on knowledge graphs and NLP. Background Technology
[0002] The Enterprise Digital Transformation Index is a comprehensive indicator system used to quantitatively measure an enterprise's overall maturity, capability level, and transformation effectiveness in the digitalization process. It reflects an enterprise's comprehensive digitalization level in areas such as digital infrastructure, intelligent business processes, data governance, organizational structure collaboration, innovation capabilities, and ecosystem integration. Constructing an Enterprise Digital Transformation Index based on Knowledge Graphs and Natural Language Processing (NLP) involves using knowledge graph technology to establish a semantic network connecting enterprise digital elements, technology nodes, business scenarios, and performance indicators. Then, NLP is used to perform semantic analysis and feature extraction on publicly available enterprise information (such as annual reports, policies, news, patents, and recruitment information) to identify key information such as digital behavior, technology investment, and transformation results from unstructured text. This results in quantifiable and interpretable digital feature vectors supported by knowledge graphs. Through dynamic reasoning and multi-dimensional weighted calculation, the Enterprise Digital Transformation Index is generated, enabling intelligent assessment and comparative analysis of an enterprise's digitalization level.
[0003] The existing technology has the following shortcomings:
[0004] During the construction of enterprise digital transformation indices, internal documents and external reports are prone to implicit semantic drift in the expression of the same concepts due to differences in writing context, expression habits, and industry semantics. When NLP models perform semantic extraction and knowledge fusion, they may misclassify these similar words from different sources but with divergent semantics as synonyms. This leads to the incorrect merging of data from different business domains during the knowledge graph fusion stage, resulting in the disruption of cross-domain feature boundaries and disordered semantic weight distribution. Once such implicit semantic drift accumulates, it can easily cause abnormal jumps or short-term distortions in the transformation index calculation results, leading to an overestimation or underestimation of the enterprise's digitalization level by the system, thereby affecting the stability and credibility of the overall assessment.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide a system for constructing an enterprise digital transformation index based on knowledge graphs and NLP, in order to solve the problems mentioned in the background technology.
[0007] To achieve the above objectives, the present invention provides the following technical solution: an enterprise digital transformation index construction system based on knowledge graph and NLP, including a context source annotation module, a semantic segmentation extraction module, a semantic redirection module, a cross-domain semantic fusion module, and a semantic weight balancing module;
[0008] Context Source Annotation Module: Based on the semantic differences between internal corporate documents and external reports, a context source annotation system is established to annotate the writing background and application field of each text data record. Context recognition basis is generated during the text input stage and used to identify the text source and domain features during subsequent semantic extraction.
[0009] Semantic segmentation extraction module: Based on the context recognition criteria formed in the context source annotation system, semantic segmentation extraction is performed on synonyms in text data. During the extraction process, the corresponding context location information and business scenario information are retained to form a set of semantic segments that can be used for subsequent semantic attribution judgment.
[0010] Semantic Redirection Module: Based on the contextual location information and business scenario information contained in the semantic segmentation set, semantic redirection is performed on high-frequency concepts in the key semantic offset area. The semantic attribution of high-frequency concepts is corrected by comparing the context before and after, forming a set of texts that have undergone semantic redirection, so that texts from different sources maintain independent semantic differences.
[0011] Cross-domain semantic fusion module: Based on the semantically redirected text set, it performs cross-domain semantic fusion on text content from different industry sources. During the fusion process, it limits the scope of contextual consistency and performs knowledge fusion operations only in areas with consistent semantic affiliation to prevent erroneous aggregation of concepts from different semantic domains during the fusion stage.
[0012] Semantic weight balancing module: Combining the semantic attribution results after cross-domain semantic fusion, the global semantic weight is dynamically balanced and adjusted. During the construction stage of the enterprise digital transformation index, semantic weights are allocated according to the stability of the context to suppress abnormal index jumps caused by semantic drift accumulation and ensure the stability and credibility of the digital level assessment results.
[0013] Preferably, the steps for establishing a contextual source annotation system based on the semantic differences between internal company documents and external reports include:
[0014] To address the differences in data sources and writing purposes between internal corporate documents and external reports, we categorized the sources of the corpus, collecting different types of text data from corporate databases, government gazettes, industry association materials, and publicly available online information sources. We then identified the sources based on the document's creator, usage scenario, information disclosure attributes, and language structure characteristics.
[0015] For text data whose sources have been segmented, the writing background and application field are labeled. By analyzing the text content structure, paragraph topic word distribution, keyword frequency, and syntactic structure, key information reflecting the text's purpose and field is extracted, and the writing background and application field information are recorded in a structured labeling table.
[0016] The writing background and application domain annotation information are structurally integrated, and a context recognition basis including context source identifiers, writing background descriptions and application domain labels is formed through semantic tag unification, context entry generation and data structure storage.
[0017] By combining context recognition with the text input process, context source identifiers, background descriptions, and application domain labels are loaded during the text input stage, and the text source and domain features are identified based on this information during the semantic extraction stage, thereby ensuring that semantic information remains continuous and consistent during the input and extraction processes.
[0018] Preferably, when context recognition is combined with the text input process, the fields are loaded sequentially according to the context source identifier, writing background description, and application domain label. During the semantic extraction stage, the corresponding semantic parsing strategy is selected based on the context source identifier, so that the text data of internal enterprise documents focuses on identifying management terms and process descriptions, while the text data of external reports focuses on identifying industry terms and market trend descriptions, thus ensuring the accuracy of domain feature identification during the semantic extraction stage.
[0019] Preferably, the steps for semantic segmentation and extraction of synonyms in text data based on context recognition criteria formed in the context source annotation system include:
[0020] Based on the context recognition criteria formed in the context source annotation system, the text data is semantically structured and preliminarily divided. Based on the context source identifier, writing background description and application domain label, the text content is semantically segmented and fragment boundary markers are established.
[0021] Based on the text data with completed semantic structure division, synonym recognition and semantic segmentation are performed within the text fragments. Words with similar semantic expressions in the same context are grouped together, and independent semantic segmentation boundaries are established for each group of words.
[0022] Based on the formation of semantic segments, the contextual location information and business scenario information corresponding to each semantic segment are retained, and the location range, hierarchical relationship and business scenario content of the semantic segments in the text are recorded in a structured form.
[0023] The semantic segments with context source identifiers, context location information and business scenario information are integrated into a semantic segment set, and the semantic segment set is hierarchically organized according to the context source identifier to form a structured semantic resource set that can be called for semantic attribution judgment.
[0024] Preferably, during the integration of the semantic segment set, when organizing the semantic segment set hierarchically according to the context source identifier, the semantic segments formed by internal enterprise documents and the semantic segments formed by external reports are respectively classified into different levels, and the hierarchical and parallel relationships between semantic segments are recorded in the set to maintain semantic coherence and ensure that the domain boundaries of each semantic segment are clear.
[0025] Preferably, the steps for semantically redirecting high-frequency concepts within the key semantic offset region based on the contextual location information and business scenario information contained in the semantic segmentation set include:
[0026] Based on the contextual location information contained in the semantic segment set, key areas that may experience semantic drift are identified and located. The range of key areas of semantic shift is determined according to the semantic segment position, semantic turning point and the connection relationship between the upper and lower segments, and an association is established with the corresponding business scenario information.
[0027] Based on the vocabulary hierarchy information and contextual location information retained in the semantic segmentation set, high-frequency concepts in the key semantic offset region are extracted and semantically aggregated, and the context range, paragraph type and business scenario information of high-frequency concepts are recorded to generate a semantic environment description.
[0028] Based on the semantic environment description of high-frequency concepts, semantic redirection is implemented by comparing the preceding and following contexts, the semantic attribution of high-frequency concepts is corrected and semantic attribution explanations are added to clarify semantic boundaries;
[0029] The revised semantic content is integrated into a new text set, which retains semantic attribution, business scenario information, and contextual location information, forming a semantically redirected text set that maintains the semantic differences between texts from different sources.
[0030] Preferably, the context comparison includes analyzing the logical relationship between the preceding and following paragraphs of the semantic segment where the high-frequency concept is located, and comparing domain feature words with business scenario information to determine the semantic change trend of the high-frequency concept in different contexts. The source type, business domain, context position and context feature words are recorded in the semantic attribution description to maintain clear semantic boundaries and accurate attribution.
[0031] Preferably, the steps for conducting cross-domain semantic fusion of text content from different industry sources based on the semantically redirected text set include:
[0032] Based on the semantically redirected text set, the text content is classified by source and the scope of contextual consistency is initially limited. Based on the semantic attribution description, the text set belonging to the same industry field or business theme is identified, and the fusion scope of semantic attribution consistency is limited.
[0033] Based on the semantically redirected text content, cross-source semantic attribution comparison is performed. By reading the semantic attribution description and business scenario information, the semantic structure of texts from different sources under the same topic is compared to determine the consistency of semantic expression.
[0034] Based on semantic segmentation groups with consistent semantic affiliation, cross-domain semantic fusion operations are carried out. Semantic descriptions of the same business theme in texts from different industries are integrated under the constraint of contextual consistency, while retaining source type, writing background, business scenario and semantic position.
[0035] The fusion results are integrated into a cross-domain semantic fusion set, which records the scope of contextual consistency, semantic attribution description and semantic boundary information to form a structured semantic fusion result and maintain the stability of semantic boundaries.
[0036] Preferably, in cross-domain semantic fusion operations, limiting the scope of contextual consistency includes limiting the upper and lower boundaries of the fusion area based on semantic attribution descriptions and business scenario information, performing fusion operations only within semantic segments with consistent semantic attribution and overlapping boundaries, and retaining source type, semantic location, and contextual scope information in the fusion results to ensure that semantic boundaries are clear and contextual independence remains stable during the semantic fusion process.
[0037] Preferably, the steps for dynamically balancing and adjusting the global semantic weights based on the semantic attribution results after cross-domain semantic fusion include:
[0038] Based on the semantic attribution results after cross-domain semantic fusion, the semantic attributes of each semantic segment are weighted and extracted. According to the semantic attribution description, source features, semantic level, contextual location information and business scenario information are extracted to form traceable semantic weight basic data.
[0039] Based on the analysis of the contextual stability in the semantic attribution results, the contextual location information and business scenario information of each semantic unit are read, and the semantic extension range and semantic boundary differences of different source texts under the same semantic theme are compared to generate contextually stable description entries.
[0040] The global semantic weights are dynamically balanced and adjusted based on the stable context description items. The global weight distribution is adjusted based on semantic units with clear semantic affiliation and stable context, while maintaining the stability of the semantic network hierarchy.
[0041] The dynamically balanced weights are applied to the construction of the enterprise digital transformation index. Semantic weights are assigned according to the stability of the context and correspond to the indicator system to ensure the stability and credibility of the evaluation results.
[0042] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0043] This invention introduces a contextual source annotation system during the text input stage and integrates it throughout the semantic segmentation, extraction, and redirection processes. This ensures that internal enterprise documents and external reports maintain clear source boundaries and domain attributes throughout the semantic processing. By continuously retaining and utilizing the contextual location information and business scenario information of synonyms, it effectively avoids the problem of incorrectly merging concepts in different contexts, making the semantic structure in the knowledge graph more stable and orderly. This, in turn, improves the accuracy of semantic understanding and the consistency of overall results during the construction of the enterprise digital transformation index.
[0044] This invention introduces a contextual consistency constraint mechanism in the cross-domain semantic fusion stage and implements dynamic balancing adjustment of semantic weights in the index construction stage, enabling the semantic weight allocation to truly reflect the stability of different semantics in their respective contexts. By suppressing the cumulative effect of semantic drift, it avoids abnormal fluctuations or short-term distortions during index calculation, thereby ensuring the continuity and credibility of the enterprise digital transformation index in long-term operation and providing a more reliable decision-making basis for assessing the enterprise's digital level. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0046] Figure 1 This is a schematic diagram of the modules of the enterprise digital transformation index construction system based on knowledge graph and NLP of the present invention. Detailed Implementation
[0047] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.
[0048] This invention provides, for example Figure 1The enterprise digital transformation index construction system based on knowledge graph and NLP shown includes a context source annotation module, a semantic segmentation extraction module, a semantic redirection module, a cross-domain semantic fusion module, and a semantic weight balancing module.
[0049] Context Source Annotation Module: Based on the semantic differences between internal corporate documents and external reports, a context source annotation system is established to annotate the writing background and application field of each text data record. Context recognition basis is generated during the text input stage and used to identify the text source and domain features during subsequent semantic extraction.
[0050] The specific implementation method for this step is as follows:
[0051] To address the differences in data sources and writing purposes between internal corporate documents and external reports, a corpus source classification was implemented. In practice, different categories of text data were first collected from corporate databases, government gazettes, industry association materials, and publicly available online information sources. Internal corporate documents included annual digital strategy plans, departmental implementation reports, technical solution specifications, and operational data summaries; external reports included media news, industry research reports, government policy documents, and third-party data analysis commentaries. After collection, source identification was performed based on the document's creator, usage scenario, information disclosure attributes, and language structure characteristics. For example, internal corporate documents are typically written by internal functional departments and used for management decisions or business improvements; therefore, the source identifier records the company name, writing department, document purpose, and publication channel information. For external reports, the source identifier records the publishing media name, report type, publication date, cited objects, and industry category. After source identification, the source attribute of each piece of text data was written into the text metadata field, maintaining a one-to-one correspondence with the original text. In this way, each text has identifiable source information, providing a basis for subsequent contextual annotation.
[0052] Subsequently, for the text data whose sources have been segmented, the writing background and application domain are labeled. In this step, semantic analysis is performed on the text content structure, paragraph topic word distribution, keyword frequency, and syntactic composition to extract key information reflecting the text's purpose and domain. For internal enterprise documents, the focus is on extracting the document type, writing purpose, relevant business processes, and application direction. For example, if words such as "system launch," "process optimization," "internal review," and "performance improvement" frequently appear in the text, it is labeled as an internal operational document in the writing background and as process management or performance management in the application domain. For external reports, the focus is on extracting the report topic, target audience, cited data sources, and industry tags. For example, if phrases such as "market forecast," "digital investment growth," and "technology application trends" appear in the report title, it is labeled as an industry analysis report in the writing background and as market analysis or technology assessment in the application domain. During the labeling process, the writing background and application domain information are stored separately in a structured labeling table, where each field corresponds to a specific semantic level, including text source, writing time, writing subject, business domain, content topic, and industry semantic classification. In this way, multi-dimensional contextual tags can be formed, so that each text has background description and application scenario information, laying the foundation for semantic-level recognition.
[0053] After completing the annotation of the writing background and application domain, this annotation information is structured and integrated to generate context recognition basis that can be used for subsequent semantic extraction. This process includes three stages: semantic tag unification, context entry generation, and data structure storage. In the semantic tag unification stage, annotation information from different sources is transformed into a unified descriptive format. For example, both the "Technical Solution Specification" in internal company documents and the "Technical Application Analysis" in external reports use "Technical Document" as the unified tag, but retain internal and external attribute information at different sub-levels. In the context entry generation stage, an independent context entry is created for each text, containing a context source identifier, writing background description, application domain tag, time attribute, regional attribute, text topic, and domain keywords. The fields of each entry are arranged in a fixed order to ensure matching consistency during context recognition. In the data structure storage stage, context entries are associated with the original text index, forming a bidirectional referencing relationship. Thus, when the text input is called, its corresponding context entry information can be directly retrieved from the index, achieving synchronous mapping between the source layer and the semantic layer. In this process, the generation of context recognition criteria includes not only source information of the text content, but also descriptions at the pragmatic and domain levels, enabling subsequent semantic extraction to accurately distinguish domain differences when identifying words and concepts.
[0054] After establishing the context recognition criteria, these criteria are integrated with the text input process, ensuring that contextual information is fully loaded during the text input stage and continuously transmitted throughout subsequent processing. Specifically, the text input process executes in the following order: first, the text content is loaded; then, the corresponding context recognition criteria are retrieved; and context source identifiers, writing background descriptions, and application domain tags are appended to the text data structure through field matching. Once loaded, when the text is read during the semantic extraction stage, the semantic processing program prioritizes identifying the context source field and selects appropriate semantic parsing strategies based on the source. For example, when the source field is an internal company document, the semantic parsing process focuses more on management terminology and internal process descriptions; when the source field is an external report, it focuses on analyzing industry terms and market trend descriptions. In parallel processing of text from multiple sources, the context recognition criteria, acting as a unified control condition, limits the semantic mapping between texts, preventing the same words in different contexts from being incorrectly classified as synonyms. Thus, during the knowledge fusion stage, the domain attributes and semantic boundaries of each text remain independent, preventing erroneous concept aggregation due to source confusion. By deeply integrating context recognition with the text input stage, it is possible to ensure the continuity and consistency of semantic information throughout the entire process from text input to semantic extraction.
[0055] Through the implementation of the above steps, the semantic source differences between internal enterprise documents and external reports are fully revealed and systematically labeled, providing each text with traceable contextual identification evidence at the input stage. This contextual identification evidence runs through the entire process of subsequent semantic extraction and knowledge fusion, fundamentally reducing word meaning misjudgment and cross-domain semantic drift caused by the mixing of semantic sources. The establishment of the contextual source labeling system ensures that various texts involved in the construction of the enterprise digital transformation index maintain independence and consistency at the source, semantic, and domain levels, thereby guaranteeing the semantic accuracy and result stability of the digital index calculation and providing a reliable data foundation and semantic constraints for assessing the enterprise's digital level.
[0056] Semantic segmentation extraction module: Based on the context recognition criteria formed in the context source annotation system, semantic segmentation extraction is performed on synonyms in text data. During the extraction process, the corresponding context location information and business scenario information are retained to form a set of semantic segments that can be used for subsequent semantic attribution judgment.
[0057] The specific implementation method for this step is as follows:
[0058] Based on the context recognition criteria formed in the context source annotation system, semantic structure recognition and preliminary segmentation are performed on the text data. In this stage, the text content is segmented semantically based on the context source identifier, writing background description, and application domain tag corresponding to each text. Specifically, the text is scanned level by level according to paragraphs, sentences, and phrases. Based on the writing background information recorded in the context recognition criteria, key areas in the text that may have semantic differences are identified. For example, in internal corporate documents, areas containing words such as "platform construction," "process optimization," and "internal integration" are identified as key segments of the internal context; in external reports, areas containing words such as "market platform," "ecosystem construction," and "industry collaboration" are identified as key segments of the external context. During the preliminary segmentation, each identified semantic segment is mapped one-to-one with a context source identifier, and segment boundary markers are established in the text data structure to distinguish content areas belonging to different contexts within the same text. This process ensures that the text structure has a basic semantic hierarchy before semantic segmentation, enabling subsequent segmentation extraction to be performed locally with context recognition criteria as a reference.
[0059] After initial semantic structure segmentation, semantic segmentation is performed on synonyms in the text data. At this stage, combined with the writing background and application domain information stored in the context recognition database, synonym identification and semantic segmentation are performed within text fragments. Specifically, words with similar semantic expressions but different word forms in the same context are grouped together according to their semantic relationships. For example, in internal corporate documents, "digital platform," "information middle platform," and "business support system" semantically express the company's internal IT development direction, and are therefore identified as a group of synonyms during semantic segmentation. In external reports, "industrial internet platform," "digital ecosystem platform," and "industry collaborative network" are extracted as another group of synonyms. During extraction, the text interval containing each group of synonyms is defined as an independent semantic segment, and segmentation boundaries are established within the text. To ensure the integrity of the segmentation, syntactic structure information, contextual connections, and semantic clue words are retained between the start and end positions of each segment, ensuring that each semantic segment not only contains core vocabulary but also retains the contextual environment in which that vocabulary appears within the sentence. The resulting semantic segments have structured features and can be directly correlated with contextual source identifiers, providing structured support for subsequent semantic attribution judgments.
[0060] After initial semantic segmentation, the contextual location information and business scenario information of each semantic segment are retained to ensure that the semantic segments maintain complete contextual relevance during subsequent processing. Contextual location information includes the semantic segment's position range in the text, its sentence / segment level, the logical relationship between adjacent semantic segments, and the relative position of the segment to the title, paragraph, and conclusion. By recording this contextual location information, the contextual level of each semantic segment can be clearly defined. For example, a semantic segment appearing in the preface of a corporate strategy document often describes the overall goals of digital construction, while a semantic segment appearing in the execution section is usually related to technology deployment or business processes. Business scenario information is extracted based on the application domain tags in the context recognition criteria to describe the business functional scope corresponding to the semantic segment, such as production operations, supply chain management, customer service, intelligent manufacturing, and data governance. During the extraction process, contextual location information and business scenario information are bound to the semantic segments to form structured description entries. Each entry includes the semantic segment content, contextual source identifier, writing background description, business scenario information, and contextual location identifier. In this way, semantic segments are no longer just text fragments, but semantic units with source attributes and contextual dependencies, which can accurately identify their domain affiliation and semantic boundaries in subsequent semantic analysis and knowledge fusion stages.
[0061] After retaining contextual location information and business scenario information, all semantic segments are integrated into a semantic segment set and archived in a unified structural format for direct retrieval in subsequent semantic attribution determination stages. The semantic segment set is a structured semantic resource set composed of multiple semantic segments possessing contextual source identifiers, contextual location information, and business scenario information. In specific implementation, the semantic segment set is hierarchically organized according to contextual source identifiers, grouping semantic segments from the same source type (such as internal company documents or external reports) into the same level; simultaneously, the hierarchical relationships between semantic segments are recorded in the set to maintain semantic coherence. For example, in the set, the semantic segment "digital infrastructure construction" from internal company documents is hierarchically related to the semantic segment "data management capability improvement"; the semantic segment "market digitalization investment growth" from external reports is parallelly related to the semantic segment "industry ecosystem collaborative development." After the semantic segment set is formed, it is associated with contextual identification criteria, enabling each semantic segment in the set to trace its textual source, writing background, and business domain. In this way, in the subsequent semantic attribution determination stage, the system can directly obtain the contextual attributes and business scenario information of each semantic segment by calling the semantic segment set, thereby determining whether the semantic attribution belongs to the same business domain or cross-domain expression.
[0062] Through the above steps, based on the context recognition criteria formed in the context source annotation system, the process of semantic segmentation and extraction of synonyms in text data combines semantic extraction with contextual association. This ensures that the semantic relationships between synonyms are not only identified during the segmentation extraction stage, but also their contextual location information and business scenario information are fully preserved. The resulting semantic segment set can provide contextual constraints and source tracing in subsequent semantic attribution judgments and cross-domain semantic fusion processes, enabling the effective differentiation and preservation of semantic differences between internal enterprise documents and external reports. This ensures the accuracy of semantic boundaries and the stability of business domain division during the knowledge graph fusion stage.
[0063] Semantic Redirection Module: Based on the contextual location information and business scenario information contained in the semantic segmentation set, semantic redirection is performed on high-frequency concepts in the key semantic offset area. The semantic attribution of high-frequency concepts is corrected by comparing the context before and after, forming a set of texts that have undergone semantic redirection, so that texts from different sources maintain independent semantic differences.
[0064] The specific implementation method for this step is as follows:
[0065] Based on the contextual location information contained in the semantic segment set, key regions where semantic drift may occur are identified and located. In this stage, the contextual location information field of each semantic segment is read, using the semantic segment set as input. The occurrence position of the semantic segment in the original text, the logical relationship between adjacent segments, and the business scenario category to which it belongs are analyzed. By comparing the semantic connection relationships between different segments from the same source, regions with frequent semantic changes in the text context are identified. For example, in internal corporate documents, "digital platform" might represent the overall architecture of the enterprise's digital construction in the strategy section, while in the execution section it refers to a specific information support platform; in external reports, this term might be used to describe a cross-enterprise collaborative ecosystem. By analyzing the semantic segment positions, semantic turning points, and connecting sentences between adjacent segments, the range of key semantic drift regions can be accurately determined, and these regions are marked in the text, enabling subsequent semantic redirection operations to focus on text segments with significant semantic fluctuations. Each identified key region is associated with its corresponding business scenario information to ensure that semantic judgment is made in conjunction with the scenario during semantic redirection.
[0066] After locating the key semantic shift regions, high-frequency concepts within these regions are extracted and semantically aggregated. In this stage, based on the lexical hierarchy and contextual location information retained in the semantic segmentation set, words within the same key region are statistically analyzed to identify concepts with high frequency of occurrence and potential multiple semantic references. For example, in internal corporate documents, terms like "intelligent system," "digital platform," "process management," and "data center" often have inconsistent semantic attributions due to their appearance in multiple scenarios; in external reports, high-frequency concepts such as "intelligentization," "industrial internet," "platform ecosystem," and "digital economy" are prone to semantic drift due to differences in industry context. To prevent these terms from being incorrectly classified as the same concept during subsequent knowledge fusion, their semantic environment within the key semantic shift regions is extracted, recording their contextual range, paragraph type, logical connectors between sentences, and business scenario information. During this process, each high-frequency concept is assigned a semantic environment description matching its contextual location and business scenario information for subsequent semantic comparison and semantic correction. This processing allows semantic redirection operations to be performed within a clear semantic context, thus avoiding cross-scenario misjudgments.
[0067] After obtaining high-frequency concepts and their semantic environment descriptions, semantic redirection is implemented through contextual comparison. This stage focuses on the key semantic offset region where the high-frequency concept is located, and performs continuous semantic content analysis on its adjacent semantic segments. For each high-frequency concept, the contextual location information of its semantic segment is first read to determine the logical relationship between its preceding and following paragraphs, such as sequential, comparative, supplementary, or causal relationships. Then, combined with the business scenario information corresponding to the semantic segment, the domain feature words in the preceding and following paragraphs are compared to identify the semantic change trend of the concept in different contexts. For example, when "intelligent platform" co-occurs with "internal process integration" and "data visualization" in internal enterprise operational documents, its semantics belong to enterprise information management; while in external reports, when it co-occurs with "ecological collaboration" and "industrial cooperation," its semantics belong to cross-enterprise ecosystem construction. Through this contextual comparison method, the semantic attribution of high-frequency concepts can be corrected, clearly distinguishing the semantic differences of the same word in different sources and scenarios. The revised high-frequency concepts will have semantic attribution descriptions added to their semantic descriptions, including source type, business domain, contextual location, and contextual feature words, thereby ensuring that the semantic boundaries of each concept are clearly defined.
[0068] After semantic redirection of high-frequency concepts, the corrected semantic content is integrated into a new text set, forming a semantically redirected text set. In this stage, each semantic segment in the original semantic segment set is updated, and the redirected high-frequency concepts are replaced with their corresponding semantic positions. Semantic attribution descriptions, business scenario information, and contextual location information are retained in the text set. For multiple expressions of the same concept in texts from different sources, their independent semantic differences are preserved, and cross-source merging is no longer performed. For example, "digital platform" in internal company documents and "industry platform" in external reports remain independent after semantic redirection; the former is marked as "internal digital infrastructure" in the semantic attribution description, while the latter is marked as "industry collaborative ecosystem structure." During the text set integration process, it is ensured that each semantic segment is consistent with its contextual source identifier, giving the semantically redirected text set a two-layer structure: one layer is the text semantic content layer, recording the corrected semantic expression; the other layer is the semantic attribute layer, recording semantic attribution descriptions, contextual location information, and business scenario information. This two-layer structure can effectively maintain the semantic independence and contextual differences between texts in the subsequent semantic fusion and knowledge association process, thereby avoiding the problem of semantic overlap or blurred boundaries between texts from different sources during the fusion stage.
[0069] Through the steps described above, relying on the contextual location information and business scenario information contained in the semantic segmentation set, the process of semantically redirecting high-frequency concepts within key semantic shift areas enables precise correction of the semantic attribution of high-frequency concepts and hierarchical expression of semantic structure. By comparing the preceding and following contexts, the semantic differences of high-frequency concepts in different contexts are clearly defined, ensuring that internal enterprise documents and external reports maintain independent semantic differences after semantic processing. The resulting semantically redirected text set not only retains the semantic features of the original context but also provides traceable semantic attribution criteria, laying a unified semantic foundation for subsequent cross-domain semantic fusion and semantic weight balancing. This effectively prevents semantic misjudgment and feature mismatch problems caused by semantic drift accumulation during the construction of the enterprise digital transformation index, ensuring clear semantic boundaries, well-defined domain divisions, and stable semantic relationships among texts from different sources.
[0070] Cross-domain semantic fusion module: Based on the semantically redirected text set, it performs cross-domain semantic fusion on text content from different industry sources. During the fusion process, it limits the scope of contextual consistency and performs knowledge fusion operations only in areas with consistent semantic affiliation to prevent erroneous aggregation of concepts from different semantic domains during the fusion stage.
[0071] The specific implementation method for this step is as follows:
[0072] Based on the semantically redirected text set, the text content is initially classified by source and its contextual consistency is defined. In this stage, text content from different sources is categorized according to contextual source identifiers, and text sets belonging to the same industry domain or business theme are identified based on the semantic attribution descriptions formed during the semantic redirection stage. For example, text fragments from internal company documents involving "intelligent manufacturing," "data governance," and "digital supply chain" are initially correlated with text fragments from external reports describing "industrial intelligence," "data element marketization," and "supply chain collaboration" to determine their potential correspondence in industry domains. Subsequently, the fusion scope of the text content is limited based on the source type, application domain, and semantic boundary descriptions in the semantic attribution descriptions. The principle of limitation is that only when the semantic attribution descriptions are consistent or similar are they allowed to enter the fusion preparation stage. In this way, before semantic fusion begins, the text content has already passed the contextual consistency screening, avoiding the simultaneous inclusion of concepts from different semantic domains in the fusion scope, and fundamentally reducing the risk of semantic conflicts. After this process, each candidate fusion region has a clear semantic attribution identifier and contextual scope definition, laying the foundation for subsequent semantic comparison and fusion operations.
[0073] After initially defining the scope of contextual consistency, cross-source semantic attribution comparison is performed on the semantically redirected text content to determine the degree of consistency in semantic expression across different source texts. In this stage, the semantic structure of internal documents and external reports on the same topic is compared by reading the semantic attribution descriptions and business scenario information for each semantic segment in the text set. Taking "digital platform" as an example, in the semantic redirection stage, "digital platform" in the internal documents is labeled as "internal digital infrastructure," while in the external reports, it is labeled as "industry collaborative ecosystem structure." Although the word forms are similar, their semantic attributions differ; therefore, they are classified as different semantic regions and not included in the same fusion scope during the comparison stage. For concepts like "data governance," if the internal documents describe it as "data security control system," while the external reports describe it as "enterprise data compliance management mechanism," both are marked as different semantic expressions within the same business scenario in their semantic attribution descriptions, and are therefore considered semantically consistent during the comparison process. The comparison results at this stage group semantically consistent segments together, providing clear regional boundaries for subsequent semantic fusion. During the comparison process, the original contextual location information of the semantic segments is preserved, so that the semantic content from different sources can still be traced back to the original context after fusion, thus ensuring semantic coherence and interpretability.
[0074] After semantic attribution comparison is completed, cross-domain semantic fusion is performed on text regions that meet the semantic consistency criteria. In this stage, semantic segments with consistent semantic attribution are integrated from texts from different industries regarding the same business theme. The fusion process is constrained by contextual consistency, performing knowledge fusion only within the same business domain, the same semantic attribution, and consistent semantic boundaries. For example, in the semantic fusion process of the "intelligent manufacturing" domain, semantic segments from internal company documents concerning "digitalization of production processes," "equipment interconnection," and "process data monitoring" are semantically aligned and integrated with semantic segments from external reports concerning "intelligent factory construction," "production automation upgrades," and "industrial data sharing." Each fusion region retains the contextual description of the source text, including source type, writing background, business scenario, and semantic location, to ensure that the fused semantic expression reflects the multidimensional information from different sources without destroying the original semantic differences. For semantic segments with subtle differences in semantic boundaries, the upper and lower boundaries of the fusion region are limited to prevent erroneous semantic extensions between different domains. For example, when the term "customer management platform" in internal company documents overlaps semantically with "user interaction platform" in external reports, only the content related to business data processing is integrated, while the parts involving user experience design remain separate. This controlled semantic fusion method maintains the independence of each semantic domain while integrating cross-domain knowledge.
[0075] After semantic fusion is completed, the results are integrated into a new cross-domain semantic fusion set, which records the scope of contextual consistency, semantic attribution, and semantic boundary information. Structurally, this fusion set is organized along a business domain framework, hierarchically organizing semantic information from different sources based on semantic attribution consistency. For example, under the theme of "data governance," the set includes semantic content from internal company documents regarding "standardized data collection processes" and semantic content from external reports regarding "data management policy guidelines." These two elements occupy independent sub-layers within the fusion set but are semantically linked across sources. To maintain the traceability of the fusion results, each semantic fusion record retains its source identifier, semantic segment identifier, and contextual location information, allowing any fused semantic content to be traced back to its original context. Once the fusion set is formed, it can serve as input for subsequent semantic weight adjustments and digital index calculations. In this way, cross-domain semantic fusion not only achieves a unified expression of semantics across different industries but also ensures the contextual controllability and semantic boundary stability of the semantic fusion process. The entire fusion process strictly follows the principle of contextual consistency, and fusion is only performed in areas with consistent semantic affiliation, thus avoiding erroneous aggregation of concepts from different semantic domains during the fusion stage.
[0076] Through the implementation of the above steps, the cross-domain semantic fusion process based on the semantically redirected text set can establish semantic connections between text content from different industry sources, while maintaining the independent semantic differences of each source text. By limiting the scope of contextual consistency and strictly controlling the semantic fusion boundary, the fusion result possesses the characteristics of semantic clarity, hierarchical structure, and traceable source. This method can effectively prevent semantic mismatch and conceptual confusion in the construction of enterprise digital transformation indices, ensuring the accuracy and consistency of knowledge fusion, and providing reliable cross-domain semantic support for assessing enterprise digitalization levels.
[0077] Semantic weight balancing module: Combining the semantic attribution results after cross-domain semantic fusion, the global semantic weight is dynamically balanced and adjusted. During the construction phase of the enterprise digital transformation index, semantic weights are allocated according to the stability of the context to suppress abnormal index jumps caused by semantic drift accumulation and ensure the stability and credibility of the digital level assessment results.
[0078] The specific implementation method for this step is as follows:
[0079] Based on the semantic attribution results after cross-domain semantic fusion, the semantic attributes of each semantic segment are weighted. In this stage, the fused semantic content is classified according to business themes and semantic attribution, and the source features, semantic level, contextual location information, and business scenario information of each semantic unit are extracted. The weight extraction process for each semantic unit is centered on its semantic attribution description, combined with its participation frequency and fusion depth in the cross-domain semantic fusion stage, to identify the relative importance of the semantic in the global semantic network. For example, the semantic segment "digital infrastructure construction" in internal enterprise documents is cross-referenced in multiple business themes, and its frequency of occurrence in the field of digital construction is high; therefore, it is assigned a high semantic importance value during weight extraction. Conversely, although "market collaboration mechanism" described in external reports is identified as a relevant concept in the semantic fusion stage, its coverage in the actual business domain of the enterprise is narrow, and therefore, it is assigned a lower importance value during extraction. During the weight extraction process, the source type, occurrence location, and contextual description information of the semantic unit are recorded simultaneously, making the extraction results traceable and interpretable. In this way, each semantic unit has a clear source attribute, semantic importance, and contextual dependency relationship before entering the dynamic equilibrium adjustment.
[0080] After semantic weight extraction, the contextual stability of the semantic attribution results is analyzed to determine the direction and magnitude of weight adjustments. This stage identifies the consistency and fluctuation of semantic expression by comparing semantic changes in different source texts under the same semantic theme. In practice, the contextual location information and business scenario information of each semantic unit in the semantic attribution results are read, and its semantic extension range, semantic boundary differences, and logical connection with adjacent semantic units are compared in different sources. If a semantic unit shows high consistency in multiple source texts, for example, both internal corporate documents and external reports define "intelligent manufacturing" as the digitalization and automation improvement of the production process, then the semantic unit is judged to have a high degree of contextual stability. If the semantic direction of a semantic unit deviates in different source texts, for example, "data platform" in internal corporate documents refers to an internal information integration system, while in external reports it refers to an industrial data service ecosystem, then the semantic unit is judged to have a low degree of contextual stability. During the analysis, each semantic unit generates a contextual stability description entry, which includes semantic attribution, source type, semantic consistency level, contextual fluctuation range, and contextual dependency. This process allows for the identification and labeling of unstable semantic regions in the semantic network, providing a basis for subsequent dynamic balance adjustments.
[0081] After completing the contextual stability analysis, the global semantic weights are dynamically adjusted based on the semantic stability results. In this stage, using context-stable descriptions as a reference, the weight distribution in the global semantic network is rebalanced, giving higher weights to semantic units with clear semantic affiliations and stable contexts, and relatively lower weights to semantic units with ambiguous semantic affiliations or large contextual fluctuations. During the adjustment process, the semantic affiliation relationships formed during the cross-domain semantic fusion stage remain unchanged; only the weight ratios are adjusted. For example, for semantics like "digital management system," which are consistently described in both internal company documents and external reports, their weight is moderately increased after adjustment to enhance their contribution to the index calculation. Conversely, for semantics like "platform ecosystem collaboration," whose meaning varies significantly across different industry contexts, their weight is moderately decreased after adjustment to prevent this type of semantic from having a biasing effect on the index calculation. The dynamic balancing process also considers the structural position of semantic units in the semantic network, i.e., the relationship between core and peripheral semantics, ensuring that the adjustment results do not disrupt the overall semantic hierarchy. The weight distribution, after dynamic balancing, forms a weight system that reflects the stability of the context, enabling the semantic network to maintain balance when semantics interact with different sources.
[0082] After completing the dynamic balancing adjustment of global semantic weights, the adjusted weight results are applied to the construction phase of the enterprise digital transformation index, and semantic weights are allocated and applied according to the stability of context. In this phase, the weight adjustment results are correlated with the indicator system required for constructing the enterprise digital index, and semantic units are assigned to different indicator dimensions according to their weight values. For semantic units with high semantic stability, such as "digital infrastructure," "intelligent manufacturing capabilities," and "data management system," their weights are allocated to the core dimension of the index to reflect the fundamental and long-term level of enterprise digitalization. For semantic units with low semantic stability, such as "digital ecosystem collaboration" and "intelligent marketing innovation," their weights are allocated to the extensional dimension of the index to reflect the dynamic and innovative aspects of digital transformation. During the weight allocation process, the binding relationship between semantic units and their contextual sources is maintained, ensuring that each indicator in the index calculation can be traced back to its semantic source. In this way, semantic weights not only reflect the importance of the semantics themselves but also the stability of semantics in different contexts, thereby suppressing abnormal jumps caused by the accumulation of semantic drift. When the enterprise's text data is updated, the new semantic attribution results can be automatically adjusted to correspond with the original weight system, maintaining the balance and consistency of the index results. In this way, the final enterprise digital transformation index has dynamic balance and structural stability at the semantic level, ensuring that the evaluation results truly reflect the actual development level of enterprises in the process of digital transformation.
[0083] Through the implementation of the above steps, and the dynamic balancing adjustment of global semantic weights based on the semantic attribution results after cross-domain semantic fusion, precise weight allocation of semantic information is achieved globally, ensuring that the semantic weight distribution remains consistent with contextual stability. In the enterprise digital transformation index construction stage, this method, through semantic weight layering and dynamic allocation, suppresses abnormal index jumps caused by accumulated semantic drift, ensuring the stability and reliability of the index calculation results. This process not only maintains the structural consistency of cross-domain semantic fusion but also strengthens the scientific rigor and reliability of enterprise digitalization level assessment through the dynamic balancing mechanism of semantic weights.
[0084] Beneficial effect 1:
[0085] This invention introduces a contextual source annotation system during the text input stage and integrates it throughout the semantic segmentation, extraction, and redirection processes. This ensures that internal enterprise documents and external reports maintain clear source boundaries and domain attributes throughout the semantic processing. By continuously retaining and utilizing the contextual location information and business scenario information of synonyms, it effectively avoids the problem of incorrectly merging concepts in different contexts, making the semantic structure in the knowledge graph more stable and orderly. This, in turn, improves the accuracy of semantic understanding and the consistency of overall results during the construction of the enterprise digital transformation index.
[0086] Benefit 2:
[0087] This invention introduces a contextual consistency constraint mechanism in the cross-domain semantic fusion stage and implements dynamic balancing adjustment of semantic weights in the index construction stage, enabling the semantic weight allocation to truly reflect the stability of different semantics in their respective contexts. By suppressing the cumulative effect of semantic drift, it avoids abnormal fluctuations or short-term distortions during index calculation, thereby ensuring the continuity and credibility of the enterprise digital transformation index in long-term operation and providing a more reliable decision-making basis for assessing the enterprise's digital level.
[0088] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A system for constructing an enterprise digital transformation index based on knowledge graphs and NLP, characterized in that: It includes a context source annotation module, a semantic segmentation and extraction module, a semantic redirection module, a cross-domain semantic fusion module, and a semantic weight balancing module; Contextual Source Annotation Module: Based on the semantic differences between internal corporate documents and external reports, a contextual source annotation system is established to annotate the writing background and application field of each text data record, generating contextual recognition basis during the text input stage; Semantic segmentation extraction module: Based on the context recognition criteria formed in the context source annotation system, semantic segmentation extraction is performed on synonyms in text data. During the extraction process, the corresponding context location information and business scenario information are retained to form a semantic segment set for semantic attribution judgment. Semantic Redirection Module: Based on the contextual location information and business scenario information contained in the semantic segmentation set, semantic redirection is performed on high-frequency concepts in the key semantic offset area. The semantic attribution of high-frequency concepts is corrected by comparing the context before and after, forming a set of texts that have undergone semantic redirection. Cross-domain semantic fusion module: Based on the semantically redirected text set, it performs cross-domain semantic fusion on text content from different industry sources. During the fusion process, it limits the scope of contextual consistency and performs knowledge fusion operations only in areas with consistent semantic affiliation. Semantic weight balancing module: Combining the semantic attribution results after cross-domain semantic fusion, the global semantic weight is dynamically balanced and adjusted. During the construction phase of the enterprise digital transformation index, semantic weights are allocated according to the stability of the context to suppress abnormal index jumps caused by the accumulation of semantic drift.
2. The enterprise digital transformation index construction system based on knowledge graphs and NLP as described in claim 1, characterized in that, The steps to establish a contextual source annotation system based on the semantic differences between internal company documents and external reports include: To address the differences in data sources and writing purposes between internal corporate documents and external reports, we categorized the sources of the corpus, collecting different types of text data from corporate databases, government gazettes, industry association materials, and publicly available online information sources. We then identified the sources based on the document's creator, usage scenario, information disclosure attributes, and language structure characteristics. For text data whose sources have been segmented, the writing background and application field are labeled. By analyzing the text content structure, paragraph topic word distribution, keyword frequency, and syntactic structure, key information reflecting the text's purpose and field is extracted, and the writing background and application field information are recorded in a structured labeling table. The writing background and application domain annotation information are structurally integrated, and a context recognition basis including context source identifiers, writing background descriptions and application domain labels is formed through semantic tag unification, context entry generation and data structure storage. By combining context recognition with the text input process, context source identifiers, background descriptions, and application domain labels are loaded during the text input stage, and the text source and domain features are identified based on this information during the semantic extraction stage.
3. The enterprise digital transformation index construction system based on knowledge graphs and NLP as described in claim 2, characterized in that, When context recognition is combined with the text input process, the fields are loaded sequentially according to the context source identifier, writing background description, and application domain label. During the semantic extraction stage, the corresponding semantic parsing strategy is selected based on the context source identifier, so that the text data of internal enterprise documents focuses on identifying management terms and process descriptions, while the text data of external reports focuses on identifying industry terms and market trend descriptions.
4. The enterprise digital transformation index construction system based on knowledge graphs and NLP according to claim 2, characterized in that, The steps for semantic segmentation and extraction of synonyms in text data based on context recognition criteria formed in the context source annotation system include: Based on the context recognition criteria formed in the context source annotation system, the text data is semantically structured and preliminarily divided. Based on the context source identifier, writing background description and application domain label, the text content is semantically segmented and fragment boundary markers are established. Based on the text data with completed semantic structure division, synonym recognition and semantic segmentation are performed within the text fragments. Words with similar semantic expressions in the same context are grouped together, and independent semantic segmentation boundaries are established for each group of words. Based on the formation of semantic segments, the contextual location information and business scenario information corresponding to each semantic segment are retained, and the location range, hierarchical relationship and business scenario content of the semantic segments in the text are recorded in a structured form. The semantic segments with context source identifiers, context location information and business scenario information are integrated into a semantic segment set, and the semantic segment set is hierarchically organized according to the context source identifier to form a structured semantic resource set that can be called for semantic attribution judgment.
5. The enterprise digital transformation index construction system based on knowledge graphs and NLP according to claim 4, characterized in that, During the integration of the semantic segment set, when organizing the semantic segment set hierarchically according to the context source identifier, the semantic segments formed by internal enterprise documents and the semantic segments formed by external reports are respectively classified into different levels, and the hierarchical and parallel relationships between semantic segments are recorded in the set.
6. The enterprise digital transformation index construction system based on knowledge graphs and NLP according to claim 4, characterized in that, The steps for semantic redirection of high-frequency concepts within key semantic offset regions, based on contextual location information and business scenario information contained in the semantic segmentation set, include: Based on the contextual location information contained in the semantic segment set, the key areas that may cause semantic drift are identified and located. The range of the key areas of semantic shift is determined according to the semantic segment position, semantic turning point and the connection relationship between the upper and lower segments, and an association is established with the corresponding business scenario information. Based on the vocabulary hierarchy information and contextual location information retained in the semantic segmentation set, high-frequency concepts in the key semantic offset region are extracted and semantically aggregated, and the context range, paragraph type and business scenario information of high-frequency concepts are recorded to generate a semantic environment description. Based on the semantic environment description of high-frequency concepts, semantic redirection is implemented by comparing the preceding and following contexts, the semantic attribution of high-frequency concepts is corrected and semantic attribution explanations are added to clarify semantic boundaries; The revised semantic content is integrated into a new text set, which retains semantic attribution, business scenario information, and contextual location information, forming a semantically redirected text set that maintains the semantic differences between texts from different sources.
7. The enterprise digital transformation index construction system based on knowledge graphs and NLP according to claim 6, characterized in that, The context comparison includes analyzing the logical relationship between the preceding and following paragraphs of the semantic segment containing the high-frequency concept, and comparing domain feature words with business scenario information to determine the semantic change trend of the high-frequency concept in different contexts. The source type, business domain, context position and context feature words are recorded in the semantic attribution description.
8. The enterprise digital transformation index construction system based on knowledge graphs and NLP according to claim 6, characterized in that, The steps for cross-domain semantic fusion of text content from different industry sources based on semantically redirected text sets include: Based on the semantically redirected text set, the text content is classified by source and the scope of contextual consistency is initially limited. Based on the semantic attribution description, the text set belonging to the same industry field or business theme is identified, and the fusion scope of semantic attribution consistency is limited. Based on the semantically redirected text content, cross-source semantic attribution comparison is performed. By reading the semantic attribution description and business scenario information, the semantic structure of texts from different sources under the same topic is compared to determine the consistency of semantic expression. Based on semantic segmentation groups with consistent semantic affiliation, cross-domain semantic fusion operations are carried out. Semantic descriptions of the same business theme in texts from different industries are integrated under the constraint of contextual consistency, while retaining source type, writing background, business scenario and semantic position. The fusion results are integrated into a cross-domain semantic fusion set, which records the scope of contextual consistency, semantic attribution description and semantic boundary information to form a structured semantic fusion result.
9. The enterprise digital transformation index construction system based on knowledge graphs and NLP according to claim 8, characterized in that, In cross-domain semantic fusion operations, limiting the scope of contextual consistency includes limiting the upper and lower boundaries of the fusion area based on semantic attribution descriptions and business scenario information, performing fusion operations only within semantic segments with consistent semantic attribution and overlapping boundaries, and retaining source type, semantic location, and contextual scope information in the fusion results.
10. The enterprise digital transformation index construction system based on knowledge graphs and NLP according to claim 8, characterized in that, The steps for dynamically balancing and adjusting the global semantic weights based on the semantic attribution results after cross-domain semantic fusion include: Based on the semantic attribution results after cross-domain semantic fusion, the semantic attributes of each semantic segment are weighted and extracted. According to the semantic attribution description, source features, semantic level, contextual location information and business scenario information are extracted to form traceable semantic weight basic data. Based on the analysis of the contextual stability in the semantic attribution results, the contextual location information and business scenario information of each semantic unit are read, and the semantic extension range and semantic boundary differences of different source texts under the same semantic theme are compared to generate contextually stable description entries. The global semantic weights are dynamically balanced and adjusted based on the context-stable description items, and the global weight distribution is adjusted based on semantic units with clear semantic affiliation and stable context. The dynamically balanced weights are applied to the construction phase of the enterprise digital transformation index, and semantic weights are assigned according to the stability of the context and correspond to the indicator system.