Business and purchase recombination whole-process information modeling system and method based on announcement content analysis

By modularly processing the text of merger and acquisition announcements and constructing a directed process graph, the problem of lacking full-process information modeling for mergers and acquisitions in existing technologies is solved, thereby improving accuracy and efficiency and supporting investors' decision-making.

CN121935374APending Publication Date: 2026-04-28HAINAN INTERNATIONAL MEDIA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HAINAN INTERNATIONAL MEDIA CO LTD
Filing Date
2026-01-15
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies lack automated methods for modeling the entire M&A process, making it difficult for investors to clearly understand the progress of corporate M&A and hindering comprehensive analysis and investment decisions.

Method used

Design a full-process information modeling system for mergers and acquisitions based on announcement content parsing. The system includes modules such as data acquisition and preprocessing, merger and acquisition process and stage division, keyword extraction and stage mapping, model fine-tuning and attribute extraction, flowchart construction and attribute data export. The system constructs a directed process graph for structured modeling by modularly processing the merger and acquisition announcement text.

Benefits of technology

It enables the automatic extraction of key attribute information of merger and acquisition events from announcement texts, improving the accuracy and consistency of information extraction, enhancing the efficiency of processing massive amounts of announcement data, and providing support for data analysis and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935374A_ABST
    Figure CN121935374A_ABST
Patent Text Reader

Abstract

The invention provides an announcement content analysis-based merger and purchase recombination whole-process information modeling system and method, and relates to the technical field of artificial intelligence. The method comprises the following steps: collecting merger and purchase recombination announcement texts in batches and preprocessing to generate a text abstract; dividing the standardized process of merger-purchase recombination into a plurality of stages; by matching the keyword extraction template of each merger-purchase recombination stage, determining the merger-purchase recombination stage to which each text abstract belongs; performing fine tuning on a pre-trained information extraction model based on a unified information extraction framework, and performing information extraction on each text abstract by using the fine-tuned information extraction model; and clustering and grouping all the text abstracts according to an extraction result, performing multi-level sorting on each group of text abstracts, converting a sorting result into a process directed graph, and storing the process directed graph, thereby realizing structured modeling of the whole merger-purchase recombination process. According to the method, the accuracy and consistency of information extraction are ensured, the processing efficiency of mass announcements is greatly improved, and support is provided for data analysis and decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a merger and acquisition restructuring information modeling system and method based on the parsing of announcement content. Background Technology

[0002] With the rapid development of the market economy, mergers and acquisitions (M&A) activities among listed companies are increasing, involving strategic restructuring and capital operations, and are of great significance to corporate resource optimization, market expansion, and investor decision-making. Against this backdrop, systematic information extraction and process modeling of M&A announcements have become particularly crucial. Traditionally, the analysis of M&A information often relies on case studies, lacking a comprehensive analysis of the event flow. However, by using event extraction technology to automatically extract and model key information from the announcement text, it is possible to reconstruct the full picture of corporate M&A in terms of process, providing more comprehensive data support and insights for investment analysis.

[0003] While current event extraction techniques have made significant progress in the field of natural language processing, they still have considerable shortcomings in modeling process events. Existing announcements typically only indicate that they fall under the category of mergers and acquisitions, without a systematic division of processes or attribute analysis. Currently, there is a lack of technology capable of automatically extracting and modeling the entire process of corporate mergers and acquisitions. Investors cannot clearly understand the progress of corporate mergers and acquisitions, and therefore cannot achieve a comprehensive analysis of each process within the mergers and acquisitions process. The data support that investors hope to use to make investment decisions through merger and acquisition process analysis remains unmet. Summary of the Invention

[0004] To address the shortcomings of the existing technologies, this invention proposes a full-process information modeling system and method for mergers and acquisitions based on the analysis of announcement content by extracting characteristic attributes from merger and acquisition events. The aim is to model the complex relationships between various processes and events in corporate restructuring, providing support for systematic analysis and investment decisions.

[0005] On the one hand, this invention proposes a full-process information modeling system for mergers and acquisitions based on the parsing of announcement content. The system includes: a data acquisition and preprocessing module, a mergers and acquisitions process and stage division module, a keyword extraction and stage mapping module, a model fine-tuning and attribute extraction module, and a flowchart construction and attribute data export module.

[0006] The data acquisition and preprocessing module is used to collect and preprocess merger and acquisition announcement texts in batches, and generate a text summary for each merger and acquisition announcement text.

[0007] The merger and acquisition process and stage division module is used to obtain the standardized process of merger and acquisition, and divide the standardized process into several merger and acquisition stages according to the legal procedures, decision-making body and performance nodes of the event. At the same time, it defines the announcement type and keyword set for each merger and acquisition stage.

[0008] The keyword extraction and stage mapping module is used to match the text summary to the corresponding merger and acquisition stage based on the announcement type and keyword set of all merger and acquisition stages.

[0009] The model fine-tuning and attribute extraction module is used to perform supervised fine-tuning of the pre-trained information extraction model based on the unified information extraction framework, to obtain the fine-tuned information extraction model and to extract information from the text summary, thereby obtaining the text summary extraction result.

[0010] The flowchart construction and attribute data export module is used to cluster and group all the extracted text summary results, and sort the text summaries in each group in multiple levels according to the preset sorting rules. The sorting results of the text summaries are converted into a directed flowchart and saved, thereby realizing the structured modeling of the entire merger and acquisition process.

[0011] Furthermore, the data acquisition and preprocessing module includes: a data import unit, a text cleaning unit, and a text summarization unit;

[0012] The data import unit is used to collect merger and reorganization announcement texts from multiple specified data sources based on preset keywords;

[0013] The text cleaning unit is used to clean up noise in the collected merger and acquisition announcement text; wherein the noise cleaning includes: deleting irrelevant paragraphs and redundant information;

[0014] The text summarization unit is used to extract a text summary from the noise-removed merger and reorganization announcement text using an extractive text summarization method.

[0015] Furthermore, the keyword extraction and stage mapping module includes: a keyword template unit and a stage mapping unit;

[0016] The keyword template unit is used to formulate a keyword extraction template for each merger and acquisition stage based on the keyword set for each stage and a preset set of characteristic keywords.

[0017] The stage mapping unit is used to extract the attribute values ​​of each feature keyword in the text summary according to the preset feature keyword set, generate the attribute set of the text summary, and match the attribute set of the text summary with the keyword extraction template of each M&A and restructuring stage in turn to determine the M&A and restructuring stage to which the text summary belongs.

[0018] Furthermore, the model fine-tuning and attribute extraction module includes: a data annotation and model fine-tuning unit and an attribute extraction unit;

[0019] The data annotation and model fine-tuning unit is used to generate structured labeled data samples for each merger and reorganization stage, and to use all the structured labeled data samples to perform supervised fine-tuning of the pre-trained information extraction model to obtain a fine-tuned information extraction model specifically for the merger and reorganization scenario.

[0020] The attribute extraction unit is used to construct a structured attribute extraction template for each stage of merger and acquisition, and based on all the structured attribute extraction templates, to extract information from each text summary using a fine-tuned information extraction model, thereby obtaining the extraction results for each text summary.

[0021] Furthermore, the flowchart construction and attribute data export module includes: a flowchart construction unit and a data export unit;

[0022] The flowchart construction unit is used to divide all text summaries into several groups based on the event subjects in the extraction results, with each group corresponding to an independent M&A business process; for any M&A business process, the text summaries contained in the M&A business process are sorted according to a preset sorting rule to obtain a text summary sorting result; each text summary in the text summary sorting result is regarded as a graph node, and directed edges are established between adjacent graph nodes according to the order in the text summary sorting result to convert the text summary sorting result into a directed process graph;

[0023] The data export unit is used to convert all text summary extraction results and process directed graphs into a standard structured data format for storage and visualization.

[0024] On the other hand, this invention proposes a method for information modeling of the entire process of mergers and acquisitions based on the parsing of announcement content. This method includes the following steps:

[0025] Collect and preprocess merger and acquisition announcement texts in batches to generate a text summary for each merger and acquisition announcement text;

[0026] Obtain the standardized process for mergers and acquisitions, and divide the standardized process into several merger and acquisition stages according to the legal procedures, decision-making entities and performance nodes of the event. At the same time, define the announcement type and keyword set for each merger and acquisition stage.

[0027] Based on the keyword set of each M&A and restructuring stage, a keyword extraction template for each M&A and restructuring stage is formulated. By matching the text summary with the keyword extraction template of each stage in turn, the M&A and restructuring stage to which each text summary belongs is determined.

[0028] The pre-trained information extraction model is invoked, and the model is supervisedly fine-tuned based on the unified information extraction framework. The fine-tuned information extraction model is then used to extract information from each text summary, resulting in the extraction results for each text summary.

[0029] Based on the event subject, all extracted text summaries are clustered and grouped, and the text summaries in each group are sorted in multiple levels according to the preset sorting rules. The sorting results of the text summaries are converted into a directed process graph and saved, thereby realizing the structured modeling of the entire merger and acquisition process.

[0030] Furthermore, the specific content of the step of batch collecting and preprocessing merger and acquisition announcement texts to generate a text summary for each merger and acquisition announcement text is as follows:

[0031] Based on preset keywords, batch collect merger and reorganization announcement texts from multiple specified data sources;

[0032] The collected merger and acquisition announcement texts are subjected to noise cleanup; wherein the noise cleanup includes: deleting irrelevant paragraphs and redundant information;

[0033] An extractive text summarization method was used to extract text summaries from the noise-removed merger and reorganization announcement texts.

[0034] Furthermore, based on the keyword sets for each M&A and restructuring stage, a keyword extraction template is formulated for each M&A and restructuring stage. Then, by sequentially matching the text summary with the keyword extraction templates for each stage, the specific content of each text summary belonging to the M&A and restructuring stage is determined as follows:

[0035] For any text summary, based on a preset set of feature keywords, the attribute values ​​of each feature keyword in the text summary are extracted, and the attribute set of the text summary is generated.

[0036] Based on the keyword set for each stage of mergers and acquisitions and the preset feature keyword set, keyword extraction templates are developed for each stage of mergers and acquisitions.

[0037] The attribute set of the text summary is matched sequentially with the keyword extraction template of each M&A stage. If the intersection of the attribute set of the text summary with the keyword extraction template of any M&A stage is not empty, then the text summary belongs to that M&A stage.

[0038] Furthermore, the pre-trained information extraction model is invoked, and based on the unified information extraction framework, the information extraction model is subjected to supervised fine-tuning. The fine-tuned information extraction model then extracts information from each text summary, and the specific content of the extraction results for each text summary is as follows:

[0039] For each stage of merger and acquisition, a predetermined number of data samples are generated, and each data sample is structured and annotated using a labeling tool; wherein the labeling information of the structured annotation includes at least: event category, trigger words, event attributes, and the position of the event attributes in the data sample;

[0040] Call the pre-trained information extraction model and the corresponding word segmenter;

[0041] All structured and labeled data samples are divided into training and validation datasets according to a preset ratio.

[0042] A data format adaptation function is used, and a word segmenter is used to convert all training and validation datasets into a standardized format that conforms to the input specification of the information extraction model.

[0043] Set training hyperparameters, and use the transformed training dataset and validation dataset to perform supervised parameter fine-tuning training on the information extraction model to obtain the fine-tuned information extraction model.

[0044] Based on the defined merger and acquisition (M&A) stages, construct a structured attribute extraction template for each M&A stage;

[0045] Using the fine-tuned information extraction model, the structured attribute extraction template corresponding to the merger and reorganization stage to which each text summary belongs is called to extract structured information from each text summary, and the extraction results of each text summary are obtained.

[0046] The extraction results include: the merger and reorganization stage, the event subject, the event category, the trigger words, and the set of event attributes.

[0047] Furthermore, the specific content of clustering and grouping all extracted text summaries based on the event subject, and performing multi-level sorting of the text summaries within each group according to preset sorting rules, converting the text summaries into a directed process graph and saving it, thereby realizing the structured modeling of the entire merger and acquisition process, is as follows:

[0048] Based on the extraction results of all text summaries, text summaries with the same event subject in the extraction results are identified as belonging to the same M&A business process, and then all text summaries are divided into several groups, each group corresponding to an independent M&A business process;

[0049] For any merger and acquisition business process, the text summaries contained in the merger and acquisition business process are sorted according to the preset sorting rules to obtain the text summaries sorting results.

[0050] The sorting rules include: if a text summary with an event status of "terminated", "interrupted", or "stopped" appears, then the text summary is marked as the end point of the process, and all text summaries after the text summary are deleted; for all the remaining text summaries, they are sorted according to the order of the merger and reorganization stages; the obtained stage sorting results are logically sorted according to the preset characteristic keyword order; in the obtained logical sorting results, for text summaries containing the same characteristic keywords, or text summaries that do not contain any characteristic keywords, they are sorted in ascending order according to the announcement date to obtain the final text summary sorting result;

[0051] Each text summary in the text summary ranking result is regarded as a graph node. Each graph node encapsulates the extraction result of the text summary as a node attribute. Directed edges are established between adjacent graph nodes according to the order in the text summary ranking result, and the text summary ranking result is converted into a process directed graph.

[0052] The edge weights of the directed edges are determined by the time interval between announcements of adjacent text summaries and the event priority.

[0053] All text summary extraction results are stored in a non-relational database, and the directed process graph is converted into a common graph file format for storage and visualization.

[0054] The beneficial effects of adopting the above technical solution are as follows:

[0055] This invention's system, through modular data acquisition and preprocessing, stage division, keyword mapping, model fine-tuning, and flowchart construction, enables the automatic extraction of key attribute information from merger and acquisition (M&A) events from announcement texts. This method ensures the accuracy and consistency of information extraction, significantly improves the efficiency of processing massive amounts of announcement data, and provides support for subsequent data analysis and decision-making.

[0056] This invention proposes a phased M&A process segmentation and specific keyword template matching strategy, enabling the system to accurately identify the M&A process stage of the announcement, construct a flowchart showing the progress of each stage, and graphically display the progress status of the M&A process, intuitively reflecting the information at each stage of the M&A event. This invention fills a gap in the research on M&A process announcements from the perspective of event extraction, improving the systematicness and depth of information extraction from such announcements. Attached Figure Description

[0057] Figure 1This is a schematic diagram of the structure of the M&A restructuring full-process information modeling system based on announcement content parsing in this embodiment;

[0058] Figure 2 This is a schematic diagram illustrating the general process of mergers and acquisitions of listed companies in this embodiment;

[0059] Figure 3 This is a flowchart illustrating the information modeling method for the entire process of mergers and acquisitions based on the parsing of announcement content in this invention. Detailed Implementation

[0060] To facilitate understanding of this application, specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and embodiments. The following embodiments are illustrative of the invention but are not intended to limit its scope. Rather, these embodiments are provided to provide a more thorough and complete understanding of the disclosure of this application.

[0061] Example 1:

[0062] This embodiment presents a merger and acquisition restructuring full-process information modeling system based on announcement content parsing, such as... Figure 1 As shown, the system includes: a data acquisition and preprocessing module, a merger and acquisition process and stage division module, a keyword extraction and stage mapping module, a model fine-tuning and attribute extraction module, and a flowchart construction and attribute data export module.

[0063] The data acquisition and preprocessing module is used to collect and preprocess merger and reorganization announcement texts in batches, and generate a text summary for each merger and reorganization announcement text.

[0064] The data acquisition and preprocessing module includes: a data import unit, a text cleaning unit, and a text summarization unit.

[0065] The data import unit is used to collect merger and reorganization announcement texts from multiple specified data sources based on preset keywords.

[0066] In this embodiment, the data import unit filters out content in the Cailian Press announcements whose titles contain the keywords "merger and acquisition" or "restructuring," and searches for and downloads announcements related to mergers and acquisitions from the announcements of listed companies on the Shanghai Stock Exchange and Shenzhen Stock Exchange.

[0067] The text cleaning unit is used to clean up noise in the collected merger and acquisition announcement text; wherein the noise cleaning includes: deleting irrelevant paragraphs and redundant information.

[0068] In this embodiment, noise is removed from the collected merger and reorganization announcement texts to extract key information from the announcements and ensure the effectiveness of subsequent processing.

[0069] The text summarization unit is used to extract a text summary from the noise-removed merger and reorganization announcement text using an extractive text summarization method.

[0070] In this embodiment, a text summary is generated for long announcement texts to retain key content and reduce text length, thereby improving system processing efficiency.

[0071] The merger and acquisition process and stage division module is used to obtain the standardized process of merger and acquisition, and divide the standardized process into several merger and acquisition stages according to the legal procedures, decision-making body and performance nodes of the event. At the same time, it defines the announcement type and keyword set for each merger and acquisition stage.

[0072] In this embodiment, as Figure 2 As shown, the M&A process and stage division module breaks down the M&A process into multiple stages by studying legal documents and announcements. Based on relevant legal documents and announcement standards, it extracts and defines standardized M&A processes to ensure the system can identify and classify key nodes in the announcement process. According to the content of the announcements and a predefined set of keywords, the M&A process is divided into six typical stages: preliminary plan stage, inquiry response stage, M&A committee review stage, restructuring implementation stage, ongoing supervision stage, and termination of restructuring stage. Furthermore, by matching keywords, the system achieves preliminary classification of the announcements.

[0073] The keyword extraction and stage mapping module is used to match text summaries to the corresponding merger and acquisition stage based on the announcement type and keyword set of all merger and acquisition stages.

[0074] The keyword extraction and stage mapping module includes a keyword template unit and a stage mapping unit.

[0075] The keyword template unit is used to formulate a keyword extraction template for each merger and reorganization stage based on the keyword set for each stage and a preset set of characteristic keywords.

[0076] In this embodiment, specific keyword extraction templates are developed for different stages of mergers and acquisitions to ensure that key attributes at each stage can be extracted more accurately. The system performs preliminary attribute extraction on the announcement text using a predefined set of keywords, forming a basic attribute set for each stage.

[0077] The stage mapping unit is used to extract the attribute values ​​of each feature keyword in the text summary according to the preset feature keyword set, generate the attribute set of the text summary, and match the attribute set of the text summary with the keyword extraction template of each M&A and restructuring stage in turn to determine the M&A and restructuring stage to which the text summary belongs.

[0078] In this embodiment, the stage mapping unit accurately maps the announcement text to the corresponding merger and reorganization stage by matching the characteristic keywords in the announcement, so as to form the stage attribution relationship of the announcement.

[0079] The model fine-tuning and attribute extraction module is used to perform supervised fine-tuning of the pre-trained information extraction model based on the unified information extraction framework, to obtain the fine-tuned information extraction model and to extract information from the text summary, thus obtaining the text summary extraction result.

[0080] The model fine-tuning and attribute extraction module includes: a data annotation and model fine-tuning unit and an attribute extraction unit.

[0081] The data annotation and model fine-tuning unit is used to generate structured labeled data samples for each M&A and restructuring stage, and to perform supervised fine-tuning of the pre-trained information extraction model using all structured labeled data samples to obtain a fine-tuned information extraction model specifically for M&A and restructuring scenarios.

[0082] In this embodiment, announcement samples from the merger and acquisition process are labeled to provide high-quality labeled data for the model. The model is optimized by adjusting hyperparameters such as the learning rate and training epochs, ultimately forming a specific model adapted to the extraction of merger and acquisition information.

[0083] The attribute extraction unit is used to construct a structured attribute extraction template for each merger and reorganization stage, and to use the fine-tuned information extraction model to extract structured information from each text summary by calling the structured attribute extraction template corresponding to the merger and reorganization stage to which each text summary belongs, thereby obtaining the extraction results of each text summary.

[0084] In this embodiment, each event in the text summary ranking result is regarded as a graph node, and each graph node encapsulates the event.

[0085] The flowchart construction and attribute data export module is used to cluster and group all the extracted text summary results, and sort the text summaries in each group in multiple levels according to the preset sorting rules. The sorting results of the text summaries are converted into a directed flowchart and saved, thereby realizing the structured modeling of the entire merger and acquisition process.

[0086] The flowchart construction and attribute data export module includes: a flowchart construction unit and a data export unit.

[0087] The flowchart construction unit is used to divide all text summaries into several groups based on the event subjects in the extraction results, with each group corresponding to an independent M&A business process; for any M&A business process, the text summaries contained in the M&A business process are sorted according to a preset sorting rule to obtain a text summary sorting result; each text summary in the text summary sorting result is regarded as a graph node, and directed edges are established between adjacent graph nodes according to the order in the text summary sorting result to convert the text summary sorting result into a directed process graph;

[0088] In this embodiment, a directed graph of the merger and acquisition process is constructed based on the results of attribute extraction. By identifying announcements in the same merger and acquisition process and sorting them according to time and content characteristics, the stage relationships of events are constructed in the graph, and finally a flowchart with stage relationships is generated.

[0089] The data export unit is used to convert all text summary extraction results and process directed graphs into a standard structured data format for storage and visualization.

[0090] In this embodiment, the data export unit structures the flowchart and attribute data into JSON format and stores it in the database. Furthermore, the flowchart is also saved in .graphml format to support subsequent visualization and analysis.

[0091] Example 2:

[0092] This embodiment presents a merger and acquisition restructuring full-process information modeling system based on announcement content parsing, such as... Figure 3 As shown, the method includes the following steps:

[0093] Collect merger and acquisition announcements in batches and preprocess them to generate a text summary for each announcement.

[0094] The specific content of the step of batch collecting and preprocessing merger and acquisition announcement texts to generate a text summary for each merger and acquisition announcement text is as follows:

[0095] Based on preset keywords, batch collection of merger and acquisition announcement texts from multiple specified data sources.

[0096] In this embodiment, the text data of announcements related to mergers and acquisitions are obtained from Cailian Press, and announcements with keywords such as merger and acquisition or restructuring in the title are filtered. In addition, merger and acquisition announcements are searched from the announcements of listed companies on the Shanghai Stock Exchange and Shenzhen Stock Exchange, and the data is downloaded directly.

[0097] The collected merger and acquisition announcement texts are subjected to noise cleanup; wherein the noise cleanup includes: deleting irrelevant paragraphs and redundant information.

[0098] An extractive text summarization method was used to extract text summaries from the noise-removed merger and reorganization announcement texts.

[0099] In this embodiment, the announcement text is filtered and compressed by cleaning up irrelevant text paragraphs and removing interfering information. For announcement texts exceeding 2000 words, the Python sumy library is used for text summarization, removing irrelevant content and shortening and extracting key information.

[0100] Obtain the standardized process for mergers and acquisitions, and divide the standardized process into several stages based on the legal procedures, decision-making bodies, and performance milestones of the event. At the same time, define the announcement type and keyword set for each stage of mergers and acquisitions.

[0101] In this embodiment, by studying relevant legal documents and analyzing the content of actual announcements, the general process of mergers and acquisitions is obtained, such as... Figure 2 As shown. Based on the obtained process, the merger and acquisition process is divided into the following six stages.

[0102] The merger and acquisition (M&A) restructuring stages include, but are not limited to: the preliminary planning stage, the inquiry response stage, the M&A restructuring committee review stage, the restructuring implementation stage, the ongoing supervision stage, and the termination of restructuring stage.

[0103] Table 1 defines the typical announcement types and keyword sets for each stage. The preliminary plan stage includes the following typical announcements: pre-restructuring suspension announcement, preliminary plan release announcement, resumption of trading announcement, board announcement, and shareholders' meeting announcement. The inquiry response stage includes the following typical announcements: exchange inquiries, CSRC feedback, listed company responses, and independent financial advisor responses. The M&A and restructuring committee review stage includes the following typical announcements: M&A and restructuring committee review meeting and suspension announcement, M&A and restructuring committee voting results and resumption of trading announcement, and CSRC approval announcement. The restructuring implementation stage includes the following typical announcements: commencement of restructuring plan implementation, implementation report, and restructuring completion announcement. The ongoing supervision stage includes ongoing supervision announcements. The termination of restructuring stage includes termination of major asset restructuring announcements.

[0104] Table 1. Typical Announcement Types and Keywords for Six Stages of Mergers and Acquisitions

[0105]

[0106] Based on the keyword sets of each M&A and restructuring stage, keyword extraction templates for each M&A and restructuring stage are developed. By sequentially matching the text summary with the keyword extraction templates of each stage, the M&A and restructuring stage to which each text summary belongs is determined.

[0107] In this embodiment, specific extraction templates are designed for different stages of mergers and acquisitions, so that the key attributes of each stage can be extracted more accurately.

[0108] Based on the keyword sets for each M&A and restructuring stage, a keyword extraction template is formulated for each stage. Then, by sequentially matching the text summary with the keyword extraction templates for each stage, the specific content of each text summary belonging to that M&A and restructuring stage is determined as follows:

[0109] For any text summary, based on a preset set of feature keywords, the attribute values ​​of each feature keyword in the text summary are extracted, and the attribute set of the text summary is generated.

[0110] In this embodiment, the text summary As an announcement pending processing, the following core attributes are extracted: Acquiring Party The acquired party Announcement Type Recombination type Event Actions Event Status Suspension status Resumption of trading status Party in charge of the matter and the number of shares traded It also extracts the attribute values ​​of each keyword in the text summary to generate the attribute set of the text summary. Each attribute is processed through a set of keywords during the initial extraction phase. Perform the extraction. For example, if... If "suspension of trading" occurs, then =“Trading suspended”; if a “Mergers and Acquisitions Committee” is involved, then =“Mergers and Acquisitions Committee”, and so on.

[0111] Based on the keyword sets for each stage of mergers and acquisitions and the preset set of characteristic keywords, keyword extraction templates are developed for each stage of mergers and acquisitions.

[0112] In this embodiment, a keyword extraction template is formulated based on a preset set of feature keywords during the planning stage. And there are: Based on a pre-defined set of characteristic keywords, a keyword extraction template is developed for the inquiry and response phase. And there are: Based on a pre-defined set of characteristic keywords, a keyword extraction template was developed for the M&A and restructuring committee review stage. And there are: Based on a pre-defined set of characteristic keywords, a keyword extraction template was developed for the M&A and restructuring committee review stage. And there are: Based on a pre-defined set of characteristic keywords, a keyword extraction template is developed for the continuous supervision phase. And there are: Based on a pre-defined set of characteristic keywords, a keyword extraction template is developed for the termination of the reorganization phase. And there are: .

[0113] The attribute set of the text summary is matched sequentially with the keyword extraction template of each M&A stage. If the intersection of the attribute set of the text summary with the keyword extraction template of any M&A stage is not empty, then the text summary belongs to that M&A stage.

[0114] In this embodiment, a set of stage features and a mapping relationship are established. This is achieved by defining a set of feature keywords for each stage and summarizing the text through matching attribute values. This is temporarily categorized into the corresponding stage. Specifically, a mapping function is defined. ,according to With the characteristic set of each merger and acquisition stage Determine whether the intersection is empty to define the text summary. The aforementioned merger and acquisition restructuring stage is referred to as:

[0115]

[0116] This mapping function summarizes text by matching keywords. It has been initially classified into one of the six stages.

[0117] A pre-trained information extraction model is invoked, and based on a unified information extraction framework, the information extraction model is fine-tuned in a supervised manner. The fine-tuned information extraction model is then used to extract information from each text summary, resulting in the extraction results for each text summary.

[0118] In this embodiment, the event at each stage of the merger and acquisition process is extracted in a refined manner, the model is fine-tuned, and then the key event attributes in the announcement text are accurately extracted.

[0119] It should be noted that in this embodiment, each merger and reorganization announcement text corresponds to a text summary, and each merger and reorganization announcement text is regarded as an event.

[0120] The process involves calling a pre-trained information extraction model and performing supervised fine-tuning based on a unified information extraction framework to adapt the model to the domain. This fine-tuned model is then used to extract information from the text summary. The specific content of the extracted text summary is as follows:

[0121] For each stage of merger and acquisition, a preset number of data samples are generated, and each data sample is structured and annotated using an annotation tool; wherein the annotation information of the structured annotation includes at least: event category, trigger words, event attributes, and the position of the event attributes in the data sample.

[0122] In this embodiment, 10-shot samples are generated for each of the six process stages of mergers and acquisitions, and the data is labeled using the Docano platform. The labeling includes event category, trigger words, event attributes, and their location information, and is finally saved in JSON format.

[0123] Call the pre-trained information extraction model and the corresponding word segmenter.

[0124] All structured and labeled data samples are divided into training and validation datasets according to a preset ratio.

[0125] A data format adaptation function is used, and a word segmenter is employed to convert all training and validation datasets into a standardized format that conforms to the input specification of the information extraction model.

[0126] Set training hyperparameters, and use the transformed training dataset and validation dataset to perform supervised parameter fine-tuning training on the information extraction model to obtain the fine-tuned information extraction model.

[0127] In this embodiment, supervised full fine-tuning is performed using the UIE module from the Python PaddleNLP library. The AutoTokenizer is used to load the tokenizer from the pre-trained model, and the load_dataset loads the training and validation datasets, which are then converted using the convert_example function to adapt to the model input. Appropriate hyperparameters, such as the learning rate and training epochs, are set for efficient fine-tuning.

[0128] Based on the defined merger and acquisition (M&A) stages, a structured attribute extraction template is constructed for each M&A stage.

[0129] By using the fine-tuned information extraction model, the structured attribute extraction template corresponding to the merger and reorganization stage to which each text summary belongs is called to extract structured information from each text summary, and the extraction results of each text summary are obtained.

[0130] The extraction results include: the stage of merger and reorganization, the subject of the event, the event category, the trigger words, and the set of attributes of the event.

[0131] In this embodiment, based on the merger and acquisition (M&A) process, six event schemas are created to define attribute information for the preliminary planning stage, inquiry response stage, M&A committee review stage, restructuring implementation stage, ongoing supervision stage, and termination of restructuring stage. Using PaddleNLP's Taskflow library, a finely tuned Ernie.3.0 model is invoked, and the information_extraction task type is selected to combine the schema information and announcement text. Input the data into the model for extraction.

[0132] Based on the event subject, all extracted text summaries are clustered and grouped, and the text summaries in each group are sorted in multiple levels according to the preset sorting rules. The sorting results of the text summaries are converted into a directed process graph and saved, thereby realizing the structured modeling of the entire merger and acquisition process.

[0133] In this embodiment, based on the M&A and restructuring stage to which each text summary belongs and the extracted structured event attribute data, a complete M&A and restructuring process is further organized and constructed to determine the stage attribution of events and derive the final flowchart and data structure.

[0134] The specific content of clustering and grouping all extracted text summaries based on the event subject, and performing multi-level sorting of the text summaries within each group according to preset sorting rules, converting the text summaries into a directed process graph and saving it, thereby realizing the structured modeling of the entire merger and acquisition process, is as follows:

[0135] Based on the extraction results of all text summaries, text summaries with completely identical event subjects are identified as belonging to the same M&A business process. Then, all text summaries are divided into several groups, with each group corresponding to an independent M&A business process.

[0136] In this embodiment, announcements containing the same pair of "acquiring party" and "acquired party" in the extraction results are considered to belong to different stages of the same merger and acquisition process. The preceding steps have already roughly divided these announcements into six corresponding stages, representing the merger and acquisition process.

[0137] For any merger and acquisition business process, the text summaries contained in the merger and acquisition business process are sorted according to the preset sorting rules to obtain the text summaries sorting results.

[0138] The sorting rules include: if a text summary with an event status of "terminated", "interrupted", or "stopped" appears, then the text summary is marked as the end point of the process, and all text summaries after that text summary are deleted; for all the remaining text summaries, they are sorted according to the order of their respective merger and reorganization stages; the obtained stage sorting results are logically sorted according to a preset order of feature keywords; in the obtained logical sorting results, for text summaries containing the same feature keywords, or text summaries that do not contain any feature keywords, they are sorted in ascending order according to the announcement date to obtain the final text summary sorting result.

[0139] In this embodiment, the sorting rules are as follows: 1. Announcements containing the same keywords, or announcements without any keywords, are sorted according to the announcement date (publish_time); 2. Announcements indicating "termination," "interruption," or "stop" of the restructuring process are marked as process termination events. After this event, all subsequent events are deleted to interrupt the process and ensure the uniqueness of the process termination; 3. Events are first sorted according to the stage sequence, then events within each stage are arranged according to keyword order, and finally events with the same keyword are sorted according to publish_time. Specifically, for each M&A restructuring stage, the events within it are sorted according to preset rules to construct a more refined event sequence. For example: the preliminary plan stage is arranged in the order of suspension, preliminary plan, resumption of trading, board of directors, and shareholders' meeting; the inquiry and response stage is arranged in the order of inquiry, feedback, and response; the M&A restructuring committee review stage is arranged in the order of suspension, review, approval, and resumption of trading; the restructuring implementation stage is arranged in the order of start of implementation, implementation status, and completion of implementation; the continuous supervision stage only includes continuous supervision events.

[0140] Each text summary in the text summary ranking result is regarded as a graph node. Each graph node encapsulates the extraction result of the text summary as a node attribute. Directed edges are established between adjacent graph nodes according to the order in the text summary ranking result, and the text summary ranking result is converted into a process directed graph.

[0141] The weight of the directed edge is determined by the time interval between announcements of adjacent text summaries and the event priority.

[0142] In this embodiment, a directed graph is constructed based on the organized event sequence, with each event considered as a node in the graph. Node attributes include: event type. Trigger words Argument Announcement Date and related extraction attributes, such as:

[0143]

[0144] in Represents graph nodes Event types; Represents graph nodes Trigger words; Represents graph nodes Arguments; Represents graph nodes The date of the announcement.

[0145] Directed edges are established between each node in the order of events, and the edge weights are defined as follows:

[0146]

[0147] in Represents graph nodes Graph Nodes The edge weights between them; Represents graph nodes The date of the announcement; and All are adjustable parameters; To prioritize events and show the sequential relationship between them, an adjacency matrix can be used to represent the structure of the directed graph of the entire merger and acquisition process, reflecting the stage progress and event connection of the merger and acquisition process.

[0148] All text summary extraction results are stored in a non-relational database, and the directed process graph is converted into a common graph file format for storage and visualization.

[0149] In this embodiment, for attribute extraction data, Python's json library is used to output it in a hierarchical JSON structure. The pymongo library is used to establish a connection with MongoDB, and the insert operation is used to insert structured JSON event information into the database for efficient retrieval and management. For the directed graph of the process, the Python NetworkX library is used to implement the graph structure and output it in .graphml and .dot formats. The generated directed graph can be saved as a file to the database or displayed directly on the system front end, providing visual process tracing.

[0150] Example 3:

[0151] This embodiment proposes an electronic device, including: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the aforementioned method for modeling the entire process of mergers and acquisitions based on the parsing of announcement content.

[0152] The electronic device may be a mobile phone, computer, or tablet computer, etc., and includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the M&A restructuring information modeling method based on announcement content parsing as described in the embodiments. It is understood that the electronic device may also include input / output (I / O) interfaces and communication components.

[0153] The processor is used to execute all or part of the steps in the M&A restructuring process information modeling method based on announcement content parsing as described in the above embodiments. The memory is used to store various types of data, which may include, for example, instructions for any application or method in an electronic device, as well as application-related data.

[0154] The processor can be implemented as an Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), controller, microcontroller, microprocessor, or other electronic components, and is used to execute the M&A restructuring full-process information modeling method based on announcement content parsing described in the above embodiments.

[0155] Example 4:

[0156] This embodiment proposes a computer-readable storage medium that stores executable instructions. When these instructions are executed, if they are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0157] The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the merger and reorganization full-process information modeling method based on announcement content parsing described in various embodiments of this application.

[0158] The aforementioned storage media include: flash memory, hard disks, multimedia cards, card-type memory (e.g., SD (Secure Digital Memory Card) or DX (Memory Data Register, MDR) memory), random access memory (RAM), static random-access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, disks, optical discs, servers, APP (Application) app stores, and other media capable of storing program verification codes. These media store computer programs, which, when executed by a processor, can implement the various steps of the aforementioned merger and acquisition restructuring full-process information modeling method based on announcement content parsing.

[0159] Example 5:

[0160] This embodiment proposes a computer program product, including a computer program or instructions, which, when executed by a processor, implements the aforementioned method for information modeling the entire process of mergers and acquisitions based on the parsing of announcement content.

[0161] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a computer program product.

[0162] The various embodiments in this application are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0163] The scope of protection of this application is not limited to the embodiments described above. Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from the scope and spirit of this disclosure. If such modifications and variations fall within the scope of this disclosure and its equivalents, then the intent of this disclosure also includes these modifications and variations.

Claims

1. A merger and acquisition restructuring full-process information modeling system based on announcement content analysis, characterized in that, The system includes: a data acquisition and preprocessing module, a merger and acquisition process and stage division module, a keyword extraction and stage mapping module, a model fine-tuning and attribute extraction module, and a flowchart construction and attribute data export module. The data acquisition and preprocessing module is used to collect and preprocess merger and acquisition announcement texts in batches, and generate a text summary for each merger and acquisition announcement text. The merger and acquisition process and stage division module is used to obtain the standardized process of merger and acquisition, and divide the standardized process into several merger and acquisition stages according to the legal procedures, decision-making body and performance nodes of the event. At the same time, it defines the announcement type and keyword set for each merger and acquisition stage. The keyword extraction and stage mapping module is used to match the text summary to the corresponding merger and acquisition stage based on the announcement type and keyword set of all merger and acquisition stages. The model fine-tuning and attribute extraction module is used to perform supervised fine-tuning of the pre-trained information extraction model based on the unified information extraction framework, to obtain the fine-tuned information extraction model and to extract information from the text summary, thereby obtaining the text summary extraction result. The flowchart construction and attribute data export module is used to cluster and group all the extracted text summary results, and sort the text summaries in each group in multiple levels according to the preset sorting rules. The sorting results of the text summaries are converted into a directed flowchart and saved, thereby realizing the structured modeling of the entire merger and acquisition process.

2. The M&A restructuring full-process information modeling system based on announcement content parsing as described in claim 1, characterized in that, The data acquisition and preprocessing module includes: a data import unit, a text cleaning unit, and a text summarization unit; The data import unit is used to collect merger and reorganization announcement texts from multiple specified data sources based on preset keywords; The text cleaning unit is used to clean up noise in the collected merger and acquisition announcement text; wherein the noise cleaning includes: deleting irrelevant paragraphs and redundant information; The text summarization unit is used to extract a text summary from the noise-removed merger and reorganization announcement text using an extractive text summarization method.

3. The M&A restructuring full-process information modeling system based on announcement content parsing according to claim 1, characterized in that, The keyword extraction and stage mapping module includes: a keyword template unit and a stage mapping unit; The keyword template unit is used to formulate a keyword extraction template for each merger and acquisition stage based on the keyword set for each stage and a preset set of characteristic keywords. The stage mapping unit is used to extract the attribute values ​​of each feature keyword in the text summary according to the preset feature keyword set, generate the attribute set of the text summary, and match the attribute set of the text summary with the keyword extraction template of each M&A and restructuring stage in turn to determine the M&A and restructuring stage to which the text summary belongs.

4. The M&A restructuring full-process information modeling system based on announcement content parsing according to claim 1, characterized in that, The model fine-tuning and attribute extraction module includes: a data annotation and model fine-tuning unit and an attribute extraction unit; The data annotation and model fine-tuning unit is used to generate structured labeled data samples for each merger and reorganization stage, and to use all the structured labeled data samples to perform supervised fine-tuning of the pre-trained information extraction model to obtain a fine-tuned information extraction model specifically for the merger and reorganization scenario. The attribute extraction unit is used to construct a structured attribute extraction template for each stage of merger and acquisition, and based on all the structured attribute extraction templates, to extract information from each text summary using a fine-tuned information extraction model, thereby obtaining the extraction results for each text summary.

5. The M&A restructuring full-process information modeling system based on announcement content parsing according to claim 1, characterized in that, The flowchart construction and attribute data export module includes: a flowchart construction unit and a data export unit; The flowchart construction unit is used to divide all text summaries into several groups based on the event subjects in the extraction results, with each group corresponding to an independent M&A business process; for any M&A business process, the text summaries contained in the M&A business process are sorted according to a preset sorting rule to obtain a text summary sorting result; each text summary in the text summary sorting result is regarded as a graph node, and directed edges are established between adjacent graph nodes according to the order in the text summary sorting result to convert the text summary sorting result into a directed process graph; The data export unit is used to convert all text summary extraction results and process directed graphs into a standard structured data format for storage and visualization.

6. A method for full-process information modeling of mergers and acquisitions based on announcement content parsing, implemented using the full-process information modeling system for mergers and acquisitions based on announcement content parsing as described in any one of claims 1-5, characterized in that, This method includes the following steps: Collect and preprocess merger and acquisition announcement texts in batches to generate a text summary for each merger and acquisition announcement text; Obtain the standardized process for mergers and acquisitions, and divide the standardized process into several merger and acquisition stages according to the legal procedures, decision-making entities and performance nodes of the event. At the same time, define the announcement type and keyword set for each merger and acquisition stage. Based on the keyword set of each M&A and restructuring stage, a keyword extraction template for each M&A and restructuring stage is formulated. By matching the text summary with the keyword extraction template of each stage in turn, the M&A and restructuring stage to which each text summary belongs is determined. The pre-trained information extraction model is invoked, and the model is supervisedly fine-tuned based on the unified information extraction framework. The fine-tuned information extraction model is then used to extract information from each text summary, resulting in the extraction results for each text summary. Based on the event subject, all extracted text summaries are clustered and grouped, and the text summaries in each group are sorted in multiple levels according to the preset sorting rules. The sorting results of the text summaries are converted into a directed process graph and saved, thereby realizing the structured modeling of the entire merger and acquisition process.

7. The method for modeling the entire process of mergers and acquisitions based on announcement content parsing as described in claim 6, characterized in that, The specific content of the step of batch collecting and preprocessing merger and acquisition announcement texts to generate a text summary for each merger and acquisition announcement text is as follows: Based on preset keywords, batch collect merger and reorganization announcement texts from multiple specified data sources; Noise removal was performed on the collected merger and acquisition announcement texts; The noise cleanup mentioned above includes: deleting irrelevant paragraphs and redundant information; An extractive text summarization method was used to extract text summaries from the noise-removed merger and reorganization announcement texts.

8. The method for modeling the entire process of mergers and acquisitions based on the parsing of announcement content as described in claim 7, characterized in that, Based on the keyword sets for each M&A and restructuring stage, a keyword extraction template is formulated for each stage. Then, by sequentially matching the text summary with the keyword extraction templates for each stage, the specific content of each text summary belonging to that M&A and restructuring stage is determined as follows: For any text summary, based on a preset set of feature keywords, the attribute values ​​of each feature keyword in the text summary are extracted, and the attribute set of the text summary is generated. Based on the keyword set for each stage of mergers and acquisitions and the preset feature keyword set, keyword extraction templates are developed for each stage of mergers and acquisitions. The attribute set of the text summary is matched sequentially with the keyword extraction template of each M&A stage. If the intersection of the attribute set of the text summary with the keyword extraction template of any M&A stage is not empty, then the text summary belongs to that M&A stage.

9. The method for modeling the entire process of mergers and acquisitions based on announcement content parsing as described in claim 8, characterized in that, The pre-trained information extraction model is invoked, and supervised fine-tuning is performed on the model based on the unified information extraction framework. The fine-tuned model then extracts information from each text summary, and the specific content of the extraction results for each text summary is as follows: For each stage of mergers and acquisitions, a predetermined number of data samples are generated, and each data sample is structured and annotated using annotation tools. The structured annotation information includes at least: event category, trigger word, event attribute, and the position of the event attribute in the data sample; Call the pre-trained information extraction model and the corresponding word segmenter; All structured and labeled data samples are divided into training and validation datasets according to a preset ratio. A data format adaptation function is used, and a word segmenter is used to convert all training and validation datasets into a standardized format that conforms to the input specification of the information extraction model. Set training hyperparameters, and use the transformed training dataset and validation dataset to perform supervised parameter fine-tuning training on the information extraction model to obtain the fine-tuned information extraction model. Based on the defined merger and acquisition (M&A) stages, construct a structured attribute extraction template for each M&A stage; Using the fine-tuned information extraction model, the structured attribute extraction template corresponding to the merger and reorganization stage to which each text summary belongs is called to extract structured information from each text summary, and the extraction results of each text summary are obtained. The extraction results include: the merger and reorganization stage, the event subject, the event category, the trigger words, and the set of event attributes.

10. The method for modeling the entire process of mergers and acquisitions based on announcement content parsing as described in claim 9, characterized in that, The specific content of clustering and grouping all extracted text summaries based on the event subject, and performing multi-level sorting of the text summaries within each group according to preset sorting rules, converting the text summaries into a directed process graph and saving it, thereby realizing the structured modeling of the entire merger and acquisition process, is as follows: Based on the extraction results of all text summaries, text summaries with the same event subject in the extraction results are identified as belonging to the same M&A business process, and then all text summaries are divided into several groups, each group corresponding to an independent M&A business process; For any merger and acquisition business process, the text summaries contained in the merger and acquisition business process are sorted according to the preset sorting rules to obtain the text summaries sorting results. The sorting rules include: if a text summary with an event status of "terminated", "interrupted", or "stopped" occurs, then that text summary is marked as the end point of the process, and all text summaries following that text summary are deleted; for all remaining text summaries, they are sorted according to the order of their respective merger and reorganization stages; the obtained stage sorting results are logically sorted according to a preset order of feature keywords; in the obtained logical sorting results, text summaries containing the same feature keywords, or text summaries that do not contain any feature keywords, are sorted in ascending order according to the announcement date to obtain the final text summary sorting result; Each text summary in the text summary ranking result is regarded as a graph node. Each graph node encapsulates the extraction result of the text summary as a node attribute. Directed edges are established between adjacent graph nodes according to the order in the text summary ranking result, and the text summary ranking result is converted into a process directed graph. The edge weights of the directed edges are determined by the time interval between announcements of adjacent text summaries and the event priority. All text summary extraction results are stored in a non-relational database, and the directed process graph is converted into a common graph file format for storage and visualization.