Transferable debt data processing system and method based on double-flow output, and storage medium
Through the convertible bond data processing system based on dual-stream output, the message queue and streaming processing architecture is used to realize high throughput transmission and abnormal data isolation of convertible bond data, solving the problem of insufficient data processing efficiency and accuracy in the existing technology, improving the quality and efficiency of data processing, and providing more accurate and timely support for investment decisions.
Patent Information
- Application Number
- CN202510465214.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-18
AI Technical Summary
It is difficult for the prior art to efficiently process the conversion of unstructured text-to-structured information of convertible bond data and outlier detection. Especially under the requirements of diversified data sources and real-time real-time, traditional manual processing methods are difficult to meet the needs of massive, real-time and high-accuracy data processing.
The convertible bond data processing system based on dual-stream output is adopted to achieve high throughput transmission through message queues, and the abnormal data is isolated using dual-output stream design. Combined with streaming processing architecture, data acquisition, transmission, cleaning and loading modules, efficient data conversion and isolation are achieved.
It improves the quality and efficiency of convertible bond data processing, realizes efficient conversion from unstructured data to structured data, reduces manual intervention and operation costs, and provides more accurate and timely investment decision support.
Smart Images

Figure CN120336413A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically to a convertible bond data processing system, method and storage medium based on dual-stream output. Background Art
[0002] The information of convertible bonds covers multiple dimensions such as issuance terms, conversion price, redemption mechanism, etc. The information sources are extensive (such as company announcements, exchange data, market research reports, etc.) and updated frequently, resulting in a significant increase in data complexity. Traditional manual processing methods are difficult to meet the requirements of processing massive, real-time and highly accurate data.
[0003] At the same time, the development of technologies such as big data, distributed computing, machine learning and artificial intelligence provides new possibilities for financial information processing. However, diverse data sources, inconsistent data quality and real-time requirements still pose severe challenges to existing technologies. In this context, how to achieve efficient conversion from unstructured text to structured information and outlier detection urgently needs to be solved. Summary of the Invention
[0004] To solve the technical problems existing in the prior art, the present invention provides a convertible bond data processing system, method and storage medium based on dual-stream output, which realizes high-throughput transmission of data through a message queue, effectively improves the conversion efficiency from unstructured text to structured information, and effectively isolates abnormal data through a dual-output stream design.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] The present invention discloses a convertible bond data processing system based on dual-stream output. The system adopts a streaming processing architecture, including: a data acquisition module, a data transmission module, a data cleaning module and a data loading module.
[0007] The data acquisition module is used to acquire the original text data of convertible bonds, and the data includes bond issuance status, bond conversion information, bond information, subscription status, winning numbers and important dates;
[0008] The data transmission module is used to transmit the original text data of convertible bonds to the data cleaning module through a specific queue;
[0009] The data cleaning module is used to perform multi-level processing on the original text data of convertible bonds. The multi-level processing includes format verification, standardization and business rule cleaning, and writes the data that meets the rules into the main output stream and abnormal data into the side output stream through a dual-output stream mechanism;
[0010] The data loading module is used to write the structured data in the main output stream into a relational database for downstream visual display.
[0011] As a further improvement of the above solution, the data cleaning module first verifies the format of the data to check the integrity and validity of the data; then performs data standardization to unify the data format; and finally cleans the data according to the set data cleaning rules and marks the output stream type of the cleaned data.
[0012] As a further improvement of the above solution, the data cleaning rules include:
[0013] Remove special characters from the basic data of string type;
[0014] Unify the conversion of floating-point type data into a standard numerical format;
[0015] Parse the date type data into a standardized time format and distinguish the start date and end date of the bond underwriting period;
[0016] Map special fields to enumerated values through a preset template; the special fields include the purchase limit field, redemption reason field, redemption type field, and put-back reason field; among them, map the purchase limit field to the amount in a unified unit; map the redemption reason field to: meet the redemption conditions, forced redemption, or maturity redemption; map the redemption type field to: full redemption or partial redemption; map the put-back reason field to: meet the put-back terms or meet the additional put-back terms.
[0017] As a further improvement of the above solution, the data cleaning module realizes dual output streams through the OutputTag marking of Flink.
[0018] As a further improvement of the above solution, the data collection module constructs a dedicated crawler based on the CrawlSpider class, configures multiple starting URLs to collect multi-source data containing convertible bond information; designs a dynamic request header constructor to automatically generate request header information that conforms to the interface specification, and realizes dynamic Token authentication management.
[0019] As a further improvement of the above solution, the data collection module constructs a dedicated crawler based on the CrawlSpider class, configures multiple starting URLs to collect multi-source data containing convertible bond information; designs a dynamic request header constructor to automatically generate request header information that conforms to the interface specification, and realizes dynamic Token authentication management.
[0020] As a further improvement of the above solution, the data collection module also adopts asynchronous request technology, realizes concurrent request processing of data through coroutines, and avoids server overload through a request queue and a flow control mechanism; assembles the collected multi-source data into a JSON format to generate the original text data of the convertible bond.
[0021] As a further improvement of the above solution, the data transmission module uses the distributed message subscription messaging system Kafka to transmit the original convertible bond text data. The upstream data acquisition module is used as the producer to create a message queue for transmitting the original convertible bond text, and the downstream data cleaning module is used as the consumer to consume the data;
[0022] Among them, a topic is designed for the original convertible bond text data, and multiple partitions are configured. Parallel processing and load balancing are achieved through the partition mechanism; the producer adopts an asynchronous sending mode; the consumer subscribes to the topic through the consumer group mechanism; during the data transmission process, the queue extrusion, consumption delay, and throughput metrics are monitored in real time, and an alarm is triggered when an abnormal situation is found.
[0023] The present invention also discloses a method for processing convertible bond data based on dual-stream output, which is applied to the convertible bond data processing system based on dual-stream output as described above; the data processing method includes:
[0024] Data acquisition: Acquire the original convertible bond text data, which includes bond issuance status, bond conversion information, bond information, subscription status, winning numbers, and important dates;
[0025] Data transmission: Transmit the original convertible bond text data to the data cleaning process through a specific queue;
[0026] Data cleaning: Perform multi-level processing on the original convertible bond text data. The multi-level processing includes format verification, standardization, and business rule cleaning, and write the data that conforms to the rules into the main output stream and the abnormal data into the side output stream through the dual-output stream mechanism;
[0027] Data loading: Write the structured data in the main output stream into a relational database for downstream visual display.
[0028] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the above-mentioned method for processing convertible bond data based on dual-stream output are implemented.
[0029] Compared with the prior art, the beneficial effects of the present invention are:
[0030] 1. The convertible bond data processing system based on dual-stream output disclosed by the present invention realizes high-throughput transmission of data through the message queue, ensures the reliability of the data, effectively isolates the abnormal data through the dual-output stream design, and realizes efficient, accurate, and structured processing of convertible bond data, fundamentally improving the quality and efficiency of convertible bond data processing, and providing more accurate and timely information support for investment decisions.
[0031] 2. The present invention realizes the efficient collection of convertible bond data by using an asynchronous interface call mechanism, uses a Kafka message queue to achieve high-throughput data transmission, uses Flink technology to ensure real-time parallel processing of convertible bond data, establishes an efficient conversion mechanism from unstructured data to structured data, constructs a complete data processing link, realizes the full-process automation of convertible bond collection, transmission, cleaning and reloading, improves data processing efficiency, and reduces manual intervention and operation costs. In addition, a loosely coupled architecture design is adopted, and each module interacts through standardized interfaces, having good scalability and maintainability. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic structural diagram of a convertible bond data processing system based on dual-stream output in Embodiment 1 of the present invention.
[0033] Figure 2 It is a flowchart of a convertible bond data processing method based on dual-stream output in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0035] Embodiment 1
[0036] Please refer to Figure 1 , this embodiment provides a convertible bond data processing system based on dual-stream output. The system adopts a streaming processing architecture and includes: a data collection module, a data transmission module, a data cleaning module, and a data loading module.
[0037] The data collection module is used to collect the original text data of convertible bonds, which includes bond issuance status, bond conversion information, bond information, subscription status, winning numbers, and important dates.
[0038] In some embodiments, the data collection module can use Python web scraping technology to read the information on the web page and use the HTTP protocol to assist in scraping the valid information on the web page. First, select the seed URL and put it into the URL queue to be obtained; secondly, capture the next-level URL that meets the conditions in the URL queue and add it to the URL queue until the final HTML or interface that meets the conditions, that is, the URL of the detailed information of the convertible bond, is obtained; finally, collect and store the text data and transmit it to the lower-level application, that is, the Kafka middleware.
[0039] In this embodiment, the data acquisition module:
[0040] (1) Use the CrawlSpider class as the base class to implement a dedicated convertible bond data crawler for downstream visual display. This crawler configures multiple starting URLs, and the data sources all come from the API interfaces encapsulated by the partners, realizing the unified acquisition of multi-source data.
[0041] (2) Design a dedicated request constructor that can automatically generate request header information that complies with the specifications according to the requirements of different interfaces, including key parameters such as User-Agent, Content-type, and Cookie. Moreover, for interfaces that require authentication, implement a dynamic Token management mechanism to ensure the legality of requests.
[0042] (3) Adopt asynchronous request technology to implement concurrent request processing through coroutines. Efficiently handle a large number of concurrent requests based on the asyncio library. At the same time, through the request queue and rate limiting mechanism, avoid causing excessive pressure on the target server, and set up timeout control and retry mechanisms to ensure data reliability.
[0043] (4) Assemble the collected data into JSON format and transmit it to Kafka, including the source institution sitename, source theme sitesort, bond code symbol, bond name Abname, source link url, bond information zqxx_content, winning number zqh_content, important date zyrq_content, where the bond information, winning number, and important date are string formats nested with JSON.
[0044] The data transmission module is used to transmit the original convertible bond text data to the data cleaning module through a specific queue.
[0045] The data transmission module uses the distributed message subscription message system Kafka to transmit the original convertible bond text data, takes the upstream data acquisition module as the producer, creates a message queue for transmitting the original convertible bond text, and takes the downstream data cleaning module as the consumer to consume the data.
[0046] In this embodiment, the data transmission module:
[0047] (1) Design a dedicated topic (Topic) for convertible bond data and configure multiple partitions to achieve parallel processing and load balancing of data through the partition mechanism. On the producer side, adopt an asynchronous sending mode and optimize the sending performance by configuring the producer buffer size, batch sending parameters, etc. On the consumer side, use the consumer group mechanism, and independent consumer groups can subscribe to this topic to achieve multiplexing of data.
[0048] (2) During the data transmission process, improve the monitoring and management mechanism, and monitor in real time metrics information such as queue backlog, consumption latency, throughput, etc. When abnormal situations are found, the alarm mechanism can be automatically triggered. At the same time, configure the log retention policy, control the historical data cycle to be 7 days, and set the consumption offset to ensure that the data is not consumed repeatedly.
[0049] The data cleaning module is used to perform multi-level processing on the original convertible bond text data. The multi-level processing includes format verification, standardization, and business rule cleaning, and writes the data that conforms to the rules into the main output stream and the abnormal data into the side output stream through the dual output stream mechanism.
[0050] The purpose of data cleaning is to convert text data into standard structured data. The convertible bond text data mainly comes from the official HTML pages of cooperative customers and interface data, and is collected through Python technology. Due to the existence of anti-crawling technology externally on the HTML page, the collected data has redundancy, errors, and inaccuracies. Data cleaning is the process of removing duplicate data and converting the remaining data into data values that meet the requirements and actual situations, and then outputting the cleaned data in the desired format through a series of cleaning steps. Data cleaning processes problems such as missing values, out-of-bounds values, inconsistent codes, and duplicate data from aspects such as the accuracy, integrity, consistency, validity, and uniqueness of the data.
[0051] In this embodiment, the data cleaning module:
[0052] (1) Configure the data source, construct the KafkeUtils tool class, configure consumer parameters, consumer starting position, parallelism, etc., encapsulate functions such as consumers and deserialization, establish a reliable connection with the upstream Kafka queue, and obtain convertible bond data.
[0053] (2) Use a multi-level processing architecture. First, perform data format verification to check the integrity and validity of the data; then, perform data standardization to unify the data format; finally, perform business rule cleaning, process the data according to the given data cleaning rules, and label it.
[0054] The data cleaning rules are mainly as follows:
[0055] 1) Basic data cleaning for string types: The basic type data mainly includes the basic information of convertible bonds, such as bond codes, abbreviations, types, rating agencies, etc. Remove specific special strings such as leading and trailing spaces.
[0056] 2) Cleaning of floating-point type data: The floating-point type data mainly includes price data such as the allotment amount per share, the issue price, the underlying stock price, the conversion price, etc., as well as information such as the price-to-book ratio of the underlying stock and the winning rate of online issuance. It is converted into data of a specified standard type.
[0057] 3) Data of important date type: First, eliminate the disclosure time of the listing date on the Shanghai Stock Exchange and the disclosure time of the listing date on the Shenzhen Stock Exchange. Second, judge the bond underwriting period data. If it is the start and end dates, change it to the start date of the bond underwriting period. If it is the end date, change it to the expiration date of the bond underwriting period.
[0058] 4) Data of some special fields: Since some of the obtained data values are different from the actual data, they need to be reprocessed to obtain the correct data. For example, special fields are mapped to enumerated values through a preset template. In this embodiment, the special fields include the subscription limit field, the redemption reason field, the redemption type field, and the put-back reason field, and their mapping rules are as follows:
[0059] a) Subscription limit field (in ten thousand yuan):
[0060] GENERAL = GENERALValue / 10
[0061] It should be noted that since there may be multiple data sources and the data source units are not unified, in order to be more standardized, a fixed unit, that is, one thousand yuan, is defined and converted.
[0062] b) Redemption reason field:
[0063]
[0064] c) Redemption type field:
[0065]
[0066] d) Put-back reason field:
[0067]
[0068] The data cleaning module realizes a dual output stream through the OutputTag of Flink, and divides the cleaned data into a main output stream and a side output stream. The data that meets all the cleaning rules enters the main output stream for downstream use; the data that does not meet the rules is written into the side output stream, enters a dedicated error handling process, and logs are recorded.
[0069] The data loading module is used to write the structured data of the main output stream into a relational database for downstream visual display.
[0070] The data loading module is also used to design the database table structure by using a standardized database design method, including source institutions, source topics, bond codes, bond names, bond portfolio codes, source links, module information, bond data names, bond data values, source IDs, status information, etc. Standardize the data format through Flink transformation operators, including field mapping, type conversion, etc. And use the built-in Sink connector of Flink to establish a connection with the SQL Server database to implement the batch writing strategy.
[0071] Convert the data into a structured format through mapping. Select the SQL Server relational database as the storage medium, and the data can be displayed visually downstream to quickly obtain relevant information about the convertible bond market.
[0072] Embodiment 2
[0073] Please refer to Figure 2 , this embodiment provides a method for processing convertible bond data based on dual-stream output, which is applied to the convertible bond data processing system based on dual-stream output as described in Embodiment 1. In addition, this embodiment also takes a certain convertible bond as an example to provide specific implementation steps. Similarly to Embodiment 1, the data processing method includes:
[0074] Data collection: Collect the original text data of convertible bonds, which includes bond issuance status, bond conversion information, bond information, subscription status, winning numbers, and important dates;
[0075] Data transmission: Transmit the original text data of convertible bonds to the data cleaning process through a specific queue;
[0076] Data cleaning: Perform multi-level processing on the original text data of convertible bonds. The multi-level processing includes format verification, standardization, and business rule cleaning, and write the data that meets the rules into the main output stream and the abnormal data into the side output stream through the dual-output stream mechanism;
[0077] Data loading: Write the structured data of the main output stream into a relational database for downstream visual display.
[0078] The above method can be implemented according to the following steps:
[0079] 1. Construct a CrawlSpider class as the base class to implement a dedicated crawler for collecting convertible bond data. This crawler realizes the unified collection of multi-source data by configuring multiple starting URLs.
[0080] 2. Build a big data platform, which includes components such as Hadoop, Zookeeper, Kafka, Flink, etc., complete the distributed deployment, and create a data target table in the database to store the final data.
[0081] 3. Build the KafkaUtils class, configure consumer parameters, consumer starting positions, and parallelism to encapsulate Kafka operations, such as consumption functions, deserialization functions, etc.
[0082] 4. Formulate cleaning rules to complete the consumption, cleaning, and transformation of convertible bond data, and use a dual output stream mechanism to isolate abnormal data.
[0083] 5. Collect data with bond code 1270xx (a certain convertible bond) and assemble it into a JSON data Message. Among them, the source institution and url involve sensitive information, which are replaced with xx. The module data is too long, and only part of the data is retained, as shown in Table 1.
[0084] Table 1. JSON data of a certain convertible bond
[0085]
[0086] 6. Transmit the Messages data to the Kafka Topic and use Flink for consumption processing. The processing is as follows:
[0087] 1). Extract zqxx_content (bond information), zqh_content (winning numbers), and zyrq_content (important dates) from the Message, and use regular expressions to remove irrelevant information such as jQuery to make the string into JSON format. Taking the bond information zqxx_content as an example, it is shown in Table 2 below.
[0088] Table 2. Results after removing irrelevant information from bond information
[0089]
[0090] 2). Classify the bond information and extract valid information. Among them, zqh_content (winning numbers) and zyrq_content (important dates) are directly classified without further subdivision; zqxx_content (bond information) needs to be further classified into issuance status, bond conversion information, bond information, and subscription status.
[0091]
[0092] When certain conditions are met, it is possible to know which module it is. For example, when the json text information field obtained is ONLINE_GENERAL_AAU, then flag = 1, which can be understood as tagging. The definition of module information is shown in Table 3.
[0093] Table 3. Definition of module information
[0094]
[0095] 3). Construct a rule mapping table as shown in Table 4:
[0096] Table 4. Rule Mapping Table
[0097]
[0098]
[0099] 4). After cleaning the bond information zqxx_content, it is converted into structured data as shown in Table 5 below.
[0100] Table 5. Structured Data of Bond Information
[0101]
[0102]
[0103] 5). For the winning number zqh_content, by obtaining TYPE and BALLOT_NUM in the data data, the last digit of the winning number is corresponded to a value, and the structured data is shown in Table 6.
[0104] Table 6. Structured Data of Winning Numbers
[0105]
[0106] 6). For the important dates zyrq_content, the disclosure times of the listing dates on the Shanghai Stock Exchange and the Shenzhen Stock Exchange are excluded through DATE_TYPE. And when DATE_TYPE is the bond underwriting period, START_DATE and END_DATE are obtained, corresponding to the start date and end date of the bond underwriting period, and the structured data is shown in Table 7.
[0107] Table 7. Structured Data of Important Dates
[0108] Important Dates Date of Determination of Online Subscription Success Rate 2023 / 10 / 18 0:00 Important Dates Start Date of Bond Underwriting Period 2023 / 10 / 16 0:00 Important Dates Date of Online Subscription Allocation Number 2023 / 10 / 18 0:00 Important Dates Date of Online Lottery Drawing 2023 / 10 / 19 0:00 Important Dates Date of Announcement of Online Subscription Result 2023 / 10 / 20 0:00 Important Dates Date of Online Roadshow Recommendation 2023 / 10 / 17 0:00 Important Dates Record Date for Original Shareholders' Preferential Allotment 2023 / 10 / 17 0:00 Important Dates End Date of Bond Underwriting Period 2023 / 10 / 24 0:00
[0109] 7). Output all the qualified data to the main data stream, and output the abnormal data to the test output stream. Finally, load the data in the main output stream into SQL Server, which can be used for downstream visual display.
[0110] Example 3
[0111] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the convertible bond data processing method based on dual-stream output as described in Example 1 are implemented.
[0112] The computer-readable storage medium may include flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the storage medium may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is generally used to store the operating system installed on the computer device and various application software. In addition, the memory may also be used to temporarily store various data that have been output or will be output.
[0113] As mentioned above, only the preferred specific embodiments of the present invention are described, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A convertible bond data processing system based on dual-stream output, characterized in that, The system adopts a streaming processing architecture, including: A data acquisition module, which is used to acquire the original text data of convertible bonds. The data includes bond issuance status, bond conversion information, bond information, subscription status, winning numbers, and important dates; A data transmission module, which is used to transmit the original text data of convertible bonds to the data cleaning module through a specific queue; A data cleaning module, which is used to perform multi-level processing on the original text data of convertible bonds. The multi-level processing includes format verification, standardization, and business rule cleaning, and writes the data that meets the rules into the main output stream and the abnormal data into the side output stream through a dual output stream mechanism; A data loading module, which is used to write the structured data of the main output stream into a relational database for downstream visual display.
2. The convertible bond data processing system based on dual-stream output according to claim 1, wherein The data cleaning module first checks the integrity and validity of the data by performing format verification on the data; then executes data standardization to unify the data format; finally, cleans the data according to the set data cleaning rules and marks the output stream type of the cleaned data.
3. The convertible bond data processing system based on dual-stream output according to claim 2, wherein, The data cleaning rules include: Removing special characters from the basic data of string type; Uniformly converting the floating-point type data into a standard numerical format; Parsing the date type data into a standardized time format and distinguishing the start date and end date of the bond underwriting period; Mapping special fields to enumeration values through a preset template; the special fields include the subscription limit field, redemption reason field, redemption type field, and putback reason field; among them, mapping the subscription limit field to an amount in a unified unit; mapping the redemption reason field to: meeting the redemption conditions, mandatory redemption, or maturity redemption; mapping the redemption type field to: full redemption or partial redemption; mapping the putback reason field to: meeting the putback terms or meeting the additional putback terms.
4. The convertible bond data processing system based on dual-stream output according to claim 2, wherein The data cleaning module realizes the dual output stream through the OutputTag marking of Flink.
5. The convertible bond data processing system based on dual-stream output according to claim 4, wherein The data acquisition module constructs a dedicated crawler based on the CrawlSpider class, configures multiple starting URLs to acquire multi-source data containing convertible bond information; designs a dynamic request header constructor to automatically generate request header information that meets the interface specifications, and realizes dynamic Token authentication management.
6. The convertible bond data processing system based on dual-stream output according to claim 5, wherein The data loading module is also used to design the database table structure, standardize the data format through Flink transformation operators, and write into the relational database using the Sink connector in Flink; among them, the database table structure includes the source institution, source theme, bond code, bond name, bond portfolio code, source link, module information, bond data name, bond data value, source ID, and status information.
7. The convertible bond data processing system based on dual-stream output according to claim 6, wherein The data acquisition module also adopts asynchronous request technology, realizes concurrent request processing of data through coroutines, and avoids server overload through a request queue and a flow control mechanism; assembles the multi-source data collected into a JSON format to generate the original text data of convertible bonds.
8. The convertible bond data processing system based on dual-stream output according to claim 7, characterized in that, The data transmission module uses the distributed message subscription messaging system Kafka to transmit the original convertible bond text data, taking the upstream data acquisition module as the producer, creating a message queue for transmitting the original convertible bond text, and taking the downstream data cleaning module as the consumer to consume the data; Among them, a topic is designed for the original convertible bond text data, and multiple partitions are configured. Parallel processing and load balancing are achieved through the partition mechanism; the producer adopts an asynchronous sending mode; the consumer subscribes to the topic through the consumer group mechanism; during the data transmission process, the queue extrusion, consumption delay, and throughput metrics are monitored in real time, and an alarm is triggered when an abnormal situation is found.
9. The convertible bond data processing method based on dual-stream output is characterized in that Applied to the convertible bond data processing system based on dual-stream output as described in any one of claims 1 to 8; the data processing method includes: Data acquisition: Acquire the original convertible bond text data, which includes bond issuance status, bond conversion information, bond information, subscription status, winning numbers, and important dates; Data transmission: Transmit the original convertible bond text data to the data cleaning process through a specific queue; Data cleaning: Perform multi-level processing on the original convertible bond text data. The multi-level processing includes format verification, standardization, and business rule cleaning, and write the data that meets the rules into the main output stream and the abnormal data into the side output stream through the dual-output stream mechanism; Data loading: Write the structured data of the main output stream into a relational database for downstream visual display.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the convertible bond data processing method based on dual-stream output as described in claim 9 are implemented.