A research report analysis method, device, equipment and storage medium
By automatically classifying investment research reports using textual and numerical feature models, and combining static and dynamic level probabilities, the problem of low efficiency and inconsistent quality in investment research report analysis is solved, achieving efficient and unified analysis standards.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2022-09-23
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies for investment research reports suffer from low analysis efficiency, limited data access, inconsistent quality, and a lack of unified standards.
The research reports are automatically graded using textual and numerical feature models. The static and dynamic grade probabilities of the textual and numerical data are combined to calculate the comprehensive probability and determine the true grade.
It enables automatic grading of investment research reports, improves classification efficiency, standardizes classification criteria, and ensures consistency in analysis quality.
Smart Images

Figure CN115470321B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology and can be used in the financial field, particularly a method, apparatus, equipment, and storage medium for analyzing investment research reports. Background Technology
[0002] Investment research, or simply investment research, involves investigating the fundamental information of a listed company, such as its products, sales, output, quality, pricing, raw materials, and corporate capital operations. An investment research report, based on the findings, is called an investment research report. In the investment research report currently displayed to users, a specific company is featured on the cover, which includes a brief description of the company and comments such as a recommendation rating, buy rating, or hold.
[0003] Therefore, the quality of investment research reports directly affects investors' investment returns. For example, when evaluating a company, different authors can give two completely different evaluations. Therefore, it is necessary to screen the investment research reports of varying quality in the market and analyze a large number of investment research reports to select those with higher credibility and recommend them to investors.
[0004] Currently, the analysis of investment research reports is mainly done manually. Specifically, traditional financial professionals rely on their company's structured databases or manually search for information through various channels. Based on their own knowledge and analytical skills, they manually analyze a large number of investment research reports to select high-quality ones.
[0005] However, this manual processing method is inefficient, has limited data acquisition, and the quality of investment research report analysis varies from person to person, with no unified standard. Summary of the Invention
[0006] To address the aforementioned problems in existing technologies, this paper aims to provide a method, apparatus, equipment, and storage medium for analyzing investment research reports. This addresses the issues of low efficiency, limited data acquisition, and inconsistent quality of investment research report analysis due to manual processing methods, which lack a unified standard.
[0007] To solve the above-mentioned technical problems, the specific technical solution presented in this paper is as follows:
[0008] On the one hand, this article provides a method for analyzing investment research reports, including:
[0009] Obtain the categorized investment research reports, which include both textual and numerical data;
[0010] The text data of the investment research report is imported into the text feature model to determine the static level probability of each preset level of the investment research report. The text feature model is trained based on the text data of historical investment research reports and the actual levels.
[0011] The digital data of the investment research report is imported into the digital feature model to determine the dynamic level probability of each preset level of the investment research report. The digital feature model is trained based on the digital data of historical investment research reports and the actual levels.
[0012] Based on the static level probability and the dynamic level probability of each preset level in the investment research report, the comprehensive probability of each preset level in the investment research report is obtained.
[0013] The preset level with the highest overall probability is selected as the true level of the investment research report.
[0014] As one embodiment of this article, before obtaining the categorized investment research reports, the following steps are included:
[0015] Each initial investment research report is preprocessed to obtain an initial investment research report with a specific theme;
[0016] Several initial investment research reports with themes are categorized according to their themes to obtain several types of investment research reports.
[0017] As an example of this article, the preprocessing of each initial investment research report to obtain an initial investment research report with a specific theme further includes:
[0018] The initial investment research report is cleaned to obtain a cleaned preliminary investment research report;
[0019] The cleaned preliminary investment research report is converted into a text matrix;
[0020] The text matrix is imported into the LDA model to obtain the probability distribution of each keyword in the preliminary investment research report;
[0021] The keyword with the highest probability distribution will be used as the topic of the initial investment research report.
[0022] As an embodiment of this article, the step of data cleaning the initial investment research report to obtain a cleaned preliminary investment research report further includes:
[0023] Remove symbols, stop words, and standardized corpora from the initial investment research report to obtain the cleaned preliminary investment research report.
[0024] As one embodiment of this article, the textual data includes the name of the producing organization and the author;
[0025] The digital data includes publication date and reader ratings;
[0026] The preset levels include a first level, a second level, and a third level.
[0027] As one example of this article, after obtaining the categorized investment research reports, the process also includes:
[0028] The text data in each investment research report is matched with the information in the blacklist. If a match is found, the investment research report is filtered.
[0029] The blacklist stores the names of production organizations and / or authors whose dishonesty rate exceeds a predetermined value.
[0030] As an embodiment of this article, the step of obtaining the comprehensive probability of each preset level of the investment research report based on the static level probability and the dynamic level probability of each preset level of the investment research report further includes:
[0031] The static and dynamic probabilities of the same level in the investment research report are weighted and averaged to obtain the comprehensive probability of the same level.
[0032] On the other hand, this article also provides a research report analysis device, including:
[0033] The acquisition unit is used to acquire the categorized investment research reports, wherein the investment research reports include textual data and numerical data;
[0034] A static grade determination unit is used to import the text data of the investment research report into a text feature model to determine the static grade probability of each preset grade of the investment research report. The text feature model is trained based on the text data of historical investment research reports and the actual grades.
[0035] The dynamic level determination unit is used to import the digital data of the investment research report into the digital feature model to determine the dynamic level probability of each preset level of the investment research report, wherein the digital feature model is trained based on the digital data of historical investment research reports and the actual levels.
[0036] The grade integration unit is used to obtain the integrated probability of each preset grade of the investment research report based on the static grade probability and the dynamic grade probability of each preset grade.
[0037] The rating determination unit is used to select the preset rating with the highest overall probability as the true rating of the investment research report.
[0038] On the other hand, this article also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the investment research report analysis methods described above.
[0039] On the other hand, this article also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements any one of the investment research report analysis methods described above.
[0040] Using the above technical solution, by acquiring categorized investment research reports, which include both textual and numerical data, it is possible to determine the textual data of each bibliographic entry and the numerical data corresponding to reader evaluations in each investment research report. By importing the textual data of the investment research reports into a text feature model, the static probability of each preset level for the investment research report is determined. This text feature model is trained based on the textual data and actual levels of historical investment research reports, enabling the determination of the static probability of the investment research report at each preset level using textual data. Furthermore, by importing the numerical data of the investment research reports into a digital feature model, the static probability of each investment research report at each preset level is determined. The dynamic level probability for each preset level is determined by the digital feature model trained on historical investment research report data and actual levels. This model can determine the dynamic level probability of an investment research report at each preset level using digital data. By combining the static and dynamic level probabilities of each preset level, a comprehensive probability for each preset level of the investment research report can be obtained, thus providing a comprehensive probability for each preset level. Finally, by selecting the preset level with the highest comprehensive probability as the actual level of the investment research report, the actual level of the evaluation report can be automatically determined, improving the efficiency of investment research report classification and standardizing the classification criteria.
[0041] To make the above and other objects, features and advantages of this document more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments or prior art described herein, the accompanying drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this article. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 This paper presents an overall system diagram of an investment research report analysis method according to an embodiment of the present invention;
[0044] Figure 2 This document illustrates the steps of an investment research report analysis method as described in an embodiment of the invention.
[0045] Figure 3 This document illustrates a schematic diagram of the investment research report classification method used in this embodiment.
[0046] Figure 4 This document shows a schematic diagram of an investment research report analysis device according to an embodiment of the invention.
[0047] Figure 5 A schematic diagram of a computer device as described in this article is shown.
[0048] Explanation of symbols in the attached drawings:
[0049] 11. Terminal;
[0050] 12. Database;
[0051] 13. Computing server;
[0052] 401. Acquisition Unit;
[0053] 402. Static level determination unit;
[0054] 403. Dynamic Level Determination Unit;
[0055] 404. Level-based integrated unit;
[0056] 405. Grade Determination Unit;
[0057] 502. Computer equipment;
[0058] 504, Processor;
[0059] 506. Memory;
[0060] 508. Drive mechanism;
[0061] 510. Input / output module;
[0062] 512. Input devices;
[0063] 514. Output devices;
[0064] 516. Presentation equipment;
[0065] 518. Graphical User Interface;
[0066] 520. Network interface;
[0067] 522. Communication link;
[0068] 524. Communication bus. Detailed Implementation
[0069] The technical solutions in the embodiments described below will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments described herein, and not all of the embodiments. Based on the embodiments described herein, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this document.
[0070] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings herein are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0071] like Figure 1 The diagram shows an overall system diagram of an investment research report analysis method, including a terminal 11, a database 12, and a computing server 13.
[0072] Terminal 11 is used to obtain the user's operation instructions, which include querying the date range and the type of investment research report. After obtaining the user's operation instructions, terminal 11 sends them to database 12.
[0073] Database 12 stores several investment research reports, each with a tag indicating its publication date and topic. Topics include military enterprises, semiconductor companies, aerospace companies, and non-ferrous metal companies. Upon receiving an operation command, database 12 sends all investment research reports within the corresponding query date range to the computing server. For example, if a user inputs "2021-10-01 to 2021-11-01," the database will send all investment research reports published between 2021-10-01 and 2021-11-01 to computing server 13.
[0074] The calculation server 13 is used to calculate the overall probability of all investment research reports at a preset level. If the preset levels are level 1, level 2, and level 3, the calculation server 13 can give the overall probability of an investment research report at the three levels. For example, if the overall probability of investment research report A at level 1 is 20%, at level 2 it is 60%, and at level 3 it is 20%, then the calculation server 13 can send the actual level of the investment research report to the terminal 11 as level 2.
[0075] Currently, the analysis of investment research reports is mainly done manually. Specifically, traditional financial professionals rely on their company's structured databases or manually search for information through various channels. Based on their own knowledge and analytical skills, they manually analyze a large number of investment research reports to select high-quality ones.
[0076] However, this manual processing method is inefficient, has limited data acquisition, and the quality of investment research report analysis varies from person to person, with no unified standard.
[0077] To address the aforementioned issues, this paper provides a method for analyzing investment research reports, which can automatically categorize investment research reports. Figure 2 This is a schematic diagram illustrating the steps of an investment research report analysis method provided in this embodiment. This specification provides the operational steps of the method described in the embodiments or flowcharts, but based on conventional or non-creative labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only execution order. In actual system or device products, the methods shown in the embodiments or accompanying drawings can be executed sequentially or in parallel. Specifically, as shown... Figure 1 As shown, the method may include:
[0078] Step 201: Obtain the categorized investment research report, which includes textual data and numerical data.
[0079] Step 202: Import the text data of the investment research report into the text feature model to determine the static level probability of each preset level of the investment research report. The text feature model is trained based on the text data of historical investment research reports and the actual levels.
[0080] Step 203: Import the digital data of the investment research report into the digital feature model to determine the dynamic level probability of each preset level of the investment research report. The digital feature model is trained based on the digital data of historical investment research reports and the actual levels.
[0081] Step 204: Based on the static level probability and the dynamic level probability of each preset level in the investment research report, obtain the comprehensive probability of each preset level in the investment research report.
[0082] Step 205: Select the preset level with the highest overall probability as the true level of the investment research report.
[0083] Using the above technical solution, by acquiring categorized investment research reports, which include both textual and numerical data, it is possible to determine the textual data of each bibliographic entry and the numerical data corresponding to reader evaluations in each investment research report. By importing the textual data of the investment research reports into a text feature model, the static probability of each preset level for the investment research report is determined. This text feature model is trained based on the textual data and actual levels of historical investment research reports, enabling the determination of the static probability of the investment research report at each preset level using textual data. Furthermore, by importing the numerical data of the investment research reports into a digital feature model... The model determines the dynamic level probability of each preset level in the investment research report. The digital feature model is trained based on historical investment research report data and actual levels, enabling the determination of the dynamic level probability of the investment research report at each preset level using digital data. By combining the static and dynamic level probabilities of each preset level, the comprehensive probability of each preset level of the investment research report is obtained, thus achieving a comprehensive probability of each preset level. Finally, by selecting the preset level with the highest comprehensive probability as the actual level of the investment research report, the true level of the evaluation report can be determined.
[0084] It should be noted that the investment research report analysis method, apparatus, equipment and storage medium disclosed herein can be used in the financial field, or in any field other than the financial field. The application field of the investment research report analysis method, apparatus, equipment and storage medium disclosed herein is not limited.
[0085] like Figure 3 The diagram illustrating the investment research report classification method is an example of this article. Before obtaining the classified investment research reports, the following steps are included:
[0086] Step 301: Preprocess each initial investment research report to obtain an initial investment research report with a specific theme.
[0087] This step specifically includes:
[0088] The initial investment research report is cleaned to obtain a cleaned preliminary investment research report. Data cleaning methods include removing symbols, stop words, and standardizing the corpus from the initial report. This approach avoids subsequent classification failures or poor classification results due to meaningless words.
[0089] The cleaned preliminary investment research report is transformed into a text matrix. This step converts textual data into a computer-readable numerical matrix, increasing the readability of the preliminary investment research report. In this paper, the text matrix can include a word matrix and a topic matrix. Text from different chapters of an investment research report is obtained, and these texts are combined using vector representations to obtain the word matrix. To avoid the influence of the locality of the word matrix on subsequent classification, this paper can also query keywords in the text and combine these keywords to obtain the topic matrix.
[0090] The text matrix is imported into the LDA model to obtain the probability distribution of each keyword in the preliminary investment research report;
[0091] In this step, the LDA model is a generative Bayesian probabilistic model with a three-layer structure including words, topics, and a corpus. The user imports the text matrix into the LDA model. In this text, each investment research report represents a probability distribution composed of several topics, and each topic represents a probability distribution composed of many words. Therefore, the LDA model's fitting result will present the core keywords and specific probabilities of each topic. The user can select at least one keyword with a high probability as the topic word for the initial investment research report.
[0092] Keywords with a higher probability distribution are used as the topics of the initial investment research report; in this article, each investment research report can have one, two, or three keywords.
[0093] Step 302: Classify the initial investment research reports with themes according to the topic to obtain several types of investment research reports.
[0094] In this step, investment research reports with similar themes can be categorized. For example, investment research report A focuses on military enterprises, report B on chip companies, report C on internet companies, report D on chip companies, report E on internet companies, and report F on military enterprises. Therefore, reports A and F can be grouped together; reports B and D together; and reports C and E together.
[0095] As an example of this article,
[0096] The textual data includes the production organization, author's name, publication region, publisher, or sales channel, etc.
[0097] The digital data includes publication date, reader ratings, classification number, or report page count, etc.
[0098] In this paper, reader ratings can be determined by the scores given by readers on different publishing channels. These reader ratings can change dynamically. The specific methods for obtaining reader ratings are conventional techniques for those skilled in the art, and will not be elaborated on here.
[0099] The preset levels include a first level, a second level, a third level, or a fourth level; for ease of explanation, the following content will use the first level, the second level, and the third level as examples.
[0100] To prevent authors from being bribed by companies to fabricate data, this paper proposes a blacklist mechanism. If an author is found to have engaged in fraudulent evaluations a certain number of times, their research reports will be suspended from rating and they will be added to the blacklist, thus protecting investors' rights. As an example, after obtaining the categorized research reports, the paper also includes:
[0101] The text data in each investment research report is matched with the information in the blacklist. If a match is found, the investment research report is filtered.
[0102] The blacklist stores the names of production institutions and / or authors whose dishonesty rate exceeds a predetermined value. In this document, the set value can be 90%, meaning that if the number of investment research reports published by the author is reported as false evaluations, reaching 90% of all evaluation reports published by the author, then the author's investment research reports will be filtered out.
[0103] As an embodiment of this article, step 202 imports the text data of the investment research report into a text feature model to determine the static level probability of each preset level of the investment research report. The text feature model is trained based on the text data of historical investment research reports and the actual levels, and includes:
[0104] In this paper, the text feature models include CharBERT (Character-aware Pre-trained Language Model), word2vec (word to vector model used to generate word vectors), and BP neural networks. The training set for this text feature model can be [(producing institution, author name), (level)], which can be used for training. The penultimate layer of the text feature model outputs the probability of each level. The final layer is a sigmoid layer, which assigns a value of 1 to the level with the highest probability and 0 to the remaining levels. Finally, the trained text feature model can output the level of the research report corresponding to the producing institution and author name.
[0105] When actually using the trained text feature model, the output of the penultimate layer can be obtained, which is the probability corresponding to each level, as shown in Table 1, which is a schematic table of output using the text feature model.
[0106] Table 1
[0107]
[0108]
[0109] As can be seen, after inputting the production organization and author name into the text feature model, the predicted probability corresponding to each level of the production organization and author name can be obtained. For example, when the production organization is A and the author name is B, the text feature model predicts that the production organization is A and the author name has a 52.18% probability of being a first-level investment research report, a 28.19% probability of being a second-level investment research report, and a 19.63% probability of being a third-level investment research report.
[0110] As an embodiment of this document, step 203 imports the digital data of the investment research report into a digital feature model to determine the dynamic level probability of each preset level of the investment research report. The digital feature model is trained based on the digital data of historical investment research reports and the actual levels, including...
[0111] In this paper, the digital feature model includes LSTM (Long-Short Term Memory) and CRF (Conditional Random Field). The training set for this digital feature model can be [(publication time, reader rating), (level)], which can be used for training. The output of the penultimate layer of the digital feature model is the probability of each level. The last layer of the digital feature model is a sigmoid layer, which assigns a value of 1 to the level with the highest probability and 0 to the other levels. Finally, the trained digital feature model can represent the level of the research report corresponding to the publication time and reader rating.
[0112] When actually using the trained digital feature model, the output of the penultimate layer can be obtained, which is the probability corresponding to each level, as shown in Table 2, which is a schematic table of output using the digital feature model.
[0113] Table 2
[0114]
[0115] As can be seen, after inputting the publication time and reader rating into the digital feature model, the predicted probability corresponding to each level of the publication time and reader rating can be obtained. For example, when the publication time is A and the reader rating is B, the digital feature model predicts that the report produced by organization A and author name has a 78.21% probability of being a first-level investment research report, a 12.02% probability of being a second-level investment research report, and a 9.77% probability of being a third-level investment research report.
[0116] As an embodiment of this article, the step of obtaining the comprehensive probability of each preset level of the investment research report based on the static level probability and the dynamic level probability of each preset level of the investment research report further includes:
[0117] The static and dynamic probabilities of the same level in the investment research report are weighted and averaged to obtain the comprehensive probability of the same level.
[0118] In this step, a weighted average is taken between the static grade probability and the dynamic grade probability of the same grade in an investment research report.
[0119] For example, it can be based on formula p ti =(1-f)p si +fp di , i∈1,2,3, f∈[0,1]
[0120] Where p ti Let f be the overall probability of the investment research report at level i, and p be the weight. si Let p be the static rank probability of the research report at rank i. di Let i be the dynamic level probability of the investment research report at level i, where i is the level.
[0121] In this paper, the weight can be 0.5. An example is given to illustrate the process of determining the overall probability of an investment research report. When the report's production institution is A and the author's name is B, the text feature model predicts that a report with production institution A and author name B has a 52.18% probability of being a first-tier report, a 28.19% probability of being a second-tier report, and a 19.63% probability of being a third-tier report. When the publication time is A and the reader rating is B, the numerical feature model predicts that a report with production institution A and author name B has a 78.21% probability of being a first-tier report, a 12.02% probability of being a second-tier report, and a 9.77% probability of being a third-tier report.
[0122] According to the formula, the overall probability of the first-level investment research report is 65.195%, the overall probability of the second-level report is 20.105%, and the overall probability of the third-level report is 14.7%.
[0123] As an example of this article, the preset level with the highest overall probability is selected as the true level of the investment research report.
[0124] In this step, the first level has the highest probability after calculation, so the first level can be taken as the true level of the investment research report.
[0125] like Figure 4 The schematic diagram shown is of an investment research report analysis device, including:
[0126] The acquisition unit 401 is used to acquire the categorized investment research report, wherein the investment research report includes textual data and numerical data.
[0127] The static grade determination unit 402 is used to import the text data of the investment research report into the text feature model to determine the static grade probability of each preset grade of the investment research report, wherein the text feature model is trained based on the text data of historical investment research reports and the actual grades.
[0128] The dynamic level determination unit 403 is used to import the digital data of the investment research report into the digital feature model to determine the dynamic level probability of each preset level of the investment research report. The digital feature model is trained based on the digital data of historical investment research reports and the actual levels.
[0129] The level integration unit 404 is used to obtain the integrated probability of each preset level of the investment research report based on the static level probability and the dynamic level probability of each preset level of the investment research report.
[0130] The rating determination unit 405 is used to select the preset rating with the highest comprehensive probability as the true rating of the investment research report.
[0131] Using the above technical solution, the acquisition unit can determine the textual data of the bibliographical items in each investment research report and the corresponding numerical data of reader evaluations; the static grade determination unit can determine the static grade probability of the investment research report at each preset grade based on the textual data; the dynamic grade determination unit can determine the dynamic grade probability of the investment research report at each preset grade based on the numerical data; the grade synthesis unit can synthesize the probability of each preset grade of the investment research report; and the grade determination unit can determine the true grade of the evaluation report.
[0132] like Figure 5As shown in the illustration, a computer device 502 provides an embodiment of this document. This computer device runs the research report analysis method described herein. The computer device 502 may include one or more processors 504, such as one or more central processing units (CPUs), each of which can implement one or more hardware threads. The computer device 502 may also include any memory 506 for storing information of any kind, such as code, settings, data, etc. Without limitation, for example, the memory 506 may include any type of RAM, any type of ROM, flash memory, hard disk, optical disk, etc. More generally, any memory can use any technology to store information. Furthermore, any memory can provide volatile or non-volatile retention of information. Furthermore, any memory can represent a fixed or removable component of the computer device 502. In one case, when the processor 504 executes associated instructions stored in any memory or combination of memories, the computer device 502 can perform any operation of the associated instructions. The computer device 502 also includes one or more drive mechanisms 508 for interacting with any memory, such as hard disk drive mechanisms, optical disk drive mechanisms, etc.
[0133] Computer device 502 may also include an input / output module 510 (I / O) for receiving various inputs (via input device 512) and providing various outputs (via output device 514). A specific output mechanism may include a presentation device 516 and an associated graphical user interface (GUI) 518. In other embodiments, the input / output module 510 (I / O), input device 512, and output device 514 may be omitted, and the device may function solely as a computer device within a network. Computer device 502 may also include one or more network interfaces 520 for exchanging data with other devices via one or more communication links 522. One or more communication buses 524 couple the components described above together.
[0134] Communication link 522 can be implemented in any way, such as via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, or any combination thereof. Communication link 522 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.
[0135] Corresponding to Figure 2 and Figure 3 In addition to the methods described above, this embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the above-described methods.
[0136] This embodiment also provides a computer-readable instruction, wherein when a processor executes the instruction, the program therein causes the processor to perform the following: Figure 2 and Figure 3 The method shown.
[0137] It should be understood that in the various embodiments of this document, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this document.
[0138] It should also be understood that, in the embodiments herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following associated objects have an "or" relationship.
[0139] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this document.
[0140] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0141] In the embodiments provided herein, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, devices, or units, or they may be electrical, mechanical, or other forms of connection.
[0142] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments described herein, depending on actual needs.
[0143] Furthermore, the functional units in the various embodiments of this document can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0144] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this paper, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this paper. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0145] This document uses specific embodiments to illustrate the principles and implementation methods of this document. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of this document. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this document. Therefore, the content of this specification should not be construed as a limitation of this document.
Claims
1. A method for analyzing investment research reports, characterized in that, include: Obtain the categorized investment research reports, which include both textual and numerical data; The text data of the investment research report is imported into the text feature model to determine the static level probability of each preset level of the investment research report. The text feature model is trained based on the text data of historical investment research reports and the actual levels. The digital data of the investment research report is imported into the digital feature model to determine the dynamic level probability of each preset level of the investment research report. The digital feature model is trained based on the digital data of historical investment research reports and the actual levels. Based on the static level probability and the dynamic level probability of each preset level in the investment research report, the comprehensive probability of each preset level in the investment research report is obtained. The preset level with the highest overall probability is selected as the true level of the investment research report. Before obtaining the categorized investment research reports, the following are included: Each initial investment research report is preprocessed to obtain an initial investment research report with a specific theme; Several initial investment research reports with themes are categorized according to their themes to obtain several types of investment research reports; The preprocessing of each initial investment research report to obtain a topic-based initial investment research report further includes: The initial investment research report is cleaned to obtain a cleaned preliminary investment research report; The cleaned preliminary investment research report is transformed into a text matrix; wherein, the text matrix includes a word matrix and a topic matrix; The text matrix is imported into the LDA model to obtain the probability distribution of each keyword in the preliminary investment research report; The keyword with the highest probability distribution will be used as the topic of the initial investment research report.
2. The investment research report analysis method according to claim 1, characterized in that, The step of cleaning the initial investment research report to obtain a cleaned preliminary investment research report further includes: Remove symbols, stop words, and standardized corpora from the initial investment research report to obtain the cleaned preliminary investment research report.
3. The investment research report analysis method according to claim 1, characterized in that, The text data includes the name of the producing organization and the author; The digital data includes publication date and reader ratings; The preset levels include a first level, a second level, and a third level.
4. The investment research report analysis method according to claim 3, characterized in that, After obtaining the categorized investment research reports, the following is also included: The text data in each investment research report is matched with the information in the blacklist. If a match is found, the investment research report is filtered. The blacklist stores the names of production organizations and / or authors whose dishonesty rate exceeds a predetermined value.
5. The investment research report analysis method according to claim 3, characterized in that, The step of obtaining the comprehensive probability of each preset level of the investment research report based on the static level probability and the dynamic level probability of each preset level of the investment research report further includes: The static and dynamic probabilities of the same level in the investment research report are weighted and averaged to obtain the comprehensive probability of the same level.
6. A research report analysis device, characterized in that, include: The acquisition unit is used to acquire the categorized investment research reports, wherein the investment research reports include textual data and numerical data; A static grade determination unit is used to import the text data of the investment research report into a text feature model to determine the static grade probability of each preset grade of the investment research report. The text feature model is trained based on the text data of historical investment research reports and the actual grades. The dynamic level determination unit is used to import the digital data of the investment research report into the digital feature model to determine the dynamic level probability of each preset level of the investment research report, wherein the digital feature model is trained based on the digital data of historical investment research reports and the actual levels. The grade integration unit is used to obtain the integrated probability of each preset grade of the investment research report based on the static grade probability and the dynamic grade probability of each preset grade. The rating determination unit is used to select the preset rating with the highest overall probability as the true rating of the investment research report. Prior to obtaining the categorized investment research reports, the process includes: Each initial investment research report is preprocessed to obtain an initial investment research report with a specific theme; Several initial investment research reports with themes are categorized according to their themes to obtain several types of investment research reports; The preprocessing of each initial investment research report to obtain a topic-based initial investment research report further includes: The initial investment research report is cleaned to obtain a cleaned preliminary investment research report; The cleaned preliminary investment research report is transformed into a text matrix; wherein, the text matrix includes a word matrix and a topic matrix; The text matrix is imported into the LDA model to obtain the probability distribution of each keyword in the preliminary investment research report; The keyword with the highest probability distribution will be used as the topic of the initial investment research report.
7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the investment research report analysis method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the investment research report analysis method as described in any one of claims 1-5.