Intelligent equity investment analysis method, device and equipment and storage medium
By combining text cleaning, word segmentation, intent classification, and large language models, the problem of quickly retrieving required information from massive amounts of equity investment information has been solved, realizing intelligent information retrieval and analysis, and improving efficiency and user satisfaction.
Patent Information
- Application Number
- CN202510923542.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies struggle to quickly retrieve required information from massive amounts of equity investment data, exhibiting low levels of intelligence, complex operation, and a high error rate.
By acquiring equity investment questions input by target users, the text is cleaned and then segmented using a word segmentation model to extract key information. In-depth analysis is then performed using an intent classification model, and effective information is retrieved from the equity investment knowledge base using a large language model. The results are then generated in the form of explanatory text, visual charts, or structured analysis reports.
It enables users to quickly find the information they need from a vast amount of equity investment data, improving business execution efficiency and user satisfaction. It also supports multiple application scenarios and covers a wider range of customers.
Smart Images

Figure CN120849449A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of equity investment technology, and in particular to intelligent analysis methods, devices, equipment and storage media for equity investment. Background Technology
[0002] Currently, the size and number of private equity funds are both experiencing continuous growth. Digitalization and platformization can be used to build equity investment information systems to provide users with the information and data they need. However, the diverse business entities, long capital chains, and complex investment relationships within the equity investment customer base result in a vast and complex amount of equity investment information stored in the system. This information is typically scattered across different modules and data sources, making it difficult for users to quickly find valuable information from the massive amounts of data, hindering accurate analysis and evaluation. Furthermore, information retrieval often requires setting complex filtering conditions, and exported data requires manual processing by the user, resulting in low levels of automation and increased operational complexity and error rates.
[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide an intelligent analysis method, device, equipment and storage medium for equity investment, which aims to solve the technical problem that existing technologies are unable to quickly retrieve the required information from massive equity investment information.
[0005] To achieve the above objectives, this application provides an intelligent analysis method for equity investment, the method comprising:
[0006] Obtain equity investment questions input by the target user, perform text cleaning on the equity investment questions, and obtain standardized text;
[0007] The standard text is segmented based on a word segmentation model to obtain segmented words, and key information is extracted from the segmented words.
[0008] Based on the intent classification model, the key information is analyzed in depth to obtain the question type and the true intent.
[0009] Based on the question type and the true intent, retrieve valid information from the equity investment knowledge base;
[0010] Based on a large language model, semantic understanding and contextual reasoning are performed on the effective information to generate answers to the equity investment questions. The answers are in the form of at least one of explanatory text, visual charts, and structured analysis reports.
[0011] In one embodiment, the step of segmenting the standard text based on a word segmentation model to obtain segmented words includes:
[0012] Based on internal business data and external publicly available data, multi-source corpus information is collected. The multi-source corpus information includes at least equity investment terminology, company names, fund product names, and investment process keywords.
[0013] Based on the multi-source corpus information, a dictionary specifically for equity investment is constructed;
[0014] The initial word segmentation model is trained based on the equity investment-specific dictionary to obtain the word segmentation model;
[0015] The standard text is input into the word segmentation model for word segmentation processing to obtain segmented words and their parts of speech. The entity types of the segmented words and the grammatical relationships between them are also labeled. The entity types include at least enterprise entities, industry terminology entities, monetary entities, and date entities.
[0016] In one embodiment, the step of extracting key information from the segmented words includes:
[0017] Based on the entity type of the segmented words, set corresponding initial extraction weights for the segmented words;
[0018] Based on the position and frequency of occurrence of the segmented words, the initial extraction weights of the segmented words are adjusted to obtain the extraction weights of the segmented words;
[0019] Based on the extraction weights of the segmented words, key information is determined in the segmented words, and the extraction weights of the key information are greater than a preset weight threshold.
[0020] In one embodiment, before the step of performing in-depth analysis of the key information based on the intent classification model to obtain the question type and true intent, the method further includes:
[0021] Obtain a preset language model, fine-tune the preset semantic model, and obtain an initial intent classification model;
[0022] Based on the training question and the corresponding training intent label, the initial intent classification model is trained to obtain training output data.
[0023] Based on the training output data, calculate the cross-entropy loss and focus loss;
[0024] The output loss is determined based on the cross-entropy loss and the focus loss.
[0025] When the output loss satisfies the preset convergence condition, the trained initial intent classification model is used as the intent classification model.
[0026] In one embodiment, the step of text cleaning the equity investment problem to obtain standardized text includes:
[0027] The equity investment question is matched with noise characters in a preset noise character rule library to identify matching noise characters. The matching noise characters in the equity investment question are then removed to obtain the noise-filtered text.
[0028] Remove redundant line breaks and tabs from the noise-filtered text, and compress consecutive whitespace characters in the noise-filtered text into single characters to obtain the formatted text;
[0029] The character encoding format of the formatted text is uniformly converted to the target encoding format, the character width of the formatted text is uniformly converted to the target character width, the date format of the formatted text is uniformly converted to the target date format, and the amount format of the formatted text is uniformly converted to the target amount format, resulting in a formatted text.
[0030] Based on the standard statement mapping table, the spoken sentence patterns of the standardized text are converted into corresponding standard sentence patterns to obtain the standardized text.
[0031] In one embodiment, after the step of performing semantic understanding and contextual reasoning on the effective information based on a large language model to generate an answer to the equity investment question, the method further includes:
[0032] Based on the target user's historical behavior data and historical transaction data, the target user's basic attribute tags, preference feature tags, and behavior prediction tags are determined, and corresponding recommendation weights are assigned to the target user's basic attribute tags, preference feature tags, and behavior prediction tags;
[0033] Based on the publication time of existing information, candidate information is determined from the existing information;
[0034] Calculate the first matching degree between the candidate information and the basic attribute label, the second matching degree between the candidate information and the preference feature label, and the third matching degree between the candidate information and the behavior prediction label, respectively;
[0035] Based on the first matching degree, the second matching degree, the third matching degree, and the recommendation weight, the recommendation score of the candidate information is determined;
[0036] Based on the recommendation scores of the candidate information, target preference information is pushed to the target user.
[0037] In one embodiment, after the step of performing semantic understanding and contextual reasoning on the effective information based on a large language model to generate an answer to the equity investment question, the method further includes:
[0038] Based on external public opinion data and the target users' historical behavior data and historical transaction data, risk indicators are calculated. The risk indicators include at least public opinion indicators, financial indicators, behavioral indicators, correlation indicators, and policy indicators.
[0039] The risk indicators are input into the risk assessment model for identification to obtain the current risk level and the confidence score corresponding to the current risk level.
[0040] When the confidence score is greater than or equal to the confidence threshold, a current risk management strategy is determined based on the current risk level, and the current risk level and the current risk management strategy are sent to the target user.
[0041] Furthermore, to achieve the above objectives, this application also proposes an intelligent analysis device for equity investment, which includes:
[0042] The intent analysis module is used to obtain equity investment questions input by the target user, and to perform text cleaning on the equity investment questions to obtain standardized text.
[0043] The intent analysis module is also used to perform word segmentation on the standard text based on the word segmentation model to obtain segmented words, and extract key information from the segmented words;
[0044] The intent analysis module is also used to perform in-depth analysis of the key information based on the intent classification model to obtain the question type and the true intent.
[0045] The information retrieval module is used to retrieve valid information from the equity investment knowledge base based on the question type and the true intent.
[0046] The intelligent answer module is used to perform semantic understanding and contextual reasoning on the effective information based on a large language model, and generate answers to the equity investment questions. The answers are in the form of at least one of explanatory text, visual charts, and structured analysis reports.
[0047] In addition, to achieve the above objectives, this application also proposes an intelligent analysis device for equity investment, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the intelligent analysis method for equity investment as described above.
[0048] In addition, to achieve the above objectives, the present invention also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the equity investment intelligent analysis method described above.
[0049] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the equity investment intelligent analysis method described above.
[0050] This application provides an intelligent analysis method for equity investment. It acquires equity investment questions input by target users, performs text cleaning on the questions to obtain standardized text, segments the standardized text using a word segmentation model to obtain segmented words, and extracts key information from these segmented words. It then performs in-depth analysis of the key information based on an intent classification model to determine the question type and true intent. Based on the question type and true intent, it retrieves effective information from an equity investment knowledge base. Finally, it performs semantic understanding and contextual reasoning on the effective information using a large language model to generate answers to the equity investment questions. The answers can take the form of at least one of explanatory text, visual charts, or structured analysis reports. This application, through a large model tool, automatically identifies user intent based on the user's input question and directly outputs accurate questions and answers. It is easy to operate, can quickly retrieve required information from massive amounts of equity investment data, and intelligently integrates it, improving business execution efficiency. Furthermore, it can meet various user needs through intelligent interaction, improving user satisfaction. It also supports multiple application scenarios, covering a wider customer base, and solves the technical problem of difficulty in quickly retrieving required information from massive amounts of equity investment data. Attached Figure Description
[0051] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0053] Figure 1 This is a flowchart illustrating an embodiment of the intelligent analysis method for equity investment in this application.
[0054] Figure 2 This is a flowchart illustrating Embodiment 2 of the intelligent analysis method for equity investment in this application;
[0055] Figure 3 This is a flowchart illustrating Embodiment 3 of the intelligent analysis method for equity investment in this application;
[0056] Figure 4 This is a schematic diagram of the overall architecture of the intelligent analysis method for equity investment provided in Embodiment 3 of this application;
[0057] Figure 5 A simplified flowchart illustrating the intelligent analysis method for equity investment provided in Embodiment 3 of this application;
[0058] Figure 6 This is a schematic diagram of the module structure of the intelligent analysis device for equity investment according to an embodiment of this application;
[0059] Figure 7 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the intelligent analysis method for equity investment in this application embodiment.
[0060] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0061] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0062] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0063] The main solution of this application embodiment is as follows: First, obtain the equity investment question input by the target user; second, perform text cleaning on the equity investment question to obtain standardized text; third, perform word segmentation on the standardized text based on a word segmentation model to obtain segmented words, and extract key information from the segmented words; fourth, perform in-depth analysis on the key information based on an intent classification model to obtain the question type and true intent; fifth, retrieve effective information from the equity investment knowledge base based on the question type and true intent; and sixth, perform semantic understanding and contextual reasoning on the effective information based on a large language model to generate an answer to the equity investment question. The answer can be in the form of at least one of explanatory text, visual charts, and structured analysis reports.
[0064] This application provides a solution that, through a large model tool, automatically identifies user intent based on user-input questions and directly outputs accurate answers. It is easy to operate, can quickly retrieve demand information from massive amounts of equity investment data, and intelligently integrates it to improve business execution efficiency. Furthermore, it can utilize intelligent interaction to meet various user needs, increasing user satisfaction. It also supports multiple application scenarios, covering a wider customer base, and solves the technical problem of difficulty in quickly retrieving demand information from massive amounts of equity investment data.
[0065] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or equity investment intelligent analysis device capable of performing the above functions. This embodiment does not specifically limit it in this regard. The following uses an equity investment intelligent analysis device as an example to describe this embodiment and the following embodiments.
[0066] This application provides an intelligent analysis method for equity investment, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the intelligent analysis method for equity investment in this application.
[0067] In this embodiment, the intelligent analysis method for equity investment includes steps S10 to S40:
[0068] Step S10: Obtain the equity investment question input by the target user, perform text cleaning on the equity investment question, and obtain standardized text;
[0069] It should be noted that the target users are those who need to obtain information related to equity investment, and the equity investment questions are those questions related to equity investment entered by the users, such as: analyzing the investment value of project A, providing investment advice for field B, and analyzing financing data. This embodiment does not make specific limitations on these.
[0070] Additionally, it should be noted that since equity investment questions are subjectively input by users, some wording may be non-standard, and there may be content that interferes with the final answer. Therefore, it is necessary to perform text cleaning on the user-input equity investment questions, automatically identify and remove noisy characters, standardize the encoding and colloquial expressions in the text, and unify the format of the input text. Standardized text refers to the text content obtained after text cleaning.
[0071] It is understood that this embodiment incorporates multilingual detection technology, supporting multiple language environments, such as Chinese and English, without specific limitations. Even if equity investment issues involve multiple languages, it can still accurately identify them.
[0072] In one feasible implementation, step S10 may include steps S101 to S104:
[0073] Step S101: Match the equity investment question with noise characters in the preset noise character rule library, determine the matching noise characters, and remove the matching noise characters in the equity investment question to obtain the noise-filtered text;
[0074] It should be noted that the preset noise character rule base is a pre-built database for storing noise characters. In this embodiment, noise characters include at least special characters, URLs (Uniform Resource Locators), and interjections. Special characters can be "#", "@", or "~", and interjections can be "ha", "oh", or "um," and can be flexibly adjusted according to actual needs. This embodiment does not impose specific limitations on these. Matching noise characters refers to characters in the equity investment problem that can match the noise characters.
[0075] Understandably, if there are matching noise characters in the equity investment question, these matching noise characters will be removed from the equity investment question. If there are no matching noise characters in the equity investment question, no processing will be performed, thus completing the noise filtering of the question text. The resulting text is the noise-filtered text.
[0076] For example, suppose the input equity investment question is "Help me check which projects Fund Manager C has invested in in the semiconductor field? And give me some policy advice as well~". If the noise characters include "#", "@", "~", "ha", "oh", and "um", then "ha" and "~" are the matching noise characters. Remove "ha" and "~" from the equity investment question, and the resulting noise-filtered text will be "Help me check which projects Fund Manager C has invested in in the semiconductor field? And give me some policy advice as well".
[0077] Step S102: Delete redundant line breaks and tabs in the noise-filtered text, and compress consecutive whitespace characters in the noise-filtered text into single characters to obtain formatted text;
[0078] It should be noted that a redundant newline character is an extra newline character (\n), a redundant tab character is an extra tab character (\t), consecutive whitespace characters are multiple consecutive whitespace characters, and a single character is a single space.
[0079] It is understood that this embodiment removes redundant line breaks and tabs, and compresses consecutive whitespace characters into a single space to ensure a clear text paragraph structure, thereby completing the standardization of text paragraphs and line breaks, and the resulting text is the formatted text.
[0080] Step S103: Convert the character encoding format of the formatted text to the target encoding format, convert the character width of the formatted text to the target character width, convert the date format of the formatted text to the target date format, and convert the amount format of the formatted text to the target amount format to obtain the formatted text;
[0081] It should be noted that the character encoding format refers to the encoding format used by the characters in the formatted text, and the target encoding format is the pre-defined standard encoding format. In this embodiment, the character encoding format in the formatted text needs to be unified to the target encoding format. The target character width is the pre-defined standard character width. If the target character width is 1 byte, then all full-width characters in the formatted text will be converted to half-width characters, thereby unifying the character width of the formatted text to 1 byte. If the target character width is 2 bytes, then all half-width characters in the formatted text will be converted to full-width characters, thereby unifying the character width of the formatted text to 2 bytes.
[0082] It is understood that the date format refers to the description format of the date, such as 2025.06.10, 2025-06-10, or June 10, 2025. The target date format is a pre-defined standard date format, such as 2025-06-10, and there is no specific limitation on it. This embodiment requires that all date formats in the format adjustment text be unified to the target date format. The amount format refers to the description format of the amount, such as 50,000 or 50,000. The target amount format is a pre-defined standard amount format, such as 50,000, and there is no specific limitation on it. This embodiment requires that all amount formats in the format adjustment text be unified to the target amount format. The text obtained by unifying the encoding format, character width, date format, and amount format is the format-unified text.
[0083] Step S104: Based on the standard statement mapping table, convert the spoken sentence patterns of the formatted text into the corresponding standard sentence patterns to obtain the standardized text.
[0084] It should be noted that colloquial sentence patterns refer to sentences expressed in everyday language, while standard sentence patterns refer to sentences that have been professionalized and standardized. The standard sentence mapping table stores colloquial sentence patterns and their corresponding standard sentence patterns. For example, the standard sentence pattern for the colloquial sentence pattern "invested in a certain company" is "to invest in a certain enterprise"; the standard sentence pattern for the colloquial sentence pattern "the money has arrived" is "funds have arrived"; the standard sentence pattern for "how about this...?" is "please analyze the investment value of..."; and the standard sentence pattern for the colloquial sentence pattern "help me take a look..." is "query...". This embodiment does not impose specific limitations on these. After converting the colloquial sentence patterns into their corresponding standard sentence patterns, the resulting text is the standardized text.
[0085] For example, suppose the standardized text is “Help me check which projects Fund Manager C has invested in in the semiconductor field? And give me some policy suggestions.” Converting the colloquial sentence into the corresponding standard sentence, the resulting standardized text is “Inquire about Fund Manager C’s investment projects in the semiconductor industry and provide relevant policy suggestions.”
[0086] Furthermore, the preset noise character rule library can be continuously updated based on user feedback. For example, if a user frequently inputs "@...", "@" can be removed from the preset noise character rule library to avoid incorrect filtering. In addition, special characters or formats that cannot be recognized can be marked, for example, using "[UNK]" to mark unknown characters and unknown formats to avoid overall text processing failure due to local noise processing failure.
[0087] Step S20: The standard text is segmented based on the word segmentation model to obtain segmented words, and key information is extracted from the segmented words;
[0088] It should be noted that word segmentation refers to the series of words obtained after word segmentation processing. The word segmentation model is a pre-trained model used for word segmentation. This embodiment utilizes the word segmentation model to achieve accurate word segmentation and syntactic annotation for subsequent semantic understanding.
[0089] In one feasible implementation, the step of segmenting the standardized text based on a word segmentation model to obtain segmented words may include steps S201 to S204:
[0090] Step S201: Based on internal business data and external publicly available data, collect multi-source corpus information, which includes at least equity investment terminology, company names, fund product names, and investment process keywords.
[0091] It should be noted that internal business data includes at least equity investment contracts, due diligence reports, and transaction records, while publicly available external data includes at least industry research reports, policy documents, and financial news. Multi-source corpus information includes at least equity investment terminology, company names (names of invested companies), fund product names, and investment process keywords (e.g., "due diligence," "fundraising," "exit"). This multi-source corpus information, including equity investment terminology, company names, fund product names, and investment process keywords, is collected by integrating internal business data and publicly available external data.
[0092] Step S202: Based on the multi-source corpus information, construct a dictionary specifically for equity investment;
[0093] Understandably, a dictionary specifically for equity investment is a corpus dictionary dedicated to the field of equity investment, constructed based on collected multi-source corpus information.
[0094] Step S203: Train the initial word segmentation model based on the equity investment-specific dictionary to obtain the word segmentation model;
[0095] It should be noted that the initial word segmentation model is the traditional word segmentation model, which can be flexibly selected according to actual needs. This embodiment does not impose specific limitations on it. Based on the initial word segmentation model, further training is performed using a dedicated equity investment dictionary to obtain the final word segmentation model.
[0096] Step S204: Input the standard text into the word segmentation model for word segmentation processing to obtain segmented words and their parts of speech, and label the entity types of the segmented words and the grammatical relationships between them. The entity types include at least enterprise entities, industry terminology entities, monetary entities, and date entities.
[0097] It should be noted that the word segmentation model is used to divide the standard text into the smallest semantic units, i.e., segmented words. For example, after segmenting "Query fund manager C's investment projects in the semiconductor industry", the resulting segmented words are "query", "fund manager C", "semiconductor industry", and "investment projects". Parts of speech typically include nouns, verbs, and adjectives. For example, "fund" is a noun, and "query" is a verb. If a segmented word contains multiple nouns / verbs / adjectives, the most important one can be marked as the core noun / core verb / core adjective. The word segmentation model is used to determine all segmented words and the part of speech of each segmented word.
[0098] It is understandable that the grammatical relationships between segmented words typically include subject-predicate relationships, verb-object relationships, etc. In this embodiment, the segmented words are marked as subject, predicate, object, and attributive according to the grammatical relationships. For example, assuming the segmented words are "fund manager C", "investment", "semiconductor industry", and "company A", then the part of speech of "fund manager C" is marked as subject, the part of speech of "investment" is marked as predicate, the part of speech of "company A" is marked as object, and the part of speech of "semiconductor industry" is marked as attributive.
[0099] It should be understood that entity types include at least corporate entities, industry terminology entities, monetary entities, and date entities. Named Entity Recognition (NER) technology identifies specific corporate names, industry terms, monetary amounts, dates, and other entities. For example, the entity type for "50 million yuan" is marked as a monetary entity, "2025" as a date entity, and "fund manager" as an industry terminology entity.
[0100] Furthermore, the step of extracting key information from the segmented words includes: setting corresponding initial extraction weights for the segmented words based on their entity types; adjusting the initial extraction weights of the segmented words based on their positions and frequencies of occurrence to obtain the extraction weights of the segmented words; and determining key information from the segmented words based on their extraction weights, wherein the extraction weight of the key information is greater than a preset weight threshold.
[0101] It should be noted that key information refers to the most critical words extracted from the segmented words, and the initial extraction weight is the weight initially set when extracting key information. In this embodiment, an initial extraction weight is first set for each segmented word according to its entity type. Generally, the initial extraction weight of industry terminology entities is usually greater than that of other entities. For example, the initial extraction weight of industry terminology entities is set to 3.5. The specific value can be set according to the actual situation and is not specifically limited.
[0102] Understandably, the position of a segmented word refers to its location within the standard text, such as the beginning, middle, or end of a sentence. Typically, the initial extraction weight of segmented words located at the beginning and end of a sentence can be increased. The frequency of a segmented word refers to the number of times it appears in the entire standard text. Generally, the higher the frequency, the greater the weight. Therefore, the initial extraction weight of segmented words with a high frequency (compared to a set threshold) can be increased. The extraction weight is the final weight used when extracting key information. After adjusting the initial extraction weight, the final extraction weight is obtained. The preset weight threshold is the pre-set threshold for the extraction weight. Based on the preset weight threshold and the extraction weight of the segmented words, suitable key information is selected. In practice, if the extraction weight of a segmented word is greater than the preset weight threshold, then that segmented word is considered key information.
[0103] In this embodiment, a dictionary specific to the equity investment field is used as a knowledge base to analyze the input text and extract representative key information. This step is the foundation for subsequent intent recognition and directly affects the accuracy of the entire semantic understanding.
[0104] Step S30: Perform in-depth analysis of the key information based on the intent classification model to obtain the question type and the true intent;
[0105] It should be noted that the intent classification model is a model used to identify intents. This embodiment introduces a fully trained intent classification model, which accurately determines the user's question type and true intent through in-depth analysis of key input information.
[0106] In one feasible implementation, before step S30, the following steps may be included: obtaining a preset language model, fine-tuning the preset semantic model to obtain an initial intent classification model; training the initial intent classification model based on the training question and the training intent label corresponding to the training question to obtain training output data; calculating cross-entropy loss and focus loss based on the training output data; determining the output loss based on the cross-entropy loss and the focus loss; and using the trained initial intent classification model as the intent classification model when the output loss satisfies a preset convergence condition.
[0107] It should be noted that the preset language model is an existing large language model, which can be selected according to actual needs. This embodiment designs a fine-tuning layer, which typically contains two fully connected layers. The fine-tuning layer is used to fine-tune the preset language model to construct an initial intent classification model, i.e., the initial intent classification model. The training question is the input data during training. The training question is assigned corresponding labels, i.e., training intent labels, to describe the intent of the question, such as: "investment project query", "risk assessment", "policy interpretation", "data statistics". The initial intent classification model is trained according to the training question and training intent labels. The training output data is the actual intent labels output during training.
[0108] Understandably, based on the training output data, the cross-entropy loss and focus loss are calculated, and then the overall loss, i.e., the output loss, is calculated using the cross-entropy loss and focus loss. The calculation relationship is shown below:
[0109] L=α·l1+β·l2
[0110] In the formula, L represents the output loss, l1 represents the cross-entropy loss, l2 represents the focus loss, and α and β represent the weights of the cross-entropy loss and the focus loss, respectively.
[0111] It should be understood that the preset convergence condition is the condition for model convergence. For example, the output loss is less than a set loss threshold, or a maximum number of training iterations can be set. This embodiment does not specifically limit this. Generally speaking, when the output loss meets the preset convergence condition, the training can be considered complete, and the initial intent classification model that has been trained can be used as an intent classification model.
[0112] Step S40: Based on the question type and the true intent, retrieve valid information from the equity investment knowledge base;
[0113] It should be noted that valid information refers to valuable information retrieved according to the question type and the true intent, which can meet the user's needs. Based on the intent parsing results, the corresponding data retrieval process is triggered to retrieve relevant terms and definitions from the structured equity investment knowledge base, and to call the pre-integrated Application Programming Interface (API) to obtain data. The API interfaces include customer information query interface, investment and financing KYC (know-your-customer) query interface, news information interface, IPO (Initial Public Offering) information interface, etc.
[0114] Step S50: Based on the large language model, perform semantic understanding and contextual reasoning on the effective information to generate an answer to the equity investment question. The answer may be in the form of at least one of explanatory text, visual charts, or structured analysis reports.
[0115] Understandably, in terms of text generation, this embodiment relies on a large language model to automatically generate intelligent answers through semantic understanding and contextual reasoning. Answers can be output as high-quality explanatory text to meet diverse expression needs; or complex data information can be transformed into intuitive visualizations, allowing answers to be output in the form of visual charts, which can help users quickly understand the data's meaning; for complex questions, structured analysis reports can be generated, and answers can be output in the form of structured analysis reports. In specific implementation, if the user does not specify a particular requirement, the answer can be output in any form among explanatory text, visual charts, and structured analysis reports; if the user specifies a particular output requirement, the appropriate form will be selected to output the answer according to the user's request.
[0116] This embodiment provides an intelligent analysis method for equity investment. It acquires equity investment questions input by target users, performs text cleaning on the questions to obtain standardized text, segments the standardized text using a word segmentation model to obtain segmented words, and extracts key information from these segmented words. It then performs in-depth analysis of the key information based on an intent classification model to determine the question type and true intent. Based on the question type and true intent, it retrieves effective information from an equity investment knowledge base. Finally, it performs semantic understanding and contextual reasoning on the effective information using a large language model to generate an answer to the equity investment question. The answer can take the form of at least one of explanatory text, visual charts, or a structured analysis report. This embodiment automatically identifies user intent based on the user's input question using a large model tool, directly outputting accurate questions and answers. It is easy to operate, can quickly retrieve required information from massive amounts of equity investment information, and intelligently integrates it, improving business execution efficiency. Furthermore, it can meet various user needs through intelligent interaction, improving user satisfaction, and supports multiple application scenarios, covering a wider customer base.
[0117] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Step S50 may be followed by steps S61 to S65:
[0118] Step S61: Based on the target user's historical behavior data and historical transaction data, determine the target user's basic attribute tags, preference feature tags, and behavior prediction tags, and assign corresponding recommendation weights to the target user's basic attribute tags, preference feature tags, and behavior prediction tags;
[0119] It should be noted that historical behavioral data refers to the target user's LP (Limited Partner) and GP (General Partner) behavioral data in the equity investment system, such as browsing history and favorite tags. This embodiment does not impose specific limitations on this. Historical transaction data refers to transaction data in the historical cooperation information database, such as investment amount and exit returns. This embodiment does not impose specific limitations on this.
[0120] Additionally, it should be noted that this embodiment sets up three levels of tags for target users based on historical behavioral data and historical transaction data: basic attribute tags, preference feature tags, and behavior prediction tags. Basic attribute tags include at least static and dynamic tags. Static tags can be user roles (e.g., LP, GP, account manager), region, and management scale. Dynamic tags can be the number of logins in the past 30 days, information reading time, and search keyword frequency. Preference feature tags include at least domain preference tags and format preference tags. Domain preference tags can be determined based on the distribution of user-read information topics, for example, "new energy" accounting for 40%. Format preference tags can be determined by statistically analyzing the proportion of user consumption of text / images / videos / reports. Behavior prediction tags can be determined based on the user's potential needs predicted by a prediction model. This embodiment achieves accurate user profiling by constructing a three-level tag system.
[0121] Step S62: Based on the publication time of existing information, determine candidate information from the existing information;
[0122] It should be noted that "existing information" refers to all existing information. Since this information may be from a long time ago and its timeliness cannot be guaranteed, this embodiment uses the publication time to perform a preliminary screening of existing information, selecting recent information as candidate information. Generally, information from the past week or the past 10 days can be screened; this embodiment does not make specific limitations in this regard.
[0123] Step S63: Calculate the first matching degree between the candidate information and the basic attribute label, the second matching degree between the candidate information and the preference feature label, and the third matching degree between the candidate information and the behavior prediction label, respectively.
[0124] It is understandable that the first degree of matching between candidate information and basic attribute labels represents the degree of matching between candidate information and basic attribute labels; the second degree of matching between candidate information and preference feature labels represents the degree of matching between candidate information and preference feature labels; and the third degree of matching between candidate information and behavior prediction labels represents the degree of matching between candidate information and behavior prediction labels. This embodiment needs to calculate the degree of matching between candidate information and basic attribute labels, the degree of matching between candidate information and preference feature labels, and the degree of matching between candidate information and behavior prediction labels.
[0125] Step S64: Determine the recommendation score of the candidate information based on the first matching degree, the second matching degree, the third matching degree, and the recommendation weight;
[0126] It should be noted that the basic attribute labels, preference feature labels, and behavior prediction labels are each assigned corresponding weights, i.e., recommendation weights. The recommendation score for candidate information is calculated by comprehensively considering the first matching degree, second matching degree, third matching degree, and recommendation weights, as shown in the following formula:
[0127] y = k1·x1 + k2·x2 + k3·x3
[0128] In the formula, y represents the recommendation score, k1, k2 and k3 represent the first matching degree, the second matching degree and the third matching degree, respectively, and x1, x2 and x3 represent the recommendation weights of the basic attribute label, the preference feature label and the behavior prediction label, respectively.
[0129] Step S65: Based on the recommendation score of the candidate information, push target preference information to the target user.
[0130] It is understood that target preference information is information determined according to user behavior and preferences. In this embodiment, candidate information is sorted according to recommendation scores, and the top 20% is recommended to the target user as target preference information.
[0131] Furthermore, industry dynamics and external public opinion data are collected to calculate public opinion sentiment vectors; based on the public opinion sentiment vectors, the public opinion sentiment trend is determined; based on the historical behavior data of the target user and the public opinion sentiment trend, the quality level of the target user is determined; when the quality level of the target user is greater than a preset quality threshold, a list of potential customers is generated based on the target user and similar users of the target user.
[0132] It should be noted that the public opinion sentiment vector, or three-dimensional sentiment vector, is denoted as [positive probability, negative probability, neutral probability]. It can be determined by statistically analyzing the number of positive, negative, and neutral keywords; a higher number corresponds to a higher probability. Alternatively, it can be determined by constructing a sentiment analysis model. This embodiment does not specifically limit this approach. Next, the public opinion sentiment tendency is determined based on the sentiment vector. If the positive probability is significantly greater than the negative and neutral probabilities, the public opinion sentiment tendency is positive. If the negative probability is significantly greater than the positive and neutral probabilities, the public opinion sentiment tendency is negative. If the neutral probability is significantly greater than the negative and positive probabilities, the public opinion sentiment tendency is neutral. If the neutral and positive probabilities are significantly greater than the negative probability, and the neutral probability is greater than the positive probability, the public opinion sentiment tendency is neutral to slightly positive. This process continues until the final public opinion sentiment tendency is determined.
[0133] Understandably, if a target user exhibits a cooperative tendency when public sentiment is positive or neutral, and withdraws at an appropriate time when public sentiment is negative, then the target user is considered to have a high quality level. The preset quality threshold is a pre-defined threshold for quality. If a target user's quality level exceeds the preset quality threshold, the user is considered a high-quality user, and users similar to this high-quality user are selected to generate a list of potential customers.
[0134] This embodiment provides an intelligent analysis method for equity investment. Based on the target user's historical behavior data and historical transaction data, it determines the target user's basic attribute tags, preference feature tags, and behavior prediction tags, and assigns corresponding recommendation weights. Based on the publication time of existing information, it identifies candidate information from the existing information. It calculates the first matching degree between the candidate information and the basic attribute tags, the second matching degree between the candidate information and the preference feature tags, and the third matching degree between the candidate information and the behavior prediction tags. Based on the first matching degree, the second matching degree, the third matching degree, and the recommendation weights, it determines the recommendation score of the candidate information. Based on the recommendation score of the candidate information, it pushes target preference information to the target user. This embodiment uses a large model tool to automatically identify the user's intent based on the user's input question and directly output accurate question and answer. It is easy to operate, can quickly find the required information from massive equity investment information, and intelligently integrate it to improve business execution efficiency. It can also meet various user needs through intelligent interaction, improve user satisfaction, and support multiple application scenarios, covering a wider customer group.
[0135] Based on the first embodiment of this application, in the third embodiment of this application, the same or similar content as the above embodiment can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 Step S50 may be followed by steps S61' to S63':
[0136] Step S61': Based on external public opinion data and the target user's historical behavior data and historical transaction data, calculate risk indicators. The risk indicators include at least public opinion indicators, financial indicators, behavioral indicators, correlation indicators, and policy indicators.
[0137] It should be noted that, in this embodiment, risk indicators include at least public opinion indicators, financial indicators, behavioral indicators, related indicators, and policy indicators. Among them, public opinion indicators can use public sentiment tendencies, financial indicators can use debt-to-equity ratio and asset-current ratio, behavioral indicators can use equity change frequency and key employee turnover rate, policy indicators can use compliance matching degree, and related indicators can use the probability of risk transmission between upstream and downstream enterprises.
[0138] Step S62': Input the risk indicator into the risk assessment model for identification to obtain the current risk level and the confidence score corresponding to the current risk level;
[0139] It should be noted that the risk levels can be divided into 1-5, where levels 1-2 are low risk, level 3 is medium risk, and levels 4-5 are high risk. The risk assessment model can be trained using the XGBoost model, but this embodiment does not impose specific limitations on it.
[0140] Understandably, risk indicators are input into the risk assessment model, which outputs the current risk level along with a confidence score, for example: the risk level is 4 and the confidence level is 91%.
[0141] Step S63': When the confidence score is greater than or equal to the confidence threshold, determine the current risk management strategy based on the current risk level, and send the current risk level and the current risk management strategy to the target user.
[0142] It should be noted that the confidence threshold is a pre-set threshold for the confidence score. A confidence score greater than or equal to the confidence threshold indicates that the result is reliable. The current risk management strategy refers to the measures taken to address the current risk level. If the current risk level is low (level 1-2), the current risk management strategy may be continuous monitoring and periodic tracking. If the current risk level is medium (level 3), the current risk management strategy may be early warning and portfolio adjustments. If the current risk level is high (level 4-5), the current risk management strategy may be activating emergency mechanisms and asset preservation.
[0143] In the specific implementation, refer to Figure 4 The system can be designed with four core modules: **Intelligent Information Recommendation and Customer Screening Module:** This module integrates internal and external equity investment information and historical cooperation information, using large-scale model fine-tuning technology to: accurately push relevant information intelligently based on user history and preferences; and screen potential customers for high-quality investment targets based on public opinion analysis and behavioral pattern recognition. **Risk Management Enhancement Module:** By integrating internal and external public opinion data and risk assessment models, it achieves negative public opinion monitoring and intelligent project risk identification based on semantic analysis; and generates a closed-loop risk management system with personalized risk warning reports and response suggestions. **Policy Knowledge Base Construction Module:** This module utilizes the powerful text processing capabilities of the large-scale model, combined with internal and external document requirements, to automatically build a dynamically updated knowledge base system; it features policy retrieval and interpretation functions; and generates standardized operation guides. **Platform Integration and Application Expansion Module:** As an important component of the investment and financing big data platform, it achieves seamless integration with existing business systems; supports core business scenarios such as investment promotion and fundraising; and promotes platform-based operation of equity business clients.
[0144] This embodiment provides an intelligent analysis method for equity investment. Based on external public opinion data and the target user's historical behavior and transaction data, risk indicators are calculated. These risk indicators are then input into a risk assessment model for identification, yielding the current risk level and its corresponding confidence score. When the confidence score is greater than or equal to a confidence threshold, a current risk management strategy is determined based on the current risk level, and this strategy is sent to the target user. This embodiment utilizes a large-scale model tool to automatically identify user intent based on input questions, directly outputting accurate answers. It is user-friendly, allowing for rapid retrieval of required information from massive amounts of equity investment data, intelligent integration, and improved business execution efficiency. Furthermore, it leverages intelligent interaction to meet diverse user needs, increasing user satisfaction, and supports multiple application scenarios, covering a wider customer base.
[0145] For example, to help understand the implementation process of the intelligent equity investment analysis method obtained by combining this embodiment with the above-described embodiment three, please refer to... Figure 5 , Figure 5 A simplified flowchart of an intelligent analysis method for equity investment is provided, specifically:
[0146] Question Reception and Preprocessing: The system receives natural language question input from users via an API interface. Multilingual detection technology is introduced, supporting multiple language environments including Chinese and English. Secondly, a text cleaning algorithm automatically identifies and processes noisy characters such as special characters and URLs, standardizing text encoding and colloquial expressions to ensure consistent input text format. Finally, a word segmentation system trained on a corpus from the equity investment field achieves accurate word segmentation and syntactic annotation.
[0147] Semantic understanding and intent recognition: The first step is keyword extraction, which uses a dictionary specific to the equity investment field as a knowledge base to analyze the input text and extract representative key information. The second step is intent parsing, which introduces a well-trained intent classification model. Through in-depth analysis of the input information, it accurately determines the user's question type and true intent.
[0148] Data retrieval and integration: Based on the results of intent parsing, the corresponding data retrieval process is triggered to retrieve relevant terms and definitions from the structured knowledge base and call pre-integrated APIs to obtain data, such as customer information query interface, investment and financing KYC query interface, news information interface, IPO information interface, etc.
[0149] Content generation and formatted output: First, in terms of text generation, relying on a large language model, it automatically generates high-quality explanatory text through semantic understanding and contextual reasoning to meet diverse expression needs. Second, in the field of data visualization, it has the ability to generate a rich variety of charts and graphs, transforming complex data information into intuitive visual displays. Third, for complex problems, it can generate structured analysis reports.
[0150] This application also provides an intelligent analysis device for equity investment; please refer to [reference needed]. Figure 6 The intelligent analysis device for equity investment includes:
[0151] The intent analysis module 10 is used to obtain the equity investment question input by the target user, and to perform text cleaning on the equity investment question to obtain standardized text;
[0152] The intent analysis module 10 is also used to perform word segmentation on the standard text based on the word segmentation model to obtain segmented words, and extract key information from the segmented words;
[0153] The intent analysis module 10 is also used to perform in-depth analysis of the key information based on the intent classification model to obtain the question type and the true intent.
[0154] Information retrieval module 20 is used to retrieve valid information from the equity investment knowledge base based on the question type and the true intent.
[0155] The intelligent answer module 30 is used to perform semantic understanding and contextual reasoning on the effective information based on a large language model to generate an answer to the equity investment question. The answer can be in the form of at least one of explanatory text, visual charts, and structured analysis reports.
[0156] In one feasible implementation, the intent analysis module 10 is further used to collect multi-source corpus information based on internal business data and external public data. The multi-source corpus information includes at least equity investment terms, company names, fund product names, and investment process keywords.
[0157] Based on the multi-source corpus information, a dictionary specifically for equity investment is constructed;
[0158] The initial word segmentation model is trained based on the equity investment-specific dictionary to obtain the word segmentation model;
[0159] The standard text is input into the word segmentation model for word segmentation processing to obtain segmented words and their parts of speech. The entity types of the segmented words and the grammatical relationships between them are also labeled. The entity types include at least enterprise entities, industry terminology entities, monetary entities, and date entities.
[0160] In one feasible implementation, the intent analysis module 10 is further configured to set corresponding initial extraction weights for the segmented words based on the entity type of the segmented words;
[0161] Based on the position and frequency of occurrence of the segmented words, the initial extraction weights of the segmented words are adjusted to obtain the extraction weights of the segmented words;
[0162] Based on the extraction weights of the segmented words, key information is determined in the segmented words, and the extraction weights of the key information are greater than a preset weight threshold.
[0163] In one feasible implementation, the intent analysis module 10 is further configured to acquire a preset language model, fine-tune the preset semantic model, and obtain an initial intent classification model.
[0164] Based on the training question and the corresponding training intent label, the initial intent classification model is trained to obtain training output data.
[0165] Based on the training output data, calculate the cross-entropy loss and focus loss;
[0166] The output loss is determined based on the cross-entropy loss and the focus loss.
[0167] When the output loss satisfies the preset convergence condition, the trained initial intent classification model is used as the intent classification model.
[0168] In one feasible implementation, the intent analysis module 10 is further configured to match the equity investment question with noise characters in a preset noise character rule library, determine the matching noise characters, and remove the matching noise characters in the equity investment question to obtain noise-filtered text;
[0169] Remove redundant line breaks and tabs from the noise-filtered text, and compress consecutive whitespace characters in the noise-filtered text into single characters to obtain the formatted text;
[0170] The character encoding format of the formatted text is uniformly converted to the target encoding format, the character width of the formatted text is uniformly converted to the target character width, the date format of the formatted text is uniformly converted to the target date format, and the amount format of the formatted text is uniformly converted to the target amount format, resulting in a formatted text.
[0171] Based on the standard statement mapping table, the spoken sentence patterns of the standardized text are converted into corresponding standard sentence patterns to obtain the standardized text.
[0172] In one feasible implementation, the intelligent recommendation module 40 is used to determine the target user's basic attribute tags, preference feature tags, and behavior prediction tags based on the target user's historical behavior data and historical transaction data, and to assign corresponding recommendation weights to the target user's basic attribute tags, preference feature tags, and behavior prediction tags;
[0173] Based on the publication time of existing information, candidate information is determined from the existing information;
[0174] Calculate the first matching degree between the candidate information and the basic attribute label, the second matching degree between the candidate information and the preference feature label, and the third matching degree between the candidate information and the behavior prediction label, respectively;
[0175] Based on the first matching degree, the second matching degree, the third matching degree, and the recommendation weight, the recommendation score of the candidate information is determined;
[0176] Based on the recommendation scores of the candidate information, target preference information is pushed to the target user.
[0177] In one feasible implementation, the risk management module 50 is used to calculate risk indicators based on external public opinion data and the target user's historical behavior data and historical transaction data. The risk indicators include at least public opinion indicators, financial indicators, behavioral indicators, correlation indicators and policy indicators.
[0178] The risk indicators are input into the risk assessment model for identification to obtain the current risk level and the confidence score corresponding to the current risk level.
[0179] When the confidence score is greater than or equal to the confidence threshold, a current risk management strategy is determined based on the current risk level, and the current risk level and the current risk management strategy are sent to the target user.
[0180] The equity investment intelligent analysis device provided in this application, employing the equity investment intelligent analysis method in the above embodiments, can solve the technical problem of difficulty in quickly finding required information from massive amounts of equity investment information. Compared with the prior art, the beneficial effects of the equity investment intelligent analysis device provided in this application are the same as those of the equity investment intelligent analysis method provided in the above embodiments, and other technical features in the equity investment intelligent analysis device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0181] All user-related data involved in this application (e.g., historical behavior data and historical transaction data) was obtained with the user's permission or consent. In other words, when this application is applied to specific products or technologies, user permission is required to acquire and process the relevant data, and the processing of the data must comply with the relevant laws, regulations, and regulatory standards of the relevant countries and regions. For example, when it is necessary to obtain a user's browsing history, a prompt to obtain browsing history can be displayed on the user's terminal. After receiving confirmation from the user regarding the prompt, the terminal can obtain the user's browsing history.
[0182] This application provides an intelligent analysis device for equity investment, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the intelligent analysis method for equity investment in Embodiment 1 described above.
[0183] The following is for reference. Figure 7 The diagram illustrates a structural schematic of an intelligent equity investment analysis device suitable for implementing embodiments of this application. The intelligent equity investment analysis device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The equity investment intelligent analysis device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0184] like Figure 7As shown, the equity investment intelligent analysis device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the equity investment intelligent analysis device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touch screens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the equity investment intelligent analysis device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows equity investment intelligent analysis devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0185] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0186] The intelligent equity investment analysis device provided in this application, employing the intelligent equity investment analysis method described in the above embodiments, can solve the technical problem of difficulty in quickly retrieving required information from massive amounts of equity investment information. Compared with the prior art, the beneficial effects of the intelligent equity investment analysis device provided in this application are the same as those of the intelligent equity investment analysis method provided in the above embodiments, and other technical features of this intelligent equity investment analysis device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0187] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0188] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0189] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the equity investment intelligent analysis method in the above embodiments.
[0190] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0191] The aforementioned computer-readable storage medium may be included in the equity investment intelligent analysis device; or it may exist independently and not be assembled into the equity investment intelligent analysis device.
[0192] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the intelligent equity investment analysis device, the intelligent equity investment analysis device performs the following actions: acquires an equity investment question input by a target user; performs text cleaning on the equity investment question to obtain standardized text; performs word segmentation on the standardized text based on a word segmentation model to obtain segmented words, and extracts key information from the segmented words; performs in-depth analysis on the key information based on an intent classification model to obtain the question type and true intent; retrieves valid information from the equity investment knowledge base based on the question type and true intent; and performs semantic understanding and contextual reasoning on the valid information based on a large language model to generate an answer to the equity investment question. The answer may be in the form of at least one of the following: explanatory text, visual charts, or a structured analysis report.
[0193] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0194] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0195] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0196] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described intelligent equity investment analysis method. This solves the technical problem of difficulty in quickly retrieving required information from massive amounts of equity investment data. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the intelligent equity investment analysis method provided in the above embodiments, and will not be elaborated upon here.
[0197] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the equity investment intelligent analysis method described above.
[0198] The computer program product provided in this application can solve the technical problem of difficulty in quickly finding the required information from massive amounts of equity investment information. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the intelligent equity investment analysis method provided in the above embodiments, and will not be repeated here.
[0199] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. An intelligent analysis method for equity investment, characterized in that, The method includes: Obtain equity investment questions input by the target user, perform text cleaning on the equity investment questions, and obtain standardized text; The standard text is segmented based on a word segmentation model to obtain segmented words, and key information is extracted from the segmented words. Based on the intent classification model, the key information is analyzed in depth to obtain the question type and the true intent. Based on the question type and the true intent, retrieve valid information from the equity investment knowledge base; Based on a large language model, semantic understanding and contextual reasoning are performed on the effective information to generate answers to the equity investment questions. The answers are in the form of at least one of explanatory text, visual charts, and structured analysis reports.
2. The method as described in claim 1, characterized in that, The step of segmenting the standardized text using a word segmentation model to obtain segmented words includes: Based on internal business data and external publicly available data, multi-source corpus information is collected. The multi-source corpus information includes at least equity investment terminology, company names, fund product names, and investment process keywords. Based on the multi-source corpus information, a dictionary specifically for equity investment is constructed; The initial word segmentation model is trained based on the equity investment-specific dictionary to obtain the word segmentation model; The standard text is input into the word segmentation model for word segmentation processing to obtain segmented words and their parts of speech. The entity types of the segmented words and the grammatical relationships between them are also labeled. The entity types include at least enterprise entities, industry terminology entities, monetary entities, and date entities.
3. The method as described in claim 2, characterized in that, The step of extracting key information from the segmented words includes: Based on the entity type of the segmented words, set corresponding initial extraction weights for the segmented words; Based on the position and frequency of occurrence of the segmented words, the initial extraction weights of the segmented words are adjusted to obtain the extraction weights of the segmented words; Based on the extraction weights of the segmented words, key information is determined in the segmented words, and the extraction weights of the key information are greater than a preset weight threshold.
4. The method as described in claim 1, characterized in that, Before the step of performing in-depth analysis of the key information based on the intent classification model to obtain the question type and true intent, the method further includes: Obtain a preset language model, fine-tune the preset semantic model, and obtain an initial intent classification model; Based on the training question and the corresponding training intent label, the initial intent classification model is trained to obtain training output data. Based on the training output data, calculate the cross-entropy loss and focus loss; The output loss is determined based on the cross-entropy loss and the focus loss. When the output loss satisfies the preset convergence condition, the trained initial intent classification model is used as the intent classification model.
5. The method as described in claim 1, characterized in that, The steps for text cleaning the equity investment issue to obtain standardized text include: The equity investment question is matched with noise characters in a preset noise character rule library to identify matching noise characters. The matching noise characters in the equity investment question are then removed to obtain the noise-filtered text. Remove redundant line breaks and tabs from the noise-filtered text, and compress consecutive whitespace characters in the noise-filtered text into single characters to obtain the formatted text; The character encoding format of the formatted text is uniformly converted to the target encoding format, the character width of the formatted text is uniformly converted to the target character width, the date format of the formatted text is uniformly converted to the target date format, and the amount format of the formatted text is uniformly converted to the target amount format, resulting in a formatted text. Based on the standard statement mapping table, the spoken sentence patterns of the standardized text are converted into corresponding standard sentence patterns to obtain the standardized text.
6. The method according to any one of claims 1 to 5, characterized in that, After the step of performing semantic understanding and contextual reasoning on the effective information based on a large language model to generate an answer to the equity investment question, the method further includes: Based on the target user's historical behavior data and historical transaction data, the target user's basic attribute tags, preference feature tags, and behavior prediction tags are determined, and corresponding recommendation weights are assigned to the target user's basic attribute tags, preference feature tags, and behavior prediction tags; Based on the publication time of existing information, candidate information is determined from the existing information; Calculate the first matching degree between the candidate information and the basic attribute label, the second matching degree between the candidate information and the preference feature label, and the third matching degree between the candidate information and the behavior prediction label, respectively; Based on the first matching degree, the second matching degree, the third matching degree, and the recommendation weight, the recommendation score of the candidate information is determined; Based on the recommendation scores of the candidate information, target preference information is pushed to the target user.
7. The method according to any one of claims 1 to 5, characterized in that, After the step of performing semantic understanding and contextual reasoning on the effective information based on a large language model to generate an answer to the equity investment question, the method further includes: Based on external public opinion data and the target users' historical behavior data and historical transaction data, risk indicators are calculated. The risk indicators include at least public opinion indicators, financial indicators, behavioral indicators, correlation indicators, and policy indicators. The risk indicators are input into the risk assessment model for identification to obtain the current risk level and the confidence score corresponding to the current risk level. When the confidence score is greater than or equal to the confidence threshold, a current risk management strategy is determined based on the current risk level, and the current risk level and the current risk management strategy are sent to the target user.
8. An intelligent analysis device for equity investment, characterized in that, The device includes: The intent analysis module is used to obtain equity investment questions input by the target user, and to perform text cleaning on the equity investment questions to obtain standardized text. The intent analysis module is also used to perform word segmentation on the standard text based on the word segmentation model to obtain segmented words, and extract key information from the segmented words; The intent analysis module is also used to perform in-depth analysis of the key information based on the intent classification model to obtain the question type and the true intent. The information retrieval module is used to retrieve valid information from the equity investment knowledge base based on the question type and the true intent. The intelligent answer module is used to perform semantic understanding and contextual reasoning on the effective information based on a large language model, and generate answers to the equity investment questions. The answers are in the form of at least one of explanatory text, visual charts, and structured analysis reports.
9. An intelligent analysis device for equity investment, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the equity investment intelligent analysis method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the equity investment intelligent analysis method as described in any one of claims 1 to 7.