Training sample data generation method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610865993.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-15
Smart Images

Figure CN122759249A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence data processing technology, specifically to the field of large language model training data construction technology, and particularly to training sample data generation methods, devices, electronic devices and storage media. Background Technology
[0002] Existing methods for generating training sample data typically involve directly scraping web text or manually annotating datasets. However, text data obtained through direct scraping often lacks a direct correspondence with the original underlying data, making it difficult for models to learn the reasoning process from factual evidence to expert opinions, and easily leading to factual errors. While manual annotation can ensure data quality, it has a high professional threshold, is time-consuming, and extremely costly, making it difficult to support the data requirements of large-scale model training. Furthermore, the use of simple templates to construct instructions results in monotonous and rigid instructions in the training data, poor model generalization ability, and an inability to adapt to complex and varied task scenarios. These methods lead to a lack of accurate correspondence between input data and target output in the training samples, a monotonous form of task instructions, and difficulty in effectively controlling data quality. Summary of the Invention
[0003] This disclosure provides a method, apparatus, electronic device, and storage medium for generating training sample data.
[0004] According to one aspect of this disclosure, a method for generating training sample data is provided, the method comprising: Obtain raw data and expert comment data, and extract metadata from the raw data and expert comment data; Based on the metadata, the original data of the same object in the same period are matched and associated with the expert comment data to obtain the association result; By analyzing the analytical dimensions, structural framework, and language style of the expert comments in the correlation results, corresponding diverse instructions are generated in reverse derivation. The diverse instructions, the raw data, and the expert commentary data are assembled into training sample data with an instruction fine-tuning format.
[0005] According to another aspect of this disclosure, a training sample data generation apparatus is provided, the apparatus comprising: The acquisition module is used to acquire raw data and expert comment data, and extract metadata from the raw data and the expert comment data; The association module is used to match and associate the original data of the same object in the same period with the expert comment data based on the metadata, so as to obtain the association result; The analysis module is used to analyze the analytical dimensions, structural framework, and language style of the expert comments in the associated results, and reverse deduce and generate corresponding diverse instructions. An assembly module is used to assemble the diverse instructions, the raw data, and the expert commentary data into training sample data with an instruction fine-tuning format.
[0006] According to a third aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described in any of the above technical solutions.
[0007] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any one of the methods described above.
[0008] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in any one of the above technical solutions.
[0009] This disclosure provides a method, apparatus, device, and storage medium for generating training sample data. By acquiring raw data and expert commentary data and extracting metadata, this application fully utilizes the information value of existing high-quality expert commentary as a benchmark, providing a rigorously aligned data foundation for subsequent training sample construction. Next, based on metadata, raw data and expert commentary data for the same object within the same period are matched and associated. Through dual constraints of object identifiers and time period identifiers, a strict correspondence between input and output is ensured, enabling the model to learn the reasoning process of "speaking based on real data," effectively reducing the illusion problem caused by data misalignment. Furthermore, by analyzing the analytical dimensions, structural framework, and language style of the expert commentary in the association results, diverse instructions are generated through reverse deduction, overcoming the bottleneck of model rigidity and insufficient generalization ability caused by fixed template instructions. This allows the training data to cover constrained scenarios of different professional roles, task organization logic, and expression patterns. Finally, diverse instructions, raw data, and expert commentary data are assembled into training sample data with a fine-tuned instruction format. This achieves the automated construction of high-quality training data with expert commentary as the output target, raw data as the input constraint, and reverse-engineered instructions as the task guide. This significantly reduces the cost of manual annotation while improving the model's factual accuracy, logical coherence, and adaptability to complex scenarios in vertical domain professional tasks. In summary, this application ensures strict correspondence between input and output through strong metadata association and avoids template-based rigidity by generating diverse instructions through reverse engineering, thus achieving the automated construction of high-quality training samples and improving the model's factual accuracy and generalization ability.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a schematic diagram of the steps of the training sample data generation method in the embodiments of this disclosure; Figure 2 This is a schematic diagram of the overall process of the training sample data generation method in the embodiments of this disclosure; Figure 3 This is a schematic block diagram of the training sample data generation device in the embodiments of this disclosure; Figure 4 This is a block diagram of an electronic device used to implement the training sample data generation method of the embodiments of this disclosure. Detailed Implementation
[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0013] This disclosure provides a method for generating training sample data; see [link to relevant documentation]. Figure 1 As shown, Figure 1 This is a schematic diagram illustrating the steps of a training sample data generation method in an embodiment of this disclosure. The method is applied to a server and specifically includes: Step S101: Obtain raw data and expert commentary data, and extract metadata from the raw data and expert commentary data.
[0014] Specifically, this step involves the parallel acquisition of dual-source data and metadata extraction. Raw data refers to unprocessed basic factual data, expert commentary data refers to high-quality text data with professional analytical attributes, and metadata refers to identifying information used to describe the characteristics of the data itself and support data association and matching. This application uses the financial vertical sector as an example to illustrate the raw data and expert commentary data. Raw data refers to original announcements and financial statements of listed companies, including structured financial data such as balance sheets, profit and loss statements, and cash flow statements, which are crawled from exchange information disclosure platforms such as the Shanghai Stock Exchange and Shenzhen Stock Exchange websites. Expert commentary data refers to financial statement commentary sections in brokerage research reports imported from a local research report database, including analysts' professional analytical opinions on revenue growth, profitability, risk factors, and future outlook. Raw data metadata refers to identifying information such as company name, stock code, year, and reporting period extracted from the raw financial statement data; expert commentary data metadata refers to identifying information such as brokerage name, company abbreviation, publication date, and reporting period extracted from the research report comments.
[0015] The specific implementation process of this scheme includes: First, obtaining raw data files from information disclosure platforms, performing structured parsing and format conversion, and then filtering by context length to obtain preprocessed raw data. Entity identifiers, unique object codes, and time period identifiers are extracted from the structured text header area as raw data metadata. Second, importing expert commentary data from a local knowledge base, performing initial screening to remove irrelevant documents using a local lightweight model, followed by named entity recognition, noise removal, and text extraction to obtain preprocessed expert commentary data. Source entity identifiers, object references, publication timestamps, and time period identifiers are extracted from the title area, abstract area, and body text as commentary data metadata, thus providing a data foundation for subsequent precise matching and association based on metadata.
[0016] Step S102: Based on metadata, match and associate the original data of the same object in the same period with the expert comment data to obtain the association result.
[0017] Specifically, this step aims to address the technical problem of a lack of strict correspondence between inputs and outputs in training samples. Matching and association refers to the data processing operation of establishing correspondences between relevant records in heterogeneous data sources based on specific identification conditions. The association result refers to the data set with a clear input-output mapping relationship formed after matching.
[0018] The specific implementation process of this scheme includes: comparing the unique object code in the original data metadata with the object identifier in the comment data metadata, and simultaneously comparing the time period identifier in the original data metadata with the time period identifier in the comment data metadata. When both the object identifier and the time period identifier match, it is determined that the original data and the expert comment data belong to the same object and the same period, thus establishing a matching association between the two. Through this dual constraint mechanism, it is ensured that the original data and the expert comment data in the training samples are strictly aligned in both the object and period dimensions, avoiding the model illusion problem caused by data misalignment, and obtaining a strictly corresponding association result between input and output.
[0019] Step S103: Analyze the analytical dimensions, structural framework, and language style of the expert comments in the correlation results, and reverse deduce the corresponding diversified instructions.
[0020] Specifically, this step aims to address the technical problem of insufficient model generalization ability caused by the single, templated nature of instructions in the training data. Here, "analysis dimension" refers to the professional analytical scope covered by the expert commentary content; "structural framework" refers to the paragraph organization logic and hierarchical relationship of the expert commentary content; and "language style" refers to the lexical sentiment and expressive patterns of the expert commentary text. "Reverse derivation" refers to the process of deriving generation conditions from existing results; and "diversified instructions" refers to the set of task instructions covering different role settings, task organization methods, and style constraints.
[0021] The specific implementation process of this scheme includes: First, performing paragraph semantic segmentation on the expert commentary data in the associated results, identifying functional blocks, and extracting at least one analytical dimension feature, including financial indicator analysis, profitability assessment, risk factor identification, and trend prediction. Then, performing syntactic feature extraction, lexical sentiment distribution statistics, and rhetorical pattern recognition on the functional blocks to obtain structural frame identifiers and language style vectors. Finally, based on the analytical dimension features, structural frame identifiers, and language style vectors, generating corresponding diversified instructions through at least one strategy among role setting reverse, task decomposition reverse, or style imitation reverse, i.e., generating role constraint instructions based on the level of professionalism represented by the analytical dimension features, generating structural constraint instructions based on the paragraph organization logic represented by the structural frame features, and generating style constraint instructions based on the expression patterns represented by the language style features. This ensures that the instructions in the training data cover different professional roles, task decomposition methods, and style imitation scenarios, avoiding model rigidity caused by single templated instructions.
[0022] Step S104: Assemble the diverse instructions, raw data, and expert commentary data into training sample data with an instruction fine-tuning format.
[0023] Specifically, this step aims to transform the aforementioned processing results into a standardized data format that can be used for supervised model fine-tuning. The instruction fine-tuning format refers to a structured data organization containing three core fields: task instructions, input context, and target output. The training sample data refers to labeled data used to guide the model in learning specific task mapping relationships during the supervised fine-tuning phase.
[0024] The specific implementation process of this scheme includes: mapping the diverse instructions generated by reverse derivation to the instruction field as a task description to be executed by the model; mapping the preprocessed raw data to the input field as the factual basis for the model to generate responses; mapping the quality-filtered expert commentary data to the output field as the benchmark truth target for model learning; and assembling the above three fields into a complete training sample through a key-value pair data structure, so that the model can learn the mapping relationship from diverse instructions and raw data to expert-level output during supervised fine-tuning, thereby possessing the ability to generate professional analysis content based on real data.
[0025] This disclosure provides a method, apparatus, electronic device, and storage medium for generating training sample data. By acquiring raw data and expert commentary data and extracting metadata, this application fully leverages the information value of existing high-quality expert commentary as a benchmark, providing a rigorously aligned data foundation for subsequent training sample construction. Next, based on metadata, raw data and expert commentary data for the same object within the same period are matched and associated. Through dual constraints of object identifiers and time period identifiers, a strict correspondence between input and output is ensured, enabling the model to learn the reasoning process of "speaking based on real data," effectively reducing the illusion problem caused by data misalignment. Furthermore, by analyzing the analytical dimensions, structural framework, and language style of expert commentary in the association results, diverse instructions are generated through reverse deduction, overcoming the bottleneck of model rigidity and insufficient generalization ability caused by fixed template instructions. This allows the training data to cover constrained scenarios of different professional roles, task organization logic, and expression patterns. Finally, diverse instructions, raw data, and expert commentary data are assembled into training sample data with a fine-tuned instruction format. This achieves the automated construction of high-quality training data with expert commentary as the output target, raw data as the input constraint, and reverse-engineered instructions as the task guide. This significantly reduces the cost of manual annotation while improving the model's factual accuracy, logical coherence, and adaptability to complex scenarios in vertical domain professional tasks. In summary, this application ensures strict correspondence between input and output through strong metadata association and avoids template-based rigidity by generating diverse instructions through reverse engineering, thus achieving the automated construction of high-quality training samples and improving the model's factual accuracy and generalization ability.
[0026] In some optional embodiments, the analytical dimensions, structural framework, and language style of expert comments in the analysis results are analyzed, and corresponding diverse instructions are derived in reverse, including: The expert comments in the associated results are subject to topic identification, and analytical dimension features are extracted; among them, the analytical dimension features include at least one of revenue analysis, profit analysis, risk warning and future outlook; The structural analysis of expert comments is performed to extract paragraph organization logic and argumentation hierarchy features, thereby obtaining structural framework features; Analyze the language features of expert comments, extract lexical sentiment and expression pattern features, and obtain language style features. Based on the analysis of dimensional features, structural framework features, and language style features, corresponding diverse instructions are generated through at least one of the following strategies: role setting reverse engineering, task decomposition reverse engineering, or style imitation reverse engineering.
[0027] Specifically, this step aims to extract multi-dimensional features from expert commentary data and reverse-engineer diverse instructions to address the problem of insufficient model generalization ability caused by single, templated instructions. Specifically, topic identification refers to the process of determining the core content category through text classification or keyword matching; dimensional feature analysis refers to the identifying information representing the professional analysis category. Structural analysis refers to the process of deconstructing the paragraph hierarchy and logical relationships of the text; structural framework features refer to the identifying information representing the content organization method. Language feature analysis refers to the process of statistically analyzing the emotional tone and expression patterns of words; language style features refer to the vector identifiers representing the text's expressive tendencies. Role setting reverse engineering refers to the strategy of generating subject identity constraints based on the content's professional level; task decomposition reverse engineering refers to the strategy of generating step constraints based on the content's organizational structure; and style imitation reverse engineering refers to the strategy of generating style constraints based on expression patterns.
[0028] The specific implementation process of this scheme includes: First, identifying the themes of expert comments in the associated results, and extracting at least one analytical dimension feature from revenue analysis, profit analysis, risk warning, and future outlook through a pre-set classification model or keyword database matching. Further, structural analysis of the expert comments is performed, extracting paragraph organization logic and argumentation hierarchy features by identifying paragraph boundaries and hierarchical relationships, thus obtaining structural framework features. Simultaneously, linguistic feature analysis is conducted on the expert comments, extracting lexical sentiment tendency and expression pattern features through statistical analysis of lexical sentiment polarity distribution and syntactic and rhetorical patterns, thus obtaining language style features. Finally, based on the analytical dimension features, structural framework features, and language style features, corresponding diversified instructions are generated through at least one strategy among role setting reverse, task decomposition reverse, or style imitation reverse. Specifically, role constraint instructions are generated based on the level of professionalism represented by the analytical dimension features, structural constraint instructions are generated based on the paragraph organization logic represented by the structural framework features, and style constraint instructions are generated based on the expression patterns represented by the language style features, thereby ensuring that the instructions in the training data cover different professional roles, task decomposition methods, and style imitation scenarios.
[0029] In this way, by identifying themes in the expert comments within the associated results to extract analytical dimension features, the professional analytical scope covered by the expert comments can be automatically identified. This allows for the generation of differentiated role constraint instructions based on the professional depth represented by different analytical dimensions, avoiding homogenized model output caused by a single role setting. Furthermore, structural analysis of the expert comments extracts paragraph organization logic and argumentation hierarchy features, enabling the understanding of the expert's analytical framework and argumentation path. This allows for the generation of differentiated task decomposition instructions based on different structural frameworks, enabling the model to organize output content according to professional logic rather than simply piling up information. Simultaneously, linguistic feature analysis of the expert comments extracts lexical sentiment and expression pattern features, capturing the different experts' expression habits and rhetorical characteristics. This allows for the generation of differentiated style constraint instructions based on different language styles, enabling the model to adapt to various expression scenarios. Finally, based on analytical dimension features, structural framework features, and language style features, at least one strategy—role setting reverse, task decomposition reverse, or style imitation reverse—is used to generate corresponding diversified instructions, achieving the technical effect of deriving multiple task instructions from a single expert comment. It significantly expands the diversity of instructions in the training data, allowing the model to be exposed to task constraints covering different roles, structures, and styles during supervised fine-tuning. This effectively breaks through the bottleneck of model rigidity caused by fixed template instructions and improves the model's generalization ability and output quality in complex professional scenarios.
[0030] In some optional embodiments, corresponding diverse instructions are generated through at least one strategy among role setting reverse engineering, task decomposition reverse engineering, or style imitation reverse engineering, including: Role attributes are mapped to the features of the analysis dimensions to generate role constraint instructions; among them, role constraint instructions are used to limit the professional identity attributes of the generating subject. The structural framework features are decomposed into task steps to generate structural constraint instructions; among them, the structural constraint instructions are used to limit the logical order of the output content. The language style features are transformed into rhetorical features to generate style constraint instructions; the style constraint instructions are used to limit the expression pattern features of the output text.
[0031] Specifically, this step aims to convert the extracted multidimensional features into specific instruction constraints to achieve fine-grained control over the model's generation behavior. Role attribute mapping refers to the process of mapping the professional depth represented by the analytical dimensional features to a predefined role attribute space; role constraint instructions are instruction fragments used to limit the identity attributes of the model's generated subject. Task step decomposition refers to the process of breaking down the paragraph organization logic represented by the structural framework features into ordered execution steps; structural constraint instructions are instruction fragments used to limit the organizational order of the model's output content. Rhetorical feature conversion refers to the process of converting the expression patterns represented by the language style features into executable style rules; style constraint instructions are instruction fragments used to limit the expression features of the model's output text.
[0032] The specific implementation process of this scheme includes: First, role attribute mapping is performed on the analytical dimension features. By matching the analytical dimension features with a preset role attribute library through similarity or rule mapping, the professional identity attributes that the generating subject should possess are determined, generating role constraint instructions. Next, task steps are decomposed on the structural framework features. By breaking down the paragraph organization logic and argumentation level features into ordered task execution steps, the organizational logical order that the output content should follow is determined, generating structural constraint instructions. Simultaneously, rhetorical feature transformation is performed on the language style features. By mapping the lexical sentiment tendency and expression mode features to a specific set of rhetorical rules, the expression mode features that the output text should adopt are determined, generating style constraint instructions. Finally, through the combination of role constraint instructions, structural constraint instructions, and style constraint instructions, diverse instructions covering the three dimensions of identity, structure, and style are formed, enabling the model to generate output content based on the corresponding professional identity, organizational logic, and expression mode when receiving a specific combination of instructions.
[0033] In this way, by mapping the analytical dimensional features to role attributes to generate role constraint instructions, the system can automatically determine the identity attributes that the generating subject should possess based on the professional analytical scope covered by the expert comments. This limits the model to generate content from a corresponding professional perspective, avoiding generalities caused by a lack of professional stance in the model output. Furthermore, by decomposing the structural framework features into task steps to generate structural constraint instructions, the system can transform the paragraph organization logic and argumentation levels of the expert comments into an executable sequence of task steps. This limits the model output content to follow a rigorous logical order, avoiding chaotic content structure and jumpy arguments. Simultaneously, by converting the language style features into rhetorical features to generate style constraint instructions, the system can transform the lexical sentiment and expression patterns of the expert comments into reusable style rules. This limits the model output text to use matching expressions, avoiding a single style and lack of adaptability. Finally, through the synergistic effect of role constraint instructions, structural constraint instructions, and style constraint instructions, a precise mapping from the multidimensional features of expert comments to multidimensional instruction constraints is achieved. This ensures that each instruction in the training data has a clear identity, structural specifications, and style requirements. During supervised fine-tuning, the model can learn the collaborative constraints of professional identity, logical organization, and expression patterns, significantly improving the professionalism, structural rationality, and stylistic diversity of the model's output. This effectively solves the technical problems of model rigidity and severe output homogenization caused by templated instructions.
[0034] In some optional embodiments, corresponding diverse instructions are generated through at least one strategy among role setting reverse engineering, task decomposition reverse engineering, or style imitation reverse engineering, including: Based on a pre-defined rule template library, the analysis dimension set, structural framework identifier, and language style vector are matched with candidate templates in the template library to generate corresponding diverse instructions; Alternatively, a large language model can be invoked, using the set of analysis dimensions, structural framework identifiers, and language style vectors as input context, and generating corresponding diverse instructions through meta-hint reasoning.
[0035] Specifically, this step aims to provide two specific and diverse implementation methods for instruction generation. The rule template library refers to a pre-built set of templates containing various instruction structure patterns. Candidate templates refer to available template entries in the template library that match the current features; matching refers to the process of selecting corresponding templates based on feature similarity or rule conditions. The large language model refers to a generative pre-trained model with a large number of parameters and strong semantic understanding capabilities; input context refers to the set of conditional information provided to the model for reasoning; meta-hint reasoning refers to a reasoning mechanism that uses guided hints to enable the model to perform specific generation tasks.
[0036] The specific implementation process of this scheme includes: The first method is to generate a matching result based on a preset rule template library: The set of analysis dimensions, structural framework identifiers, and language style vectors are compared with candidate templates in the template library using similarity calculations or rule condition comparisons to select the target template with the highest matching degree. Feature information is then filled into the corresponding placeholder positions of the target template to generate corresponding diversified instructions.
[0037] The second approach is meta-hint generation based on a large language model: This involves calling the interface of a large language model deployed in the cloud or locally, encoding the analysis dimension set, structural framework identifiers, and language style vectors into a structured text-based input context, and designing meta-hints that include a description of the generation task and output format requirements. The meta-hints and input context are then concatenated and input into the large language model, which generates corresponding diverse instructions through its autoregressive inference mechanism. Both implementation methods can be flexibly selected or combined based on system resource constraints and generation quality requirements. The rule template library approach offers advantages in high generation determinism and low computational overhead, while the large language model approach offers advantages in high generation flexibility and broad scenario coverage.
[0038] In this way, by providing two diverse instruction generation methods—rule template library matching and large language model meta-hint inference—resource adaptation and scenario coverage of the generation mechanism are achieved. When generating instructions based on a preset rule template library, the analysis dimension set, structural framework identifier, and language style vector are conditionally compared with candidate templates. Template filling quickly generates highly deterministic and formatted instructions, suitable for scenarios with high requirements for generation efficiency and output consistency, effectively reducing computational overhead and ensuring the stability of the instruction structure. When generating instructions using a large language model for meta-hint inference, the aforementioned feature vectors are encoded as input context. Guided meta-hints are designed to trigger the model's autoregressive inference capability, generating highly flexible instructions with broad scenario coverage. This is suitable for scenarios with high requirements for instruction diversity and innovation, effectively overcoming the coverage blind spots caused by the size limitation of the template library. The two implementation methods can be flexibly selected or combined according to system resource constraints and generation quality requirements. The rule template library method ensures reliable coverage of basic scenarios, while the large language model method expands the generation capabilities for complex scenarios. The synergy between the two enables the system to stably output diverse instructions under different computing power environments and quality requirements, significantly improving the adaptability and scalability of training data construction, and effectively solving the technical contradiction that a single generation mechanism cannot balance efficiency and flexibility.
[0039] In some optional embodiments, the raw data and expert comment data are obtained, and the metadata of the raw data and expert comment data is extracted, including: The raw data is processed to obtain preprocessed raw data and raw data metadata; The expert comment data is processed to obtain preprocessed expert comment data and comment data metadata.
[0040] Specifically, this step involves parallel preprocessing and metadata extraction of dual-source data, aiming to provide a structured data foundation and matching identification information for subsequent data association. Here, raw data refers to unprocessed basic factual data; preprocessing refers to the operations of format conversion, noise filtering, and structuring of the raw data; and raw data metadata refers to the identification information extracted from the raw data to characterize the data source object and time period. Expert commentary data refers to high-quality text data with professional analytical attributes; and commentary data metadata refers to the identification information extracted from the expert commentary data to characterize the commentary source, object reference, and time period.
[0041] The specific implementation process of this scheme includes: processing the raw data, which involves: obtaining raw data files from the information disclosure platform; performing structured parsing and format conversion on the raw data files to unify the data representation; and filtering by context length to remove documents with insufficient information or exceeding processing capacity, resulting in preprocessed raw data. Next, the header region of the structured text is identified from the preprocessed raw data, and entity identifiers, unique object codes, and time period identifiers are extracted from the header region as raw data metadata.
[0042] The processing of expert review data includes: importing expert review data from a local knowledge base; calling a lightweight classification model deployed on a local computing node for initial screening to remove documents irrelevant to the target domain; and then performing named entity recognition, noise removal, and text extraction on the pre-screened expert review data to obtain pre-processed expert review data. Next, the title and summary regions are identified from the pre-processed expert review data, and the source entity identifier in the title region, the object reference in the summary region, and the publication timestamp are extracted. Time period identifiers are then matched from the text based on a preset regular expression pattern as review data metadata. Through the above parallel processing, structured raw data and its metadata, as well as expert review data and its metadata, are obtained, providing a data foundation for subsequent precise matching and association based on metadata.
[0043] In this way, by processing the raw data, preprocessed raw data and its metadata are obtained, giving the raw data a unified data format and clear identification information, providing a structured input foundation for subsequent data association. Simultaneously, the expert commentary data is processed to obtain preprocessed expert commentary data and its metadata, giving the expert commentary data expert-level quality and clear identification information, providing a high-quality output benchmark for subsequent data association. Through parallel preprocessing and metadata extraction of the two source data, structured raw data and its metadata, as well as expert commentary data and its metadata, are obtained respectively. This provides a two-way data foundation and identification basis for subsequent accurate matching and association based on metadata, ensuring that the input and output data in the training sample construction process are reliably guaranteed in terms of format standardization and identification integrity.
[0044] In some optional embodiments, the original data is processed to obtain preprocessed original data and original data metadata, including: Raw data is crawled from information disclosure platforms using targeted web crawlers, and then structured and format-converted, followed by context length filtering, to obtain preprocessed raw data. The structured text header region is identified from the preprocessed raw data, and entity identifiers, object unique codes, and time period identifiers are extracted from the header region to obtain the raw data metadata.
[0045] Specifically, this step involves the automated collection, preprocessing, and metadata extraction of raw data. Specifically, targeted web crawlers refer to data collection programs configured with crawling rules for specific data sources; information disclosure platforms refer to official information release systems that publicly provide raw data files; structured parsing refers to the process of converting unstructured or semi-structured data into a standardized data format; format conversion refers to the operation of unifying data from different sources into a preset standard representation; and context length filtering refers to a mechanism that removes documents with insufficient information or exceeding processing capacity based on a preset length threshold. The header area refers to a fixed location area in the structured text containing core identifying information; entity identifiers refer to string codes used to uniquely identify the data source object; object unique codes refer to object identifiers that guarantee global uniqueness within a specific namespace; and time period identifiers refer to marking information used to characterize the time interval to which the data belongs.
[0046] The specific implementation process of this solution includes: First, a targeted web crawler is used to retrieve raw data from the information disclosure platform. The crawler accesses the target platform and downloads the raw data files based on preset URL (Uniform Resource Locator) rules, request frequency, and parsing templates. Then, the retrieved raw data undergoes structured parsing and format conversion. By identifying the document type, the corresponding parsing engine is called to uniformly convert raw data in different formats such as PDF (Portable Document Format), HTML (HyperText Markup Language), or scanned images into a standardized structured text format. Simultaneously, context length filtering is performed, calculating the number of characters or terms in the converted text, and removing documents below the minimum length threshold or exceeding the maximum length threshold, resulting in preprocessed raw data. Finally, the head region of the structured text is identified from the preprocessed raw data. The text block containing the core identification information is located by using preset positioning rules or pattern matching. The entity identifier representing the source object of the data, the unique object code that ensures global uniqueness, and the time period identifier representing the time interval to which the data belongs are extracted from the head region to obtain the raw data metadata, which provides the identification basis for subsequent matching and association with expert commentary data.
[0047] In this way, by using targeted web crawlers to extract raw data from information disclosure platforms, it is possible to automatically and scalably acquire basic factual data from authoritative sources, avoiding the problems of low efficiency and incomplete data coverage associated with manual collection. Furthermore, the raw data undergoes structured parsing and format conversion, unifying heterogeneous formats such as PDF, HTML, or scanned images into a standardized structured text representation, eliminating format barriers between different data sources and improving the compatibility of subsequent processing. Simultaneously, context length filtering is performed, removing documents with insufficient information or exceeding processing capacity based on preset thresholds, ensuring the quality control and processing feasibility of the input data. Finally, the header region is identified from the preprocessed raw data, and entity identifiers, unique object codes, and time period identifiers are extracted as raw data metadata, giving the raw data clear identity and time positioning information. This provides a reliable identification basis for subsequent accurate matching and association with expert commentary data based on metadata. This effectively solves the problem of difficulty in association caused by inconsistent formats and missing identifiers in heterogeneous data sources, significantly improving the automation level of data collection and preprocessing and the accuracy of data alignment.
[0048] In some optional embodiments, the expert comment data is processed to obtain preprocessed expert comment data and comment data metadata, including: Local preprocessing is performed on the expert commentary data to obtain preprocessed expert commentary data; The preprocessed expert review data is then subjected to cloud-based fine screening and metadata extraction to obtain the review data metadata.
[0049] Specifically, this step involves layered processing and metadata extraction of expert review data, aiming to ensure data quality and integrity of identification through collaborative processing between local and cloud platforms. Local preprocessing refers to initial data cleaning and structuring operations performed on local computing nodes; the preprocessed expert review data refers to the expert review text that has reached a basic level of quality after local processing. Cloud-based fine screening refers to the operation of using a high-performance model deployed on cloud computing nodes for in-depth quality assessment; and metadata extraction refers to the process of identifying and extracting key identifying information for data association from the text.
[0050] The specific implementation process of this solution includes: First, local preprocessing of the expert commentary data: Importing expert commentary data from the local knowledge base, calling a lightweight classification model deployed on the local computing node for initial screening, and removing documents irrelevant to the target domain through text classification or rule matching to obtain the initially screened expert commentary data. Then, named entity recognition is performed on the initially screened expert commentary data to extract object identifiers, time period identifiers, and source entity identifiers. Noise removal and text extraction are then performed on the initially screened expert commentary data, removing non-core content such as headers, footers, and disclaimers, retaining professional analysis paragraphs, to obtain the preprocessed expert commentary data. Subsequently, the preprocessed expert commentary data undergoes cloud-based fine screening and metadata extraction: a large language model deployed on cloud computing nodes is invoked, and the preprocessed expert commentary data is input into the model for quality assessment and secondary screening. Low-quality content with chaotic logic and insufficient professionalism is identified and removed through semantic understanding capabilities, resulting in finely screened expert commentary data. Then, the title and summary regions are identified from the finely screened expert commentary data, and the source entity identifier in the title region, the object referent in the summary region, and the publication timestamp are extracted. Based on a preset regular expression pattern, time period identifiers are matched from the main text to obtain the commentary data metadata, providing an identification basis for subsequent matching and association with the original data.
[0051] In this way, by preprocessing the expert review data locally, initial screening, entity recognition, and text extraction can be quickly completed using local computing resources, resulting in preprocessed expert review data. This effectively reduces cloud computing overhead and improves processing efficiency. Furthermore, the preprocessed expert review data undergoes cloud-based fine screening and metadata extraction. A high-performance cloud model is used to deeply evaluate data quality and extract identifier information from the screened text, obtaining review data metadata. This ensures the expert-level quality and completeness of the identifiers in the expert review data. Through the layered collaboration of local preprocessing and cloud-based fine screening, a balance between computational efficiency and quality control is achieved. This avoids the problem of local processing being unable to identify deep quality defects and the resource waste caused by cloud processing alone, providing a high-quality data foundation and reliable identifier basis for subsequent accurate matching and association with the original data.
[0052] In some optional embodiments, the expert comment data is preprocessed locally to obtain preprocessed expert comment data, including: Import expert review data from the local knowledge base, call the lightweight classification model deployed on the local computing node to perform initial screening of the expert review data, remove documents that are irrelevant to the target domain, and obtain the initial screening of expert review data; Named entity recognition is performed on the initially screened expert comment data to extract object identifiers, time period identifiers, and source entity identifiers, thus obtaining the metadata of the initially screened expert comment data. Noise removal and text extraction were performed on the initially screened expert comment data to obtain preprocessed expert comment data.
[0053] Specifically, this step involves the entire local preprocessing workflow of expert review data, aiming to achieve initial data screening, metadata extraction, and text purification through layered processing. Here, the local knowledge base refers to a collection of expert review documents stored on a local server or in a private environment; the lightweight classification model refers to a text classification model with fewer parameters and faster inference speed; initial screening refers to the preliminary screening operation of quickly filtering irrelevant documents based on the classification results; named entity recognition refers to natural language processing technology that automatically identifies and extracts specific types of entity information from text; object identifiers refer to coded information used to uniquely identify the analyzed object; time period identifiers refer to the marking information used to characterize the time interval to which the data belongs; source entity identifiers refer to the marking information used to characterize the source organization or author of the review; noise removal refers to the process of removing irrelevant and interfering content from the document; and text extraction refers to the process of extracting the core analytical paragraphs from the document.
[0054] The specific implementation process of this solution includes: First, importing expert review data from a local knowledge base, then using a lightweight classification model deployed on a local computing node to perform initial screening of the expert review data. This involves using a text classification algorithm to determine the relevance of documents to the target domain, removing documents irrelevant to the target domain, and obtaining the initially screened expert review data. Next, named entity recognition is performed on the initially screened expert review data. Using a pre-trained entity recognition model or rule matching method, object identifiers representing the analyzed object, time period identifiers representing the time interval to which the data belongs, and source entity identifiers representing the source of the review are identified and extracted from the text, obtaining the metadata of the initially screened expert review data. Simultaneously, noise removal and text extraction are performed on the initially screened expert review data. Non-core noise content such as headers, footers, disclaimers, and inserted advertisements are identified and removed using preset rules or models. Paragraphs containing professional analysis are located and extracted, resulting in preprocessed expert review data, providing structured foundational data for subsequent cloud-based fine screening and data association.
[0055] In this way, by importing expert commentary data from a local knowledge base and calling a lightweight classification model deployed on a local computing node for initial screening, the fast inference speed of the lightweight model is used to quickly remove documents irrelevant to the target domain, resulting in pre-screened expert commentary data. This allows for efficient initial filtering of large-scale data under local computing resource constraints, effectively reducing the computational overhead and cloud access costs of subsequent processing. Furthermore, named entity recognition is performed on the pre-screened expert commentary data to extract object identifiers, time period identifiers, and source entity identifiers as metadata. This allows for the acquisition of key identification information for subsequent matching and association at an early stage of data processing, avoiding the resource waste caused by extracting metadata after full-text processing. Simultaneously, noise removal and text extraction are performed on the pre-screened expert commentary data, removing non-core noise content such as headers, footers, and disclaimers while retaining professional analysis paragraphs, resulting in pre-processed expert commentary data. This enables the subsequent cloud-based fine-screening model to focus on core analytical text rather than redundant interference information. Through the collaborative processing of local initial screening, early metadata extraction, and text purification, both computational efficiency and data quality have been improved. This ensures the professional relevance of expert commentary data and provides structured basic data and complete identification basis for subsequent cloud-based in-depth processing and cross-source data association.
[0056] In some optional embodiments, the preprocessed expert review data undergoes cloud-based screening and metadata extraction to obtain review data metadata, including: The large language model deployed on the cloud computing node is invoked to perform a second screening of the preprocessed expert comment data, resulting in refined expert comment data. The title and abstract regions are identified from the carefully selected expert review data. The source entity identifier in the title region, the object referent in the abstract region, and the publication timestamp are extracted. Based on the preset regular expression pattern, the time period identifier is matched from the main text to obtain the review data metadata.
[0057] Specifically, this step involves in-depth cloud-based screening and refined metadata extraction of expert commentary data, aiming to ensure data quality and improve identification information through high-performance computing resources. Here, cloud computing nodes refer to computing resources deployed on remote server clusters, possessing high-performance computing and large-scale model inference capabilities. Large language models refer to pre-trained language models with a large number of parameters and deep semantic understanding and generation capabilities; secondary screening refers to in-depth quality assessment and refined filtering operations based on the initial local screening; refined expert commentary data refers to high-quality expert commentary texts retained after cloud-based model evaluation. The title area refers to the text block containing the document title; the abstract area refers to the text block containing the document abstract or core viewpoints; the source entity identifier refers to the tagging information used to represent the source institution or author of the commentary; the object referent refers to the abbreviation or alias used to refer to the object being analyzed; the publication timestamp refers to the specific time mark of the document's publication; and the regular expression pattern refers to the pattern rule expression used for text matching.
[0058] The specific implementation process of this solution includes: First, a large language model deployed on cloud computing nodes is used to perform a secondary screening of the preprocessed expert commentary data. The preprocessed expert commentary data is input into the large language model, and the model's deep semantic understanding capabilities are used to evaluate the professionalism, logical coherence, and analytical depth of the text. Low-quality content with chaotic logic, weak argumentation, or insufficient professionalism is identified and eliminated, resulting in refined expert commentary data. Next, the title and abstract regions are identified from the refined expert commentary data. The document title block and abstract block are located using preset positioning rules or layout analysis algorithms. Source entity identifiers representing the source of the commentary are extracted from the title region, and object referents representing the analyzed object and publication timestamps representing the document's publication time are extracted from the abstract region. Simultaneously, time period identifiers are matched from the body text based on preset regular expression patterns. By designing regular expression rules for matching time period expressions, time period identifiers representing the time interval to which the data belongs are searched and extracted from the body text, obtaining commentary data metadata, which provides a complete identifier basis for subsequent accurate matching and association with the original data.
[0059] In this way, by calling a large language model deployed on cloud computing nodes to perform a second screening of the preprocessed expert commentary data, and leveraging the high-performance computing resources and deep semantic understanding capabilities of the cloud, low-quality content with chaotic logic, weak argumentation, or insufficient professionalism is identified and eliminated, resulting in refined expert commentary data. This overcomes the limitations of local lightweight models in terms of semantic understanding depth, ensuring that the retained expert commentary data possesses expert-level professional quality and logical rigor. Next, the title and abstract regions are identified from the refined expert commentary data. The source entity identifier in the title region and the object referent and publication timestamp in the abstract region are extracted. Based on a preset regular expression pattern, time period identifiers are matched from the body text to obtain the commentary data metadata. This enables the accurate extraction of key identifier information for data association from high-quality text, avoiding metadata extraction errors caused by interference from low-quality text. By combining in-depth cloud-based screening with refined metadata extraction, the dual goals of controlling the quality of expert review data and ensuring the integrity of the identification were achieved. This not only ensured the professional level of the training data but also provided a reliable identification basis for the subsequent accurate matching and association with the original data based on metadata. This effectively solved the technical problems of the inability of single local processing to identify deep quality defects and inaccurate metadata extraction.
[0060] In some optional embodiments, diverse instructions, raw data, and expert commentary data are assembled into training sample data with an instruction fine-tuning format, including: Map the diverse instructions to the instruction field to obtain the value of the first field; Map the original data to the input field to obtain the value of the second field; Map the expert commentary data to the output fields to obtain the value of the third field; Based on the values of the first, second, and third fields, a key-value pair data structure is constructed to obtain the training sample data.
[0061] Specifically, this step involves assembling the processed multi-source data into a standardized training sample format, aiming to establish a structured mapping relationship between instructions, inputs, and outputs. Here, diverse instructions refer to task instructions generated through reverse engineering, covering different roles, structures, and style constraints. Instruction fields refer to the data fields in the training samples used to carry task instructions; the first field value indicates the specific content of the instruction field. Raw data refers to preprocessed basic factual data; input fields refer to the data fields in the training samples used to carry the model's input context; the second field value indicates the specific content of the input field. Expert commentary data refers to high-quality expert analysis texts after careful screening; output fields refer to the data fields in the training samples used to carry the model's learning objectives; the third field value indicates the specific content of the output field. Key-value pair data structure refers to a structured data organization form where data types are identified by keys and specific content is stored by keys.
[0062] The specific implementation process of this scheme includes: First, mapping diverse instructions to an instruction field, and writing the reverse-generated diverse instruction text into the instruction field through field assignment operations to obtain the first field value. Next, mapping the raw data to an input field, and writing the preprocessed structured raw data text into the input field through field assignment operations to obtain the second field value. Simultaneously, mapping expert commentary data to an output field, and writing the refined expert commentary text into the output field through field assignment operations to obtain the third field value. Finally, based on the first, second, and third field values, a key-value pair data structure is constructed, using preset key names to identify the three data types: instructions, input, and output. The corresponding field values are stored as keys to obtain training sample data. This allows the model to learn the mapping relationship from diverse instructions and raw data to expert comments during the supervised fine-tuning stage by reading this key-value pair data structure.
[0063] In this way, by mapping diverse instructions to the instruction field to obtain the first field value, the training samples possess clear task guidance information, enabling the model to understand and generate targets based on different roles, structures, and style constraints. Next, mapping the raw data to the input field to obtain the second field value ensures the training samples contain complete factual basis, allowing the model to learn to reason based on real data rather than generating data out of thin air. Simultaneously, mapping expert commentary data to the output field to obtain the third field value allows the training samples to learn from high-quality expert analysis, enabling the model to mimic expert-level professional expression and logical organization. Finally, a key-value pair data structure is constructed based on the first, second, and third field values to obtain the training sample data, establishing a clear one-to-one correspondence between instructions, inputs, and outputs, facilitating efficient reading and learning during supervised fine-tuning. Through field-based mapping and the assembly of key-value pair structures, standardized organization and efficient storage of training data are achieved, enabling the model to simultaneously learn the complete mapping relationship between diverse task instructions, factual input constraints, and expert output targets, significantly improving data utilization efficiency and model learning effectiveness during supervised fine-tuning.
[0064] Optionally, after obtaining the training sample data, it can be used in the supervised fine-tuning stage of the large language model. Specifically, a standard instruction fine-tuning format sample containing diverse instructions, raw data, and expert commentary data is input into the base model, enabling the model to learn the mapping relationship from diverse task instructions and raw factual data to expert-level professional output. When the trained model receives raw data or related queries input by the user, it can generate professional commentary content that conforms to expert analysis standards based on instruction constraints. It can be applied to scenarios such as intelligent investment research assistants, automated research report generation systems, and professional domain question-and-answer systems, realizing the automated output of professional opinions based on real data and improving the model's factual accuracy, logical coherence, and adaptability to complex scenarios in vertical domain tasks.
[0065] See Figure 2 , Figure 2 This is a schematic diagram of the overall process of generating training sample data in this embodiment. The flowchart illustrates the complete processing flow for constructing training samples based on a dual-stream data flow. Path A at the top is the raw data flow (corresponding to the original data): Raw announcements and financial reports are crawled from the Shanghai Stock Exchange and Shenzhen Stock Exchange websites. After table format conversion, the heterogeneous formats are unified into structured text. Then, invalid documents are filtered out using context length filtering to obtain preprocessed raw data. Path B at the bottom is the expert knowledge flow (corresponding to expert commentary data): Expert research report data is imported from the research report database. First, financial report comments are filtered based on a local model to quickly eliminate documents outside the target domain. Then, key metadata such as company name, year, reporting period, and securities firm are identified. Subsequently, text extraction and format cleaning are performed to remove noise. Finally, financial report comments are filtered based on a public cloud model to deeply evaluate their professionalism and logic, resulting in high-quality expert commentary data. The two data streams converge at the associated node. The metadata extracted from Path A and Path B is used to strictly match and associate the raw financial report data of the same company and the same reporting period with the expert research report comments. The correlated data then enters the Prompt generation stage based on reverse engineering of financial statement reviews. This stage analyzes the analytical dimensions, structural framework, and language style of the expert reviews to derive diverse instructions. Finally, these diverse instructions, the original financial statement data, and the expert reviews are assembled into training samples <instructions, financial statement reviews> for supervised fine-tuning of the large-scale financial model.
[0066] The following describes an apparatus embodiment of this application, which can be used to execute the training sample data generation method in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the training sample data generation method described above.
[0067] This disclosure also provides a training sample data generation device 300, such as... Figure 3 As shown, it includes: The acquisition module 301 is used to acquire raw data and expert comment data, and extract metadata from the raw data and expert comment data. The association module 302 is used to match and associate the original data of the same object in the same period with the expert comment data based on metadata, so as to obtain the association result; Analysis module 303 is used to analyze the analytical dimensions, structural framework and language style of expert comments in the correlation results, and reverse deduce and generate corresponding diverse instructions. Assembly module 304 is used to assemble diverse instructions, raw data, and expert commentary data into training sample data in an instruction fine-tuning format.
[0068] In some optional embodiments, the analysis module 303 analyzes the analytical dimensions, structural framework, and language style of the expert comments in the correlation results, and reverse-derives and generates corresponding diverse instructions, including: The expert comments in the associated results are subject to topic identification, and analytical dimension features are extracted; among them, the analytical dimension features include at least one of revenue analysis, profit analysis, risk warning and future outlook; The structural analysis of expert comments is performed to extract paragraph organization logic and argumentation hierarchy features, thereby obtaining structural framework features; Analyze the language features of expert comments, extract lexical sentiment and expression pattern features, and obtain language style features. Based on the analysis of dimensional features, structural framework features, and language style features, corresponding diverse instructions are generated through at least one of the following strategies: role setting reverse engineering, task decomposition reverse engineering, or style imitation reverse engineering.
[0069] In some optional embodiments, the analysis module 303 generates corresponding diversified instructions through at least one strategy among role setting reverse engineering, task decomposition reverse engineering, or style imitation reverse engineering, including: Role attributes are mapped to the features of the analysis dimensions to generate role constraint instructions; among them, role constraint instructions are used to limit the professional identity attributes of the generating subject. The structural framework features are decomposed into task steps to generate structural constraint instructions; among them, the structural constraint instructions are used to limit the logical order of the output content. The language style features are transformed into rhetorical features to generate style constraint instructions; the style constraint instructions are used to limit the expression pattern features of the output text.
[0070] In some optional embodiments, the analysis module 303 generates corresponding diversified instructions through at least one strategy among role setting reverse engineering, task decomposition reverse engineering, or style imitation reverse engineering, including: Based on a pre-defined rule template library, the analysis dimension set, structural framework identifier, and language style vector are matched with candidate templates in the template library to generate corresponding diverse instructions; Alternatively, a large language model can be invoked, using the set of analysis dimensions, structural framework identifiers, and language style vectors as input context, and generating corresponding diverse instructions through meta-hint reasoning.
[0071] In some optional embodiments, the acquisition module 301 acquires the raw data and expert comment data, and extracts the metadata of the raw data and expert comment data, including: The raw data is processed to obtain preprocessed raw data and raw data metadata; The expert comment data is processed to obtain preprocessed expert comment data and comment data metadata.
[0072] In some optional embodiments, the acquisition module 301 processes the raw data to obtain preprocessed raw data and raw data metadata, including: Raw data is crawled from information disclosure platforms using targeted web crawlers, and then structured and format-converted, followed by context length filtering, to obtain preprocessed raw data. The structured text header region is identified from the preprocessed raw data, and entity identifiers, object unique codes, and time period identifiers are extracted from the header region to obtain the raw data metadata.
[0073] In some optional embodiments, the acquisition module 301 processes the expert comment data to obtain preprocessed expert comment data and comment data metadata, including: Local preprocessing is performed on the expert commentary data to obtain preprocessed expert commentary data; The preprocessed expert review data is then subjected to cloud-based fine screening and metadata extraction to obtain the review data metadata.
[0074] In some optional embodiments, the acquisition module 301 performs local preprocessing on the expert comment data to obtain preprocessed expert comment data, including: Import expert review data from the local knowledge base, call the lightweight classification model deployed on the local computing node to perform initial screening of the expert review data, remove documents that are irrelevant to the target domain, and obtain the initial screening of expert review data; Named entity recognition is performed on the initially screened expert comment data to extract object identifiers, time period identifiers, and source entity identifiers, thus obtaining the metadata of the initially screened expert comment data. Noise removal and text extraction were performed on the initially screened expert comment data to obtain preprocessed expert comment data.
[0075] In some optional embodiments, the acquisition module 301 performs cloud-based fine screening and metadata extraction on the preprocessed expert review data to obtain review data metadata, including: The large language model deployed on the cloud computing node is invoked to perform a second screening of the preprocessed expert comment data, resulting in refined expert comment data. The title and abstract regions are identified from the carefully selected expert review data. The source entity identifier in the title region, the object referent in the abstract region, and the publication timestamp are extracted. Based on the preset regular expression pattern, the time period identifier is matched from the main text to obtain the review data metadata.
[0076] In some optional embodiments, the assembly module 304 assembles diverse instructions, raw data, and expert commentary data into training sample data with an instruction fine-tuning format, including: Map the diverse instructions to the instruction field to obtain the value of the first field; Map the original data to the input field to obtain the value of the second field; Map the expert commentary data to the output fields to obtain the value of the third field; Based on the values of the first, second, and third fields, a key-value pair data structure is constructed to obtain the training sample data.
[0077] The acquisition, storage, and application of any type of information, such as user personal information, involved in the technical solutions disclosed herein comply with relevant laws and regulations and do not violate public order and good morals.
[0078] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0079] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0080] like Figure 4 As shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 may also store various programs and data required for the operation of the device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0081] Multiple components in device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 408, such as various types of monitors, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0082] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as training sample data generation methods. For example, in some embodiments, the training sample data generation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to perform training sample data generation methods by any other suitable means (e.g., by means of firmware).
[0083] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0084] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable training sample data generation device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0085] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0086] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0087] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0088] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0089] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0090] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for generating training sample data, wherein, The method includes: Obtain raw data and expert comment data, and extract metadata from the raw data and expert comment data; Based on the metadata, the original data of the same object in the same period are matched and associated with the expert comment data to obtain the association result; By analyzing the analytical dimensions, structural framework, and language style of the expert comments in the correlation results, corresponding diverse instructions are generated in reverse derivation. The diverse instructions, the raw data, and the expert commentary data are assembled into training sample data with an instruction fine-tuning format.
2. The method according to claim 1, wherein, The analysis of the expert comments in the correlation results, including their analytical dimensions, structural framework, and language style, is used to deduce and generate corresponding diverse instructions, including: The expert comments in the associated results are subject to topic identification, and analytical dimension features are extracted; wherein, the analytical dimension features include at least one of revenue analysis, profit analysis, risk warning and future outlook; The structure of the expert comments is analyzed to extract the paragraph organization logic and argumentation hierarchy features, thus obtaining the structural framework features. The linguistic features of the expert comments were analyzed to extract lexical sentiment and expression pattern features, thereby obtaining linguistic style features; Based on the analytical dimension features, the structural framework features, and the language style features, corresponding diverse instructions are generated through at least one of the following strategies: role setting reverse engineering, task decomposition reverse engineering, or style imitation reverse engineering.
3. The method according to claim 2, wherein, The method of generating diverse instructions through at least one of the following strategies—role setting reverse engineering, task decomposition reverse engineering, or style imitation reverse engineering—includes: Role attribute mapping is performed on the analytical dimension features to generate role constraint instructions; wherein, the role constraint instructions are used to limit the professional identity attributes of the generating subject; The structural framework features are decomposed into task steps to generate structural constraint instructions; wherein, the structural constraint instructions are used to limit the logical order of the output content. The language style features are subjected to rhetorical feature transformation to generate style constraint instructions; wherein, the style constraint instructions are used to limit the expression pattern features of the output text.
4. The method according to claim 2, wherein, The method of generating diverse instructions through at least one of the following strategies—role setting reverse engineering, task decomposition reverse engineering, or style imitation reverse engineering—includes: Based on a preset rule template library, the set of analysis dimensions, the structural framework identifier, and the language style vector are matched with candidate templates in the template library to generate corresponding diverse instructions; Alternatively, a large language model can be invoked, using the set of analysis dimensions, the structural framework identifier, and the language style vector as input context, and generating corresponding diverse instructions through meta-hint reasoning.
5. The method according to claim 1, wherein, The process of acquiring raw data and expert commentary data, and extracting metadata from the raw data and expert commentary data, includes: The original data is processed to obtain preprocessed original data and original data metadata; The expert review data is processed to obtain preprocessed expert review data and review data metadata.
6. The method according to claim 5, wherein, The process of processing the original data to obtain preprocessed original data and original data metadata includes: Raw data is crawled from information disclosure platforms using a targeted web crawler. The raw data is then subjected to structured parsing and format conversion, and context length filtering is performed to obtain preprocessed raw data. The structured text header region is identified from the preprocessed raw data, and the entity identifier, object unique code, and time period identifier in the header region are extracted to obtain the raw data metadata.
7. The method according to claim 5, wherein, The process of processing the expert comment data to obtain preprocessed expert comment data and comment data metadata includes: The expert comment data is preprocessed locally to obtain preprocessed expert comment data; The preprocessed expert review data is then subjected to cloud-based fine screening and metadata extraction to obtain the review data metadata.
8. The method according to claim 7, wherein, The step of performing local preprocessing on the expert comment data to obtain preprocessed expert comment data includes: Import expert review data from the local knowledge base, call a lightweight classification model deployed on the local computing node to perform initial screening of the expert review data, remove documents that are irrelevant to the target domain, and obtain the initial screening of expert review data; Named entity recognition is performed on the initially screened expert comment data to extract object identifiers, time period identifiers, and source entity identifiers, thereby obtaining the metadata of the initially screened expert comment data. The initial screening of expert comment data is subjected to noise removal and text extraction to obtain preprocessed expert comment data.
9. The method according to claim 7, wherein, The preprocessed expert review data undergoes cloud-based fine screening and metadata extraction to obtain review data metadata, including: The large language model deployed on the cloud computing node is invoked to perform a second screening on the preprocessed expert comment data to obtain the refined expert comment data. The title and summary regions are identified from the refined expert review data. The source entity identifier in the title region, the object referent in the summary region, and the publication timestamp are extracted. Based on a preset regular expression pattern, the time period identifier is matched from the main text to obtain the review data metadata.
10. The method according to any one of claims 1 to 9, wherein, The assembly of the diverse instructions, the raw data, and the expert commentary data into training sample data with an instruction fine-tuning format includes: The diverse instructions are mapped to the instruction field to obtain the value of the first field; The original data is mapped to the input field to obtain the value of the second field; The expert comment data is mapped to the output field to obtain the value of the third field; Based on the first field value, the second field value, and the third field value, a key-value pair data structure is constructed to obtain the training sample data.
11. A training sample data generation device, wherein, The device includes: The acquisition module is used to acquire raw data and expert comment data, and extract metadata from the raw data and the expert comment data; The association module is used to match and associate the original data of the same object in the same period with the expert comment data based on the metadata, so as to obtain the association result; The analysis module is used to analyze the analytical dimensions, structural framework, and language style of the expert comments in the associated results, and reverse deduce and generate corresponding diverse instructions. An assembly module is used to assemble the diverse instructions, the raw data, and the expert commentary data into training sample data with an instruction fine-tuning format.
12. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.
14. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-10.