AI large model corpus construction method and system based on digital experience

By building a multi-source data acquisition channel to acquire and preprocess the data and digitize it to the database, the problem of poor data quality in the AI ​​big model corpus is solved and the application efficiency of AI big model is improved.

CN120069041APending Publication Date: 2025-05-30CHONGQING PAPER CLIP INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510235342.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, enterprises have low efficiency in the application of AI models due to poor data quality in the corpus.

Method used

By building a multi-source data acquisition channel, the preprocessing of target problems and digitized structured data after real-time, and hierarchically store them in the database based on the structural characteristics of the structured data to generate an AI big model corpus.

Benefits of technology

It improves the quality of data in the corpus, makes AI big models more efficient when reading data, and significantly improves the application efficiency of AI big models in enterprise decision support and business process optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069041A_ABST
    Figure CN120069041A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of corpus construction, and particularly relates to an AI large model corpus construction method and system based on digital experience, and the method comprises the steps: obtaining the structural data of a target problem in real time based on a constructed multi-source data obtaining channel after preprocessing and experience digitization; and then, according to the structural features of the structured data of the target problem, storing the structured data in a database in a hierarchical manner, and generating an AI large model corpus. According to the method and the device, the problem of low application efficiency of the AI large model caused by poor data quality in a corpus in the process of constructing the AI large model by an enterprise in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of corpus construction, and particularly relates to a method and system for constructing an AI large model corpus based on digital experience. Background Art

[0002] In the complex architecture system of today's enterprise operations, enterprise managers, leaders of various departments, and numerous employees each perform their respective duties and carry out work in their respective exclusive execution fields. However, enterprise operations are not isolated individual behaviors, and there are countless connections between departments and employees during the process of handling daily affairs, and work negotiations occur frequently. Especially when facing various work problems, a good coordination and matching mechanism becomes a key factor for the efficient operation of enterprises. With the sweeping of the digital wave, more and more enterprises have realized the importance of digital transformation for enhancing competitiveness and turned their attention to the construction of AI large models, hoping to promote the comprehensive digital transformation of enterprises with the help of their powerful data analysis and processing capabilities.

[0003] In the process of constructing an AI large model, the construction of the corpus undoubtedly plays a central role. The corpus is like the cornerstone of a building, and its quality and scale directly determine the performance of the AI large model. It is pointed out in relevant research that "a rich and high-quality corpus is the primary prerequisite for training an accurate and efficient AI large model." For an enterprise, it contains a vast amount of communication data, detailed work records, and valuable work experience, and these information constitute the unique knowledge treasure of the enterprise. However, how to effectively transform these traditional and unstructured information into digital form so that it can be fully absorbed and utilized by the AI large model has become a major challenge in constructing a high-quality corpus.

[0004] Traditional enterprise knowledge management methods often focus on the simple storage and classification of documents, and it is difficult to meet the needs of AI large models for in-depth data mining and intelligent analysis. For example, some enterprises only archive work meeting records in text form, but do not extract and structure the key information in them, resulting in these data being unable to play their due value in the training of AI large models. In contrast, successfully constructing a corpus based on digital experience can enable the AI large model to more accurately understand the enterprise business process, employee collaboration mode, and problem-solving strategy, thus significantly improving its application efficiency in enterprise decision-making support, business process optimization, etc. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and system for constructing an AI large model corpus based on digital experience to solve the problem that the application efficiency of the AI large model is low due to poor data quality in the corpus during the construction of the AI large model in the prior art.

[0006] The basic solution provided by the present invention: A method for constructing an AI large model corpus based on digital experience, including:

[0007] Based on the constructed multi-source data acquisition channels, real-time acquire the structured data of the target problem after preprocessing and experience digitization;

[0008] According to the structural characteristics of the structured data of the target problem, store it in hierarchical levels in the database to generate an AI large model corpus.

[0009] The principle and advantages of the present invention are as follows: In this application, first, aiming at the source of the data in the corpus, it is acquired through multi-source channels, and the acquired data is preprocessed and digitized to generate structured data. Subsequently, according to the structural characteristics of the structured data, it is stored in hierarchical levels in the database to generate the corpus of the AI large model. Therefore, in this solution, the data in the corpus is structured data, which greatly improves the quality of the data. At the same time, compared with the data storage method in the prior art, when the AI large model reads the data in the corpus, the efficiency is higher.

[0010] Further, the specific process of real-time acquiring the structured data of the target problem after preprocessing and experience digitization based on the constructed multi-source data acquisition channels is as follows:

[0011] Acquire problem data through the constructed multi-source data acquisition channels;

[0012] After preprocessing the acquired problem data, generate problem-processed data;

[0013] Extract the structural characteristics of the problem-processed data, perform structured association on the problem-processed data according to the structural characteristics, and perform experience digitization on the problem-processed data after structured association.

[0014] Beneficial effects: In this solution, aiming at the acquisition source of the problem data, by constructing multi-source data acquisition channels, the coverage of the problem data acquisition channels is wide, which is convenient for users to generate problem data in various ways; and the preprocessing of the problem data is convenient for unifying the formats and making the content complete among multi-source problem data, which is convenient for identification and processing; the final experience digitization operation makes the problem data change from a redundant data structure type to a data structure type that is easy to identify and extract.

[0015] Further, the specific process of generating problem-processed data after preprocessing the acquired problem data is as follows:

[0016] Extract the problem information in the problem data to generate a problem phenomenon;

[0017] Extract the cause analysis corresponding to the problem phenomenon in the problem data and output the problem cause;

[0018] Extract the solutions generated from the problem data according to the problem causes, and output the problem measures;

[0019] Extract the evaluation results of the people and the processing process of the problem measures in the problem-solving process from the problem data, and output the personnel evaluation results;

[0020] Extract the evaluation data of the results of the problem measures from the problem data, and output the measure evaluation results.

[0021] Beneficial effects: In this solution, for the preprocessing of the problem data, innovatively generate the problem phenomenon, problem cause, problem measures and personnel evaluation results from the problem data in sequence. Among them, the problem phenomenon can clearly know the question purpose of the problem initiator or the questioner, the problem cause helps to improve the efficiency of the problem handler in analyzing problems, the problem measures can timely reflect the solution measures of the problem data, and the personnel evaluation results can reflect the satisfaction of the problem initiator or the questioner with the solution of the problem data. In this way, the complete processing process and the sense of clear hierarchy of the problem data are greatly improved.

[0022] Furthermore, after preprocessing the obtained problem data, generating the problem processing data specifically is as follows:

[0023] Extract the problem measures from the problem data;

[0024] Extract the corresponding problem phenomenon in the problem measures;

[0025] Extract the problem cause analyzed for the problem phenomenon in the problem measures;

[0026] Extract the evaluation results of the people and the processing process in the problem measures, and output the personnel evaluation results;

[0027] Extract the evaluation data of the results of the problem measures, and output the measure evaluation results.

[0028] Beneficial effects: In addition to the above-mentioned personnel evaluation results, it also includes the evaluation of the results of the measures. For example, after the problem data is processed, whether this measure solves the problem phenomenon. Therefore, the evaluation of the results of the measures represents the good or bad of the problem measures, which is helpful for the construction of the corpus.

[0029] Furthermore, the specific method of obtaining the problem data through the preset multi-source data acquisition channels is as follows:

[0030] Based on the constructed AI large model and / or robot and / or knowledge base, obtain the user's question data and the data of the question-solving process;

[0031] Based on the sharing platform, obtain the problem data to be analyzed among users and the data of the problem analysis process;

[0032] Summarize the user's question data, the data of the process of solving the question data, the problem data to be analyzed, and the data of the process of analyzing the problem to generate a problem data set, and multiple groups of problem data are included in the problem data set.

[0033] Beneficial effects: For multi-source data acquisition channels, in this solution, it includes methods such as AI large models, robots, knowledge bases, and sharing platforms. The problem data obtained from these channels can cover various types of problems. At the same time, different users have different ways of asking questions and focuses, and individual differences can be captured from the user's question data; and the question data among users on the sharing platform can reflect the interactions and view collisions among different user groups, providing multiple ideas for solving problems and helping to understand and solve problems from multiple perspectives.

[0034] Further, the problem information in the problem data includes the description information of the problem itself, the description information of the problem phenomenon, and the standard description information of the problem.

[0035] Beneficial effects: In the process of extracting the problem phenomenon, the diversity of problem information can present the problem phenomenon more comprehensively. For example, the description information of the problem itself can accurately locate the core of the problem, directly point out what the problem is, and at the same time clarify the problem itself, which is beneficial to constructing the causal relationship between problem phenomena; while the description information of the problem phenomenon can directly provide the basis for generating a complete problem phenomenon, which is conducive to verifying whether the generated problem phenomenon conforms to the actual situation and supplementing new details of the problem phenomenon; and the standard description information of the problem can provide a standardized framework to make the generated problem phenomenon meet certain standards.

[0036] Further, the problem measures include the execution tasks generated when solving the problem and the proposed solution generated when solving the problem.

[0037] Beneficial effects: Among the generated problem measures, they are composed in the form of execution tasks and proposed solutions. Among them, the execution tasks can enable the problem solver to clearly know the actions to be taken to deal with the problem, and the proposed solutions can provide the problem solver with multiple ideas and choices. At the same time, the form of execution tasks and solutions can well achieve the purposes of coordination, communication, and recording, making the content of the problem measures clear and the purpose clear.

[0038] Further, the hierarchical storage of the structure data of the target problem into the database according to the structural characteristics of the structure data of the target problem is specifically as follows:

[0039] Extract the structural characteristics of the structured data of the target problem data, and the structural characteristics include problem phenomena, problem causes, problem measures, personnel evaluation results, and measure evaluation results;

[0040] According to the structural characteristics, the problem phenomenon is used as the first level, the problem cause as the second level, the problem measure as the third level, the personnel evaluation result as the fourth level, and the measure evaluation result as the fifth level to construct an AI large model corpus.

[0041] Beneficial effects: The corpus stored hierarchically with the structural characteristics of this solution can, on the one hand, improve the understanding ability of the AI large model because the hierarchical structure of this solution enables the AI large model to more comprehensively understand the essence of the problem, and at the same time, the association between each level helps the model learn the causal relationship and logical order of the problem; on the other hand, it can enhance the model training effect because the hierarchical structure is convenient for more accurate data annotation, and at the same time, diverse training samples are generated by combining data from different levels.

[0042] An AI large model corpus construction system based on digital experience, which is applied to the above-mentioned AI large model corpus construction method based on digital experience, includes:

[0043] Data acquisition module: used to obtain target problem data in real time based on building a multi-source data acquisition channel;

[0044] Data preprocessing module: used to preprocess the obtained target problem data to generate preprocessed problem data;

[0045] Experience digitization module: used to digitize the preprocessed problem data to generate structured problem data;

[0046] Corpus generation module: used to store hierarchically into the database according to the structural characteristics of the structured problem data to generate an AI large model corpus.

[0047] A computer program product, which includes a program and a development environment for the program to execute, and the program enables a computer to execute the above-mentioned AI large model corpus construction method based on digital experience. Brief Description of the Drawings

[0048] Figure 1 It is a flowchart of an embodiment of the present invention;

[0049] Figure 2 It is a schematic diagram of the corpus generation process of an embodiment of the present invention. Detailed Description of the Specific Embodiment

[0050] The following is further detailed through specific embodiments:

[0051] The embodiment is basically as shown in the appendix Figure 1 and Figure 2 shown: An AI large model corpus construction method based on digital experience includes the following steps, which are respectively:

[0052] Step 1: Obtain the structured data after preprocessing and experience digitization of the target problem in real time based on the constructed multi-source data acquisition channels; specifically, for better implementation of the detailed implementation steps of the above Step 1, the sub-steps are specifically as follows:

[0053] S1: Obtain problem data through the constructed multi-source data acquisition channels;

[0054] S2: After preprocessing the obtained problem data, generate problem processing data;

[0055] S3: Extract the structural features of the problem processing data, perform structured association on the problem processing data according to the structural features, and perform experience digitization on the structurally associated problem processing data.

[0056] Among them, S1 includes:

[0057] S1-1: Obtain user question data and question data solving process data based on the constructed AI large model and / or robot and / or knowledge base;

[0058] S1-2: Obtain user-to-user problem data to be analyzed and problem analysis process data based on the sharing platform;

[0059] S1-3: Aggregate the user question data, question data solving process data, problem data to be analyzed, and problem analysis process data to generate a problem data set, and the problem data set includes multiple groups of problem data.

[0060] In this embodiment, the constructed AI large model and robot already have the ability to interact with users initially. The knowledge base is the knowledge base called by the AI large model or robot during the human-computer interaction process. Therefore, at the user interaction front end of the AI large model and robot, various types of text inputs are received through methods such as input boxes, such as natural language questions, files, pictures, etc., so as to receive user question data. Subsequently, the AI large model and robot route it to the corresponding processing module according to the type and field of the question. The processing module generates answers according to the pre-written rule scripts (such as knowledge graphs, expert systems). The process of generating answers and the subsequent continuous questioning process is the process of solving question data. During the solving process, by recording the key intermediate steps, such as which knowledge entries in the knowledge base are called by the rule script, which neuron layers the AI large model has passed through during answer generation, and the output results of each neuron layer, etc., these solving processes are recorded, thereby generating question data solving process data. Finally, the user question data and the question data solving process data are associated to establish a one-to-one or one-to-many relationship.

[0061] The process of obtaining the problem data to be analyzed and the problem analysis process data among users through the sharing platform is as follows: First, call the internal sharing platform of the enterprise or the sharing platform of the enterprise and external collaborators. These sharing platforms include user registration and login modules, problem publishing modules, problem classification modules, comment and reply modules, etc. Users log in to the sharing platform through the registration and login module and are assigned different permissions. Subsequently, they publish the problem to be analyzed through the problem publishing module. During the publishing process, by guiding users to fill in the necessary information of the problem, such as the field to which the problem belongs, the type of the problem, the title of the problem, the description of the problem, etc., the problem classification module then classifies the content of the problem to be analyzed to be published according to the preset classification rules, and collects the relevant attachment content and detailed information of the problem to be analyzed, such as video files, picture files, etc., which can facilitate other users to view the problem scenario more intuitively. Subsequently, through the comment and reply module, obtain the comments and replies of other users on the problem to be analyzed published, as well as the adoption records of the publisher. In this way, a complete set of problem data to be analyzed and problem analysis process data is formed.

[0062] Finally, summarize the above user question data, the process data for solving the question data, the problem data to be analyzed, and the problem analysis process data to generate a problem dataset. Therefore, in this solution, the established problem data acquisition channels include methods such as AI large models, robots, knowledge bases, sharing platforms, etc. The problem data obtained from these channels can cover various types of problems. At the same time, different users have different ways of asking questions and focuses, and individual differences can be captured from the user question data; while the problem data among users on the sharing platform can reflect the interactions and view collisions among different user groups, providing multiple ideas for solving problems and helping to understand and solve problems from multiple perspectives.

[0063] In S2, after preprocessing the obtained problem data, the specific generation of problem processing data is as follows:

[0064] Extract the problem information in the problem data to generate a problem phenomenon;

[0065] Extract the cause analysis corresponding to the problem phenomenon in the problem data and output the problem cause;

[0066] Extract the solution generated according to the problem cause in the problem data and output the problem measure;

[0067] Extract the evaluation results of the people and the processing process in the problem-solving process of the problem measure in the problem data and output the personnel evaluation results;

[0068] Extract the evaluation data of the result of the problem measure in the problem data and output the measure evaluation results.

[0069] In this embodiment, for the preprocessing of the obtained problem data, it is first necessary to extract the problem information in the problem data and generate a problem phenomenon based on the problem information. The problem information includes the description information of the problem itself, the description information of the problem phenomenon, and the standard description information of the problem. Among them, the description information of the problem itself refers to explaining what the problem is in the problem data through concise language. For example, problem descriptions such as "The printer cannot print documents"; the description information of the problem phenomenon refers to the description of various specific situations that can be observed when the problem occurs. For example, for the problem of "The printer cannot print documents", the problem phenomenon description is "The power indicator of the printer flashes, the paper tray has a paper feeding action, but the paper is not printed"; the standard description information of the problem refers to the description of the normal state of things or the problem judgment criteria. Taking the printer as an example, the standard description of the problem may include the range of the normal printing speed of the printer (such as printing 10-20 pages per minute), the normal paper feeding angle, the printing quality standard (such as a resolution of 1200 dpi, etc.). These standards are the reference basis for measuring whether things are in a normal state. When the actual situation deviates from these standards, problems may occur.

[0070] Subsequently, the cause analysis corresponding to the problem phenomenon in the problem data is extracted, and the problem cause is output. In this embodiment, since the problem phenomenon in the problem data is proposed by the problem originator, the corresponding cause analysis is obtained by the problem receiver analyzing the possible causes of the problem phenomenon using the existing knowledge and experience. For example, taking the printer as an example, when the problem phenomenon is that the printer jams, the cause analysis given by the problem receiver may be improper user operation, printer part failure, etc. In this way, the problem cause corresponding to the problem phenomenon can be obtained.

[0071] Next, extract the solutions generated according to the problem causes from the problem data and output the problem measures. In this embodiment, the problem measures include the execution tasks generated when solving the problem and the proposed solution plans generated when solving the problem. Among them, the advantage of taking the execution task as one of the problem measures is that, firstly, clearly defining the specific execution task can enable the problem solver to clearly know the actions to be taken. For example, if the problem is the decline in product quality and the reason is the aging of production equipment, at this time, the execution task can be "conduct a comprehensive overhaul of the aging equipment and replace the severely worn parts". Therefore, the specific execution task can directly point to the root cause of the problem and avoid the ambiguity of the solution measures; secondly, after listing the execution tasks in detail, team members can better divide the work and cooperate. For example, when solving a problem, one operation can be performed by person A and another operation can be performed by person B. In this way, the clear division of labor in the execution task can avoid work delays caused by unclear responsibilities and make the entire problem-solving process more efficient. Moreover, the detailed task content helps to track and evaluate the progress. After solving each task item by item, problems in the process of executing the task can be found and adjusted in a timely manner; finally, the detailed execution task is very important in internal team communication and cross-departmental communication. When it is necessary to explain to other members or departments how to solve the problem, the other party can quickly understand the key points and requirements of the work through the clear execution task list.

[0072] For the proposed solution plans, firstly, the proposed solution plans provide multiple ideas and options for problem-solving. The person who proposes the suggestions can put forward different solution plans based on their work experience, and the executor can choose the most suitable plan according to the situation, which also provides a reference for further optimizing the problem-solving strategy; secondly, reasonable proposed solution plans can guide the problem-solving towards a more optimized direction. For example, when solving the problem of the decline in the enterprise's market share, the proposed solution plans may include "1. Conduct market research to understand the new product features and market strategies of competitors; 2. Optimize the product packaging to make it more in line with the aesthetics and usage habits of current consumers; 3. Strengthen online marketing and cooperate with well-known internet celebrities for product promotion". These detailed suggestions can enable the enterprise to start from multiple perspectives and comprehensively consider various factors, thereby more effectively increasing the market share; finally, the detailed proposed solution plans help in communication in meetings and discussions at different levels. In the high-level decision-making meetings of the enterprise, being able to provide proposed solution plans with depth and details can enable decision-makers to better evaluate the risks and benefits of each plan. For example, when discussing the problem of the company expanding into new business, the detailed proposed solution plans include market size prediction, estimated capital investment required, talent demand analysis, etc., all of which can provide strong evidence for decision-making.

[0073] After the above problem measures are output, extract the evaluation results of the people and the handling process during the problem-solving process, and output the personnel evaluation results; as well as the evaluation data of the results of the problem measures, and output the measure evaluation results. Therefore, in this embodiment, the evaluation data includes two dimensions, one is the personnel evaluation result, and the other is the measure evaluation result. Among them, the personnel evaluation result is obtained through self-evaluation, peer evaluation by colleagues, superior evaluation, and customer feedback. In self-evaluation, the personnel participating in the problem-solving are required to evaluate their own performance during the task execution according to the preset evaluation criteria. The evaluation criteria may include the timeliness of task completion, work quality, cooperation and communication with team members, etc., and are evaluated using scores; peer evaluation by colleagues can be mutual evaluation among team members. For example, if multiple people participate in the problem-solving process and are responsible for different aspects, they will evaluate each other's problem-solving quality in the corresponding aspects; superior evaluation can be carried out by the superior leader in charge of the problem-solving project or the relevant work field. Based on a clear understanding of the overall project goals, resource allocation, and team member responsibilities, the superior comprehensively evaluates the performance of the subordinates during the problem-solving process; customer feedback is for the problem-solving process involving external customers, and the customer's feedback can indirectly reflect the work effectiveness of the relevant personnel.

[0074] The measure evaluation result includes quantitative index evaluation and qualitative index evaluation. Among them, the quantitative index evaluation includes time dimension evaluation, cost dimension evaluation, and result quantity dimension evaluation. The time dimension evaluation is the time spent from the implementation of the problem measure to the solution of the problem. The cost dimension evaluation is to calculate the capital cost consumed by implementing the problem measure. The result quantity dimension evaluation is to obtain the evaluation by counting relevant data for some problems with clearly quantifiable results.

[0075] The qualitative index evaluation includes effectiveness evaluation, adaptability evaluation, and sustainability evaluation. Among them, the effectiveness evaluation is to judge whether the problem measure targets the core of the problem; the adaptability evaluation is whether the problem measure is applicable to the internal rules and regulations of the enterprise or unit, etc. The sustainability evaluation is whether the problem measure can play a role in the long term and maintain a stable effect.

[0076] Therefore, after obtaining the structured data after the above-mentioned experience digitization, step two is executed, specifically:

[0077] Step two: Store the structured data of the target problem in layers according to the structural characteristics of the structured data of the target problem in the database to generate an AI large model corpus. Among them, storing the structured data of the target problem in layers in the database according to the structural characteristics of the structured data of the target problem is specifically:

[0078] Extract the structural characteristics of the structured data of the target problem data, and the structural characteristics include problem phenomenon, problem cause, problem measure, personnel evaluation result, and measure evaluation result;

[0079] According to the structural characteristics, the problem phenomenon is taken as the first level, the problem cause as the second level, the problem measures as the third level, the personnel evaluation result as the fourth level, and the measure evaluation result as the fifth level to construct the AI large model corpus.

[0080] In this embodiment, the constructed AI large model corpus is stored in a hierarchical manner. On the one hand, it can improve the understanding ability of the AI large model because the hierarchical structure of this solution enables the AI large model to comprehensively understand the essence of the problem, and the association between levels helps the model learn the causal relationship and logical order of the problem. On the other hand, it can enhance the model training effect because the hierarchical structure facilitates more accurate data annotation and generates diverse training samples by combining data from different levels.

[0081] In another embodiment of this embodiment, there is also an AI large model corpus construction system based on digital experience, which is applied to the above-mentioned method for constructing an AI large model corpus based on digital experience, and includes:

[0082] Data acquisition module: used to obtain target problem data in real time based on building a multi-source data acquisition channel;

[0083] Data preprocessing module: used to preprocess the obtained target problem data to generate preprocessed problem data;

[0084] Experience digitization module: used to digitize the preprocessed problem data to generate structured problem data;

[0085] Corpus generation module: used to store the structured problem data in a hierarchical manner in the database according to its structural characteristics to generate the AI large model corpus.

[0086] There is also a computer program product, which includes a program and a development environment for the program to execute. The program enables the computer to execute the above-mentioned method for constructing an AI large model corpus based on digital experience.

[0087] The computer program product can be written in any combination of one or more programming languages for the program code to execute the operations of the embodiments of the present invention. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed completely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or completely on a remote computing device or server.

[0088] In addition, an embodiment of the present invention may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps of a method for constructing an AI large model corpus based on digital experience provided by any embodiment of the present invention.

[0089] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0090] Embodiment 2:

[0091] The difference between Embodiment 2 and Embodiment 1 is that in Embodiment 2, after preprocessing the obtained problem data, generating the problem processing data specifically includes:

[0092] Extracting the problem measures in the problem data;

[0093] Extracting the corresponding problem phenomena in the problem measures;

[0094] Extracting the problem causes analyzed for the problem phenomena in the problem measures;

[0095] Extracting the evaluation results of people and the processing process in the problem measures and outputting the personnel evaluation results;

[0096] Extracting the evaluation data of the results of the problem measures and outputting the measure evaluation results.

[0097] In this embodiment, in addition to structuring the problem data in the manner of problem phenomena, problem causes, problem measures, and problem evaluations disclosed in Embodiment 1 above, it also includes a way of directly creating an execution task. For example, the user directly creates an execution task on the sharing platform and distributes it to the corresponding problem solvers. The content of the execution task generated in this type of way includes problem phenomena, problem causes, etc. Therefore, by directly evaluating the problem measures, that is, the above-mentioned execution tasks, the problem measures, problem phenomena, problem causes, and problem evaluations thus formed are also a way of structured data.

[0098] The above are only embodiments of the present invention. Common knowledge such as specific structures and characteristics known in the art is not described in detail herein. Those of ordinary skill in the art know all the common general technical knowledge in the technical field to which the invention pertains before the filing date or the priority date, are able to learn all the prior art in this field, and have the ability to apply the conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given in this application, complete and implement this solution in combination with their own abilities. Some typical well-known structures or well-known methods should not become an obstacle for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can still be made, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be based on the content of its claims, and the specific implementation manners and the like recorded in the specification can be used to interpret the content of the claims.

Claims

1. A method for constructing an AI large model corpus based on digital experience, characterized in that: include: Based on the constructed multi-source data acquisition channel, the structured data of the target problem after preprocessing and empirical digitization is acquired in real time; According to the structural characteristics of the structured data of the target problem, it is stored hierarchically in the database to generate an AI large model corpus.

2. The method for constructing a large AI model corpus based on digital experience according to claim 1, characterized in that: The structured data of the target problem obtained in real time based on the constructed multi-source data acquisition channel after preprocessing and empirical digitization are specifically: Acquire problem data through the constructed multi-source data acquisition channel; After preprocessing the acquired problem data, problem processing data is generated; The structural features of the problem processing data are extracted, the problem processing data are structurally associated according to the structural features, and the problem processing data after the structural association is empirically digitized.

3. The method for constructing a large AI model corpus based on digital experience according to claim 2, characterized in that: After the acquired problem data is preprocessed, the problem processing data is generated specifically as follows: Extract problem information from problem data and generate problem phenomena; Extract the cause analysis corresponding to the problem phenomenon in the problem data and output the cause of the problem; Extract solutions generated from problem data based on the causes of the problem and output problem measures; Extract the evaluation results of people and the processing process of problem measures in the problem solving process from the problem data, and output the personnel evaluation results; Evaluation data on the results of measures taken against the problem are extracted from the problem data, and the measure evaluation results are output.

4. The method for constructing a large AI model corpus based on digital experience according to claim 2, characterized in that: After the acquired problem data is preprocessed, the problem processing data is generated specifically as follows: Extract problem measures from problem data; Extract the corresponding problem phenomena from the problem measures; Extract the causes of the problem from the analysis of the problem phenomenon in the problem measures; Extract the evaluation results of people and processing processes in problem measures and output the personnel evaluation results; The evaluation data of the results of the problem measures are extracted, and the measure evaluation results are output.

5. The method for constructing a large AI model corpus based on digital experience according to claim 2, characterized in that: The specific method of obtaining problem data through the preset multi-source data acquisition channel is as follows: Acquire user question data and question data solution process data based on the built AI big model and / or robot and / or knowledge base; Based on the sharing platform, we can obtain the problem data to be analyzed and the problem analysis process data between users; The user question data, question data solving process data, problem data to be analyzed and problem analysis process data are aggregated to generate a problem data set, wherein the problem data set includes multiple groups of problem data.

6. A method for constructing an AI large model corpus based on digital experience according to claim 3 or 4, characterized in that: The problem information in the problem data includes description information of the problem itself, description information of the problem phenomenon and description information of the problem standard.

7. The method for constructing a large AI model corpus based on digital experience according to claim 6, characterized in that: The problem measures include execution tasks generated when solving the problem and suggested solutions generated when solving the problem.

8. The method for constructing a large AI model corpus based on digital experience according to claim 7, characterized in that: The structure characteristics of the structure data of the target problem are stored hierarchically in the database as follows: Extracting structural features of structured data of target problem data, wherein the structural features include problem phenomenon, problem cause, problem measures, personnel evaluation results and measure evaluation results; According to the structural characteristics, the problem phenomenon is taken as the first level, the cause of the problem is taken as the second level, the problem measures are taken as the third level, the personnel evaluation results are taken as the fourth level, and the measure evaluation results are taken as the fifth level to construct the AI ​​large model corpus.

9. A system for constructing a large AI model corpus based on digital experience, applied to a method for constructing a large AI model corpus based on digital experience as described in any one of claims 1 to 8 above, characterized in that: include: Data acquisition module: used to acquire target problem data in real time based on building multi-source data acquisition channels; Data preprocessing module: used to preprocess the acquired target problem data and generate preprocessed problem data; Experience digitization module: used to digitize the preprocessed problem data to generate structured problem data; Corpus generation module: used to generate AI large model corpus by storing it hierarchically in the database according to the structural characteristics of structured problem data.

10. A computer program product, characterized in that: It includes a program and a development environment for executing the program, and the program enables a computer to execute a method for constructing an AI large model corpus based on digital experience as described in any one of claims 1 to 8 above.

Citation Information

Cited By

  • Intelligent corpus construction method and system based on digital experience

    CN120973925A