Assessment object evaluation method, evaluation system and electronic equipment
By combining multi-source data processing and a large language model, evaluation prompts are generated and results are output, solving the consistency and efficiency problems of evaluation of assessment subjects in existing technologies, and realizing efficient and accurate evaluation of assessment subjects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA SOUTHERN AIRLINES DIGITAL TECHNOLOGY (GUANGDONG) CO LTD
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-08
AI Technical Summary
In existing technologies, enterprises rely on manual qualitative judgment for subjective evaluation of assessment subjects, resulting in poor consistency of evaluation results, inability to process large-scale assessment materials in batches, long evaluation cycles, and low efficiency.
By acquiring multi-source data, extracting target text information, combining evaluation knowledge base and large language model to generate evaluation prompt words, using large language model to output evaluation results, integrating multi-dimensional data to improve the comprehensiveness and accuracy of evaluation basis, and using evaluation knowledge base to ensure the consistency and rationality of evaluation logic.
It has automated and intelligentized the evaluation of assessment subjects, improved the accuracy and efficiency of evaluation results, ensured the consistency and rationality of evaluation logic, and enhanced the comprehensiveness and efficiency of evaluation work.
Smart Images

Figure CN121998489A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to an evaluation method, evaluation system and electronic device for an evaluation object. Background Technology
[0002] As enterprises continue their digital transformation, performance evaluation of managers is gradually shifting from a traditional "human experience-driven" model to a "data-driven" model. Subjective evaluations, as the core vehicle for depicting the non-quantifiable characteristics of those being evaluated, directly impact the accuracy of profile building and efficient screening. However, current evaluation methods still have many shortcomings, failing to meet the demands of performance evaluation for high-precision, high-efficiency, and highly consistent labeling. Summary of the Invention
[0003] This invention aims to provide an evaluation method, system, and electronic device for assessing subjects, capable of integrating multi-source data for analysis and evaluation, thereby improving the accuracy and efficiency of evaluation results. Firstly, this application provides an evaluation method for an assessment subject, characterized by the following steps: acquiring multi-source data and subjective evaluation tags of the target assessment subject; extracting target text information from the multi-source data; using the target text information to reflect the work performance of the target assessment subject; generating evaluation prompt words for the target assessment subject based on the target text information, the target subjective evaluation tags, and an evaluation knowledge base; wherein the evaluation knowledge base is used to store evaluation rules and evaluation cases corresponding to different subjective evaluation tags; and obtaining the evaluation result of the target assessment subject by calling a large language model based on the evaluation prompt words.
[0004] The evaluation method for assessment subjects provided in this application combines multi-source data with target subjective evaluation tags to extract target text information reflecting the work performance of the assessment subjects. It generates evaluation prompt words using an evaluation knowledge base that stores evaluation rules and cases, and then calls a large language model to output the evaluation results. This not only improves the comprehensiveness and accuracy of the evaluation basis by combining multi-dimensional data, but also ensures the consistency and rationality of the evaluation logic through the standardized guidance of the evaluation knowledge base. At the same time, it effectively improves the overall efficiency of the evaluation work by leveraging the automated generation capability of the large language model.
[0005] In some embodiments, based on target text information, target subjective evaluation tags, and an evaluation knowledge base, evaluation prompt words for the target assessment object are generated, including: retrieving evaluation rules and evaluation cases corresponding to the target subjective evaluation tags from the evaluation knowledge base; integrating the target text information, the evaluation rules and evaluation cases corresponding to the target subjective evaluation tags into a preset prompt word template to generate evaluation prompt words for the target assessment object.
[0006] In some embodiments, extracting target text information from multi-source data includes: removing redundant data from the multi-source data to obtain initial data; converting the initial data into unified text data; and extracting target text information from the text data through semantic analysis.
[0007] In some embodiments, after obtaining the evaluation results of the target assessment object, the method further includes: pushing the evaluation results of the target assessment object to an auditing terminal so that the auditing personnel can audit the evaluation results of the target assessment object; receiving the auditing results fed back by the auditing terminal; the auditing results are used to indicate whether the evaluation results of the target assessment object have passed the audit; if the auditing results indicate that the evaluation results of the target assessment object have passed the audit, determining that the evaluation results of the target assessment object are effective, and pushing the evaluation results of the target assessment object to the user terminal; if the auditing results indicate that the evaluation results of the target assessment object have not passed the audit, determining that the evaluation results of the target assessment object are invalid, and iteratively correcting the evaluation results of the target assessment object until the evaluation results of the target assessment object have passed the audit.
[0008] In some embodiments, the evaluation results of the target assessment object are iteratively corrected, including: in each iteration, the evaluation prompt words are corrected based on the modification opinions of the assessors to obtain the corrected evaluation prompt words; and the corrected evaluation results of the target assessment object are obtained based on the corrected evaluation prompt words.
[0009] In some embodiments, the method further includes: adjusting the review priority of the evaluation results of the target assessment object based on the confidence level of the evaluation results of the target assessment object; wherein the confidence level is determined based on the semantic matching degree between the target text information corresponding to the evaluation prompt words of the target assessment object and the target subjective evaluation label; wherein the semantic matching degree is determined based on a large language model.
[0010] In some embodiments, the method further includes: upon detecting an update operation of the evaluation knowledge base, generating a new prompt word template based on the updated content of the evaluation knowledge base; wherein the updated content includes newly added subjective evaluation tags, as well as the evaluation rules and evaluation cases corresponding to the newly added subjective evaluation tags.
[0011] Secondly, this application provides an evaluation system for assessment subjects, which includes: a multi-source data processing module, an intelligent agent, and a large language model invocation module; the multi-source data processing module is used to acquire multi-source data of the target assessment subject, extract target text information from the multi-source data, and send it to the intelligent agent; the target text information is used to reflect the work performance of the target assessment subject; the intelligent agent is used to acquire the target subjective evaluation tags and target text information of the target assessment subject, and generate evaluation prompt words for the target assessment subject based on the target text information, target subjective evaluation tags, and evaluation knowledge base; the large language model invocation module is used to invoke the large language model, input the evaluation prompt words of the target assessment subject into the large language model, obtain the evaluation result of the target assessment subject output by the large language model, and return the evaluation result of the target assessment subject to the intelligent agent.
[0012] The evaluation system for assessment subjects provided in this application embodiment automates and intelligentizes the assessment process through the collaborative work of a multi-source data processing module, an intelligent agent, and a large language model invocation module. Specifically, the multi-source data processing module accurately extracts target text information reflecting the performance of the assessment subject, providing solid data support for the evaluation; the intelligent agent, relying on an evaluation knowledge base, generates targeted evaluation prompts by combining target text information and subjective evaluation tags, ensuring the standardization and consistency of the evaluation logic; and the large language model invocation module outputs evaluation results by invoking a large language model, effectively improving the efficiency of the evaluation work. The overall system integrates the objective dimension of multi-source data and the subjective dimension of subjective evaluation tags, and leverages the natural language processing capabilities of the large language model to automatically generate evaluation results, significantly improving the comprehensiveness, accuracy, and efficiency of the evaluation work.
[0013] In some embodiments, the system further includes: an evaluation knowledge base module, a result review module, and a result output module; the evaluation knowledge base module is used to store evaluation rules and evaluation cases corresponding to different subjective evaluation tags; the result review module is used to review the evaluation results of the target assessment object; and the result output module is used to push the evaluation results of the assessment object to the user terminal.
[0014] Thirdly, this application provides an evaluation device, which includes: an acquisition module and a processing module; the acquisition module is used to acquire multi-source data of the target assessment object and target subjective evaluation tags; the processing module is used to extract target text information from the multi-source data; the target text information is used to reflect the work performance of the target assessment object; based on the target text information, target subjective evaluation tags and an evaluation knowledge base, evaluation prompt words for the target assessment object are generated; wherein, the evaluation knowledge base is used to store evaluation rules and evaluation cases corresponding to different subjective evaluation tags; based on the evaluation prompt words, the evaluation result of the target assessment object is obtained by calling a large language model.
[0015] In some embodiments, the processing module is specifically used to retrieve the evaluation rules and evaluation cases corresponding to the target subjective evaluation tags from the evaluation knowledge base; integrate the target text information, the evaluation rules and evaluation cases corresponding to the target subjective evaluation tags into a preset prompt word template, and generate evaluation prompt words for the target assessment object.
[0016] In some embodiments, the processing module is specifically used to remove redundant data from multi-source data to obtain initial data; to convert the initial data into unified text data; and to extract target text information from the text data through semantic analysis.
[0017] In some embodiments, after obtaining the evaluation results of the target assessment object, the processing module is further configured to push the evaluation results of the target assessment object to the review terminal so that the reviewer can review the evaluation results of the target assessment object; receive the review results fed back by the review terminal; the review results are used to indicate whether the evaluation results of the target assessment object have passed the review; if the review results indicate that the evaluation results of the target assessment object have passed the review, determine that the evaluation results of the target assessment object are effective and push the evaluation results of the target assessment object to the user terminal; if the review results indicate that the evaluation results of the target assessment object have not passed the review, determine that the evaluation results of the target assessment object are invalid and iteratively correct the evaluation results of the target assessment object until the evaluation results of the target assessment object have passed the review.
[0018] In some embodiments, the processing module is specifically used to revise the evaluation prompts based on the modification opinions of the evaluators during each iteration, and obtain the revised evaluation prompts; based on the revised evaluation prompts, obtain the revised evaluation result of the target evaluation object.
[0019] In some embodiments, the processing module is further configured to adjust the review priority of the evaluation results of the target assessment object based on the confidence level of the evaluation results of the target assessment object; wherein, the confidence level is determined based on the semantic matching degree between the target text information corresponding to the evaluation prompt words of the target assessment object and the target subjective evaluation label; wherein, the semantic matching degree is determined based on a large language model.
[0020] In some embodiments, the processing module is further configured to generate a new prompt word template based on the updated content of the evaluation knowledge base when an update operation of the evaluation knowledge base is detected; wherein the updated content includes newly added subjective evaluation tags, as well as the evaluation rules and evaluation cases corresponding to the newly added subjective evaluation tags.
[0021] Fourthly, this application provides an electronic device comprising: a processor and a memory; the memory storing processor-executable instructions; when the processor is configured to execute the instructions, causing the electronic device to implement the method of the first aspect described above.
[0022] Fifthly, this application provides a computer-readable storage medium comprising: computer software instructions; which, when executed in an electronic device, cause the electronic device to implement the method described in the first aspect.
[0023] Sixthly, this application provides a computer program product comprising a computer program; when the computer program is run in an electronic device, it causes the electronic device to implement the method described in the first aspect.
[0024] The beneficial effects of the second to sixth aspects mentioned above can be referred to the corresponding descriptions of the first or second aspects, and will not be repeated here. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 A schematic diagram of the architecture of an evaluation system for assessment subjects provided in this application; Figure 2 A schematic diagram of the architecture of an evaluation system for another assessment object provided in this application; Figure 3 A flowchart of an evaluation method for assessment subjects provided in this application; Figure 4 A flowchart of an alternative evaluation method for the assessment object provided in this application; Figure 5 A flowchart of another evaluation method for assessment subjects provided in this application; Figure 6 A flowchart of another evaluation method for assessment subjects provided in this application; Figure 7 A flowchart of another evaluation method for assessment subjects provided in this application; Figure 8 A flowchart of another evaluation method for assessment subjects provided in this application; Figure 9 This is a schematic diagram of the structure of an evaluation device provided in an embodiment of this application; Figure 10This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] It should be noted that in the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0029] In the embodiments of this application, the terms "first," "second," "third," "fourth," "fifth," and "sixth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," "third," "fourth," "fifth," and "sixth" may explicitly or implicitly include one or more of that feature.
[0030] In embodiments of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0031] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.
[0032] Currently, companies rely heavily on manual qualitative judgments based on assessment materials to evaluate candidates' subjective performance. There is a lack of unified evaluation standards for subjective evaluations such as the overall performance of candidates. This not only makes them susceptible to human bias, resulting in inconsistent evaluation results, but also makes it impossible to process large-scale assessment materials in batches, leading to long evaluation cycles and low efficiency.
[0033] To address the aforementioned technical issues, this application provides an evaluation method, system, and electronic device for assessing candidates. This system integrates multi-source data for analysis and evaluation, improving the accuracy and efficiency of the evaluation results. The method's approach is as follows: First, acquire multi-source data and subjective evaluation tags of the target candidate. Second, extract target text information from the multi-source data. Third, use the target text information to reflect the target candidate's work performance. Fourth, based on the target text information, subjective evaluation tags, and an evaluation knowledge base, generate evaluation prompts for the target candidate. The evaluation knowledge base stores evaluation rules and cases corresponding to different subjective evaluation tags. Fifth, based on the evaluation prompts, obtain the evaluation results for the target candidate by calling a large language model.
[0034] The following description, in conjunction with the accompanying drawings, introduces an evaluation method, evaluation system, and electronic equipment for assessing test subjects provided in this application.
[0035] Figure 1 This application provides a schematic diagram of the architecture of an evaluation system for assessment objects, as shown in the embodiments of this application. Figure 1 As shown, the evaluation system 100 for the assessment object includes a multi-source data processing module 101, an intelligent agent 102, and a large language model calling module 103. The modules are connected to each other through communication links.
[0036] The multi-source data processing module 101 is used to acquire multi-source data of the target assessment object, extract target text information from the multi-source data, and send it to the intelligent agent. The target text information reflects the work performance of the target assessment object.
[0037] In some embodiments, the multi-source data processing module 101 is deployed on an application server, and establishes a communication connection with the server where the data management system of the target assessment object is located through an API interface to obtain the multi-source data of the target object.
[0038] Intelligent agent 102 is used to obtain the target subjective evaluation tags and target text information of the target assessment object, and generate evaluation prompt words for the target assessment object based on the target text information, target subjective evaluation tags and evaluation knowledge base.
[0039] In some embodiments, the intelligent agent 102 is deployed on an application server and can retrieve the evaluation rules and case information of subjective evaluation tags from the evaluation knowledge base module 104 to provide data support for the generation of evaluation prompt words.
[0040] In some embodiments, agent 102 sends evaluation prompts to the large language model invocation module 103. Simultaneously, this module can receive evaluation results returned by the large language model invocation module 104, and, based on human review feedback from the tag review module 105, correct and optimize the evaluation prompts to continuously improve the adaptability of subsequent prompts to the large model.
[0041] The large language model calling module 103 is used to call the large language model, input the evaluation prompt words of the target assessment object provided by the intelligent agent 102 into the large language model, obtain the evaluation result of the target assessment object output by the large language model, and return the evaluation result of the target assessment object to the intelligent agent 102.
[0042] In some embodiments, the large language model invocation module 104 is deployed on a GPU server and has built-in mainstream large language models (such as Transformer architecture-derived models) and standardized model invocation interfaces to provide basic support for model invocation and data interaction.
[0043] The evaluation system for assessment subjects provided in this application, through the collaborative work of a multi-source data processing module, an intelligent agent, and a large language model invocation module, automates and intelligentizes the assessment process. Specifically, the multi-source data processing module accurately extracts target text information reflecting the performance of the assessment subjects, providing solid data support for the evaluation; the intelligent agent, relying on an evaluation knowledge base, generates targeted evaluation prompts by combining target text information and subjective evaluation tags, ensuring the standardization and consistency of the evaluation logic; and the large language model invocation module outputs evaluation results by invoking a large language model, effectively improving the efficiency of the evaluation work. The overall system integrates the objective dimension of multi-source data and the subjective dimension of subjective evaluation tags, and leverages the natural language processing capabilities of the large language model to automatically generate evaluation results, significantly improving the comprehensiveness, accuracy, and efficiency of the evaluation work.
[0044] Figure 2 A schematic diagram of the architecture of an evaluation system for another assessment object provided in an embodiment of this application, such as... Figure 2 As shown, in Figure 1 Based on this, the evaluation system 100 for the assessed subjects also includes an evaluation knowledge base module 104, a result review module 105, and a result output module 106.
[0045] The evaluation knowledge base module 104 is used to store evaluation rules and evaluation cases corresponding to different subjective evaluation tags. The evaluation cases include excellent cases and incorrect cases. The excellent case library stores typical behavioral cases that match each subjective evaluation tag, providing positive reference for tag matching. The incorrect case library stores typical behavioral cases that do not match each subjective evaluation tag, providing negative reference for tag matching.
[0046] In some embodiments, the evaluation knowledge base module 104 is deployed on a data server to provide standard basis and case support for the identification of subjective evaluation tags of assessment subjects, and supports administrators to maintain and update the data in the database through terminal devices.
[0047] The results review module 105 is used to review the evaluation results of the target assessment subjects.
[0048] In some embodiments, the result review module 105 is deployed on the application server and establishes a communication connection with the administrator terminal and the data server. Its core function is to realize the manual verification and feedback of the evaluation results.
[0049] In some embodiments, this module allows administrators to view the subjective evaluation labels, label recognition reasons, and corresponding original assessment materials generated by the large language model calling module 104 through terminal devices, providing comprehensive information support for manual review and facilitating accurate judgment by administrators.
[0050] The result output module 106 is used to push the evaluation results of the assessed individuals to the user terminal.
[0051] In some embodiments, the result output module 106 is deployed on an application server and establishes a communication connection with the user terminal through an HTTP interface, enabling the display, export, and traceability of evaluation results.
[0052] In some embodiments, the system's supporting hardware resources include application servers, data servers, GPU servers, administrator terminals, and user terminals. Each hardware resource is adapted to its corresponding functional module through standardized data interfaces, which can ensure the stability and security of data transmission between modules.
[0053] It should be noted that the system architecture and application scenarios of the embodiments in this application are not limited. The system architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of communication technology and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0054] Figure 3 A flowchart illustrating an evaluation method for an assessment object provided in this application embodiment, the method being applicable to, for example... Figure 1 or Figure 2 The evaluation system for the assessed individuals is shown, such as Figure 3 As shown, the method includes the following steps S101-S104: S101. Obtain multi-source data and subjective evaluation labels of the target assessment subjects.
[0055] In some embodiments, the target of performance evaluation may be the company's managers, ordinary employees, etc.
[0056] In some embodiments, subjective evaluation labels are pre-defined evaluation markers used to quantitatively or qualitatively describe the subjective performance of the target assessment subject. For example, subjective evaluation labels may take the form of: "Good work attitude" or "Strong innovation ability."
[0057] In some embodiments, multi-source data refers to various data sets originating from different business systems and used to characterize the work-related situations of the target assessment object. For example, multi-source data includes, but is not limited to, work report texts in office systems, project participation records in project management systems, and communication and interaction information in collaborative office platforms.
[0058] In some embodiments, multi-source data is acquired by calling the interface of the server where the multi-source data is located, and may include various types of material data. After acquisition, the data is transmitted to the data server for redundant data removal.
[0059] S102. Extract target text information from multi-source data.
[0060] The target text information is used to reflect the work performance of the target assessment object.
[0061] In some embodiments, the acquired multi-source data is subjected to structured parsing and text filtering to extract target text information; wherein, the target text information is text content that can directly or indirectly reflect the work performance of the target assessment object, including but not limited to descriptions of work results, descriptions of task execution processes, and explanations of problem-solving approaches.
[0062] In some embodiments, non-textual data and redundant text unrelated to work performance in multi-source data are filtered out through keyword matching, semantic association analysis, etc., to ensure that the extracted target text information has evaluation relevance and effectiveness.
[0063] S103. Based on the target text information, target subjective evaluation tags, and evaluation knowledge base, generate evaluation prompt words for the target assessment object.
[0064] The evaluation knowledge base is used to store evaluation rules and evaluation cases corresponding to different subjective evaluation labels.
[0065] In some embodiments, the evaluation knowledge base is a structured database pre-set within the data server for storing evaluation rules and evaluation cases corresponding to various subjective evaluation tags. The evaluation rules are used to determine whether the assessment object meets the evaluation conditions of the corresponding subjective evaluation tag, and the evaluation cases are historical assessment instances that have been reviewed and approved and are matched with the corresponding subjective evaluation tags.
[0066] S104. Based on the evaluation prompts, the evaluation results of the target assessment object are obtained by calling the large language model.
[0067] It should be understood that evaluation prompts refer to the instructional text generated and transmitted by the agent module to instruct the large language model to complete the evaluation and analysis task. The core content includes evaluation dimension requirements, label matching rules, etc. The large language model is an artificial intelligence model with deep semantic understanding capabilities, and its operation relies on the computing power provided by GPU servers. The evaluation results of the target assessment object refer to the structured data containing subjective evaluation labels and corresponding recognition reasons, which are used to characterize the subjective performance characteristics of the target assessment object.
[0068] In some embodiments, evaluation prompts transmitted by the intelligent agent are received through a preset interface, and target text information corresponding to the target assessment object is retrieved from the data server. Then, the received evaluation prompts and target text information are converted into an input format that can be recognized by the large language model, so as to ensure that the data is effectively parsed and processed by the large language model.
[0069] In some embodiments, the format-adapted input data is sent to the GPU server via a communication link, triggering the large language model to start running. Based on the instructions in the evaluation prompts, the large language model performs deep semantic understanding of the target text information, compares the preset label evaluation criteria with the case features in the text, identifies the core information related to the evaluation dimensions in the target text information, and generates subjective evaluation labels matching the target assessment object and corresponding identification reasons. The identification reasons are explanatory content explaining the semantic correlation between the subjective evaluation labels and the target text information, used to support the rationality of the labels.
[0070] The evaluation method for assessment subjects provided in this application combines multi-source data with target subjective evaluation tags to extract target text information reflecting the work performance of the assessment subjects. It generates evaluation prompt words using an evaluation knowledge base that stores evaluation rules and cases, and then calls a large language model to output the evaluation results. This not only improves the comprehensiveness and accuracy of the evaluation basis by combining multi-dimensional data, but also ensures the consistency and rationality of the evaluation logic through the standardized guidance of the evaluation knowledge base. At the same time, it effectively improves the overall efficiency of the evaluation work by leveraging the automated generation capability of the large language model.
[0071] Since the acquired multi-source data is raw data, it has problems such as redundant data, inconsistent formats, and hidden effective information, and cannot be directly used for subsequent extraction of target text information. Therefore, it is necessary to process the multi-source data to prepare for accurate extraction of target text information.
[0072] like Figure 4 As shown, S102 includes the following steps S201-S203: S201. Remove redundant data from the multi-source data to obtain the initial data.
[0073] It should be understood that redundant data refers to data from multiple sources that lacks substantial information value and does not support the subsequent extraction of target text information. Specifically, it includes duplicate data resulting from repeated uploading of materials and invalid data without substantial content.
[0074] In some embodiments, for duplicate data, a data fingerprint comparison algorithm can be used to calculate a unique data fingerprint for each piece of material data. Materials with the same data fingerprint are identified as duplicate uploads, and only one valid copy is retained, while the remaining duplicate copies are discarded. For invalid data without substantial content, a substantial content determination threshold is set. By detecting the amount of valid character information in the material (the number of characters after removing spaces, line breaks, and meaningless placeholders), materials that do not meet the determination threshold, as well as materials that only contain titles or have no body content and lack analytical value, are identified as invalid data and discarded. After the above elimination operation, the remaining data with substantial information value is the initial data.
[0075] S202. Transform the initial data into unified text data.
[0076] It should be understood that converting materials of different formats in the initial data into standard format text data can eliminate the impact of format differences on subsequent semantic analysis.
[0077] For example, the original format of the initial data may include various non-uniform text formats such as Word format and PDF format.
[0078] In some embodiments, the initial data can be converted to TXT format with UTF-8 encoding. UTF-8 encoding offers good compatibility and versatility, ensuring the complete preservation of different types of text information and smooth subsequent processing. The specific conversion process is as follows: First, for initial data in different original formats, the corresponding format parsing tool (such as a DOCX parsing component for Word format and a PDF parsing component for PDF format) is called to extract the text content from each format and remove format control characters (such as font style tags, page layout tags, etc.). Then, the extracted plain text content is encoded according to the UTF-8 encoding standard to generate a TXT format text file. Finally, the converted text data is encoded and verified to ensure there are no garbled characters, missing characters, or other issues. The text data that passes the verification is considered to be in a unified format.
[0079] S203. Extract target text information from text data through semantic analysis.
[0080] It should be understood that target text information specifically refers to key text paragraphs that can reflect the subjective performance of the assessment subject, and can provide key information support for the subsequent tag matching process.
[0081] In some embodiments, by invoking a built-in deep learning-based semantic analysis model (such as the BERT semantic understanding model), semantic modeling is performed on the text data after pre-format standardization, deep semantic features of the text content are mined and extracted, and the extracted semantic features of the text data are matched with a key topic thesaurus for similarity, and text fragments containing expressions related to key topics are selected.
[0082] In some embodiments, the matching process uses a semantic similarity algorithm to calculate the feature matching degree. When the matching degree reaches a preset threshold, the text fragment is determined to be a candidate text fragment.
[0083] Figure 5 A flowchart illustrating another evaluation method for an assessment object provided in this application embodiment, combined with... Figure 3 ,like Figure 5 As shown, generating evaluation prompts for the target assessment object includes the following steps S301-S302: S301. Retrieve the evaluation rules and evaluation cases corresponding to the target subjective evaluation tags from the evaluation knowledge base.
[0084] In some embodiments, the intelligent agent performs targeted retrieval of the evaluation knowledge base based on the unique identification information of the target subjective evaluation label. During the retrieval process, the retrieval module extracts all evaluation rules and evaluation cases corresponding to the target subjective evaluation label by matching the label identifier with the data in the knowledge base, and temporarily stores the retrieved results in the cache module of the data server for subsequent steps.
[0085] S302. Integrate the target text information, the evaluation rules corresponding to the target subjective evaluation tags, and the evaluation cases into the preset prompt word template to generate evaluation prompt words for the target assessment object.
[0086] In some embodiments, the preset prompt word template is an instructional text template with a fixed logical framework and populated fields, including target text information field, evaluation rule field and evaluation case field.
[0087] In some embodiments, the agent reads the retrieved evaluation rules, evaluation cases, and pre-stored target text information of the target assessment object from the cache; secondly, the agent calls the preset prompt word template inside the data server through the built-in prompt word generation module; finally, the prompt word generation module fills the target text information, evaluation rules, and evaluation cases into the corresponding fields of the preset prompt word template to complete information integration and text generation, forming the evaluation prompt word of the target assessment object, and pushes the evaluation prompt word to the agent module.
[0088] To ensure the accuracy and reasonableness of the evaluation results, after obtaining the evaluation results of the target assessment object, the evaluation results can be pushed to the review terminal for review by the reviewers.
[0089] like Figure 6 As shown, after obtaining the evaluation results of the target assessment subjects, the method also includes the following steps: S401. Push the evaluation results of the target assessment object to the audit terminal so that the auditors can audit the evaluation results of the target assessment object.
[0090] In some embodiments, the evaluation results of the target assessment object are generated based on the target text information, which includes the subjective evaluation tags of the corresponding target assessment object, the reasons for tag identification, and the matching target text information. The relevant data of the evaluation results can be stored on a data server.
[0091] In some embodiments, the audit terminal is a dedicated terminal device for auditors (such as administrators) to carry out audit work. Auditors can access relevant data of evaluation results stored in the data server through this terminal.
[0092] In some embodiments, the push operation is completed through the communication link between the data server and the review terminal. The specific process is as follows: First, the relevant data of the evaluation results of the target assessment object are standardized and encapsulated in the data server. The encapsulated evaluation results are pushed to the review terminal through a preset communication protocol (such as HTTP / HTTPS protocol). The complete content of the evaluation results is displayed to the reviewers through a visual interface. At the same time, the review operation entry is set up to provide support for the reviewers to carry out the review work.
[0093] S402. Receive the audit results from the audit terminal.
[0094] The audit result is used to indicate whether the evaluation results of the target assessment object have passed the audit.
[0095] In some embodiments, the audit result is in the form of a structured audit instruction generated by the audit terminal, which includes the audit conclusion (pass / fail), the auditor's identifier and the audit timestamp. If the audit fails, corresponding modification opinions (such as insufficient label matching basis, insufficient identification reasons, etc.) must also be attached.
[0096] In some embodiments, the receiving process is as follows: After the auditor completes the audit through the visual operation interface of the audit terminal, he submits the audit result. The audit terminal encodes the audit result according to the preset data format and sends it back to the data server through the communication link. The data server starts the receiving and listening program to receive and decode the data packets sent back by the audit terminal, parse out the specific content of the audit result, associate the audit result with the evaluation result of the corresponding target assessment object, and update the audit status field of the corresponding evaluation result in the "cadre tag table" in the data server to realize the synchronous recording of the audit result.
[0097] In some embodiments, the modification suggestions refer to the specific optimization directions and problem descriptions proposed by the reviewers for the evaluation results that failed the review, which are used for subsequent iterations and corrections.
[0098] S4031. If the review result indicates that the evaluation result of the target assessment object has passed the review, the evaluation result of the target assessment object shall be determined to be effective and the evaluation result of the target assessment object shall be pushed to the user terminal.
[0099] In some embodiments, when the audit result indicates that the evaluation result of the target assessment object has passed, the data server updates the evaluation result of the corresponding subjective evaluation tag in the internally stored tag table of the target assessment object to "effective" and persists it; then the effective evaluation result is optimized and encapsulated, and pushed to the user terminal via a secure communication link. The user terminal receives and displays it for subsequent operations.
[0100] S4032. If the audit results indicate that the evaluation results of the target assessment object have not passed the audit, the evaluation results of the target assessment object shall be determined to be invalid, and the evaluation results of the target assessment object shall be iteratively corrected until the evaluation results of the target assessment object pass the audit.
[0101] In some embodiments, if the review result indicates that the evaluation result of the target assessment object has failed, the data server determines that the evaluation result is invalid and initiates an iterative correction process until the evaluation result passes the review. Specifically: after the data server determines that the review has failed, the evaluation result of the corresponding subjective evaluation tag in the internally stored tag table of the target assessment object is updated to "invalid" and modification suggestions are added; the invalid result and modification suggestions are pushed to the intelligent agent, which optimizes the prompt words, calls the target text information of the target assessment object to regenerate a new evaluation result, and the data server pushes the new evaluation result to the review terminal to repeat the review process. After the review is passed, the activation and push operations are performed.
[0102] In some embodiments, the data server records correction information to form a traceability record.
[0103] Figure 7A flowchart of another evaluation method for an assessment object is provided for embodiments of this application, such as... Figure 7 As shown, iteratively correcting the evaluation results of the target assessment object can be specifically implemented through the following steps S501-S502: S501. In each iteration, the evaluation prompts are revised based on the feedback from the evaluators to obtain the revised evaluation prompts.
[0104] In some embodiments, during each iteration of the correction process, modification comments transmitted from the review terminal and entered by the assessors are received through a preset interface. These modification comments are the assessors' requests for adjustments and optimization suggestions regarding invalid evaluation results, based on the review outcomes.
[0105] In some embodiments, the data server pushes modification suggestions to the intelligent agent. Based on the core logic of the modification suggestions, the intelligent agent makes targeted adjustments to the original evaluation prompts. The adjustments include, but are not limited to, strengthening the retrieval logic of specific evaluation dimensions, supplementing the judgment criteria for tag matching, and optimizing the instructions for semantic analysis. The revised evaluation prompts are obtained and cached for subsequent use.
[0106] S502. Based on the revised evaluation prompts, obtain the revised evaluation results for the target assessment object.
[0107] In some embodiments, according to the description in S104, the corrected evaluation prompts are re-input into the large language model to obtain the corrected evaluation results of the target assessment object.
[0108] Based on the above scheme, in order to further improve the efficiency and pertinence of the evaluation result review work, avoid the waste of resources caused by indiscriminate review, and ensure that low confidence evaluation results are given priority verification, the method can also assign priorities to different evaluation results and schedule the review process by adjusting the priorities.
[0109] In practice, the review priority of the evaluation results of the target assessment subjects can be adjusted based on the confidence level of their evaluation results.
[0110] The confidence level is determined based on the semantic matching degree between the target text information corresponding to the evaluation prompts of the target assessment object and the target subjective evaluation labels.
[0111] The semantic matching degree is determined based on a large language model.
[0112] It should be understood that, in order to ensure the accuracy of the evaluation results for the target assessment subjects, a confidence threshold can be set for the confidence level of the evaluation results to filter the evaluation results, distinguish the results that need to be reviewed in detail, and transmit them to the data server.
[0113] In some embodiments, after the large language model completes the semantic analysis between the target text information and the target subjective evaluation label, the output evaluation result of the target assessment object also includes the target subjective evaluation label and the corresponding confidence level result data. Then, the review priority of the evaluation result of the target assessment object is adjusted according to the relationship between the confidence level and the preset threshold.
[0114] In some embodiments, the confidence level is related to the semantic matching degree between the target text information and the target subjective evaluation label. The higher the matching degree, the higher the confidence level, and vice versa.
[0115] In some embodiments, if the confidence level is less than a preset threshold, it indicates that the semantic matching degree between the target text information and the target subjective evaluation label is low. This may be caused by a program malfunction. In this case, it is necessary to adjust the review priority of the evaluation results of the target assessment object to priority review, determine whether the evaluation results are normal as soon as possible, and improve efficiency.
[0116] In some embodiments, if the confidence level is greater than or equal to a preset threshold, it indicates that the semantic matching degree between the target text information and the target subjective evaluation label is low, and the review priority of the evaluation results of the target assessment object can be reduced.
[0117] For example, the preset threshold can be set to 0.7. This application does not limit the specific value of the preset threshold, which can be determined according to the evaluation requirements in actual applications.
[0118] Figure 8 A flowchart of another evaluation method for an assessment object provided in the embodiments of this application is shown below. Figure 8 As shown, the method further includes the following step S601: S601. Upon detecting an update operation of the evaluation knowledge base, generate a new prompt word template based on the updated content of the evaluation knowledge base.
[0119] The updated content includes newly added subjective evaluation tags, as well as the evaluation rules and evaluation examples corresponding to the newly added subjective evaluation tags.
[0120] In some embodiments, the update operation of the evaluation knowledge base originates from the addition and modification operations of the evaluation knowledge base administrator on the terminal to the data in the evaluation knowledge base, and the update operation of the knowledge base is detected in real time through a preset detection program.
[0121] In some embodiments, the update operation specifically refers to the structured data supplementation and modification operation initiated by the administrator through the terminal to the evaluation knowledge base when management business needs change.
[0122] In some embodiments, the detection of update operations is achieved through a preset lightweight listening subroutine. When an update operation is detected, a start signal is sent to the agent, and the relevant information of the update operation is stored and recorded for subsequent traceability and review.
[0123] For example, a lightweight listening subroutine can use a polling mechanism with a polling interval of 10 seconds.
[0124] In some embodiments, new prompt word templates are generated based on the updated content. Prompt word templates matching the type of the newly added subjective evaluation label can be added. Alternatively, the prompt word templates can be modified according to the evaluation rules and evaluation cases corresponding to the newly added subjective evaluation labels, so that the large language model can more accurately understand the intent expressed by the prompt words.
[0125] As can be seen, the above mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the embodiments of this application provide corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the modules and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.
[0126] This application embodiment can divide the evaluation device into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. Optionally, the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0127] Figure 9 A schematic diagram illustrating the composition of an evaluation device provided in this application. Figure 9As shown, the evaluation device 700 includes: an acquisition module 701 and a processing module 702; the acquisition module 701 is used to acquire multi-source data of the target assessment object and target subjective evaluation tags; the processing module 702 is used to extract target text information from the multi-source data; the target text information is used to reflect the work performance of the target assessment object; based on the target text information, target subjective evaluation tags, and evaluation knowledge base, evaluation prompt words for the target assessment object are generated; wherein, the evaluation knowledge base is used to store evaluation rules and evaluation cases corresponding to different subjective evaluation tags; based on the evaluation prompt words, the evaluation result of the target assessment object is obtained by calling a large language model.
[0128] In some embodiments, the processing module 702 is specifically used to retrieve the evaluation rules and evaluation cases corresponding to the target subjective evaluation tags from the evaluation knowledge base; integrate the target text information, the evaluation rules and evaluation cases corresponding to the target subjective evaluation tags into a preset prompt word template, and generate evaluation prompt words for the target assessment object.
[0129] In some embodiments, the processing module 702 is specifically used to remove redundant data from multi-source data to obtain initial data; convert the initial data into unified text data; and extract target text information from the text data through semantic analysis.
[0130] In some embodiments, after obtaining the evaluation results of the target assessment object, the processing module 702 is further configured to push the evaluation results of the target assessment object to the review terminal so that the reviewer can review the evaluation results of the target assessment object; receive the review results fed back by the review terminal; the review results are used to indicate whether the evaluation results of the target assessment object have passed the review; if the review results indicate that the evaluation results of the target assessment object have passed the review, determine that the evaluation results of the target assessment object are effective and push the evaluation results of the target assessment object to the user terminal; if the review results indicate that the evaluation results of the target assessment object have not passed the review, determine that the evaluation results of the target assessment object are invalid and iteratively correct the evaluation results of the target assessment object until the evaluation results of the target assessment object have passed the review.
[0131] In some embodiments, the processing module 702 is specifically used to revise the evaluation prompts based on the modification opinions of the evaluators during each iteration, and obtain the revised evaluation prompts; based on the revised evaluation prompts, obtain the revised evaluation result of the target evaluation object.
[0132] In some embodiments, the processing module 702 is further configured to adjust the review priority of the evaluation results of the target assessment object based on the confidence level of the evaluation results of the target assessment object; wherein, the confidence level is determined based on the semantic matching degree between the target text information corresponding to the evaluation prompt words of the target assessment object and the target subjective evaluation label; wherein, the semantic matching degree is determined based on a large language model.
[0133] In some embodiments, the processing module 702 is further configured to generate a new prompt word template based on the updated content of the evaluation knowledge base when an update operation of the evaluation knowledge base is detected; wherein the updated content includes newly added subjective evaluation tags, as well as the evaluation rules and evaluation cases corresponding to the newly added subjective evaluation tags.
[0134] In the case of implementing the functions of the integrated modules described above in hardware, this embodiment of the invention provides a possible structural schematic diagram of the electronic device involved in the above embodiments. For example... Figure 10 As shown, the electronic device 800 includes: a processor 802, a communication interface 803, and a bus 804. Optionally, the electronic device 800 may also include a memory 801.
[0135] Processor 802 may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 802 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 802 may also be a combination that implements computing functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0136] The communication interface 803 is used to connect to other devices via a communication network. This communication network can be Ethernet, wireless access network, wireless local area network (WLAN), etc.
[0137] The memory 801 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.
[0138] In one possible implementation, the memory 801 can exist independently of the processor 802. The memory 801 can be connected to the processor 802 via a bus 804 and is used to store instructions or program code. When the processor 802 calls and executes the instructions or program code stored in the memory 801, it can implement the evaluation method for the assessment object provided in this embodiment of the invention.
[0139] In another possible implementation, the memory 801 can also be integrated with the processor 802.
[0140] The 804 bus can be an extended industry standard architecture (EISA) bus, etc. The 804 bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0141] Through the above description of the implementation methods, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the service calling device can be divided into different functional modules to complete all or part of the functions described above.
[0142] This application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be instructed by computer program instructions to be implemented by related hardware. This program can be stored in the aforementioned computer-readable storage medium. When executed on a computer, the computer program instructions cause the computer to perform the evaluation method for the assessment object as described in any of the above embodiments.
[0143] Exemplary examples of computer-readable storage media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes), optical discs (e.g., compact disks (CDs), digital versatile disks (DVDs), etc.), smart cards, and flash memory devices (e.g., erasable programmable read-only memory (EPROMs), cards, sticks, or key drives, etc.). The various computer-readable storage media described in this disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.
[0144] This application also provides a computer program product, which includes a computer program that, when run on a computer, causes the computer to execute any of the evaluation methods for assessment objects provided in the above embodiments.
[0145] In the description of the embodiments of this application, specific features, structures, materials or characteristics may be combined in any suitable manner in one or more embodiments or examples.
[0146] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An evaluation method for assessment subjects, characterized in that, The method includes: Obtain multi-source data and subjective evaluation tags of the target assessment subjects; Target text information is extracted from the multi-source data; the target text information is used to reflect the work performance of the target assessment object. Based on the target text information, the target subjective evaluation tags, and the evaluation knowledge base, evaluation prompt words for the target assessment object are generated; wherein, the evaluation knowledge base is used to store evaluation rules and evaluation cases corresponding to different subjective evaluation tags; Based on the evaluation prompts, the evaluation results of the target assessment object are obtained by calling the large language model.
2. The method according to claim 1, characterized in that, The step of generating evaluation prompts for the target assessment subject based on the target text information, the target subjective evaluation tags, and the evaluation knowledge base includes: Retrieve the evaluation rules and evaluation cases corresponding to the target subjective evaluation tags from the evaluation knowledge base; The target text information, the evaluation rules and evaluation cases corresponding to the target subjective evaluation tags are integrated into a preset prompt word template to generate evaluation prompt words for the target assessment object.
3. The method according to claim 1, characterized in that, The extraction of target text information from the multi-source data includes: Redundant data is removed from the multi-source data to obtain the initial data; The initial data is then converted into uniform text data. The target text information is extracted from the text data through semantic analysis.
4. The method according to claim 1, characterized in that, After obtaining the evaluation results of the target assessment object, the method further includes: The evaluation results of the target assessment object are pushed to the review terminal so that the reviewers can review the evaluation results of the target assessment object; The system receives the audit results from the audit terminal; the audit results are used to indicate whether the evaluation results of the target assessment object have passed the audit. If the review result indicates that the evaluation result of the target assessment object has passed the review, the evaluation result of the target assessment object is determined to be effective, and the evaluation result of the target assessment object is pushed to the user terminal; If the audit result indicates that the evaluation result of the target assessment object has not passed the audit, the evaluation result of the target assessment object is determined to be invalid, and the evaluation result of the target assessment object is iteratively corrected until the evaluation result of the target assessment object passes the audit.
5. The method according to claim 4, characterized in that, The iterative correction of the evaluation results of the target assessment object includes: In each iteration, the evaluation prompts are revised based on the feedback from the evaluators to obtain the revised evaluation prompts. Based on the revised evaluation prompts, the revised evaluation results for the target assessment object are obtained.
6. The method according to claim 4, characterized in that, The method further includes: Based on the confidence level of the evaluation results of the target assessment object, the review priority of the evaluation results of the target assessment object is adjusted; wherein, the confidence level is determined based on the semantic matching degree between the target text information corresponding to the evaluation prompt words of the target assessment object and the target subjective evaluation tag; wherein, the semantic matching degree is determined based on the large language model.
7. The method according to claim 1, characterized in that, The method further includes: Upon detecting an update operation to the evaluation knowledge base, a new prompt word template is generated based on the updated content of the evaluation knowledge base; wherein, the updated content includes newly added subjective evaluation tags, as well as the evaluation rules and evaluation examples corresponding to the newly added subjective evaluation tags.
8. An evaluation system for assessment subjects, characterized in that, The system includes: a multi-source data processing module, an intelligent agent, and a large language model invocation module; The multi-source data processing module is used to acquire multi-source data of the target assessment object, extract target text information from the multi-source data and send it to the intelligent agent; the target text information is used to reflect the work performance of the target assessment object. The intelligent agent is used to acquire the target subjective evaluation tags and the target text information of the target assessment object, and generate evaluation prompt words for the target assessment object based on the target text information, the target subjective evaluation tags and the evaluation knowledge base; The large language model invocation module is used to invoke the large language model, input the evaluation prompt words of the target assessment object into the large language model, obtain the evaluation result of the target assessment object output by the large language model, and return the evaluation result of the target assessment object to the intelligent agent.
9. The evaluation system for assessment subjects according to claim 8, characterized in that, The system also includes: an evaluation knowledge base module, a result review module, and a result output module; The evaluation knowledge base module is used to store evaluation rules and evaluation cases corresponding to different subjective evaluation tags; The result review module is used to review the evaluation results of the target assessment object; The result output module is used to push the evaluation results of the assessment subjects to the user terminal.
10. An electronic device, characterized in that, The electronic device includes: a processor and a memory; The memory stores instructions that the processor can execute; When the processor is configured to execute the instructions, the electronic device implements the method as described in any one of claims 1-7.