Text labeling method and device, equipment, storage medium and computer program product

By dividing the text set to be marked into multiple subtext sets and annotating based on the matching degree and occurrence of text, the problem of insufficient reliability and consistency of text annotation results in the prior art is solved, and automated labeling and refined labeling of text continuous value attributes are realized.

CN120067788APending Publication Date: 2025-05-30MICRO DREAM TECHTRONIC NETWORK TECH CHINACO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510064989.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When facing high-difficulty text labeling tasks, the reliability and consistency of the labeling results are insufficient, and traditional manual labeling is inefficient and costly, which limits the breadth and depth of data labeling.

Method used

By obtaining the text set to be marked, dividing it into multiple subtext sets, and identifying text attributes for each subtext set, marking the text with the highest and lowest matching degree with the preset target attributes. Based on the number of times the text appears in multiple subtext sets and the frequency of being marked, the matching degree of the text is determined and marked.

Benefits of technology

It realizes automatic annotation of text continuous value attributes, provides more refined attribute annotation, improves the fineness and efficiency of the annotation, and enhances the reliability and consistency of the annotation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067788A_ABST
    Figure CN120067788A_ABST
Patent Text Reader

Abstract

The invention provides a text labeling method and device, equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a to-be-labeled text set, and dividing the text set into a plurality of groups of sub-text sets; for each group of sub-text sets, identifying the attribute of each text in the sub-text sets, marking the text with the maximum matching degree with the preset target attribute in the sub-text sets as a first text, and marking the text with the minimum matching degree with the preset target attribute as a second text; aiming at each to-be-labeled text in the to-be-labeled text set, based on a first total number of times that the text appears in the plurality of groups of sub-text sets, a second total number of times that the text is marked as a first text in different sub-text sets, and a third total number of times that the text is marked as a second text in different sub-text sets, marking the text as the first text in the plurality of groups of sub-text sets; determining a first matching degree between the text and a preset target attribute, and marking the matching degree of the text. By means of the method and device, automatic annotation of the text continuous value attributes can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to text processing technology, and in particular to a text annotation method, device, equipment, storage medium and computer program product. Background Art

[0002] With the rapid enhancement of computer computing power, deep learning methods have firmly occupied the dominant position in the field of natural language processing. The efficient operation of deep learning models relies on a large-scale labeled dataset as support, and crowdsourcing annotation has become the preferred strategy for data annotation due to its high efficiency and economy. This method widely recruits online users to participate in annotation tasks through online platforms, making full use of the wisdom and strength of the public. However, although crowdsourcing annotation has significant advantages in terms of efficiency and cost, when facing some high-difficulty tasks, the reliability and consistency of its annotation results face severe challenges. In addition, the traditional manual annotation method is not only inefficient but also further increases the annotation cost, which to a certain extent limits the breadth and depth of data annotation. Summary of the Invention

[0003] Embodiments of this application provide a text annotation method, device, storage medium and computer program product, which can realize the automatic annotation of continuous value attributes of text.

[0004] The technical solution of the embodiments of this application is implemented as follows:

[0005] Embodiments of this application provide a text annotation method, including: obtaining a text set to be annotated, and dividing the text set into multiple groups of sub-text sets; wherein, the sub-text sets of different groups are different, and the texts to be annotated in each group of sub-text sets are different; for each group of sub-text sets, identifying the attributes of each text in the sub-text set, and marking the text with the maximum matching degree with the preset target attribute in the sub-text set as the first text, and marking the text with the minimum matching degree with the preset target attribute as the second text; for each text to be annotated in the text set to be annotated, based on the first total number of times the text appears in the multiple groups of sub-text sets, the second total number of times the text is marked as the first text in different sub-text sets, and the third total number of times the text is marked as the second text in different sub-text sets, determining the matching degree of the text with the preset target attribute, and performing matching degree annotation on the text.

[0006] An embodiment of the present application provides a text annotation device, including: a first acquisition module, configured to acquire a text set to be annotated and divide the text set into multiple groups of sub-text sets; wherein, the sub-text sets of different groups are different, and the texts to be annotated in each group of sub-text sets are different; a marking module, configured to, for each group of sub-text sets, identify the attributes of each text in the sub-text set, and mark the text with the highest matching degree with a preset target attribute in the sub-text set as the first text, and mark the text with the lowest matching degree with the preset target attribute as the second text; a first determination module, configured to, for each text to be annotated in the text set to be annotated, determine a first matching degree of the text with the preset target attribute based on a first total number of times the text appears in the multiple groups of sub-text sets, a second total number of times the text is marked as the first text in different sub-text sets, and a third total number of times the text is marked as the second text in different sub-text sets, and perform matching degree annotation on the text.

[0007] An embodiment of the present application provides an electronic device, including:

[0008] A memory, configured to store executable instructions;

[0009] A processor, configured to, when executing the executable instructions stored in the memory, implement the text annotation method provided by the embodiment of the present application.

[0010] An embodiment of the present application provides a computer-readable storage medium, on which a computer program or executable instructions are stored, and when the computer program or executable instructions are executed by a processor, the text annotation method provided by the embodiment of the present application is implemented.

[0011] An embodiment of the present application provides a computer program product, including a computer program or instructions, and when the computer program or instructions are executed by a processor, the text annotation method provided by the embodiment of the present application is implemented.

[0012] The embodiment of the present application has the following beneficial effects:

[0013] In the embodiments of the present disclosure, multiple groups of sub-text sets are generated from the text set to be annotated, and for each group of sub-text sets, the first text closest to the preset target attribute and the second text most deviated from the preset target attribute within the sub-text set are determined. By counting the second total number of times each text to be annotated is marked as the first text and the third total number of times it is marked as the second text in different sub-text sets, the overall closeness of the text to the target attribute is reflected by the second total number, and the overall remoteness of the text from the target attribute is reflected by the third total number; and the text is marked with a matching degree based on the first matching degree determined by the first total number, the second total number, and the third total number of times the text is evaluated in multiple groups of sub-text sets. Since the first matching degree of the text is a quantitative representation of the closeness of the text to the target attribute, and the first matching degrees of different texts with the preset target attribute may vary, then marking the text based on the first matching degree with the preset target attribute provides a more refined evaluation index for the annotation of the text. Therefore, a more refined attribute annotation can be obtained for the text, thereby realizing the automatic annotation of the continuous value attribute of the text. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a schematic structural diagram of a text marking system provided by an embodiment of the present application;

[0015] Figure 2 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0016] Figure 3 is an optional flowchart of a text annotation method provided by an embodiment of the present application;

[0017] Figure 4 is a flowchart of a method for generating sub-text sets provided by an embodiment of the present application;

[0018] Figure 5 is a flowchart of a method for determining a target large language model provided by an embodiment of the present application;

[0019] Figure 6 is a flow example diagram of a method for generating sub-text sets provided by an embodiment of the present application;

[0020] Figure 7 is a flow example diagram of a method for determining a target large language model provided by an embodiment of the present application;

[0021] Figure 8 is a schematic diagram of the principle of a method for batch annotating texts provided by an embodiment of the present application;

[0022] Figure 9 is a flowchart of a method for evaluating the consistency of text annotation results provided by an embodiment of the present application;

[0023] Figure 10 It is a schematic diagram of a text annotation method provided by an embodiment of the present application. Detailed implementation manners

[0024] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0025] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0026] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0028] In the related art, the automated annotation method for text has undergone significant evolution. Initially, it mainly relied on heuristic annotation methods, which based on empirically set rules or intuitive judgments by humans to identify data. These rules often originated from the knowledge and experience of experts, but due to the subjectivity of humans and the limitations of the rules, there were bottlenecks in the annotation efficiency and accuracy. With the rise of pre-trained language models, a completely new annotation method emerged. This method utilizes the massive knowledge embedded in the model and realizes the automated annotation of text through the method of zero-shot or few-shot prompt completion. This method has greatly improved the annotation efficiency and flexibility and reduced the dependence on artificial rules. And in recent years, large language models represented by ChatGPT have pushed the automated annotation of text to a new height. These models can accurately understand and follow detailed human instructions, thus realizing the precise annotation of text data. However, most of the above automated annotation methods are more suitable for data annotation of discrete value attributes of text (such as the ternary classification of "positive / negative / neutral" for sentiment tendency).

[0029] For the data annotation of text continuous value attributes, such as quantifying the positive sentiment intensity expressed in text, a common method is to first discretize the continuous values, that is, construct a rating scale (such as a Likert scale or an additive scale), and divide the sentiment intensity into several discrete levels or categories. However, simply through this discretization transformation and using a large language model for automated annotation similar to discrete value attributes is actually an over-simplification and approximation of the continuous value annotation task. Specifically, attributes such as "the positive sentiment intensity expressed in text" are essentially a continuously varying spectrum. If it is forced to be discretized into forms such as "extremely strong, strong, medium, weak, extremely weak, none" or numerical levels "5, 4, 3, 2, 1, 0", it will undoubtedly lose the original richness and precision of the continuous values and cannot comprehensively and accurately capture the subtle differences in sentiment intensity. In addition, during the process of discretizing the attribute values, the selection of the granularity of the discrete values is crucial. If the granularity is set too fine, although theoretically it can more accurately reflect the differences in sentiment intensity, correspondingly, it poses higher requirements for the design of data annotation instructions, and instructions with very high precision and elaboration must be used. This not only greatly increases the burden of instruction design but also requires the large language model to have excellent long text modeling capabilities to understand and execute these complex instructions. On the contrary, if the granularity is set too coarse, although it simplifies the instruction design, it limits the model's ability to finely evaluate the text sentiment, resulting in the model being able to only make a rough evaluation according to the established granularity even if it has higher evaluation accuracy. In addition, the preference for specific discrete scales that may form in the large language model during the training process is also an issue that cannot be ignored. This preference may lead to systematic biases in the annotation results and affect the accuracy and reliability of the annotation.

[0030] In view of the above technical problems, the embodiments of the present application provide a text annotation method, device, equipment, storage medium and computer program product, which can realize the automated annotation of text continuous value attributes.

[0031] The following describes the electronic device provided by the embodiments of the present application for implementing the above text annotation method. The device provided by the embodiments of the present application can be implemented as various types of user terminals such as laptops, tablets, desktop computers, set-top boxes, mobile devices (such as mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), or can also be implemented as the background server of a network platform (hereinafter referred to as the server). Next, an exemplary application will be described when the electronic device is implemented as a server.

[0032] See Figure 1 , Figure 1It is a schematic structural diagram of the text annotation system 100 provided by an embodiment of the present application. To implement the text annotation method of the embodiment of the present application, terminal devices (exemplarily showing terminal 400-1 and terminal 400-2, Figure 1 only two terminal devices are shown, but it does not limit the number of terminal devices included in the text annotation system of the embodiment of the present application) are connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.

[0033] In some possible implementation manners, user A can log in to the application corresponding to the network platform through the terminal device 400-1, and initiate a request A for annotating text data on the graphical interface 410-1 of the terminal device 400-1. User B can also log in to the application corresponding to the network platform through the terminal device 400-2, and initiate a request B for annotating text data on the graphical interface 410-2 of the terminal device 400-2. The request A and the request B perform information interaction with the server 200 through the network 300; the server 200 obtains the text set to be annotated through the network 300; after the server 200 annotates the text data based on the text annotation method in the embodiment of the present disclosure, it can send the annotation result to the terminal device 400-1 and the terminal device 400-2 through the network 300, and display the annotation result on the application program interfaces of the terminal device 400-1 and the terminal device 400-2.

[0034] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms. The terminal device 400 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and are not limited in the embodiment of the present invention.

[0035] See Figure 2 , Figure 2 It is a schematic structural diagram of the electronic device 200 provided by an embodiment of the present application. It should be noted that, Figure 2 the shown electronic device 200 can be Figure 1 any one of the two electronic devices shown in Figure 2The electronic device 200 shown includes: at least one processor 210, a memory 250, at least one network interface 220 and a user interface 230. The various components in the electronic device 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 240 is not described in detail. Figure 2 Various buses are labeled as bus system 240 .

[0036] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0037] The user interface 230 includes one or more output devices 231 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 230 also includes one or more input devices 232, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0038] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 250 may optionally include one or more storage devices that are physically remote from the processor 210.

[0039] The memory 250 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0040] In some embodiments, memory 250 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.

[0041] Operating system 251, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0042] A network communication module 252 for reaching other computing devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.

[0043] A presentation module 253 for enabling the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more output devices 231 associated with the user interface 230 (such as a display screen, a speaker, etc.).

[0044] An input processing module 254 for detecting and translating one or more user inputs or interactions from one of one or more input devices 232.

[0045] In some embodiments, the device provided by the embodiments of the present application may be implemented in software. Figure 2 A device 255 stored in the memory 250 is shown, which may be software in the form of a program and a plug-in, etc., including the following software modules: a first acquisition module 2551, a marking module 2552, and a first determination module 2553. These modules are logical, so they can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.

[0046] In other embodiments, the device provided by the embodiments of the present application may be implemented in hardware. As an example, the device provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the text annotation method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor may employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0047] In some embodiments, an electronic device or a server may implement the text annotation method provided in the embodiments of the present application by running a computer program. For example, the computer program may be a native program or a software module in an operating system; it may be a local application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a text annotation APP; it may also be a small program, that is, a program that only needs to be downloaded to the browser environment to run; it may also be a small program that can be embedded into any APP. In short, the above computer program may be any form of application program, module or plug-in.

[0048] Next, the text annotation method provided in the embodiments of the present application will be described in conjunction with the exemplary applications and implementations of the terminal provided in the embodiments of the present application.

[0049] See Figure 3 , Figure 3 is a flowchart of the text annotation method provided in the embodiments of the present application, and will be described in conjunction with Figure 3 the steps S301 - S303 shown.

[0050] In step S301, a text set to be annotated is obtained, and the text set is divided into multiple groups of sub - text sets; wherein, the sub - text sets of different groups are different, and the texts to be annotated in each group of sub - text sets are different.

[0051] The text annotation method in the embodiments of the present disclosure can be applied to a wide range of text annotation scenarios such as text classification, sentiment analysis, medical record information processing, comment analysis, learning resource annotation, and product recommendation tagging.

[0052] It should be noted that the text annotation method proposed in the present disclosure can be applied to an electronic device or a server. Here, the electronic device may include: a terminal device, for example, a mobile terminal or a fixed terminal. Among them, the mobile terminal may include: devices such as mobile phones, tablet computers, and laptop computers. The fixed terminal may include: desktop computers, smart TVs, etc. As a type of computer, a server can provide computing or application services for other client machines (such as terminal devices such as computers, smartphones, ATMs, and even large devices such as train systems) in the network.

[0053] The execution subject of the embodiments of the present disclosure may, in terms of hardware, be, for example, a central processing unit (CPU) in a server or an electronic device, and in terms of software, may be, for example, the relevant background service in a server or an electronic device, and this is not limited.

[0054] The comment generation method in the embodiments of the present disclosure can also be configured in a text annotation device, which can be set in a server or in an electronic device. The embodiments of the present disclosure do not limit this, and the embodiments of the present disclosure will be described by taking the text annotation device as the execution subject as an example.

[0055] In the embodiments of the present disclosure, the text annotation device obtains a text set to be annotated. Among them, the text set to be annotated includes a plurality of texts to be annotated, and the text to be annotated is original text data that needs to be annotated in a certain form (such as part-of-speech annotation, semantic role annotation, named entity recognition, sentiment analysis, etc.). Here, the text to be annotated can be an original text, for example, an original blog post, etc.; the text to be annotated can also be a newly generated text based on the original, for example, a newly generated original blog post obtained from an application, real-time generated original information, etc. The content of the text to be annotated is usually one sentence or one paragraph of text, including characters, numbers, etc. It should be noted that the text set can carry the text to be annotated in various forms, including but not limited to database files, log files, spreadsheets, and web page files, etc.; at the same time, the text set can be stored on a remote server for multi-user collaborative processing, centralized management, and data sharing; it can also be stored locally to meet the usage requirements in cases such as single-machine processing, high requirements for data privacy, or limited network environment.

[0056] In some embodiments, the text set is stored in a database on a remote server, and the text annotation device has a data receiving function and can receive the text set to be annotated sent by the server; in other embodiments, the user specifies text data related to text annotation based on the input module on the text annotation device, and then the text annotation module can directly obtain the text data to be annotated specified by the user, so as to obtain the text set to be annotated.

[0057] In the embodiments of the present disclosure, the text annotation device can implement a random sampling strategy multiple times on the text set to be annotated, so as to obtain multiple groups of sub-text sets based on the text set. Here, the texts within each group of sub-text sets are highly representative and non-repetitive, and different groups of sub-text sets are also different to maintain significant differences between the sub-text sets. During the construction process of each group of sub-text sets, the text annotation device can adopt the method of random sampling without replacement, so that the texts within each group of sub-text sets will not be repeated. In addition, to reduce the situation where different groups of sub-text sets are the same, the text annotation device can set a minimum difference threshold between the sub-text sets. After randomly sampling to generate multiple sub-text sets, by calculating the similarity between each sub-text set, industry-recognized metrics such as cosine similarity and Jaccard similarity can be used to measure here. Then, compare the similarity between each sub-text set with the minimum difference threshold. Only when the similarity between all sub-text sets is lower than this minimum difference threshold, it is considered that the multiple groups of sub-text sets generated by this random sampling meet the requirements of diversification. Exemplarily, if the text set to be annotated is [data 1, data 2, data 3, data 4, data 5, data 6,...], then the text set can be sampled to obtain sub-text set 1, sub-text set 2, etc. Among them, sub-text set 1 is [data 1, data 2, data 3, data 4], and sub-text set 2 is [data 1, data 2, data 4, data 5]. Of course, the content of the sub-text set is not limited to this, as long as it meets the limited requirements in the embodiments of the present disclosure.

[0058] It should be noted that the number of texts included in different groups of sub-text sets can be the same or different, as long as there are enough texts in each group for subsequent analysis.

[0059] In step S302, for each group of sub-text sets, identify the attributes of each text in the sub-text set, and mark the text with the maximum matching degree with the preset target attribute in the sub-text set as the first text, and mark the text with the minimum matching degree with the preset target attribute as the second text.

[0060] In the embodiments of the present disclosure, the text annotation device processes each group of sub-text sets separately. For each group of sub-text sets, identify the attributes of each text in the group of sub-text sets. Among them, these attributes can be keywords referring to the theme, sentiment tendency, etc. of the text. For example, the attribute of the emotional expression in the text can be positive emotion, negative emotion, and neutral emotion, etc.

[0061] In the embodiments of the present disclosure, the text annotation device also marks the text with the highest matching degree with the preset target attribute in the subset of texts as the first text, which means that this first text is the closest to the target attribute in the subset of texts; at the same time, the text with the lowest matching degree with the preset target attribute is marked as the second text, which also means that this second text is the farthest from the target attribute in the subset of texts, that is, the difference degree between the text and the preset target attribute is represented according to the matching degree between the text and the preset target attribute. Here, the preset target attribute can be set according to the specific annotation scenario. For example, when marking the positive sentiment degree of the text or the text theme, the preset target attribute can be set as the positive situation or a specific theme. In some embodiments, the text annotation device can define a series of relevant keywords according to the preset target attribute, where these keywords should be able to accurately reflect the characteristics of the preset target attribute. Then, by traversing each text in the subset of texts, it is checked whether the text contains the predefined keywords, and the matching degree between the text and the keywords is determined by calculating the number or frequency of the keywords appearing in the text. The matching degree is positively correlated with the number or frequency of the keywords appearing in the text. According to the matching degree, the text annotation device can mark the text with the highest matching degree with the preset target attribute in the subset of texts as the first text. Similarly, the text with the lowest matching degree with the preset target attribute can be marked as the second text. Exemplarily, the subset of texts 1 is [data 1, data 2, data 3, data 4]. Assuming that the subset of texts 1 is text data related to emotional expression, then each text in the subset of texts 1 can be identified, and the data 1 closest to the positive emotion and the data 3 farthest from the positive emotion can be identified. Then, the data 1 can be determined as the first text in the subset of texts 1, and the data 3 can be determined as the second text in the subset of texts 1.

[0062] In some other embodiments, the above text annotation method further includes:

[0063] Determining a target large language model from multiple alternative large language models;

[0064] For each group of subsets of texts, identifying the attributes of each text in the subset of texts, and marking the text with the highest matching degree with the preset target attribute in the subset of texts as the first text, and marking the text with the lowest matching degree with the preset target attribute as the second text, includes:

[0065] For each group of subsets of texts, based on the target large language model, identifying the attributes of each text in the subset of texts, and marking the text with the highest matching degree with the preset target attribute in the subset of texts as the first text, and marking the text with the lowest matching degree with the preset target attribute as the second text.

[0066] In the disclosed embodiment, the text annotation device can use the large prediction model to implement annotation of each text in the sub-text set. The large language model refers to a deep neural network model that has been pre-trained with massive text data and has obtained the ability to follow human instructions and align human values ​​or preferences through instruction fine-tuning and human feedback reinforcement learning. Before performing text annotation, the text annotation device determines a target large language model from multiple candidate large language models. The candidate large language model refers to a preset large language model that can be used for text annotation. In some embodiments, factors such as specific application scenarios, performance requirements, and computing resources can be considered, and multiple candidate large language models can be compared for each factor to determine the most suitable target large language model. Exemplarily, due to the different degrees of dependence of different models on computing resources, some models require high-performance graphics processing units (GPUs) to support, while others can run on ordinary central processing units (CPUs). Then, by comparing the performance of the processing efficiency of each candidate large language model on the current hardware resources in the system, the target large language model that is most suitable for the current hardware resources can be selected according to the performance results of each candidate large language model.

[0067] In other embodiments, the text annotation device can also determine the most suitable target large language model based on the prediction capabilities of multiple candidate large prediction models for the attributes of the text, where the prediction capability can be characterized by the accuracy or stability of identifying the attributes of the text.

[0068] In the disclosed embodiment, for each group of sub-text sets, the text annotation device initiates a text annotation request, calls the determined target large language model, and performs attribute identification on each text in the sub-text set. Exemplarily, the semantic content of the text can be analyzed based on the target large language model to obtain the attribute label or attribute vector of the text. After performing attribute identification, the target large language model marks the text in the sub-text set with the greatest matching degree with the preset target attribute as the first text, and marks the text with the least matching degree with the preset target attribute as the second text, and outputs the first text and the second text to the text annotation device. The above method of determining the first text and the second text in the sub-text set through the large prediction model can significantly improve the efficiency of annotation because the large language model has efficient reasoning ability and can quickly process large-scale text data. In addition, the target large language model is usually trained with a large amount of data, has good generalization performance, and can cope with diverse text data and complex annotation requirements.

[0069] It should be noted that in the embodiments of the present disclosure, the attributes of the texts in multiple groups of sub-text sets can also be batch recognized based on the target large prediction model.

[0070] In step S303, for each text to be annotated in the text set to be annotated, based on the first total number of times the text appears in the multiple subsets of texts, the second total number of times the text is marked as the first text in different subsets of texts, and the third total number of times the text is marked as the second text in different subsets of texts, determine the first matching degree of the text with the preset target attribute, and perform matching degree annotation on the text.

[0071] In the embodiments of the present disclosure, for each text to be annotated in the text set to be annotated, the text annotation device counts the total number of times the text appears in multiple subsets of texts, and records this total number of times as the first total number of times. The first total number of times represents the total number of times the text is used for attribute recognition, that is, the total number of times it is evaluated. Then, count the number of times the text is marked as the first text in different subsets of texts, and record it as the second total number of times. The second total number of times reflects the overall closeness of the text to the target attribute. Next, count the number of times the text is marked as the second text in different subsets of texts, and record it as the third total number of times. The third total number of times reflects the overall remoteness of the text from the target attribute. Finally, the text annotation device determines the first matching degree of the text with the preset target attribute based on the first total number of times, the second total number of times, and the third total number of times. In some embodiments, the text annotation device may, based on the matching degree calculation model, input the first total number of times, the second total number of times, and the third total number of times of the text into the calculation model to obtain the result output of the first matching degree with the preset target attribute. Among them, the matching degree calculation model can be obtained by training and tuning parameters of networks such as Convolutional Neural Networks (CNN) and Deep Neural Networks (DNN) based on a large amount of training sample data and label values. Among them, the training sample data may include the first total number of times, the second total number of times, and the third total number of times samples of multiple groups of texts, and the label value is the preset matching degree, and the preset matching degree can be determined based on experience.

[0072] In some other embodiments, Figure 3 the determination process of the first matching degree in step S303 shown above can also be implemented by the following solution:

[0073] Determine the ratio of the difference between the second total number of times and the third total number of times to the first total number of times as the first matching degree of the text with the preset target attribute.

[0074] In the embodiments of the present disclosure, since the second total number reflects the overall proximity of the text to the preset target attribute, and the third total number reflects the overall remoteness of the text from the target attribute. When calculating the matching degree between the text and the preset target attribute, the text annotation device first calculates the difference between the second total number and the third total number. This difference can be regarded as the net effect of the content in the text that is relatively relevant to the preset target attribute. The larger this difference is, the more the content related to the preset target attribute in the text is relative to the irrelevant content. Then, the text annotation device divides this difference by the first total number to obtain a ratio, and determines this ratio as the first matching degree between the text and the preset target attribute. Among them, the first matching degree is a relative measure used to measure the overall association degree between the text and the target attribute. In this way, the matching degree between the text and the target attribute can be simply and intuitively represented. Exemplarily, the text set to be annotated is [data 1, data 2, data 3, data 4, data 5, data 6,...], and sub-text sets 1, 2, and 3 are obtained by sampling this text set. Among them, sub-text set 1 is [data 1, data 2, data 3, data 4], sub-text set 2 is [data 1, data 2, data 4, data 6], and sub-text set 3 is [data 1, data 5, data 6, data 7]. Assume that data 1 is marked as the first text relative to the positive sentiment in text sets 1 and 2, and is marked as the second text relative to the positive sentiment in text set 3. Then, based on the above method for calculating the first matching degree, the first matching degree of data 1 relative to the positive sentiment can be obtained as (2 - 1) / 3 = 1 / 3.

[0075] In the embodiments of the present disclosure, after the text annotation device determines the first matching degree between the text and the preset target attribute, it can perform matching degree annotation on the text based on the first matching degree. Since the first matching degree of the text essentially reflects the degree of closeness of the text to the preset target attribute. When there are differences in the first matching degrees of different texts with the preset target attribute, it can be considered that these texts are different in terms of their proximity to the preset target attribute. Exemplarily, assume that the first matching degree of text 1 with the preset target attribute is 0.85, and the first matching degree of text 2 with the preset target attribute is 0.9. Then it can be judged that text 2 is closer to the preset target attribute. Therefore, more refined and different attribute annotations can be assigned to these two texts respectively based on the first matching degree and the preset target attribute.

[0076] In the embodiments of the present disclosure, multiple groups of sub-text sets are generated from the text set to be labeled, and for each group of sub-text sets, the first text closest to the preset target attribute and the second text most deviated from the preset target attribute within the sub-text set are determined. By counting the second total number of times each text to be labeled is marked as the first text and the third total number of times it is marked as the second text in different sub-text sets, the overall proximity degree of the text to the target attribute is reflected by the second total number, and the overall remoteness degree of the text from the target attribute is reflected by the third total number. Moreover, the text is labeled with a matching degree based on the first matching degree determined by the first total number, the second total number, and the third total number of times the text is evaluated in multiple groups of sub-text sets. Since the first matching degree of the text is a quantitative representation of the closeness of the text to the target attribute, and the first matching degrees of different texts with the preset target attribute may vary, then labeling the text based on the first matching degree with the preset target attribute provides a more refined evaluation index for the labeling of the text on the basis of the preset target attribute. Therefore, a more refined attribute labeling can be obtained for the text, thereby realizing the automatic labeling of the continuous value attribute of the text.

[0077] Figure 4 is a schematic flowchart of a method for generating a sub-text set provided by an embodiment of the present application. In the following text, the generation process of the above-mentioned multiple groups of sub-text sets will be described in conjunction with Figure 4 the steps S401 - S403 shown.

[0078] In step S401, the text set is randomly sampled to generate multiple groups of first sample sub-text sets; wherein, different groups of first sample sub-text sets are different, and the texts to be labeled in each group of first sample sub-text sets are different;

[0079] In step S402, based on the multiple groups of first sample sub-text sets, the frequency of occurrence of the first combined text of each first preset combined length in the multiple groups of first sample sub-text sets is determined; wherein, the first combined text includes multiple texts arbitrarily combined in any group of first sample sub-text sets, and the number of texts indicated by the combination of the first preset combined length is less than the total number of texts included in the corresponding first sample sub-text set;

[0080] In step S403, based on the degree of dispersion of the frequencies of the first combined texts, the multiple groups of sub-text sets are determined.

[0081] In the embodiments of the present disclosure, the text labeling device can generate multiple groups of different first sample sub-text sets by implementing a multiple random sampling strategy on the text set to be labeled. Here, the generated first sample sub-text sets are different, and the texts in each group of first sample sub-text sets are also different, that is, the texts cannot be repeated within each group of first sample sub-text sets. The random sampling method adopted in the embodiments of the present disclosure can be simple random sampling, stratified sampling, etc.

[0082] In an embodiment of the present disclosure, the text annotation device determines a first preset combination length, which indicates how many texts are to be selected from each group of first sample sub-text sets to form a first combined text. Then, for each group of first sample sub-text sets, texts therein are randomly sampled without replacement, and all first combined texts are generated according to the first preset combination length. And the frequency of occurrence of each first combined text in multiple groups of first sample sub-text sets is counted. Here, the frequency can refer to the total number of times a single first combined text appears in all first sample sub-text sets. In this way, by counting the frequencies of occurrence of each first combined text in multiple groups of first sample sub-text sets, a statistical result of the frequencies of all first combined texts can be obtained. Here, to satisfy the traversal of the first combined text in the first sample sub-text set to obtain the frequency of occurrence of the first combined text, the first preset combination length needs to be less than the total number of texts included in the first sample sub-text set corresponding to the first combined text. Exemplarily, the first sample sub-text set 1 is [data 1, data 2, data 3, data 4]. Assuming the first preset combination length is 2, then (data 1, data 2) can be determined as the first combined text 1. For the first combined text 1, the frequency of occurrence of this combined text in all first sample sub-text sets is traversed. This step is repeatedly executed in this way to obtain the frequencies corresponding to all first combined texts.

[0083] It should be noted that the composition of the first combined text is closely related to the arrangement order of its internal text elements. For example, the sub-text set 1 is [data 1, data 2, data 3, data 4], and the sub-text set 2 is [data 2, data 1, data 3, data 4]. Here, since the arrangement orders of data 1 and data 2 are different in these sub-text sets, (data 1, data 2) and (data 2, data 1) are regarded as different first combined texts.

[0084] In the embodiments of the present disclosure, the text annotation device calculates the degree of dispersion of the frequencies of each first combined text. Here, the degree of dispersion can be measured by statistical quantities such as the variance and standard deviation of the frequencies. The degree of dispersion here reflects the fluctuation of the frequencies of different first combined texts in all first sample subsets of texts. Among them, the smaller the degree of dispersion, the smaller the frequency fluctuation of different first combined texts, which means that the frequencies of these first combined texts in all first sample subsets of texts are more balanced, that is, the first combined texts present a relatively stable distribution pattern. Then, the first sample subset of texts may include information that is more valuable for the text annotation process. On the contrary, if the degree of dispersion is larger, it indicates that the frequency fluctuation of different first combined texts is larger, which means that the occurrence times of different combined texts are more unbalanced, that is, some first combined texts may dominate in the text set, while other first combined texts appear less frequently. Then, the first sample subset of texts may contain more irrelevant information and can be excluded. In actual usage requirements, the smaller the degree of dispersion, that is, the more balanced the occurrence frequencies of the first combined texts, the more beneficial it is to the text standardization process. Then, the text annotation device determines multiple subsets of texts according to the magnitudes of multiple degrees of dispersion. Exemplarily, the text annotation device may determine multiple first sample subsets of texts corresponding to each first combined text whose degree of dispersion satisfies being less than a preset degree-of-dispersion threshold as multiple subsets of texts.

[0085] In the embodiments of the present disclosure, by analyzing the degree of dispersion of the frequencies of the first combined texts in multiple first sample subsets of texts to determine multiple subsets of texts for subsequent text annotation processing, the range of the subsets of texts can be more precisely defined, which is beneficial to improving the efficiency and accuracy of subsequent text annotation tasks.

[0086] In some embodiments, the degree of dispersion is characterized by the standard deviation. The step S403 shown above Figure 4 can also be implemented by the following solution:

[0087] In the case where the standard deviation is less than a preset standard-deviation threshold, determine the multiple first sample subsets of texts as the multiple subsets of texts;

[0088] In the case where the standard deviation is greater than or equal to the preset standard-deviation threshold, randomly resample to generate multiple first sample subsets of texts until the standard deviation of the frequencies of each first combined text determined based on the regenerated multiple first sample subsets of texts is less than the preset standard-deviation threshold, and determine the regenerated multiple first sample subsets of texts as the multiple subsets of texts.

[0089] In the embodiments of the present disclosure, after calculating the standard deviation of the frequencies of each first combined text, the text annotation device compares the obtained standard deviation with a preset standard deviation threshold, where the preset standard deviation threshold can be set according to specific circumstances. Here, the standard deviation is used to characterize the degree of dispersion of the frequencies of the first combined texts. If the standard deviation of these frequencies is small, it indicates that the fluctuations in the frequencies of different first combined texts are smaller, meaning that the frequencies of these first combined texts in all the first sample subsets are more balanced. On the contrary, if the standard deviation of these frequencies is large, it indicates that the fluctuations in the frequencies of different first combined texts are larger, meaning that the occurrences of different combined texts are more unbalanced. In the embodiments of the present disclosure, if the standard deviation is less than the preset standard deviation threshold, it is considered that the texts in the current multiple groups of first sample subsets include information that is more valuable for the text annotation process, and the text annotation device determines them as the final multiple groups of subsets.

[0090] In the embodiments of the present disclosure, if the standard deviation is greater than or equal to the preset standard deviation threshold, it is considered that the degree of dispersion of the current multiple groups of first sample subsets is relatively high, and it may not be suitable to implement the text annotation method in this embodiment. Then the text annotation device re-performs the aforementioned random sampling operation on the text set to be annotated, regenerates multiple groups of first sample subsets, and repeats the above steps (including calculating the standard deviation of the frequencies of each first combined text and comparing the magnitudes of the standard deviations), until the standard deviation of the frequencies determined based on the regenerated multiple groups of first sample subsets is less than the preset threshold. Only at this time will the regenerated multiple groups of first sample subsets be determined as multiple groups of subsets.

[0091] In the embodiments of the present disclosure, by introducing the statistic of the standard deviation to quantify the degree of dispersion of the frequencies of the first combined texts, and then quantitatively describing the balance degree of the frequencies of the first combined texts in multiple groups of first sample subsets, it helps to accurately screen out multiple groups of subsets that include information more valuable for the text annotation process. And the screening process continues until the requirements are met, thereby further strengthening the consistency of the texts in the multiple groups of subsets and greatly improving the reliability and accuracy of the annotation results based on the multiple groups of subsets.

[0092] In the embodiments of the present disclosure, as described above, the text annotation device can also select a target large language model so that the text annotation device implements the text annotation method based on the target large language model. Refer to Figure 5 , Figure 5 is a flowchart showing a method for determining a target large language model provided by an embodiment of the present application. The following will be described in conjunction with Figure 5 the steps S501 - S504 shown to illustrate the determination process of the target large language model.

[0093] In step S501, multiple groups of test sample text sets are obtained, as well as the attribute labels of each test sample text set.

[0094] In step S502, the multiple alternative large language models are used to process the multiple groups of test sample text sets, and the predicted attributes determined for each group of test sample text sets based on each alternative large language model are obtained.

[0095] In step S503, for each alternative large language model, based on the attribute labels of each group of test sample text sets and the predicted attributes determined for each group of test sample text sets based on the alternative large language model, the prediction ability of the alternative large language model is determined.

[0096] In step S504, based on the prediction ability of each alternative large language model, the target large language model is determined.

[0097] In the embodiments of the present disclosure, the text annotation device obtains multiple groups of test sample text sets. Here, the multiple groups of test sample text sets can be preset sample text sets. The text annotation device also obtains the attribute labels of each group of test sample text sets. Among them, the attribute labels can be the attribute labels obtained by manually annotating the test sample text sets, or the attribute labels obtained by annotating the test sample text sets based on a certain known text annotation model with relatively high accuracy. Here, the attribute labels can be a set of attribute sorting results determined based on the attributes of each text in the test sample text set.

[0098] In the embodiments of the present disclosure, the text annotation device will use multiple pre-prepared alternative large language models to process each group of test sample text sets. Each alternative large language model will generate predicted attributes for each group of test sample text sets according to its own algorithm and training data. Among them, the predicted attributes can be a set of attribute sorting results obtained by the alternative large language model for attribute prediction of multiple groups of test sample text sets.

[0099] In the embodiments of the present disclosure, for each alternative large language model, the text annotation device determines the prediction ability of the large language model by comparing the predicted attributes obtained by the alternative large language model for attribute prediction of each group of test sample text sets with the attribute labels of the group of test sample text sets. Exemplarily, if the predicted attributes of a certain group of test sample text sets can match the attribute labels of a certain group of test sample text sets, then it is considered that the large language model accurately predicts the attributes of this group of test sample text sets. Then, the prediction accuracy rate or other relevant evaluation indicators of the large language model can be calculated, so as to quantify the prediction ability of the large language model based on these evaluation indicators and understand the performance differences of different large language models in the text annotation function.

[0100] In the embodiments of the present disclosure, the text annotation device ranks or scores multiple alternative large language models according to the prediction ability evaluation index calculated in step S503, so that the text annotation device can select the large language model with the strongest prediction ability and the best performance as the target large language model for subsequent text annotation tasks.

[0101] In the embodiments of the present disclosure, by comparing the prediction abilities of multiple alternative large language models on multiple groups of test sample sub-text sets, the target large language model can be quickly and accurately determined, which helps to improve the accuracy of text annotation.

[0102] In some embodiments, Figure 5 The shown step S501 can be implemented by the following solution:

[0103] Randomly sample a preset test text set to generate multiple groups of second sample sub-text sets; wherein, different groups of second sample sub-text sets are different, and the test texts in each group of second sample sub-text sets are different;

[0104] Based on the multiple groups of second sample sub-text sets, determine the frequency of occurrence of each second combined text with a second preset combined length in the multiple groups of second sample sub-text sets; wherein, the second combined text includes multiple test texts arbitrarily combined in any group of second sample sub-text sets, and the number of texts in the combination indicated by the second preset combined length is less than or equal to the total number of texts included in the corresponding second sample sub-text set;

[0105] Based on the dispersion degree of the frequencies of each second combined text, determine multiple groups of intermediate sub-text sets;

[0106] For each group of intermediate sub-text sets, select third texts outside the test texts included in the intermediate sub-text set from the test text set multiple times, and insert the third texts into the intermediate sub-text set to obtain multiple groups of expanded text sets with expanded text quantities; wherein, the number of test texts included in each group of expanded text sets is the same; the multiple groups of expanded text sets corresponding to the multiple intermediate sub-text sets form the multiple groups of test sample text sets.

[0107] In the embodiments of the present disclosure, the text annotation device uses the method of random sampling to extract texts from a preset test text set to generate multiple groups of second sample sub-text sets. Each group of second sample sub-text sets contains different test texts, and different groups of second sample sub-text sets are also different from each other. Among them, for the specific process of generating multiple groups of second sample sub-text collections, please refer to the specific description of obtaining multiple groups of sub-text sets by random sampling based on a text set in the foregoing embodiments.

[0108] In the embodiments of the present disclosure, the text annotation device defines a second preset combination length, which indicates how many texts are to be selected from each group of second sample sub-text sets to form a second combined text. Similarly, to satisfy the traversal of the second combined text in the second sample sub-text sets to obtain the frequency of occurrence of the second combined text, the second preset combination length needs to be less than or equal to the total number of texts included in the second sample sub-text set corresponding to the second combined text. Here, the process of determining the frequency of occurrence of each second combined text in multiple groups of second sample sub-text sets can refer to the specific description of the frequency of occurrence of each first combined text in multiple groups of first sample sub-text sets in the foregoing embodiments.

[0109] In the embodiments of the present disclosure, the text annotation device calculates the degree of dispersion based on the frequency distribution of each second combined text to evaluate the balance of the frequencies of occurrence of different second combined texts in multiple groups of second sample sub-text sets. Here, the degree of dispersion can also be measured by statistical quantities such as the variance and standard deviation of the frequencies. Specifically, reference can be made to the description of the degree of dispersion of the frequencies of the first combined text. In the embodiments of the present disclosure, the text annotation device may determine the multiple groups of second sample sub-text sets corresponding to the second combined texts with a smaller degree of dispersion as multiple groups of intermediate sub-text sets.

[0110] In the embodiments of the present disclosure, for each group of intermediate sub-text sets, the text annotation device performs multiple sampling operations. Each time, a third text that does not belong to the current intermediate sub-text set is randomly selected from a preset test text set. And the selected multiple third texts are inserted into the intermediate sub-text set to form an expanded text set with an increased number of texts. The number of test texts included in each group of expanded text sets is the same to maintain the consistency of the number of the data sets. Here, multiple insertions can be performed on the intermediate sub-text set based on multiple third texts. During the insertion processes of different batches, the insertion positions of the third texts may be different. In this way, based on one intermediate sub-text set, multiple different expanded text sets can be obtained by expanding the number of the intermediate sub-text set using multiple third texts, that is, multiple groups of expanded text sets can be obtained by expanding the number of one group of intermediate sub-text sets. Exemplarily, assume that the preset test text set is [data 1, data 2, data 3, data 4, data 5, data 6,...], and one group of intermediate sub-text sets is [data 1, data 2]. Then, multiple data can be sampled from the test text set multiple times, and the intermediate sub-text set can be inserted and expanded based on the sampled multiple data. For example, the expanded text set 1 [data 1, data 2, data 3, data 4], the expanded text set 2 [data 1, data 3, data 2, data 5] can be obtained.

[0111] In the embodiments of the present disclosure, if the above process is implemented once for each group of intermediate sub-text sets, then by repeating the above process multiple times, multiple groups of intermediate sub-text sets can generate corresponding multiple groups of expanded text sets, and finally multiple groups of test sample text sets are formed.

[0112] In the embodiments of the present disclosure, by performing frequency analysis and dispersion degree evaluation on the second sample sub-text set, multiple groups of intermediate sub-text sets with specific diversity characteristics are determined. Moreover, by inserting the third text, the flexible expansion of the intermediate sub-text sets is realized, meeting the diverse requirements for text data in the actual scenario of text annotation.

[0113] In some embodiments, Figure 5 The shown step S503 can be implemented through the following solution:

[0114] For each group of intermediate sub-text sets, use the alternative large language model to determine the matching degree between each test text in each group of extended text sets associated with the intermediate sub-text sets and the preset target attribute, and obtain the matching degree ranking of the test texts in each group of extended text sets;

[0115] For each group of intermediate sub-text sets, if the matching degree rankings of multiple test texts in the intermediate sub-text sets are the same in multiple groups of extended text sets, determine that the multiple test texts in the intermediate sub-text sets belong to the first target text group for evaluating the prediction ability;

[0116] Based on the total number of the first target text groups determined by multiple intermediate sub-text sets and the total number of multiple intermediate sub-text sets, determine the prediction ability of the alternative large language model; wherein, the prediction ability is positively correlated with the total number of the first target text groups.

[0117] In the embodiments of the present disclosure, after obtaining multiple groups of extended text sets generated corresponding to all groups of intermediate sub-text sets, for each group of intermediate sub-text sets, the text annotation device determines the attributes of each test text in the extended text of this group according to the extended text sets corresponding to this group of intermediate sub-text sets, and uses the alternative large language model, so as to determine the matching degree between the attributes of each test text in the extended text of this group and the preset target attribute based on the attributes of each test text in the extended text of this group and the preset target attribute.

[0118] In the embodiments of the present disclosure, for each group of intermediate sub-text sets, the text annotation device also sorts the test texts in each group of extended text sets corresponding to this group of intermediate sub-text sets according to the matching degree, and obtains the matching degree ranking of the test texts in the extended text sets corresponding to this group. Exemplarily, a group of intermediate sub-text sets is [Data 1, Data 2]. Based on this intermediate sub-text set for expansion, an extended text set 1 [Data 1, Data 2, Data 3, Data 4] is obtained. Then, the matching degrees of the 4 data in the extended text set 1 with the positive sentiment can be recognized, and the matching degree ranking within the extended text set 1 can be obtained by sorting the matching degrees of the 4 data in the extended text set 1.

[0119] In the embodiments of the present disclosure, for each set of intermediate sub-text sets, the text annotation device checks whether the matching degree rankings of multiple test texts in the intermediate sub-text sets with respect to the test texts in multiple sets of augmented text sets are consistent. If the matching degree rankings of multiple test texts in the intermediate sub-text set are consistent in multiple sets of augmented text sets, then these test texts are determined as the first target text group for evaluating the prediction ability. That is, only when the matching degree rankings of these test texts are consistent in all augmented text sets, will this intermediate sub-text set be classified into the first target text group to determine the prediction ability of the alternative large language model based on the stability of the prediction results. Exemplarily, assume that a set of intermediate sub-text sets is [Data 1, Data 2]. Based on this set of intermediate text sets, augmented text set 1 [Data 1, Data 2, Data 3, Data 4] and augmented text set 2 [Data 1, Data 3, Data 2, Data 5] are obtained. The matching degree ranking in augmented text set 1 is Data 3 > Data 1 > Data 2 > Data 4, and the matching degree ranking in augmented text set 2 is Data 5 > Data 4 > Data 1 > Data 2. Then, for Data 1 and Data 2, since their matching degree rankings are consistent in the two sets of augmented text sets, it can be considered that this set of intermediate sub-text sets is a stable prediction pair for this alternative large language model, and it is determined as the first target text group.

[0120] In the embodiments of the present disclosure, the text annotation device counts the total number of the first target text groups determined in multiple intermediate sub-text sets, and determines the prediction ability of the alternative large language model based on the total number of the first target text groups and the total number of multiple intermediate sub-text sets. The prediction ability is positively correlated with the total number of the first target text groups. That is, the larger the number of the first target text groups, the stronger the prediction ability of this alternative large language model on specific target attributes. Exemplarily, the prediction ability of the alternative large language model can be determined according to the ratio between the total number of the first target text groups and the total number of multiple intermediate sub-text sets; or the prediction ability of the alternative large language model can be determined according to the difference between the total number of the first target text groups and the total number of multiple intermediate sub-text sets.

[0121] In the embodiments of the present disclosure, by comparing the consistency of the matching degree rankings of multiple test texts in the intermediate sub-text sets with respect to multiple sets of augmented text sets to determine the prediction ability of the alternative large language model based on the stability of the prediction results, this method can measure the consistency and reliability of the model in different scenarios, thereby effectively improving the effectiveness and accuracy of selecting the target large language model.

[0122] In some embodiments, the method for determining the prediction ability of the above-mentioned large language model further includes:

[0123] For each first target text group, when the matching degree rankings of multiple test texts in the first target text group in the multiple groups of augmented text sets are consistent with the preset matching degree rankings of the multiple test texts, determine the first target text group as a second target text group;

[0124] Based on the total number of second target text groups determined from multiple first target text groups and the total number of multiple intermediate sub-text sets, determine the prediction ability of the alternative large language model; wherein, the prediction ability is positively correlated with the total number of second target text groups.

[0125] In the embodiments of the present disclosure, for each first target text group, the text annotation device compares whether the matching degree rankings of multiple test texts in the first target text group in the multiple groups of augmented text sets are consistent with the preset matching degree rankings of the multiple test texts. If they are consistent, then determine the first target text group as a second target text group. That is, only when the matching degree rankings of multiple test texts in the first target text group in all augmented text sets are consistent and are consistent with the preset matching degree rankings of the multiple test texts, will the first target text group be classified into the second target text group for evaluating the prediction ability, so as to determine the prediction ability of the alternative large language model based on the accuracy of the prediction results. Exemplarily, assume that a group of intermediate sub-text sets is [Data 1, Data 2], and the preset matching degree ranking of the two test texts in the intermediate sub-text set is Data 1 > Data 2. Based on this group of intermediate sub-text sets, data augmentation is performed to augmented text set 1 [Data 1, Data 2, Data 3, Data 4] and augmented text set 2 [Data 1, Data 3, Data 2, Data 5]. The matching degree ranking in augmented text set 1 predicted by the alternative large language model is Data 3 > Data 1 > Data 2 > Data 4, and the matching degree ranking in augmented text set 2 predicted by the alternative large language model is Data 5 > Data 4 > Data 1 > Data 2. Then for Data 1 and Data 2, their matching degree rankings in the two augmented text sets are Data 1 > Data 2, which are both consistent with the preset matching degree rankings of the two test texts in the intermediate sub-text set. It can be considered that this group of intermediate sub-text sets is a correct prediction pair relative to the alternative large language model, and determine it as a second target text group.

[0126] In the embodiments of the present disclosure, the text annotation device counts the total number of the second target text groups determined in multiple intermediate sub-text sets and the total number of the multiple intermediate sub-text sets, and determines the prediction ability of the alternative large language model. The prediction ability is positively correlated with the total number of the second target text groups, that is, the larger the total number of the second target text groups, the stronger the prediction ability of the alternative large language model on a specific task. Exemplarily, the prediction ability of the alternative large language model can be determined according to the ratio between the total number of the second target text groups and the total number of the multiple intermediate sub-text sets; or the prediction ability of the alternative large language model can be determined according to the difference between the total number of the second target text groups and the total number of the multiple intermediate sub-text sets.

[0127] In the embodiments of the present disclosure, by comparing the consistency between the matching degree ranking of the test text and the preset matching degree ranking, the second target text group is determined, and the prediction ability of the alternative large language model is determined based on the accuracy of the prediction result. This method can measure the accuracy of the alternative large language model in different scenarios, thereby further improving the effectiveness and accuracy of selecting the target large language model.

[0128] In some embodiments, the above text annotation method further includes:

[0129] Combining the first text and the second text determined based on each group of sub-text sets into a standby text set, and randomly dividing the standby text set into multiple groups of standby sub-text sets; wherein, the number of texts to be annotated in each group of standby sub-text sets is the same;

[0130] For each text to be annotated in each group of standby sub-text sets, count the fourth total number of times that each text to be annotated in the text set to be annotated is marked as the first text in the standby sub-text set, and the fifth total number of times that the text is marked as the second text in the standby sub-text set, and determine the second matching degree between the text and the preset target attribute based on the first total number of times the text appears in the multiple groups of sub-text sets, the fourth total number of times corresponding to the text, and the fifth total number of times;

[0131] For each group of standby sub-text sets, perform a matching degree ranking based on the second matching degree of each text;

[0132] Based on the matching degree ranking results of the texts in each group of standby sub-text sets, determine the correlation of the ranking positions of the same text in different standby sub-text sets;

[0133] The determining the first matching degree between the text and the preset target attribute based on the first total number of times the text appears in the multiple groups of sub-text sets, the second total number of times the text is marked as the first text in different sub-text sets, and the third total number of times the text is marked as the second text in different sub-text sets includes:

[0134] When the relevance corresponding to the text is greater than or equal to a preset relevance threshold, based on the first total number, the second total number, and the third total number corresponding to the text, determine the first matching degree between the text and the preset target attribute.

[0135] In an embodiment of the present disclosure, the text annotation device determines the first text and the second text corresponding to each group of sub-text sets based on each group of sub-text sets, and then forms a backup text set from multiple groups of first texts and second texts as the input for subsequent processing.

[0136] In an embodiment of the present disclosure, the text annotation device also randomly divides the backup text set into multiple groups of backup sub-text sets. Here, the number of texts to be annotated in each group of backup sub-text sets is the same, that is, the backup text set is divided into multiple groups of backup sub-text sets with equal lengths.

[0137] In an embodiment of the present disclosure, for each group of backup sub-text sets, the text annotation device counts the fourth total number of times that each text to be annotated in the aforementioned text set to be annotated is marked as the first text and the fifth total number of times that it is marked as the second text in this group of backup sub-text sets. And based on the first total number of times that this text appears in the aforementioned multiple groups of sub-text sets, the fourth total number of times and the fifth total number of times in this group of backup sub-text sets, calculate the second matching degree between this text and the preset target attribute in this group of backup sub-text sets. For the specific calculation method of the second matching degree, reference can be made to the detailed description of the aforementioned first matching degree calculation method.

[0138] In an embodiment of the present disclosure, after determining the second matching degree of each text, for each group of backup sub-texts, the text annotation device sorts each text according to the second matching degree of each text to obtain the matching degree sorting result of each text in each group of backup sub-text sets. Exemplarily, the backup sub-text set 1 is [data 1, data 2, data 3, data 4]. Assume that the calculated second matching degree between data 1 and the preset target attribute is 1 / 3, data 2 is 1 / 2, data 3 is 1 / 5, and data 4 is 1 / 4. Then the descending matching degree sorting result of the backup sub-text set 1 is data 2 > data 1 > data 4 > data 3. It should be noted that the order of the matching degree sorting result in the embodiment of the present disclosure is not limited, and it can be from large to small or from small to large based on the second matching degree.

[0139] In the embodiments of the present disclosure, based on the sorting results of the matching degrees of the texts in each set of alternative sub-text sets, if there is a certain text that is the same in each alternative sub-text set, then this certain text is marked as "the same text". For these same texts, based on their sorting positions in different alternative sub-text sets, the relevance of the texts is calculated. According to the calculated relevance of the texts, where the larger the relevance value indicates that the sorting results of the texts in each alternative sub-text set are closer, it can be considered that the marking result of the text is more reliable. Exemplarily, the relevance of the texts can be determined based on methods such as the Spearman correlation coefficient and the Pearson correlation coefficient.

[0140] It should be noted that the alternative text set only includes the text data marked as the first text or the second text, so the alternative sub-text sets generated according to the alternative text set also only include the text data marked as the first text or the second text. And the text set to be annotated includes all the text data to be annotated, which may result in the situation that some text data in the text set to be annotated do not exist in the alternative text set or the alternative sub-text set. In the embodiments of the present disclosure, for each set of alternative sub-text sets, the total fourth number of times that each text to be annotated in the text set to be annotated is marked as the first text in the alternative sub-text set, and the total fifth number of times that the text is marked as the second text in the alternative sub-text set are counted. Here, since some texts may not exist in the alternative sub-text set, the total fourth number of times and the total fifth number of times that the text is counted are 0, that is, the second matching degree of the text with the preset target attribute calculated based on the total first number, the total fourth number, and the total fifth number is also 0.

[0141] In the embodiments of the present disclosure, when the relevance of the text at the corresponding sorting position is greater than or equal to the preset relevance threshold, it is considered that the marking result obtained based on multiple sets of sub-text sets is relatively reliable. Therefore, the first matching degree of the text with the preset target attribute can be comprehensively determined based on the total first number of times the text appears in multiple sets of sub-text sets, the total second number of times the text is marked as the first text, and the total third number of times the text is marked as the second text.

[0142] In the embodiments of the present disclosure, by grouping the annotation results composed of multiple sets of the first text and the second text, and for each set of alternative sub-text sets, calculating the second matching degree of each text to be annotated in the text set to be annotated for sorting based on the matching degree, so as to determine the relevance of the sorting positions of the same text in different sets of alternative sub-text sets, and only when the relevance is greater than or equal to the preset relevance threshold, the first matching degree of each text in the text set with the preset target attribute is calculated. Through this method, before calculating the first matching degree, by grouping for comparing the matching degree differences, the relevance and differences between different texts are comprehensively considered, which is beneficial to improving the stability and reliability of the results.

[0143] In some embodiments, the above text annotation method further includes:

[0144] Performing a preprocessing operation on the text set to be annotated to obtain a processed text set; wherein, the preprocessing operation includes at least one of the following at least:

[0145] For each text to be annotated in the text set, removing the preset strings in the text and removing the repeated strings with a length greater than the preset length threshold in the text;

[0146] Filtering the texts in the text set with a similarity greater than the preset similarity threshold;

[0147] The dividing the text set into multiple groups of sub-text sets includes:

[0148] Dividing the processed text set into the multiple groups of sub-text sets.

[0149] In the embodiments of the present disclosure, before performing text annotation on the text set to be annotated, the text annotation device also performs a preprocessing operation on the text set to be annotated to optimize the data to be annotated. Among them, the preprocessing operation includes various processing methods for the text set.

[0150] In the embodiments of the present disclosure, the preprocessing operation includes removing the preset strings. Specifically, the text annotation device removes the preset strings in each text to be annotated in the text set. These preset strings may include insignificant content or content without natural semantics, such as punctuation marks, symbols, special characters, links, etc. They may have no practical value for subsequent text analysis, so they need to be removed.

[0151] In the embodiments of the present disclosure, the preprocessing operation includes removing the repeated strings with a length greater than the preset length threshold. In the text, there may be a large number of repeated strings. To improve the annotation efficiency, the text annotation device detects the repeated strings in each text to be annotated in the text set and removes these repeated strings in the text, only retaining the unique strings among them.

[0152] In the embodiments of the present disclosure, the preprocessing operation includes filtering similar texts. Here, in addition to performing exact matching on the texts and filtering out the completely identical texts in the text set, the text annotation device will also filter out the texts in the text set with a similarity greater than the threshold according to the preset similarity threshold. This helps to reduce redundant annotations and improve the annotation efficiency.

[0153] In the embodiments of the present disclosure, after completing the preprocessing operation, the processed text set is divided into multiple groups of sub-text sets. Thus, by generating sub-text sets from the text set after the above preprocessing operation and then performing text annotation, the annotation efficiency can be improved.

[0154] In the embodiments of the present disclosure, through multiple preprocessing operations on the text set to be annotated, redundant and irrelevant annotation data as well as highly similar annotation data in the text set are removed, and a sub-text set is generated based on the text set after the preprocessing operations, so as to perform text annotation, which can significantly improve the annotation efficiency.

[0155] Figure 6 It is a flow example diagram of a method for generating a sub-text set provided by an embodiment of the present application. As Figure 6 shown, it includes the following steps:

[0156] S601. Remove irrelevant content and content without natural semantics.

[0157] In the embodiments of the present disclosure, the text annotation device removes irrelevant content and content without natural semantics in the text set to be annotated. For example, it can remove preset strings such as punctuation marks, symbols, special characters, links, etc. or content irrelevant to the text in the text set.

[0158] S602. Exact match deduplication, repeated match deduplication, and fuzzy deduplication.

[0159] In the embodiments of the present disclosure, the text annotation device performs multiple deduplication operations on the text set to be annotated. Among them, exact match deduplication refers to removing completely identical texts in the text set, repeated match deduplication refers to, for each text to be annotated in the text set to be annotated, removing repeated strings with a length greater than a preset length threshold in the text, and fuzzy deduplication refers to filtering texts with a similarity greater than a preset similarity threshold in the text set to be annotated.

[0160] S603. Obtain the data to be annotated [data 1, data 2,...].

[0161] In the embodiments of the present disclosure, the text annotation device obtains the data to be annotated. Among them, the data to be annotated refers to the text set after being processed in the above steps S601 and S602, that is, the text set to be annotated in the embodiments of the present disclosure. The data to be annotated includes multiple text data, such as data 1, data 2, etc.

[0162] S604. Generate 4-tuples [(data 1, data 2, data 3, data 4),...] with the set parameters.

[0163] In the embodiments of the present disclosure, the text annotation device generates multiple 4-tuples [(data 1, data 2, data 3, data 4),...] based on the data to be annotated in step S603 according to the set parameters. Here, the parameters refer to the number of 4-tuples to be generated, and the 4-tuples refer to the first sample sub-text set in the embodiments of the present disclosure.

[0164] S605. Statistically calculate the co-occurrence frequency of any two randomly selected pieces of data, i.e., {(data 1, data 2): n_1, (data 2, data n): n_2, …}.

[0165] In an embodiment of the present disclosure, the text annotation device sets a first preset combination length, which is set to 2 here; and statistically calculates the frequency of occurrence of the first combined text formed by any combination of two text data in all 4-tuples among all 4-tuples. Here, the first combined text is (data 1, data 2), (data 2, data n), etc., and the corresponding frequencies of occurrence are n_1, n_2.

[0166] S606. Calculate the sample standard deviation of the co-occurrence frequencies.

[0167] In an embodiment of the present disclosure, the text annotation device calculates the sample standard deviation of the occurrence frequencies of the first combined text in multiple groups of 4-tuples based on the frequency of each first combined text in step S605. Here, the sample standard deviation refers to the standard deviation used to characterize the degree of dispersion of the frequencies of the foregoing first combined texts.

[0168] S607. Determine whether the sample standard deviation is lower than the current sample standard deviation? If so, execute step S608; if not, execute step S604.

[0169] In an embodiment of the present disclosure, the text annotation device determines whether the sample standard deviation is lower than the current sample standard deviation, where the sample standard deviation refers to a preset standard deviation threshold. If the sample standard deviation is lower than the sample standard deviation, it is considered that the degree of dispersion of the first combined text in all 4-tuples is low, and then step S608 is executed; while if the sample standard deviation is greater than or equal to the sample standard deviation, it is considered that the degree of dispersion of the first combined text in all 4-tuples is large, and this batch of generated multiple 4-tuples does not meet the data requirements for text annotation, so it returns to step S604 to continue generating new 4-tuples.

[0170] S608. Include the 4-tuples to be annotated.

[0171] In an embodiment of the present disclosure, since the degree of dispersion of the first combined text in all 4-tuples is low, multiple 4-tuples [(data 1, data 2, data 3, data 4), …] are retained for subsequent processing.

[0172] Figure 7 This is a flowchart example of a method for determining a target large language model provided by an embodiment of the present application. As Figure 7 shown, it includes the following steps:

[0173] S701. Obtain data [data 1, data 2, …].

[0174] In the embodiments of the present disclosure, a text annotation device obtains data [data 1, data 2,...], where the data refers to a preset test text set in the embodiments of the present disclosure.

[0175] S702. Data preprocessing, duplicate removal, and organization.

[0176] In the embodiments of the present disclosure, the text annotation device performs preprocessing and duplicate removal operations on the data in S701. Specifically, for each text to be annotated in the annotation data, it removes the preset strings in the text. The preset strings can be, for example, punctuation marks, symbols, special characters, links, etc., filters the duplicate strings in the text whose length is greater than the preset length threshold, and filters the texts in the annotation data whose similarity is greater than the preset similarity threshold. And based on the processed data to be annotated, it is organized to obtain multiple groups of second sample sub-text sets. Here, the organization refers to randomly sampling the data to be annotated as described above to obtain multiple groups of second sample sub-text sets, and multiple groups of intermediate sub-text sets determined based on the dispersion degree of the occurrence frequency of the second combined texts with the second preset combined length in the multiple groups of second sample sub-text sets.

[0177] S703. Obtain [(data 1, data 2), (data 3, data 4),...].

[0178] In the embodiments of the present disclosure, the text annotation device obtains multiple groups of intermediate sub-text sets [(data 1, data 2), (data 3, data 4),...].

[0179] S704. Voting by multiple people.

[0180] In the embodiments of the present disclosure, the text annotation device obtains multiple groups of intermediate sub-text sets and the attribute labels of each group of intermediate sub-text sets. Among them, the attribute labels of each group of intermediate sub-text sets are determined by voting by multiple people.

[0181] S705. Obtain an evaluation set [(data 1, data 2, label 1), (data 3, data 4, label 2),...].

[0182] In the embodiments of the present disclosure, the text annotation device obtains an evaluation set, where the evaluation set includes multiple groups of intermediate sub-text sets and the attribute labels of each group of intermediate sub-text sets. For example, (data 1, data 2) is a group of intermediate sub-text sets, and label 1 is the attribute label of this group of intermediate sub-text sets.

[0183] S706. Obtain additional data to be annotated.

[0184] In the embodiments of the present disclosure, for each group of intermediate sub-text sets, the text annotation device repeatedly selects a third text from the test text set that is not included in this intermediate sub-text set. Here, the third text refers to the additional data to be annotated.

[0185] S707. Obtain the standard test set.

[0186] In the embodiments of the present disclosure, for each group of intermediate sub - text sets, the text annotation device inserts the third text into the intermediate sub - text set to obtain multiple groups of augmented text sets with an increased number of texts. Here, the multiple groups of augmented text sets with an increased number of texts are the standard test set. For example, [(data1, data2, data3_newly added, data4_newly added), (data5_newly added, data6_newly added, data2, data1), (data1, data2, data7_newly added, data8_newly added), ((data1, data9_newly added, data2, data10_newly added), (data2, data11_newly added, data1, data12_newly added),...).

[0187] S708. The large language model to be tested.

[0188] In the embodiments of the present disclosure, the large language model to be tested refers to multiple alternative large language models.

[0189] S709. Obtain the test results.

[0190] In the embodiments of the present disclosure, after predicting multiple groups of augmented text sets based on multiple groups of alternative large language models, the test results are obtained. Here, the test results refer to the matching degree sorting results of multiple test texts in the intermediate sub - text set in each group of augmented text sets, including the first text that best conforms to the target attribute and the second text that least conforms to the target attribute. Here, the target attribute refers to the preset target attribute in the foregoing embodiments. Here, for the augmented text set 1 (data1, data2, data3_newly added, data4_newly added), the text that best conforms to the target attribute is data3, and the text that least conforms to the target attribute is data1. That is, the first text of the augmented text set 1 is data3, and the second text is data1.

[0191] S710. Result processing.

[0192] In the embodiments of the present disclosure, the text annotation device processes the text marking results of the augmented text set 1. For the augmented text set 1 (data1, data2, data3_newly added, data4_newly added), according to the first text predicted by the alternative large language model as data3 and the second text as data1, it can be deduced that there are 5 cases of the distance from the target attribute for the augmented text set 1: data3 > data1, data3 > data2, data3 > data4, data2 > data1, data4 > data1. If the label 1 is data2 > data1, then it can be considered that the preset matching degree sorting in this group of intermediate sub - text sets is consistent with the matching degree sorting in the augmented text set 1.

[0193] S711. Ability analysis.

[0194] In the embodiments of the present disclosure, the text annotation device analyzes the prediction ability of alternative large language models from two aspects: prediction stability and prediction accuracy. For each set of intermediate sub-text sets, when the text annotation device determines that the matching degree rankings of multiple test texts in this set of intermediate sub-text sets in multiple associated extended text sets are all consistent, it determines this set of intermediate sub-text sets as a stable prediction pair, that is, the first target text group, and determines the prediction ability of this alternative large language model based on the ratio of the total number of the first target text groups determined based on multiple intermediate sub-text sets to the total number of multiple intermediate sub-text sets.

[0195] For each set of intermediate sub-text sets, when the matching degree ranking of the test texts in the first target text group is consistent with the preset matching degree ranking of the test texts, the text annotation device determines the first target text group as an accurate prediction pair, that is, the second target text group, and determines the prediction ability of this alternative large language model based on the ratio of the total number of the second target text groups determined based on multiple intermediate sub-text sets to the total number of multiple intermediate sub-text sets.

[0196] Figure 8 is a schematic diagram of a text batch annotation method provided by an embodiment of the present application, as Figure 8 shown, including the following steps:

[0197] S801. Obtain the quadruple to be annotated.

[0198] In the embodiments of the present disclosure, the text annotation device obtains the quadruple to be annotated, and here the quadruple refers to the sub-text set. [(data1, data2, data3, data4), …] are multiple sets of sub-text sets, and the quadruple (data1, data2, data3, data4) is a sub-text set.

[0199] S802. Utilize the large language model and the inference engine.

[0200] In the embodiments of the present disclosure, the text annotation device performs attribute recognition on the texts in each sub-text set based on the large language model and the inference engine to obtain the first text and the second text in each sub-text set. Here, the large language model refers to the target large language model determined based on the prediction abilities of multiple alternative large language models in the embodiments of the present disclosure.

[0201] S803. The batch inference interface is used for batch automated annotation.

[0202] In the embodiments of the present disclosure, the text annotation device batch-executes the inference interface to achieve batch automated annotation of the sub-text sets.

[0203] Figure 9 is a schematic flowchart of a text annotation result consistency evaluation method provided by an embodiment of the present application, as Figure 9As shown, it includes the following steps:

[0204] S901. Obtain the annotation result.

[0205] In the embodiment of the present disclosure, the text annotation device obtains the annotation result. Here, the annotation result refers to a set of alternative texts composed of the first text and the second text determined based on each set of sub-text sets.

[0206] S902. Post-processing operation.

[0207] In the embodiment of the present disclosure, the text annotation device eliminates unreasonable texts from the set of alternative texts. For example, it eliminates the results in the case where the text that best matches the target attribute and the text that least matches the target attribute cannot be determined, and eliminates the results where the text that best matches the target attribute and the text that least matches the target attribute are selected as the same item.

[0208] S903. Determine the annotation result subset 1.

[0209] In the embodiment of the present disclosure, the text annotation device divides the set of alternative texts into two sets of alternative sub-texts. For example, the alternative sub-text set 1 is the annotation result subset 1.

[0210] S904. Determine the annotation result subset 2.

[0211] In the embodiment of the present disclosure, the text annotation device divides the set of alternative texts into two sets of alternative sub-texts. For another example, the alternative sub-text set 2 is the annotation result subset 2.

[0212] S905. Maximize the difference metric model.

[0213] In the embodiment of the present disclosure, for each set of alternative sub-text sets, the text annotation device uses the maximize difference metric model to calculate the second matching degree of each text to be annotated in the text set to be annotated with the preset target attribute in each set of alternative sub-text sets.

[0214] S906. Result 1.

[0215] In the embodiment of the present disclosure, the text annotation device obtains the second matching degree of each text with the preset target attribute in the first set of alternative sub-texts, that is, Result 1.

[0216] S907. Result 2.

[0217] In the embodiment of the present disclosure, the text annotation device obtains the second matching degree of each text with the preset target attribute in the second set of alternative sub-texts, that is, Result 2.

[0218] S908. Spearman correlation coefficient and Pearson correlation coefficient.

[0219] In the embodiments of the present disclosure, the text annotation device determines the correlation of the sorting positions of two sets of spare sub-texts in the same text based on the second matching degree in two ways: the Spearman correlation coefficient and the Pearson correlation coefficient.

[0220] Figure 10 It is a schematic diagram of a text annotation method provided by an embodiment of this application. As Figure 10 shown, the following steps may be included:

[0221] S1001. Preprocess and deduplicate the annotation data.

[0222] In the embodiments of the present disclosure, the text annotation method can be applied to a text annotation device, and the annotation data refers to the text set to be annotated. The text annotation device preprocesses and deduplicates the annotation data. Among them, for each text to be annotated in the annotation data, it includes removing the preset string in the text. The preset string can be, for example, punctuation marks, symbols, special characters, links, etc., removing the repeated strings with a length greater than the preset length threshold in the text, and filtering the texts with a similarity greater than the preset similarity threshold in the annotation data.

[0223] S1002. Organize the annotation data.

[0224] In the embodiments of the present disclosure, the text annotation device organizes the processed data to be annotated to obtain multiple groups of first sample sub-text sets. Here, the organization refers to randomly sampling the data to be annotated as described above to obtain multiple groups of first sample sub-text sets, and determining multiple groups of sub-text sets based on the dispersion degree of the frequency of occurrence of the first combined text with the first preset combined length in the multiple groups of first sample sub-text sets.

[0225] S1003. Design the annotation instructions.

[0226] In the embodiments of the present disclosure, the text annotation device determines the annotation instructions for annotating the first text and the second text on the sub-text sets.

[0227] S1004. Select the large language model.

[0228] In the embodiments of the present disclosure, the text annotation device selects from a preset multiple large language models to determine the target large language model. Here, the multiple large language models refer to the multiple alternative large language models in the embodiments of the present disclosure.

[0229] S1005. Batch automated annotation.

[0230] In the embodiments of the present disclosure, the text annotation device performs batch automated annotation operations on multiple sub-text sets based on the target large language model.

[0231] S1006. Post-process and evaluate the annotation results.

[0232] In the embodiments of the present disclosure, the text annotation device performs post-processing on the annotation. Here, the annotation result refers to the first text and the second text determined based on each group of sub-text sets. Among them, the post-processing operation includes eliminating the result in the case where the text that best meets the target attribute and the text that least meets the target attribute cannot be determined, and eliminating the result where the text that best meets the target attribute and the text that least meets the target attribute are selected as the same item.

[0233] Next, continue to describe the exemplary structure of the software module implementation of the text annotation device 255 provided in the embodiments of the present application. In some embodiments, as Figure 2 shown, the software module in the text annotation device 255 stored in the memory 250 may include:

[0234] The first acquisition module 2551 is used to acquire the text set to be annotated and divide the text set into multiple groups of sub-text sets; where different groups of sub-text sets are different, and the texts to be annotated in each group of sub-text sets are different; the marking module 2552 is used to identify the attributes of each text in the sub-text set for each group of sub-text sets, mark the text with the highest matching degree with the preset target attribute in the sub-text set as the first text, and mark the text with the lowest matching degree with the preset target attribute as the second text; the first determination module 2553 is used to determine the first matching degree of the text with the preset target attribute for each text to be annotated in the text set to be annotated based on the first total number of times the text appears in the multiple groups of sub-text sets, the second total number of times the text is marked as the first text in different sub-text sets, and the third total number of times the text is marked as the second text in different sub-text sets, and perform matching degree annotation on the text.

[0235] In some possible implementation manners, the first acquisition module 2551 is used to randomly sample the text set to generate multiple groups of first sample sub-text sets; where different groups of first sample sub-text sets are different, and the texts to be annotated in each group of first sample sub-text sets are different; based on the multiple groups of first sample sub-text sets, determine the frequency of occurrence of the first combined text of each first preset combination length in the multiple groups of first sample sub-text sets; where the first combined text includes any combination of multiple texts in any group of first sample sub-text sets, and the number of texts indicated by the first preset combination length is less than the total number of texts included in the corresponding first sample sub-text set; based on the dispersion degree of the frequencies of each first combined text, determine the multiple groups of sub-text sets.

[0236] In some possible implementation manners, the degree of dispersion is characterized by the standard deviation; the first acquisition module 2551 is configured to, when the standard deviation is less than a preset standard deviation threshold, determine the multiple groups of first sample sub-text sets as the multiple groups of sub-text sets; when the standard deviation is greater than or equal to the preset standard deviation threshold, randomly sample again to generate multiple groups of first sample sub-text sets until the standard deviation of the frequencies of each first combined text determined based on the regenerated multiple groups of first sample sub-text sets is less than the preset standard deviation threshold, and determine the regenerated multiple groups of first sample sub-text sets as the multiple groups of sub-text sets.

[0237] In some possible implementation manners, the apparatus further includes: a second determination module, configured to determine a target large language model from multiple alternative large language models; the marking module 2552 is configured to, for each group of sub-text sets, based on the target large language model, identify the attributes of each text in the sub-text set, mark the text with the maximum matching degree with the preset target attribute in the sub-text set as the first text, and mark the text with the minimum matching degree with the preset target attribute as the second text.

[0238] In some possible implementation manners, the apparatus further includes: a second acquisition module, configured to acquire multiple groups of test sample text sets and the attribute labels of each test sample text set; the second determination module is configured to process the multiple groups of test sample text sets by using the multiple alternative large language models to obtain the predicted attributes of each test sample text set determined based on each alternative large language model; for each alternative large language model, determine the prediction ability of the alternative large language model based on the attribute labels of each group of test sample text sets and the predicted attributes of each group of test sample text sets determined based on the alternative large language model; and determine the target large language model based on the prediction ability of each alternative large language model.

[0239] In some possible implementation manners, the second acquisition module is configured to randomly sample a preset test text set to generate multiple groups of second sample sub-text sets; wherein, different groups of second sample sub-text sets are different, and the test texts in each group of second sample sub-text sets are different; based on the multiple groups of second sample sub-text sets, determine the frequency of occurrence of each second combined text with a second preset combined length in the multiple groups of second sample sub-text sets; wherein, the second combined text includes multiple test texts combined arbitrarily in any group of second sample sub-text sets, and the number of texts indicated by the second preset combined length is less than or equal to the total number of texts included in the corresponding second sample sub-text set; based on the degree of dispersion of the frequencies of each second combined text, determine multiple groups of intermediate sub-text sets; for each group of intermediate sub-text sets, repeatedly select third texts outside the test texts included in the intermediate sub-text set from the test text set, and insert the third texts into the intermediate sub-text set to obtain multiple groups of expanded text sets with the number of texts expanded; wherein, the number of test texts included in each group of expanded text sets is the same; the multiple groups of expanded text sets corresponding to the multiple groups of intermediate sub-text sets form the multiple groups of test sample text sets.

[0240] In some possible implementation manners, the second determination module is configured to, for each group of intermediate sub-text sets, use the alternative large language model to determine the matching degree between each test text in each group of expanded text sets associated with the intermediate sub-text set and the preset target attribute, and obtain the matching degree ranking of the test texts in each group of expanded text sets; for each group of intermediate sub-text sets, if the matching degree rankings of the multiple test texts in the intermediate sub-text set are the same in the multiple groups of expanded text sets, determine that the multiple test texts in the intermediate sub-text set belong to the first target text group for evaluating the prediction ability; based on the total number of the first target text groups determined by the multiple intermediate sub-text sets and the total number of the multiple intermediate sub-text sets, determine the prediction ability of the alternative large language model; wherein, the prediction ability is positively correlated with the total number of the first target text groups.

[0241] In some possible implementation manners, the apparatus further includes: a third determination module, configured to, for each first target text group, when the matching degree ranking of the multiple test texts in the first target text group in the multiple groups of expanded text sets is the same as the preset matching degree ranking of the multiple test texts, determine the first target text group as the second target text group; a fourth determination module, configured to, based on the total number of the second target text groups determined by the multiple first target text groups and the total number of the multiple intermediate sub-text sets, determine the prediction ability of the alternative large language model; wherein, the prediction ability is positively correlated with the total number of the second target text groups.

[0242] In some possible embodiments, the first determination module 2553 is configured to determine the ratio of the difference between the second total number of times and the third total number of times to the first total number of times as the first matching degree of the text with the preset target attribute.

[0243] In some possible embodiments, the apparatus further includes:

[0244] A grouping module, configured to form a backup text set by combining a first text and a second text determined based on each set of sub-text sets, and randomly divide the backup text set into multiple sets of backup sub-text sets; wherein the number of texts to be labeled in each set of backup sub-text sets is the same; a fifth determination module, configured to, for each set of backup sub-text sets, count the fourth total number of times that each text to be labeled in the text set to be labeled is marked as the first text in the backup sub-text set, and the fifth total number of times that the text is marked as the second text in the backup sub-text set, and determine the second matching degree of the text with the preset target attribute based on the first total number of times that the text appears in the multiple sets of sub-text sets, the fourth total number of times corresponding to the text, and the fifth total number of times; a sorting module, configured to, for each set of backup sub-text sets, perform a matching degree sorting based on the second matching degrees of the texts; a sixth determination module, configured to determine the correlation of the sorting positions of the same text in different backup sub-text sets based on the matching degree sorting results of the texts in each set of backup sub-text sets; the first determination module 2553 is configured to, when the correlation corresponding to the text is greater than or equal to a preset correlation threshold, determine the first matching degree of the text with the preset target attribute based on the first total number of times, the second total number of times, and the third total number of times corresponding to the text.

[0245] The embodiments of the present application provide a computer program product, which includes: a computer program or executable instructions, and the computer program or executable instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer program or executable instructions from the computer-readable storage medium, and the processor executes the computer program or executable instructions, so that the computer device executes the text annotation method described above in the embodiments of the present application.

[0246] The embodiments of the present application provide a computer-readable storage medium, on which a computer program or executable instructions are stored. When the computer program or executable instructions are executed by a processor, the processor will be caused to execute the text annotation method provided by the embodiments of the present application. For example, Figure 3 the text annotation method shown.

[0247] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0248] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0249] As an example, the executable instructions may or may not correspond to a file in a file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or, stored in multiple cooperating files (such as files that store one or more modules, subroutines, or portions of code).

[0250] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one site, or, on multiple computing devices distributed across multiple sites and interconnected by a communication network.

[0251] As described above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are included within the protection scope of the present application.

Claims

1. A text annotation method, characterized in that: include: Acquire a text set to be annotated, and divide the text set into a plurality of sub-text sets; wherein the sub-text sets of different groups are different, and the text to be annotated in each sub-text set is different; For each sub-text set, identifying the attributes of each text in the sub-text set, marking the text in the sub-text set that has the greatest matching degree with the preset target attribute as the first text, and marking the text that has the least matching degree with the preset target attribute as the second text; For each text to be annotated in the text set to be annotated, based on a first total number of times the text appears in the multiple groups of sub-text sets, a second total number of times the text is marked as the first text in different sub-text sets, and a third total number of times the text is marked as the second text in different sub-text sets, a first matching degree of the text with the preset target attribute is determined, and the text is annotated with the matching degree.

2. The method according to claim 1, characterized in that The step of dividing the text set into a plurality of sub-text sets comprises: Randomly sampling the text set to generate multiple groups of first sample sub-text sets; wherein different groups of first sample sub-text sets are different, and the text to be annotated in each group of first sample sub-text sets is different; Based on the multiple groups of first sample subtext sets, determining the frequency of occurrence of each first combination text of the first preset combination length in the multiple groups of first sample subtext sets; wherein the first combination text includes a plurality of texts in any combination in any group of first sample subtext sets, and the number of texts in the combination indicated by the first preset combination length is less than the total number of texts included in the corresponding first sample subtext set; The plurality of subtext sets are determined based on the discrete degrees of the frequencies of the first combined texts.

3. The method according to claim 2, characterized in that The degree of dispersion is characterized by a standard deviation, and the plurality of sub-text sets are determined based on the degree of dispersion of the frequencies of the first combined texts, including: In a case where the standard deviation is less than a preset standard deviation threshold, determining the plurality of groups of first sample subtext sets as the plurality of groups of subtext sets; When the standard deviation is greater than or equal to the preset standard deviation threshold, multiple groups of first sample sub-text sets are randomly re-sampled until the standard deviation of the frequency of each first combined text determined based on the regenerated multiple groups of first sample sub-text sets is less than the preset standard deviation threshold, and the regenerated multiple groups of first sample sub-text sets are determined as the multiple groups of sub-text sets.

4. The method according to claim 1, characterized in that: The method further comprises: Determine a target large language model from multiple candidate large language models; The method of identifying the attributes of each text in the subtext set for each group of subtext sets, marking the text in the subtext set with the greatest matching degree with the preset target attribute as the first text, and marking the text with the least matching degree with the preset target attribute as the second text, includes: For each group of sub-text sets, based on the target large language model, the attributes of each text in the sub-text set are identified, and the text in the sub-text set that matches the preset target attribute the most is marked as the first text, and the text that matches the preset target attribute the least is marked as the second text.

5. The method according to claim 4, characterized in that The method further comprises: Obtain multiple groups of test sample text sets and attribute labels of each test sample text set; The step of determining a target large language model from a plurality of candidate large language models comprises: Processing the plurality of test sample text sets using the plurality of candidate large language models to obtain a prediction attribute determined for each test sample text set based on each candidate large language model; For each candidate large language model, based on the attribute labels of each group of test sample text sets and the prediction attributes determined by each group of test sample text sets based on the candidate large language model, determining the prediction capability of the candidate large language model; Based on the prediction capability of each candidate large oracle model, the target large language model is determined.

6. The method according to claim 5, characterized in that The step of obtaining multiple test sample text sets includes: Randomly sampling a preset test text set to generate multiple groups of second sample sub-text sets; wherein different groups of second sample sub-text sets are different, and the test texts in each group of second sample sub-text sets are different; Based on the plurality of sets of second sample subtexts, determining the frequency of occurrence of each second combination text of the second preset combination length in the plurality of sets of second sample subtexts; wherein the second combination text comprises a plurality of test texts in any combination in any set of second sample subtexts, and the number of texts in the combination indicated by the second preset combination length is less than or equal to the total number of texts included in the corresponding second sample subtext set; Determining a plurality of intermediate subtext sets based on the discrete degrees of the frequencies of the second combined texts; For each group of intermediate sub-text sets, a third text other than the test text included in the intermediate sub-text set is selected from the test text set multiple times, and the third text is inserted into the intermediate sub-text set to obtain multiple groups of expanded text sets with expanded text quantity; wherein the quantity of test text included in each group of expanded text sets is the same; and the multiple groups of expanded text sets corresponding to the multiple groups of intermediate sub-text sets form the multiple groups of test sample text sets.

7. The method according to claim 6, characterized in that The step of determining the prediction capability of the candidate large language model based on the attribute labels of each group of test sample text sets and the prediction attributes determined by each group of test sample text sets based on the candidate large language model comprises: For each group of intermediate sub-text sets, the candidate large language model is used to determine the matching degree between each test text in each group of extended text sets associated with the intermediate sub-text set and the preset target attribute, and the matching degree ranking of the test texts in each group of extended text sets is obtained; For each intermediate sub-text set, if the matching rankings of the multiple test texts in the intermediate sub-text set in the multiple expanded text sets are consistent, it is determined that the multiple test texts in the intermediate sub-text set belong to the first target text set for evaluating the prediction ability; Based on the total number of the first target text group determined by the multiple intermediate sub-text sets and the total number of the multiple intermediate sub-text sets, the prediction ability of the candidate large language model is determined; wherein the prediction ability is positively correlated with the total number of the first target text group.

8. The method according to claim 7, characterized in that The method further comprises: For each first target text group, if the matching degree ranking of the plurality of test texts in the first target text group in the plurality of expanded text sets is consistent with the preset matching degree ranking of the plurality of test texts, determining the first target text group as the second target text group; Based on the total number of second target text groups determined by the plurality of first target text groups and the total number of the plurality of intermediate subtext sets, the prediction capability of the candidate large language model is determined; wherein the prediction capability is positively correlated with the total number of the second target text groups.

9. The method according to claim 1, characterized in that: The determining a first degree of match between the text and the preset target attribute based on a first total number of times the text appears in the plurality of subtext sets, a second total number of times the text is marked as the first text in different subtext sets, and a third total number of times the text is marked as the second text in different subtext sets includes: The ratio of the difference between the second total number of times and the third total number of times to the first total number of times is determined as a first matching degree between the text and the preset target attribute.

10. The method according to claim 1, characterized in that The method further comprises: The first text and the second text determined based on each sub-text set form a standby text set, and the standby text set is randomly divided into a plurality of standby sub-text sets; wherein the number of texts to be annotated in each standby sub-text set is the same; For each group of spare sub-text sets, counting the fourth total number of times each text to be annotated in the text set to be annotated is marked as the first text in the spare sub-text set, and the fifth total number of times the text is marked as the second text in the spare sub-text set, and determining the second matching degree between the text and the preset target attribute based on the first total number of times the text appears in the multiple groups of sub-text sets, the fourth total number of times the text corresponds to, and the fifth total number of times; For each set of backup sub-text sets, sorting the matching degrees based on the second matching degrees of each text; Based on the matching ranking results of the texts in each group of backup sub-text sets, determining the relevance of the ranking positions of the same text in different backup sub-text sets; The determining a first degree of match between the text and the preset target attribute based on a first total number of times the text appears in the plurality of subtext sets, a second total number of times the text is marked as the first text in different subtext sets, and a third total number of times the text is marked as the second text in different subtext sets includes: When the relevance corresponding to the text is greater than or equal to a preset relevance threshold, a first matching degree between the text and the preset target attribute is determined based on a first total number of times, a second total number of times, and a third total number of times corresponding to the text.

11. A text annotation device, characterized in that: The device comprises: An acquisition module is used to acquire a text set to be annotated, and divide the text set into a plurality of sub-text sets; wherein different sub-text sets are different, and the text to be annotated in each sub-text set is different; A marking module, for identifying the attributes of each text in each sub-text set, marking the text in the sub-text set with the greatest matching degree with a preset target attribute as a first text, and marking the text with the least matching degree with the preset target attribute as a second text; A determination module is used to determine, for each text to be annotated in the text set to be annotated, a first matching degree between the text and the preset target attribute based on a first total number of times the text appears in the multiple sub-text sets, a second total number of times the text is marked as the first text in different sub-text sets, and a third total number of times the text is marked as the second text in different sub-text sets, and to annotate the text with the matching degree.