Label Optimization Method, Device, Electronic Device and Readable Medium

By annotating and dividing the tag collection and generating optimized tags, the resource consumption problem caused by the increase in the number of tags is solved, and the accuracy of the tags and the operation efficiency of the data platform are improved.

CN115048999BActive Publication Date: 2025-05-27PING AN BANK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210664326.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2025-05-27
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

In the prior art, the number of tags increases with the use of the data platform, resulting in the need of additional resources for classification and sorting, which reduces the operation efficiency of the data platform.

Method used

By obtaining the set of tags to be optimized, annotating them to obtain scene information, dividing the tag subset, and selecting the target tag subset from it to generate optimization tags, realizing automatic classification, sorting and merging.

Benefits of technology

By automatically classifying and organizing labels, the accuracy of labels is improved, resource consumption is reduced, and the operation efficiency of the data platform is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115048999B_ABST
    Figure CN115048999B_ABST
Patent Text Reader

Abstract

The present application provides a label optimization method, apparatus, electronic device, and readable medium. The method includes: obtaining a set of labels to be optimized, where the set of labels to be optimized includes N labels to be optimized; annotating each label to be optimized in the set of labels to be optimized to obtain the scenario information corresponding to each label to be optimized, where the scenario information includes a scenario type label and a scenario relevance score; dividing the set of labels to be optimized into M label subsets according to the scenario information corresponding to each label to be optimized; selecting a target label subset from the M label subsets, where the target label subset includes at least two labels to be optimized; and generating an optimized label according to at least two labels to be optimized in the target label subset. This method can automatically classify, organize, and merge a large number of accumulated labels to generate optimized labels, thereby improving the accuracy of the labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method, device, electronic device, and readable medium for label optimization. Background Art

[0002] Risk labels have always been an important function in the risk field. Risk labels are the essence extracted from the experience of risk strategy personnel. Based on labels for risk judgment can greatly reduce the resource consumption of the online system, and at the same time can be simply and conveniently applied to more scenarios.

[0003] Currently, the labels applied in each scenario are usually set and marked manually, and then provided to each platform for further processing and application.

[0004] However, the process of manually defining labels is usually independent. With the long-term use of the data platform, the number of labels in the platform also increases accordingly. Additional resources need to be invested in classification and sorting for subsequent use, which reduces the overall operation efficiency of the data platform. Summary of the Invention

[0005] Based on the above technical problems, this application provides a method, device, electronic device, and readable medium for label optimization to automatically classify, sort, and merge a large number of accumulated labels to generate optimized labels, thereby improving the accuracy of labels.

[0006] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.

[0007] According to one aspect of the embodiments of this application, a method for label optimization is provided, including:

[0008] Obtain a set of labels to be optimized, where the set of labels to be optimized includes N labels to be optimized, and N is an integer greater than 1;

[0009] Annotate each label to be optimized in the set of labels to be optimized to obtain the scenario information corresponding to each label to be optimized, where the scenario information includes a scenario type label and a scenario correlation score, and the scenario correlation score is used to indicate the strength of the correlation between the label to be optimized and the scenario type label;

[0010] According to the scenario information corresponding to each label to be optimized, divide the set of labels to be optimized into M label subsets, where M is an integer greater than or equal to 1 and less than or equal to N;

[0011] Select a target label subset from the M label subsets, where the target label subset includes at least two labels to be optimized;

[0012] Generate an optimized tag based on at least two tags to be optimized in the target tag subset.

[0013] In some embodiments of the present application, based on the above technical solutions, each tag to be optimized in the tag set to be optimized is labeled to obtain the scenario information corresponding to each tag to be optimized, including:

[0014] Identify the tag to be optimized according to the tag labeling rule to obtain a scenario type tag;

[0015] Calculate the relevance score of the tag to be optimized through a scenario relevance model to obtain a scenario relevance score;

[0016] Generate the scenario information corresponding to the tag to be optimized according to the scenario type tag and the scenario relevance score.

[0017] In some embodiments of the present application, based on the above technical solutions, the tag set to be optimized is divided into M tag subsets according to the scenario information corresponding to each tag to be optimized, including:

[0018] Group the tags to be optimized according to the scenario type tags in the scenario information to obtain M tag groups;

[0019] Filter the tags to be optimized in each tag group among the M tag groups according to the scenario relevance score in the scenario information to obtain the M tag subsets.

[0020] In some embodiments of the present application, based on the above technical solutions, the tags to be optimized in each tag group among the M tag groups are filtered according to the scenario relevance score in the scenario information to obtain the M tag subsets, including:

[0021] For the tags to be optimized in each tag group, if the scenario relevance score in the corresponding scenario information is lower than the relevance score threshold, select the tag to be optimized from the tag group.

[0022] In some embodiments of the present application, based on the above technical solutions, selecting a target tag subset from the M tag subsets includes:

[0023] Calculate the average value of the scenario relevance scores according to the scenario relevance scores in each of the M tag subsets;

[0024] Determine the tag subset corresponding to the largest average value of the scenario relevance scores as the target tag subset.

[0025] In some embodiments of the present application, based on the above technical solutions, generating an optimized tag according to at least two tags to be optimized in the target tag subset includes:

[0026] Determining the semantic similarity between at least two tags to be optimized in the target tag subset through a semantic parsing model;

[0027] If the semantic similarity is greater than the similarity threshold, determining the tag to be optimized corresponding to the maximum scene relevance score among the at least two tags to be optimized as the optimized tag corresponding to the target tag subset;

[0028] If the semantic similarity is less than or equal to the similarity threshold, concatenating the at least two tags to be optimized to generate the optimized tag corresponding to the target tag subset.

[0029] According to one aspect of the embodiments of the present application, there is provided a tag optimization device, including:

[0030] A set acquisition module, configured to acquire a set of tags to be optimized, where the set of tags to be optimized includes N tags to be optimized, and N is an integer greater than 1;

[0031] A labeling module, configured to label each tag to be optimized in the set of tags to be optimized to obtain the scene information corresponding to each tag to be optimized, where the scene information includes a scene type tag and a scene relevance score, and the scene relevance score is used to indicate the strength of the correlation between the tag to be optimized and the scene type tag;

[0032] A set partitioning module, configured to partition the set of tags to be optimized into M tag subsets according to the scene information corresponding to each tag to be optimized, where M is an integer greater than or equal to 1 and less than or equal to N;

[0033] A subset selection module, configured to select a target tag subset from the M tag subsets, where the target tag subset includes at least two tags to be optimized;

[0034] A tag generation module, configured to generate an optimized tag according to at least two tags to be optimized in the target tag subset.

[0035] In some embodiments of the present application, based on the above technical solutions, the labeling module includes:

[0036] A tag marking unit, configured to mark the tag to be optimized according to a tag marking rule to obtain a scene type tag;

[0037] A relevance calculation unit for calculating the relevance of the to-be-optimized tag through a scenario relevance model to obtain a scenario relevance score;

[0038] A scenario information generation unit for generating scenario information corresponding to the to-be-optimized tag according to the scenario type tag and the scenario relevance score.

[0039] According to one aspect of the embodiments of the present application, an electronic device is provided, which includes: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the tag optimization method in the above technical solution by executing the executable instructions.

[0040] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the tag optimization method in the above technical solution is implemented.

[0041] In the embodiments of the present application, the tags are divided into multiple tag subsets according to the scenario information obtained by annotating the existing tags, and then optimized tags are generated from at least two to-be-optimized tags in the tag subsets. Through the above method, a large number of accumulated tags can be automatically classified, sorted, and merged to generate optimized tags, thereby improving the accuracy of the tags.

[0042] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:

[0044] Figure 1 Schematically shows an exemplary system architecture diagram of the technical solution of the present application in an application scenario;

[0045] Figure 2 Shows a schematic diagram of the process of generating optimized tags in the embodiments of the present application;

[0046] Figure 3 Is a flowchart of a tag optimization method provided by the embodiments of the present application;

[0047] Figure 4 Schematically shows a block diagram of the composition of the tag optimization device in the embodiments of the present application;

[0048] Figure 5 The figure shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners

[0049] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0050] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0051] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0052] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0053] It should be understood that the solution of the present application can be applied to the field of artificial intelligence, and specifically applied to the field of risk control based on artificial intelligence. Specifically, in this field, various risk labels will be marked for the risk items to be evaluated, so as to provide auxiliary information for risk assessment. The solution of the present application can label and classify the existing labels according to information such as scenarios and results, and then optimize the combination of the labels according to the classification results, so as to obtain more accurate label content.

[0054] Figure 1 Schematically shows an exemplary system architecture diagram of the technical solution of the present application in an application scenario. As Figure 1As shown in the figure, this application scenario includes an application server 110 and a label server 120, and there is a network connection between the two. The application server 110 is used to actually use the labels and send information such as the application scenarios and application results of the labels to the label server 120 for label analysis and optimization. It can be understood that the above servers can be independent physical servers, or server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. A software application for actual risk assessment based on labels is installed on the application server 110. A machine learning model for label analysis and optimization is installed on the label server 120. The application server 110 uses the labels and feeds back the application results to the label server 120. The label server 120 then generates new optimized labels based on the existing labels combined with the application results.

[0055] Figure 2 The figure shows a schematic diagram of the optimized label generation process in an embodiment of the present application. As Figure 2 shown, the application system will use the existing Labels 1 to 4. The recommendation algorithm of the machine model will analyze and calculate these labels, and select Labels 1 and 3 to be optimized in combination with the application conclusions fed back by the application system, and generate a new optimized label 5 based on the labels to be optimized.

[0056] The following will make a detailed description of the technical solution provided by the present application in combination with specific embodiments. For the convenience of introduction, please refer to Figure 3 , Figure 3 is a flowchart of a label optimization method provided by an embodiment of the present application. This method can be applied to the above label server for execution. The label server can be regarded as a computer device. In the embodiment of the present application, the computer device is used as the execution subject to introduce the label optimization method. The label optimization method may include the following steps S310 to S350:

[0057] Step S310, obtain a set of labels to be optimized, where the set of labels to be optimized includes N labels to be optimized, and N is an integer greater than 1.

[0058] This solution can be executed by the data platform itself or by a dedicated server. The set of tags to be optimized, i.e., the tags that need to be optimized, can usually be directly obtained from the data platform. The tags covered in the set of tags to be optimized can be specified directly, for example, by specifying the scenario and domain to which the tag belongs, the time when the tag was generated, the entity that defined the tag, etc. The set of tags to be optimized can also be sourced from data synchronization of an external system. Specifically, the server can send a request to the data platform to obtain the tags to be optimized, and the request includes the conditions that the tags to be optimized need to meet, such as the scenario and domain they belong to, length, generation time, etc. The data platform then filters out the tags that meet the conditions according to the request for obtaining the tags to be optimized and sends them to the server, thereby obtaining the set of tags to be optimized.

[0059] Step S320: Annotate each tag to be optimized in the set of tags to be optimized to obtain the scenario information corresponding to each tag to be optimized, where the scenario information includes a scenario type tag and a scenario relevance score, and the scenario relevance score is used to indicate the strength of the relevance between the tag to be optimized and the scenario type tag.

[0060] Annotate each tag to be optimized to obtain the scenario information corresponding to each tag to be optimized. Specifically, the annotation process includes generating a scenario type tag and a scenario relevance score. The scenario type tag is used to represent the scenario to which the tag to be optimized applies. This process can match the scenario corresponding to the tag to be optimized in the data platform with the scenario type tag according to a pre-agreed scenario mapping table, thereby determining the scenario to which the tag to be optimized should belong, so as to integrate the tags to be optimized in similar scenarios. The scenario relevance score indicates the strength of the relevance between the tag to be optimized and the scenario type tag. This process can be carried out by a trained machine learning model or by preset rules to determine the score.

[0061] In an embodiment of the present application, the process of annotating each tag to be optimized in the set of tags to be optimized to obtain the scenario information corresponding to each tag to be optimized may include the following steps: identify the tag to be optimized according to the tag annotation rules to obtain a scenario type tag; perform a relevance calculation on the tag to be optimized through a scenario relevance model to obtain a scenario relevance score; generate the scenario information corresponding to the tag to be optimized according to the scenario type tag and the scenario relevance score. The tag annotation rules usually depend on the specific field and scenario applied, such as the data meeting specific conditions, having specific attributes, and reaching a certain trigger condition, etc. For example, an amount or an average value being lower than a certain specific threshold can be used as a tag annotation rule.

[0062] Step S330: Divide the set of tags to be optimized into M tag subsets according to the scenario information corresponding to each tag to be optimized, where M is an integer greater than or equal to 1 and less than or equal to N.

[0063] The tags to be optimized are divided into different tag subsets according to the content in the scenario information, thus obtaining M tag subsets. The number of M usually equals or is greater than the number of scenario type tags covered in the entire set of tags to be optimized. The scenario correlation score is used to filter out tags with insufficiently high correlation from the tag subsets. Tags can also be clustered according to the scenario correlation score. In subsequent processing, the processing methods for tags in the same cluster can be different from those in different clusters.

[0064] In an embodiment of the present application, the process of dividing the set of tags to be optimized into M tag subsets according to the scenario information corresponding to each tag to be optimized may include the following steps: Group the tags to be optimized according to the scenario type tags in the scenario information to obtain M tag groups; Filter the tags to be optimized in each of the M tag groups according to the scenario correlation score in the scenario information to obtain the M tag subsets.

[0065] In an embodiment of the present application, the process of filtering the tags to be optimized in each of the M tag groups according to the scenario correlation score in the scenario information to obtain the M tag subsets may be: For the tags to be optimized in each tag group, if the scenario correlation score in the corresponding scenario information is lower than the correlation score threshold, then select the tag to be optimized from the tag group.

[0066] Step S340: Select a target tag subset from the M tag subsets, where the target tag subset includes at least two tags to be optimized;

[0067] The target label subset identifies the set of labels that need to be optimized, which can be determined according to the scope targeted by label optimization. For example, if optimization is targeted at a specific application scenario, the target label subset should be the set of labels for that application scenario. If integration optimization is performed on all labels in the data platform, any label subset can be randomly selected from the M label subsets as the target label subset. Since a single label does not meet the premise of label optimization, the target label subset includes at least two labels to be optimized. Selecting the target label subset can also be determined according to the characteristics of the labels to be optimized in the label subset. For example, the label subset with the largest sum or average of scene relevance scores is used as the target label subset, or the label subset with the largest number of labels to be optimized is used as the target label subset, etc.

[0068] In an embodiment of the present application, the process of selecting the target label subset from the M label subsets may include the following steps:

[0069] Calculate the mean of the scene relevance scores according to the scene relevance scores in each of the M label subsets;

[0070] Determine the label subset corresponding to the largest mean of the scene relevance scores as the target label subset.

[0071] Step S350, generate optimized labels according to at least two labels to be optimized in the target label subset.

[0072] The optimization method can be calculated according to a preset optimization recommendation algorithm, or can be determined by performing semantic analysis through a machine learning model. Specifically, several labels to be optimized with relatively high scene relevance scores can be directly used as candidates. Subsequently, the several labels to be optimized are arranged and combined to obtain multiple groups of combined labels. The scene relevance model is used to calculate the relevance scores of the combined labels, and the combined label with the maximum score is used as the optimized label. Or, the semantic relevance between each of the labels to be optimized in the target label subset can be compared and analyzed, so as to replace and integrate the labels to be optimized with similar semantics, and combine the labels with relatively large semantic differences to generate optimized labels.

[0073] In one embodiment of the present application, the step of generating an optimized tag based on at least two tags to be optimized in the target tag subset may include the following steps: determining the semantic similarity between at least two tags to be optimized in the target tag subset through a semantic parsing model; if the semantic similarity is greater than a similarity threshold, determining the tag to be optimized corresponding to the maximum scene correlation score among the at least two tags to be optimized as the optimized tag corresponding to the target tag subset; if the semantic similarity is less than or equal to the similarity threshold, concatenating the at least two tags to be optimized to generate the optimized tag corresponding to the target tag subset.

[0074] In an embodiment of the present application, the tags are divided into multiple tag subsets according to the scene information obtained by annotating the existing tags, and then optimized tags are generated from at least two tags to be optimized in the tag subsets. Through the above method, a large number of accumulated tags can be automatically classified, sorted, and merged to generate optimized tags, thereby improving the accuracy of the tags.

[0075] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0076] The following introduces the apparatus embodiment of the present application, which can be used to execute the tag optimization method in the above embodiments of the present application. Figure 4 The block diagram of the tag optimization apparatus according to an embodiment of the present application is schematically shown. As Figure 4 shown, the tag optimization apparatus 400 mainly may include:

[0077] A set acquisition module 410, configured to acquire a set of tags to be optimized, where the set of tags to be optimized includes N tags to be optimized, and N is an integer greater than 1;

[0078] An annotation module 420, configured to annotate each tag to be optimized in the set of tags to be optimized to obtain the scene information corresponding to each tag to be optimized, where the scene information includes a scene type tag and a scene correlation score, and the scene correlation score is used to indicate the strength of the correlation between the tag to be optimized and the scene type tag;

[0079] A set partitioning module 430, configured to partition the set of tags to be optimized into M tag subsets according to the scene information corresponding to each tag to be optimized, where M is an integer greater than or equal to 1 and less than or equal to N;

[0080] A subset selection module 440, configured to select a target label subset from the M label subsets, where the target label subset includes at least two labels to be optimized;

[0081] A label generation module 450, configured to generate an optimized label according to at least two labels to be optimized in the target label subset.

[0082] In an embodiment of the present application, based on the above solution, the annotation module 420 includes:

[0083] A label identification unit, configured to identify the label to be optimized according to a label annotation rule to obtain a scenario type label;

[0084] A relevance calculation unit, configured to calculate the relevance of the label to be optimized through a scenario relevance model to obtain a scenario relevance score;

[0085] A scenario information generation unit, configured to generate scenario information corresponding to the label to be optimized according to the scenario type label and the scenario relevance score.

[0086] In an embodiment of the present application, based on the above solution, the set partitioning module 430 includes:

[0087] A grouping unit, configured to group the labels to be optimized according to the scenario type labels in the scenario information to obtain M label groups;

[0088] A filtering unit, configured to filter the labels to be optimized in each label group among the M label groups according to the scenario relevance scores in the scenario information to obtain the M label subsets.

[0089] In an embodiment of the present application, based on the above solution, the filtering unit includes:

[0090] A label selection subunit, configured to, for the labels to be optimized in each label group, if the scenario relevance score in the corresponding scenario information is lower than a relevance score threshold, select the label to be optimized from the label group.

[0091] In an embodiment of the present application, based on the above solution, the subset selection module 440 includes:

[0092] An average value calculation unit, configured to calculate an average value of scenario relevance scores according to the scenario relevance scores in each label subset among the M label subsets;

[0093] A target determination unit, configured to determine the label subset corresponding to the maximum average value of scenario relevance scores as the target label subset.

[0094] In an embodiment of the present application, based on the above solution, the tag generation module 450 includes:

[0095] A similarity determination unit, configured to determine the semantic similarity between at least two tags to be optimized in the target tag subset through a semantic parsing model;

[0096] If the semantic similarity is greater than a similarity threshold, determine the tag to be optimized corresponding to the maximum scenario relevance score among the at least two tags to be optimized as the optimized tag corresponding to the target tag subset;

[0097] A tag splicing unit, configured to splice the at least two tags to be optimized if the semantic similarity is less than or equal to the similarity threshold to generate the optimized tag corresponding to the target tag subset.

[0098] It should be noted that the device provided in the above embodiment and the method provided in the above embodiment belong to the same concept. The specific manners in which each module performs operations have been described in detail in the method embodiment and will not be elaborated here.

[0099] Figure 5 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.

[0100] It should be noted that Figure 5 The computer system 500 of the electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0101] As Figure 5 shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for system operation are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.

[0102] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 510 as needed so that a computer program read therefrom is installed into the storage section 508 as needed.

[0103] Specifically, according to an embodiment of the present application, the processes described in each of the method flowcharts can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by a central processing unit (CPU) 501, various functions defined in the system of the present application are executed.

[0104] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0106] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0107] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0108] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.

[0109] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for label optimization, characterized in that, comprising: Obtain a set of labels to be optimized, where the set of labels to be optimized includes N labels to be optimized, and N is an integer greater than 1; Annotate each label to be optimized in the set of labels to be optimized to obtain the scenario information corresponding to each label to be optimized, where the scenario information includes a scenario type label and a scenario relevance score, and the scenario relevance score is used to indicate the strength of the relevance between the label to be optimized and the scenario type label; Group the labels to be optimized according to the scenario type labels in the scenario information to obtain M label groups; Filter the labels to be optimized in each label group among the M label groups according to the scenario relevance scores in the scenario information to obtain the M label subsets, where M is an integer greater than or equal to 1 and less than or equal to N; Select a target label subset from the M label subsets, where the target label subset includes at least two labels to be optimized; Determine the semantic similarity between at least two labels to be optimized in the target label subset through a semantic parsing model; If the semantic similarity is greater than the similarity threshold, determine the label to be optimized corresponding to the maximum scenario relevance score among the at least two labels to be optimized as the optimized label corresponding to the target label subset; If the semantic similarity is less than or equal to the similarity threshold, splice the at least two labels to be optimized to generate the optimized label corresponding to the target label subset.

2. The method according to claim 1, characterized in that, The annotating each label to be optimized in the set of labels to be optimized to obtain the scenario information corresponding to each label to be optimized includes: Identifying the label to be optimized according to the label annotation rule to obtain a scenario type label; Calculating the relevance of the label to be optimized through a scenario relevance model to obtain a scenario relevance score; Generating the scenario information corresponding to the label to be optimized according to the scenario type label and the scenario relevance score.

3. The method according to claim 1, characterized in that, The filtering the labels to be optimized in each label group among the M label groups according to the scenario relevance scores in the scenario information to obtain the M label subsets includes: For the labels to be optimized in each label group, if the scenario relevance score in the corresponding scenario information is lower than the relevance score threshold, select the label to be optimized from the label group.

4. The method according to claim 1, characterized in that, The selecting a target label subset from the M label subsets includes: Calculating the average value of the scenario relevance scores according to the scenario relevance scores in each label subset among the M label subsets; Determining the label subset corresponding to the maximum average value of the scenario relevance scores as the target label subset.

5. A label optimization device, characterized in that, comprising: A set acquisition module, configured to acquire a set of tags to be optimized, where the set of tags to be optimized includes N tags to be optimized, and N is an integer greater than 1; A labeling module, configured to label each tag to be optimized in the set of tags to be optimized, so as to obtain the scenario information corresponding to each tag to be optimized, where the scenario information includes a scenario type tag and a scenario relevance score, and the scenario relevance score is used to indicate the strength of the relevance between the tag to be optimized and the scenario type tag; A set partitioning module, configured to group the tags to be optimized according to the scenario type tags in the scenario information to obtain M tag groups, and filter the tags to be optimized in each of the M tag groups according to the scenario relevance scores in the scenario information to obtain the M tag subsets, where M is an integer greater than or equal to 1 and less than or equal to N; A subset selection module, configured to select a target tag subset from the M tag subsets, where the target tag subset includes at least two tags to be optimized; A tag generation module, configured to determine the semantic similarity between at least two tags to be optimized in the target tag subset through a semantic parsing model. If the semantic similarity is greater than a similarity threshold, the tag to be optimized corresponding to the maximum scenario relevance score among the at least two tags to be optimized is determined as the optimized tag corresponding to the target tag subset. If the semantic similarity is less than or equal to the similarity threshold, the at least two tags to be optimized are concatenated to generate the optimized tag corresponding to the target tag subset.

6. The apparatus according to claim 5, wherein, the labeling module includes: a tag marking unit, configured to mark the tag to be optimized according to a tag marking rule to obtain a scenario type tag; a relevance calculation unit, configured to calculate the relevance of the tag to be optimized through a scenario relevance model to obtain a scenario relevance score; a scenario information generation unit, configured to generate the scenario information corresponding to the tag to be optimized according to the scenario type tag and the scenario relevance score.

7. An electronic device, wherein, it includes: a processor; a memory, configured to store executable instructions of the processor; wherein, the processor is configured to execute the tag optimization method according to any one of claims 1 to 4 by executing the executable instructions.

8. A computer-readable medium, on which a computer program is stored, wherein, the computer program, when executed by a processor, implements the tag optimization method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Intelligent label association method and device, computer equipment and storage medium

    CN111752958A

  • Content tag generation method and device, electronic equipment and storage medium

    CN114021577A