Business scene determination method and device, medium and equipment

By determining the entity words and attribute words in the text information in the automatic audit method, generating tag representation data, and combining text representation data for feature extraction, the problem of low accuracy in inference in the existing technology is solved, and higher business scenario classification accuracy and audit efficiency are achieved.

CN120218222APending Publication Date: 2025-06-27TENPAY PAID TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311759152.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The automatic audit method based on artificial intelligence in the prior art has low accuracy when inferring business scenarios, and it is easy to hit multiple business scenarios, affecting the accuracy of audits.

Method used

By obtaining the pending text information, determining entity words and attribute words, generating tag representation data, combining text representation data for feature extraction, and then performing scene classification to improve the accuracy of business scenarios.

Benefits of technology

It improves the accuracy of business scenario classification, enhances the feature representation of entity words and attribute words, and improves the efficiency of subsequent audit tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218222A_ABST
    Figure CN120218222A_ABST
Patent Text Reader

Abstract

The invention discloses a business scene determination method and device, a medium and equipment, and relates to the field of artificial intelligence, and the method comprises the steps: determining at least one entity word and at least one attribute word in to-be-processed text information; performing word segmentation processing on the to-be-processed text information to obtain a text character sequence; performing embedding characterization processing on the text character sequence to obtain text characterization data; according to the text character sequence, the at least one entity word and the at least one attribute word, generating mark representation data; performing feature extraction processing on the text representation data and the mark representation data to obtain text semantic feature data corresponding to the to-be-processed text information; and performing scene classification processing according to the text semantic feature data to obtain a target business scene category corresponding to the to-be-processed text information. According to the method, the representation of the entity features and the attribute features in the text semantic feature data is enhanced, so that the accuracy of deducing the business scene based on the text semantic feature data can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of artificial intelligence, and specifically relates to a method, apparatus, medium and device for determining a business scenario. Background Art

[0002] Artificial Intelligence (AI) is a comprehensive technology in computer science. By studying the design principles and implementation methods of various intelligent machines, machines are enabled to have functions of perception, reasoning and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, such as several major directions including natural language processing, machine learning, and deep learning. With the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0003] In some Internet applications, users need to upload specified information or file materials of a specified type to prove their identity or the compliance and legality of the interaction. In related technologies, an automatic auditing method based on artificial intelligence simply infers a business scenario through the matching of scenario keywords, and then uses the inferred business scenario to assist in auditing. However, the accuracy of the inferred business scenario is low, and the situation of hitting multiple business scenarios easily occurs, which instead affects the accuracy of the auditing. Summary of the Invention

[0004] In order to improve the accuracy of business scenario classification, the present application provides a method, apparatus, medium and device for determining a business scenario. The technical solutions are as follows:

[0005] In a first aspect, the present application provides a method for determining a business scenario, the method including:

[0006] Obtain text information to be processed;

[0007] Determine at least one entity word in the text information to be processed and at least one attribute word in the text information to be processed; the at least one attribute word is index data of a business dimension;

[0008] Perform word segmentation processing on the text information to be processed to obtain a text character sequence;

[0009] Perform embedding representation processing on the text character sequence to obtain text representation data;

[0010] Generate marker representation data according to the text character sequence, the at least one entity word and the at least one attribute word, where the marker representation data is used to indicate a first character corresponding to each entity word and a second character corresponding to each attribute word in the text character sequence;

[0011] Perform feature extraction processing on the text representation data and the tag representation data to obtain text semantic feature data corresponding to the text information to be processed;

[0012] Perform scenario classification processing according to the text semantic feature data to obtain the target business scenario category corresponding to the text information to be processed.

[0013] Optionally, the method further includes:

[0014] Perform feature extraction processing on the text representation data and the tag representation data to obtain text semantic feature data corresponding to the text information to be processed, entity feature data corresponding to the at least one entity word, and attribute feature data corresponding to the at least one attribute word;

[0015] The performing scenario classification processing according to the text semantic feature data to obtain the target business scenario category corresponding to the text information to be processed includes:

[0016] Perform scenario classification processing on the text semantic feature data, the entity feature data, and the attribute feature data to obtain the target business scenario category corresponding to the text information to be processed.

[0017] Optionally, the performing scenario classification processing on the text semantic feature data, the entity feature data, and the attribute feature data to obtain the target business scenario category corresponding to the text information to be processed includes:

[0018] Perform average pooling processing on the entity feature data to obtain target entity feature data;

[0019] Perform average pooling processing on the attribute feature data to obtain target attribute feature data;

[0020] Concatenate the text semantic feature data, the target entity feature data, and the target attribute feature data to obtain concatenated feature data;

[0021] Input the concatenated feature data into a classification network to perform category prediction processing of the business scenario to obtain the target business scenario category.

[0022] Optionally, the tag representation data includes role representation data, and the generating tag representation data according to the text character sequence, the at least one entity word, and the at least one attribute word includes:

[0023] Traverse each text character in the text character sequence;

[0024] When the currently traversed text character is the first character, determine that the role representation data corresponding to the currently traversed text character is the preset first role representation data;

[0025] When the currently traversed text character is the second character, determine that the role representation data corresponding to the currently traversed text character is the preset second role representation data;

[0026] When the currently traversed text character is neither the first character nor the second character, determine that the role representation data corresponding to the currently traversed text character is the preset third role representation data;

[0027] Generate the role representation data according to the role representation data corresponding to each text character.

[0028] Optionally, the marker representation data includes distance representation data, and generating the marker representation data according to the text character sequence, the at least one entity word, and the at least one attribute word includes:

[0029] Traverse each text character in the text character sequence;

[0030] When the currently traversed text character is the first character, determine that the distance representation data corresponding to the currently traversed text character is the preset first distance representation data;

[0031] When the currently traversed text character is the second character, determine at least one text distance data, and each text distance data in the at least one text distance data represents the text similarity between the first character corresponding to each entity word and the currently traversed text character;

[0032] Based on the at least one text distance data, determine the target distance representation data corresponding to the target text distance data, and use the target distance representation data as the distance representation data corresponding to the currently traversed text character; the target text distance data is the minimum value in the at least one text distance data;

[0033] When the currently traversed text character is neither the first character nor the second character, determine that the distance representation data corresponding to the currently traversed text character is the preset second distance representation data;

[0034] Generate the distance representation data according to the distance representation data corresponding to each text character.

[0035] Optionally, performing feature extraction processing on the text representation data and the marker representation data to obtain the text semantic feature data corresponding to the text information to be processed includes:

[0036] Concatenate the text representation data and the marker representation data to obtain target representation data;

[0037] Input the target representation data into a feature transformation model for feature extraction processing to obtain target feature data;

[0038] Based on a preset semantic feature identifier, determine the text semantic feature data in the target feature data.

[0039] Optionally, the determining at least one entity word and at least one attribute word in the text information to be processed includes:

[0040] Obtain the regular library corresponding to each entity in a variety of pre-constructed entities;

[0041] Perform regular matching according to the regular library corresponding to each entity and the text information to be processed to obtain at least one original entity word;

[0042] Perform entity completion processing or entity conversion processing on the at least one original entity word to obtain the at least one entity word;

[0043] Obtain a business attribute keyword dictionary constructed in advance;

[0044] Perform mapping matching processing according to the business attribute keyword dictionary and the text information to be processed to obtain the at least one attribute word.

[0045] Optionally, the obtaining the text information to be processed includes:

[0046] Obtain at least one picture to be reviewed;

[0047] Perform text recognition on the at least one picture to be reviewed to obtain initial text information;

[0048] Perform text preprocessing on the initial text information to obtain the text information to be processed.

[0049] Optionally, the method further includes:

[0050] According to the at least one entity word and the target business scenario category, review the at least one picture to be reviewed to obtain the review result corresponding to the at least one picture to be reviewed.

[0051] In a second aspect, the present application provides a business scenario determination device, and the device includes:

[0052] An information acquisition module, configured to acquire text information to be processed;

[0053] An entity extraction and attribute extraction module, which is used to determine at least one entity word and at least one attribute word in the to-be-processed text information; the at least one attribute word is index data in a business dimension;

[0054] A text word segmentation module, which is used to perform word segmentation processing on the to-be-processed text information to obtain a text character sequence;

[0055] A first representation module, which is used to perform embedding representation processing on the text character sequence to obtain text representation data;

[0056] A second representation module, which is used to generate marked representation data according to the text character sequence, the at least one entity word, and the at least one attribute word, and the marked representation data is used to indicate the first character corresponding to each entity word and the second character corresponding to each attribute word in the text character sequence;

[0057] A feature extraction module, which is used to perform feature extraction processing on the text representation data and the marked representation data to obtain text semantic feature data corresponding to the to-be-processed text information;

[0058] A scenario classification module, which is used to perform scenario classification processing according to the text semantic feature data to obtain a target business scenario category corresponding to the to-be-processed text information.

[0059] In a third aspect, the present application provides a computer-readable storage medium, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement a business scenario determination method as described in the first aspect.

[0060] In a fourth aspect, the present application provides a computer device, which includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement a business scenario determination method as described in the first aspect.

[0061] In a fifth aspect, the present application provides a computer program product, which includes computer instructions, and when the computer instructions are executed by a processor, a business scenario determination method as described in the first aspect is implemented.

[0062] The business scenario determination method, device, medium, and equipment provided by the present application have the following technical effects:

[0063] This application provides a solution for efficiently and accurately determining the business scenario corresponding to text information. In the technical solution provided by this application, the text information to be processed is segmented to obtain a text character sequence, and the text character sequence is subjected to an embedding representation process to obtain text representation data. At the same time, entity words and attribute words are extracted from the text information to be processed, obtaining at least one entity word and at least one attribute word in the text information to be processed, where the attribute word is indicator data of a business dimension. In the solution provided by this application, the entity words and attribute words are associated with the business scenario; combining the obtained text character sequence, at least one entity word, and at least one attribute word, marker representation data is generated, and the marker representation data is used to indicate the first character corresponding to each entity word in the text character sequence, that is, each attribute word; by performing feature extraction processing on the text representation data and the marker representation data, text semantic feature data corresponding to the text information to be processed is obtained. Due to the indication of the marker representation data, the feature representation of the entity words and attribute words is strengthened in the obtained text semantic feature data. Therefore, when performing scenario classification processing based on the text semantic feature data, the accuracy of the target business scenario category corresponding to the text information to be processed is higher.

[0064] The technical solution provided by this application infers the business scenario by extracting entity words and attribute words related to the business scenario from the text information to be processed. During the inference process of the business scenario, the extraction of entity word features and attribute word features is strengthened through the indication of the marker representation data. Then, the business scenario is classified according to the text semantic feature data with strengthened entity word features and attribute word features, which can effectively improve the accuracy of business scenario classification, thereby better serving subsequent tasks such as auditing.

[0065] Additional aspects and advantages of this application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of this application. Brief Description of the Drawings

[0066] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the accompanying drawings required for describing the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0067] Figure 1 It is a schematic diagram of the implementation environment of a business scenario determination method provided by an embodiment of this application;

[0068] Figure 2 It is a schematic flowchart of a business scenario determination method provided by an embodiment of this application;

[0069] Figure 3 It is a schematic flow diagram of text recognition provided by an embodiment of the present application;

[0070] Figure 4 It is a schematic flow diagram of entity word extraction and attribute word extraction provided by an embodiment of the present application;

[0071] Figure 5 It is a schematic diagram of text word segmentation provided by an embodiment of the present application;

[0072] Figure 6 It is a schematic flow diagram of the generation process of role representation data provided by an embodiment of the present application;

[0073] Figure 7 It is a schematic flow diagram of the generation process of distance representation data provided by an embodiment of the present application;

[0074] Figure 8 It is a schematic diagram of the composition of marker representation data provided by an embodiment of the present application;

[0075] Figure 9 It is a schematic diagram of the structure of the Transformer model provided by an embodiment of the present application;

[0076] Figure 10 It is a schematic flow diagram of another method for determining a business scenario provided by an embodiment of the present application;

[0077] Figure 11 It is a schematic flow diagram of a method for scenario classification processing provided by an embodiment of the present application;

[0078] Figure 12 It is a schematic diagram of data for scenario classification processing provided by an embodiment of the present application;

[0079] Figure 13 It is a schematic diagram of a device for determining a business scenario provided by an embodiment of the present application;

[0080] Figure 14 It is a schematic diagram of the hardware structure of a device for implementing a method for determining a business scenario provided by an embodiment of the present application. Detailed implementation manners

[0081] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics.

[0082] The solution provided in the embodiments of this application relates to technologies such as Deep Learning (DL) and Natural Language Processing (NLP) in artificial intelligence.

[0083] Deep Learning (DL) is a major research direction in the field of Machine Learning (ML). It is introduced into machine learning to make it closer to the original goal - artificial intelligence. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability to analyze and learn like humans and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed those of previous related technologies. Deep learning has achieved many results in search technology, data mining, machine learning, machine translation, natural language processing, multimedia learning, speech, recommendation and personalization technology, and other related fields. Deep learning enables machines to imitate human activities such as seeing, hearing, and thinking, solves many complex pattern recognition problems, and makes great progress in artificial intelligence-related technologies.

[0084] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0085] The solution provided by the embodiments of this application can be deployed in the cloud, which also involves cloud technology and so on.

[0086] Cloud technology: It refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data calculation, storage, processing, and sharing. It can also be understood as the general term for network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. The back-end services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system. Therefore, cloud technology needs to be supported by cloud computing. Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services according to needs. The network that provides resources is called the "cloud". The resources in the "cloud" seem to users to be infinitely expandable, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. As a basic capability provider of cloud computing, a cloud computing resource pool platform will be established, abbreviated as the cloud platform, generally referred to as Infrastructure as a Service (IaaS). Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (which can be virtual machines, including operating systems), storage devices, and network devices.

[0087] To improve the accuracy of business scenario classification, embodiments of the present application provide a method, apparatus, medium, and device for determining a business scenario. Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end.

[0088] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used in appropriate cases can be interchanged so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0089] It can be understood that in the specific implementation of the present application, data related to pictures to be reviewed, text information to be processed, etc. are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0090] Please refer to Figure 1 , which is a schematic diagram of the implementation environment of a method for determining a business scenario provided by an embodiment of the present application. As Figure 1 shown, the implementation environment may at least include a client 01 and a server 02.

[0091] Specifically, the client 01 may include devices such as smartphones, desktop computers, tablet computers, laptop computers, in-vehicle terminals, digital assistants, smart wearable devices, and voice interaction devices, and may also include software running on the devices. For example, some web pages provided by service providers to users, or applications provided by these service providers to users. Specifically, the client 01 may be used to input text or pictures and upload the text or pictures to the server 02, so that the server 02 determines the text information to be processed from the text or pictures and infers the business scenario category corresponding to the text information to be processed.

[0092] Specifically, the server 02 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), big data, and artificial intelligence platforms. The server 02 can include a network communication unit, a processor, a memory, and so on. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this. Specifically, the server 02 can be used to receive the text or picture uploaded by the client 01, and determine the text information to be processed according to the text or picture; extract at least one entity word and at least one attribute word from the text information to be processed, and at least one attribute word is the index data of the business dimension; perform word segmentation processing on the text information to be processed to obtain a text character sequence; perform embedding representation processing on the text character sequence to obtain text representation data; generate marked representation data according to the text character sequence, at least one entity word, and at least one attribute word, and the marked representation data is used to indicate the first character corresponding to each entity word and the second character corresponding to each attribute word in the text character sequence; perform feature extraction processing on the text representation data and the marked representation data to obtain the text semantic feature data corresponding to the text information to be processed; perform scenario classification processing according to the text semantic feature data to obtain the target business scenario category corresponding to the text information to be processed.

[0093] The embodiments of this application can also be implemented in combination with cloud technology. Cloud technology (Cloud technology) refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing. It can also be understood as the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. Cloud technology requires cloud computing as the support. Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". Specifically, the server 02 and the database are located in the cloud, and the server 02 can be a physical machine or a virtualized machine.

[0094] The following introduces a business scenario determination method provided by this application. Figure 2It is a flowchart of a method for determining a service scenario provided by an embodiment of the present application. The present application provides the method operation steps as described in the embodiment or flowchart, but based on routine or non-creative labor, it may include more or fewer operation steps. The step sequence listed in the embodiment is only one way among the execution sequences of numerous steps and does not represent the only execution sequence. When the actual system or server product executes, it can be executed in the order of the method shown in the embodiment or the accompanying drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). Please refer to Figure 2 , a method for determining a service scenario provided by an embodiment of the present application may include the following steps:

[0095] S210: Obtain the text information to be processed.

[0096] In the embodiment of the present application, the text information to be processed is the text data of the service scenario to be determined, and the service scenario corresponding to the text data can be determined by the method provided by the embodiment of the present application. The text information to be processed can be determined according to the text or picture submitted and uploaded by the user through the client in a certain service scenario.

[0097] In an embodiment of an audit application of the present application, in combination with Figure 3 as shown, the text information to be processed is derived from at least one picture to be audited uploaded by the client. There can be various types of at least one picture to be audited, and the types of at least one picture to be audited can include, but are not limited to, identity authentication picture A, face picture B, document scan picture C, etc.; by performing text recognition on at least one picture to be audited, such as calling the program interface of Optical Character Recognition (OCR) to recognize each type of at least one picture to be audited, the initial text information is obtained. The initial text information can include all text characters contained in at least one picture to be audited; the initial text information is preprocessed to obtain the text information to be processed. Specifically, the text preprocessing can include text rectification, text noise reduction, etc. Exemplarily, there will be certain text errors after OCR recognition. For example, the comma in the amount "40.000.00" is recognized as a dot, and it can be corrected to "40,000.00" after text rectification; text noise reduction is to perform processing such as removing stop words and removing noise on the text information to be processed, such as removing characters like "7ct" that have no actual physical meaning or business indicator meaning. In the above embodiment, the original picture to be audited is converted into a text format to implement the method provided by the present application. At the same time, during the conversion process, text noise reduction and text rectification processing are used to improve the normativity of the text information to be processed, so as to more accurately determine the service scenario corresponding to the picture to be audited / text information to be processed.

[0098] In another embodiment of the present application, image recognition and image semantics extraction can be performed on at least one picture to be reviewed, and the recognition result and extraction result are represented in text form to obtain text information to be processed.

[0099] In another embodiment of the present application, the text information to be processed can be sourced from a video to be reviewed uploaded by a client. Automatic Speech Recognition (ASR) processing is performed on the video to be reviewed to obtain the text information to be processed, or text extraction processing of the images at the frame level in the video to be reviewed can be performed to obtain the text information to be processed.

[0100] S220: Determine at least one entity word in the text information to be processed and at least one attribute word in the text information to be processed; the at least one attribute word is index data in the business dimension.

[0101] In the embodiment of the present application, the entity words and attribute words involved in the text information to be processed can help determine the business scenario corresponding to the text information to be processed. Among them, the entity words indicate the atomic information elements involved in the text information to be processed, such as personal names, organization / institution names, geographical locations, events / dates, character values, amount values, etc. The attribute words are index data in the business dimension, which can be index values corresponding to various business indicators. The business indicators can include, but are not limited to, account identifiers, account levels, account permissions, product / service types, interaction protocol types, interaction protocol version numbers, interaction resource types, interaction resource quantities, etc. The entity words and attribute words implicitly contain the description information of the business scenario, and can be regarded as having an implicit association relationship between the entity words, attribute words and the business scenario. Entity extraction (also known as named entity recognition, abbreviated as NER for short) and attribute extraction are performed on the text information to be processed to obtain at least one entity word and at least one attribute word, so that the entity features and attribute features can be explicitly strengthened and introduced into the scenario classification, which can improve the inference accuracy of the business scenario category.

[0102] In one embodiment of the present application, specifically, as Figure 4 shown, step S220 can be implemented as:

[0103] S221: Obtain the regular library corresponding to each type of entity in a variety of pre-constructed entities.

[0104] The regular library corresponding to each type of entity can include at least one regular matching rule. The regular matching rule uses special characters and syntax to construct a regular expression for matching specific patterns in the text, such as matching specific characters, quantifiers or character combination methods.

[0105] S222: Perform regular matching based on the regular expression library corresponding to each type of entity and the text information to be processed, to obtain at least one original entity word.

[0106] Feasibly, the regular expression library corresponding to each type of entity among multiple entities can have corresponding priorities, and the matching between the text information to be processed and the regular expression library corresponding to each type of entity is performed in sequence according to the priority level from high to low.

[0107] S223: Perform entity completion processing or entity conversion processing on at least one original entity word to obtain at least one entity word.

[0108] Exemplarily, for the entity of date, the standard text format is "year-month-day". For the original entity word such as "2023.8.2", perform digit completion and symbol conversion to obtain the final entity word "2023-08-02".

[0109] Exemplarily, for the entity of date, if the extracted original entity word is "today", it can be converted according to the current date. Assuming the current date is 20230802, it can be converted to "2023-08-02".

[0110] In the above embodiments, based on the regular matching method, modular and diversified extraction of entity words can be realized, which can improve the utilization rate of the code program for performing regular matching, and also facilitate the expansion of entity types and the supplementation of matching rules. At the same time, the text preprocessing is used to unify the text forms corresponding to the same type of entity, ensuring the accuracy and effectiveness of data in the subsequent processing process.

[0111] S224: Obtain a pre-constructed business attribute keyword dictionary.

[0112] The business attribute keyword dictionary can include business indicators involved in various products / services or identifiers of business indicators. The business attribute keyword dictionary can be a data table that is updated in real time in the database, and the business indicators are the data fields in the data table.

[0113] S225: Perform mapping matching processing based on the business attribute keyword dictionary and the text information to be processed to obtain at least one attribute word.

[0114] Feasibly, according to the data fields corresponding to the business indicators in the business attribute keyword dictionary, determine the field values in the text information to be processed corresponding to the data fields as the attribute words matching the business indicators.

[0115] In the above embodiments, based on the mapping matching method, effective and accurate extraction of attribute words is realized.

[0116] S230: Perform word segmentation processing on the text information to be processed to obtain a text character sequence.

[0117] In the embodiments of the present application, in order to convert the text information to be processed into a numerical form that can be processed by an artificial intelligence model, it is first necessary to perform word segmentation processing on the text information to be processed, which can also be called Tokenization processing. It is to split the text information to be processed into a string sequence (the smallest element among them is usually called a token or a word), so as to facilitate subsequent processing and analysis work.

[0118] Exemplarily, as Figure 5 shown, for the text "Changjiang Automobile has entered the bankruptcy liquidation process", it can be split into a string sequence with individual Chinese characters as the smallest elements, and additional delimiters [CLS] and [SEP] are added at the beginning and end of the string sequence respectively. Among them, [CLS] is the abbreviation of "classification". In text classification tasks, it usually represents the beginning of a sentence or a document. In the BERT (Bidirectional Encoder Representation from Transformers) model, [CLS] corresponds to the word vector of the first word in the input text, and the first word vector output by the corresponding BERT model is usually used to predict the category of the text. [SEP] is the abbreviation of "separator", which usually represents the end of a sentence or a document. In the BERT model, [SEP] corresponds to the word vector of the last word in the input text, and its function is to separate different sentences. For example, when processing sentence pairs in the BERT model, a [SEP] is usually inserted between the two sentences to represent their demarcation point.

[0119] S240: Perform embedding representation processing on the text character sequence to obtain text representation data.

[0120] In the embodiments of the present application, in order to convert the text information to be processed into a data form that can be processed by an artificial intelligence model, such as a vector form, it is first necessary to perform word segmentation processing on the text information to be processed, and then perform embedding representation processing after the word segmentation processing, which can also be called Embedding processing. The obtained text representation data is the data form that can be directly processed by the model or network in the scene classification process.

[0121] In an embodiment of the present application, embedding representation processing is performed in the BERT model. Specifically, for the text character sequence, such as Figure 5As shown, for each segmented character, word vector representation, paragraph representation, and position representation are performed respectively, and the corresponding token embeddings, segment embeddings, and position embeddings of each character are obtained. The token embeddings are the vectorized representations of the corresponding characters, the segment embeddings are used to indicate which sentence the corresponding character belongs to, and the position embeddings are used to indicate the position information of the corresponding character in the entire text character sequence.

[0122] In another embodiment of the present application, the ALBert (Alite BERT, lightweight BERT) model can also be used, which will not be elaborated here.

[0123] S250: Generate labeled representation data according to the text character sequence, at least one entity word, and at least one attribute word. The labeled representation data is used to indicate the first character corresponding to each entity word and the second character corresponding to each attribute word in the text character sequence.

[0124] In the embodiments of the present application, considering the implicit association relationship between entity words, attribute words, and business scenarios, if the data used for scene classification can more prominently feature entity words and attribute words, the accuracy of scene classification can be improved. Therefore, in the embodiments of the present application, by generating labeled representation data that can indicate the first character corresponding to each entity word and the second character corresponding to each attribute word in the text character sequence, more attention can be paid to the feature representation data corresponding to entity words and the feature representation data corresponding to attribute words during subsequent analysis and processing of the text representation data.

[0125] In the embodiments of the present application, the labeled representation data can indicate the first character corresponding to each entity word and the second character corresponding to each attribute word in the text character sequence, and correspondingly can also indicate the third character corresponding to other words except entity words and attribute words. Other words will also be referred to as ordinary words later. The entity words, attribute words, and ordinary words can be distinguished according to the labeled representation data.

[0126] In the embodiments of the present application, there is a mapping relationship between the text character sequence, the text representation data, and the labeled representation data. Specifically, the text character sequence can be expressed as: [CLS]T1 T2 T3…T a …T b …T m …T n …[SEP], where T1 T2 T3… are the third characters corresponding to at least one ordinary word, T a …T bis the first character corresponding to at least one entity word, T m …T n is the second character corresponding to at least one attribute word. The above character distribution can be simply described. In actual situations, common words, entity words, and attribute words can be interleaved. Correspondingly, the first character, the second character, and the third character are also interleaved. The text representation data can be correspondingly represented as: E [CLS] E1 E2 E3…E a …E b …E m …E n …E [SEP] , where E1E2 E3… is the text representation data corresponding to each third character, E a …E b is the text representation data corresponding to each first character, E m …E n is the text representation data corresponding to each second character. The marker representation data can be correspondingly identified as: F [CLS] F1 F2 F3…F a …F b …F m …F n …F [SEP] , where F1 F2 F3… is the marker representation data corresponding to each third character, F a …F b is the marker representation data corresponding to each first character, F m …F n is the marker representation data corresponding to each second character. The order of each third character in the text character sequence, the order of the text representation data corresponding to this third character in the text representation data, and the order of the marker representation data corresponding to this third character in the marker representation data are all the same. The order of each second character in the text character sequence, the order of the text representation data corresponding to this second character in the text representation data, and the order of the marker representation data corresponding to this second character in the marker representation data are all the same. The order of each first character in the text character sequence, the order of the text representation data corresponding to this first character in the text representation data, and the order of the marker representation data corresponding to this first character in the marker representation data are all the same. Further, the dimension of the marker representation data corresponding to each character is the same as the dimension of the text representation data corresponding to each character.

[0127] In an embodiment of the present application, the marker representation data includes role representation data. The role representation data can indicate the first character corresponding to each entity word and the second character corresponding to each attribute word in the text character sequence, that is, the type marking of entity words, attribute words, and common words in the text character sequence. Specifically, as Figure 6As shown, step S250 can be implemented as follows:

[0128] S2501: Traverse each text character in the text character sequence.

[0129] S2502: When the currently traversed text character is the first character, determine that the role representation data corresponding to the currently traversed text character is the preset first role representation data.

[0130] Feasibly, when performing text tokenization, the first characters corresponding to at least one entity word can be aggregated into a first character set, the second characters corresponding to at least one attribute word can be aggregated into a second character set, and the third characters corresponding to at least one common word can be aggregated into a third character set.

[0131] By performing character matching between the currently traversed text character and the first character set, the second character set, and the third character set respectively, it can be determined which character set the currently traversed text character belongs to, and correspondingly, it can be determined whether the currently traversed text character is the first character, the second character, or the third character.

[0132] When the currently traversed text character is the first character, determine that the role representation data corresponding to the currently traversed text character is the preset first role representation data, that is, the first role representation data is a mark for the first character corresponding to the entity word. Feasibly, a preset value or symbol can be selected and the data or symbol can be vectorized to obtain the first role representation data.

[0133] S2503: When the currently traversed text character is the second character, determine that the role representation data corresponding to the currently traversed text character is the preset second role representation data.

[0134] When the currently traversed text character is the second character, determine that the role representation data corresponding to the currently traversed text character is the preset second role representation data, that is, the second role representation data is a mark for the second character corresponding to the attribute word. Feasibly, a preset value or symbol can be selected and the data or symbol can be vectorized to obtain the second role representation data.

[0135] S2504: When the currently traversed text character is neither the first character nor the second character, determine that the role representation data corresponding to the currently traversed text character is the preset third role representation data.

[0136] When the currently traversed text character is neither the first character nor the second character, that is, the third character, it is determined that the role representation data corresponding to the currently traversed text character is the preset third role representation data. That is, the third role representation data is a mark for the third character corresponding to the common word. Feasibly, a preset value or symbol can be selected and the data or symbol can be vectorized to obtain the third role representation data.

[0137] S2505: Generate role representation data according to the role representation data corresponding to each text character.

[0138] It can be understood that the first role representation data, the second role representation data, and the third role representation data are different from each other. In a feasible embodiment, the first role representation data, the second role representation data, and the third role representation data are learnable variables. After randomly initializing the parameters in the role representation network, the model parameters can be trained together with the downstream tasks, so that targeted role representation processing can be performed on different text characters in the actual application stage. Further, the role representation data can not only indicate the first character corresponding to the entity word, the second character corresponding to the attribute word, and the third character corresponding to the common word, but also indicate the entity feature corresponding to the entity word and the attribute feature corresponding to the attribute word.

[0139] The generated role representation data can be in matrix form, with dimensions of max_length*hidden_size, where max_length corresponds to the number of text characters in the text character sequence, and hidden_size represents the number of hidden layer nodes, corresponding to the dimension of the role representation data corresponding to each text character.

[0140] In the above embodiment, by marking the text characters corresponding to the entity word, the attribute word, and the common word, the marked representation data that can indicate and distinguish these three types of words is obtained, so as to strengthen the attention to the entity word and the attribute word in the subsequent processing and analysis.

[0141] In an embodiment of the present application, the marked representation data includes distance representation data. In addition to marking the types of entity words, attribute words, and common words in the text character sequence, the distance representation data can also indicate the distance between the attribute word and the entity word in the text character sequence. Specifically, as Figure 7 shown, step S250 can be implemented as:

[0142] S2506: Traverse each text character in the text character sequence.

[0143] S2507: When the currently traversed text character is the first character, determine that the distance representation data corresponding to the currently traversed text character is the preset first distance representation data.

[0144] The method for determining whether the currently traversed text character is the first character, the second character, or the third character can refer to the foregoing embodiments and will not be elaborated here.

[0145] When the currently traversed text character is the first character, determine that the distance representation data corresponding to the currently traversed text character is the preset first distance representation data. That is, the first distance representation data is a mark for the first character corresponding to the entity word. Feasibly, a preset value such as 1 can be selected and the data can be vectorized to obtain the first distance representation data.

[0146] S2508: When the currently traversed text character is the second character, determine at least one text distance data.

[0147] Each text distance data in the at least one text distance data represents the text similarity between the first character corresponding to each entity word and the currently traversed text character.

[0148] Feasibly, the distance between the text representation data corresponding to the first character and the text representation data corresponding to the currently traversed text character in the vector space can be used as the text distance data between the first character and the currently traversed text character. The distance in the vector space also represents the text similarity between the first character and the currently traversed text character.

[0149] S2509: Based on the at least one text distance data, determine the target distance representation data corresponding to the target text distance data, and use the target distance representation data as the distance representation data corresponding to the currently traversed text character.

[0150] The target text distance data is the minimum value among the at least one text distance data.

[0151] That is, the target distance representation data corresponding to the second character is determined according to the target text distance data between the second character and the closest first character. For example, the target text distance data between the first character closest to the second character and the second character is 0.8, and 0.8 is vectorized to obtain the target distance representation data.

[0152] S2510: When the currently traversed text character is neither the first character nor the second character, determine that the distance representation data corresponding to the currently traversed text character is the preset second distance representation data.

[0153] When the currently traversed text character is neither the first character nor the second character, that is, the third character, determine that the distance representation data corresponding to the currently traversed text character is the preset third distance representation data. That is, the third distance representation data is a mark for the third character corresponding to the common word. Feasibly, a preset value such as 0 can be selected and the data can be vectorized to obtain the second distance representation data. The second distance representation data can also indicate that the vector distance between the third character corresponding to the common word and the first character corresponding to the entity word is 0.

[0154] S2511: Generate distance representation data according to the distance representation data corresponding to each text character.

[0155] In a feasible embodiment, the first distance representation data and the second distance representation data can be fixed data, and the target distance needs to be determined according to the distance between the second character corresponding to each attribute word and the first character corresponding to each entity word in the text character sequence.

[0156] In another feasible embodiment, the first distance representation data, the target distance representation data, and the second distance representation data are learnable variables. After randomly initializing the parameters in the distance representation network, the model parameters can be trained together with the downstream task, so that targeted distance representation processing can be performed on different text characters in the actual application stage. Further, in addition to indicating the first character corresponding to the entity word, the second character corresponding to the attribute word, and the third character corresponding to the common word, the distance representation data can also indicate the entity feature corresponding to the entity word, the attribute feature corresponding to the attribute word, and the similarity degree between the first character corresponding to the entity word and the second character corresponding to the attribute word.

[0157] The generated distance representation data can be in matrix form, with a dimension of max_length*hidden_size. Among them, max_length corresponds to the number of text characters in the text character sequence, and hidden_size represents the number of hidden layer nodes, corresponding to the dimension of the distance representation data corresponding to each text character.

[0158] In the above embodiments, by determining the distances of the attribute word and the common word from the entity word in the vector space, the text characters corresponding to the entity word, the attribute word, and the common word are marked to obtain the marked representation data that can indicate and distinguish these three types of words, so as to strengthen the attention to the entity word and the attribute word in the subsequent processing and analysis. In addition, the marked representation data can also indicate the degree of association between the attribute word and the entity word, and can pay more attention to the learning and extraction of the features of the attribute word with a high degree of association with the entity word in the subsequent processing and analysis.

[0159] In one embodiment of the present application, the labeled representation data may include role representation data and distance representation data, and the role representation data and the distance representation data may be represented in the form of a matrix. As Figure 8 shown, for the text character sequence: [CLS]T1 T2 T3…Ta…Tb…Tm…Tn…[SEP], the text representation data is the superposition of Token Embeddings, Segment Embeddings, and Position Embeddings, and the role representation data is Figure 8 the Role Embeddings shown in Figure 8 , and the distance pointer data is the Distance Embeddings shown in EAMT . When T1 T2 T3… are the third characters corresponding to at least one ordinary word, Ta…Tb are the first characters corresponding to at least one entity word, and Tm…Tn are the second characters corresponding to at least one attribute word, the first role representation data (indicating the entity word) is denoted as ATT , the second role representation data (indicating the attribute word) is E O , the third role representation data (indicating the ordinary word) is E 99 , the second distance representation data (indicating the ordinary word) is denoted as E ji , the first distance representation data is E0, and the target distance representation data is E ji , where E

[0160] S260: Perform feature extraction processing on the text representation data and the labeled representation data to obtain the text semantic feature data corresponding to the text to be processed.

[0161] In the embodiment of the present application, during the feature extraction processing, the labeled representation data can be used to strengthen the feature learning of the representation data corresponding to the entity word and the representation data corresponding to the attribute word in the text representation data, so that the feature data corresponding to the entity word and the attribute word can be more prominently reflected in the output text semantic feature data.

[0162] In one embodiment of the present application, feature extraction processing is performed on the text representation data and the labeled representation data based on the BERT model. Specifically, step S260 can be implemented as:

[0163] S261: Concatenate the text representation data and the labeled representation data to obtain the target representation data.

[0164] As Figure 8 shown, the text representation data corresponding to each character in the text character sequence and the labeled representation data corresponding to the same character are horizontally concatenated in vectors to obtain the target representation data.

[0165] S262: Input the target representation data into the feature transformer model for feature extraction processing to obtain target feature data.

[0166] As Figure 9 shown, BERT uses a bidirectional feature transformer model for feature extraction. With the multi-head self-attention mechanism and the encoding and decoding process of the Transformer model, it is possible to achieve deep interaction between the target representation data corresponding to each text character in the input sequence, so as to obtain the syntactic relationship encoding data and semantic relationship encoding data between the target representation data corresponding to the text character and the target representation data corresponding to all other text characters in the input. In addition, the encoding of the Transformer model supports parallel input of the target representation data corresponding to multiple text character sequences, thus greatly improving the processing efficiency.

[0167] S263: Determine the text semantic feature data in the target feature data based on a preset semantic feature identifier.

[0168] In the sequence output by the Transformer model, the semantic feature identifier [CLS] corresponds to the first word vector in the sequence and is a semantic feature vector that can represent the entire text character sequence, which can be directly used for classification tasks.

[0169] In the above embodiments, the Transformer model is used for deep and effective feature extraction processing, and it is possible to strengthen the extraction of entity word features and attribute word features by means of marked representation data during the feature extraction processing, so that the expression of entity word features and attribute word features is strengthened in the text semantic feature data.

[0170] S270: Perform scene classification processing based on the text semantic feature data to obtain the target business scene category corresponding to the text information to be processed.

[0171] In the embodiments of the present application, the text semantic feature data can be input into a classification model or a classification network for scene classification processing, that is, to predict the business scene category corresponding to the text information to be processed. The classification model or classification network can be constructed based on the softmax layer. Softmax is a common multi-classifier used to predict the probability that the input belongs to each category, and its working principle can be shown in formula (1), where C represents the number of business scene categories, z i represents the output of the i-th node of the softmax layer, z j represents the output of the j-th node of the softmax layer, j = 1,...C, and e is the natural constant:

[0172]

[0173] Through formula (1), the output values of multi-classification can be converted into a probability distribution with a range of [0, 1] and a sum of 1. The probability distribution can indicate the probabilities of the text information to be processed belonging to each business scenario category respectively. According to the probability value with the largest value in the probability distribution, the target business scenario category corresponding to the text information to be processed can be determined.

[0174] In the embodiments of the present application, the business scenario can indicate a specific business type, an application interaction type, and can also indicate the interaction intention or interaction requirement of the user in the current application environment. Thus, after clarifying the target business scenario category corresponding to the text information to be processed, targeted and effective review processing can be performed on the text information to be processed.

[0175] Please refer to Figure 3 As Figure 10 shown, another method for determining a business scenario provided by the embodiments of the present application may include the following steps:

[0176] S210: Obtain the text information to be processed.

[0177] S220: Determine at least one entity word in the text information to be processed and at least one attribute word in the text information to be processed; the at least one attribute word is indicator data of a business dimension.

[0178] S230: Perform word segmentation processing on the text information to be processed to obtain a text character sequence.

[0179] S240: Perform embedding representation processing on the text character sequence to obtain text representation data.

[0180] S250: Generate marker representation data according to the text character sequence, at least one entity word, and at least one attribute word, where the marker representation data is used to indicate the first character corresponding to each entity word and the second character corresponding to each attribute word in the text character sequence.

[0181] Steps S210 to S250 may refer to the foregoing embodiments and will not be elaborated herein.

[0182] S280: Perform feature extraction processing on the text representation data and the marker representation data to obtain text semantic feature data corresponding to the text information to be processed, entity feature data corresponding to at least one entity word, and attribute feature data corresponding to at least one attribute word.

[0183] According to the foregoing embodiments, the text representation data and the marker representation data are concatenated to obtain target representation data, and the target representation data is input into a feature converter model for feature extraction processing to obtain target feature data. Based on a preset semantic feature identifier, the text semantic feature data in the target feature data can be determined; entity feature data corresponding to at least one entity word and attribute feature data corresponding to at least one attribute word can also be extracted from the target feature data.

[0184] S290: Perform scenario classification processing on the text semantic feature data, entity feature data, and attribute feature data to obtain the target business scenario category corresponding to the text information to be processed.

[0185] In an embodiment of the present application, the text semantic feature data, entity feature data, and attribute feature data can be concatenated and then input into a classification model or classification network for scenario classification processing.

[0186] In the foregoing embodiments, in addition to strengthening the expression of the semantic features of entity words and attribute words in the text semantic feature data, the entity feature data corresponding to the entity words and the attribute feature data corresponding to the attribute words are also explicitly input into the classification model or classification network, and further, the entity words and attribute words are used to improve the effectiveness of scenario classification.

[0187] In an embodiment of the present application, as Figure 11 shown, step S290 can be implemented as:

[0188] S291: Perform average pooling processing on the entity feature data to obtain target entity feature data.

[0189] S292: Perform average pooling processing on the attribute feature data to obtain target attribute feature data.

[0190] Specifically, considering that the entity feature data and attribute feature data are feature data represented in vector form, the average pooling processing on the entity feature data can be performed by calculating the average value of the entity feature data according to each component data in the entity feature data to obtain the target entity feature data, and the average pooling processing on the attribute feature data is the same.

[0191] S293: Concatenate the text semantic feature data, target entity feature data, and target attribute feature data to obtain concatenated feature data.

[0192] S294: Input the concatenated feature data into a classification network for class prediction processing of the business scenario to obtain the target business scenario category.

[0193] The process of class prediction processing of the business scenario can refer to step S270 in the foregoing embodiments and will not be elaborated here.

[0194] In the above embodiments, on the one hand, average pooling processing can reduce the dimension to splice the target entity feature data, target attribute feature data, and text semantic feature data, and on the other hand, it can retain the main feature data.

[0195] In a specific embodiment of the present application, as Figure 12 shown, for the text character sequence: [CLS]T1T2 T3…T a …T b …T m …T n …[SEP], where T1 T2 T3… are the third characters corresponding to at least one ordinary word, T a …T b are the first characters corresponding to at least one entity word, T m …T n … are the second characters corresponding to at least one attribute word. After tokenization processing, embedding representation processing, and feature extraction processing by a BERT model including multiple Transformers, the target feature data [CLS]E1 E2 E3…E a …E b …E m …E n …[SEP] is obtained. The text semantic feature data [CLS] is extracted from the target feature data, and at the same time, the feature data of the entity words and attribute words in the target feature data are averaged and pooled and then concatenated (Concat) with [CLS]. The concatenated feature data is input into the softmax layer, and finally, the target business scenario category corresponding to the inferred text information or text character sequence to be processed is output.

[0196] In an embodiment of the present application, at least one to-be-reviewed picture is reviewed according to at least one entity word and the target business scenario category, and the review result corresponding to at least one to-be-reviewed picture is obtained. Compared with the manual review method, the review cost can be reduced and the review efficiency can be improved. Compared with the review method based on artificial intelligence in the related technology, since more accurate target business scenario categories and entity words are introduced, the review accuracy is also effectively improved.

[0197] As can be seen from the above embodiments, in a method for determining a service scenario provided by the present application, word segmentation is performed on the text information to be processed to obtain a text character sequence, and embedding representation processing is performed on the text character sequence to obtain text representation data. At the same time, entity words and attribute words are extracted from the text information to be processed, and at least one entity word and at least one attribute word in the text information to be processed are obtained. Among them, the attribute words are index data of the service dimension. In the solution provided by the present application, the entity words and attribute words are associated with the service scenario; combining the obtained text character sequence, at least one entity word and at least one attribute word, marker representation data is generated, and the marker representation data is used to indicate the first character corresponding to each entity word in the text character sequence, that is, each attribute word; by performing feature extraction processing on the text representation data and the marker representation data, text semantic feature data corresponding to the text information to be processed is obtained. Due to the indication of the marker representation data, the feature representation of the entity words and attribute words is strengthened in the obtained text semantic feature data. Therefore, when performing scenario classification processing based on the text semantic feature data, the accuracy of the target service scenario category corresponding to the text information to be processed is higher.

[0198] The method provided by the present application infers the service scenario by extracting entity words and attribute words related to the service scenario from the text information to be processed. During the inference process of the service scenario, the extraction of entity word features and attribute word features is strengthened through the indication of the marker representation data, and then the service scenario is classified according to the text semantic feature data with strengthened entity word features and attribute word features, which can effectively improve the accuracy of service scenario classification, thereby better serving subsequent tasks such as auditing.

[0199] The embodiment of the present application also provides a service scenario determination device 1300, as Figure 13 shown, the device may include:

[0200] An information acquisition module 1310, configured to acquire text information to be processed;

[0201] An entity extraction and attribute extraction module 1320, configured to determine at least one entity word in the text information to be processed and at least one attribute word in the text information to be processed; the at least one attribute word is index data of the service dimension;

[0202] A text word segmentation module 1330, configured to perform word segmentation on the text information to be processed to obtain a text character sequence;

[0203] A first representation module 1340, configured to perform embedding representation processing on the text character sequence to obtain text representation data;

[0204] The second characterization module 1350 is configured to generate labeled characterization data according to the text character sequence, the at least one entity word, and the at least one attribute word, where the labeled characterization data is used to indicate the first characters corresponding to each entity word and the second characters corresponding to each attribute word in the text character sequence;

[0205] The feature extraction module 1360 is configured to perform feature extraction processing on the text characterization data and the labeled characterization data to obtain text semantic feature data corresponding to the text information to be processed;

[0206] The scenario classification module 1370 is configured to perform scenario classification processing according to the text semantic feature data to obtain the target business scenario category corresponding to the text information to be processed.

[0207] In an embodiment of the present application, the feature extraction module 1360 may include:

[0208] The first feature extraction unit is configured to perform feature extraction processing on the text characterization data and the labeled characterization data to obtain text semantic feature data corresponding to the text information to be processed, entity feature data corresponding to the at least one entity word, and attribute feature data corresponding to the at least one attribute word;

[0209] The scenario classification module 1370 may include:

[0210] The first scenario classification unit is configured to perform scenario classification processing on the text semantic feature data, the entity feature data, and the attribute feature data to obtain the target business scenario category corresponding to the text information to be processed.

[0211] In an embodiment of the present application, the first scenario classification unit may include:

[0212] The first pooling subunit is configured to perform average pooling processing on the entity feature data to obtain target entity feature data;

[0213] The second pooling subunit is configured to perform average pooling processing on the attribute feature data to obtain target attribute feature data;

[0214] The first splicing subunit is configured to splice the text semantic feature data, the target entity feature data, and the target attribute feature data to obtain spliced feature data;

[0215] The classification subunit is configured to input the spliced feature data into a classification network to perform category prediction processing of the business scenario to obtain the target business scenario category.

[0216] In an embodiment of the present application, the second characterization module 1350 includes:

[0217] A first traversal unit for traversing each text character in the text character sequence;

[0218] A first determination unit for character representation data, which is used to determine that the character representation data corresponding to the currently traversed text character is a preset first character representation data when the currently traversed text character is the first character;

[0219] A second determination unit for character representation data, which is used to determine that the character representation data corresponding to the currently traversed text character is a preset second character representation data when the currently traversed text character is the second character;

[0220] A third determination unit for character representation data, which is used to determine that the character representation data corresponding to the currently traversed text character is a preset third character representation data when the currently traversed text character is not the first character and not the second character;

[0221] A character representation data generation unit for generating the character representation data according to the character representation data corresponding to each text character.

[0222] In an embodiment of the present application, the second representation module 1350 includes:

[0223] A second traversal unit for traversing each text character in the text character sequence;

[0224] A first determination unit for distance representation data, which is used to determine that the distance representation data corresponding to the currently traversed text character is a preset first distance representation data when the currently traversed text character is the first character;

[0225] A text distance data determination unit for determining at least one text distance data when the currently traversed text character is the second character, and each text distance data in the at least one text distance data represents the text similarity between the first character corresponding to each entity word and the currently traversed text character;

[0226] A second determination unit for distance representation data, which is used to determine the target distance representation data corresponding to the target text distance data based on the at least one text distance data, and use the target distance representation data as the distance representation data corresponding to the currently traversed text character; the target text distance data is the minimum value in the at least one text distance data;

[0227] A third distance representation data determination unit, configured to determine that the distance representation data corresponding to the currently traversed text character is a preset second distance representation data when the currently traversed text character is neither the first character nor the second character;

[0228] A distance representation data generation unit, configured to generate the distance representation data according to the distance representation data corresponding to each text character.

[0229] In an embodiment of the present application, the feature extraction module 1360 includes:

[0230] A splicing unit, configured to splice the text representation data and the marker representation data to obtain target representation data;

[0231] A second feature extraction unit, configured to input the target representation data into a feature converter model for feature extraction processing to obtain target feature data

[0232] A text semantic feature data determination unit, configured to determine text semantic feature data in the target feature data based on a preset semantic feature identifier.

[0233] In an embodiment of the present application, the entity extraction and attribute extraction module 1320 includes:

[0234] A regular library acquisition unit, configured to acquire a regular library corresponding to each entity in a plurality of pre-constructed entities;

[0235] A regular matching unit, configured to perform regular matching according to the regular library corresponding to each entity and the text information to be processed to obtain at least one original entity word;

[0236] An entity word determination unit, configured to perform entity completion processing or entity conversion processing on the at least one original entity word to obtain the at least one entity word;

[0237] A dictionary acquisition unit, configured to acquire a business attribute keyword dictionary constructed in advance;

[0238] A mapping matching unit, configured to perform mapping matching processing according to the business attribute keyword dictionary and the text information to be processed to obtain the at least one attribute word.

[0239] In an embodiment of the present application, the information acquisition module 1310 includes:

[0240] A picture acquisition unit, configured to acquire at least one picture to be audited;

[0241] A text recognition unit, configured to perform text recognition on the at least one picture to be audited to obtain initial text information;

[0242] A text preprocessing unit for preprocessing the initial text information to obtain the text information to be processed.

[0243] In an embodiment of the present application, the device 1300 further includes:

[0244] A picture review unit for reviewing the at least one picture to be reviewed according to the at least one entity word and the target business scenario category to obtain the review result corresponding to the at least one picture to be reviewed.

[0245] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0246] It should be noted that when the device provided in the above embodiment realizes its functions, only the division of the above-mentioned function modules is used for illustration. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept. For the specific implementation process, please refer to the method embodiment, which will not be elaborated here.

[0247] The embodiments of the present application provide a computer device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement a business scenario determination method as provided in the above method embodiment.

[0248] Figure 14 The figure shows a schematic hardware structure diagram of a device for implementing a business scenario determination method provided in the embodiments of the present application. The device can participate in forming or include the device or system provided in the embodiments of the present application. As Figure 14As shown, the device 10 may include one or more processors 1002 (illustrated as 1002a, 1002b, ……, 1002n in the figure) (the processor 1002 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1004 for storing data, and a transmission device 1006 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 14 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the device 10 may further include more or fewer components than Figure 14 shown in, or have a different configuration from Figure 14 that shown.

[0249] It should be noted that the above one or more processors 1002 and / or other data processing circuits may generally be referred to as "data processing circuits" herein. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the device 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0250] The memory 1004 may be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the methods described in the embodiments of the present application. The processor 1002 executes various functional applications and data processing by running the software programs and modules stored in the memory 1004, that is, implements the above-mentioned method for determining a service scenario. The memory 1004 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory 1004 may further include a memory remotely set relative to the processor 1002, and these remote memories may be connected to the device 10 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0251] The transmission device 1006 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of device 10. In one example, the transmission device 1006 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 1006 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0252] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of device 10 (or the mobile device).

[0253] The embodiment of the present application also provides a computer-readable storage medium, which can be arranged in a server to store at least one instruction or at least one program segment related to a service scenario determination method in the method embodiment. The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the service scenario determination method provided in the above method embodiment.

[0254] Optionally, in this embodiment, the above storage medium can be located in at least one of multiple network servers in a computer network. Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0255] The embodiment of the present invention also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes a service scenario determination method provided in the above various optional implementation manners.

[0256] It should be noted that the above-mentioned order of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above-mentioned specific embodiments of the present application have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0257] The various embodiments in the present application are all described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiments of the device, equipment, and storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0258] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing the relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like.

[0259] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.

Claims

1. A method for determining a service scenario, characterized in that, The method includes: Obtain the text information to be processed; Determine at least one entity word in the text information to be processed and at least one attribute word in the text information to be processed; the at least one attribute word is index data in the business dimension; Perform word segmentation on the text information to be processed to obtain a text character sequence; Perform embedding representation processing on the text character sequence to obtain text representation data; Generate marker representation data according to the text character sequence, the at least one entity word, and the at least one attribute word, where the marker representation data is used to indicate the first character corresponding to each entity word and the second character corresponding to each attribute word in the text character sequence; Perform feature extraction processing on the text representation data and the marker representation data to obtain text semantic feature data corresponding to the text information to be processed; Perform scenario classification processing according to the text semantic feature data to obtain the target business scenario category corresponding to the text information to be processed.

2. The method according to claim 1, wherein The method further includes: Perform feature extraction processing on the text representation data and the marker representation data to obtain text semantic feature data corresponding to the text information to be processed, entity feature data corresponding to the at least one entity word, and attribute feature data corresponding to the at least one attribute word; The performing scenario classification processing according to the text semantic feature data to obtain the target business scenario category corresponding to the text information to be processed includes: Perform scenario classification processing on the text semantic feature data, the entity feature data, and the attribute feature data to obtain the target business scenario category corresponding to the text information to be processed.

3. The method according to claim 2, wherein The performing scenario classification processing on the text semantic feature data, the entity feature data, and the attribute feature data to obtain the target business scenario category corresponding to the text information to be processed includes: Perform average pooling processing on the entity feature data to obtain target entity feature data; Perform average pooling processing on the attribute feature data to obtain target attribute feature data; Concatenate the text semantic feature data, the target entity feature data, and the target attribute feature data to obtain concatenated feature data; Input the concatenated feature data into a classification network to perform category prediction processing of the business scenario to obtain the target business scenario category.

4. The method according to claim 1, wherein The marker representation data includes role representation data, and the generating marker representation data according to the text character sequence, the at least one entity word, and the at least one attribute word includes: Traverse each text character in the text character sequence; When the currently traversed text character is the first character, determine that the role representation data corresponding to the currently traversed text character is a preset first role representation data; When the currently traversed text character is the second character, determine that the role representation data corresponding to the currently traversed text character is a preset second role representation data; When the currently traversed text character is neither the first character nor the second character, determine that the role representation data corresponding to the currently traversed text character is a preset third role representation data; Generate the character representation data according to the character representation data corresponding to each text character.

5. The method according to claim 1, wherein The token representation data includes distance representation data. Generating the token representation data according to the text character sequence, the at least one entity word, and the at least one attribute word includes: Traverse each text character in the text character sequence; When the currently traversed text character is the first character, determine the distance representation data corresponding to the currently traversed text character as a preset first distance representation data; When the currently traversed text character is the second character, determine at least one text distance data, where each text distance data in the at least one text distance data represents the text similarity between the first character corresponding to each entity word and the currently traversed text character; Based on the at least one text distance data, determine the target distance representation data corresponding to the target text distance data, and use the target distance representation data as the distance representation data corresponding to the currently traversed text character; the target text distance data is the minimum value in the at least one text distance data; When the currently traversed text character is neither the first character nor the second character, determine the distance representation data corresponding to the currently traversed text character as a preset second distance representation data; Generate the distance representation data according to the distance representation data corresponding to each text character.

6. The method according to claim 1, characterized in that Performing feature extraction processing on the text representation data and the token representation data to obtain the text semantic feature data corresponding to the text information to be processed includes: Concatenate the text representation data and the token representation data to obtain target representation data; Input the target representation data into a feature transformation model for feature extraction processing to obtain target feature data; Based on a preset semantic feature identifier, determine the text semantic feature data in the target feature data.

7. The method according to claim 1, characterized in that Determining at least one entity word in the text information to be processed and at least one attribute word in the text information to be processed includes: Obtain the regular expression library corresponding to each entity in a pre-constructed multiple entities; Perform regular matching according to the regular expression library corresponding to each entity and the text information to be processed to obtain at least one original entity word; Perform entity completion processing or entity conversion processing on the at least one original entity word to obtain the at least one entity word; Obtain a pre-constructed business attribute keyword dictionary; Perform mapping matching processing according to the business attribute keyword dictionary and the text information to be processed to obtain the at least one attribute word.

8. The method according to claim 1, wherein Obtaining the text information to be processed includes: Obtain at least one picture to be reviewed; Perform text recognition on the at least one picture to be reviewed to obtain initial text information; Perform text preprocessing on the initial text information to obtain the text information to be processed.

9. The method according to claim 8, wherein The method further includes: Review the at least one picture to be reviewed according to the at least one entity word and the target business scenario category to obtain the review result corresponding to the at least one picture to be reviewed.

10. A service scenario determination device, characterized in that, The device includes: An information acquisition module, configured to acquire text information to be processed; An entity extraction and attribute extraction module, configured to determine at least one entity word and at least one attribute word in the text information to be processed; the at least one attribute word is indicator data in a business dimension; A text word segmentation module, configured to perform word segmentation processing on the text information to be processed to obtain a text character sequence; A first representation module, configured to perform embedding representation processing on the text character sequence to obtain text representation data; A second representation module, configured to generate marker representation data according to the text character sequence, the at least one entity word, and the at least one attribute word, where the marker representation data is used to indicate a first character corresponding to each entity word and a second character corresponding to each attribute word in the text character sequence; A feature extraction module, configured to perform feature extraction processing on the text representation data and the marker representation data to obtain text semantic feature data corresponding to the text information to be processed; A scenario classification module, configured to perform scenario classification processing according to the text semantic feature data to obtain a target business scenario category corresponding to the text information to be processed.

11. A computer-readable storage medium, characterized in that, At least one instruction or at least one program segment is stored in the computer-readable storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement a business scenario determination method according to any one of claims 1 to 9.

12. A computer device, characterized in that, The computer device includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement a business scenario determination method according to any one of claims 1 to 9.